Software Engineer, Data Infrastructure
$216k - $300k • New York, NY; Washington, DC
Posted 1mo ago
About the job
Scale AI is seeking a highly skilled and motivated Software Engineer to join our dynamic Public Sector Engineering team. You will play a critical role in supporting Scale’s government customers by scoping and developing onsite solutions. Your expertise will be instrumental in designing and implementing systems that can handle interactions with existing customer systems to help our products integrate into existing customer workflows. We are looking for an exceptional Senior Software Engineer to architect and build the foundational data infrastructure that will serve as the brain of a project ecosystem. You will be responsible for designing highly novel data models and processing pipelines capable of handling massive quantities of output data from complex simulations. At the core of this role is the challenge of building a foundational data ensemble—a unified architecture that seamlessly aggregates, structures, and stages diverse sources of simulation outputs and user inputs. Your systems will manage enormous batch throughput jobs with strict, minimal latency requirements, ensuring that downstream AI systems and language models have the exact context they need to actionably reason over complex, multi-dimensional scenarios.
Responsibilities
- Architect and implement the data ensemble architecture to unify various sources of context for LLM consumption.
- Build highly scalable, resilient data architectures from scratch, optimizing for massive batch jobs and minimal latency.
- Design sophisticated, highly relational data models for massive, state-based simulation environments.
- Design custom, ground-up systems where existing tools cannot handle complexity or scale.
- Set the technical standard for the data infrastructure team, driving code quality, system performance, and architectural clarity.
Requirements
- 5+ years of backend or data infrastructure experience at a Senior, Staff, or Principal level.
- Expert-level proficiency in systems languages (e.g., Rust, Go, C++, or highly optimized Python/Java, Spark, PySpark).
- Fundamental understanding of memory management, compute limits, and distributed systems architecture.
- Proven track record of processing massive datasets with optimized massive batch jobs and parallel processing.
- Expertise in surfacing the right information to feed decision-making engines.
- Experience building complex information retrieval systems or recommendation engines.
- Experience managing complex state, data relationships, and telemetry for massive simulations.
- Experience processing disparate, massive streams of data for algorithmic decision-making.
Benefits
- Base salary
- Equity
- Comprehensive health, dental and vision coverage
- Retirement benefits
- Learning and development stipend
- Generous PTO
- Commuter stipend