Software Engineer, ML Infrastructure
San Francisco • FullTime
Posted 7mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
The ML Infrastructure team builds large-scale compute, storage, and software infrastructure to support the company's work building the world's best agentic coding model. This role works closely with ML researchers and engineers to enable their work through improvements to our training framework, systems reliability/performance, and developer experience. We are looking for strong engineers interested in building high-performance infrastructure and the software to support it.
Responsibilities
- Collaborate with ML researchers to improve the throughput and reliability of training.
- Work with OEMs, cloud service providers, and others to plan and build cutting-edge GPU infrastructure.
- Improve the density and scalability of compute environments to enable increasingly large RL workloads.
- Create software and systems to automate building, monitoring, and running GPU clusters.
- Build workload scheduling and data movement systems to support the company's growing training footprint.
Requirements
- Strong background in systems and infrastructure-focused software engineering.
- Experience with distributed storage and networking infrastructure, particularly on Linux systems across cloud and bare metal environments.
- Experience with large-scale systems and their unique challenges, ideally across thousands of nodes with significant resource footprints.
- Production use of infrastructure-as-code and configuration management, across hosts and Kubernetes.