Research Engineer, Discovery
San Francisco, CA
Posted 16d ago
About the job
As a Research Engineer on our team, you will work end-to-end across the entire model stack, identifying and addressing key infrastructure blockers on the path to scientific AGI. You should have familiarity with elements of language model training, evaluation, and inference, and be eager to quickly dive into and get up to speed in areas where you are not yet an expert. This may include performance optimization, distributed systems, VM/sandboxing/container deployment, and large-scale data pipelines. Join us in our mission to develop advanced AI systems that push the frontiers of science and benefit humanity.
Responsibilities
- Design and implement large-scale infrastructure systems for AI scientist training, evaluation, and deployment in distributed environments.
- Identify and resolve infrastructure bottlenecks hindering progress toward scientific capabilities.
- Develop robust and reliable evaluation frameworks to measure progress towards scientific AGI.
- Build scalable and performant VM/sandboxing/container architectures for safe execution of long-horizon AI tasks and scientific workflows.
- Collaborate to translate experimental requirements into production-ready infrastructure.
- Develop large-scale data pipelines for advanced language model training requirements.
- Optimize large-scale training and inference pipelines for stable and efficient reinforcement learning.
Requirements
- 6+ years of highly-relevant experience in infrastructure engineering with demonstrated expertise in large-scale distributed systems.
- Strong communication and collaboration skills.
- Deep knowledge of performance optimization techniques and system architectures for high-throughput ML workloads.
- Experience with containerization technologies (Docker, Kubernetes) and orchestration at scale.
- Proven track record of building large-scale data pipelines and distributed storage systems.
- Ability to diagnose and resolve complex infrastructure challenges in production environments.
- Effectively work across the full ML stack from data pipelines to performance optimization.
- Experience collaborating with researchers to scale experimental ideas.
- Ability to thrive in fast-paced environments and rapidly iterate from experimentation to production.
Benefits
- Annual compensation range: $350,000 - $850,000 USD
- Visa sponsorship available
- Hybrid work policy (minimum 25% in office)