Staff + Sr. Software Engineer, Scaling
New York City, NY; San Francisco, CA | Seattle, WA
Posted 13d ago
About the job
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. The Inference team is responsible for building and scaling the critical systems that serve Claude to millions of users worldwide. This involves managing the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators, with a dual mandate of maximizing compute efficiency for customer growth and enabling breakthrough research by providing high-performance inference infrastructure. The team tackles complex, distributed systems challenges across multiple accelerator families and emerging AI hardware on various cloud platforms.
Responsibilities
- Design, build, and maintain distributed systems for serving Claude to millions of users.
- Develop resilient, flexible systems that adapt in real-time to real-world events.
- Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators and multiple cloud providers.
- Maximize compute efficiency and optimize cost across the fleet by autoscaling and orchestrating production, research, and experimental workloads.
- Build and operate production-grade deployment pipelines for releasing new models.
- Provide high-performance inference infrastructure to enable researchers to develop next-generation models.
- Integrate new AI accelerator platforms and support inference for new model architectures.
Requirements
- Significant software engineering experience, particularly with distributed systems.
- Results-oriented with a bias towards flexibility and impact.
- Willingness to learn and adapt to new challenges.
- Desire to learn more about machine learning systems and infrastructure.
- Ability to thrive in environments where technical excellence drives business results and research breakthroughs.
- Care about the societal impacts of your work.
- Experience with high-performance, large-scale distributed systems.
- Experience implementing and deploying machine learning systems at scale.
- Experience with load balancing, request routing, or traffic management systems.
- Familiarity with LLM inference optimization, batching, and caching strategies.
- Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure).
- Proficiency in Python or Rust.
- Bachelor’s degree or an equivalent combination of education, training, and/or experience.
Benefits
- Annual compensation range: $320,000 - $485,000 USD
- Visa sponsorship available for some roles.