Staff + Senior Software Engineer, Inference Deployment
San Francisco, CA | New York City, NY | Seattle, WA
Posted 16d ago
About the job
Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. Our growing team of researchers, engineers, and policy experts collaborates to create these advanced AI systems. The Inference team is central to this mission, responsible for building and maintaining the critical systems that serve Claude to millions of users globally. We manage the entire stack, from intelligent request routing to fleet-wide orchestration across diverse AI accelerators, ensuring Claude is brought to life efficiently and reliably.
Responsibilities
- Design, build, and maintain distributed systems for serving Claude to millions of users.
- Develop resilient and flexible systems that adapt to real-world events.
- Create intelligent request routing, load balancing, and traffic management systems for thousands of accelerators.
- Maximize compute efficiency through autoscaling and orchestration of production, research, and experimental workloads.
- Build and operate production-grade deployment pipelines for releasing new models.
- Provide high-performance inference infrastructure to enable next-generation model development.
- Integrate new AI accelerator platforms and support inference for new model architectures.
Requirements
- Significant software engineering experience, particularly with distributed systems.
- Results-oriented with a bias towards flexibility and impact.
- Willingness to learn and adapt to tasks outside of the immediate job description.
- Desire to learn about machine learning systems and infrastructure.
- Ability to thrive in environments where technical excellence drives business results and research breakthroughs.
- Care about the societal impacts of AI work.
- Experience with high-performance, large-scale distributed systems.
- Experience implementing and deploying machine learning systems at scale.
- Experience with load balancing, request routing, or traffic management systems.
- Familiarity with LLM inference optimization, batching, and caching strategies.
- Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure).
- Proficiency in Python or Rust.
Benefits
- Annual compensation range: $320,000 - $485,000 USD
- Visa sponsorship available for eligible candidates.