Staff + Sr. Software Engineer, Cloud Inference Launch Engineering
San Francisco, CA
Posted 16d ago
About the job
Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. The Cloud Inference team is responsible for scaling and optimizing Claude to serve massive audiences across AWS, GCP, Azure, and other cloud providers. This involves end-to-end ownership of Claude on each platform, from API integration and request routing to inference execution and capacity management. The model & inference launch team specifically focuses on the validation pipeline for the inference server and load balancer, ensuring every inference change, including model launches and performance improvements, is deployed with correctness, performance, and reliability. This high-leverage infrastructure work is critical for fast and cost-effective deployment of frontier models and features, directly impacting compute capacity and the speed at which innovations reach production.
Responsibilities
- Be on the critical path for frontier model launches, bringing up inference for new model architectures and shipping them to cloud platforms.
- Integrate new inference features (e.g., structured sampling, prompt caching) into cloud platforms.
- Identify and fix cross-platform inference differences (config drift, observability, deployment patterns, bugs).
- Design, build, and own CI/CD infrastructure for inference servers and load balancers across cloud platforms.
- Drive down merge-to-production cycle time by making validation faster, more parallel, and cost-effective.
- Analyze observability data across providers to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation.
Requirements
- Strong interest in LLM serving; prior inference or ML experience not required.
- Significant software engineering experience with a strong background in high-performance, large-scale distributed systems.
- Track record of building automation or test infrastructure that improved release velocity or reliability.
- Experience building or operating services on AWS, GCP, or Azure, with exposure to Kubernetes, Infrastructure as Code, or container orchestration.
- Ability to thrive in cross-functional collaboration.
- Fast learner capable of quickly ramping up on new technologies and platforms.
- Highly autonomous with end-to-end ownership of problems.
Benefits
- Annual compensation range: $320,000 - $485,000 USD
- Visa sponsorship available for some roles/candidates
- Hybrid work policy (at least 25% in office)