Staff Software Engineer, Inference
London, UK
Posted 14d ago
About the job
Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team has a dual mandate: maximizing compute efficiency to serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms. As a Staff Software Engineer on our Inference team, you will work end to end, identifying and addressing key infrastructure blockers to serve Claude to millions of users while enabling breakthrough AI research.
Responsibilities
- Build and maintain critical systems that serve Claude to millions of users.
- Manage the entire inference stack, from intelligent request routing to fleet-wide orchestration.
- Maximize compute efficiency to support customer growth.
- Enable breakthrough research by providing high-performance inference infrastructure.
- Address complex, distributed systems challenges across diverse AI accelerators and hardware.
- Identify and address key infrastructure blockers.
- Design intelligent routing algorithms to optimize request distribution.
- Autoscale compute fleets to match supply with demand.
- Build production-grade deployment pipelines for new models.
- Integrate new AI accelerator platforms.
- Contribute to new inference features.
- Support inference for new model architectures.
- Analyze observability data to tune performance.
- Manage multi-region deployments and geographic routing.
Requirements
- Significant software engineering experience, particularly with distributed systems.
- Familiarity with performance optimization.
- Familiarity with large-scale service orchestration.
- Familiarity with intelligent request routing.
- Experience implementing and deploying machine learning systems at scale.
- Experience with load balancing, request routing, or traffic management systems.
- Experience with LLM inference optimization, batching, and caching strategies.
- Experience with Kubernetes and cloud infrastructure (AWS, GCP).
- Proficiency in Python or Rust.
- Bachelor’s degree or equivalent combination of education, training, and/or experience.
Benefits
- Annual compensation range: £325,000 — £390,000 GBP
- Visa sponsorship available