Senior Software Engineer, Inference
Dublin, IE
Posted 16d ago
About the job
Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team has a dual mandate: maximizing compute efficiency to serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms.
Responsibilities
- Build and maintain critical systems for serving AI models.
- Manage the entire inference stack, from request routing to fleet orchestration.
- Maximize compute efficiency for customer growth.
- Enable breakthrough research by providing high-performance inference infrastructure.
- Address complex, distributed systems challenges across diverse AI hardware and cloud platforms.
- Design intelligent routing algorithms for request distribution.
- Implement autoscaling for compute fleets.
- Build production-grade deployment pipelines for new models.
- Integrate new AI accelerator platforms.
- Contribute to new inference features like structured sampling and prompt caching.
- Support inference for new model architectures.
- Analyze observability data to tune performance.
- Manage multi-region deployments and geographic routing.
Requirements
- Significant software engineering experience, particularly with distributed systems.
- Results-oriented with a bias towards flexibility and impact.
- Proactive in taking on tasks outside of the immediate job description.
- Desire to learn about machine learning systems and infrastructure.
- Thrive in environments where technical excellence drives business results and research breakthroughs.
- Care about the societal impacts of work.
- Experience with high-performance, large-scale distributed systems.
- Experience implementing and deploying machine learning systems at scale.
- Experience with load balancing, request routing, or traffic management systems.
- Experience with LLM inference optimization, batching, and caching strategies.
- Experience with Kubernetes and cloud infrastructure (AWS, GCP).
- Proficiency in Python or Rust.
- Bachelor's degree or equivalent combination of education, training, and/or experience.