Staff + Senior Software Engineer, Inference
Ontario, CAN
Posted 20h ago
Job Location
Ontario, CAN
Tech Stack
Remote Work Policy
On-site
Categories
AI Infrastructure Engineer
About the job
The Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. This role involves managing the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team's dual mandate is to maximize compute efficiency for explosive customer growth while enabling breakthrough research by providing high-performance inference infrastructure for next-generation models. This involves tackling complex, distributed systems challenges across multiple accelerator families and emerging AI hardware on various cloud platforms.
Responsibilities
- Design, build, and maintain distributed systems serving Claude to millions of users.
- Develop resilient, flexible systems that adapt in real-time to real-world events.
- Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators.
- Maximize compute efficiency across the fleet by autoscaling and orchestrating production, research, and experimental workloads.
- Build and operate production-grade deployment pipelines for releasing new models.
- Provide high-performance inference infrastructure to enable researchers to develop next-generation models.
- Integrate new AI accelerator platforms and support inference for new model architectures.
Requirements
- Significant software engineering experience, particularly with distributed systems.
- Results-oriented with a bias towards flexibility and impact.
- Willingness to learn and contribute beyond the immediate job description.
- Desire to learn more about machine learning systems and infrastructure.
- Ability to thrive in environments where technical excellence drives business results and research breakthroughs.
- Care about the societal impacts of your work.
- Experience with high-performance, large-scale distributed systems.
- Experience implementing and deploying machine learning systems at scale.
- Experience with load balancing, request routing, or traffic management systems.
- Familiarity with LLM inference optimization, batching, and caching strategies.
- Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure).
- Proficiency in Python or Rust.