Staff + Senior Software Engineer, Inference

Ontario, CAN

Posted 20h ago

Remote Work Policy

On-site

Categories

AI Infrastructure Engineer

About the job

The Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. This role involves managing the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team's dual mandate is to maximize compute efficiency for explosive customer growth while enabling breakthrough research by providing high-performance inference infrastructure for next-generation models. This involves tackling complex, distributed systems challenges across multiple accelerator families and emerging AI hardware on various cloud platforms.

Responsibilities

  • Design, build, and maintain distributed systems serving Claude to millions of users.
  • Develop resilient, flexible systems that adapt in real-time to real-world events.
  • Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators.
  • Maximize compute efficiency across the fleet by autoscaling and orchestrating production, research, and experimental workloads.
  • Build and operate production-grade deployment pipelines for releasing new models.
  • Provide high-performance inference infrastructure to enable researchers to develop next-generation models.
  • Integrate new AI accelerator platforms and support inference for new model architectures.

Requirements

  • Significant software engineering experience, particularly with distributed systems.
  • Results-oriented with a bias towards flexibility and impact.
  • Willingness to learn and contribute beyond the immediate job description.
  • Desire to learn more about machine learning systems and infrastructure.
  • Ability to thrive in environments where technical excellence drives business results and research breakthroughs.
  • Care about the societal impacts of your work.
  • Experience with high-performance, large-scale distributed systems.
  • Experience implementing and deploying machine learning systems at scale.
  • Experience with load balancing, request routing, or traffic management systems.
  • Familiarity with LLM inference optimization, batching, and caching strategies.
  • Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure).
  • Proficiency in Python or Rust.

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.