Staff + Senior Software Engineer, Inference

San Francisco, CA | New York City, NY | Seattle, WA

Posted 16d ago

Job Location

San Francisco, CA | New York City, NY | Seattle, WA

Tech Stack

Remote Work Policy

On-site

Categories

Applied AI Engineer

About the job

Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide, managing the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team's dual mandate is to maximize compute efficiency for explosive customer growth while enabling breakthrough research by providing scientists with high-performance inference infrastructure for next-generation models. This involves tackling complex, distributed systems challenges across multiple accelerator families and emerging AI hardware on various cloud platforms.

Responsibilities

  • Design, build, and maintain distributed systems for serving Claude globally.
  • Develop intelligent request routing, load balancing, and traffic management systems.
  • Maximize compute efficiency through autoscaling and orchestrating diverse workloads.
  • Build and operate production-grade deployment pipelines for new model releases.
  • Provide high-performance inference infrastructure for research and model development.
  • Integrate new AI accelerator platforms and support inference for new model architectures.
  • Tune and improve performance using observability data from production workloads.

Requirements

  • Significant software engineering experience, particularly with distributed systems.
  • Results-oriented with a bias towards flexibility and impact.
  • Willingness to adapt and take on tasks outside of the immediate job description.
  • Enjoy pair programming.
  • Desire to learn about machine learning systems and infrastructure.
  • Thrive in environments where technical excellence drives business results and research breakthroughs.
  • Care about the societal impacts of your work.
  • Experience with high-performance, large-scale distributed systems.
  • Experience implementing and deploying machine learning systems at scale.
  • Experience with load balancing, request routing, or traffic management systems.
  • Familiarity with LLM inference optimization, batching, and caching strategies.
  • Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure).
  • Proficiency in Python or Rust.
  • Bachelor's degree or equivalent combination of education, training, and/or experience.

Benefits

  • Annual compensation range: $320,000 - $485,000 USD
  • Visa sponsorship available for some roles and candidates.

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.