Staff + Sr. Software Engineer, Cloud Inference Launch Engineering

San Francisco, CA

Posted 16d ago

Job Location

San Francisco, CA

Tech Stack

Remote Work Policy

On-site

Categories

Applied AI Engineer

About the job

Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. The Cloud Inference team is responsible for scaling and optimizing Claude to serve massive audiences across AWS, GCP, Azure, and other cloud providers. This involves end-to-end ownership of Claude on each platform, from API integration and request routing to inference execution and capacity management. The model & inference launch team specifically focuses on the validation pipeline for the inference server and load balancer, ensuring every inference change, including model launches and performance improvements, is deployed with correctness, performance, and reliability. This high-leverage infrastructure work is critical for fast and cost-effective deployment of frontier models and features, directly impacting compute capacity and the speed at which innovations reach production.

Responsibilities

  • Be on the critical path for frontier model launches, bringing up inference for new model architectures and shipping them to cloud platforms.
  • Integrate new inference features (e.g., structured sampling, prompt caching) into cloud platforms.
  • Identify and fix cross-platform inference differences (config drift, observability, deployment patterns, bugs).
  • Design, build, and own CI/CD infrastructure for inference servers and load balancers across cloud platforms.
  • Drive down merge-to-production cycle time by making validation faster, more parallel, and cost-effective.
  • Analyze observability data across providers to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation.

Requirements

  • Strong interest in LLM serving; prior inference or ML experience not required.
  • Significant software engineering experience with a strong background in high-performance, large-scale distributed systems.
  • Track record of building automation or test infrastructure that improved release velocity or reliability.
  • Experience building or operating services on AWS, GCP, or Azure, with exposure to Kubernetes, Infrastructure as Code, or container orchestration.
  • Ability to thrive in cross-functional collaboration.
  • Fast learner capable of quickly ramping up on new technologies and platforms.
  • Highly autonomous with end-to-end ownership of problems.

Benefits

  • Annual compensation range: $320,000 - $485,000 USD
  • Visa sponsorship available for some roles/candidates
  • Hybrid work policy (at least 25% in office)

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.