Staff+ Software Engineer, ML Sampling Path

San Francisco, CA

Posted 4d ago

Job Location

San Francisco, CA

Tech Stack

Remote Work Policy

On-site

Categories

Machine Learning Engineer

About the job

The Safeguards ML Sampling Path team builds and operates the production services that power Claude's safety systems. These services are critical, sitting on the token generation path across every platform Claude runs on. You will be responsible for maintaining low latency and high reliability as traffic grows, ensuring robustness against dependency failures, and safely deploying changes to a system that cannot afford downtime. This role involves designing, building, and operating backend systems that process every token on the generation path, managing the streaming contract with the API and inference engines, and owning latency and reliability end-to-end.

Responsibilities

  • Design, build, and operate backend systems processing every token on Claude's generation path.
  • Manage the streaming contract with the API and inference engines.
  • Own latency and reliability end-to-end, defining and maintaining SLOs and error budgets.
  • Lead incident response and postmortem follow-through.
  • Ship changes to the hot path rapidly and safely using canaried and gradual rollouts, error budget, and latency gating.
  • Drive per-token performance, managing tail latency and cost.
  • Set technical direction for the sampling path, leading design reviews and making trade-off calls.
  • Mentor engineers and raise the operational bar for the Safeguards organization.

Requirements

  • Experience designing, building, and operating high QPS systems at global scale with accountability for production incidents and remediation.
  • Strong foundation in distributed systems, including replication, consistency tradeoffs, failure modes, and SLO management under load.
  • Experience designing systems for graceful degradation, planning for dependency failures and partial rollouts.
  • Proven success in shipping broad or all-encompassing changes to mission-critical systems.
  • 8+ years of industry software engineering experience.
  • Familiarity with LLM inference systems and transformer-based models (a plus).

Benefits

  • Annual compensation range: $320,000 - $485,000 USD
  • Visa sponsorship available
  • Hybrid work policy (at least 25% in office)

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.