Staff+ Site Reliability Engineer, Safeguards ML Infra

Remote Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY

Posted 10d ago

Job Location

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY

Tech Stack

Remote Work Policy

Fully remote

Categories

AI Infrastructure Engineer

About the job

The Safeguards ML Infra team is responsible for designing, building, and operating the production infrastructure that ensures the safety of Claude's AI systems. This role is central to ensuring safeguards are correctly configured and deployed for model launches and managing the rollout of new safety classifiers. You will be responsible for verifying that the correct safeguards are active across all platforms, detecting and resolving configuration drift, and automating manual processes to improve efficiency and reliability. The team plays a critical role in every frontier model release, owning the operational aspects of deploying and verifying safety measures across various platforms and leading incident response when issues arise.

Responsibilities

  • Serve as launch captain for model releases, configuring and verifying safeguards, and acting as the primary contact during release windows.
  • Manage the deployment of new safety classifiers, including canarying rollouts, conducting post-deploy validations, and investigating discrepancies.
  • Ensure the accuracy of safeguards across all deployment platforms (1P, AWS Bedrock, GCP Vertex, etc.) and eliminate configuration drift.
  • Automate operational tasks by transforming runbooks into tooling, manual checks into continuous validation, and one-off deploys into repeatable pipelines.
  • Develop and maintain a safeguards registry detailing production configurations, model associations, platform deployment, and deployment history.
  • Participate in on-call rotations for incident response, model provisioning, and time-sensitive launches.

Requirements

  • Experience owning production change management at scale, including deploy pipelines, config management systems, and canary analysis.
  • Proven ability to manage high-stakes releases, serving in roles like launch captain or incident commander for critical systems.
  • Significant on-call experience for production systems, including incident response and implementing improvements based on postmortems.
  • Demonstrated ability to address gaps and manually manage processes until automation and tooling can be developed.
  • Hands-on experience deploying and operating systems on cloud platforms (AWS, GCP) at scale.
  • Proficiency in Python; experience with Rust is a plus.

Benefits

  • Annual compensation range: $320,000 - $485,000 USD

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.