Staff+ Site Reliability Engineer, Safeguards ML Infra
Remote • Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY
Posted 10d ago
About the job
The Safeguards ML Infra team is responsible for designing, building, and operating the production infrastructure that ensures the safety of Claude's AI systems. This role is central to ensuring safeguards are correctly configured and deployed for model launches and managing the rollout of new safety classifiers. You will be responsible for verifying that the correct safeguards are active across all platforms, detecting and resolving configuration drift, and automating manual processes to improve efficiency and reliability. The team plays a critical role in every frontier model release, owning the operational aspects of deploying and verifying safety measures across various platforms and leading incident response when issues arise.
Responsibilities
- Serve as launch captain for model releases, configuring and verifying safeguards, and acting as the primary contact during release windows.
- Manage the deployment of new safety classifiers, including canarying rollouts, conducting post-deploy validations, and investigating discrepancies.
- Ensure the accuracy of safeguards across all deployment platforms (1P, AWS Bedrock, GCP Vertex, etc.) and eliminate configuration drift.
- Automate operational tasks by transforming runbooks into tooling, manual checks into continuous validation, and one-off deploys into repeatable pipelines.
- Develop and maintain a safeguards registry detailing production configurations, model associations, platform deployment, and deployment history.
- Participate in on-call rotations for incident response, model provisioning, and time-sensitive launches.
Requirements
- Experience owning production change management at scale, including deploy pipelines, config management systems, and canary analysis.
- Proven ability to manage high-stakes releases, serving in roles like launch captain or incident commander for critical systems.
- Significant on-call experience for production systems, including incident response and implementing improvements based on postmortems.
- Demonstrated ability to address gaps and manually manage processes until automation and tooling can be developed.
- Hands-on experience deploying and operating systems on cloud platforms (AWS, GCP) at scale.
- Proficiency in Python; experience with Rust is a plus.
Benefits
- Annual compensation range: $320,000 - $485,000 USD