Staff+ Site Reliability Engineer, Safeguards ML Infra
Remote • Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY
Posted 4d ago
About the job
The Safeguards ML Infra team is responsible for designing, building, and operating the production infrastructure that ensures Claude's safety systems function correctly. This critical role involves managing backend services for token generation safety, overseeing the deployment of safeguards for new model launches, and deploying new safety classifiers. You will be at the forefront of ensuring safeguards are properly configured and deployed for model launches, owning the off-cycle deployment of new safety classifiers, and leading incident response when issues arise. The goal is to automate manual processes, transforming launch runbooks into tooling and hand-built checks into continuous validation systems.
Responsibilities
- Serve as launch captain for model releases, configuring and verifying safeguards, and acting as the primary contact during release windows.
- Manage the deployment of new safety classifiers, including canarying rollouts, conducting post-deploy validations, and investigating discrepancies.
- Ensure the correct safeguards are deployed across all platforms (1P, AWS Bedrock, GCP Vertex) and prevent configuration drift.
- Automate operational tasks by converting launch runbooks into tooling, manual checks into continuous validation, and one-off deploys into repeatable pipelines.
- Develop and maintain a safeguards registry detailing production deployments, including model, platform, and deployment history.
- Participate in on-call rotations for service incidents, model provisioning, and time-sensitive launches.
Requirements
- Experience owning production change management at scale, including deploy pipelines, config management systems, and canary analysis.
- Experience running high-stakes releases, serving as a launch captain, incident commander, or release owner for critical systems.
- Meaningful on-call experience for production systems, including incident response and implementing process/tooling improvements based on postmortems.
- A proactive approach to identifying and closing gaps, even if it requires manual intervention until automation is built.
- Hands-on experience deploying and operating on cloud platforms (AWS, GCP) at scale.
- Proficiency in Python; experience with Rust is a plus.