Engineering Manager, Safeguards Interventions

San Francisco, CA

Posted 16d ago

Job Location

San Francisco, CA

Tech Stack

Remote Work Policy

On-site

Categories

Applied AI Engineer

About the job

Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. The Safeguards team ensures our models and products are developed and deployed safely. This role manages the Interventions team, which is responsible for the systems that operate between our detection stack and the user across all Anthropic surfaces. The team's work is crucial for enabling products to grow safely by evolving and driving the quality of interventions for areas like bio, cyber, acceptable usage, child safety, and copyright.

Responsibilities

  • Lead and grow a team of engineers, owning the roadmap, OKRs, and execution.
  • Drive cross-functional collaboration with ML Infra, Research, Product, Policy, Legal, and cloud partners.
  • Define and measure intervention quality, representing safety and product tradeoffs to leadership and stakeholders.
  • Ensure production reliability of intervention and compliance systems, including incident response and postmortems.

Requirements

  • Managed engineering teams shipping production ML or safety-enforcement systems with direct user impact.
  • Experience running high-stakes, compliance-adjacent production systems, including on-call and incident management.
  • Proficiency in building and utilizing evaluation metrics to prove system effectiveness.
  • Ability to drive multi-stakeholder tradeoffs (safety, UX, latency, cost) to a decision.
  • Deep commitment to AI safety and enabling the deployment of advanced models.

Benefits

  • Annual compensation range: $405,000 - $485,000 USD
  • Visa sponsorship available
  • Hybrid work policy (at least 25% in office)

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.