Technical Program Manager, Safeguards (Infrastructure & Evals)

San Francisco, CA | New York City, NY | Seattle, WA

Posted 13d ago

Job Location

San Francisco, CA | New York City, NY | Seattle, WA

Tech Stack

Remote Work Policy

On-site

Categories

AI Infrastructure Engineer

About the job

Safeguards Engineering builds and operates the infrastructure that keeps Anthropic's AI systems safe in production, including classifiers, detection pipelines, evaluation platforms, and monitoring systems. This infrastructure must be reliable, as failures in safety-critical pipelines can have serious and often invisible consequences. As a Technical Program Manager for Safeguards Infrastructure and Evals, you will own the operational health and progress of this stack. Your main focus will be driving reliability through incident response, post-mortem processes, ensuring Service Level Objectives (SLOs) are met, and managing platform investments like migrations and evaluation platform improvements.

Responsibilities

  • Own the Safeguards Engineering operations review, including surfacing incidents, visibility into reliability trends, and facilitating decision-making.
  • Drive incident tracking and post-mortem execution, ensuring follow-through on action items from incidents across the organization.
  • Establish and maintain SLOs with partner teams for safety-critical pipelines and build reporting to track their performance.
  • Maintain runbook quality and clarify incident ownership for safety-critical systems to prevent issues from falling through the cracks.
  • Manage program execution for infrastructure projects, including platform and cloud system migrations, ensuring sequencing and cross-team coordination.
  • Coordinate improvements to the evaluation platform, including scoping work, tracking dependencies, and ensuring alignment with team needs.

Requirements

  • Solid technical program management experience, especially in operational or infrastructure-heavy environments.
  • Comfortable owning both ongoing operational cadences and discrete project work.
  • Understanding of production ML systems sufficient for incident triage and technical discussions with engineers.
  • Ability to drive processes and follow-ups to ensure action items are completed and SLOs are met.
  • Skilled at working across team boundaries, coordinating with partner teams without direct authority, and driving work through influence and clear communication.
  • Ability to context-switch effectively between operational "keep the lights on" tasks and new development projects.

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.