Staff+ Software Engineer, Safeguards Evals

San Francisco, CA | New York City, NY

Posted 7d ago

Job Location

San Francisco, CA | New York City, NY

Tech Stack

Remote Work Policy

On-site

Categories

Applied AI Engineer

About the job

Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. This role focuses on developing the evaluation infrastructure for AI-powered misuse investigation systems. You will design experiments to measure the performance of investigative agents, build datasets representing real-world abuse, and integrate these methods into production pipelines. Your work will directly influence the trust Anthropic places in its automated abuse detection and guide improvements.

Responsibilities

  • Build and own the evaluation harness for an agentic investigation system, defining metrics, test cases, and grading approaches.
  • Construct high-quality evaluation datasets representing real-world misuse across various harm areas.
  • Measure agent performance end-to-end and drive improvements in detection precision, recall, investigation quality, and robustness.
  • Analyze coverage to identify measurement gaps and evolve evaluations as agent capabilities advance.
  • Productionize research into regression and release pipelines for agent changes, prompt updates, and model upgrades.
  • Build tooling to enable policy experts to author, run, and iterate on evaluations independently.
  • Construct RL environments to enhance Claude's safety investigation capabilities.

Requirements

  • Proficiency in Python and comfort working across the full stack.
  • Experience building and maintaining data pipelines.
  • Experience working with LLMs, understanding their capabilities and failure modes, especially agentic systems with tool use and multi-step reasoning.
  • Strong data analysis skills to derive insights from large datasets.
  • Ability to transition between research prototyping and production-quality code.
  • Ability to translate ambiguous problems into concrete, testable experiments.

Benefits

  • Annual compensation range: $320,000 - $485,000 USD
  • Visa sponsorship available

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.