[Expression of Interest] Research Manager, Interpretability

San Francisco, CA

Posted 16d ago

Job Location

San Francisco, CA

Tech Stack

Remote Work Policy

On-site

Categories

AI Research Engineer

About the job

Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. The Interpretability team focuses on mechanistic interpretability, aiming to reverse engineer how trained models work by discovering how neural network parameters map to meaningful algorithms. This role supports a team of expert researchers and engineers dedicated to understanding the internal workings of large language models at a deep, mechanistic level. The manager will play a critical role in ensuring the team meets its ambitious safety research goals by partnering with an individual contributor research lead to drive team success, translate research ideas into tangible goals, and oversee execution. Responsibilities include managing team execution, careers, performance, facilitating inter-team relationships, and driving the hiring pipeline.

Responsibilities

  • Manage team execution, careers, and performance.
  • Facilitate relationships within and across teams.
  • Drive the hiring pipeline for the Interpretability team.
  • Partner with an individual contributor research lead to drive team success.
  • Translate cutting-edge research ideas into tangible goals and oversee their execution.

Requirements

  • Experience managing research or engineering teams.
  • Experience in AI safety research or related fields.
  • Familiarity with mechanistic interpretability research.
  • Understanding of large language models.
  • Ability to translate research ideas into actionable goals.
  • Strong leadership and people management skills.
  • Experience in hiring and talent development.

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.