[Expression of Interest] Research Manager, Interpretability
San Francisco, CA
Posted 16d ago
About the job
Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. The Interpretability team focuses on mechanistic interpretability, aiming to reverse engineer how trained models work by discovering how neural network parameters map to meaningful algorithms. This role supports a team of expert researchers and engineers dedicated to understanding the internal workings of large language models at a deep, mechanistic level. The manager will play a critical role in ensuring the team meets its ambitious safety research goals by partnering with an individual contributor research lead to drive team success, translate research ideas into tangible goals, and oversee execution. Responsibilities include managing team execution, careers, performance, facilitating inter-team relationships, and driving the hiring pipeline.
Responsibilities
- Manage team execution, careers, and performance.
- Facilitate relationships within and across teams.
- Drive the hiring pipeline for the Interpretability team.
- Partner with an individual contributor research lead to drive team success.
- Translate cutting-edge research ideas into tangible goals and oversee their execution.
Requirements
- Experience managing research or engineering teams.
- Experience in AI safety research or related fields.
- Familiarity with mechanistic interpretability research.
- Understanding of large language models.
- Ability to translate research ideas into actionable goals.
- Strong leadership and people management skills.
- Experience in hiring and talent development.