Research Scientist, Interpretability
San Francisco, CA
Posted 17d ago
About the job
The Interpretability team at Anthropic is dedicated to reverse-engineering how trained language models work, believing that a deep mechanistic understanding is crucial for ensuring the safety of advanced AI systems. This role focuses on mechanistic interpretability, aiming to map neural network parameters to meaningful algorithms, akin to performing 'biology' or 'neuroscience' on neural networks. The team's work involves developing tools and methods to understand complex models, addressing challenges like 'superposition' and decomposing models into interpretable components. Collaborations with other Anthropic teams, such as Alignment Science and Societal Impacts, are integral to applying this research to enhance AI safety.
Responsibilities
- Reverse-engineer how trained language models work.
- Discover how neural network parameters map to meaningful algorithms.
- Develop tools and methods for mechanistic interpretability of neural networks.
- Address challenges like 'superposition' in neural networks.
- Decompose models into more interpretable components.
- Collaborate with other teams to use interpretability research for AI safety.
Requirements
- Strong research and engineering background.
- Interest in understanding how neural networks function.
- Focus on mechanistic interpretability.
- Experience or interest in reverse-engineering neural networks.
- Familiarity with concepts like superposition in neural networks.
- Ability to collaborate effectively with research and engineering teams.