Research Engineer, Interpretability
San Francisco, CA
Posted 17d ago
About the job
The Interpretability team at Anthropic is dedicated to understanding how large language models work, believing that a mechanistic understanding is key to making advanced AI systems safe and reliable. This role involves building and maintaining the specialized infrastructure for interpretability research, akin to performing 'neuroscience' on neural networks. The work spans the entire lifecycle of a production language model, from pretraining and inference to performance optimization, pushing the boundaries of hardware and software to address critical bottlenecks. As interpretability research matures and is applied to safety audits on frontier models, engineering and infrastructure have become crucial, making this role directly impactful on one of AI's most significant open problems.
Responsibilities
- Build and maintain specialized inference and training infrastructure for interpretability research, including instrumented passes, activation extraction, and steering vector application.
- Resolve scaling and efficiency bottlenecks through profiling, optimization, and collaboration with infrastructure teams.
- Design tools, abstractions, and platforms to enable rapid researcher experimentation.
- Contribute to bringing interpretability research into production safety audits with high reliability expectations.
- Work across the full stack, from model internals and accelerator optimization to user-facing research tooling.
Requirements
- 5-10+ years of software building experience.
- High proficiency in at least one programming language (e.g., Python, Rust, Go, Java) and productivity with Python.
- High curiosity and ability to quickly learn and apply knowledge in unfamiliar domains.
- Strong ability to prioritize impactful work and comfort with ambiguity.
- Preference for fast-moving collaborative projects.
- Curiosity about interpretability research and AI safety (research experience not required).
- Awareness of societal impacts and ethics of AI work.
- Comfort working closely with researchers and translating needs into engineering solutions.