Software Engineer, Infrastructure, Interpretability
San Francisco, CA
Posted 1d ago
About the job
Anthropic is building reliable, interpretable, and steerable AI systems to ensure AI is safe and beneficial for society. The Interpretability team focuses on understanding the inner workings of trained models and applying techniques to enhance AI safety. This role is an early hire on a new infrastructure effort within the Interpretability team, focused on defining and building the systems that enable secure, private, and low-friction access to frontier models for researchers. The work involves designing and implementing solutions across security, privacy, data and compute management, and developer experience to support cutting-edge AI research and its application to safety decisions.
Responsibilities
- Design, build, and own shared infrastructure for Interpretability, including research environments, data systems, and compute tooling.
- Lead cross-team efforts with agentic engineering, security, compute, and storage platform teams to align company-wide solutions with research needs.
- Identify and resolve organization-wide developer experience issues.
- Transition interpretability methods from research code to dependable audit pipelines.
Requirements
- Proficiency in at least one programming language (e.g., Python, Rust, Go, Java) and productivity with Python.
- Significant experience building and operating secure and scalable software infrastructure, cloud systems, distributed systems, or developer tooling.
- Strong cross-functional communication skills, comfortable working with both researchers and platform/security teams.
- High curiosity about unfamiliar domains.
- Strong ability to prioritize impactful work and comfort operating with ambiguity.
- Curiosity about interpretability research and its role in AI safety.
- Commitment to the societal impacts and ethics of AI work.