Research Engineer, Knowledge Foundations
San Francisco, CA
Posted 17d ago
Remote Work Policy
On-site
Categories
AI Research Engineer
About the job
The Knowledge Work team at Anthropic builds the training environments and evaluations that empower Claude to excel in professional workflows, including searching, analyzing, and creating content across various tools and documents. As this work expands, the underlying systems require the same rigor as the research itself. In this role, you will design and execute experiments to enhance Claude's ability to search, retrieve, and reason over information at scale. Your responsibilities will encompass environment design, data curation, RL training, evaluation, and supporting infrastructure, requiring flexibility to address progress blockers. You will collaborate closely with researchers and other RL teams to implement capabilities that directly influence Claude's performance. We believe that true ownership and impact stem from both hardening existing environments and creating new ones, ensuring the quality of the entire stack that drives superhuman epistemics.
Responsibilities
- Design, build, and iterate on training environments and data pipelines to improve Claude's reasoning over knowledge-intensive tasks.
- Conduct end-to-end ML experiments, including hypothesis formation, infrastructure building, model training, result analysis, and iteration planning.
- Develop evaluations to accurately measure progress in search, retrieval, and reasoning quality.
- Identify model failure modes and translate them into actionable training signals.
- Collaborate with researchers across RL Data, post-training, and product teams on priorities and feature delivery.
- Contribute to shared infrastructure and tooling to enhance team velocity.
- Maintain a clean, canonical set of evaluation tools and processes for Knowledge Work capabilities, including those for model releases.
- Build and automate observability, dashboards, and operational tooling for training environments and evaluation systems, focusing on high signal-to-noise metrics.
Requirements
- Highly experienced Python engineer with a track record of shipping reliable, well-instrumented production code.
- Experience designing, running, and analyzing ML experiments.
- Ability to work across the full stack, from data pipelines to model training and evaluation.
- 5+ years of experience operating ML or distributed systems at scale.
- Comfort working with ambiguity and prioritizing impactful problems.
- Clear written and verbal communication skills, particularly for cross-time zone collaboration.
- Satisfaction derived from making critical systems dependable.