Research Engineer, Model Evaluations
$40k - $57k • Remote • Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NY
Posted 16d ago
About the job
We are seeking Research Engineers to develop evaluations that will define and measure Claude's capabilities and personality. Your work will translate abstract concepts of intelligence into concrete, verifiable metrics for researchers, leadership, and the public. You will design and implement evaluations across Claude's full spectrum of abilities and build the scalable infrastructure to run them. This role involves close collaboration with researchers throughout the development of new capabilities, from defining measurement criteria to interpreting results against live training checkpoints. The ultimate goal is to establish Anthropic as the leader in highly characterized AI systems, with performance rigorously measured and validated across critical tasks.
Responsibilities
- Design and execute new evaluations for Claude's reasoning, agentic behavior, knowledge, and safety properties, creating visualizations for clear communication.
- Build and maintain a distributed evaluation platform to reliably run hundreds of evaluations against production RL training checkpoints.
- Develop and improve dashboards for monitoring model health during training, enhancing signal-to-noise, reducing latency, and preventing regressions.
- Debug anomalous evaluation results during training runs, identifying the root cause (model change or infrastructure issue) and communicating findings under pressure.
- Enhance tooling, libraries, and workflows for researchers implementing and iterating on evaluations.
- Collaborate with research teams on the full lifecycle of new capabilities, from defining measurement goals to interpreting results.
- Conduct experiments to characterize the impact of prompting, sampling, and scaffolding on performance across benchmarks.
- Communicate evaluation findings and results to internal stakeholders and external audiences as appropriate.
Requirements
- Strong Python programming skills, including experience with production or research infrastructure.
- Experience building or operating reliable, large-scale distributed systems or data pipelines.
- Excellent written and verbal communication skills, particularly in explaining technical results to non-specialists.
- Willingness to operate in an on-call or production-support capacity during live training runs.
- A commitment to understanding the societal impacts of AI and a desire to steer AI towards safety and benefit.