Researcher, Agent Safety, Training and Evaluations
San Francisco • FullTime
Posted 10d ago
About the job
The Agent Safety team is dedicated to ensuring that advanced AI agents operate safely, make sound judgments, and align with user intentions. Our mission is to minimize the risk of severe unintended consequences from increasingly capable AI agents while maintaining their effectiveness and autonomy. This role involves training and evaluating frontier models to reduce harmful or misaligned agent actions, developing clear hypotheses, and executing independently in ambiguous situations. You will also mine incidents and build scalable systems for measurement, data processing, and evaluation to transform real failures into repeatable safety signals. Collaboration with post-training, capabilities, oversight, and pre-training teams is essential to integrate research-backed mitigations into large-scale training and agent systems.
Responsibilities
- Train and evaluate frontier models to reduce harmful or misaligned agent actions.
- Formulate clear hypotheses and execute independently through ambiguity.
- Mine incidents and build scalable measurement, data-processing, and evaluation systems.
- Collaborate with various teams to ship research-backed mitigations into training and agent systems.
Requirements
- Demonstrated strength in research engineering, ML engineering, quantitative research, or applied model research.
- Ability to own ambiguous projects end to end.
- Excellent technical execution across experimentation, data, evaluation, and/or infrastructure.
- Strong intuition for modern frontier-model research.
- Motivation by agent safety and eagerness to work on urgent, practical problems.
Benefits
- Hybrid work model (3 days in office per week)
- Relocation assistance for new employees