Research Engineer, Frontier Evals & Environments
San Francisco • FullTime
Posted 1y ago
About the job
OpenAI is seeking a Research Engineer for Frontier Evals & Environments to help build north star model environments that drive progress towards safe AGI/ASI. This role will directly guide the research programs of ambitious training runs, with past open-sourced evaluations including GDPval, SWE-bench Verified, MLE-bench, PaperBench, and SWE-Lancer. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to define model run objectives, measure outcomes, and ship improvements into widely used products. This is a high-agency role for individuals passionate about seeing their work directly impact frontier models and steering them towards beneficial outcomes.
Responsibilities
- Create ambitious RL environments to push models to their limits and measure frontier model capabilities, skills, and behaviors.
- Develop new methodologies for automatically exploring model behavior.
- Dive deep into the science of measurement, focusing on scalability, reliability, and variance of evaluation methodologies.
- Help steer training for large training runs and gain early insights into future model capabilities.
- Design scalable systems and processes to support continuous evaluation.
- Build self-improvement loops to automate model understanding.
Requirements
- Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, with the ability to learn quickly.
- Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.
- Excitement for open-ended problems with unclear paths and noisy signals, requiring both research taste and engineering execution.
- Focus on product impact and model behavior, with opinions on what makes an agent useful, reliable, honest, tasteful, and easy to work with.
- Ability to translate vague behavioral problems into concrete experiments, including hypothesis definition, pipeline building, model execution, result analysis, and decision-making.
- Comfort working across research, product, infrastructure, data, evals, and safety boundaries, with clear communication skills.
- Willingness to build load-bearing systems and processes when needed, even if not glamorous.
- Desire to train and ship models that make agents genuinely useful for developers, enterprises, researchers, and everyday users.