Research Engineer, Frontier Evals & Environments

San Francisco FullTime

Posted 1y ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Research Engineer

About the job

OpenAI is seeking a Research Engineer for Frontier Evals & Environments to help build north star model environments that drive progress towards safe AGI/ASI. This role will directly guide the research programs of ambitious training runs, with past open-sourced evaluations including GDPval, SWE-bench Verified, MLE-bench, PaperBench, and SWE-Lancer. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to define model run objectives, measure outcomes, and ship improvements into widely used products. This is a high-agency role for individuals passionate about seeing their work directly impact frontier models and steering them towards beneficial outcomes.

Responsibilities

  • Create ambitious RL environments to push models to their limits and measure frontier model capabilities, skills, and behaviors.
  • Develop new methodologies for automatically exploring model behavior.
  • Dive deep into the science of measurement, focusing on scalability, reliability, and variance of evaluation methodologies.
  • Help steer training for large training runs and gain early insights into future model capabilities.
  • Design scalable systems and processes to support continuous evaluation.
  • Build self-improvement loops to automate model understanding.

Requirements

  • Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, with the ability to learn quickly.
  • Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.
  • Excitement for open-ended problems with unclear paths and noisy signals, requiring both research taste and engineering execution.
  • Focus on product impact and model behavior, with opinions on what makes an agent useful, reliable, honest, tasteful, and easy to work with.
  • Ability to translate vague behavioral problems into concrete experiments, including hypothesis definition, pipeline building, model execution, result analysis, and decision-making.
  • Comfort working across research, product, infrastructure, data, evals, and safety boundaries, with clear communication skills.
  • Willingness to build load-bearing systems and processes when needed, even if not glamorous.
  • Desire to train and ship models that make agents genuinely useful for developers, enterprises, researchers, and everyday users.

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.