Agent Post-Training, Frontier Evals and Environments Research

San Francisco FullTime

Posted 2mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Agent Engineer

About the job

The Agent Post-Training team at OpenAI is responsible for developing the frontier agents that are shipped to the world. This role focuses on building "north star" model environments to drive progress towards safe AGI/ASI, directly guiding ambitious research programs. You will collaborate with researchers, engineers, product, infrastructure, and safety teams to define model run objectives, measure outcomes, and integrate improvements into widely used products. This is a high-agency position for individuals eager to directly influence cutting-edge AI models and witness their rapid advancement.

Responsibilities

  • Create ambitious RL environments to test and measure frontier model capabilities, skills, and behaviors.
  • Develop new methodologies for automatically exploring model behavior.
  • Investigate the science of measurement, focusing on scalability, reliability, and variance of evaluation methodologies.
  • Assist in steering training for large-scale training runs.
  • Design scalable systems and processes to support continuous evaluation.
  • Build self-improvement loops to automate model understanding.

Requirements

  • Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, with the ability to learn quickly.
  • Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.
  • Excitement for open-ended problems requiring research taste and engineering execution.
  • Focus on product impact and model behavior, not just benchmark scores.
  • Ability to translate vague behavioral problems into concrete experiments, including hypothesis definition, pipeline building, model execution, and result analysis.
  • Comfort working across research, product, infrastructure, data, evals, and safety boundaries.
  • Clear communication skills with diverse teams.
  • Willingness to build load-bearing systems and processes when needed.
  • Desire to train and ship models that make agents genuinely useful for various user groups.

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.