Agent Post-Training, Context Research

San Francisco FullTime

Posted 2mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Agent Engineer

About the job

The Agent Post-Training team is responsible for developing the frontier agents that OpenAI releases. This role focuses on scaling compute spent on context, enabling a new paradigm of model training with a clear product interface for iterative deployment. You will collaborate with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to define model run content, measure outcomes, and integrate improvements into widely used products. This is a high-agency position for individuals eager to directly influence frontier models.

Responsibilities

  • Design and execute experiments to enhance the scaling of compute on context.
  • Lead end-to-end improvements for the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model behavior analysis.
  • Develop evals and environments to identify model failures, converting these into training data, product fixes, or new research avenues.
  • Collaborate with product teams to understand user needs and translate product signals into model enhancements.
  • Work on early-training and alignment interventions, such as data mixtures, objectives, synthetic data, and eval loops.
  • Contribute to decisions on integrating capabilities and fixes into major model runs.
  • Enhance the machinery for large-scale training and launch, focusing on experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
  • Undertake cross-functional projects involving model training, product infrastructure, and agent production harnesses.
  • Debug complex failures in shipped or near-shipped models, transforming qualitative behavior into concrete hypotheses, experiments, and fixes.

Requirements

  • Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, with the ability to learn quickly.
  • Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.
  • Excitement for open-ended problems with unclear paths and noisy signals, requiring both research taste and engineering execution.
  • Focus on product impact and model behavior, with opinions on agent usefulness, reliability, honesty, tastefulness, and ease of use.
  • Ability to translate vague behavioral problems into concrete experiments, including hypothesis definition, pipeline building, model execution, result analysis, and decision-making.
  • Comfort working across research, product, infrastructure, data, evals, and safety boundaries, with clear communication skills.
  • Enjoyment in building load-bearing systems and processes when needed.
  • Desire to train and ship models that make agents genuinely useful for developers, enterprises, researchers, and everyday users.

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.