Agent Post-Training, Artifacts Research

San Francisco FullTime

Posted 2mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Agent Engineer

About the job

The Agent Post-Training team is responsible for developing the frontier agents that OpenAI ships to the world, including models for Codex, ChatGPT, and the API. This role focuses on training models to produce polished, useful work products such as documents, spreadsheets, and reports, transforming vague user goals into finished artifacts with strong structure, visual taste, and correctness. The position involves owning improvements across the post-training stack, including RL, data pipelines, graders, reward signals, and evaluations. You will collaborate with researchers, engineers, product teams, and safety partners to shape major model runs and ship improvements into widely used products. This is a high-agency role for individuals eager to directly impact frontier models.

Responsibilities

  • Design and execute experiments to enhance agentic model behavior for complex software and plugins.
  • Lead end-to-end improvements in the post-training stack, encompassing RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis.
  • Develop evaluation metrics and environments to identify model failures and convert these into training data, product fixes, or research directions.
  • Collaborate with product teams to understand user needs and translate product feedback into model enhancements.
  • Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and eval loops.
  • Contribute to decisions regarding the inclusion of integrations, capabilities, and fixes in major model runs.
  • Enhance the infrastructure for large-scale training and deployment, focusing on experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
  • Undertake cross-functional projects involving model training, product infrastructure, and agent deployment systems.
  • Debug complex failures in deployed or near-deployment models, converting qualitative observations into concrete hypotheses, experiments, and solutions.

Requirements

  • Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, with the ability to learn quickly.
  • Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.
  • Ability to thrive in open-ended problem spaces requiring research taste and engineering execution.
  • Focus on product impact and model behavior, with clear opinions on agent usefulness, reliability, honesty, taste, and usability.
  • Capability to translate vague behavioral problems into concrete experiments, including hypothesis definition, pipeline building, model execution, result analysis, and decision-making.
  • Comfort working across research, product, infrastructure, data, evals, and safety domains, with clear communication skills.
  • Experience building load-bearing systems and processes.
  • Desire to train and ship models that make agents genuinely useful for developers, enterprises, researchers, and everyday users.
  • Prior background in consulting, finance, marketing, operations, or data science is beneficial.

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.