Agent Post-Training, API & Power Users

San Francisco FullTime

Posted 2mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Agent Engineer

About the job

The Agent Post-Training team is responsible for developing the frontier agents that OpenAI ships to the world, including models for Codex, ChatGPT, and the API. This role focuses on improving the capabilities, reliability, and product fit of OpenAI's agentic models specifically for power users and API developers. You will tackle ambiguous model behavior problems, translating them into concrete progress by enhancing aspects like tool use, planning, instruction following, and error recovery. The position involves close collaboration across research, engineering, data, evals, and product teams to refine model behavior for real-world workflows and API integrations, ultimately shaping the next generation of AI agents.

Responsibilities

  • Design and run experiments to improve model behavior in API and power-user workflows, including function calling, tool use, coding, planning, and error recovery.
  • Build evaluation frameworks, graders, and environments based on real developer workflows to identify failures and convert them into training data or interventions.
  • Partner with API and power-users to identify key behavior gaps and translate product signals into post-training improvements.
  • Enhance model composition and system integration, ensuring reliable tool use, adherence to developer intent, and effective handling of multi-step tasks.
  • Own end-to-end model behavior projects, from initial analysis to data generation, training, evaluation, and launch readiness.
  • Develop feedback loops using power-user traces and API usage patterns to discover new agentic model failures.
  • Contribute to decisions on which agentic capabilities and behavioral fixes are ready for major model runs.
  • Debug complex failures in models by analyzing traces, evaluations, training data, and product context.
  • Work on early-training and alignment interventions, including data mixtures, objectives, and synthetic data generation.
  • Improve the infrastructure for large-scale training and launch, focusing on velocity, reliability, and production readiness.
  • Undertake cross-functional projects involving model training, product infrastructure, and agent harnesses, such as multi-agent systems.

Requirements

  • Strong technical fundamentals in ML, software engineering, systems, statistics, or applied research, with the ability to learn across unfamiliar areas.
  • Hands-on experience with LLMs, post-training techniques, RL/RLHF/RLAIF, evaluations, graders, synthetic data, coding agents, tool-using agents, API products, or production ML systems.
  • A keen sense for model behavior, capable of forming hypotheses from transcripts, traces, or API interactions.
  • Excitement for ambiguous capability problems with noisy signals and qualitative failures.
  • Deep care for developer and expert-user experience, particularly how models perform in real workflows and API products.
  • Comfort working across research, product, infrastructure, data, evals, and safety boundaries, with clear communication skills.
  • Willingness to build load-bearing systems and processes when needed.
  • Desire to train and ship models that make agents genuinely useful for a wide range of users.

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.