Agent Post-Training, Connectors Research

San Francisco FullTime

Posted 2mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Agent Engineer

About the job

The Agent Post-Training team is responsible for developing the frontier agents that OpenAI ships to the world, including models for Codex, ChatGPT, and the API. This role focuses on teaching models to interface with professional software using code, APIs, and tools. You will enable agents to operate across applications like Slack, Google Workspace, GitHub, and Salesforce, taking actions within a user's digital context to find information, update systems, coordinate work, and complete multi-step workflows. The position involves training models to leverage productivity and enterprise software, turning connected tools into a powerful action surface for agents. You will collaborate with researchers, engineers, product teams, and safety partners to define model capabilities, measure performance, and ship improvements into products.

Responsibilities

  • Design and run experiments to improve agentic model behavior for complex software and plugins.
  • Own end-to-end improvements to the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis.
  • Build evals and environments to identify model failures, and convert these into training data, product fixes, or research directions.
  • Partner with Codex and ChatGPT product teams to translate product signal into model improvements.
  • Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and eval loops.
  • Help decide which integrations, capabilities, and fixes are ready for major model runs.
  • Improve the machinery for large-scale training and launch, focusing on experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
  • Undertake cross-functional projects involving model training, product infrastructure, and agent harness, such as multi-agent systems or training against production-like environments.
  • Debug failures in shipped or near-shipped models and translate qualitative behavior into concrete hypotheses, experiments, and fixes.

Requirements

  • Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, with the ability to learn quickly.
  • Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.
  • Excitement for open-ended problems with unclear paths and noisy signals, requiring both research taste and engineering execution.
  • Focus on product impact and model behavior, with opinions on agent usefulness, reliability, honesty, taste, and ease of use.
  • Ability to move from a vague behavioral problem to a concrete experiment, including hypothesis definition, pipeline building, model running, result analysis, and decision-making.
  • Comfort working across research, product, infrastructure, data, evals, and safety boundaries, with clear communication skills.
  • Enjoyment in building load-bearing systems and processes when needed.
  • Desire to train and ship models that make agents genuinely useful for developers, enterprises, researchers, and everyday users.

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.