Agent Post-Training Research
San Francisco • FullTime
Posted 2mo ago
About the job
OpenAI is seeking a highly motivated individual to join the Agent Post-Training team, responsible for developing frontier agents that can operate computers, collaborate with humans and other agents, and expand human potential. This role is intentionally broad, requiring individuals who can tackle ambiguous capability problems across research, engineering, data, evals, and product. You will work on models that write and debug code, use tools, call functions, operate computers, and complete valuable work for users. The position involves collaborating with researchers, engineers, product teams, and safety partners to define model capabilities, measure performance, and ship improvements into widely used products. This is a high-agency role for those who want their work to directly impact frontier models.
Responsibilities
- Design and execute experiments to enhance agentic model behavior in areas like coding, tool use, function calling, computer operation, multi-agent collaboration, long-horizon tasks, factuality, instruction following, and calibrated reasoning.
- Lead end-to-end improvements for the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis.
- Develop evals and environments to identify model failures, and convert these failures into training data, product fixes, or new research avenues.
- Collaborate with product teams (Codex, API/platform, ChatGPT) to understand user needs and translate product feedback into model enhancements.
- Work on early-training and alignment interventions, such as data mixtures, objectives, synthetic data, and eval loops, to shape agent behavior.
- Contribute to decisions regarding the inclusion of integrations, capabilities, and fixes in major model runs.
- Improve the infrastructure for large-scale training and launch, focusing on experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
- Undertake cross-functional projects involving model training, product infrastructure, and the production agent harness, including multi-agent systems and training in production-like environments.
- Debug complex failures in shipped or near-shipped models, transforming qualitative behavior into concrete hypotheses, experiments, and fixes.
Requirements
- Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, with the ability to learn quickly in new areas.
- Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.
- Excitement for open-ended problems with unclear paths and noisy signals, requiring both research taste and engineering execution.
- Focus on product impact and model behavior, with clear opinions on agent usefulness, reliability, honesty, tastefulness, and ease of use.
- Ability to translate vague behavioral problems into concrete experiments, including hypothesis definition, pipeline building, model execution, result analysis, and decision-making.
- Comfort working across research, product, infrastructure, data, evals, and safety boundaries, with clear communication skills.
- Adept at building load-bearing systems and processes when needed, even if the work is not glamorous.
- Desire to train and ship models that make agents genuinely useful for developers, enterprises, researchers, and everyday users.