Agent Post-Training, Computer Use Research
San Francisco • FullTime
Posted 2mo ago
About the job
The Agent Post-Training team is responsible for developing the frontier agents that OpenAI ships to the world. This role focuses on training models to operate computers, enabling them to navigate browsers and desktops, utilize tools, reason through complex workflows, and collaborate with users and other agents. The work involves a blend of frontier model training, product behavior, evaluation, and systems engineering, directly influencing the computer-use capabilities of OpenAI's next-generation agents. You will collaborate with researchers, engineers, product teams, and safety partners to define model training runs, measure outcomes, and ship improvements to products used by millions.
Responsibilities
- Design and execute experiments to enhance agentic model behavior for complex computer use, including desktop and browser interactions.
- Lead end-to-end improvements for the post-training stack, encompassing RL, data pipelines, graders, reward signals, evaluations, diagnostics, and model behavior analysis.
- Develop evaluations and environments to identify model failures, converting these into training data, product fixes, or research directions.
- Collaborate with product teams to understand user needs and translate product signals into model enhancements.
- Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and eval loops.
- Contribute to decisions regarding the inclusion of integrations, capabilities, and fixes in major model runs.
- Enhance the infrastructure for large-scale training and launch, focusing on experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
- Undertake cross-functional projects involving model training, product infrastructure, and agent production systems, such as multi-agent systems or training in production-like environments.
- Debug complex failures in shipped or near-shipped models, transforming qualitative behavior into concrete hypotheses, experiments, and solutions.
Requirements
- Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, with the ability to learn quickly in new areas.
- Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.
- Ability to tackle open-ended problems with unclear paths and noisy signals, requiring both research taste and engineering execution.
- Focus on product impact and model behavior, with clear opinions on agent usefulness, reliability, honesty, tastefulness, and ease of use.
- Proficiency in translating vague behavioral problems into concrete experiments, including hypothesis definition, pipeline building, model execution, result analysis, and decision-making.
- Comfort working across research, product, infrastructure, data, evals, and safety boundaries, with clear communication skills for each group.
- Experience building load-bearing systems and processes when required by the team.
- Desire to train and ship models that make agents genuinely useful for developers, enterprises, researchers, and everyday users.