Agent Post-Training, Context Research
San Francisco • FullTime
Posted 2mo ago
About the job
The Agent Post-Training team is responsible for developing the frontier agents that OpenAI releases. This role focuses on scaling compute spent on context, enabling a new paradigm of model training with a clear product interface for iterative deployment. You will collaborate with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to define model run content, measure outcomes, and integrate improvements into widely used products. This is a high-agency position for individuals eager to directly influence frontier models.
Responsibilities
- Design and execute experiments to enhance the scaling of compute on context.
- Lead end-to-end improvements for the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model behavior analysis.
- Develop evals and environments to identify model failures, converting these into training data, product fixes, or new research avenues.
- Collaborate with product teams to understand user needs and translate product signals into model enhancements.
- Work on early-training and alignment interventions, such as data mixtures, objectives, synthetic data, and eval loops.
- Contribute to decisions on integrating capabilities and fixes into major model runs.
- Enhance the machinery for large-scale training and launch, focusing on experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
- Undertake cross-functional projects involving model training, product infrastructure, and agent production harnesses.
- Debug complex failures in shipped or near-shipped models, transforming qualitative behavior into concrete hypotheses, experiments, and fixes.
Requirements
- Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, with the ability to learn quickly.
- Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.
- Excitement for open-ended problems with unclear paths and noisy signals, requiring both research taste and engineering execution.
- Focus on product impact and model behavior, with opinions on agent usefulness, reliability, honesty, tastefulness, and ease of use.
- Ability to translate vague behavioral problems into concrete experiments, including hypothesis definition, pipeline building, model execution, result analysis, and decision-making.
- Comfort working across research, product, infrastructure, data, evals, and safety boundaries, with clear communication skills.
- Enjoyment in building load-bearing systems and processes when needed.
- Desire to train and ship models that make agents genuinely useful for developers, enterprises, researchers, and everyday users.