Research Engineer, Codex
San Francisco • FullTime
Posted 1y ago
About the job
The Codex Research team at OpenAI is responsible for developing frontier AI agents, including the models behind Codex, ChatGPT, and other advanced products. This role focuses on improving the capabilities, reliability, and product fit of these agentic models. You will have the opportunity to own research directions, build critical training infrastructure, create evaluation methods, or drive capabilities from initial concept through experimentation and launch. This is a broad, high-agency role for individuals who can tackle ambiguous problems across research, engineering, data, evals, and product, and who are excited about building AI that can act in the world, write code, use tools, and collaborate with users and other agents.
Responsibilities
- Design and run experiments to improve agentic model behavior in areas like coding, tool use, function calling, and multi-agent collaboration.
- Own end-to-end improvements for the post-training stack, including RL, data pipelines, graders, and reward signals.
- Build evaluation environments to identify model failures and convert these into training data or product fixes.
- Partner with product teams to translate user needs and product signals into model improvements.
- Work on early-training and alignment interventions, including data mixtures and synthetic data generation.
- Help decide which integrations and capabilities are ready for major model runs.
- Improve the infrastructure for large-scale training and launch, focusing on velocity, reliability, and cost.
- Undertake cross-functional projects involving model training, product infrastructure, and agent systems.
- Debug failures in shipped models and translate qualitative behavior into concrete hypotheses and experiments.
Requirements
- Strong technical fundamentals in machine learning, software engineering, systems, or statistics.
- Ability to learn quickly across different technical domains.
- Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, or production ML systems.
- Excitement for open-ended problems requiring research taste and engineering execution.
- Focus on product impact and model behavior, not just benchmarks.
- Ability to translate vague behavioral problems into concrete experiments and analyses.
- Comfort working across research, product, infrastructure, data, evals, and safety boundaries.
- Experience building load-bearing systems and processes.
- Desire to train and ship models that make agents useful for developers, enterprises, and users.