AI Systems Engineer, Codex Agents
San Francisco • FullTime
Posted 4mo ago
About the job
We are seeking engineers to build the AI systems that make Codex agents dependable in production. This role involves hands-on work across low-level systems and ML workflows, with the ability to debug Codex behavior end-to-end across the harness, model behavior, inference/runtime stack, GPU fleet, and product surface. You will collaborate with research, infrastructure, and product teams to design agent harness capabilities, run experiments, build frameworks for assessing production agent performance, and implement durable improvements from observed failures.
Responsibilities
- Design and build the core agent harness and execution loop for agents to interpret model outputs, use tools, execute code, and complete long-horizon tasks safely.
- Build sandboxing, isolation, orchestration, state, and workflow infrastructure for agents operating in real development environments.
- Develop evaluation, experimentation, and debugging systems to differentiate between harness issues, model behavior, inference/runtime issues, and product failures.
- Run ablations across prompts, model-facing interfaces, context construction, tool-use strategies, and harness behavior to enhance solve rate, reliability, latency, and cost.
- Improve observability, profiling, and diagnostics across the agent stack, from backend systems to inference, GPUs, and fleet capacity.
- Collaborate with research to make the harness trainable, measurable, and useful for improving frontier agentic models.
- Build shared primitives to make Codex faster, safer, more reliable, and easier for other teams and open-source users to build upon.
Requirements
- Experience building or operating production systems in distributed systems, infrastructure, developer tooling, sandboxing, virtualization, cloud platforms, or ML systems.
- Ability to work across layers including Rust systems code, Python configuration, APIs, agent orchestration, evals, logs/traces, inference behavior, runtime constraints, and user outcomes.
- Hands-on experience with LLM applications, coding agents, evals, model deployment, inference, compiler/runtime performance, or developer platforms.
- Strong focus on reliability, safety, performance, debuggability, and clean abstractions.
- Ability to debug from evidence and quickly resolve ambiguous production failures with practical, durable fixes.
- Willingness to work closely with research while shipping production changes.
- Demonstrated ability to write meaningful code, show strong ownership, and lead scoped or multi-team AI systems work.