Systems Generalist, GPT Infrastructure
Remote • San Francisco • FullTime
Posted 1mo ago
About the job
OpenAI is seeking an experienced systems generalist to join the GPT Infrastructure team. This role involves building an automated inference optimization platform that handles workload analysis, target hardware profiling, compiler and runtime context, and trusted verification. You will design and implement both the OpenAI-hosted control plane and the partner-side software responsible for evaluating candidates on real accelerator hardware. The goal is to create a reliable system that ensures reproducible performance results and maintains clear trust boundaries for sensitive information, turning a powerful research workflow into a scalable product.
Responsibilities
- Design, build, and operate durable APIs and control-plane services for optimization campaigns, including scheduling, retries, budgets, checkpoints, artifact lineage, and observability.
- Build secure partner-side runner and grader software to compile, execute, verify, and benchmark artifacts on third-party accelerator hardware.
- Integrate hardware profiles, ISA and toolchain context, compilers, runtimes, and inference-serving engines into a repeatable optimization workflow.
- Transform research prototypes into reliable product surfaces with clear contracts, debuggable failure modes, reproducible outputs, and excellent developer ergonomics.
- Develop correctness and performance evaluation systems covering latency, throughput, memory use, utilization, and cost efficiency.
- Build artifact, provenance, and qualification workflows for safe review and deployment of optimized kernels, binaries, configurations, and reports.
- Collaborate with Research, Inference Engineering, Infrastructure, Security, Product, and Strategic Partnerships to deliver production-ready solutions.
- Drive technical architecture and execution for ambiguous, cross-functional initiatives connecting OpenAI systems with partner environments.
Requirements
- 8+ years of professional software engineering experience in large-scale distributed systems, infrastructure platforms, or cloud services, or equivalent depth of experience.
- Strong programming skills in C++, Python, Go, or Rust.
- Experience designing and operating highly available backend systems, APIs, job orchestration systems, or durable workflows for production workloads.
- Strong understanding of distributed systems, Linux, networking, storage, containers, and modern cloud architectures.
- Experience debugging complex systems and using measurement, profiling, and benchmarks for engineering decisions.
- Proven ability to lead complex technical initiatives as a senior individual contributor and work effectively across organizational boundaries.