Software Engineer, Model Runtime
Remote • San Francisco • FullTime
Posted 21d ago
About the job
OpenAI's Hardware organization is developing AI-native silicon and system-level solutions for advanced AI workloads, aiming to power the next generation of frontier models. As a Software Engineer on this team, you will build the model runtime within the inference engine that executes complex models at scale on OpenAI's custom silicon. This runtime will bridge the gap between models and the serving software stack, translating inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. Your work will involve designing a production-grade runtime comparable to systems like vLLM and SGLang, customized for OpenAI's AI accelerators, and will directly influence how new model capabilities map onto the platform and how quickly custom silicon delivers performance.
Responsibilities
- Design and implement the LLM inference runtime for frontier models on custom silicon.
- Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference.
- Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization.
- Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization for diverse model architectures and serving workloads.
- Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks.
- Enable new model features, execution patterns, numerical formats, and hardware capabilities in a reliable production runtime.
- Create profiling, observability, benchmarking, and performance-modeling tools.
- Debug complex correctness, performance, and reliability issues across the hardware-software stack.
- Translate workload insights into requirements for future silicon and system architecture generations.
Requirements
- Strong systems programming experience in C++, Rust, Python, or comparable performance-oriented environments.
- Experience building or optimizing runtimes, distributed systems, compilers, kernels, model-serving infrastructure, or adjacent systems software.
- Understanding of modern LLM inference, including prefill and decode behavior, batching, KV-cache tradeoffs, and model parallelism.
- Ability to quantitatively reason about latency, throughput, compute intensity, memory bandwidth, communication, and utilization.
- Comfort profiling and debugging performance across multiple layers of a hardware-software stack.
- Ability to design clean abstractions while retaining low-level control for specialized hardware performance.
- Effectiveness in collaborating across model, systems, compiler, kernel, and hardware teams to resolve ambiguous technical problems.
- Commitment to production quality, including correctness, observability, reliability, maintainability, and graceful behavior at scale.