Software Engineer, GPT Infrastructure
Remote • San Francisco • FullTime
Posted 4mo ago
About the job
We are seeking a software engineer to help build the platform that qualifies and optimizes inference workloads across heterogeneous compute environments. You will develop both OpenAI-hosted services and secure partner-side software for running long-lived optimization workflows. These workflows generate candidate kernels, runtime configurations, and serving-stack changes; compile and execute them on target hardware; verify their correctness; measure their performance; and use the results to guide further optimization. You will work across model architecture, distributed execution, compilers, runtimes, networking, and accelerator systems. A central part of the role is turning research prototypes and one-off hardware bring-up efforts into reliable, reusable infrastructure with clear contracts, reproducible results, strong observability, and well-defined security boundaries.
Responsibilities
- Design, build, and operate APIs and control-plane services for workload qualification and optimization campaigns.
- Build secure partner-side execution and evaluation software for compiling, running, verifying, profiling, and benchmarking artifacts on accelerator hardware.
- Integrate model workloads, hardware profiles, compiler toolchains, runtimes, serving engines, and distributed-execution backends into a repeatable platform.
- Develop correctness and performance evaluation systems for output fidelity, latency, throughput, memory footprint, accelerator utilization, communication efficiency, scaling behavior, and cost efficiency.
- Automate the generation, evaluation, and improvement of kernels, runtime configurations, parallelization strategies, and serving-stack changes.
- Diagnose performance and correctness issues across model code, kernels, compilers, runtimes, memory systems, networking, collective communication, and hardware.
- Build artifact-management, provenance, regression-testing, and qualification workflows.
- Turn experimental research workflows into reliable product surfaces with clear interfaces and actionable failure modes.
- Collaborate with Research, Inference Engineering, Runtime and Compiler teams, Infrastructure, Security, Product, and Strategic Partnerships.
- Drive technical architecture and execution across ambiguous initiatives spanning OpenAI systems and partner environments.
Requirements
- Strong software engineering experience building distributed systems, infrastructure platforms, production services, developer platforms, or orchestration systems.
- Proficiency in systems-oriented languages such as Python, C++, Go, or Rust.
- Experience designing and operating APIs, job orchestration systems, durable workflows, or large-scale backend services.
- Strong understanding of Linux, networking, storage, containers, distributed execution, and modern infrastructure architectures.
- Ability to reason about model execution and diagnose problems across software and hardware boundaries.
- Experience using profiling, tracing, benchmarking, and measurement to guide engineering decisions.
- Strong ownership and ability to work effectively across research, engineering, security, product, and external-partner teams.