Software Engineer, Workload Enablement
Remote • San Francisco • FullTime
Posted 5mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
OpenAI is seeking a Software Engineer to join the Scaling team, which builds the architectural and engineering backbone for the company's infrastructure. This role focuses on enabling production workloads and end-to-end testing on new platforms. You will be responsible for creating test harnesses, developing platform stress benchmarks, and porting existing AI model training and inference workloads to new systems and hardware. A key aspect of this position involves analyzing performance, identifying bottlenecks, and characterizing the behavior of new compute, communication, storage, and control plane systems, including their failure modes.
Responsibilities
- Port and validate key inference and training workloads on new platforms, ensuring correctness, performance, and stability.
- Build benchmarks and stress tests to capture real end-to-end workload behavior across all system aspects (CPU, GPU, memory, networking, storage, thermals).
- Deep-dive into performance issues for distributed training/inference, including collective performance tuning, compute/communication overlap, kernel bottlenecks, and memory bandwidth.
- Create repeatable test harnesses for CI/lab environments that provide actionable outputs for pass/fail, performance scoring, and regression detection.
- Partner with systems and fleet bring-up engineers to ensure platforms are stable, performant, operationally usable, and scalable.
- Work cross-functionally with vendors and internal stakeholders by providing clear bug reports, minimal reproductions, and prioritized issue lists.
Requirements
- BS in CS/EE or equivalent practical experience.
- 5+ years in ML systems, performance engineering, distributed systems, or HPC.
- Strong hands-on experience with PyTorch and modern LLM training/inference stacks.
- Experience with large-scale distributed training concepts (data/model/pipeline parallel, collective comms).
- Experience with RDMA and debugging/optimizing comms libraries (NCCL or RCCL) and their hardware/network interaction.
- Proficiency in Python.
- Comfort reading/writing performance-critical code (C++/CUDA/HIP is a plus).
- Strong profiling and debugging skills (e.g., Nsight, rocprof, perf, flamegraphs).