Software Engineer, Workload Enablement

Remote San Francisco FullTime

Posted 5mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

OpenAI is seeking a Software Engineer to join the Scaling team, which builds the architectural and engineering backbone for the company's infrastructure. This role focuses on enabling production workloads and end-to-end testing on new platforms. You will be responsible for creating test harnesses, developing platform stress benchmarks, and porting existing AI model training and inference workloads to new systems and hardware. A key aspect of this position involves analyzing performance, identifying bottlenecks, and characterizing the behavior of new compute, communication, storage, and control plane systems, including their failure modes.

Responsibilities

  • Port and validate key inference and training workloads on new platforms, ensuring correctness, performance, and stability.
  • Build benchmarks and stress tests to capture real end-to-end workload behavior across all system aspects (CPU, GPU, memory, networking, storage, thermals).
  • Deep-dive into performance issues for distributed training/inference, including collective performance tuning, compute/communication overlap, kernel bottlenecks, and memory bandwidth.
  • Create repeatable test harnesses for CI/lab environments that provide actionable outputs for pass/fail, performance scoring, and regression detection.
  • Partner with systems and fleet bring-up engineers to ensure platforms are stable, performant, operationally usable, and scalable.
  • Work cross-functionally with vendors and internal stakeholders by providing clear bug reports, minimal reproductions, and prioritized issue lists.

Requirements

  • BS in CS/EE or equivalent practical experience.
  • 5+ years in ML systems, performance engineering, distributed systems, or HPC.
  • Strong hands-on experience with PyTorch and modern LLM training/inference stacks.
  • Experience with large-scale distributed training concepts (data/model/pipeline parallel, collective comms).
  • Experience with RDMA and debugging/optimizing comms libraries (NCCL or RCCL) and their hardware/network interaction.
  • Proficiency in Python.
  • Comfort reading/writing performance-critical code (C++/CUDA/HIP is a plus).
  • Strong profiling and debugging skills (e.g., Nsight, rocprof, perf, flamegraphs).

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.