Software Engineer, Compute Infrastructure

Remote San Francisco FullTime

Posted 1mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

We are seeking engineers to build the compute platform that powers OpenAI's research and products. This role involves designing, provisioning, scheduling, operating, and optimizing systems that connect accelerators, CPUs, networks, storage, data centers, and orchestration software into a cohesive experience for researchers and product teams. You will work across the entire stack, from capacity planning and cluster lifecycle to deep system optimization and developer experience, aiming to improve research velocity through enhancements in communication, scheduling, hardware efficiency, and debugging workflows. The specific focus will be matched to your strengths and interests, whether that's close to hardware, users, CaaS, agent infrastructure, or the control and data planes in between.

Responsibilities

  • Build and deeply optimize reliable system software for large-scale compute systems running demanding AI workloads.
  • Design and operate infrastructure across accelerators, CPUs, NICs, switches, networking protocols, storage, data centers, cluster orchestration, scheduling, and fleet health.
  • Profile, benchmark, and optimize training workloads across compute, memory, storage, networking, NCCL, collective communication, and cluster scheduling bottlenecks.
  • Create hardware-aware automation for faster and less error-prone provisioning, firmware/driver upgrades, incident response, and daily operations.
  • Build CaaS, agent infrastructure, profiling, observability, benchmarking, and platform tools to reduce friction for launching, debugging, and optimizing workloads.
  • Translate operational lessons into improved systems, stronger abstractions, and clearer ownership boundaries.
  • Collaborate with research, engineering, security, networking, hardware, and data center teams to enhance compute capacity and usability.

Requirements

  • Ability to reason carefully about complex systems.
  • Proficiency in writing durable software.
  • Skill in raising the quality and velocity of surrounding teams.
  • Experience building or operating distributed systems, infrastructure platforms, high-performance computing environments, large-scale networking systems, Kubernetes clusters, developer tools, or production systems with demanding reliability requirements.
  • Comfort working across software, hardware, networking, systems performance, reliability, and user needs.
  • A desire to make complex infrastructure understandable, observable, and usable.
  • Ability to diagnose difficult problems under operational pressure while investing in long-term engineering quality.
  • Skill in building leverage for others through APIs, automation, debugging tools, CaaS/agent infrastructure primitives, workflow improvements, or platform abstractions.

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.