Software Engineer, Compute Infrastructure
Remote • San Francisco • FullTime
Posted 1mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
We are seeking engineers to build the compute platform that powers OpenAI's research and products. This role involves designing, provisioning, scheduling, operating, and optimizing systems that connect accelerators, CPUs, networks, storage, data centers, and orchestration software into a cohesive experience for researchers and product teams. You will work across the entire stack, from capacity planning and cluster lifecycle to deep system optimization and developer experience, aiming to improve research velocity through enhancements in communication, scheduling, hardware efficiency, and debugging workflows. The specific focus will be matched to your strengths and interests, whether that's close to hardware, users, CaaS, agent infrastructure, or the control and data planes in between.
Responsibilities
- Build and deeply optimize reliable system software for large-scale compute systems running demanding AI workloads.
- Design and operate infrastructure across accelerators, CPUs, NICs, switches, networking protocols, storage, data centers, cluster orchestration, scheduling, and fleet health.
- Profile, benchmark, and optimize training workloads across compute, memory, storage, networking, NCCL, collective communication, and cluster scheduling bottlenecks.
- Create hardware-aware automation for faster and less error-prone provisioning, firmware/driver upgrades, incident response, and daily operations.
- Build CaaS, agent infrastructure, profiling, observability, benchmarking, and platform tools to reduce friction for launching, debugging, and optimizing workloads.
- Translate operational lessons into improved systems, stronger abstractions, and clearer ownership boundaries.
- Collaborate with research, engineering, security, networking, hardware, and data center teams to enhance compute capacity and usability.
Requirements
- Ability to reason carefully about complex systems.
- Proficiency in writing durable software.
- Skill in raising the quality and velocity of surrounding teams.
- Experience building or operating distributed systems, infrastructure platforms, high-performance computing environments, large-scale networking systems, Kubernetes clusters, developer tools, or production systems with demanding reliability requirements.
- Comfort working across software, hardware, networking, systems performance, reliability, and user needs.
- A desire to make complex infrastructure understandable, observable, and usable.
- Ability to diagnose difficult problems under operational pressure while investing in long-term engineering quality.
- Skill in building leverage for others through APIs, automation, debugging tools, CaaS/agent infrastructure primitives, workflow improvements, or platform abstractions.