Software Engineer, AI accelerator Runtime

Remote San Francisco FullTime

Posted 20d ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

OpenAI's Hardware organization is developing AI-native silicon and system-level solutions for advanced AI workloads. As a Software Engineer on this team, you will build the low-level device runtime that executes compiled programs efficiently on OpenAI's custom AI accelerators. This software is crucial for scheduling kernel launches, managing device memory, coordinating synchronization, and providing reliable abstractions to higher-level frameworks. You will collaborate closely with compiler, kernel, architecture, verification, and silicon teams, working at the intersection of software and hardware. Utilizing event-based, cycle-accurate simulation, you will develop runtime capabilities, diagnose performance and correctness issues, and contribute to hardware-software co-design before and after silicon is available.

Responsibilities

  • Design and implement the low-level device runtime for custom AI silicon.
  • Build kernel-launch scheduling, command submission, queueing, dependency tracking, and completion handling.
  • Manage device memory spaces, allocation, virtual-to-physical mappings, data movement, and lifetime across concurrent workloads.
  • Implement synchronization primitives, events, barriers, streams, and ordering guarantees.
  • Define interfaces between the runtime, drivers, firmware, compiler-generated code, kernels, and higher-level execution systems.
  • Use cycle-accurate simulators to develop, validate, debug, and performance-tune runtime behavior.
  • Diagnose concurrency, memory-ordering, deadlock, race, correctness, and performance issues.
  • Build tests, tracing, profiling, observability, and reproducible workloads for runtime correctness and performance.
  • Partner with architecture and silicon teams to improve hardware-software interfaces based on workload and simulator insights.

Requirements

  • Strong low-level systems programming experience in C, C++, Rust, or comparable environments.
  • Experience building runtimes, drivers, firmware, operating-system components, accelerator software, or adjacent systems infrastructure.
  • Understanding of concurrency, synchronization, asynchronous execution, queues, events, and memory-ordering semantics.
  • Understanding of memory management, address spaces, DMA, caching, coherency, and hardware-software interfaces.
  • Hands-on experience with event-based, cycle-accurate simulators or closely related architectural and performance models.
  • Ability to reason quantitatively about scheduling, latency, throughput, utilization, contention, and resource tradeoffs.
  • Skilled at debugging failures that span software abstractions, device interfaces, and hardware behavior.
  • Ability to work effectively across compiler, kernel, architecture, verification, firmware, and silicon teams.

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.