Hardware / Software CoDesign Engineer - 3P

Remote San Francisco FullTime

Posted 5mo ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

OpenAI's Hardware organization is developing next-generation AI-native silicon and system-level solutions tailored for advanced AI workloads. As an Engineer on our hardware optimization and co-design team, you will collaborate with vendors to co-design future hardware for programmability and performance. You will work closely with kernel, compiler, and machine learning engineers to understand their needs regarding ML techniques, algorithms, and programming expressivity. Your role will involve evangelizing these constraints to vendors to influence future hardware architectures for efficient AI model training and inference, and optimizing large language models across devices, networking bottlenecks, and compute pipelines.

Responsibilities

  • Co-design future hardware for programmability and performance with hardware vendors.
  • Assist hardware vendors in developing optimal kernels and integrate support into our compiler.
  • Develop performance estimates for critical kernels across different hardware configurations to guide decisions on compute core and memory hierarchy features.
  • Build system performance models at various abstraction levels to inform decisions on scale-up, scale-out, and front-end networking.
  • Collaborate with machine learning engineers, kernel engineers, and compiler developers to understand their requirements for high-performance accelerators.
  • Manage communication and coordination with internal and external partners.
  • Influence the roadmaps of hardware partners to optimize for OpenAI's specific workloads.
  • Evaluate potential partners' accelerators and platforms.
  • Understand and influence roadmaps for hardware partners concerning datacenter networks, racks, and buildings as the team's scope expands.

Requirements

  • 4+ years of industry experience, including harnessing compute at scale and optimizing ML platform code for target hardware.
  • Strong experience in software/hardware co-design.
  • Deep understanding of GPU and/or other AI accelerators.
  • Experience with CUDA, Triton, or a similar accelerator programming language.
  • Experience driving Machine Learning accuracy with low-precision formats.
  • Experience with system performance modeling and analysis for ML model deployment optimization.
  • Strong coding skills in C/C++ and Python.
  • Familiarity with the fundamentals of deep learning computing and chip architecture/microarchitecture.
  • Ability to collaborate effectively with ML engineers, kernel writers, compiler developers, system engineers, and chip architects/microarchitects.

Benefits

  • Hybrid work model (3 days in office per week).
  • Relocation assistance for new employees.

About OpenAI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.