Software Engineer, Data Infrastructure - Research
San Francisco • FullTime
Posted 1y ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
We are looking for an engineer to design and implement the dataset infrastructure that powers OpenAI’s next-generation training stack. You will be responsible for building standardized dataset interfaces, scaling pipelines across thousands of GPUs, and proactively testing performance bottlenecks. In this role, you will collaborate closely with the multimodal researchers, and other infra groups to ensure datasets are unified, efficient, and easy to consume.
Responsibilities
- Design and maintain standardized dataset APIs, including for multimodal data that cannot fit in memory.
- Build proactive testing and scale validation pipelines for dataset loading at GPU scale.
- Collaborate with teammates to integrate datasets seamlessly into training and inference pipelines.
- Document and maintain dataset interfaces for discoverability, consistency, and ease of adoption.
- Establish safeguards and validation systems to ensure datasets remain reproducible and unchanged.
- Debug and resolve performance bottlenecks in distributed dataset loading.
- Provide visualization and inspection tools to surface errors, bugs, or bottlenecks in datasets.
Requirements
- Strong engineering fundamentals with experience in distributed systems, data pipelines, or infrastructure.
- Experience building APIs, modular code, and scalable abstractions.
- Comfortable debugging bottlenecks across large fleets of machines.
- Experience building infrastructure that is reliable and scalable.
- Collaborative, humble, and excited to own a foundational part of the ML stack.
- Background knowledge in data math, probability, or distributed data theory (bonus).
- Experience with GPU-scale distributed systems or dataset scaling for real-time data (bonus).