Staff AI Infrastructure Engineer
$30k - $60k • Remote • Redwood City, CA • FullTime
Posted 1mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Luma is seeking a Staff AI Infrastructure Engineer to own the reliability of its extensive GPU fleet, focusing on scheduling, efficiency, and resilience. This role requires deep systems knowledge to drive company-wide reliability and technical leadership. The position involves close-to-the-metal work with kernels, containers, schedulers, networking, storage, and GPU behavior, addressing challenges posed by high demand. You will also be responsible for setting technical standards and growing the engineering team.
Responsibilities
- Architect and operate large, heterogeneous GPU environments under extreme demand, improving utilization and performance.
- Resolve failures spanning hardware, OS, runtimes, and orchestration, and eliminate classes of instability.
- Define infrastructure and workload evolution as cluster size and concurrency grow, focusing on scheduling, placement, and resource management.
- Collaborate with research teams to build systems for new model capabilities and scale inference reliably and with low latency.
- Hire and develop exceptional systems and reliability engineers, setting high standards for depth, judgment, and production ownership.
- Shape product and research architecture through strong partnerships.
Requirements
- Deep expertise in Linux and distributed systems.
- Experience operating GPU or accelerator clusters in real production environments.
- Strong fluency in Kubernetes and modern open-source infrastructure.
- Comfort debugging across hardware, kernel, runtime, and orchestration, understanding system behavior under contention and at scale.
- Proficiency in writing code and building automation, with a focus on bottlenecks, failure modes, and trade-offs.
- Demonstrated judgment that engineers trust, especially during critical incidents.