Performance Engineer (Inference, Training & GPU)
$200k - $300k • San Francisco
Posted 20d ago
About the job
World Labs is a frontier AI research and product company advancing spatial intelligence. We are seeking a Performance Engineer to optimize our large generative world models for both training and serving, ensuring they run as fast as the hardware allows. This hands-on, individual-contributor role involves identifying and eliminating performance bottlenecks across the entire stack, from low-level kernel optimization to fleet-wide serving efficiency. You will work closely with researchers to accelerate their models and productionize them for efficient deployment.
Responsibilities
- Optimize inference and serving for latency, throughput, batching, caching, and scheduling.
- Write and tune GPU kernels (CUDA, Triton) for performance-critical paths.
- Optimize training throughput and GPU utilization, including parallelism strategies and communication/compute overlap.
- Develop performance models, profiling workflows, and observability tools.
- Ensure numerical correctness across precision, kernel, and hardware changes.
- Partner with researchers to productionize models and accelerate experiments.
- Contribute to distributed systems for training and inference, focusing on maximizing GPU efficiency.
Requirements
- Strong performance engineering fundamentals, including profiling, roofline analysis, and root-cause investigation.
- Deep GPU programming and optimization experience with CUDA and/or Triton.
- Hands-on experience optimizing inference and serving for large models (batching, caching, quantization).
- Hands-on experience optimizing training performance (parallelism, distributed communication, mixed precision).
- Working knowledge of ML framework internals (PyTorch, JAX) and compiler paths.
- Strong proficiency in Python, with ability to use C++/CUDA.
- High-ownership mindset focused on tangible performance improvements.