Tech Lead Manager, Inference
$30k - $60k • Remote • Redwood City, CA • FullTime
Posted 1mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Luma is seeking a Tech Lead Manager for its Inference team to own the entire inference serving stack, encompassing routing, scheduling, and fleet-wide orchestration across thousands of GPUs, multiple clouds, and hardware vendors. This is a hands-on role where at least half of your time will be dedicated to architecting and building core platform components, making critical design decisions, and debugging complex incidents. You will also be responsible for leading, growing, and developing the inference engineering team, including hiring, coaching, and managing on-call rotations. The role involves setting the technical roadmap for serving infrastructure, owning platform SLOs and economics, and partnering with research to deploy new architectures and integrate serving into online RL and evaluation loops. The ideal candidate has extensive experience operating large-scale inference fleets and a genuine desire to remain hands-on in building and improving the serving stack.
Responsibilities
- Spend at least half your time hands-on architecting and building core platform components, owning design decisions, and debugging incidents.
- Lead, grow, and develop the inference engineering team, including hiring, coaching, and managing on-call, incident response, capacity planning, and postmortems.
- Set the technical roadmap for serving engines, routing, scheduling, autoscaling, caching, observability, and deployment.
- Own the platform's SLOs and economics, focusing on latency, availability, GPU utilization, and cost per generation.
- Partner with research to ship new architectures to production and integrate serving into online RL and evaluation loops.
- Build scheduling and queueing systems to efficiently utilize GPU resources against live traffic, cluster availability, and user priority.
Requirements
- 8+ years in large-scale distributed systems or ML infrastructure, with several years building and operating model-serving or inference platforms in production.
- Experience running inference platforms at the thousands-of-GPUs scale across multiple clusters or clouds.
- Technical leadership experience through rapid growth, with a desire to stay at least half hands-on.
- Deep expertise in LLM and foundation-model serving engines (e.g., vLLM, SGLang, TensorRT-LLM), ideally with experience modifying engine internals.
- Strong command of continuous batching, KV-cache management, quantization, speculative decoding, and parallelism strategies.
- Strong Python and PyTorch skills.
- Experience with Kubernetes at scale.
- Experience with queues, scheduling, traffic control, and fleet management.