Performance Engineer, Inference
Bengaluru • FullTime
Posted 1mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Sarvam is seeking a Performance Engineer specializing in Inference to join their Performance Engineering team. This role focuses on integrating and optimizing model serving stacks for large, distributed models across a fleet of GPUs. You will be responsible for the end-to-end production serving path, modifying and extending existing serving runtimes like SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM. The position involves building and training custom speculative decoding models and ensuring the performance and cost-efficiency of the serving infrastructure. You will collaborate closely with model, kernel, and SRE teams to achieve critical performance metrics such as latency, throughput, and GPU utilization.
Responsibilities
- Own Sarvam's production serving path for large distributed models end to end.
- Integrate artifacts from model and kernel teams into a running multi-node, multi-tenant stack.
- Build and train custom speculative decoding models (draft models, distillation, acceptance-rate tuning).
- Produce and defend latency and throughput numbers for company planning.
- Collaborate with architecture, kernels, model, and SRE teams on co-design and optimization.
- Monitor and manage key performance indicators including TTFT, TPOT, throughput, GPU utilization, and cost per million tokens.
- Take on-call ownership of inference SLOs.
Requirements
- 5+ years in ML systems, with 2+ years on inference serving at production scale.
- Experience serving 100B+ parameter models in production across multi-node tensor, pipeline, or expert parallelism.
- Source-level fluency in SGLang, vLLM, Dynamo, or TensorRT-LLM, with ability to modify scheduler, KV allocator, or disaggregation path.
- Experience with distributed serving at operating-and-extending depth (disaggregated prefill-decode, distributed KV/cache transfer, cross-node routing and scheduling).
- Competency in building and training speculative decoding models.
- Deep understanding of KV cache internals (block tables, copy-on-write, prefix sharing, fragmentation).
- Working command of TP/PP/EP, NCCL primitives, and their interaction with the scheduler.
- Experience with multi-tenant serving (model co-location, MIG/MPS isolation).
- Profiling fluency with Nsight Systems, framework tracing, and py-spy/perf.
- C++ and CUDA at a read-and-modify level.