Performance Engineer, Inference

Bengaluru FullTime

Posted 1mo ago

Job Location

Bengaluru

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Sarvam is seeking a Performance Engineer specializing in Inference to join their Performance Engineering team. This role focuses on integrating and optimizing model serving stacks for large, distributed models across a fleet of GPUs. You will be responsible for the end-to-end production serving path, modifying and extending existing serving runtimes like SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM. The position involves building and training custom speculative decoding models and ensuring the performance and cost-efficiency of the serving infrastructure. You will collaborate closely with model, kernel, and SRE teams to achieve critical performance metrics such as latency, throughput, and GPU utilization.

Responsibilities

  • Own Sarvam's production serving path for large distributed models end to end.
  • Integrate artifacts from model and kernel teams into a running multi-node, multi-tenant stack.
  • Build and train custom speculative decoding models (draft models, distillation, acceptance-rate tuning).
  • Produce and defend latency and throughput numbers for company planning.
  • Collaborate with architecture, kernels, model, and SRE teams on co-design and optimization.
  • Monitor and manage key performance indicators including TTFT, TPOT, throughput, GPU utilization, and cost per million tokens.
  • Take on-call ownership of inference SLOs.

Requirements

  • 5+ years in ML systems, with 2+ years on inference serving at production scale.
  • Experience serving 100B+ parameter models in production across multi-node tensor, pipeline, or expert parallelism.
  • Source-level fluency in SGLang, vLLM, Dynamo, or TensorRT-LLM, with ability to modify scheduler, KV allocator, or disaggregation path.
  • Experience with distributed serving at operating-and-extending depth (disaggregated prefill-decode, distributed KV/cache transfer, cross-node routing and scheduling).
  • Competency in building and training speculative decoding models.
  • Deep understanding of KV cache internals (block tables, copy-on-write, prefix sharing, fragmentation).
  • Working command of TP/PP/EP, NCCL primitives, and their interaction with the scheduler.
  • Experience with multi-tenant serving (model co-location, MIG/MPS isolation).
  • Profiling fluency with Nsight Systems, framework tracing, and py-spy/perf.
  • C++ and CUDA at a read-and-modify level.

About sarvam

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.