Performance Engineer, Kernels
Bengaluru • FullTime
Posted 1mo ago
Job Location
Bengaluru
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Sarvam is seeking a highly specialized Performance Engineer focused on the kernel layer for its advanced AI infrastructure. This role is critical for optimizing the performance of multiple model families, including LLMs, Mixture-of-Experts, and multimodal models, across a large fleet of H100/H200/B200 GPUs. You will be responsible for authoring custom CUDA, DSL-based, and PTX kernels to bridge performance gaps left by stock libraries. This is a high-leverage position where your contributions will directly impact production p99 latency and cost efficiency, requiring engineers who have a proven track record of shipping kernels that outperform established baselines on real-world workloads.
Responsibilities
- Author custom CUDA, DSL-based, and PTX kernels to optimize performance beyond stock libraries.
- Identify and address performance bottlenecks at the kernel layer.
- Ensure production p99 latency improvements and explain the underlying reasons.
- Collaborate with the inference team to integrate optimized kernels into the serving stack.
Requirements
- 5+ years of experience in ML systems.
- 2+ years of experience authoring production CUDA kernels that have measurably improved performance.
- Expertise in CUDA at the kernel-authoring level, including thread-block sizing, shared-memory layout, warp primitives, async copies (cp.async, TMA), and MMA selection.
- Proficiency in modifying and extending CUTLASS / CuTe DSL, including layout algebra.
- Experience debugging and modifying PTX code.
- Fluency in Nsight Compute and Systems for performance analysis and optimization.
- Experience authoring or modifying attention kernels (e.g., FlashAttention-family, paged, MLA, sliding-window, sparse).
- Understanding of multi-architecture differences (e.g., Hopper to Blackwell).