ML Engineer (Training Infra), Foundational Models

Bengaluru • FullTime

Posted 9h ago

Job Location

Bengaluru

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Sarvam is building the bedrock of Sovereign AI for India, developing a full-stack AI platform focused on research, models, infrastructure, and applications. This role owns the infrastructure for training foundational models at scale, involving deep systems work such as distributed training, parallelism strategies, GPU kernel optimization, and reliability engineering. The position demands a high level of expertise to identify and resolve complex issues that can significantly impact training timelines and costs.

Responsibilities

  • Build, maintain, and optimize the distributed training stack across large GPU clusters.
  • Design and implement various parallelism strategies (data, tensor, pipeline, sequence, expert) for efficient model training.
  • Profile and optimize end-to-end training throughput, including kernel performance, communication overlap, memory layout, checkpointing, and data loading.
  • Write and tune custom GPU kernels (CUDA, Triton) to enhance performance.
  • Ensure the reliability of long-running training jobs through fault tolerance, checkpoint integrity, deterministic restarts, and automated detection of issues.
  • Collaborate with researchers to ensure architectural ideas are trainable efficiently and with the data team to prevent pipeline bottlenecks.

Requirements

  • BS or MS in Computer Science or a related technical field, or equivalent demonstrated experience.
  • 3+ years of experience building ML training infrastructure or large-scale distributed systems.
  • Hands-on experience training large models with distributed training frameworks (e.g., Megatron-LM, DeepSpeed, FSDP, NeMo).
  • Deep working knowledge of GPU architecture, the CUDA programming model, and GPU workload profiling tools (e.g., Nsight, PyTorch profiler).
  • Strong PyTorch internals knowledge, including comfort with low-level training code.
  • Meaningful open-source contributions in the training infrastructure ecosystem (e.g., Megatron, DeepSpeed, PyTorch, vLLM, Triton, NCCL).

Benefits

  • High ownership and high impact from day one.
  • Opportunity to work on problems that could change how an entire country learns, works, and communicates.
  • Work alongside a fast-moving, high talent-density team.

About sarvam

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.