Platform Engineer - AI Infrastructure

Bengaluru FullTime

Posted 2mo ago

Job Location

Bengaluru

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Sarvam is building the bedrock of Sovereign AI for India, developing a full-stack AI platform across research, models, infrastructure, and applications. This role focuses on building the platform that sits on top of a large, multi-vendor GPU fleet, serving demanding training and inference workloads. You will design and ship control-plane services, scheduler integrations, autoscaling controllers, inference-serving platforms, RBAC and quota systems, observability and cost tooling, and the CLI and APIs that ML engineers use daily. The work is heavy software engineering, treating the platform as a product with internal users as customers.

Responsibilities

  • Build the serving platform, including control plane for scalable, multi-tenant endpoints with intelligent routing, load balancing, and rollout machinery.
  • Develop the scaling and elasticity layer for both training and serving, including autoscaling, capacity pooling, and efficient bin-packing across GPUs.
  • Implement scheduling and orchestration layers, managing gang scheduling, priority, preemption, queue fairness, quota enforcement, and topology-aware placement.
  • Design and implement multi-tenancy, RBAC, and isolation mechanisms, including tenant models, namespacing, policy enforcement, and secrets management.
  • Develop networking components for platform-level plumbing, CNI configuration, ingress/service routing, and tenant network policy.
  • Build observability and cost tooling as a platform feature, managing metrics, logging, and tracing at fleet scale.
  • Create abstractions over storage and data paths, including parallel filesystems, caching tiers, and volume provisioning.
  • Enhance developer experience and self-service through CLI, SDK, and APIs for ML engineers.
  • Develop provisioning and infrastructure-as-code capabilities for reproducible cluster setup and management.

Requirements

  • 5+ years of experience building infrastructure or platform software with a track record of shipping services and control planes.
  • Strong software engineering skills in Go or Python, with the ability to build and debug maintainable systems.
  • Deep understanding of Kubernetes at the controller and internals level, including experience with operators or controllers.
  • Working literacy in GPU-specific platform constraints such as MIG, GPU sharing, gang scheduling, and topology-aware placement.
  • A product mindset focused on internal users, designing APIs and abstractions for adoption and self-service.
  • Ability to own a capability end-to-end, from design through rollout and documentation.

About sarvam

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.