Platform Engineer - AI Infrastructure
Bengaluru • FullTime
Posted 2mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Sarvam is building the bedrock of Sovereign AI for India, developing a full-stack AI platform across research, models, infrastructure, and applications. This role focuses on building the platform that sits on top of a large, multi-vendor GPU fleet, serving demanding training and inference workloads. You will design and ship control-plane services, scheduler integrations, autoscaling controllers, inference-serving platforms, RBAC and quota systems, observability and cost tooling, and the CLI and APIs that ML engineers use daily. The work is heavy software engineering, treating the platform as a product with internal users as customers.
Responsibilities
- Build the serving platform, including control plane for scalable, multi-tenant endpoints with intelligent routing, load balancing, and rollout machinery.
- Develop the scaling and elasticity layer for both training and serving, including autoscaling, capacity pooling, and efficient bin-packing across GPUs.
- Implement scheduling and orchestration layers, managing gang scheduling, priority, preemption, queue fairness, quota enforcement, and topology-aware placement.
- Design and implement multi-tenancy, RBAC, and isolation mechanisms, including tenant models, namespacing, policy enforcement, and secrets management.
- Develop networking components for platform-level plumbing, CNI configuration, ingress/service routing, and tenant network policy.
- Build observability and cost tooling as a platform feature, managing metrics, logging, and tracing at fleet scale.
- Create abstractions over storage and data paths, including parallel filesystems, caching tiers, and volume provisioning.
- Enhance developer experience and self-service through CLI, SDK, and APIs for ML engineers.
- Develop provisioning and infrastructure-as-code capabilities for reproducible cluster setup and management.
Requirements
- 5+ years of experience building infrastructure or platform software with a track record of shipping services and control planes.
- Strong software engineering skills in Go or Python, with the ability to build and debug maintainable systems.
- Deep understanding of Kubernetes at the controller and internals level, including experience with operators or controllers.
- Working literacy in GPU-specific platform constraints such as MIG, GPU sharing, gang scheduling, and topology-aware placement.
- A product mindset focused on internal users, designing APIs and abstractions for adoption and self-service.
- Ability to own a capability end-to-end, from design through rollout and documentation.