ML Ops Engineer, Chanakya
Delhi • FullTime
Posted 4mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Machine Learning Engineer
About the job
Sarvam is building India's full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. The MLOps Engineer will own the model lifecycle across all deployments, ensuring systems are always operational, accurate, and auditable. This role involves supporting field engineers and managing deployment infrastructure for new products, with an uncompromising standard for reliability, as model failures are considered operational risks.
Responsibilities
- Design and operate model serving infrastructure for on-prem and cloud deployments.
- Build and maintain CI/CD pipelines for model updates, rollbacks, and evaluation-gated deployments.
- Monitor model performance in production (latency, accuracy drift, throughput, failure modes) and build systems to proactively surface issues.
- Develop evaluation infrastructure, including harnesses, A/B testing, and model comparison tooling.
- Manage containerized model serving in constrained, air-gapped, and edge environments.
- Collaborate with Data Scientists on evaluation pipelines and own the underlying infrastructure.
- Create runbooks and operational playbooks for field deployment engineers.
- Own incident response for model-layer failures across all active deployments.
Requirements
- 3-5 years of experience in ML engineering or MLOps, with at least one production LLM or ML system in continuous operation.
- Deep expertise in model serving technologies (vLLM, TGI, Triton Inference Server, or equivalent).
- Experience with quantized model formats (GGUF, AWQ, GPTQ).
- Experience fine-tuning and adapting models in constrained, on-prem, or air-gapped environments, including managing data pipelines and compute limitations.
- Containerization experience with Docker, Kubernetes, or lightweight alternatives (K3s, K0s).
- Familiarity with deploying across heterogeneous hardware and infrastructure configurations.
- Experience with monitoring and observability tools (Prometheus, Grafana, or equivalent) and building custom eval dashboards.
- Proficiency in Python.
- Familiarity with fine-tuning workflows and model evaluation frameworks.
- Hands-on experience with CI/CD tooling for ML pipelines (GitHub Actions, ArgoCD, DVC, or similar).