Senior Backend Engineer, Vision
Bengaluru • FullTime
Posted 27d ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Sarvam is building India's full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. This role is for a Senior Backend Engineer who will own the architecture of the serving harness for Sarvam's vision models. The system needs to deliver frontier-grade extraction quality from in-house models at a national scale, while managing cost and latency budgets. This involves making critical trade-offs between accuracy, latency, and cost, and setting high standards for reliability, durability, and idempotency in document processing pipelines.
Responsibilities
- Own the end-to-end architecture of the OCR and extraction serving harness, including API layer, orchestration, inference, post-processing, and delivery.
- Design and implement the accuracy harness, incorporating multi-pass extraction, ensembling, cross-verification, and confidence calibration.
- Architect durable and resumable document workflows using Temporal, handling fan-out, partial failure recovery, and long-running jobs.
- Manage the inference serving layer, including batching, GPU pool management, autoscaling, and multi-model routing.
- Drive down latency, throughput, and unit economics by profiling, measuring, and defending cost-per-page targets.
- Build the observability substrate with distributed tracing, per-stage metrics, SLOs, and alerting.
- Design for multi-tenancy, isolation, rate limiting, and fair scheduling for enterprise customers.
- Support on-prem and constrained deployments within customer environments.
- Set technical direction and mentor junior engineers through design reviews.
Requirements
- 5-6+ years in backend engineering with experience operating high-throughput production systems and on-call responsibilities.
- Deep proficiency in Go and/or Python.
- Strong distributed systems design skills, including queues, workflow orchestration, idempotency, backpressure, and consistency trade-offs.
- Production experience with Temporal or a similar durable execution engine.
- Experience with Kubernetes in production, including autoscaling and GPU workload scheduling.
- Demonstrable experience serving ML or LLM inference in production, including batching, caching, and latency budgeting.
- Rigorous approach to observability and reliability, including designing SLOs and managing incidents.
- Proven ability to reduce system costs without compromising quality.