Staff Engineer, Product Infrastructure
Bengaluru • FullTime
Posted 4mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Sarvam is building India's full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. The company partners with leading enterprises and public institutions. At the core of Sarvam's conversational AI platform is a system that enables customers to build, configure, deploy, and operate voice agents and campaigns at scale, currently handling over 50 million minutes of conversations monthly and experiencing rapid growth. We are seeking a Staff Engineer to take end-to-end ownership of this critical system, including its data model, APIs, orchestration, reliability, performance, and engineering standards. This is an individual contributor role focused on systems and technical excellence, without people management responsibilities.
Responsibilities
- Own the end-to-end backend for agent, campaign, and user management, including data modeling, API design, and deployment using Python and FastAPI.
- Implement workflow orchestration with Temporal for complex, long-running, stateful operations.
- Ensure performance and reliability for a multi-tenant SaaS platform with sub-second latency SLOs.
- Develop and maintain observability systems (logging, metrics, tracing) for efficient production debugging.
- Build and manage integration testing infrastructure to enable rapid, safe deployments.
- Collaborate closely with the frontend team (Next.js) to deliver a high-quality user experience.
- Establish a design-first and documentation-first engineering culture, emphasizing RFCs and written decision-making.
Requirements
- 5-10 years of experience building production backend systems, with deep expertise in Python and FastAPI (or Golang).
- Proven track record of full-stack ownership, including system design, trade-off analysis, and production operation.
- Hands-on experience with large-scale workflow orchestration, preferably using Temporal.
- Experience building and operating high-scale distributed systems (multi-tenant, low-latency, resilient).
- Strong fundamentals in PostgreSQL and Redis, including schema design, query performance, and caching strategies.
- Demonstrated judgment in making build vs. buy, optimize vs. ship, and abstract vs. inline decisions.