Data Scientist - Evaluations, Chanakya
Bengaluru • FullTime
Posted 5mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Sarvam is building India's full-stack sovereign AI platform, focusing on making AI genuinely work for India across research, models, infrastructure, and applications. This intellectually demanding role anchors the evaluations function for the AI vertical. You will design, build, and maintain evaluation frameworks to measure model and system quality in operational contexts, focusing on domain-specific, high-stakes use cases where accuracy is critical. You will collaborate closely with MLOps Engineers, Product Managers, and the deployment team to ensure deployed AI systems meet stringent quality standards.
Responsibilities
- Design and build evaluation frameworks for AI outputs across domain-specific requirements like document comprehension, command summarization, and geospatial reasoning.
- Define quality metrics in collaboration with domain experts and clients, translating operational requirements into measurable signals.
- Conduct structured evaluation cycles pre- and post-deployment and build dashboards to monitor model quality in production.
- Identify failure modes, edge cases, and distribution shifts with a critical perspective.
- Collaborate with MLOps Engineers to operationalize evaluation pipelines, ensuring they are automated, versioned, and reproducible.
- Build and manage domain-specific datasets for fine-tuning, evaluation, and benchmarking, including human annotation workflows.
- Publish internal findings and quality reports to inform the product and engineering roadmap.
Requirements
- 3-6 years in data science, ML research, or applied AI.
- At least 2 years working with LLMs in production contexts.
- Strong fundamentals in statistics and probability for valid evaluation design.
- Experience designing evaluation frameworks from scratch, including custom metrics, inter-rater reliability, and red-teaming methodologies.
- Proficiency in Python and comfort with libraries like pandas, NumPy, HuggingFace datasets, RAGAS, EleutherAI Eval Harness, or LangSmith.
- Experience with prompt engineering, model fine-tuning, or RLHF in applied settings.
- Ability to work with unstructured domain data such as PDFs, doctrine documents, transcripts, and field reports.