Embedded Data Scientist, Chanakya
Delhi • FullTime
Posted 5mo ago
About the job
Sarvam is building India's full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. As an Embedded Data Scientist, you will be deployed alongside Strategic Deployment Engineers at client sites. Your primary role is to transform complex client data into structures that AI systems can reliably reason over. This involves working directly with client data environments, understanding and structuring heterogeneous, multimodal datasets (including documents, images, audio, geospatial data, and structured records), and designing the semantic structures that enable AI interpretation and reasoning. You will define data representation within AI systems, including document segmentation, metadata, entity relationships, and cross-modal connections, by designing ontologies, tagging systems, and knowledge graph structures. You will own the quality of the data layer, ensuring the system is built on a foundation for reliable, scaled reasoning, often working with sensitive datasets where standard tooling may be absent.
Responsibilities
- Understand client data landscapes across various modalities and formats.
- Design domain ontologies representing entities, relationships, and concepts.
- Define document segmentation and chunking strategies for semantic preservation.
- Determine how heterogeneous datasets should be indexed, embedded, and linked.
- Collaborate with engineers to translate semantic structures into data ingestion pipelines.
- Evaluate and refine AI system performance based on data retrieval and reasoning.
- Define benchmarks and evaluation criteria with models and other teams.
- Translate client data insights into structured signals for product and engineering teams.
Requirements
- 2-5 years in data science, applied machine learning, or large-scale data analysis.
- Strong Python skills with pandas, NumPy, and modern NLP/LLM tooling.
- Solid grounding in ML fundamentals for understanding model behavior and evaluation.
- Experience with large unstructured datasets (documents, transcripts, reports).
- Familiarity with LLM-based systems, retrieval pipelines, or vector search.
- Experience designing or working with data schemas, metadata frameworks, or semantic data structures.