ML Engineer (Data), Foundational Models
Bengaluru • FullTime
Posted 3mo ago
About the job
Sarvam is building India's sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. This role is critical in owning the data infrastructure that fuels our foundational models. You will be responsible for building petabyte-scale curation and filtering pipelines, designing systems for training data selection and proportioning, and ensuring data quality with research-level rigor. This is an engineering- and research-intensive position involving large-scale deduplication, quality modeling, contamination detection, mixture design, curriculum learning, attribution, and debugging.
Responsibilities
- Design and build large-scale data pipelines for pre-training and post-training (ingestion, parsing, normalization, filtering, deduplication, tokenization, packing) at petabyte scale.
- Develop and improve quality filtering systems, including model-based quality classifiers and contamination detection.
- Own data mixture design, curriculum, and annealing strategies in partnership with the research team, ensuring precise tracking of data seen by models.
- Build tooling for researchers and engineers to analyze, slice, attribute, and debug data.
- Scale pipelines to handle multilingual corpora, code, math, multi-source web data, and licensed datasets, while tracking provenance and licensing.
- Partner with the training infrastructure team to ensure data is never a bottleneck in production training runs.
Requirements
- BS or MS in Computer Science or a closely related technical field, or equivalent demonstrated experience.
- 3+ years of experience building large-scale data systems (petabyte-scale processing, distributed data pipelines, or comparable).
- Hands-on experience with data curation and filtering for LLM training, including defending choices made for pre-training corpora.
- Deep familiarity with distributed data processing frameworks (Spark, Ray, Beam, Dask, or equivalent) and underlying storage systems.
- Strong Python skills, with comfort in low-level data path components (tokenization, sharding, packing, IO patterns) and performance tradeoffs.
- Meaningful open-source contributions in the data tooling ecosystem (datasets, dedup libraries, filtering frameworks, or work on open data releases).
Benefits
- High ownership and high impact from day one.
- Opportunity to work on problems that could change how an entire country learns, works, and communicates.
- Work alongside researchers, engineers, builders, and business leaders.
- AI-first approach to building, shipping, and problem-solving.