ML Researcher, Foundational Models
Bengaluru • FullTime
Posted 3mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Research Engineer
About the job
Sarvam is building the bedrock of Sovereign AI for India, developing a full-stack AI platform focused on making AI genuinely work for India. This role is for a researcher who will tackle open-ended questions about the architecture, optimization, data composition, and training dynamics of our next generation of foundational models. You will have direct access to large compute resources and a tight feedback loop with engineers, driving research from initial hunches to production-ready decisions. This is a hands-on role requiring independent research design, execution at scale, and the ability to translate findings into concrete proposals for production training runs.
Responsibilities
- Drive open-ended research on architecture, optimization, scaling behavior, training stability, and post-training recipes for foundational models.
- Design and execute ablations at scales that inform large-run decisions, including end-to-end pre-training experiments.
- Translate research findings into concrete proposals for the next training run and own them through to production.
- Collaborate closely with infrastructure and data teams on research questions at their intersection.
- Publish research findings internally and externally when appropriate.
Requirements
- PhD in Machine Learning, Computer Science, or a closely related field (or in the final stages of completion).
- 3+ years of research experience post-PhD (or equivalent depth).
- First-author publications at top-tier ML venues (NeurIPS, ICML, ICLR, ACL, EMNLP, COLM).
- Hands-on experience pre-training transformer-based language models from scratch (ideally 7B+ parameters), with the ability to describe an end-to-end training run and debugging process.
- Meaningful contributions to the open-source LLM ecosystem (research code, model releases, datasets, or substantive contributions to widely-used projects).
- Fluency in PyTorch and comfort with distributed training, including identifying performance or stability issues in training loops.
- Strong intuition for experimental design, including what to measure, ablate, and at what scale results need to hold.
Benefits
- Direct access to large compute.
- Tight feedback loop with engineers building training stack and data pipelines.
- Autonomy and compute to make architectural and training-recipe decisions on frontier-scale models.
- High ownership and high impact from day one.
- Opportunity to work on problems with population-scale impact in India.