Research Engineer, Mid-Training
San Francisco • FullTime
Posted 1mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
AI Research Engineer
About the job
We are an applied AI lab building end-to-end software agents, known for creating Devin, the first AI software engineer. Our team is composed of highly talented individuals with backgrounds in competitive programming and leadership roles at cutting-edge AI companies. We are tackling significant global challenges and developing AI capable of real-world reasoning. This role focuses on the critical 'mid-training' phase, bridging pre-training and post-training to refine raw model capabilities. You will be instrumental in shaping our models' fundamental abilities by owning late-stage training decisions, including data mix and quality, annealing schedules, context length extension, capability injection, and synthetic data strategies.
Responsibilities
- Design and iterate on high-quality data mixtures for late-stage and annealing training runs.
- Develop principled methods for sourcing, filtering, and weighting data to sharpen model capabilities.
- Drive targeted improvements in coding, mathematics, and long-horizon reasoning through curated data strategies.
- Develop and evaluate synthetic data pipelines for generating training signal at scale.
- Research and optimize multi-stage learning rate schedules, warmup strategies, and compute allocation.
- Research and implement methods for extending effective context length without degrading performance.
- Build evaluations to distinguish real capability improvements from benchmark overfitting.
- Measure how mid-training interventions scale with compute and data.
- Develop new approaches when existing methods hit ceilings.
Requirements
- Deep familiarity with the LLM training pipeline end to end.
- Hands-on experience with continual pre-training, annealing, or late-stage data mixing for large models.
- Strong intuition for data quality, including filtering and curation at scale.
- Experience developing or evaluating synthetic data pipelines for capability improvement.
- Proficiency in Python and deep learning frameworks like PyTorch.
- Comfort debugging distributed training at scale.
- Strong fundamentals in optimization, statistics, and ML theory.
- Ability to distinguish real effects from noise, instability, and overfitting.
- A track record of original contributions (publications, open-source impact, or internal results).
- Comfort operating in ambiguous, fast-moving environments.