Research Intern RL & Post-Training Systems, Turbo (Fall 2026)
Remote • San Francisco
Posted 1mo ago
About the job
The Turbo Research team focuses on making post-training and reinforcement learning for large language models efficient, scalable, and reliable. This work intersects RL algorithms, inference systems, and large-scale experimentation, where inference costs significantly impact training efficiency and the practicality of learning algorithms. As a research intern, you will investigate RL and post-training methods whose performance and scalability are closely tied to inference behavior, co-designing algorithms and systems. Projects aim to enable new experimental regimes, including larger models, longer rollouts, and more complex evaluations, by re-evaluating the interaction between inference, scheduling, and training.
Responsibilities
- Study RL and post-training methods whose performance and scalability are tightly coupled to inference behavior.
- Co-design algorithms and systems rather than treating them independently.
- Design controlled experiments and interpret results to draw principled conclusions.
- Work across abstraction layers, including modifying inference or training systems.
- Design rigorous benchmarks and diagnostics for post-training and RL efficiency.
- Study failure modes in long-horizon training and how system constraints shape outcomes.
Requirements
- Pursuing a PhD or MS in Computer Science, EE, or a related field (exceptional undergraduates considered).
- Research experience in RL or post-training for large models (e.g., RLHF, RLAIF, GRPO, preference optimization).
- Research experience in ML systems (inference engines, runtimes, distributed systems).
- Research experience in large-scale empirical ML research or evaluation.
- Comfortable with empirical research, designing controlled experiments, and interpreting noisy results.
- Ability to work across abstraction layers.
- Strong Python skills for experimentation.
- Willingness to modify inference or training systems (experience with C++, CUDA, or similar is a plus).
Benefits
- Competitive compensation
- Housing stipends
- Other competitive benefits