GRPO Jobs
9 open roles mentioning GRPO
Research Intern RL & Post-Training Systems, Turbo (Fall 2026)
Together AI
The Turbo Research team focuses on making post-training and reinforcement learning for large language models efficient, scalable, and reliable. This work intersects RL algorithms, inference systems, and large-scale experimentation, where inference costs significantly impact training efficiency and the practicality of learning algorithms. As a research intern, you will investigate RL and post-training methods whose performance and scalability are closely tied to inference behavior, co-designing algorithms and systems. Projects aim to enable new experimental regimes, including larger models, longer rollouts, and more complex evaluations, by re-evaluating the interaction between inference, scheduling, and training.
AI Researcher, Core ML (Turbo)
Together AI
The Turbo team operates at the intersection of efficient inference (algorithms, architectures, engines) and post-training/RL systems. We are responsible for building and managing the systems that power Together's API, focusing on high-performance inference and RL/post-training engines capable of operating at production scale. Our core mission is to advance the frontiers of efficient inference and RL-driven training, aiming to make models significantly faster and more cost-effective to run, while simultaneously enhancing their capabilities through RL-based post-training methods. This role involves working across the entire stack, from RL algorithms and training engines to kernels and serving systems, to develop and refine state-of-the-art models using RL pipelines. We value individuals with deep expertise in one area and a strong willingness to collaborate and grow across others.
$200k - $280k
Research Scientist, Agent Robustness
Scale AI
Scale Labs is seeking talented researchers to join a new team focused on policy research, bridging the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. This role will tackle fundamental challenges in building AI agents that are safe and aligned with humans, researching agent capabilities, designing evaluation harnesses, building exploits and mitigations for failure modes, and characterizing risks of multi-agent systems. The team collaborates broadly across industry, the public sector, and academia, regularly publishing findings.
$216k - $270k
Research Scientist, AI Controls and Monitoring
Scale AI
Scale Labs is seeking a Research Scientist focused on AI Controls and Monitoring to join a new team dedicated to policy research. This role will bridge the gap between AI research and policymakers, focusing on scientific decisions about AI risks and capabilities. The team tackles challenges in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. You will design methods, systems, and experiments to ensure advanced AI models and agents remain aligned with intended goals, even in high-stakes or adversarial environments. This role involves collaboration across industry, the public sector, and academia, with regular publication of findings.
$216k - $270k
Research Scientist, Safety Post Training
Scale AI
Scale Labs is seeking talented researchers to join a new team focused on policy research, bridging the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. This role will develop and apply post-training methods and interpretability techniques to make frontier AI systems safer and better understood. You will design and run post-training pipelines, develop interpretability-informed evaluations, and collaborate with policymakers, engineers, and other researchers to translate findings into actionable safety standards and best practices.
$216k - $270k
Machine Learning Research Engineer, Agents - Enterprise GenAI
Scale AI
Scale is seeking a Machine Learning Research Engineer focused on Agents for its Enterprise GenAI team. This role will be instrumental in accelerating the development of AI applications by working on state-of-the-art post-training algorithms for complex enterprise agents. You will apply proprietary Agent RL Training + Building algorithms to real-world enterprise datasets and benchmarks, aiming to create best-in-class Agents that achieve state-of-the-art results. If you are passionate about shaping the future of GenAI, this is an opportunity to contribute to cutting-edge research and development.
$265k - $331k
Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI
Scale AI
Scale is seeking a Machine Learning Systems Research Engineer to join their Enterprise ML Research Lab. This role will focus on building algorithms for a next-generation Agent RL training platform, supporting large-scale training, and integrating state-of-the-art technologies to optimize ML systems. You will collaborate with other ML researchers and engineers who apply these algorithms to client use cases, including AI cybersecurity firewalls and healthtech search models. If you are passionate about shaping the future of AI, this is an exciting opportunity to contribute to cutting-edge advancements in enterprise GenAI.
$265k - $331k
Staff Machine Learning Research Engineer, Agent Post-training - Enterprise GenAI
Scale AI
Scale is seeking a Staff Agent Post-Training ML Research Engineer to join our Enterprise ML Research Lab. This role will focus on building out our next-generation Agent RL training platform, integrating cutting-edge research to train best-in-class Agents for real enterprise use-cases. You will contribute to the development of AI applications that are becoming vital across all sectors, from cybersecurity LLMs to foundation healthtech search models, shaping the future of the modern GenAI movement.
$265k - $331k
Tech Lead Manager- MLRE, ML Systems
Scale AI
Scale's LLM post-training platform team builds our internal distributed framework for large language model training, powering MLEs, researchers, data scientists, and operators for fast and automatic training and evaluation of LLMs. This platform also serves as the underlying training framework for the data quality evaluation pipeline. You will work closely with Scale’s ML teams and researchers to build the foundation platform which supports all our ML research and development works, optimizing it to enable next generation LLM training, inference, and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you!
$265k - $331k