DPO Jobs

5 open roles mentioning DPO

Research Intern RL & Post-Training Systems, Turbo (Fall 2026)

1mo ago
Together AI

Together AI

The Turbo Research team focuses on making post-training and reinforcement learning for large language models efficient, scalable, and reliable. This work intersects RL algorithms, inference systems, and large-scale experimentation, where inference costs significantly impact training efficiency and the practicality of learning algorithms. As a research intern, you will investigate RL and post-training methods whose performance and scalability are closely tied to inference behavior, co-designing algorithms and systems. Projects aim to enable new experimental regimes, including larger models, longer rollouts, and more complex evaluations, by re-evaluating the interaction between inference, scheduling, and training.

San Francisco remote
PythonC#NLP +7 more

AI Researcher, Core ML (Turbo)

1mo ago
Together AI

Together AI

The Turbo team operates at the intersection of efficient inference (algorithms, architectures, engines) and post-training/RL systems. We are responsible for building and managing the systems that power Together's API, focusing on high-performance inference and RL/post-training engines capable of operating at production scale. Our core mission is to advance the frontiers of efficient inference and RL-driven training, aiming to make models significantly faster and more cost-effective to run, while simultaneously enhancing their capabilities through RL-based post-training methods. This role involves working across the entire stack, from RL algorithms and training engines to kernels and serving systems, to develop and refine state-of-the-art models using RL pipelines. We value individuals with deep expertise in one area and a strong willingness to collaborate and grow across others.

$200k - $280k

San Francisco remote
PythonTransformersRLHF +7 more

Research Scientist, Agent Robustness

2mo ago
Scale AI

Scale AI

Scale Labs is seeking talented researchers to join a new team focused on policy research, bridging the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. This role will tackle fundamental challenges in building AI agents that are safe and aligned with humans, researching agent capabilities, designing evaluation harnesses, building exploits and mitigations for failure modes, and characterizing risks of multi-agent systems. The team collaborates broadly across industry, the public sector, and academia, regularly publishing findings.

$216k - $270k

San Francisco, CA; New York, NY onsite
RLHFGRPODPO +7 more

Research Scientist, AI Controls and Monitoring

2mo ago
Scale AI

Scale AI

Scale Labs is seeking a Research Scientist focused on AI Controls and Monitoring to join a new team dedicated to policy research. This role will bridge the gap between AI research and policymakers, focusing on scientific decisions about AI risks and capabilities. The team tackles challenges in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. You will design methods, systems, and experiments to ensure advanced AI models and agents remain aligned with intended goals, even in high-stakes or adversarial environments. This role involves collaboration across industry, the public sector, and academia, with regular publication of findings.

$216k - $270k

San Francisco, CA; New York, NY remote
RLHFGRPODPO +7 more

Research Scientist, Safety Post Training

2mo ago
Scale AI

Scale AI

Scale Labs is seeking talented researchers to join a new team focused on policy research, bridging the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. This role will develop and apply post-training methods and interpretability techniques to make frontier AI systems safer and better understood. You will design and run post-training pipelines, develop interpretability-informed evaluations, and collaborate with policymakers, engineers, and other researchers to translate findings into actionable safety standards and best practices.

$216k - $270k

San Francisco, CA; New York, NY onsite
RLHFGRPODPO +4 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.