RLHF Jobs

36 open roles mentioning RLHF

Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI

2mo ago
Scale AI

Scale AI

Scale is seeking a Machine Learning Systems Research Engineer to join their Enterprise ML Research Lab. This role will focus on building algorithms for a next-generation Agent RL training platform, supporting large-scale training, and integrating state-of-the-art technologies to optimize ML systems. You will collaborate with other ML researchers and engineers who apply these algorithms to client use cases, including AI cybersecurity firewalls and healthtech search models. If you are passionate about shaping the future of AI, this is an exciting opportunity to contribute to cutting-edge advancements in enterprise GenAI.

$265k - $331k

San Francisco, CA; New York, NY onsite
LLMPyTorchTransformers +7 more

Staff Machine Learning Research Engineer, Agent Post-training - Enterprise GenAI

2mo ago
Scale AI

Scale AI

Scale is seeking a Staff Agent Post-Training ML Research Engineer to join our Enterprise ML Research Lab. This role will focus on building out our next-generation Agent RL training platform, integrating cutting-edge research to train best-in-class Agents for real enterprise use-cases. You will contribute to the development of AI applications that are becoming vital across all sectors, from cybersecurity LLMs to foundation healthtech search models, shaping the future of the modern GenAI movement.

$265k - $331k

San Francisco, CA; New York, NY onsite
LLMRLHFAI +4 more

Machine Learning Research Engineer, Agents - Enterprise GenAI

2mo ago
Scale AI

Scale AI

Scale is seeking a Machine Learning Research Engineer focused on Agents for its Enterprise GenAI team. This role will be instrumental in accelerating the development of AI applications by working on state-of-the-art post-training algorithms for complex enterprise agents. You will apply proprietary Agent RL Training + Building algorithms to real-world enterprise datasets and benchmarks, aiming to create best-in-class Agents that achieve state-of-the-art results. If you are passionate about shaping the future of GenAI, this is an opportunity to contribute to cutting-edge research and development.

$265k - $331k

San Francisco, CA; New York, NY onsite
RLHFLLMsGRPO +3 more

Forward Deployed Engineer, GenAI

2mo ago
Scale AI

Scale AI

Scale AI is seeking a Forward Deployed Engineer, GenAI to join their Data Engine team. This role is at the forefront of providing critical data infrastructure that powers advanced AI models, directly influencing how humanity interacts with AI. You will work with the world's leading AI companies and government agencies to solve their most complex AI data-related problems, contributing to the advancement of AI by delivering critical data solutions for leading AI innovators and government agencies. You will interact daily with technical customers, understand their unique challenges, and translate them into impactful solutions, while also designing, building, and deploying features across the entire stack. This position offers a unique opportunity to lead critical projects, shape engineering culture, and accelerate career growth in the rapidly evolving field of Generative AI.

$179k - $224k

San Francisco, CA; New York, NY hybrid
Reinforcement LearningRLHFAI +8 more

Senior Software Engineer, GenAI

2mo ago
Scale AI

Scale AI

Scale AI is seeking a Senior Software Engineer to join our Generative AI Data Engine team. This role is crucial in accelerating the development of AI applications by powering the world's most advanced LLMs and generative models. You will work on high-impact datasets, optimize contributor onboarding, and ensure data integrity through advanced trust, safety, and security measures. This position operates at the intersection of ML, operations, and analytics to deliver high-quality data at scale, contributing to the future of human-AI interaction.

$216k - $270k

San Francisco, CA; New York, NY hybrid
PythonTypeScriptReinforcement Learning +6 more

Tech Lead Manager- MLRE, ML Systems

2mo ago
Scale AI

Scale AI

Scale's LLM post-training platform team builds our internal distributed framework for large language model training, powering MLEs, researchers, data scientists, and operators for fast and automatic training and evaluation of LLMs. This platform also serves as the underlying training framework for the data quality evaluation pipeline. You will work closely with Scale’s ML teams and researchers to build the foundation platform which supports all our ML research and development works, optimizing it to enable next generation LLM training, inference, and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you!

$265k - $331k

San Francisco, CA; New York, NY onsite
LLMPyTorchTransformers +7 more

ML Research Engineer, ML Systems

2mo ago
Scale AI

Scale AI

Scale's ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. This platform powers MLEs, researchers, data scientists, and operators for fast and automatic training and evaluation of LLMs, as well as data quality evaluation. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development, optimizing it to enable the next generation of LLM training, inference, and data curation. If you are excited about shaping the future of AI via fundamental innovations, we would love to hear from you!

$190k - $237k

San Francisco, CA; Seattle, WA; New York, NY onsite
LLMPyTorchTransformers +7 more

GenAI Strategic Projects Lead, Public Sector

2mo ago
Scale AI

Scale AI

Scale AI is seeking a Strategic Projects Lead for its Public Sector team to own high-impact projects focused on Generative AI and Large Language Models. This role involves working across operations, engineering, and customer engagement to produce high-quality training and test data for LLMs, particularly for Public Sector customers. You will be instrumental in building Generative AI data-labeling pipelines, creating operational processes for an expert data workforce, and developing novel technology-driven approaches to enhance data quality. This is a unique opportunity to contribute at the intersection of AI and national security, partnering with internal ML experts and external stakeholders to ensure data supports mission-critical AI applications.

$170k - $212k

Washington, DC onsite
Prompt EngineeringFine-TuningReinforcement Learning +9 more

Machine Learning Research Scientist, Post-Training

2mo ago
Scale AI

Scale AI

Scale works with leading AI labs to accelerate progress in GenAI research, focusing on optimizing data curation and evaluation to enhance LLM capabilities in text and multimodal modalities. This role involves developing novel methods to improve the alignment and generalization of large-scale generative models, collaborating with researchers and engineers on best practices in data-driven AI development, and providing technical and strategic input to foundation model labs for the next generation of AI models.

$252k - $315k

San Francisco, CA; Seattle, WA; New York, NY onsite
LLMFine-TuningReinforcement Learning +7 more

Research, Post-Training Data

2mo ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking researchers to bridge the gap between raw AI intelligence and useful, safe, and collaborative systems. This role focuses on post-training data research, combining human insight and machine learning techniques to capture and steer model behavior based on human preferences. You will be responsible for translating research ideas into actionable data through labeling and collection campaigns, understanding data quality science, and developing metrics to measure the impact of data and training interventions. The position also involves exploring new paradigms for human-AI interaction and scalable oversight, blending research, data operations, and technical implementation to advance human-centered AI systems. This role requires both fundamental research and practical engineering, making it ideal for individuals who enjoy deep theoretical exploration and hands-on experimentation.

$350k - $475k

San Francisco onsite
OpenAIMistralPython +5 more

Code Data Annotation Quality Specialist

3mo ago
Mistral AI

Mistral AI

Mistral is seeking highly motivated Data Quality Specialists to join our Human Data Annotation team. This hybrid role involves reviewing and auditing code annotations against rubrics to ensure high-quality data for AI model training and evaluation. You will also be responsible for building, maintaining, and troubleshooting the internal tooling that annotators use daily. This position offers the opportunity to collaborate closely with annotators, technical program managers, and engineering stakeholders, contributing to the refinement of guidelines and processes that shape data production.

Paris hybrid Full-time
MistralPythonJavaScript +5 more

Forward Deployed Engineer - LLM Post-training

3mo ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. We build open-weight models that empower users to control their AI and shape its future. As a core member of the Applied AI team, you will drive model fine-tuning and evaluations for enterprise customers. This role involves adapting our open-weight models for specific customer domains, tasks, and constraints, working hands-on with customer data, running fine-tuning workflows, building evaluation harnesses, and deploying adapted models to production. You will collaborate directly with customers to understand their needs and with research teams to advance the possibilities of AI.

San Francisco, CA onsite FullTime
PythonFine-TuningRLHF

Research Engineer

5mo ago
Cohere

Cohere

Cohere Labs is seeking Research Engineers to join their dedicated research arm, focused on pushing machine learning forward through open, collaborative research and hands-on experimentation. This is a highly practical role where you will work closely with scientists and engineers to implement new methods, run large-scale experiments, and help shape the infrastructure supporting our research programs. You will be responsible for building experiments, debugging models, scaling training pipelines, and turning research ideas into working systems. We value curiosity, strong fundamentals, and a willingness to learn quickly in a fast-moving research environment, with a focus on practical impact.

Toronto remote FullTime
CohereFine-TuningPyTorch +4 more

Member of Technical Staff - Safety

6mo ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower users to control their intelligence and shape the future of AI. As a Member of Technical Staff - Safety, you will be instrumental in ensuring the safety and reliability of our AI models. This role involves owning the red-teaming and adversarial evaluation pipeline, translating safety findings into concrete guardrails, and validating that every release meets our risk thresholds before deployment. You will develop scalable, automated safety benchmarks and research state-of-the-art jailbreaking techniques and defenses to proactively address potential vulnerabilities.

San Francisco, CA onsite FullTime
Reinforcement LearningRLHF

Product Manager, Forge

7mo ago
Mistral AI

Mistral AI

Mistral AI is seeking a talented and experienced Product Manager to define and execute the strategy for Forge, a product that empowers customers to build, fine-tune, and deploy custom AI models at scale. Forge transforms cutting-edge research into enterprise-ready capabilities by supporting model fine-tuning, reinforcement learning, and post-training workflows. This role operates at the intersection of research and product, enabling customers to train specialized models for real-world business value. You will collaborate closely with applied AI scientists and research engineers to translate frontier techniques into scalable and reliable solutions, shaping a 0-1 product with significant business impact and defining the future of how organizations train and deploy AI models.

Paris hybrid Full-time
MistralKubernetesGo +9 more

Researcher, Post Training

9mo ago
c

cartesia

Cartesia is seeking a Researcher for their Post-Training team to develop methods and systems that make multimodal models adaptive, aligned, and grounded in human intent. This role involves working at the intersection of machine learning research, alignment, and infrastructure, focusing on preference optimization, model evaluation, and feedback-driven learning. You will explore how feedback signals can guide models to reason more effectively across modalities and build the infrastructure to measure and improve these behaviors at scale. Your work will directly influence how Cartesia's foundation models learn, improve, and connect with people.

*HQ - San Francisco, CA onsite FullTime
Fine-TuningRLHF

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.