Reinforcement Learning Jobs
50 open roles mentioning Reinforcement Learning
Senior Software Engineer, GenAI
Scale AI
Scale AI is seeking a Senior Software Engineer to join our Generative AI Data Engine team. This role is crucial in accelerating the development of AI applications by powering the world's most advanced LLMs and generative models. You will work on high-impact datasets, optimize contributor onboarding, and ensure data integrity through advanced trust, safety, and security measures. This position operates at the intersection of ML, operations, and analytics to deliver high-quality data at scale, contributing to the future of human-AI interaction.
$216k - $270k
Forward Deployed Engineer, GenAI
Scale AI
Scale AI is seeking a Forward Deployed Engineer, GenAI to join their Data Engine team. This role is at the forefront of providing critical data infrastructure that powers advanced AI models, directly influencing how humanity interacts with AI. You will work with the world's leading AI companies and government agencies to solve their most complex AI data-related problems, contributing to the advancement of AI by delivering critical data solutions for leading AI innovators and government agencies. You will interact daily with technical customers, understand their unique challenges, and translate them into impactful solutions, while also designing, building, and deploying features across the entire stack. This position offers a unique opportunity to lead critical projects, shape engineering culture, and accelerate career growth in the rapidly evolving field of Generative AI.
$179k - $224k
GenAI Strategic Projects Lead, Public Sector
Scale AI
Scale AI is seeking a Strategic Projects Lead for its Public Sector team to own high-impact projects focused on Generative AI and Large Language Models. This role involves working across operations, engineering, and customer engagement to produce high-quality training and test data for LLMs, particularly for Public Sector customers. You will be instrumental in building Generative AI data-labeling pipelines, creating operational processes for an expert data workforce, and developing novel technology-driven approaches to enhance data quality. This is a unique opportunity to contribute at the intersection of AI and national security, partnering with internal ML experts and external stakeholders to ensure data supports mission-critical AI applications.
$170k - $212k
Machine Learning Research Scientist, Post-Training
Scale AI
Scale works with leading AI labs to accelerate progress in GenAI research, focusing on optimizing data curation and evaluation to enhance LLM capabilities in text and multimodal modalities. This role involves developing novel methods to improve the alignment and generalization of large-scale generative models, collaborating with researchers and engineers on best practices in data-driven AI development, and providing technical and strategic input to foundation model labs for the next generation of AI models.
$252k - $315k
Member of Technical Staff - Mid-Training Infra
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower individuals to control their intelligence and shape the future of AI. As a Member of Technical Staff focused on Mid-Training Infrastructure, you will be instrumental in designing, building, and operating large-scale GPU infrastructure crucial for high-throughput model inference and mid-training workloads. This role involves developing systems that support synthetic data generation and reinforcement learning pipelines at scale, as well as building high-performance inference platforms capable of serving and evaluating models across thousands of GPUs.
Research Scientist, Multimodal Alignment, Safety, and Fairness
Google DeepMind
Google DeepMind's Frontier AI unit is seeking experienced Research Scientists to join a multimodal safety research effort. This role focuses on interdisciplinary sociotechnical modeling and requires a passion for understanding AI-society interactions, a strong awareness of AI alignment and safety, and a drive to develop novel ideas, methods, interfaces, and tools. You will contribute to advancing the state of the art in AI research and Google DeepMind's mission towards Artificial General Intelligence (AGI), with a focus on leading new breakthrough research directions in areas like AI behavior exploration, assessment, and steering, particularly for subjective and creative tasks. The work involves tackling fundamental research questions to improve alignment objectives, assess adherence to desired behaviors, and enable AI agents to monitor real-world social context and evolve system behaviors over long time-horizons. You will develop new paradigms for human+AI rating that are adaptive and context-aware, driving breakthroughs within Google DeepMind, Google products, and the broader AI alignment community.
Research Scientist, Gemini Safety
Google DeepMind
The Gemini Safety team at Google DeepMind is responsible for the safety and fairness of the latest Gemini models. As a Research Scientist/Engineer, you will apply and develop cutting-edge data and algorithmic solutions to advance these user-facing models. This is a fast-paced, highly collaborative role within a supportive team dedicated to pushing the boundaries of AI for public benefit and scientific discovery, with safety and ethics as the highest priorities.
Member of Technical Staff - Safety
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower users to control their intelligence and shape the future of AI. As a Member of Technical Staff - Safety, you will be instrumental in ensuring the safety and reliability of our AI models. This role involves owning the red-teaming and adversarial evaluation pipeline, translating safety findings into concrete guardrails, and validating that every release meets our risk thresholds before deployment. You will develop scalable, automated safety benchmarks and research state-of-the-art jailbreaking techniques and defenses to proactively address potential vulnerabilities.
Product Manager, Forge
Mistral AI
Mistral AI is seeking a talented and experienced Product Manager to define and execute the strategy for Forge, a product that empowers customers to build, fine-tune, and deploy custom AI models at scale. Forge transforms cutting-edge research into enterprise-ready capabilities by supporting model fine-tuning, reinforcement learning, and post-training workflows. This role operates at the intersection of research and product, enabling customers to train specialized models for real-world business value. You will collaborate closely with applied AI scientists and research engineers to translate frontier techniques into scalable and reliable solutions, shaping a 0-1 product with significant business impact and defining the future of how organizations train and deploy AI models.
Forward Deployed Engineer, Lead - LLM Post-training
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible. We are seeking an exceptional technical leader to build and scale our post-training and evaluation capabilities within the Applied AI team. This role involves taking our open-weight models and adapting them for specific customer domains, tasks, and constraints. You will own the end-to-end technical strategy for model customization, from synthetic data generation and reward modeling through training and production deployment, working directly with customers and research teams.