Reinforcement Learning Jobs

50 open roles mentioning Reinforcement Learning

Research Engineer / Scientist, Alignment

16d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer/Scientist for its Alignment Science team. This role involves designing and executing machine learning experiments to understand and steer the behavior of advanced AI systems, with a focus on AI safety and potential risks from future human-level AI. You will collaborate with other teams on exploratory research, contributing to Anthropic's mission of creating reliable, interpretable, and steerable AI systems that are helpful, honest, and harmless.

San Francisco, CA onsite
AnthropicKubernetesPython +4 more

Research Engineer/Research Scientist, Pre-training

16d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer to join its Pre-training team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, contributing to the creation of safe, steerable, and trustworthy AI systems. The team is dedicated to ensuring that transformative AI systems are aligned with human interests and societal benefit.

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY remote
AnthropicKubernetesPython +4 more

Research Engineer, Pretraining

16d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer to join its Pretraining team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, aiming to build safe, steerable, and trustworthy AI systems. The mission is to ensure that transformative AI systems are aligned with human interests and are beneficial for society. The team is dedicated to pushing the boundaries of AI while prioritizing safety and ethics.

London, UK onsite
AnthropicKubernetesPython +4 more

Research Engineer, Performance RL (Reinforcement Learning)

16d ago
Anthropic

Anthropic

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the RL Teams Our Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.6 and Opus 4.6. Our work spans several key areas: - Developing systems that enable models to use computers effectively - Advancing code generation through reinforcement learning - Pioneering fundamental RL research for large language models - Building scalable RL infrastructure and training methodologies - Enhancing model reasoning capabilities We collaborate closely with Anthropic's alignment and frontier red teams to ensure our systems are both capable and safe. We partner with the applied production training team to bring research innovations into deployed models, and are dedicated to implement our research at scale. Our Reinforcement Learning teams sit at the intersection of cutting-edge research and engineering excellence, with a deep commitment to building high-quality, scalable systems that push the boundaries of what AI can accomplish. About the Role We're hiring for the Code RL team within the RL organization. As a Research Engineer, you'll advance our models' ability to safely write correct, fast code for accelerators. You'll need to know accelerator performance well to turn it into tasks and signals models can learn from. Specifically, you will: - Invent, design and implement RL environments and evaluations. - Conduct experiments and shape our research roadmap. - Deliver your work into training runs. - Collaborate with other researchers, engineers, and performance engineering specialists across and outside Anthropic. You may be a good fit if you: - Have expertise with accelerators (CUDA, ROCm, Triton, Pallas), ML framework programming (JAX or PyTorch). - Have worked across the stack – kernels, model code, distributed systems. - Know how to balance research exploration with engineering implementation. - Are passionate about AI's potential and committed to developing safe and beneficial systems. Strong candidates may also have: - Experience with reinforcement learning. - Experience porting ML workloads between different types of accelerators. - Familiarity with LLM training methodologies. The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: $350,000—$850,000 USD Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings. How we're different We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills. The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Come work with us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.

San Francisco, CA onsite
AnthropicPyTorchClaude +2 more

Research Engineer, Knowledge Team

16d ago
Anthropic

Anthropic

Anthropic is seeking Research Engineers to reimagine how Claude interacts with external data sources. This role involves designing novel architectures for organizing information and training language models to effectively utilize these architectures. The goal is to move beyond traditional data paradigms to accommodate the capabilities of Large Language Models (LLMs).

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY remote
AnthropicPythonRAG +3 more

Research Engineer, Cybersecurity RL (Reinforcement Learning)

16d ago
Anthropic

Anthropic

The Cybersecurity RL team within Anthropic is seeking a Research Engineer to advance the capabilities of AI models in secure coding, vulnerability remediation, and defensive cybersecurity. This role combines research and engineering, requiring the design and implementation of RL environments, conducting experiments, delivering work into production training runs, and collaborating with cross-functional teams. The ideal candidate will have domain expertise in cybersecurity and an interest or experience in training safe AI models, potentially coming from backgrounds like white hat hacking, security engineering, or detection and response.

San Francisco, CA | New York City, NY onsite
AnthropicClaudeReinforcement Learning

Research Engineer, Code RL (Reinforcement Learning)

16d ago
Anthropic

Anthropic

We are seeking a Research Engineer for our Code RL team, focused on advancing AI models' capabilities in writing, editing, testing, debugging, and shipping real software. This role involves designing RL environments, coding tasks, and reward signals, as well as running training experiments on frontier models. You will diagnose model performance, improve pipeline speed and reliability, and contribute to areas like agentic coding behaviors, code correctness, and autonomous engineering. The position blends cutting-edge research with practical engineering to build high-quality, scalable AI systems.

San Francisco, CA | New York City, NY onsite
AnthropicPythonFine-Tuning +5 more

Research Engineer, Chip Design RL (Reinforcement Learning)

16d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer for its Code RL team to advance AI models' ability to design silicon. This role involves leveraging chip design expertise to create tasks and signals for models, focusing on hardware design domains. The position is at the intersection of cutting-edge research and engineering excellence, with a commitment to building high-quality, scalable systems that push the boundaries of AI capabilities.

San Francisco, CA | New York City, NY onsite
AnthropicClaudeReinforcement Learning

Member of Technical Staff - Post-Training

1mo ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible for everyone to use, customize, and build on. We are building open models that empower individuals to control their intelligence and shape the future of AI. As a Member of Technical Staff - Post-Training, you will play a crucial role in transforming powerful pre-trained models into aligned and general agents. This position involves driving research and engineering initiatives at the forefront of post-training techniques, from data curation to large-scale optimization, and contributing to the advancement of large model reasoning and instruction following capabilities.

San Francisco, CA onsite FullTime
Reinforcement Learning

Senior AI Product Manager, Healthcare Agents

1mo ago
Scale AI

Scale AI

Scale is seeking an AI Product Manager to lead the Healthcare vertical within the Agents Data & Reinforcement Learning Environments team. This role involves owning the development of realistic RL environments for training and evaluating AI agents in healthcare software and workflows, as well as defining the "data as a product" strategy that supports them. The ideal candidate will possess deep understanding of the Healthcare industry and its workflows, combined with insight into AI research and current agent capabilities. You will translate this expertise into environments and datasets that enable AI agents to perform real healthcare tasks, serving as the domain expert for Scale's key customers and researchers. A strong entrepreneurial and go-to-market mindset is essential for success in this position.

$206k - $257k

San Francisco, CA; Seattle, WA; New York, NY remote
GoReinforcement LearningAI +8 more

Senior AI Product Manager, Finance Agents

1mo ago
Scale AI

Scale AI

Scale is seeking an AI Product Manager to lead the Finance vertical within the Agents Data & Reinforcement Learning Environments team. This role involves owning the development of realistic RL environments for training and evaluating AI agents in financial workflows, as well as defining the "data as a product" strategy that supports them. You will leverage your deep understanding of the Finance industry and AI capabilities to identify valuable financial tasks for AI modeling, determine data sourcing and structuring strategies, and translate domain expertise into a defensible product. The ideal candidate will possess a strong entrepreneurial and go-to-market mindset, coupled with the ability to pair finance industry experience with an understanding of AI research and current agent capabilities in financial workflows.

$205k - $257k

San Francisco, CA; Seattle, WA; New York, NY onsite
GoReinforcement LearningAI +5 more

Research Intern RL & Post-Training Systems, Turbo (Fall 2026)

1mo ago
Together AI

Together AI

The Turbo Research team focuses on making post-training and reinforcement learning for large language models efficient, scalable, and reliable. This work intersects RL algorithms, inference systems, and large-scale experimentation, where inference costs significantly impact training efficiency and the practicality of learning algorithms. As a research intern, you will investigate RL and post-training methods whose performance and scalability are closely tied to inference behavior, co-designing algorithms and systems. Projects aim to enable new experimental regimes, including larger models, longer rollouts, and more complex evaluations, by re-evaluating the interaction between inference, scheduling, and training.

San Francisco remote
PythonC#NLP +7 more

Software Engineer, Identity

1mo ago
Scale AI

Scale AI

Scale is seeking a Software Engineer, Identity to join our Platform Engineering team. In this role, you will be instrumental in designing and developing core platforms and software systems, with a specific focus on identity, access management, authorization, and authentication. You will gain broad exposure to the cutting edge of the AI industry as Scale supports enterprises, startups, and governments. This position offers the opportunity to contribute to the foundational elements of products that power advanced LLMs and generative models, playing a crucial role in how humanity interacts with AI.

$216k - $270k

San Francisco, CA; New York, NY remote
AWSAzurePython +14 more

AI research scientist

1mo ago
W

Writer

AI research at WRITER focuses on building the scientific foundation for ambitious enterprise AI deployments. As a staff AI research scientist, you will drive a high-impact research agenda centered on large language models, agentic reasoning, and system-level capabilities essential for enterprise-scale AI. This role offers a unique opportunity to advance the field while directly contributing to products used by hundreds of thousands daily. You will work on post-training, planning, multi-step reasoning, and agentic workflows, directly shaping the future of enterprise AI performance and scalability. The role provides resources, infrastructure, and cross-functional support to pursue and implement ambitious ideas rapidly.

San Francisco, CA hybrid FullTime
PythonAI AgentsFine-Tuning +5 more

Research Intern, Model Shaping (Fall 2026)

1mo ago
Together AI

Together AI

As a Research Intern in the Model Shaping team, you will work on advanced post-training methods, new techniques for efficient neural network training, and robust evaluation of foundation model capabilities. The Model Shaping team at Together AI focuses on tailoring open foundation models for downstream applications, building services for machine learning developers, and developing new methods for efficient model training and evaluation. This role offers the opportunity to contribute to cutting-edge research and potentially influence open-source projects.

San Francisco hybrid
PyTorchNLPReinforcement Learning +6 more

Technical Program Manager, Engineering

1mo ago
Scale AI

Scale AI

Scale is at the forefront of the AI revolution, developing data engines and technologies that power the world's leading LLMs. This role focuses on leading critical programs within the Platform and Security Engineering teams, overseeing the design and development of core data storage systems and security initiatives. You will drive company-wide programs, improve processes, and ensure alignment with industry standards, gaining exposure to the cutting edge of AI adoption across various sectors. The work is crucial for making AI models safe, aligned, and useful through human evaluation and reinforcement learning.

$181k - $226k

San Francisco, CA; New York, NY onsite
AWSFine-TuningSQL +2 more

Senior/Staff Machine Learning Research Engineer, General Agents, Enterprise GenAI

1mo ago
Scale AI

Scale AI

Scale AI is seeking a Senior/Staff Machine Learning Engineer for its General Agents team. This role is crucial in designing, building, and deploying production-ready AI agents to address high-impact enterprise challenges. You will be involved in the entire agent lifecycle, from conceptualization and system design to evaluation, deployment, and ongoing iteration. The position requires bridging cutting-edge agentic techniques with the practical demands of real-world customer environments, focusing on creating scalable, reliable, and generalizable agent systems.

$265k - $331k

San Francisco, CA; New York, NY onsite
OpenAIPythonAI Agents +10 more

Software Engineer, Platform

1mo ago
Scale AI

Scale AI

Scale is at the forefront of the AI revolution, building the Generative AI Data Engine and other products that power the world's most advanced LLMs. The Platform Engineering team is foundational to these efforts, responsible for designing and developing shared platforms, architecting core cloud infrastructure, and redefining software development processes. This role offers exposure to the cutting edge of AI development across various sectors, from startups to governments. You will drive the design and implementation of critical platforms, collaborate with cross-functional teams, and proactively improve engineering practices. This is an opportunity to shape the future of AI infrastructure and contribute to some of the most important work in how humanity interacts with AI.

$216k - $270k

San Francisco, CA; New York, NY remote
AWSDockerKubernetes +11 more

Technical Lead Manager, Physical AI

2mo ago
Scale AI

Scale AI

Scale AI is seeking a Technical Lead Manager for its Physical AI team, focusing on the development of general AI that can reason and act in the physical world. This role bridges cutting-edge Machine Learning research with physical robot deployment, leading a team of Research Engineers while remaining a hands-on technical contributor. The primary focus is on developing and evaluating Large-Scale Foundation Models, such as VLAs and World models, to enable robots and autonomous vehicles to generalize across diverse tasks and environments. The team leverages Scale's extensive data infrastructure to help build Foundation Models for Physical AI, aiming to redefine the future of automation.

$249k - $311k

San Francisco, CA remote
Fine-TuningPyTorchReinforcement Learning +9 more

Staff Software Engineer, Data Platform

2mo ago
Scale AI

Scale AI

Scale is at the forefront of the AI revolution, developing data engines and technologies that power the world's leading LLMs and generative models. This role is on the Platform Engineering team, responsible for the foundational data infrastructure that supports these cutting-edge AI products. You will lead the design and development of core data storage, streaming, caching, and indexing platforms, gaining exposure to the rapidly evolving AI landscape across various industries. The work involves driving architecture, implementation, and reliability of these critical systems, collaborating with stakeholders, and mentoring junior engineers.

$252k - $315k

San Francisco, CA; New York, NY remote
KubernetesPythonFine-Tuning +10 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.