Reinforcement Learning Jobs

50 open roles mentioning Reinforcement Learning

Staff Software Engineer, Code RL

6d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. This role focuses on the engineering aspects of reinforcement learning for Claude's coding capabilities, involving the creation and scaling of agentic coding environments. You will have significant influence on technical direction and standards, embedding with research teams to understand their needs, build supporting frameworks and infrastructure, and then transfer ownership of these well-maintained systems. The role also includes ensuring the ongoing health and maintainability of production RL runs, including monitoring and triage tooling. The team's work spans client-side sandboxed execution for agentic RL, large-scale data processing, dataset lifecycle management, and the frameworks researchers use to build environments. You will focus on areas where your deep expertise is most valuable, particularly if you have strong Python skills, a keen eye for API and framework design, and experience with complex system failures.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicPythonClaude +2 more

Staff Software Engineer, Environments Infrastructure

6d ago
Anthropic

Anthropic

Anthropic's Environments organization is responsible for building and maintaining the infrastructure that enhances Claude's capabilities through reinforcement learning. This includes developing frameworks for researchers to create environments and managing the infrastructure that runs them, with a core mission to productionize research. The role involves embedding with research teams to understand their workflows and designing frameworks and APIs that accelerate their progress, ensuring these systems are understandable, ownable, and maintainable. A key aspect is also ensuring the health, maintainability, monitoring, and ease of triage for production RL runs.

San Francisco, CA | New York City, NY onsite
AnthropicPythonClaude +2 more

Research Engineer, Developer Experience, Tinker

16d ago
thinkingmachines

thinkingmachines

Thinking Machines Lab is seeking a Research Engineer focused on developer experience to build and enhance their Tinker platform. This role involves working hands-on with users to understand their challenges and translate them into product improvements. You will be responsible for creating and updating documentation, adding library features, prototyping integrations, and ensuring users can smoothly customize frontier AI models. This position acts as a crucial link between Tinker users and the internal research and infrastructure teams, surfacing user patterns to inform product and infrastructure priorities and sharing learnings through various channels.

$350k - $475k

San Francisco onsite
OpenAIMistralFine-Tuning +2 more

Model Performance Software Engineer, Claude Code

16d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. We are a growing team of researchers, engineers, policy experts, and business leaders dedicated to this mission. We are seeking a Staff Software Engineer to lead technical direction at the intersection of engineering and research for the Claude Code team. In this role, you will collaborate with researchers and engineering leadership to define how we measure, understand, and enhance Claude's coding abilities. You will architect the systems, tooling, and evaluation infrastructure that accelerate our research progress and be responsible for technical decisions impacting the team and beyond. This senior individual contributor position is for someone with a proven track record of building and owning large-scale systems, ready to take on a technical leadership role by driving architecture, mentoring engineers, and influencing the future of Claude Code.

San Francisco, CA | New York City, NY onsite
AnthropicPythonTypeScript +2 more

Research Engineer, Domain Scaling

16d ago
Anthropic

Anthropic

The Domain Scaling team aims to make Claude world-class at real-world knowledge work in domains like finance, healthcare, and legal. This role combines direct applied research with data sourcing (real-world and synthetic) to improve our models. You will own the end-to-end process of creating RL environments for new capabilities, which includes identifying high-value tasks, designing reward signals, managing vendor relationships, and measuring impact on model performance.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicFine-TuningClaude +1 more

Research Lead, Training Insights

16d ago
Anthropic

Anthropic

As a Research Lead on the Training Insights team, you will develop the strategy for, and lead execution on, how we measure and characterize model capabilities across training and deployment. This is a hands-on leadership role where you will drive original research into new evaluation methodologies while leading a small team of researchers and research engineers. Your work will span the full lifecycle of model development, from researching and building new long-horizon evaluations to developing novel approaches for measuring emerging capabilities and deepening our understanding of how those capabilities develop. You will also take a cross-organizational view, working across various teams to map the landscape of model evaluations and identify critical gaps. This role carries significant visibility and impact, helping to shape the evaluation narrative for model releases and contributing directly to how Anthropic communicates about its models. Done well, you will change how the industry measures and understands model capabilities, significantly furthering our safety mission.

Remote-Friendly (Travel Required) | San Francisco, CA; San Francisco, CA | New York City, NY remote
AnthropicReinforcement Learning

Software Engineer, RL Data

16d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. This senior, foundational role on a new team involves making key architectural decisions and shaping initial development. The work is hands-on and varied, encompassing pipeline and infrastructure engineering, prompt tuning, and supporting research teams. The RL Data team focuses on building systems for high-quality reinforcement learning data for Claude, including data collection pipelines, human feedback tooling, execution environments, and quality assurance to ensure trustworthy training data at scale. The goal is to enhance Claude's capabilities in real-world tasks, particularly in AI safety research and beneficial AI deployments.

San Francisco, CA | New York City, NY onsite
AnthropicDockerKubernetes +4 more

Anthropic Fellows Program, The Anthropic Institute (Economics & Policy)

16d ago
Anthropic

Anthropic

Anthropic is seeking talented individuals for its Fellows Program, designed to foster AI research and engineering talent. The program provides funding and mentorship to promising technical individuals, regardless of prior experience, to work on empirical projects aligned with Anthropic's research priorities. The goal is to produce public outputs, such as paper submissions. Fellows will primarily utilize external infrastructure like open-source models and public APIs. The program emphasizes AI safety and beneficial AI development for society.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +1 more

Research Engineer, Life Sciences

16d ago
Anthropic

Anthropic

Anthropic is seeking an exceptional Research Engineer to join its Life Sciences team. This role focuses on accelerating progress in life sciences through AI, from early discovery to translation. You will leverage deep expertise in machine learning engineering to develop novel evaluation frameworks and training strategies, pushing the boundaries of AI in biology. Working at the intersection of AI and biological sciences, you will develop rigorous methods to measure and improve model performance on complex scientific tasks, collaborating with researchers and engineers to build AI systems for all phases of research and development, while upholding Anthropic's commitment to safety and beneficial impact. Previous experience in life sciences is welcome but not required.

San Francisco, CA onsite
AnthropicDockerKubernetes +2 more

Research Engineer, Computer Use

16d ago
Anthropic

Anthropic

The Computer Use team focuses on teaching Claude to see, use, and understand computer interfaces. As a Research Engineer on the team, you'll work on advancing our models' ability to reliably and safely operate real software. We're looking for someone who's genuinely excited about both the research and the product sides of computer use. Your work will translate directly into model improvements in our own and our customers' products. You can try Claude's computer use capabilities today through the Claude in Chrome extension and Claude Cowork.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicPythonFine-Tuning +2 more

Research Engineer, Visual Knowledge Work

16d ago
Anthropic

Anthropic

We are seeking research engineers with a strong computer vision background to enhance the visual and spatial reasoning capabilities of our state-of-the-art Claude models. This role involves research, development, and evaluation, taking a full-stack approach across pretraining, RL, and runtime techniques. You will collaborate closely with the product organization to ensure that vision improvements directly impact Claude's performance on real-world tasks and address customer challenges.

New York City, NY; San Francisco, CA; Seattle, WA onsite
AnthropicFine-TuningClaude +3 more

Research Engineer, Universes

16d ago
Anthropic

Anthropic

The Universes team within Research is responsible for training AI models to perform complex, difficult, long-horizon agentic tasks in ultra-realistic settings. We design and implement novel training environments that go far beyond what models can do today — environments where models learn to navigate ambiguity, handle interruptions, maintain context over extended interactions, and exercise judgment in open-ended scenarios. We're looking for Research Engineers to help us build the next generation of training environments for capable and safe agentic AI. This role blends research and engineering responsibilities, requiring you to both implement novel approaches and contribute to research direction. You'll work on fundamental research in reinforcement learning, designing training environments and methodologies that push the state of the art, and building evaluations that measure genuine capability.

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY remote
AnthropicGoFine-Tuning +1 more

Anthropic Fellows Program, Reinforcement Learning

16d ago
Anthropic

Anthropic

Anthropic is seeking talented individuals for its Fellows Program, focusing on AI research and engineering. This program provides funding and mentorship to promising technical talent, regardless of prior experience, to work on empirical projects aligned with Anthropic's research priorities. The goal is to produce public outputs, such as research papers, contributing to the development of reliable, interpretable, and steerable AI systems that are safe and beneficial for society. Fellows will engage in full-time research for four months, with opportunities for extension, and will be mentored by Anthropic researchers.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +3 more

Anthropic Fellows Program, ML Systems & Performance

16d ago
Anthropic

Anthropic

Anthropic is seeking talented individuals for its Fellows Program, focusing on AI research and engineering to ensure AI systems are reliable, interpretable, and steerable. This program provides funding and mentorship to promising technical talent, regardless of prior experience, to work on empirical projects aligned with Anthropic's research priorities. The goal is to produce public outputs, such as paper submissions, contributing to the development of safe and beneficial AI for society. Fellows will engage in full-time research for four months, with direct mentorship from Anthropic researchers and access to a shared workspace in either Berkeley, California, or London, UK. The program also offers connections to the broader AI safety and security research community.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +3 more

Anthropic Fellows Program, AI Security

16d ago
Anthropic

Anthropic

Anthropic is seeking talented individuals for the Anthropic Fellows Program, focusing on AI Security. This program is designed to foster AI research and engineering talent by providing funding and mentorship to promising individuals, regardless of prior experience. Fellows will work on empirical projects using external infrastructure, aiming to produce public outputs like research papers. The program offers a structured 4-month full-time research period with direct mentorship from Anthropic researchers, access to a shared workspace in Berkeley or London, and connections to the AI safety and security research community.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +3 more

Anthropic Fellows Program, AI Safety

16d ago
Anthropic

Anthropic

Anthropic is seeking talented individuals for its Fellows Program, focusing on AI Safety. This program aims to foster AI research and engineering talent by providing funding and mentorship to promising individuals, regardless of prior experience. Fellows will engage in empirical projects using external infrastructure and open-source models, with the goal of producing public outputs like research papers. The program offers a structured 4-month full-time research period, direct mentorship from Anthropic researchers, access to shared workspaces in Berkeley or London, and connections to the AI safety community. The program is designed to encourage diverse perspectives and applications, even from those who may not meet every single qualification.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +1 more

Anthropic Fellows Program

16d ago
Anthropic

Anthropic

Anthropic is seeking candidates for its Fellows Program, designed to cultivate AI research and engineering talent. This program provides funding and mentorship to promising individuals, irrespective of prior experience, to work on empirical projects aligned with Anthropic's research priorities. The goal is to produce public outputs, such as paper submissions, with a strong track record of fellows achieving this in previous cohorts. The program offers a 4-month full-time research period, direct mentorship from Anthropic researchers, access to a shared workspace in Berkeley or London, and connections to the AI safety and security research community. Applications are reviewed on a rolling basis for cohorts starting in July 2026 and beyond.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +5 more

Staff Research Engineer, Discovery Team

16d ago
Anthropic

Anthropic

Anthropic is dedicated to building reliable, interpretable, and steerable AI systems that are safe and beneficial for society. As a Research Engineer on the Discovery Team, you will work end-to-end to identify and address key blockers on the path to scientific Artificial General Intelligence (AGI). This role involves improving models' abilities to use computers, acting as a laboratory for long-horizon tasks and a crucial component for scientific workflows. You will collaborate with a team of researchers and engineers focused on pushing the scientific frontier.

San Francisco, CA onsite
AnthropicDockerKubernetes +2 more

ML/Research Engineer, Safeguards

16d ago
Anthropic

Anthropic

Anthropic is seeking ML Engineers and Research Engineers to join the Safeguards ML team. The primary focus of this role is to develop systems that detect and mitigate misuse of AI systems, ranging from individual policy violations to sophisticated coordinated attacks. You will build defenses to ensure product safety as AI capabilities advance, protect user well-being, and guarantee appropriate model behavior across various contexts. This work is crucial for Anthropic's Responsible Scaling Policy commitments.

San Francisco, CA | New York City, NY onsite
AnthropicPythonReinforcement Learning +1 more

[Expression of Interest] Research Engineer / Scientist, Alignment - London

16d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to ensure AI is safe and beneficial for society. As a Research Engineer on the Alignment Science team in London, you will design and execute machine learning experiments to understand and steer the behavior of advanced AI systems. You will focus on AI safety, particularly risks from future human-level AI systems, collaborating with teams like Interpretability and Frontier Red Team. The role involves exploratory research in areas such as AI Control and Alignment Stress-testing, aiming to make AI helpful, honest, and harmless.

London, UK onsite
AnthropicKubernetesPython +3 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.