Reinforcement Learning Jobs

93 open roles mentioning Reinforcement Learning

Research Engineer / Scientist (Robot Learning)

20d ago
W

World-labs

World Labs is a frontier AI research and product company focused on spatial intelligence, co-founded by Dr. Fei-Fei Li, Justin Johnson, and Ben Mildenhall. The company is developing world models that can perceive, generate, reason, and interact with virtual and physical worlds, with its flagship product Marble transforming text, images, and video into navigable 3D worlds. Backed by leading investors, World Labs is building a world-class team at the intersection of AI research and real-world deployment.

$250k - $350k

San Francisco onsite
PythonC#PyTorch +2 more

Technical Program Manager, RL Research

24d ago
Anthropic

Anthropic

As a Technical Program Manager on the reinforcement learning team, you will own the systems and programs that determine how fast our research moves. This involves providing a trustworthy read on the state of RL research and managing the review and prioritization processes that translate that understanding into critical decisions for production RL runs. You will need strong technical depth, including the ability to debug data pipelines, analyze RL transcripts for issues, and make real-time allocation and quality decisions when research or production runs encounter problems. Equally important is organizational effectiveness: navigating a fast-growing organization, identifying key individuals and teams across research, infrastructure, product, and data operations, and coordinating their efforts to maintain velocity. Join us in our mission to build AI systems that are safe, reliable, and beneficial to humanity.

San Francisco, CA | New York City, NY onsite
AnthropicClaudeReinforcement Learning +1 more

Staff Software Engineer, Environments Infrastructure

24d ago
Anthropic

Anthropic

Anthropic's Environments organization is responsible for building and maintaining the infrastructure that enhances Claude's capabilities through reinforcement learning. This includes developing frameworks for researchers to create environments and managing the infrastructure that runs them, with a core mission to productionize research. The role involves embedding with research teams to understand their workflows and designing frameworks and APIs that accelerate their progress, ensuring these systems are understandable, ownable, and maintainable. A key aspect is also ensuring the health, maintainability, monitoring, and ease of triage for production RL runs.

San Francisco, CA | New York City, NY onsite
AnthropicPythonClaude +2 more

Staff Software Engineer, Code RL

24d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. This role focuses on the engineering aspects of reinforcement learning for Claude's coding capabilities, involving the creation and scaling of agentic coding environments. You will have significant influence on technical direction and standards, embedding with research teams to understand their needs, build supporting frameworks and infrastructure, and then transfer ownership of these well-maintained systems. The role also includes ensuring the ongoing health and maintainability of production RL runs, including monitoring and triage tooling. The team's work spans client-side sandboxed execution for agentic RL, large-scale data processing, dataset lifecycle management, and the frameworks researchers use to build environments. You will focus on areas where your deep expertise is most valuable, particularly if you have strong Python skills, a keen eye for API and framework design, and experience with complex system failures.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicPythonClaude +2 more

Research Engineer, Computer Use

24d ago
Anthropic

Anthropic

The Computer Use team focuses on teaching Claude to see, use, and understand computer interfaces. As a Research Engineer on the team, you'll work on advancing our models' ability to reliably and safely operate real software. We're looking for someone who's genuinely excited about both the research and the product sides of computer use. Your work will translate directly into model improvements in our own and our customers' products. You can try Claude's computer use capabilities today through the Claude in Chrome extension and Claude Cowork.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicPythonFine-Tuning +2 more

Research Engineer, Domain Scaling

24d ago
Anthropic

Anthropic

The Domain Scaling team aims to make Claude world-class at real-world knowledge work in domains like finance, healthcare, and legal. This role combines direct applied research with data sourcing (real-world and synthetic) to improve our models. You will own the end-to-end process of creating RL environments for new capabilities, which includes identifying high-value tasks, designing reward signals, managing vendor relationships, and measuring impact on model performance.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicFine-TuningClaude +1 more

Research Engineer, Life Sciences

24d ago
Anthropic

Anthropic

Anthropic is seeking an exceptional Research Engineer to join its Life Sciences team. This role focuses on accelerating progress in life sciences through AI, from early discovery to translation. You will leverage deep expertise in machine learning engineering to develop novel evaluation frameworks and training strategies, pushing the boundaries of AI in biology. Working at the intersection of AI and biological sciences, you will develop rigorous methods to measure and improve model performance on complex scientific tasks, collaborating with researchers and engineers to build AI systems for all phases of research and development, while upholding Anthropic's commitment to safety and beneficial impact. Previous experience in life sciences is welcome but not required.

San Francisco, CA onsite
AnthropicDockerKubernetes +2 more

Research Engineer, Code RL (Reinforcement Learning)

24d ago
Anthropic

Anthropic

We are seeking a Research Engineer for our Code RL team, focused on advancing AI models' capabilities in writing, editing, testing, debugging, and shipping real software. This role involves designing RL environments, coding tasks, and reward signals, as well as running training experiments on frontier models. You will diagnose model performance, improve pipeline speed and reliability, and contribute to areas like agentic coding behaviors, code correctness, and autonomous engineering. The position blends cutting-edge research with practical engineering to build high-quality, scalable AI systems.

San Francisco, CA | New York City, NY onsite
AnthropicPythonFine-Tuning +5 more

Research Engineer, Chip Design RL (Reinforcement Learning)

24d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer for its Code RL team to advance AI models' ability to design silicon. This role involves leveraging chip design expertise to create tasks and signals for models, focusing on hardware design domains. The position is at the intersection of cutting-edge research and engineering excellence, with a commitment to building high-quality, scalable systems that push the boundaries of AI capabilities.

San Francisco, CA | New York City, NY onsite
AnthropicClaudeReinforcement Learning

Research Engineer, Performance RL (Reinforcement Learning)

24d ago
Anthropic

Anthropic

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the RL Teams Our Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.6 and Opus 4.6. Our work spans several key areas: - Developing systems that enable models to use computers effectively - Advancing code generation through reinforcement learning - Pioneering fundamental RL research for large language models - Building scalable RL infrastructure and training methodologies - Enhancing model reasoning capabilities We collaborate closely with Anthropic's alignment and frontier red teams to ensure our systems are both capable and safe. We partner with the applied production training team to bring research innovations into deployed models, and are dedicated to implement our research at scale. Our Reinforcement Learning teams sit at the intersection of cutting-edge research and engineering excellence, with a deep commitment to building high-quality, scalable systems that push the boundaries of what AI can accomplish. About the Role We're hiring for the Code RL team within the RL organization. As a Research Engineer, you'll advance our models' ability to safely write correct, fast code for accelerators. You'll need to know accelerator performance well to turn it into tasks and signals models can learn from. Specifically, you will: - Invent, design and implement RL environments and evaluations. - Conduct experiments and shape our research roadmap. - Deliver your work into training runs. - Collaborate with other researchers, engineers, and performance engineering specialists across and outside Anthropic. You may be a good fit if you: - Have expertise with accelerators (CUDA, ROCm, Triton, Pallas), ML framework programming (JAX or PyTorch). - Have worked across the stack – kernels, model code, distributed systems. - Know how to balance research exploration with engineering implementation. - Are passionate about AI's potential and committed to developing safe and beneficial systems. Strong candidates may also have: - Experience with reinforcement learning. - Experience porting ML workloads between different types of accelerators. - Familiarity with LLM training methodologies. The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: $350,000—$850,000 USD Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings. How we're different We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills. The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Come work with us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.

San Francisco, CA onsite
AnthropicPyTorchClaude +2 more

Model Performance Software Engineer, Claude Code

24d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to be safe and beneficial for users and society. We are a growing team of researchers, engineers, policy experts, and business leaders dedicated to this mission. We are seeking a Staff Software Engineer to lead technical direction at the intersection of engineering and research for the Claude Code team. In this role, you will collaborate with researchers and engineering leadership to define how we measure, understand, and enhance Claude's coding abilities. You will architect the systems, tooling, and evaluation infrastructure that accelerate our research progress and be responsible for technical decisions impacting the team and beyond. This senior individual contributor position is for someone with a proven track record of building and owning large-scale systems, ready to take on a technical leadership role by driving architecture, mentoring engineers, and influencing the future of Claude Code.

San Francisco, CA | New York City, NY onsite
AnthropicPythonTypeScript +2 more

Research Engineer, Visual Knowledge Work

24d ago
Anthropic

Anthropic

We are seeking research engineers with a strong computer vision background to enhance the visual and spatial reasoning capabilities of our state-of-the-art Claude models. This role involves research, development, and evaluation, taking a full-stack approach across pretraining, RL, and runtime techniques. You will collaborate closely with the product organization to ensure that vision improvements directly impact Claude's performance on real-world tasks and address customer challenges.

New York City, NY; San Francisco, CA; Seattle, WA onsite
AnthropicFine-TuningClaude +3 more

Research Engineer, Universes

24d ago
Anthropic

Anthropic

The Universes team within Research is responsible for training AI models to perform complex, difficult, long-horizon agentic tasks in ultra-realistic settings. We design and implement novel training environments that go far beyond what models can do today — environments where models learn to navigate ambiguity, handle interruptions, maintain context over extended interactions, and exercise judgment in open-ended scenarios. We're looking for Research Engineers to help us build the next generation of training environments for capable and safe agentic AI. This role blends research and engineering responsibilities, requiring you to both implement novel approaches and contribute to research direction. You'll work on fundamental research in reinforcement learning, designing training environments and methodologies that push the state of the art, and building evaluations that measure genuine capability.

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY remote
AnthropicGoFine-Tuning +1 more

Research Engineer, Cybersecurity RL (Reinforcement Learning)

24d ago
Anthropic

Anthropic

The Cybersecurity RL team within Anthropic is seeking a Research Engineer to advance the capabilities of AI models in secure coding, vulnerability remediation, and defensive cybersecurity. This role combines research and engineering, requiring the design and implementation of RL environments, conducting experiments, delivering work into production training runs, and collaborating with cross-functional teams. The ideal candidate will have domain expertise in cybersecurity and an interest or experience in training safe AI models, potentially coming from backgrounds like white hat hacking, security engineering, or detection and response.

San Francisco, CA | New York City, NY onsite
AnthropicClaudeReinforcement Learning

Research Engineer, RL Engineering

24d ago
Anthropic

Anthropic

Anthropic's mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. As an ML Systems Engineer on the Reinforcement Learning Engineering team, you will build and improve the critical algorithms and infrastructure that researchers use to train AI models like Claude. Your work will directly enable breakthroughs in AI capabilities and safety, focusing on enhancing the performance, robustness, and usability of these systems to accelerate research progress. You will support and empower the research team in their mission to build beneficial AI systems, specifically by building, maintaining, and improving the algorithms and systems used for finetuning production and research models with methods like RLHF.

San Francisco, CA | New York City, NY | Seattle, WA onsite
AnthropicPythonFine-Tuning +3 more

ML/Research Engineer, Safeguards

24d ago
Anthropic

Anthropic

Anthropic is seeking ML Engineers and Research Engineers to join the Safeguards ML team. The primary focus of this role is to develop systems that detect and mitigate misuse of AI systems, ranging from individual policy violations to sophisticated coordinated attacks. You will build defenses to ensure product safety as AI capabilities advance, protect user well-being, and guarantee appropriate model behavior across various contexts. This work is crucial for Anthropic's Responsible Scaling Policy commitments.

San Francisco, CA | New York City, NY onsite
AnthropicPythonReinforcement Learning +1 more

Research Engineer/Research Scientist, Pre-training

24d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer to join its Pre-training team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, contributing to the creation of safe, steerable, and trustworthy AI systems. The team is dedicated to ensuring that transformative AI systems are aligned with human interests and societal benefit.

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY remote
AnthropicKubernetesPython +4 more

Research Engineer / Scientist, Alignment

24d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer/Scientist for its Alignment Science team. This role involves designing and executing machine learning experiments to understand and steer the behavior of advanced AI systems, with a focus on AI safety and potential risks from future human-level AI. You will collaborate with other teams on exploratory research, contributing to Anthropic's mission of creating reliable, interpretable, and steerable AI systems that are helpful, honest, and harmless.

San Francisco, CA onsite
AnthropicKubernetesPython +4 more

Staff Research Engineer, Discovery Team

24d ago
Anthropic

Anthropic

Anthropic is dedicated to building reliable, interpretable, and steerable AI systems that are safe and beneficial for society. As a Research Engineer on the Discovery Team, you will work end-to-end to identify and address key blockers on the path to scientific Artificial General Intelligence (AGI). This role involves improving models' abilities to use computers, acting as a laboratory for long-horizon tasks and a crucial component for scientific workflows. You will collaborate with a team of researchers and engineers focused on pushing the scientific frontier.

San Francisco, CA onsite
AnthropicDockerKubernetes +2 more

Research Engineer, Knowledge Team

24d ago
Anthropic

Anthropic

Anthropic is seeking Research Engineers to reimagine how Claude interacts with external data sources. This role involves designing novel architectures for organizing information and training language models to effectively utilize these architectures. The goal is to move beyond traditional data paradigms to accommodate the capabilities of Large Language Models (LLMs).

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY remote
AnthropicPythonRAG +3 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.