Reinforcement Learning Jobs

93 open roles mentioning Reinforcement Learning

Engineering Manager (TLM, Agents)

22h ago
P

Perplexity AI

Perplexity is looking for a Tech Lead Manager (TLM) to lead and grow its Agents engineering team. This team comprises AI/ML, backend, and full-stack engineers focused on building agentic experiences within the Comet ecosystem. The vision is to create AI agents that can fulfill user intentions through open-ended interactions with the world. As the Agents TLM, you will combine AI expertise, product intuition, and engineering management skills to push the boundaries of what AI agents can achieve for millions of users. You will lead, develop, and support a team in tackling significant challenges in AI, such as designing agents for digital navigation, training models for complex decision-making, ensuring excellent user experiences across various platforms, developing secure agentic capabilities, and optimizing agent-environment interactions.

San Francisco onsite FullTime
PythonTypeScriptGo +3 more

Member of Technical Staff (AI Software Engineer, Agents)

22h ago
P

Perplexity AI

Perplexity is seeking energetic engineers to join our highly driven Agents engineering team. The Agents team collaborates to build harnesses and AI systems powering delightful agentic experiences, including Perplexity Computer, the Comet ecosystem, and our Agent API. Our vision is to empower users with agents that faithfully actualize their intent through open-ended interactions with the world. As an engineer on our Agents team, you will bring AI expertise, sharp product intuition, and a tinkerer's mindset to advance the frontier of what agents can accomplish for our millions of devoted users, working across applied research and engineering to solve open problems in AI.

San Francisco onsite FullTime
PythonTypeScriptGo +3 more

Researcher, Safety Training, National Security

2d ago
OpenAI

OpenAI

We are seeking a researcher to train and evaluate models for U.S. government use, with a focus on national security applications. You will advance safety post-training and robustness, helping models follow nuanced policies while preserving their usefulness and capabilities. This role involves researching and implementing methods for safety training, reinforcement learning, and adversarial robustness, as well as developing evaluations, identifying model failure modes, and using findings to improve training. You will also collaborate with research, engineering, security, and policy partners to support safe, reliable deployment.

San Francisco onsite FullTime
OpenAIReinforcement LearningDeep Learning +1 more

Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

3d ago
P

Perplexity AI

Perplexity is building the foundations for agentic AI, transforming knowledge into action for millions of users. The Agent Capabilities team is at the forefront of AI research and product innovation, responsible for translating frontier AI breakthroughs into reusable product capabilities. This role offers broad ownership at the intersection of AI research, agent systems, platform engineering, and product innovation, allowing you to evaluate emerging models, identify user value, and build reliable, scalable experiences for users and agents.

San Francisco onsite FullTime
AWSPythonTypeScript +4 more

Vice President, Product Management, Managed AI

3d ago
c

crusoe

Crusoe is seeking a Vice President of Product Management for its rapidly scaling Managed AI Services portfolio. Launched in late 2025, this business has already secured significant ARR commitments and is expanding to include Managed Inference, Serverless Fine Tuning, Reinforcement Learning, and Agent Sandbox. As VP, you will be responsible for the vision, strategy, execution, and roadmap of this business, reporting to the SVP of Product Management. You will lead the portfolio as a business unit, fostering strong cross-functional partnerships and aligning with Crusoe's vertically integrated cloud strategy.

$345k - $385k

San Francisco, CA - US onsite FullTime
GoFine-TuningReinforcement Learning

Software Engineer, AI for Chip Design

5d ago
OpenAI

OpenAI

We are seeking a Software Engineer to develop the research infrastructure and tooling that will enable OpenAI models to design silicon. This role involves transforming chip-design workflows into robust environments for reinforcement learning and evaluation, and empowering researchers to conduct experiments and iterate on new ideas efficiently. You will bridge software engineering, tool integration, and open research challenges, leveraging strong coding fundamentals, sound technical judgment, and the ability to work independently. While prior chip-design experience is beneficial, domain knowledge can be acquired alongside our hardware specialists.

San Francisco hybrid FullTime
OpenAIPythonReinforcement Learning

Software Engineer - Voice Model

5d ago
x

xAI

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. The Grok Voice Model team is building the world's best voice AI, delivering smooth, natural, low-latency spoken interactions that are expressive, multilingual, and reliable across devices and real-time scenarios. The team owns the full training pipeline, from massive data curation and premium audio processing to frontier speech-language pre-training and intensive post-training to push quality, speed, and stability to the limit. The goal is to make talking to AI feel like conversing with a charming, kind, and knowledgeable person, and exceptionally smart, execution-oriented engineers are sought to achieve this.

$150k - $450k

Palo Alto, CA onsite
KubernetesPythonFine-Tuning +5 more

Member of Technical Staff - Sandbox Service

5d ago
x

xAI

The Sandbox service team at SpaceXAI builds and maintains a secure, scalable system that gives our models safe, controlled access to computational environments. This infrastructure powers critical workloads across training and product, enabling models to run code, build software, interact with tools, and even control applications with user interfaces. We provision containers and virtual machines on large-scale clusters, granting models interactive control over these remote environments. Our work spans the full stack: from orchestrating massive jobs and resource scheduling at the cluster level, to fine-tuning filesystem performance on nodes. The Sandbox service enables Grok to safely run and test code in real-time for user queries, and supports reinforcement learning in training, where models interactively explore tools ranging from compilers to productivity apps.

£107k - £262k

London, England, United Kingdom onsite
PythonGoRust +3 more

Member of Technical Staff - RL Training Framework

5d ago
x

xAI

SpaceXAI is seeking an engineer to join the RL infrastructure team to help develop our RL training framework. The team is small, highly motivated, and focused on engineering excellence, operating with a flat organizational structure where all employees are expected to be hands-on and contribute directly to the company's mission. This role requires strong initiative, curiosity, work ethic, prioritization, and communication skills.

$180k - $440k

Palo Alto, CA onsite
PythonRustC# +2 more

Anthropic Fellows Program, ML Systems & Reinforcement Learning

5d ago
Anthropic

Anthropic

Anthropic is seeking talented individuals for its Fellows Program, focusing on AI research and engineering to ensure AI systems are reliable, interpretable, and steerable. This program provides funding and mentorship to promising technical talent, regardless of prior experience, to work on empirical projects aligned with Anthropic's research priorities. The goal is to produce public outputs, such as paper submissions, contributing to the development of safe and beneficial AI for society. Fellows will engage in full-time research for four months, with direct mentorship from Anthropic researchers and access to a shared workspace in either Berkeley, California, or London, UK. The program also offers connections to the broader AI safety and security research community.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +4 more

Anthropic Fellows Program, AI Safety & Security

7d ago
Anthropic

Anthropic

Anthropic is seeking talented individuals for its Fellows Program, focusing on AI Safety. This program aims to foster AI research and engineering talent by providing funding and mentorship to promising individuals, regardless of prior experience. Fellows will engage in empirical projects using external infrastructure and open-source models, with the goal of producing public outputs like research papers. The program offers a structured 4-month full-time research period, direct mentorship from Anthropic researchers, access to shared workspaces in Berkeley or London, and connections to the AI safety community. The program is designed to encourage diverse perspectives and applications, even from those who may not meet every single qualification.

$25k - $50k

London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA remote
AnthropicPythonGo +3 more

Research Engineer, Safety

9d ago
d

decagon

Decagon is seeking a Research Engineer focused on Safety to ensure the reliability and controllability of their AI agents from evaluation through production. This role involves identifying real-world failure modes and developing the necessary models, evaluations, and safeguards to prevent them. The ideal candidate is a strong engineer passionate about advancing applied AI safety in production, with the autonomy to own their work end-to-end, ship impactful improvements, and make high-stakes technical decisions.

$200k - $400k

San Francisco onsite FullTime
PythonAI AgentsReinforcement Learning

Senior Research Engineer, Safety

9d ago
d

decagon

Decagon is seeking a Senior Research Engineer focused on Safety to join their conversational AI platform team. This role is crucial for ensuring the reliability and controllability of Decagon's AI agents from evaluation through production. You will be responsible for identifying real-world failure modes and developing the necessary models, evaluations, and safeguards to prevent them. This is an opportunity for strong engineers to advance applied AI safety in a production environment, owning their work end-to-end and making high-impact technical decisions.

$200k - $400k

San Francisco onsite FullTime
PythonAI AgentsReinforcement Learning

Applied Machine Learning Engineer, EMEA

11d ago
f

fireworks ai

Fireworks is seeking an Applied Machine Learning Engineer for the EMEA region. In this role, you will be the technical owner of customer engagements, embedding within client teams to understand their specific needs and challenges. You will be responsible for the entire lifecycle of a customer's deployment on the Fireworks platform, from initial scoping and model selection to ensuring production readiness, performance, and cost-efficiency. This position emphasizes first principles thinking and requires a deep understanding of software engineering, machine learning techniques, and infrastructure optimization.

London onsite FullTime
PythonFine-TuningPyTorch +4 more

Research Engineer, Pretraining

16d ago
Anthropic

Anthropic

Anthropic is seeking a Research Engineer to join its Pretraining team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, aiming to build safe, steerable, and trustworthy AI systems. The mission is to ensure that transformative AI systems are aligned with human interests and are beneficial for society. The team is dedicated to pushing the boundaries of AI while prioritizing safety and ethics.

London, UK onsite
AnthropicKubernetesPython +4 more

[Expression of Interest] Research Engineer / Scientist, Alignment - London

16d ago
Anthropic

Anthropic

Anthropic is building reliable, interpretable, and steerable AI systems to ensure AI is safe and beneficial for society. As a Research Engineer on the Alignment Science team in London, you will design and execute machine learning experiments to understand and steer the behavior of advanced AI systems. You will focus on AI safety, particularly risks from future human-level AI systems, collaborating with teams like Interpretability and Frontier Red Team. The role involves exploratory research in areas such as AI Control and Alignment Stress-testing, aiming to make AI helpful, honest, and harmless.

London, UK onsite
AnthropicKubernetesPython +3 more

Software Engineer - Multi-Agent UAV Autonomy (Defense)

18d ago
A

Applied Intuition

Applied Intuition is seeking a Software Engineer to join our Korea office and build the company's multi-agent autonomy development tools. This on-site role is crucial for developing capabilities that defense customers rely on to design, train, and validate coordinated Unmanned Aerial Vehicle (UAV) behavior. You will collaborate closely with the U.S. product engineering team and directly engage with Korean defense customers to translate their operational requirements into tangible product features. The position involves contributing to system-level design and architecture, ensuring scalability, reliability, and performance, while also guiding customers on optimizing their use of our tools for developing their own multi-agent unmanned systems.

$60k - $300k

Seoul onsite FullTime
PythonGoC# +1 more

Product Lead, Foundational Models and Post-Training

18d ago
Abridge

Abridge

Abridge is seeking a Product Lead to drive the strategy for its foundational models and post-training efforts. This role sits at the intersection of model science, platform strategy, and clinical product delivery, focusing on leveraging Abridge's unique corpus of clinical conversations and clinician feedback to enhance model capabilities. You will own the product strategy, connecting research initiatives to tangible product outcomes, and determining where in-house models offer a competitive advantage. The ideal candidate will have a deep understanding of the modern model-development lifecycle and the technical judgment to collaborate effectively with scientists and engineers, while always anchoring the work in clinician value, patient safety, and business impact.

SF Office hybrid FullTime
Fine-TuningReinforcement Learning

Member of Technical Staff, Data Flywheel

19d ago
Reflection ai

Reflection ai

Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower users to control their intelligence and shape the future of AI. The Data Flywheel team plays a crucial role in bridging the gap between theoretical benchmark performance and practical, real-world utility. We achieve this by identifying and developing the signals, data, and feedback mechanisms that transform model usage into rigorous evaluations, targeted training data, and measurable advancements for future AI generations. This is a hands-on technical position situated at the nexus of research and deployment, where you will guide ambiguous model behaviors from initial observation through measurement, intervention, and validated improvement. You will collaborate across evaluation, human and synthetic data, infrastructure, post-training, and live deployment aspects, working closely with internal researchers and engineers, as well as external customers, partners, vendors, and the open-source community.

New York, NY onsite FullTime
Reinforcement Learning

Member of Technical Staff (AI Researcher)

19d ago
P

Perplexity AI

Perplexity is seeking top-tier AI Research Scientists and Engineers to advance our AI products and capabilities, focusing on building the future of AI-powered search and agent experiences. You will contribute to SOTA experiences that handle hundreds of millions of queries and continue to scale rapidly. Depending on your interests and expertise, you can join one of three specialized teams: the Core Research Team focusing on foundational models, the Agent Products Team fine-tuning models for agent and product experiences, or the Comet Agent Team dedicated to developing and enhancing the Comet Agent product.

San Francisco onsite FullTime
PythonC#Fine-Tuning +3 more

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.