Reinforcement Learning Jobs
93 open roles mentioning Reinforcement Learning
Engineering Manager (TLM, Agents)
Perplexity AI
Perplexity is looking for a Tech Lead Manager (TLM) to lead and grow its Agents engineering team. This team comprises AI/ML, backend, and full-stack engineers focused on building agentic experiences within the Comet ecosystem. The vision is to create AI agents that can fulfill user intentions through open-ended interactions with the world. As the Agents TLM, you will combine AI expertise, product intuition, and engineering management skills to push the boundaries of what AI agents can achieve for millions of users. You will lead, develop, and support a team in tackling significant challenges in AI, such as designing agents for digital navigation, training models for complex decision-making, ensuring excellent user experiences across various platforms, developing secure agentic capabilities, and optimizing agent-environment interactions.
Member of Technical Staff (AI Software Engineer, Agents)
Perplexity AI
Perplexity is seeking energetic engineers to join our highly driven Agents engineering team. The Agents team collaborates to build harnesses and AI systems powering delightful agentic experiences, including Perplexity Computer, the Comet ecosystem, and our Agent API. Our vision is to empower users with agents that faithfully actualize their intent through open-ended interactions with the world. As an engineer on our Agents team, you will bring AI expertise, sharp product intuition, and a tinkerer's mindset to advance the frontier of what agents can accomplish for our millions of devoted users, working across applied research and engineering to solve open problems in AI.
Researcher, Safety Training, National Security
OpenAI
We are seeking a researcher to train and evaluate models for U.S. government use, with a focus on national security applications. You will advance safety post-training and robustness, helping models follow nuanced policies while preserving their usefulness and capabilities. This role involves researching and implementing methods for safety training, reinforcement learning, and adversarial robustness, as well as developing evaluations, identifying model failure modes, and using findings to improve training. You will also collaborate with research, engineering, security, and policy partners to support safe, reliable deployment.
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Perplexity AI
Perplexity is building the foundations for agentic AI, transforming knowledge into action for millions of users. The Agent Capabilities team is at the forefront of AI research and product innovation, responsible for translating frontier AI breakthroughs into reusable product capabilities. This role offers broad ownership at the intersection of AI research, agent systems, platform engineering, and product innovation, allowing you to evaluate emerging models, identify user value, and build reliable, scalable experiences for users and agents.
Vice President, Product Management, Managed AI
crusoe
Crusoe is seeking a Vice President of Product Management for its rapidly scaling Managed AI Services portfolio. Launched in late 2025, this business has already secured significant ARR commitments and is expanding to include Managed Inference, Serverless Fine Tuning, Reinforcement Learning, and Agent Sandbox. As VP, you will be responsible for the vision, strategy, execution, and roadmap of this business, reporting to the SVP of Product Management. You will lead the portfolio as a business unit, fostering strong cross-functional partnerships and aligning with Crusoe's vertically integrated cloud strategy.
$345k - $385k
Software Engineer, AI for Chip Design
OpenAI
We are seeking a Software Engineer to develop the research infrastructure and tooling that will enable OpenAI models to design silicon. This role involves transforming chip-design workflows into robust environments for reinforcement learning and evaluation, and empowering researchers to conduct experiments and iterate on new ideas efficiently. You will bridge software engineering, tool integration, and open research challenges, leveraging strong coding fundamentals, sound technical judgment, and the ability to work independently. While prior chip-design experience is beneficial, domain knowledge can be acquired alongside our hardware specialists.
Software Engineer - Voice Model
xAI
SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. The Grok Voice Model team is building the world's best voice AI, delivering smooth, natural, low-latency spoken interactions that are expressive, multilingual, and reliable across devices and real-time scenarios. The team owns the full training pipeline, from massive data curation and premium audio processing to frontier speech-language pre-training and intensive post-training to push quality, speed, and stability to the limit. The goal is to make talking to AI feel like conversing with a charming, kind, and knowledgeable person, and exceptionally smart, execution-oriented engineers are sought to achieve this.
$150k - $450k
Member of Technical Staff - Sandbox Service
xAI
The Sandbox service team at SpaceXAI builds and maintains a secure, scalable system that gives our models safe, controlled access to computational environments. This infrastructure powers critical workloads across training and product, enabling models to run code, build software, interact with tools, and even control applications with user interfaces. We provision containers and virtual machines on large-scale clusters, granting models interactive control over these remote environments. Our work spans the full stack: from orchestrating massive jobs and resource scheduling at the cluster level, to fine-tuning filesystem performance on nodes. The Sandbox service enables Grok to safely run and test code in real-time for user queries, and supports reinforcement learning in training, where models interactively explore tools ranging from compilers to productivity apps.
£107k - £262k
Member of Technical Staff - RL Training Framework
xAI
SpaceXAI is seeking an engineer to join the RL infrastructure team to help develop our RL training framework. The team is small, highly motivated, and focused on engineering excellence, operating with a flat organizational structure where all employees are expected to be hands-on and contribute directly to the company's mission. This role requires strong initiative, curiosity, work ethic, prioritization, and communication skills.
$180k - $440k
Anthropic Fellows Program, ML Systems & Reinforcement Learning
Anthropic
Anthropic is seeking talented individuals for its Fellows Program, focusing on AI research and engineering to ensure AI systems are reliable, interpretable, and steerable. This program provides funding and mentorship to promising technical talent, regardless of prior experience, to work on empirical projects aligned with Anthropic's research priorities. The goal is to produce public outputs, such as paper submissions, contributing to the development of safe and beneficial AI for society. Fellows will engage in full-time research for four months, with direct mentorship from Anthropic researchers and access to a shared workspace in either Berkeley, California, or London, UK. The program also offers connections to the broader AI safety and security research community.
$25k - $50k
Anthropic Fellows Program, AI Safety & Security
Anthropic
Anthropic is seeking talented individuals for its Fellows Program, focusing on AI Safety. This program aims to foster AI research and engineering talent by providing funding and mentorship to promising individuals, regardless of prior experience. Fellows will engage in empirical projects using external infrastructure and open-source models, with the goal of producing public outputs like research papers. The program offers a structured 4-month full-time research period, direct mentorship from Anthropic researchers, access to shared workspaces in Berkeley or London, and connections to the AI safety community. The program is designed to encourage diverse perspectives and applications, even from those who may not meet every single qualification.
$25k - $50k
Research Engineer, Safety
decagon
Decagon is seeking a Research Engineer focused on Safety to ensure the reliability and controllability of their AI agents from evaluation through production. This role involves identifying real-world failure modes and developing the necessary models, evaluations, and safeguards to prevent them. The ideal candidate is a strong engineer passionate about advancing applied AI safety in production, with the autonomy to own their work end-to-end, ship impactful improvements, and make high-stakes technical decisions.
$200k - $400k
Senior Research Engineer, Safety
decagon
Decagon is seeking a Senior Research Engineer focused on Safety to join their conversational AI platform team. This role is crucial for ensuring the reliability and controllability of Decagon's AI agents from evaluation through production. You will be responsible for identifying real-world failure modes and developing the necessary models, evaluations, and safeguards to prevent them. This is an opportunity for strong engineers to advance applied AI safety in a production environment, owning their work end-to-end and making high-impact technical decisions.
$200k - $400k
Applied Machine Learning Engineer, EMEA
fireworks ai
Fireworks is seeking an Applied Machine Learning Engineer for the EMEA region. In this role, you will be the technical owner of customer engagements, embedding within client teams to understand their specific needs and challenges. You will be responsible for the entire lifecycle of a customer's deployment on the Fireworks platform, from initial scoping and model selection to ensuring production readiness, performance, and cost-efficiency. This position emphasizes first principles thinking and requires a deep understanding of software engineering, machine learning techniques, and infrastructure optimization.
Research Engineer, Pretraining
Anthropic
Anthropic is seeking a Research Engineer to join its Pretraining team, focusing on developing the next generation of large language models. This role operates at the intersection of cutting-edge research and practical engineering, aiming to build safe, steerable, and trustworthy AI systems. The mission is to ensure that transformative AI systems are aligned with human interests and are beneficial for society. The team is dedicated to pushing the boundaries of AI while prioritizing safety and ethics.
[Expression of Interest] Research Engineer / Scientist, Alignment - London
Anthropic
Anthropic is building reliable, interpretable, and steerable AI systems to ensure AI is safe and beneficial for society. As a Research Engineer on the Alignment Science team in London, you will design and execute machine learning experiments to understand and steer the behavior of advanced AI systems. You will focus on AI safety, particularly risks from future human-level AI systems, collaborating with teams like Interpretability and Frontier Red Team. The role involves exploratory research in areas such as AI Control and Alignment Stress-testing, aiming to make AI helpful, honest, and harmless.
Software Engineer - Multi-Agent UAV Autonomy (Defense)
Applied Intuition
Applied Intuition is seeking a Software Engineer to join our Korea office and build the company's multi-agent autonomy development tools. This on-site role is crucial for developing capabilities that defense customers rely on to design, train, and validate coordinated Unmanned Aerial Vehicle (UAV) behavior. You will collaborate closely with the U.S. product engineering team and directly engage with Korean defense customers to translate their operational requirements into tangible product features. The position involves contributing to system-level design and architecture, ensuring scalability, reliability, and performance, while also guiding customers on optimizing their use of our tools for developing their own multi-agent unmanned systems.
$60k - $300k
Product Lead, Foundational Models and Post-Training
Abridge
Abridge is seeking a Product Lead to drive the strategy for its foundational models and post-training efforts. This role sits at the intersection of model science, platform strategy, and clinical product delivery, focusing on leveraging Abridge's unique corpus of clinical conversations and clinician feedback to enhance model capabilities. You will own the product strategy, connecting research initiatives to tangible product outcomes, and determining where in-house models offer a competitive advantage. The ideal candidate will have a deep understanding of the modern model-development lifecycle and the technical judgment to collaborate effectively with scientists and engineers, while always anchoring the work in clinician value, patient safety, and business impact.
Member of Technical Staff, Data Flywheel
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible. We build open models that empower users to control their intelligence and shape the future of AI. The Data Flywheel team plays a crucial role in bridging the gap between theoretical benchmark performance and practical, real-world utility. We achieve this by identifying and developing the signals, data, and feedback mechanisms that transform model usage into rigorous evaluations, targeted training data, and measurable advancements for future AI generations. This is a hands-on technical position situated at the nexus of research and deployment, where you will guide ambiguous model behaviors from initial observation through measurement, intervention, and validated improvement. You will collaborate across evaluation, human and synthetic data, infrastructure, post-training, and live deployment aspects, working closely with internal researchers and engineers, as well as external customers, partners, vendors, and the open-source community.
Member of Technical Staff (AI Researcher)
Perplexity AI
Perplexity is seeking top-tier AI Research Scientists and Engineers to advance our AI products and capabilities, focusing on building the future of AI-powered search and agent experiences. You will contribute to SOTA experiences that handle hundreds of millions of queries and continue to scale rapidly. Depending on your interests and expertise, you can join one of three specialized teams: the Core Research Team focusing on foundational models, the Agent Products Team fine-tuning models for agent and product experiences, or the Comet Agent Team dedicated to developing and enhancing the Comet Agent product.