Reinforcement Learning Jobs
93 open roles mentioning Reinforcement Learning
AI Engineer
Baseten
Baseten is seeking an AI Engineer to join their Training Product team. This role involves building AI-driven product features for customers training frontier models and enhancing Baseten's internal AI capabilities by transforming manual workflows into agentic ones. You will collaborate directly with research engineers, taking concepts from internal research to customer-ready products. This is a hands-on position offering significant autonomy, where you will identify key problems, develop reliable AI systems with appropriate harnesses and guardrails, and take ownership of the outcomes. The ideal candidate has a proven track record of shipping agents and is eager to make that their primary focus.
Member of Technical Staff - RL Environments
Cohere
Cohere is seeking a Member of Technical Staff to focus on Reinforcement Learning (RL) Environments. In this role, you will be instrumental in building and refining the AI agents that power enterprise solutions. This involves creating realistic work environments, populating them with challenging tasks, and defining clear reward mechanisms. Your work will directly contribute to the evaluation and training of these agents, with customer feedback closing the loop for continuous improvement. You will collaborate across teams to identify performance gaps and enhance both agents and environments, ensuring the delivery of cutting-edge AI capabilities to our clients.
Research Engineer, LangSmith Engine
Langchain
LangChain is seeking an experienced Research Engineer to join the LangSmith Engine team. This role focuses on enhancing the capabilities and efficiency of a proactive agent engineer that analyzes production traces, identifies failures, and implements fixes to prevent recurrence. You will study agent failures, build benchmarks, run experiments to improve performance, and translate successful ideas into production. This involves optimizing prompting, agent harnesses, model selection, fine-tuning, and post-training custom models, with a strong emphasis on measurable improvements to the overall agent. The role also requires an understanding of production engineering and system-level trade-offs, including cost, latency, reliability, and scalability, working closely with production engineers to ensure reliable real-world performance.
Software Engineer
Applied Intuition
Applied Intuition is a rapidly growing company focused on powering the future of physical AI. We are building the digital infrastructure necessary to bring intelligence to every moving machine across various industries like automotive, defense, and construction. Our solutions are trusted by leading global automakers and military organizations. We are seeking a Software Engineer to join our team and contribute to the design, implementation, and deployment of software and machine learning components, with a particular focus on behavior prediction and environmental interactions. This role involves building scalable software for inference and decision-making in dynamic environments, optimizing data pipelines, and developing robust testing frameworks to ensure system performance and model accuracy.
$60k - $300k
Staff AI Scientist
fiddler-ai
Fiddler is building trust into AI, especially with the rise of Generative AI and Agents. Our platform helps organizations deploy trustworthy and transparent AI solutions by monitoring, evaluating, securing, analyzing, and improving AI applications. We partner with AI-first organizations to establish responsible AI practices, enabling engineering teams and business stakeholders to understand AI outcomes. Joining Fiddler means making an impact by ensuring AI applications at production scale have operational transparency and security. This is an opportunity to be a trailblazer in the rapidly innovating AI and ML industry, contributing to AI Observability.
$220k - $260k
Machine Learning Engineer, API Multicloud
OpenAI
OpenAI is seeking Machine Learning Engineers to join its API Multicloud team, focusing on extending OpenAI's API platform into strategic cloud environments, starting with AWS. This role involves building and improving AI systems that help strategic partners adapt OpenAI models for cloud-native use cases. You will operate at the intersection of production ML systems, developer platforms, model behavior, and large-scale infrastructure, spanning post-training workflows, evaluation, data pipelines, and API/infrastructure integration. The ideal candidate will enjoy working with external technical partners, diagnosing issues, and translating learnings into platform improvements, collaborating closely with Research, Applied, Safety Systems, and infrastructure teams.
Member of Technical Staff - Post-Training
Reflection ai
Reflection is a research lab dedicated to making intelligence open and accessible for everyone to use, customize, and build on. We are building open models that empower individuals to control their intelligence and shape the future of AI. As a Member of Technical Staff - Post-Training, you will play a crucial role in transforming powerful pre-trained models into aligned and general agents. This position involves driving research and engineering initiatives at the forefront of post-training techniques, from data curation to large-scale optimization, and contributing to the advancement of large model reasoning and instruction following capabilities.
Strategic Projects Lead, Deeptune
Mercor
Mercor is a leading AI data company building the layer between human expertise and frontier models. The Deeptune lab within Mercor builds training gyms for AI agents using reinforcement learning. We are seeking a Strategic Projects Lead to own and scale the data operations that power frontier AI training. This is a high-ownership role for an individual with strong agency, comfortable making decisions with incomplete information, and capable of building processes that hold up under pressure. The role involves owning projects end-to-end, managing distributed teams, forecasting timelines, meeting throughput targets, and maintaining a high bar for quality. Your work will directly shape frontier models and how they are deployed.
$120k - $150k
Member of Technical Staff, Deeptune Environments
Mercor
Mercor is seeking a Member of Technical Staff for their Deeptune Environments team. This role involves building the core systems that power high-fidelity simulation environments for training AI agents through reinforcement learning. You will develop the APIs and tool interfaces agents interact with, the grading layer for evaluating agent success, pipelines for transforming human demonstrations into training environments, and the orchestration and sandboxing infrastructure for large-scale operation. You will collaborate closely with RL researchers, translating their ideas into functional, scalable systems, and will be responsible for the end-to-end development of these critical components.
Researcher, Alignment CoT Monitorability
OpenAI
OpenAI is seeking a Researcher focused on Chain-of-Thought (CoT) Monitorability to join their Alignment team. This role involves studying and improving the monitorability of frontier reasoning models, particularly their chain-of-thought processes, to enable scalable oversight. The team's work is crucial for ensuring AI safety and trustworthiness as models become more capable. You will design and execute experiments to understand how various training interventions impact monitorability, develop evaluation methods to measure it, and translate research findings into practical recommendations for monitoring and training. This position is ideal for someone who can bridge ambiguous research questions with concrete experimental designs, moving from hypothesis formulation to analysis and actionable insights.
Machine Learning Engineer, Assistant Quality
Glean
Glean is seeking a Machine Learning Engineer to enhance the quality of its AI Assistant and autonomous agents. This role focuses on production machine learning, LLM-powered systems, and product engineering, with an emphasis on building, evaluating, and refining assistant experiences to be useful, reliable, and grounded in enterprise workflows. The ideal candidate is driven by shipping production systems rather than pure research and is eager to improve Glean's assistant through stronger signals, tighter feedback loops, and better end-to-end execution quality.
$180k - $205k
Research Engineer / Research Scientist - Personal AGI, Personalization
OpenAI
The Personalization-Memory team at OpenAI is focused on developing agents that learn from past interactions to become more helpful and efficient. This role involves researching and developing improvements to memory usage and personalization in frontier models, working on reinforcement learning, dataset creation, evaluations, and post-training methods. The team collaborates with research and product teams across the company to realize the vision of a personalized ChatGPT. We are seeking individuals passionate about product-driven research with a background in frontier model post-training and the ability to iterate quickly.
Research Engineer / Research Scientist - Personal AGI, Memory
OpenAI
The Personalization-Memory team at OpenAI is developing agents capable of learning from past interactions to improve helpfulness and efficiency over time. This role focuses on building general-purpose memory and personalization capabilities that can be applied across various agentic products, including ChatGPT. You will be instrumental in advancing memory architecture, post-training techniques, and creating long-horizon tasks for training and evaluation. We are seeking individuals with a strong background in reinforcement learning research who can translate scientific rigor into tangible product impact.
ML Research Intern
Modal
AI needs a new infrastructure layer, and we're building it at Modal. Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno, who rely on Modal for instant GPU access, sub-second container starts, and native storage for tasks like low-latency inference, model fine-tuning, and accessing production-ready sandboxes at scale. We are seeking PhD research interns with strong research experience in reinforcement learning, machine learning, and foundation models, including large language and multimodal models, to join our research team. This internship is ideal for candidates interested in improving existing methods and developing new techniques for large-scale model training, optimization, and inference, extending models to long-context and long-horizon tasks, and enhancing inference-time efficiency, reliability, and robustness in high-stakes real-world deployments.
Research Scientist / Engineer – Reinforcement Learning Infrastructure
lumalabs
Luma is seeking a Research Scientist / Engineer to build the systems that enable reinforcement learning (RL) at frontier scale. This role involves coupling policy optimization with large fleets of inference workers, agentic environments, and reward/verification systems to transform model behavior into learning signals. RL is crucial for Luma's models to evolve from capable to useful. Operating RL at scale is a complex systems challenge, encompassing training, rollout generation, environment execution, and reward computation across thousands of GPUs, demanding speed, stability, and correctness. This position is ideal for someone with hands-on experience in post-training LLMs with RL, building environments and verifiers, and debugging large-scale asynchronous rollout pipelines.
$30k - $60k
Applied Machine Learning Engineer, Singapore
fireworks ai
As an Applied Machine Learning Engineer, you will serve as a vital bridge between cutting-edge AI research and practical, real-world applications. Your work will focus on developing, fine-tuning, and operationalizing machine learning models that drive business value and enhance user experiences. This is a hands-on engineering role that combines deep technical expertise with a strong customer focus to deliver scalable AI solutions.
Research Scientist - Multimodal Agent, Consumer Devices
OpenAI
We are seeking a Research Engineer / Scientist to join our applied research team focused on developing new methods, models, and evaluation frameworks for the future of computing. This role will concentrate on building the learning and evaluation foundations that enable AI models to become more context-aware, adaptive, and useful over time. You will tackle challenges such as reward modeling, preference learning, long-horizon evaluation, and policy improvement for systems that require high-quality behavioral decisions in realistic user settings. The work is deeply product-grounded, aiming for improved model behavior in real-world use rather than just benchmark performance. The ideal candidate is passionate about advancing AI systems beyond simple assistant interactions towards those that improve through feedback, learn from richer signals, and are trained against meaningful notions of user value.
Research Engineer, Developer Experience, Tinker
thinkingmachines
Thinking Machines Lab is seeking a Research Engineer focused on developer experience to build and enhance their Tinker platform. This role involves working hands-on with users to understand their challenges and translate them into product improvements. You will be responsible for creating and updating documentation, adding library features, prototyping integrations, and ensuring users can smoothly customize frontier AI models. This position acts as a crucial link between Tinker users and the internal research and infrastructure teams, surfacing user patterns to inform product and infrastructure priorities and sharing learnings through various channels.
$350k - $475k
Member of Engineering (Inference Infrastructure)
poolside
Poolside is building a world where AI drives economically valuable work and scientific progress, aiming to accelerate software development with agentic systems, coding assistants, and frontier models. This role is on the compute team, focusing on optimizing GPU workload scheduling and inference serving. You will partner with the inference team to enhance throughput and latency for evaluations and reinforcement learning, collaborate with the scalability team to stabilize large-scale fault-tolerant training, and work closely with the infrastructure team to ensure GPU nodes are healthy and fully utilized. Your work will directly impact research velocity and contribute to the company's mission of building frontier models.
Member of Technical Staff - Research, Post-Training
Modal
We are building a platform that covers the entire lifecycle of Large Language Models (LLMs), from training to deployment and production observation. Our existing infrastructure supports multi-node training, elastic inference, sandboxes, and distributed volumes, with full control over the underlying systems. We are seeking individuals with deep research expertise in post-training techniques to complement our existing systems and product development efforts. This role is ideal for candidates passionate about improving current methods and creating novel techniques for large-scale model training, optimization, and inference. You will focus on extending models to handle long-context and long-horizon tasks, and enhancing inference-time efficiency, reliability, and robustness for critical real-world applications.