RLHF Jobs
68 open roles mentioning RLHF
Agent Post-Training, Artifacts Research
OpenAI
The Agent Post-Training team is responsible for developing the frontier agents that OpenAI ships to the world, including models for Codex, ChatGPT, and the API. This role focuses on training models to produce polished, useful work products such as documents, spreadsheets, and reports, transforming vague user goals into finished artifacts with strong structure, visual taste, and correctness. The position involves owning improvements across the post-training stack, including RL, data pipelines, graders, reward signals, and evaluations. You will collaborate with researchers, engineers, product teams, and safety partners to shape major model runs and ship improvements into widely used products. This is a high-agency role for individuals eager to directly impact frontier models.
Agent Post-Training, Computer Use Research
OpenAI
The Agent Post-Training team is responsible for developing the frontier agents that OpenAI ships to the world. This role focuses on training models to operate computers, enabling them to navigate browsers and desktops, utilize tools, reason through complex workflows, and collaborate with users and other agents. The work involves a blend of frontier model training, product behavior, evaluation, and systems engineering, directly influencing the computer-use capabilities of OpenAI's next-generation agents. You will collaborate with researchers, engineers, product teams, and safety partners to define model training runs, measure outcomes, and ship improvements to products used by millions.
Agent Post-Training, Connectors Research
OpenAI
The Agent Post-Training team is responsible for developing the frontier agents that OpenAI ships to the world, including models for Codex, ChatGPT, and the API. This role focuses on teaching models to interface with professional software using code, APIs, and tools. You will enable agents to operate across applications like Slack, Google Workspace, GitHub, and Salesforce, taking actions within a user's digital context to find information, update systems, coordinate work, and complete multi-step workflows. The position involves training models to leverage productivity and enterprise software, turning connected tools into a powerful action surface for agents. You will collaborate with researchers, engineers, product teams, and safety partners to define model capabilities, measure performance, and ship improvements into products.
Agent Post-Training, Context Research
OpenAI
The Agent Post-Training team is responsible for developing the frontier agents that OpenAI releases. This role focuses on scaling compute spent on context, enabling a new paradigm of model training with a clear product interface for iterative deployment. You will collaborate with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to define model run content, measure outcomes, and integrate improvements into widely used products. This is a high-agency position for individuals eager to directly influence frontier models.
Agent Post-Training, Frontier Evals and Environments Research
OpenAI
The Agent Post-Training team at OpenAI is responsible for developing the frontier agents that are shipped to the world. This role focuses on building "north star" model environments to drive progress towards safe AGI/ASI, directly guiding ambitious research programs. You will collaborate with researchers, engineers, product, infrastructure, and safety teams to define model run objectives, measure outcomes, and integrate improvements into widely used products. This is a high-agency position for individuals eager to directly influence cutting-edge AI models and witness their rapid advancement.
Agent Post-Training, Personality
OpenAI
The Agent Post-training Personality team at OpenAI is responsible for shaping the collaborative capabilities of frontier AI agents. This role focuses on understanding and enhancing agent "personality," which encompasses thoughtfulness, clarity, perceptiveness, appropriate proactivity, and ease of interaction. The goal is to create agents that are not only effective but also exceptional collaborators, understanding user intent, communicating with good judgment, adapting to context, and taking initiative when needed. This position involves a blend of behavioral research, product thinking, and communication expertise, requiring close collaboration with product teams, human experts, and researchers across the organization to ensure these improvements are integrated into widely used AI models.
Agent Post-Training, API & Power Users
OpenAI
The Agent Post-Training team is responsible for developing the frontier agents that OpenAI ships to the world, including models for Codex, ChatGPT, and the API. This role focuses on improving the capabilities, reliability, and product fit of OpenAI's agentic models specifically for power users and API developers. You will tackle ambiguous model behavior problems, translating them into concrete progress by enhancing aspects like tool use, planning, instruction following, and error recovery. The position involves close collaboration across research, engineering, data, evals, and product teams to refine model behavior for real-world workflows and API integrations, ultimately shaping the next generation of AI agents.
Agent Post-Training Research
OpenAI
OpenAI is seeking a highly motivated individual to join the Agent Post-Training team, responsible for developing frontier agents that can operate computers, collaborate with humans and other agents, and expand human potential. This role is intentionally broad, requiring individuals who can tackle ambiguous capability problems across research, engineering, data, evals, and product. You will work on models that write and debug code, use tools, call functions, operate computers, and complete valuable work for users. The position involves collaborating with researchers, engineers, product teams, and safety partners to define model capabilities, measure performance, and ship improvements into widely used products. This is a high-agency role for those who want their work to directly impact frontier models.
Research Intern RL & Post-Training Systems, Turbo (Fall 2026)
Together AI
The Turbo Research team focuses on making post-training and reinforcement learning for large language models efficient, scalable, and reliable. This work intersects RL algorithms, inference systems, and large-scale experimentation, where inference costs significantly impact training efficiency and the practicality of learning algorithms. As a research intern, you will investigate RL and post-training methods whose performance and scalability are closely tied to inference behavior, co-designing algorithms and systems. Projects aim to enable new experimental regimes, including larger models, longer rollouts, and more complex evaluations, by re-evaluating the interaction between inference, scheduling, and training.
AI research scientist
Writer
AI research at WRITER focuses on building the scientific foundation for ambitious enterprise AI deployments. As a staff AI research scientist, you will drive a high-impact research agenda centered on large language models, agentic reasoning, and system-level capabilities essential for enterprise-scale AI. This role offers a unique opportunity to advance the field while directly contributing to products used by hundreds of thousands daily. You will work on post-training, planning, multi-step reasoning, and agentic workflows, directly shaping the future of enterprise AI performance and scalability. The role provides resources, infrastructure, and cross-functional support to pursue and implement ambitious ideas rapidly.
Forward Deployed Engineer (Inference & Post-Training)
Together AI
As a Forward Deployed Engineer (FDE) focused on Inference & Post-Training, you will be a hands-on technical partner to strategic customers, assisting production AI teams with leveraging high-quality models and performing inference at scale. You will act as a deep-domain specialist in inference optimization, fine-tuning pipelines, and production deployment, partnering with Solutions Architects. FDEs add significant value by ensuring complex Proofs of Concept (POCs) are met, facilitating platform adoption, and guiding tailored optimization efforts, directly impacting customer success and company growth.
$270k - $300k
AI Researcher, Core ML (Turbo)
Together AI
The Turbo team operates at the intersection of efficient inference (algorithms, architectures, engines) and post-training/RL systems. We are responsible for building and managing the systems that power Together's API, focusing on high-performance inference and RL/post-training engines capable of operating at production scale. Our core mission is to advance the frontiers of efficient inference and RL-driven training, aiming to make models significantly faster and more cost-effective to run, while simultaneously enhancing their capabilities through RL-based post-training methods. This role involves working across the entire stack, from RL algorithms and training engines to kernels and serving systems, to develop and refine state-of-the-art models using RL pipelines. We value individuals with deep expertise in one area and a strong willingness to collaborate and grow across others.
$200k - $280k
Research Engineer, Core ML
Together AI
This research engineering role focuses on translating new Reinforcement Learning (RL) algorithms, scheduling methods, and inference optimizations into production-grade systems that power Together's API. The Core ML team operates at the intersection of efficient inference (algorithms, architectures, engines) and post-training/RL systems, building and maintaining high-performance inference and RL engines at production scale. The goal is to significantly improve model speed, cost-efficiency, and capabilities through RL-based post-training. This position requires a blend of algorithmic understanding and systems engineering, with opportunities to work across the entire stack from RL algorithms and training engines to kernels and serving systems, ultimately driving measurable improvements in latency, throughput, cost, and model quality at scale.
$200k - $280k
Technical Program Manager, Engineering
Scale AI
Scale is at the forefront of the AI revolution, developing data engines and technologies that power the world's leading LLMs. This role focuses on leading critical programs within the Platform and Security Engineering teams, overseeing the design and development of core data storage systems and security initiatives. You will drive company-wide programs, improve processes, and ensure alignment with industry standards, gaining exposure to the cutting edge of AI adoption across various sectors. The work is crucial for making AI models safe, aligned, and useful through human evaluation and reinforcement learning.
$181k - $226k
Software Engineer, Platform
Scale AI
Scale is at the forefront of the AI revolution, building the Generative AI Data Engine and other products that power the world's most advanced LLMs. The Platform Engineering team is foundational to these efforts, responsible for designing and developing shared platforms, architecting core cloud infrastructure, and redefining software development processes. This role offers exposure to the cutting edge of AI development across various sectors, from startups to governments. You will drive the design and implementation of critical platforms, collaborate with cross-functional teams, and proactively improve engineering practices. This is an opportunity to shape the future of AI infrastructure and contribute to some of the most important work in how humanity interacts with AI.
$216k - $270k
Technical Program Manager, Enterprise
Scale AI
As a Technical Program Manager, you will partner with our Frontier Agent Engineering teams on enterprise customer engagements, owning operational execution and delivery of technical work by managing timelines, milestones, risks, and dependencies. You will drive the strategic alignment and end-to-end execution of critical Enterprise initiatives, serving as the core communication backbone between engineering, product, and executive leadership. Operating in a demanding AI environment, you will translate technical complexity into clear execution strategies, proactively mitigate risks, and ensure engineering teams deliver reliable, high-value solutions at scale.
$211k - $264k
Research Scientist, Agent Robustness
Scale AI
Scale Labs is seeking talented researchers to join a new team focused on policy research, bridging the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. This role will tackle fundamental challenges in building AI agents that are safe and aligned with humans, researching agent capabilities, designing evaluation harnesses, building exploits and mitigations for failure modes, and characterizing risks of multi-agent systems. The team collaborates broadly across industry, the public sector, and academia, regularly publishing findings.
$216k - $270k
Research Scientist, AI Controls and Monitoring
Scale AI
Scale Labs is seeking a Research Scientist focused on AI Controls and Monitoring to join a new team dedicated to policy research. This role will bridge the gap between AI research and policymakers, focusing on scientific decisions about AI risks and capabilities. The team tackles challenges in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. You will design methods, systems, and experiments to ensure advanced AI models and agents remain aligned with intended goals, even in high-stakes or adversarial environments. This role involves collaboration across industry, the public sector, and academia, with regular publication of findings.
$216k - $270k
Research Scientist, Safety Post Training
Scale AI
Scale Labs is seeking talented researchers to join a new team focused on policy research, bridging the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. This role will develop and apply post-training methods and interpretability techniques to make frontier AI systems safer and better understood. You will design and run post-training pipelines, develop interpretability-informed evaluations, and collaborate with policymakers, engineers, and other researchers to translate findings into actionable safety standards and best practices.
$216k - $270k
Forward Deployed Product Manager, Enterprise
Scale AI
Scale is seeking a Forward Deployed Product Manager (FDPM) to drive the success of enterprise AI deployments. This role is embedded with customers, focusing on achieving real production outcomes and translating operational realities into actionable product insights. Unlike a traditional roadmap PM or solutions engineer, the FDPM owns product outcomes within a portfolio of enterprise accounts, building trust with senior stakeholders, guiding deployments to production, and identifying product versus execution bottlenecks. The ideal candidate can differentiate between stated customer needs, underlying problems, and optimal platform solutions, having previously shipped products into large organizations.
$206k - $300k