Research Scientist / Engineer – Reinforcement Learning Infrastructure
Remote • Remote, EU • FullTime
Posted 4h ago
Job Location
Remote, EU
Tech Stack
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Machine Learning Engineer
About the job
This role focuses on building and scaling the infrastructure that powers reinforcement learning (RL) at a frontier scale. You will be responsible for coupling policy optimization with large fleets of inference workers, agentic environments, and reward/verification systems to generate learning signals from model behavior. This is a full-loop systems problem involving training, rollout generation, environment execution, and reward computation across thousands of GPUs, requiring a focus on speed, stability, and correctness. The ideal candidate has direct experience operating RL at scale, including post-training LLMs with RL, building environments and verifiers, and debugging large-scale asynchronous rollout pipelines.
Responsibilities
- Design, build, and scale distributed RL post-training systems orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs.
- Build high-throughput rollout generation, integrating inference engines, weight synchronization, and asynchronous/off-policy schemes.
- Design RL environments for agentic, multi-step tasks including sandboxed code execution, tool use, computer use, and multimodal interaction, ensuring reproducibility and scalability.
- Build reward infrastructure encompassing verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking.
- Develop evaluation, monitoring, and debugging tooling to maintain stability of large RL runs.
- Advance training efficiency and stability, and collaborate with researchers to implement new post-training ideas in production.
Requirements
- Hands-on experience post-training LLMs with RL (PPO/GRPO-family, RLHF, RLVR) at meaningful scale.
- Extensive distributed PyTorch training and parallelism (FSDP, Tensor/Pipeline/Expert Parallel) for foundation models.
- Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents, including sandboxed execution and multi-turn tool use.
- Deep familiarity with RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang).
- Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference workloads.