Research Scientist / Engineer – Reinforcement Learning Infrastructure
$30k - $60k • Remote • Redwood City, CA • FullTime
Posted 1mo ago
Job Location
Redwood City, CA
Tech Stack
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Luma is seeking a Research Scientist / Engineer to build the systems that enable reinforcement learning (RL) at frontier scale. This role involves coupling policy optimization with large fleets of inference workers, agentic environments, and reward/verification systems to transform model behavior into learning signals. RL is crucial for Luma's models to evolve from capable to useful. Operating RL at scale is a complex systems challenge, encompassing training, rollout generation, environment execution, and reward computation across thousands of GPUs, demanding speed, stability, and correctness. This position is ideal for someone with hands-on experience in post-training LLMs with RL, building environments and verifiers, and debugging large-scale asynchronous rollout pipelines.
Responsibilities
- Design, build, and scale distributed RL post-training systems, orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs.
- Build high-throughput rollout generation, integrating inference engines, weight synchronization, and asynchronous/off-policy schemes.
- Design RL environments for agentic, multi-step tasks, ensuring reproducibility and scalability.
- Build reward infrastructure, including verifiable rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking.
- Develop evaluation, monitoring, and debugging tooling for stable large RL runs.
- Advance training efficiency and stability, and implement new post-training ideas with researchers.
Requirements
- Hands-on experience post-training LLMs with RL (PPO/GRPO-family, RLHF, RLVR) at meaningful scale.
- Extensive distributed PyTorch training and parallelism (FSDP, Tensor/Pipeline/Expert Parallel) for foundation models.
- Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents, including sandboxed execution and multi-turn tool use.
- Deep familiarity with RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang).
- Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference workloads.