1 open role mentioning MPI
lumalabs
This role focuses on building and scaling the infrastructure that powers reinforcement learning (RL) at a frontier scale. You will be responsible for coupling policy optimization with large fleets of inference workers, agentic environments, and reward/verification systems to generate learning signals from model behavior. This is a full-loop systems problem involving training, rollout generation, environment execution, and reward computation across thousands of GPUs, requiring a focus on speed, stability, and correctness. The ideal candidate has direct experience operating RL at scale, including post-training LLMs with RL, building environments and verifiers, and debugging large-scale asynchronous rollout pipelines.