Research Engineer / Performance Engineer, RL Distributed Systems
San Francisco, CA | New York City, NY | Seattle, WA
Posted 20h ago
Job Location
San Francisco, CA | New York City, NY | Seattle, WA
Tech Stack
Remote Work Policy
On-site
Categories
AI Research Engineer
About the job
Reinforcement learning is crucial for training AI models like Claude to reason, write code, and act autonomously. At scale, this process demands a highly sophisticated distributed system. This role focuses on optimizing and maintaining this complex system, which involves concurrent training, sampling, and environment execution across a large fleet of accelerators. The system must be resilient to hardware failures, shifting loads, and evolving research needs. As a Research Engineer on the Distributed Systems team, you will identify and address performance bottlenecks within this system, potentially working on scheduling, data movement, environment execution, storage, networking, fault tolerance, autoscaling, or observability. We seek adaptable engineers who can reason from first principles about unfamiliar systems and prioritize the most impactful problems.
Responsibilities
- Design, build, and operate distributed systems for large-scale RL training, sampling, and environment execution.
- Identify and resolve system limitations in areas like scheduling, data movement, storage, networking, or coordination.
- Implement fault tolerance mechanisms for failure detection, isolation, and recovery to ensure continuous job progress.
- Develop resource management and autoscaling solutions to dynamically allocate compute resources based on demand.
- Build observability tools to monitor run performance, identify slowdowns, and diagnose unexpected behavior.
- Create automation for detecting and resolving common issues, and design safe operational interfaces for engineers and tools.
- Collaborate with researchers and performance engineers to ensure system changes maintain training correctness and avoid introducing nondeterminism.
- Eliminate failure sources through incident analysis, testing, and redesign, and document system designs clearly.
Requirements
- Strong software engineering skills in Python and at least one systems language (e.g., Rust, C++, Go).
- Experience designing, building, and operating large-scale distributed systems in production.
- Deep understanding of distributed systems fundamentals (consistency, coordination, consensus, failure modes, recovery).
- Ability to quantitatively analyze throughput, latency, and resource costs across compute, memory, storage, and network.
- Experience debugging complex, multi-host failures, including those difficult to reproduce locally.
- Strong written communication skills for design documents and incident reports.