Research Engineer / Performance Engineer, RL Distributed Systems

San Francisco, CA | New York City, NY | Seattle, WA

Posted 20h ago

Job Location

San Francisco, CA | New York City, NY | Seattle, WA

Tech Stack

Remote Work Policy

On-site

Categories

AI Research Engineer

About the job

Reinforcement learning is crucial for training AI models like Claude to reason, write code, and act autonomously. At scale, this process demands a highly sophisticated distributed system. This role focuses on optimizing and maintaining this complex system, which involves concurrent training, sampling, and environment execution across a large fleet of accelerators. The system must be resilient to hardware failures, shifting loads, and evolving research needs. As a Research Engineer on the Distributed Systems team, you will identify and address performance bottlenecks within this system, potentially working on scheduling, data movement, environment execution, storage, networking, fault tolerance, autoscaling, or observability. We seek adaptable engineers who can reason from first principles about unfamiliar systems and prioritize the most impactful problems.

Responsibilities

  • Design, build, and operate distributed systems for large-scale RL training, sampling, and environment execution.
  • Identify and resolve system limitations in areas like scheduling, data movement, storage, networking, or coordination.
  • Implement fault tolerance mechanisms for failure detection, isolation, and recovery to ensure continuous job progress.
  • Develop resource management and autoscaling solutions to dynamically allocate compute resources based on demand.
  • Build observability tools to monitor run performance, identify slowdowns, and diagnose unexpected behavior.
  • Create automation for detecting and resolving common issues, and design safe operational interfaces for engineers and tools.
  • Collaborate with researchers and performance engineers to ensure system changes maintain training correctness and avoid introducing nondeterminism.
  • Eliminate failure sources through incident analysis, testing, and redesign, and document system designs clearly.

Requirements

  • Strong software engineering skills in Python and at least one systems language (e.g., Rust, C++, Go).
  • Experience designing, building, and operating large-scale distributed systems in production.
  • Deep understanding of distributed systems fundamentals (consistency, coordination, consensus, failure modes, recovery).
  • Ability to quantitatively analyze throughput, latency, and resource costs across compute, memory, storage, and network.
  • Experience debugging complex, multi-host failures, including those difficult to reproduce locally.
  • Strong written communication skills for design documents and incident reports.

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.