Research Engineer, Pretraining Scaling

San Francisco, CA

Posted 16d ago

Job Location

San Francisco, CA

Tech Stack

Remote Work Policy

On-site

Categories

AI Research Engineer

About the job

Anthropic's ML Performance and Scaling team is responsible for training the company's production pretrained models, a critical function that directly shapes the future of Anthropic and its mission to build safe, beneficial AI systems. As a Research Engineer on this team, you will ensure that frontier models train reliably, efficiently, and at scale. This role bridges the gap between research and engineering, involving work across the entire production training stack, including performance optimization, hardware debugging, experimental design, and launch coordination. During model launches, the team operates in close collaboration, addressing production issues that require immediate attention.

Responsibilities

  • Own critical aspects of the production pretraining pipeline, including model operations, performance optimization, observability, and reliability.
  • Debug and resolve complex issues across the full stack, from hardware and networking to training dynamics and evaluation infrastructure.
  • Design and execute experiments to enhance training efficiency, reduce step time, increase uptime, and improve model performance.
  • Respond to on-call incidents during model launches, quickly diagnosing problems and coordinating solutions across teams.
  • Build and maintain production logging, monitoring dashboards, and evaluation infrastructure.
  • Add new capabilities to the training codebase, such as long context support or novel architectures.
  • Collaborate with teammates across different locations and with other specialized teams.
  • Contribute to the team's institutional knowledge through documentation of systems, debugging approaches, and lessons learned.

Requirements

  • Hands-on experience training large language models, or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems.
  • A balance of research and engineering work, with a roughly 50/50 split.
  • Willingness to be on-call for production systems, work extended hours during launches, and solve hard problems under pressure.
  • Ability to thrive when working on the most impactful tasks, adapting priorities as needed.
  • Excellence in debugging complex, ambiguous problems across multiple layers of the stack.
  • Clear communication and effective collaboration skills, especially across time zones and during high-stress incidents.
  • Passion for the work and a desire to refine craft as a research engineer.
  • Commitment to the societal impacts of AI and responsible scaling.

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.