Staff Software Engineer, AI Reliability Engineering

Dublin, IE

Posted 16d ago

Job Location

Dublin, IE

Tech Stack

Remote Work Policy

On-site

Categories

Applied AI Engineer

About the job

AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths, from the SDK through our network, API layers, serving infrastructure, and accelerators. This role involves jumping into the trenches alongside partner teams to make the systems that deliver Claude more robust and resilient, be it during an incident or collaborating on projects. Reliability here is an emergent phenomenon that transcends any single team's boundaries, requiring someone to zoom out and look at the whole picture, offering dynamic, cross-cutting exposure to the systems that matter most.

Responsibilities

  • Develop Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity.
  • Design and implement monitoring and observability systems across the token path.
  • Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers.
  • Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
  • Support the reliability of safeguard model serving, critical for both site reliability and Anthropic's safety commitments.

Requirements

  • Strong distributed systems, infrastructure, or reliability backgrounds.
  • Reliability-minded software engineers and SREs.
  • Curiosity and bravery to jump into unfamiliar systems during an incident and help drive resolution.
  • Holistic thinking about system composition and seams.
  • Ability to build lasting relationships across teams.
  • Ownership over outcomes, even for systems not directly owned.
  • Excellent communication and collaboration skills.
  • Diverse experience in product stacks, scaled databases, or massive distributed systems.
  • Experience as an SRE, Production Engineer, or similar reliability-focused roles on large-scale systems.
  • Experience operating large-scale model serving or training infrastructure (>1000 GPUs).
  • Experience with ML hardware accelerators (GPUs, TPUs, Trainium).
  • Understanding of ML-specific networking optimizations like RDMA and InfiniBand.
  • Expertise in AI-specific observability tools and frameworks.
  • Experience with chaos engineering and systematic resilience testing.
  • Contribution to open-source infrastructure or ML tooling.

About Anthropic

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.