Member of Technical Staff - Reliability Engineering

Remote San Mateo FullTime

Posted 27d ago

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Fireworks AI is seeking a Member of Technical Staff focused on Reliability Engineering to ensure the dependable operation of their AI platform. This role involves working across cloud infrastructure, AI systems, and product teams to guarantee seamless integration, graceful failure handling, and robust performance under load. You will be instrumental in defining reliability standards, owning the reliability toolchain, and ensuring a positive customer experience by addressing failures that span across systems. The position requires a proactive approach to incident management, a commitment to reducing operational toil through automation, and strong collaboration with various engineering teams.

Responsibilities

  • Define reliability standards including SLOs, error budgets, and production readiness criteria, informed by system instrumentation and telemetry.
  • Own and manage the reliability toolchain, encompassing logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tools.
  • Ensure customer experience is not compromised by addressing cross-system failures and driving fixes through owning teams.
  • Identify and resolve complex failures that occur at the intersection of different systems.
  • Coordinate live production incident management, conduct blameless postmortems, and track follow-up actions to completion.
  • Automate repetitive operational tasks to prevent unsustainable on-call burdens.
  • Partner with cloud infrastructure, inference and training, performance, and product teams to address capacity, risk, failure modes, rollouts, and customer-facing reliability.

Requirements

  • 5+ years of experience with Linux internals, system performance troubleshooting, and networking fundamentals (TCP/IP, HTTP, gRPC).
  • 5+ years of experience in software engineering with Python, Go, C++, or Rust, focusing on production-grade tools and systems code.
  • Experience operating and debugging Kubernetes, Terraform, and Docker in high-throughput production environments.
  • Experience with distributed systems, including high-throughput control planes, microservices, or multi-region setups.
  • Solid understanding of reliability fundamentals such as fault-tolerant design, SLO/SLA management, automated failover, and high-availability architecture.
  • Ability to influence without authority, driving adoption of standards through credibility and useful tooling.
  • Willingness to investigate unfamiliar parts of the technology stack when problems cross system boundaries.
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, or equivalent practical experience.

About fireworks ai

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.