Staff Site Reliability Engineer

Remote Remote - United States FullTime

Posted 21h ago

Job Location

Remote - United States

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Join our Site Reliability Engineering (SRE) team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust observability solutions, lead incident response, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit.

Responsibilities

  • Architect and implement comprehensive monitoring, logging, and tracing solutions.
  • Define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
  • Lead incident management and response for high-impact incidents.
  • Automate operational tasks and infrastructure as code using tools like Terraform or Pulumi.
  • Performance-tune and optimize large-scale cloud deployments, focusing on Kubernetes, Docker, and GCP.
  • Debug and harden distributed systems, implementing long-term fixes.
  • Review feature and system designs for reliability, scalability, security, and operational integrity.
  • Educate and mentor the broader engineering team on reliability best practices.
  • Write high-quality, well-tested code in Python or Go for internal tools and integrations.

Requirements

  • 8-10 years of experience in Site Reliability Engineering or similar roles.
  • Strong programming skills in Python or Go, with experience writing high-quality, well-tested code.
  • Deep understanding of distributed systems, including designing, building, scaling, and maintaining production services.
  • Deep experience with container orchestration platforms, specifically Kubernetes, and cloud-native technologies.
  • Proven track record of designing, implementing, and maintaining sophisticated monitoring and observability solutions.
  • Strong incident management skills with experience leading incident response for complex systems.
  • Experience with infrastructure as code (e.g., Terraform, Pulumi) and configuration management tools.
  • Excellent written and verbal communication skills.
  • Strong interpersonal skills, with experience working with and mentoring engineers.
  • Willingness to debug and improve any layer of the stack.
  • Passion for making software creation accessible.

Benefits

  • Competitive Salary & Equity
  • 401(k) Program with a 4% match (US Only)
  • Health, Dental, Vision and Life Insurance
  • Short Term and Long Term Disability
  • Paid Parental, Medical, Caregiver Leave
  • Flexible Time Off (FTO) + Holidays

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.