Senior Site Reliability Engineer

Remote Remote - Europe FullTime

Posted 2mo ago

Job Location

Remote - Europe

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to design and implement robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability and performance.

Responsibilities

  • Design and implement comprehensive monitoring and alerting systems using modern observability tools.
  • Create dashboards and metrics for real-time visibility into system health and performance.
  • Implement logging strategies for quick problem identification and resolution.
  • Architect and implement infrastructure automation solutions using tools like Terraform, Ansible, or Pulumi.
  • Design and maintain CI/CD pipelines for reliable and consistent deployments.
  • Create self-healing systems that automatically respond to common failure scenarios.
  • Define and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs) with product and engineering teams.
  • Build systems to track and report on SLOs/SLIs to ensure high reliability standards.
  • Lead incident response efforts, conduct post-mortems, and implement preventative improvements.
  • Develop and maintain runbooks for critical services.
  • Build tools and processes to reduce Mean Time To Recovery (MTTR).
  • Identify and resolve performance bottlenecks across infrastructure.
  • Implement capacity planning strategies and optimize resource utilization.
  • Work on reducing latency and improving system efficiency across global regions.

Requirements

  • 4-8 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering, Infrastructure Engineering).
  • Strong programming skills in languages commonly used for automation (Python, Go, or similar).
  • Deep understanding of distributed systems.
  • Experience with container orchestration platforms (Kubernetes) and cloud-native technologies.
  • Proven track record of implementing and maintaining monitoring/observability solutions.
  • Strong incident management skills with experience leading incident response.
  • Experience with infrastructure as code and configuration management tools.

Benefits

  • Competitive Salary & Equity
  • 401(k) Program with a 4% match (US Only)
  • Health, Dental, Vision and Life Insurance
  • Short Term and Long Term Disability
  • Paid Parental, Medical, Caregiver Leave
  • Flexible Time Off (FTO) + Holidays
  • Commuter Benefits (In-Office Only)
  • Monthly Wellness Stipend
  • Autonomous Work Environment
  • In Office Set-Up Reimbursement (In-Office Only)
  • Quarterly Team Gatherings
  • In Office Amenities (In-Office Only)

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.