Staff Site Reliability Engineer
Remote • Remote - Europe • FullTime
Posted 2mo ago
About the job
Join our Site Reliability Engineering (SRE) team to ensure the reliability, scalability, and performance of Replit's infrastructure serving millions of developers. As a Staff Site Reliability Engineer, you will bridge development and operations, implementing automation and best practices for efficient scaling and high availability. We seek passionate individuals to proactively identify and analyze reliability issues, then design and implement solutions for significant improvements. Your mission includes designing observability, leading incident response, automating tasks, and mentoring the engineering team to embed reliability as a core value.
Responsibilities
- Architect and implement comprehensive monitoring, logging, and tracing solutions, creating dashboards for real-time system visibility.
- Define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs), building systems to monitor and report on these metrics.
- Lead incident management and response for high-impact incidents, guiding the team to rapid resolution and conducting blameless post-mortems.
- Drive automation to eliminate toil and operational work, designing and maintaining CI/CD pipelines and infrastructure as code.
- Optimize performance on Kubernetes, Docker, and GCP, identifying and resolving bottlenecks and implementing capacity planning.
- Debug and harden distributed systems, designing and implementing long-term fixes for robustness and operability.
- Review feature and system designs across the company, owning reliability, scalability, security, and operational integrity.
- Educate and mentor the broader engineering team to improve system reliability and foster a reliability-centric culture.
- Write high-quality, well-tested code in Python or Go for internal tools and third-party integrations.
Requirements
- 8-10 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering, Infrastructure Engineering).
- Strong programming skills in Python or Go, with experience writing high-quality, well-tested code.
- Deep understanding of distributed systems, including designing, building, scaling, and maintaining production services and service-oriented architectures.
- Deep experience with container orchestration platforms, specifically Kubernetes, and cloud-native technologies.
- Proven track record of designing, implementing, and maintaining sophisticated monitoring and observability solutions.
- Strong incident management skills with experience leading incident response for complex systems.
- Experience with infrastructure as code (e.g., Terraform, Pulumi) and configuration management tools.
- Excellent written and verbal communication skills, with the ability to explain complex technical concepts clearly.
- Strong interpersonal skills, with experience working with and mentoring engineers at various levels.
- A willingness to dive into understanding, debugging, and improving any layer of the stack.
- Passion for making software creation accessible.
Benefits
- Competitive Salary & Equity
- 401(k) Program with a 4% match (US Only)
- Health, Dental, Vision and Life Insurance
- Short Term and Long Term Disability
- Paid Parental, Medical, Caregiver Leave
- Flexible Time Off (FTO)