Senior Site Reliability Engineer
Remote • Remote - Europe • FullTime
Posted 2mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to design and implement robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability and performance.
Responsibilities
- Design and implement comprehensive monitoring and alerting systems using modern observability tools.
- Create dashboards and metrics for real-time visibility into system health and performance.
- Implement logging strategies for quick problem identification and resolution.
- Architect and implement infrastructure automation solutions using tools like Terraform, Ansible, or Pulumi.
- Design and maintain CI/CD pipelines for reliable and consistent deployments.
- Create self-healing systems that automatically respond to common failure scenarios.
- Define and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs) with product and engineering teams.
- Build systems to track and report on SLOs/SLIs to ensure high reliability standards.
- Lead incident response efforts, conduct post-mortems, and implement preventative improvements.
- Develop and maintain runbooks for critical services.
- Build tools and processes to reduce Mean Time To Recovery (MTTR).
- Identify and resolve performance bottlenecks across infrastructure.
- Implement capacity planning strategies and optimize resource utilization.
- Work on reducing latency and improving system efficiency across global regions.
Requirements
- 4-8 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering, Infrastructure Engineering).
- Strong programming skills in languages commonly used for automation (Python, Go, or similar).
- Deep understanding of distributed systems.
- Experience with container orchestration platforms (Kubernetes) and cloud-native technologies.
- Proven track record of implementing and maintaining monitoring/observability solutions.
- Strong incident management skills with experience leading incident response.
- Experience with infrastructure as code and configuration management tools.
Benefits
- Competitive Salary & Equity
- 401(k) Program with a 4% match (US Only)
- Health, Dental, Vision and Life Insurance
- Short Term and Long Term Disability
- Paid Parental, Medical, Caregiver Leave
- Flexible Time Off (FTO) + Holidays
- Commuter Benefits (In-Office Only)
- Monthly Wellness Stipend
- Autonomous Work Environment
- In Office Set-Up Reimbursement (In-Office Only)
- Quarterly Team Gatherings
- In Office Amenities (In-Office Only)