Engineering Manager, Site Reliability Engineering
Remote • Foster City, CA • FullTime
Posted 2h ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
Machine Learning Engineer
About the job
Replit is seeking an Engineering Manager to lead Site Reliability Engineering (SRE) across critical areas including observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure. This role involves leading and growing an existing team that builds and operates production platforms, working hands-on across application and infrastructure boundaries. It's a software-building leadership position focused on enabling teams to ship safely, understand production behavior, and implement engineering improvements to address performance bottlenecks. The ideal candidate will be comfortable diving deep into technical issues while also developing technical leaders and fostering sustainable ownership within a distributed team.
Responsibilities
- Build and operate metrics, logs, traces, and alerting capabilities, helping teams establish SLOs and use telemetry for problem diagnosis and improvement verification.
- Own incident tooling and practices, coordinate cross-team responses, and translate incident reviews into engineering improvements to reduce recovery time and prevent repeat failures.
- Build and maintain load/failure testing capabilities to validate critical paths under demand, quantify headroom, and test recovery and production readiness.
- Lead engagements with internal teams on SLOs and end-to-end performance, using profiling, telemetry, and load tests to identify and implement improvements.
- Stay technically engaged by reviewing designs and production changes, debugging failures, and using AI coding tools to prototype and automate.
- Build and grow a high-ownership engineering team through coaching, developing technical leaders, managing performance, and hiring.
- Measure outcomes such as rollout safety, recovery time, repeat incidents, and latency/throughput, and track improvements from cost/capacity analysis.
- Ensure deliberate distributed collaboration, mentoring, and backup coverage within the team.
Requirements
- Demonstrated engineering management experience, including leading and developing engineers, making prioritization and performance decisions, hiring, and delivering through a team.
- Deep understanding of software-oriented production systems, including building and operating distributed systems or reliability platforms.
- Experience reasoning across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms.
- Experience leading consequential migrations or incidents and using measurement to diagnose reliability or performance problems.
- Ability to distinguish symptoms from causes and validate fixes under realistic conditions.
- Experience building capabilities that other teams adopt and leading hands-on engagements without absorbing all service operations.
- Ability to make clear tradeoffs among reliability, performance, engineering effort, and cost.
Benefits
- Competitive Salary & Equity
- 401(k) Program with a 4% match (US Only)
- Health, Dental, Vision and Life Insurance
- Short Term and Long Term Disability
- Paid Parental, Medical, Caregiver Leave
- Flexible Time Off (FTO) + Holidays
- Commuter Benefits (In-Office & US Only)
- Monthly Wellness Stipend
- Autonomous Work Environment
- In Office Set-Up Reimbursement (In-Office Only)
- Quarterly Team Gatherings
- In Office Amenities (In-Office Only)