Site Reliability Engineer
$150k - $200k • Remote • Remote - USA • FullTime
Posted 2mo ago
About the job
Runpod is seeking a Site Reliability Engineer to join their remote-first team. This role is central to ensuring the stability, performance, and operational excellence of Runpod's global AI Developer Cloud platform. You will partner with engineering teams to enhance system design, strengthen observability, and proactively prevent incidents. This position blends software engineering with production operations, focusing on reliability frameworks, SLO design, automation, and production hardening to reduce errors and improve performance across various services and infrastructure. Your work will directly impact platform uptime, incident frequency, and overall system resilience, playing a key role in maintaining trust with developers running critical AI workloads.
Responsibilities
- Define and enforce reliability standards across engineering.
- Design incident response processes and improve recovery times.
- Build observability systems and reliability tooling.
- Drive SLO adoption and production readiness reviews.
- Reduce operational toil through automation.
- Define and implement SLIs/SLOs for critical services.
- Lead incident response and coordinate cross-team mitigation efforts.
- Conduct blameless postmortems and ensure corrective actions are completed.
- Perform production readiness reviews for new services and features.
- Identify systemic risks and drive preventative improvements.
- Design and improve monitoring, alerting, and dashboards.
- Improve signal-to-noise ratio in alerts and reduce alert fatigue.
- Build internal tooling for reliability tracking and reporting.
- Improve visibility into GPU performance and distributed systems health.
- Automate recurring operational workflows.
- Build tools and scripts to eliminate manual processes.
- Improve deployment safety through automation and guardrails.
- Strengthen CI/CD reliability and release processes.
- Partner with engineering teams to improve system resilience.
- Provide guidance on fault tolerance, scalability, and failure handling.
- Contribute to architectural discussions with a reliability-first mindset.
Requirements
- 5+ years of experience in SRE, Reliability Engineering, or Production Engineering.
- Strong Linux systems and Networking expertise.
- Experience managing containerized production systems.
- Strong understanding of distributed systems and failure modes.
- Experience defining and managing SLIs/SLOs.
- Proven incident response and postmortem leadership experience.
- Strong scripting or programming skills.
- Experience with monitoring and alerting systems.
- Excellent written communication skills.
- Experience with GPU infrastructure or AI/ML platforms (preferred).
- Experience improving reliability in high-growth or large scale environments (preferred).
- Familiarity with GPU observability tooling (preferred).
- Experience with Infrastructure as Code (preferred).
- Experience working in startup environments (preferred).
- Experience building internal reliability platforms or frameworks (preferred).
Benefits
- Competitive base pay ranging from $150,000- $200,000 usd.
- Meaningful equity in a fast-growing company.
- Generous medical, dental & vision plans.
- Flexible PTO.