Senior Site Reliability Engineer
Remote • Bengaluru • FullTime
Posted 2mo ago
About the job
As a Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability, and performance of our legal AI platform. You’ll join a high-leverage team that sits at the intersection of infrastructure and product, owning the systems that keep our platform fast, secure, and always on. From scaling across 50+ regions to automating mission-critical operations, your work will ensure that Harvey remains resilient as we grow. If you’re passionate about building robust systems and reducing complexity through automation, we’d love to work with you.
Responsibilities
- Design, implement, and manage monitoring, alerting, and infrastructure resources across global regions
- Lead incident management processes, including postmortems and root cause analyses
- Automate operational tasks and workflows to maintain high reliability and reduce manual intervention
- Collaborate across teams to drive reliability, security, and compliance throughout the software lifecycle
- Optimize infrastructure costs while maintaining system performance and reliability
Requirements
- 3+ years of experience in Site Reliability Engineering or similar roles
- Expertise in infrastructure as code (IaC) tools
- Deep familiarity with observability tools and incident response practices
- Proficiency with cloud infrastructure platforms
- Strong programming skills in Python, Bash, Go, or similar languages
- Proven track record of diagnosing complex system problems and implementing durable solutions
- Solid understanding of CI/CD, Kubernetes, containerization, networking, databases, and cloud security principles
- Excellent problem-solving skills and meticulous attention to detail