Staff Site Reliability Engineer

Remote Bengaluru FullTime

Posted 4mo ago

Job Location

Bengaluru

Tech Stack

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Harvey is transforming legal and professional services by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. This is a rare opportunity to help build a generational company at an inflection point, with strong product-market fit and world-class investor support. The team moves fast, takes ownership, and is deeply committed to the mission, operating with intensity and pushing for excellence. As a Staff Software Engineer on the Site Reliability team, you will ensure the reliability, scalability, and performance of our legal AI platform, owning the systems that keep our platform fast, secure, and always on. Your work will be critical in scaling across 50+ regions and automating mission-critical operations to ensure Harvey remains resilient as we grow. If you are passionate about building robust systems and reducing complexity through automation, we encourage you to apply.

Responsibilities

  • Design, implement, and manage monitoring, alerting, and infrastructure resources across global regions.
  • Lead incident management processes, including postmortems and root cause analyses.
  • Automate operational tasks and workflows to maintain high reliability and reduce manual intervention.
  • Establish and drive best practices for security, compliance, and reliability across teams.
  • Optimize infrastructure costs while maintaining system performance and reliability.
  • Provide technical mentorship and leadership to foster team growth.

Requirements

  • 12+ years of experience in Site Reliability Engineering or similar roles.
  • Proven ability to mentor and guide technical teams.
  • Expertise in infrastructure as code (IaC) tools like Pulumi, Terraform, or CloudFormation.
  • Deep familiarity with observability tools (Datadog, Sentry) and incident response practices (PagerDuty).
  • Proficiency with cloud infrastructure platforms (Azure, GCP, AWS).
  • Strong programming skills in Python, Bash, Go, or similar languages.
  • Proven track record of diagnosing complex system problems and implementing durable solutions.
  • Solid understanding of CI/CD, Kubernetes, containerization, networking, databases, and cloud security principles.
  • Excellent problem-solving skills and meticulous attention to detail.

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.