Software Engineer, Site Reliability (SRE)

San Francisco, CA FullTime

Posted 10mo ago

Job Location

San Francisco, CA

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

Sierra is building a platform to enable companies to create better, more human customer experiences with AI. As a Software Engineer on the Site Reliability team, you will be responsible for establishing the foundation of reliability, observability, and scalability for Sierra's AI-driven infrastructure. You will collaborate closely with core engineering and product teams to ensure our systems are highly available, efficient, and designed for growth.

Responsibilities

  • Own the observability stack, including monitoring, alerting, logging, and tracing, to provide clear system health and performance visibility.
  • Partner with product and platform engineers to design reliable and scalable systems from inception.
  • Design and implement scalable, reliable, and secure cloud infrastructure on AWS using Terraform and modern DevOps tooling.
  • Enhance the reliability and scalability of LLM deployments for robust, performant, and cost-effective operation.
  • Lead improvements in deployment pipelines, CI/CD tooling, and incident management processes to minimize downtime and response times.
  • Define and influence SRE practices, tooling, and best practices across the engineering organization.

Requirements

  • 5+ years of hands-on experience in Site Reliability or Infrastructure engineering for complex SaaS or cloud-based systems.
  • Experience designing for availability, scalability, and reliability at both infrastructure and application layers.
  • Deep experience with Terraform, AWS services, container orchestration, and cloud networking (including IAM and VPC architecture).
  • Strong background in observability systems (e.g., Prometheus, Grafana, Datadog, or similar).
  • Experience working with enterprise customers, understanding their compliance and networking needs, and integration patterns.
  • Comfortable working in fast-moving environments and collaborating across product, ML, and core engineering teams.
  • Degree in Computer Science or a related field, or equivalent professional experience.
  • Experience with LLM infrastructure (optimizing inference, managing fine-tuned models, or large-scale model deployment) is a plus.
  • Past experience in an early-stage startup, especially defining SRE culture and tooling, is a plus.
  • Familiarity with incident management automation or self-healing infrastructure patterns is a plus.

Benefits

  • Flexible (unlimited) paid time off
  • Medical, dental, and vision benefits for you and your family
  • Life insurance and disability benefits
  • Retirement plan dependent on country of employment
  • Parental leave
  • Fertility and family building benefits
  • Lunch, snacks, and coffee
  • Discretionary benefit stipend
  • Free alphorn lessons
  • Equity plans

About sierra.ai

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.