Software Engineer, Site Reliability (SRE)
San Francisco, CA • FullTime
Posted 10mo ago
About the job
Sierra is building a platform to enable companies to create better, more human customer experiences with AI. As a Software Engineer on the Site Reliability team, you will be responsible for establishing the foundation of reliability, observability, and scalability for Sierra's AI-driven infrastructure. You will collaborate closely with core engineering and product teams to ensure our systems are highly available, efficient, and designed for growth.
Responsibilities
- Own the observability stack, including monitoring, alerting, logging, and tracing, to provide clear system health and performance visibility.
- Partner with product and platform engineers to design reliable and scalable systems from inception.
- Design and implement scalable, reliable, and secure cloud infrastructure on AWS using Terraform and modern DevOps tooling.
- Enhance the reliability and scalability of LLM deployments for robust, performant, and cost-effective operation.
- Lead improvements in deployment pipelines, CI/CD tooling, and incident management processes to minimize downtime and response times.
- Define and influence SRE practices, tooling, and best practices across the engineering organization.
Requirements
- 5+ years of hands-on experience in Site Reliability or Infrastructure engineering for complex SaaS or cloud-based systems.
- Experience designing for availability, scalability, and reliability at both infrastructure and application layers.
- Deep experience with Terraform, AWS services, container orchestration, and cloud networking (including IAM and VPC architecture).
- Strong background in observability systems (e.g., Prometheus, Grafana, Datadog, or similar).
- Experience working with enterprise customers, understanding their compliance and networking needs, and integration patterns.
- Comfortable working in fast-moving environments and collaborating across product, ML, and core engineering teams.
- Degree in Computer Science or a related field, or equivalent professional experience.
- Experience with LLM infrastructure (optimizing inference, managing fine-tuned models, or large-scale model deployment) is a plus.
- Past experience in an early-stage startup, especially defining SRE culture and tooling, is a plus.
- Familiarity with incident management automation or self-healing infrastructure patterns is a plus.
Benefits
- Flexible (unlimited) paid time off
- Medical, dental, and vision benefits for you and your family
- Life insurance and disability benefits
- Retirement plan dependent on country of employment
- Parental leave
- Fertility and family building benefits
- Lunch, snacks, and coffee
- Discretionary benefit stipend
- Free alphorn lessons
- Equity plans