Software Engineer, Infrastructure
San Francisco, CA • FullTime
Posted 3mo ago
About the job
Sierra is building a platform to enable companies to create better, more human customer experiences with AI. As a Software Engineer, Infrastructure, you will be responsible for designing, building, and maintaining the core systems that power our AI platform. Your focus will be on ensuring Sierra's infrastructure is secure, reliable, and scalable, empowering product teams to deliver with speed and confidence.
Responsibilities
- Ensure the reliability, scalability, and performance of the platform and LLM inference serving.
- Build and maintain cloud infrastructure using Terraform for scalable, secure, and reproducible environments.
- Create and maintain a self-serve infrastructure platform for engineering teams.
- Own and evolve CI/CD pipelines and release management for fast, reliable deployments.
- Architect and operate distributed systems leveraging distributed databases, retrieval systems, and ML models.
- Develop and maintain core data serving abstractions, authentication, and security features (SSO, RBAC).
- Integrate the stack with enterprise customer environments in scalable and maintainable ways.
- Enhance observability tooling (metrics, logging, tracing) for platform health visibility.
- Lead and participate in incident management, improving system resilience through monitoring, root cause analysis, and postmortems.
Requirements
- 5-7+ years of hands-on software engineering experience in highly technical products.
- Strong inclination towards building automation, tooling, and platforms.
- Proven experience with cloud platforms (AWS, GCP, or Azure).
- Hands-on expertise with infrastructure as code (Terraform preferred).
- Hands-on expertise in CI/CD systems, release management, and container orchestration (e.g., Docker, Kubernetes).
- Experience with observability tools (Prometheus, Grafana, Datadog, OpenTelemetry, etc.).
- Experience in incident response and operating distributed systems in production.
- Degree in Computer Science or related field, or equivalent professional experience.
- Production experience working with LLMs and machine learning models (preferred).
- Background in distributed systems, running SaaS services at scale, and agentic architecture (preferred).
- Familiarity with security and authentication protocols (OAuth, SSO, mTLS) (preferred).
- Previous experience in a fast-paced startup environment or platform/infra-focused team (preferred).
Benefits
- Flexible (unlimited) paid time off
- Medical, dental, and vision benefits for you and your family
- Life insurance and disability benefits
- Retirement plan dependent on country of employment
- Parental leave
- Fertility and family building benefits
- Lunch, snacks, and coffee
- Discretionary benefit stipend
- Free alphorn lessons