Staff Security Reliability Engineer
Remote • San Francisco • FullTime
Posted 2mo ago
About the job
We are seeking an experienced Site Reliability Engineer to focus on security infrastructure. This role involves designing, building, and operating reliable, secure, and scalable infrastructure for identity, access, endpoint, and shared platform services across the company. You will be a senior technical owner, responsible for the end-to-end lifecycle of infrastructure and identity systems, from architecture and implementation to policy enforcement, upgrades, recovery, and ongoing operations. The goal is to build durable, production-grade platforms that reduce operational friction, enforce security by default, and empower teams to move forward with confidence. This is a hands-on role for a senior engineer who excels in ambiguous environments, enjoys end-to-end system ownership, and elevates reliability and security by replacing fragile systems with standardized, repeatable infrastructure.
Responsibilities
- Design, build, and operate reliable infrastructure across on-prem, hybrid, shared, and product-adjacent environments.
- Establish standardized infrastructure patterns to replace bespoke implementations with repeatable, auditable, secure-by-default systems.
- Own the lifecycle of critical infrastructure platforms, including provisioning, deployment, upgrades, patching, recovery, and long-term reliability.
- Build infrastructure-as-code and configuration management using tools like Terraform, Chef, and Ansible.
- Mature identity-adjacent and policy-enforced infrastructure, including Microsoft Entra and Azure management patterns.
- Build observability, alerting, and incident response mechanisms to improve availability, recoverability, and operational confidence.
- Automate high-toil and high-risk workflows with guardrails, progressive rollout patterns, and safe rollback paths.
- Translate incidents, design reviews, and operational learnings into durable fixes, reusable patterns, and stronger technical standards.
- Influence cross-functional partners across security, identity, network, and platform teams through architecture, implementation, operational data, and clear technical writing.
Requirements
- 10+ years of hands-on experience operating and architecting mission-critical infrastructure in high-reliability environments.
- Experience running Security Infrastructure.
- Served as the senior technical owner for the design and maturation of complex on-prem, hybrid, or cloud-integrated systems, establishing durable architectural patterns used by multiple teams.
- Applied Site Reliability Engineering principles at scale, using observability, automation, and incident learnings to materially reduce risk and operational toil.
- Comfortable operating in ambiguity, making sound architectural decisions under pressure while staying close to technical detail.
- Experience operating infrastructure for R&D or specialized labs, manufacturing, or other safety-critical environments where uptime and recoverability are essential.
- Experience with fleet, endpoint, or virtual desktop platforms such as FleetDM, Chef, or Azure Virtual Desktop.
- Experience partnering closely with identity or security engineering teams on hardened, policy-enforced infrastructure at scale.