Site Reliability Engineer - Memphis
Southaven, MS; Memphis, TN
Posted 5d ago
About the job
SpaceXAI is seeking a Site Reliability Engineer focused on campus reliability to design and implement systems that ensure the availability and trustworthiness of our campus infrastructure. This role involves technically leading cross-discipline incident responses and building mechanisms to minimize future incidents. You will act as the central point of contact across compute, network, storage, power, and cooling systems, requiring strong incident leadership, fleet-scale observability expertise, and the ability to drive reliability improvements across both software and facility domains.
Responsibilities
- Own monitoring architecture and signal quality, including alert suppression and trust, using NOC feedback to improve designs.
- Provide SEV command support, including technical incident leadership, bridge coordination, and timeline/severity management.
- Conduct blameless postmortems and ensure corrective actions are fully implemented.
- Lead cross-functional reliability projects involving compute, network, storage, and facility telemetry.
- Develop and maintain playbooks, conduct game days, and keep dependency maps up-to-date.
- Define error budgets and availability objectives at campus and service levels.
- Participate in on-call rotations and incident response for SEV-class events in the Memphis data center campus.
Requirements
- Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or equivalent experience.
- 5+ years of experience in site reliability, systems engineering, or large-scale production operations.
- Proven experience in large-scale incident command and technical leadership during incidents.
- Demonstrated experience in designing monitoring and observability systems at fleet or campus scale.
- Experience working with at least two of the following: compute, network, storage, power, and cooling telemetry.
- Experience writing and operating playbooks or runbooks with a 24/7 operations team.
- Proficiency in scripting (Python, Bash) and experience in at least one systems language (C, C++, Java, Go, Rust, or similar).
- Excellent problem-solving skills with a data-driven approach.
- Ability to collaborate effectively with cross-functional teams.