Staff Cloud Support Engineer
$156k - $190k • San Francisco, CA - US • FullTime
Posted 7mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
As a Staff Cloud Support Engineer at Crusoe, you will be a technical authority within Crusoe Cloud, acting as a force multiplier for Customer Experience, SRE, Networking, Fleet, and Product teams. Your role extends beyond simple ticket resolution; you will design reliability guardrails, influence architectural decisions, mentor other engineers, and directly contribute to revenue protection by preventing large-scale incidents. This position requires deep expertise in Linux systems, Kubernetes, networking, and AI/ML infrastructure, applied with a strong customer focus. You should be comfortable operating in ambiguous environments, leading incident response efforts, and shaping the scalability of Crusoe's high-performance AI infrastructure globally.
Responsibilities
- Serve as the highest-level escalation point for complex P1/P0 incidents.
- Lead cross-functional root cause investigations involving compute, networking, storage, and orchestration layers.
- Partner with SRE and Software teams to design systemic fixes for recurring issues.
- Design and improve node validation, burn-in processes, performance baselining, and release readiness.
- Influence Kubernetes architecture, workload orchestration, and AI/ML cluster stability.
- Reduce Mean Time To Recover (MTTR) and incident recurrence through structural improvements.
- Troubleshoot NCCL, Infiniband, GPU driver/firmware issues, and distributed training failures.
- Support complex AI workloads with performance tuning and observability improvements.
- Act as a technical advisor during high-risk customer incidents.
- Deliver executive-ready Root Cause Analyses (RCAs).
- Mentor junior engineers.
- Define Standard Operating Procedures (SOPs) and technical standards for support excellence.
- Partner with Enablement teams to raise the technical bar across the organization.
Requirements
- 8+ years of experience in SRE, DevOps, HPC, or Cloud Infrastructure roles.
- Advanced Linux systems expertise.
- Deep Kubernetes operational experience (CKA-level or higher).
- Strong networking knowledge, including Infiniband, RDMA, RoCE, and SDN.
- Experience supporting AI/ML workloads at scale, specifically GPU clusters.
- Proven track record of resolving multi-layer, distributed system failures.
- Strong customer communication skills and executive-facing presence.
Benefits
- Competitive compensation
- Restricted Stock Units (RSUs)
- Paid time off
- Paid holidays
- Comprehensive health, dental, and vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance
- Short-term and long-term disability insurance
- Professional development opportunities
- Tuition reimbursement
- Mental health and wellness support
- Commuter benefits (parking and transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off