Staff Cloud Support Engineer

$156k - $190k San Francisco, CA - US FullTime

Posted 7mo ago

Job Location

San Francisco, CA - US

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

Applied AI Engineer

About the job

As a Staff Cloud Support Engineer at Crusoe, you will be a technical authority within Crusoe Cloud, acting as a force multiplier for Customer Experience, SRE, Networking, Fleet, and Product teams. Your role extends beyond simple ticket resolution; you will design reliability guardrails, influence architectural decisions, mentor other engineers, and directly contribute to revenue protection by preventing large-scale incidents. This position requires deep expertise in Linux systems, Kubernetes, networking, and AI/ML infrastructure, applied with a strong customer focus. You should be comfortable operating in ambiguous environments, leading incident response efforts, and shaping the scalability of Crusoe's high-performance AI infrastructure globally.

Responsibilities

  • Serve as the highest-level escalation point for complex P1/P0 incidents.
  • Lead cross-functional root cause investigations involving compute, networking, storage, and orchestration layers.
  • Partner with SRE and Software teams to design systemic fixes for recurring issues.
  • Design and improve node validation, burn-in processes, performance baselining, and release readiness.
  • Influence Kubernetes architecture, workload orchestration, and AI/ML cluster stability.
  • Reduce Mean Time To Recover (MTTR) and incident recurrence through structural improvements.
  • Troubleshoot NCCL, Infiniband, GPU driver/firmware issues, and distributed training failures.
  • Support complex AI workloads with performance tuning and observability improvements.
  • Act as a technical advisor during high-risk customer incidents.
  • Deliver executive-ready Root Cause Analyses (RCAs).
  • Mentor junior engineers.
  • Define Standard Operating Procedures (SOPs) and technical standards for support excellence.
  • Partner with Enablement teams to raise the technical bar across the organization.

Requirements

  • 8+ years of experience in SRE, DevOps, HPC, or Cloud Infrastructure roles.
  • Advanced Linux systems expertise.
  • Deep Kubernetes operational experience (CKA-level or higher).
  • Strong networking knowledge, including Infiniband, RDMA, RoCE, and SDN.
  • Experience supporting AI/ML workloads at scale, specifically GPU clusters.
  • Proven track record of resolving multi-layer, distributed system failures.
  • Strong customer communication skills and executive-facing presence.

Benefits

  • Competitive compensation
  • Restricted Stock Units (RSUs)
  • Paid time off
  • Paid holidays
  • Comprehensive health, dental, and vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance
  • Short-term and long-term disability insurance
  • Professional development opportunities
  • Tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits (parking and transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off

About crusoe

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.