Staff Software Engineer, DC Infrastructure
$215k - $260k • San Francisco, CA - US • FullTime
Posted 1mo ago
About the job
Crusoe is seeking a highly skilled and motivated Software Engineer to join its Data Center Infrastructure Engineering team. This role focuses on developing software for managing a fleet of GPU servers and the data centers that house them. The position involves creating and implementing advanced diagnostic, observability, automation, and repair tools for high-performance GPU compute clusters. The ideal candidate will be a hands-on problem solver, comfortable working independently, and will play a crucial role in maintaining the health and scalability of Crusoe's rapidly expanding GPU fleet.
Responsibilities
- Develop and implement deep-level diagnostics and troubleshooting for hardware faults in GPU racks and high-density compute systems.
- Develop troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.
- Develop automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware.
- Develop innovative tooling and AI agents for managing critical environments in conjunction with data center operations.
- Develop tooling for post-repair validation and testing, such as burn-in, Pytorch, and NVIDIA NCCL, to ensure system stability and performance.
- Own the deployment, monitoring, and operational support of developed tooling to maximize GPU fleet availability and performance.
- Develop automation and operational tooling for facilities management power and direct liquid cooling hardware systems.
Requirements
- Software engineering experience.
- Ability to identify problems, rapidly develop scalable solutions, and ship them.
- Ability to assist team members with critical or complex technical initiatives.
- Ability to set technical direction for specific projects and execute.
- Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.).
- Proficiency in at least one programming language: Go, Python, Java, or Rust.
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration skills.
- Ability to work independently and within a team.
- Experience with Temporal and Kubernetes (nice to have).
- Experience working directly with hardware vendors (nice to have).
- Background in large-scale GPU fleet operations or hyperscale data center environments (nice to have).
Benefits
- Industry competitive pay
- Restricted Stock Units
- Health insurance package options (HDHP and PPO)
- Vision insurance
- Dental insurance
- Employer contributions to HSA accounts
- Paid Parental Leave
- Paid life insurance
- Short-term disability
- Long-term disability
- Teladoc
- 401(k) with a 100% match up to 4% of salary
- Generous paid time off
- Holiday schedule
- Cell phone reimbursement
- Tuition reimbursement
- Subscription to the Calm app
- MetLife Legal
- Company paid commuter benefit ($300 per month)
- Bonus