Staff Software Engineer, DC Infrastructure

$215k - $260k San Francisco, CA - US FullTime

Posted 1mo ago

Job Location

San Francisco, CA - US

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

AI Infrastructure Engineer

About the job

Crusoe is seeking a highly skilled and motivated Software Engineer to join its Data Center Infrastructure Engineering team. This role focuses on developing software for managing a fleet of GPU servers and the data centers that house them. The position involves creating and implementing advanced diagnostic, observability, automation, and repair tools for high-performance GPU compute clusters. The ideal candidate will be a hands-on problem solver, comfortable working independently, and will play a crucial role in maintaining the health and scalability of Crusoe's rapidly expanding GPU fleet.

Responsibilities

  • Develop and implement deep-level diagnostics and troubleshooting for hardware faults in GPU racks and high-density compute systems.
  • Develop troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.
  • Develop automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware.
  • Develop innovative tooling and AI agents for managing critical environments in conjunction with data center operations.
  • Develop tooling for post-repair validation and testing, such as burn-in, Pytorch, and NVIDIA NCCL, to ensure system stability and performance.
  • Own the deployment, monitoring, and operational support of developed tooling to maximize GPU fleet availability and performance.
  • Develop automation and operational tooling for facilities management power and direct liquid cooling hardware systems.

Requirements

  • Software engineering experience.
  • Ability to identify problems, rapidly develop scalable solutions, and ship them.
  • Ability to assist team members with critical or complex technical initiatives.
  • Ability to set technical direction for specific projects and execute.
  • Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.).
  • Proficiency in at least one programming language: Go, Python, Java, or Rust.
  • Strong analytical and problem-solving skills.
  • Excellent communication and collaboration skills.
  • Ability to work independently and within a team.
  • Experience with Temporal and Kubernetes (nice to have).
  • Experience working directly with hardware vendors (nice to have).
  • Background in large-scale GPU fleet operations or hyperscale data center environments (nice to have).

Benefits

  • Industry competitive pay
  • Restricted Stock Units
  • Health insurance package options (HDHP and PPO)
  • Vision insurance
  • Dental insurance
  • Employer contributions to HSA accounts
  • Paid Parental Leave
  • Paid life insurance
  • Short-term disability
  • Long-term disability
  • Teladoc
  • 401(k) with a 100% match up to 4% of salary
  • Generous paid time off
  • Holiday schedule
  • Cell phone reimbursement
  • Tuition reimbursement
  • Subscription to the Calm app
  • MetLife Legal
  • Company paid commuter benefit ($300 per month)
  • Bonus

About crusoe

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.