Senior Staff Software Engineer, Cloud Availability Platform

$250k - $300k • Remote • San Francisco, CA - US • FullTime

Posted 4h ago

Remote Work Policy

Fully remote

Employment Type

FullTime

Categories

Machine Learning Engineer

About the job

Crusoe is building the operating system for AI datacenters, powering the world's most ambitious AI workloads with an energy-first approach. We are seeking problem-solving teammates to join our mission to accelerate the abundance of energy and intelligence. This role focuses on building a true platform, not just internal tooling, with an API-first approach, SDKs, and a self-serve portal. The platform consists of agents, a distributed infra graph, a reconciliation core, and domain services, operating on a continuous autonomy loop.

This is a core distributed-systems engineering role where you will design and build foundational platform services like RBAC, tenancy, workflow and policy engines, and state reconcilers. You will also define the public face of the platform through its API gateway, resource model, and versioned API contracts. The work involves creating SDKs, workflow templates, and a developer portal to enable self-service for internal teams, and building the inventory and topology graph to maintain fleet integrity. You will also develop agents and an event bus for reliable telemetry and command execution, and contribute to delivering a roadmap that includes full platform deployment, automated operations, and scaling to over 100,000 GPUs with aggressive performance targets.

Responsibilities

  • Design and build core platform services including RBAC, tenancy, workflow engine, policy engine, and state reconciler.
  • Design the public interface of the platform, including the API gateway, resource model, and versioned API contracts.
  • Develop SDKs, workflow templates, and golden paths for self-service, along with a developer portal and micro frontend framework.
  • Build the inventory and topology graph and associated pipelines to ensure fleet data accuracy.
  • Develop site, GPU, and network agents and the event bus for reliable fleet telemetry and command execution.
  • Deliver the platform roadmap, including pilot site deployments, fully automated operations, and scaling to large GPU fleets with strict performance metrics.
  • Collaborate with embedded engineers.

Requirements

  • Experience building platforms, not just integrating them.
  • Proficiency in core distributed systems engineering concepts like event buses, graph models, reconciliation loops, and policy evaluation.
  • Ability to treat internal engineers as customers, focusing on API ergonomics, documentation, and onboarding.
  • Comfort owning and defending a hard abstraction.
  • Interest in problems at the intersection of physical infrastructure and software.
  • Focus on metrics such as reducing human paging and enabling fast shipping of new features by other teams.

About crusoe

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.