Principal Systems Software Engineer
$260k - $340k • San Francisco, CA - US • FullTime
Posted 1mo ago
Remote Work Policy
On-site
Employment Type
FullTime
Categories
Applied AI Engineer
About the job
Crusoe is seeking a Principal Systems Software Engineer to lead the vision for its next-generation AI infrastructure. This role is for an industry expert with hyperscale experience, tasked with redefining the I/O path for generative AI. You will design the core fabric that unifies Bare-Metal-as-a-Service, Intelligent IaaS, and Elastic CaaS into a high-performance pool of intelligence. This position involves bridging silicon and software, advising executive leadership on hardware/software co-design, and leading R&D teams in shipping production-grade kernel and orchestration code. The ideal candidate is a master of the I/O path, capable of pushing massive-scale training workloads to hardware limits.
Responsibilities
- Architect systems for raw GPU throughput via zero-latency InfiniBand/RDMA fabrics for massive-scale training.
- Design optimized, thin virtualization layers using KVM or custom micro-VMs for enterprise-grade isolation.
- Build a high-performance container substrate (Kubernetes or Slurm) for AI workloads to burst and scale across heterogeneous GPU nodes.
- Lead the architectural design of the internal cloud fabric, driving the technical roadmap for SR-IOV, RDMA, and virtualized GPU scheduling.
- Lead workstreams to prototype and productionize novel methods for managing memory, networking, and compute.
- Draft white papers and RFCs defining the compute and networking stack roadmap.
- Resolve complex race conditions in the I/O path and optimize kernel-level memory pinning for GPU clusters.
- Represent Crusoe in open-source communities and industry forums to influence cloud-native AI infrastructure direction.
Requirements
- 12+ years of experience designing and shipping core infrastructure at a major hyperscaler or specialized HPC cloud.
- Authoritative knowledge of the Linux kernel, virtualization internals (KVM, QEMU, Firecracker), and high-performance networking (RoCE v2, InfiniBand).
- Proven ability to design software that maximizes the performance of NVIDIA/AMD GPUs and high-speed NICs.
- Experience leading cross-functional teams through high-ambiguity projects and delivering production-ready systems.
- A portfolio of significant contributions to the field (patents, major open-source contributions, or published research).
- Ability to explain complex technical concepts to both engineers and executive leadership.
- Bachelor's or Master's degree in Computer Science, Computer Engineering, or related field (or equivalent professional experience).
Benefits
- Competitive compensation
- Restricted Stock Units
- Paid time off & paid holidays
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off