Staff Engineer, Distributed Storage and HPC & AI Infrastructure

$250k - $300k Remote San Francisco

Posted 1mo ago

Remote Work Policy

Fully remote

Categories

AI Infrastructure Engineer

About the job

Together AI is seeking a Staff Engineer to design and deliver multi-petabyte storage systems optimized for large-scale AI training and inference workloads. You will architect high-performance parallel filesystems and object stores, integrate cutting-edge technologies, and drive significant cost optimization. The role involves building Kubernetes-native storage operators and self-service platforms for automated provisioning and multi-tenancy. You will focus on optimizing data paths, designing multi-tier caching architectures, and tuning parallel filesystems for AI applications. This is a research-driven role within a company focused on lowering the cost of modern AI systems through co-design of software, hardware, algorithms, and models.

Responsibilities

  • Design multi-petabyte AI/ML storage systems and integrate technologies like WekaFS and Ceph.
  • Lead capacity planning and cost optimization efforts to achieve significant savings.
  • Design and optimize RDMA, InfiniBand, and 400GbE networks for maximum throughput and minimal latency.
  • Implement NVMe-oF/iSCSI and troubleshoot storage bottlenecks.
  • Build Kubernetes storage operators and controllers for automated provisioning, self-service, multi-tenancy, and quotas.
  • Deliver high data throughput per GPU node and optimize caching, parallel filesystems, and data paths.
  • Build and optimize multi-tier caches, data locality, and model-weight distribution.
  • Implement monitoring, alerting, SLOs, and design disaster recovery and backup strategies.
  • Run chaos engineering and ensure high uptime through proactive remediation.
  • Partner with ML/SRE teams, mentor on storage best practices, contribute to open-source, and document learnings.

Requirements

  • 8+ years in storage engineering, with 3+ years managing distributed storage at multi-petabyte scale.
  • Proven track record deploying and operating high-performance storage for GPU/HPC clusters.
  • Deep Kubernetes and cloud-native storage experience in production environments.
  • Strong coding skills in Go and Python for building production-grade tools.
  • BS/MS in Computer Science, Engineering, or equivalent practical experience.
  • History of technical leadership, designing systems that improved performance, reliability, or cost efficiency.
  • Deep expertise in parallel filesystems (WekaFS, Lustre, GPFS, BeeGFS) at multi-petabyte scale.
  • Production experience with object storage (S3, MinIO, Ceph, R2) including performance and cost management.
  • Experience with Kubernetes storage concepts (CSI drivers, StatefulSets, PersistentVolumes, storage operators).
  • Expertise in storage optimization for GPU workloads, RDMA/InfiniBand networking, and parallel filesystem optimization.
  • Proficiency in Infrastructure as Code tools like Terraform, Ansible, Helm, and GitOps.
  • Advanced knowledge of the Linux storage stack (filesystems, LVM, NVMe optimization).
  • Experience with observability tools like Prometheus, Grafana, and Thanos.

Benefits

  • Competitive compensation
  • Startup equity
  • Health insurance
  • Flexibility in terms of remote work

About Together AI

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.