Staff Engineer, Distributed Storage and HPC & AI Infrastructure
$250k - $300k • Remote • San Francisco
Posted 1mo ago
Job Location
San Francisco
Tech Stack
Remote Work Policy
Fully remote
Categories
AI Infrastructure Engineer
About the job
Together AI is seeking a Staff Engineer to design and deliver multi-petabyte storage systems optimized for large-scale AI training and inference workloads. You will architect high-performance parallel filesystems and object stores, integrate cutting-edge technologies, and drive significant cost optimization. The role involves building Kubernetes-native storage operators and self-service platforms for automated provisioning and multi-tenancy. You will focus on optimizing data paths, designing multi-tier caching architectures, and tuning parallel filesystems for AI applications. This is a research-driven role within a company focused on lowering the cost of modern AI systems through co-design of software, hardware, algorithms, and models.
Responsibilities
- Design multi-petabyte AI/ML storage systems and integrate technologies like WekaFS and Ceph.
- Lead capacity planning and cost optimization efforts to achieve significant savings.
- Design and optimize RDMA, InfiniBand, and 400GbE networks for maximum throughput and minimal latency.
- Implement NVMe-oF/iSCSI and troubleshoot storage bottlenecks.
- Build Kubernetes storage operators and controllers for automated provisioning, self-service, multi-tenancy, and quotas.
- Deliver high data throughput per GPU node and optimize caching, parallel filesystems, and data paths.
- Build and optimize multi-tier caches, data locality, and model-weight distribution.
- Implement monitoring, alerting, SLOs, and design disaster recovery and backup strategies.
- Run chaos engineering and ensure high uptime through proactive remediation.
- Partner with ML/SRE teams, mentor on storage best practices, contribute to open-source, and document learnings.
Requirements
- 8+ years in storage engineering, with 3+ years managing distributed storage at multi-petabyte scale.
- Proven track record deploying and operating high-performance storage for GPU/HPC clusters.
- Deep Kubernetes and cloud-native storage experience in production environments.
- Strong coding skills in Go and Python for building production-grade tools.
- BS/MS in Computer Science, Engineering, or equivalent practical experience.
- History of technical leadership, designing systems that improved performance, reliability, or cost efficiency.
- Deep expertise in parallel filesystems (WekaFS, Lustre, GPFS, BeeGFS) at multi-petabyte scale.
- Production experience with object storage (S3, MinIO, Ceph, R2) including performance and cost management.
- Experience with Kubernetes storage concepts (CSI drivers, StatefulSets, PersistentVolumes, storage operators).
- Expertise in storage optimization for GPU workloads, RDMA/InfiniBand networking, and parallel filesystem optimization.
- Proficiency in Infrastructure as Code tools like Terraform, Ansible, Helm, and GitOps.
- Advanced knowledge of the Linux storage stack (filesystems, LVM, NVMe optimization).
- Experience with observability tools like Prometheus, Grafana, and Thanos.
Benefits
- Competitive compensation
- Startup equity
- Health insurance
- Flexibility in terms of remote work