HPC Storage Engineer - West Coast
$180k - $260k • Remote • Remote - USA • FullTime
Posted 4d ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
RAG Engineer
About the job
Runpod is seeking a Senior Storage Engineer to join their remote-first Infrastructure team. This role is critical to the AI Developer Cloud, focusing on the design, scaling, and reliability of Runpod's multi-region storage ecosystem, which includes network volumes, local NVMe, and S3-compatible object storage. You will be responsible for implementing and maintaining distributed storage deployments, writing code and automation to manage them, and contributing to strategic decisions about future storage systems. This hands-on position requires a deep understanding of storage internals, performance tuning, and a proactive approach to problem-solving and continuous improvement, directly impacting the performance and reliability of AI workloads for over a million developers.
Responsibilities
- Own capacity, durability, availability, and performance of network volumes, local NVMe, and S3-compatible object storage.
- Tune the full I/O path, including device and filesystem configuration, caching, replication, and client-side mount behavior.
- Diagnose and resolve hard performance problems end-to-end.
- Lead capacity expansions, hardware refreshes, migrations, and rebalances with minimal customer disruption.
- Collaborate with networking teams to design and tune storage-dependent network paths.
- Understand and optimize RDMA/RoCE and high-speed IB/Ethernet fabrics for storage traffic.
- Write production code (Go, Python, or similar) for storage control-plane services, provisioning, data movement, and monitoring.
- Build against and extend various APIs, including control plane, S3, CSI drivers, and Kubernetes APIs.
- Automate operational tasks, moving beyond manual runbooks.
- Treat infrastructure as code, participating in code review, testing, and CI.
- Instrument the storage fleet to monitor behavior like IOPS, throughput, latency, errors, and utilization.
- Build dashboards, SLOs, and alerts to proactively identify degradation.
- Participate in an on-call rotation and drive blameless post-incident follow-through.
Requirements
- 8+ years in infrastructure, storage, or systems engineering with substantial ownership of production storage at scale.
- Deep, practical experience with at least one distributed storage system (e.g., Ceph, MinIO, Lustre, GPFS, WekaFS, VAST, ZFS-based).
- Strong Linux internals and storage-stack knowledge (block layer, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, iSCSI/NVMe-oF).
- Experience building and/or operating S3-compatible object storage services.
- Solid networking fundamentals with specific experience tuning networks for storage workloads.
- Proficiency in writing and shipping production code in Go, Python, Rust, or similar.
- Hands-on experience with observability tooling (e.g., Prometheus, Grafana, Datadog) and designing metrics.
- Proven track record of performance analysis and debugging under production pressure.
- Self-starter capable of taking a high-level goal and developing a diagnosis and plan.
- Commitment to continuous improvement and eliminating recurring toil.
- Strong sense of ownership, following problems across team boundaries.