Director of Infrastructure Engineering
$225k - $325k • Remote • Remote - USA • FullTime
Posted 2mo ago
Remote Work Policy
Fully remote
Employment Type
FullTime
Categories
AI Infrastructure Engineer
About the job
Runpod is seeking a Director of Infrastructure Engineering to lead and scale its core cloud and bare-metal environments. This pivotal role will oversee critical foundational layers including Site Reliability Engineering (SRE), global networking, High-Performance Computing (HPC) networks, and distributed storage engines. The ideal candidate will establish the operating rhythm, culture, and technical direction necessary to ensure Runpod's platform remains highly available, performant, and capable of meeting massive GPU computing demands. This position requires close partnership with Product Engineering, Product, and GTM leadership to support enterprise customers, focusing on delivering the fastest, most reliable, and lowest-latency infrastructure for large-scale AI workloads.
Responsibilities
- Lead multiple engineering teams responsible for SRE, networking, and storage, establishing rigorous SRE practices, SLA/SLO definitions, incident response, observability, and automated remediation.
- Oversee the design, scaling, and operation of Runpod’s global network backbone and ultra-low-latency HPC cluster networks, driving the implementation and optimization of InfiniBand and RoCE for massive GPU training workloads.
- Direct the architecture and performance tuning of highly scalable, distributed storage systems to deliver massive IOPS and throughput for deep learning tasks.
- Hire, mentor, and grow engineering managers and senior ICs, fostering a culture of ownership, operational excellence, and craft in a remote-first environment.
- Partner with Program Management and Product to forecast capacity requirements, shape technical roadmaps, and translate scale challenges into clear technical scopes and measurable outcomes.
- Drive measurable improvements in infrastructure reliability and delivery metrics, including deployment frequency, MTTR, IaC coverage, and system uptime.
- Provide architectural oversight for bare-metal provisioning, virtualization layers, network fabrics, and storage clusters to ensure seamless scalability.
- Coordinate with product delivery and platform teams to ensure infrastructure primitives are robust, well-documented, and highly available.
Requirements
- 7+ years of engineering leadership experience, including managing managers and multiple squads, with a proven record of scaling high-availability cloud environments.
- 8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms.
- Proven hands-on background or strong architectural understanding of ultra-low latency networking, including InfiniBand and/or RoCE, spine-leaf architectures, and global WAN routing protocols (BGP).
- Experience building, operating, or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF) for AI/ML I/O loads.
- Strong foundation in reliability engineering, infrastructure-as-code (Terraform, Ansible), container orchestration (Kubernetes), and modern observability stacks.
- Experience building culture, accountability, and momentum across distributed technical teams in a remote-first setting.
- Clear written and verbal communication, strong stakeholder management, and decisive leadership during high-stakes operational incidents.
- Successful completion of a background check.
Benefits
- Competitive base pay ranging from $225,000 - $325,000
- Meaningful equity in a fast-growing company
- Stock options