Member of Technical Staff - Lead, Storage

San Francisco FullTime

Posted 11h ago

Job Location

San Francisco

Tech Stack

Remote Work Policy

On-site

Employment Type

FullTime

Categories

RAG Engineer

About the job

We are seeking a strong technical lead to guide engineers in designing, building, and maintaining the novel, high-performance systems that comprise our serverless platform. You will lead the team responsible for the distributed object storage system, which underpins every container image, volume, and checkpoint on Modal. This system handles hundreds of petabytes of data, replicated across multiple cloud object stores and a CDN, cached on local NVMe across a large fleet of workers, and shared peer-to-peer within datacenters. You will set technical direction for the primitives used by other teams, balancing durability, latency, throughput, and cost, and own the roadmap for current challenges like petabyte-scale garbage collection and active-active replication, as well as future architectural bets.

Responsibilities

  • Guide engineers in designing, building, and maintaining high-performance distributed systems.
  • Lead the team responsible for the distributed object storage system.
  • Set technical direction for storage primitives, balancing durability, latency, throughput, and cost.
  • Own the roadmap for current and future storage system challenges.
  • Manage a team of 3-8 engineers while remaining hands-on across the stack.
  • Guide observability, automation, and on-call practices for the storage system.
  • Participate in the on-call rotation and respond to production incidents.

Requirements

  • 7+ years of experience writing high-quality production code.
  • 3+ years of direct people management experience, including project planning, growth, and performance conversations.
  • Experience building high-performance distributed storage or caching systems at a large scale.
  • Strong cloud skills, including deep familiarity with object storage (S3 or similar), CDNs, and their consistency, throughput, and cost characteristics.
  • Strong knowledge of low-level operating system foundations (Linux kernel, file systems, page cache, containers).
  • Experience with replication, content addressing, and consistency models in multi-region or multi-cloud systems.
  • Experience operating storage systems at scale (petabyte-scale datasets, high-throughput read/write paths, large-scale garbage collection or data migration).
  • Proven track record of setting technical direction and driving architectural decisions.
  • Willingness to participate in on-call rotations and respond to production incidents.

Get new AI jobs in your inbox

A weekly digest of the newest AI engineering roles.

© 2026 AI Job Board. All rights reserved.