Senior Distributed Systems Engineer - Cache
$194k - $266k • Hybrid
Posted 5d ago
Remote Work Policy
On-site
Categories
Applied AI Engineer
About the job
Cloudflare is seeking a Senior Distributed Systems Engineer to join the Cache team. This team builds and operates the high-performance reverse proxy and caching data plane at Cloudflare's edge, built on Pingora, Cloudflare's open-source Rust framework. The role involves working on backend routing, load balancing, cache storage, globally distributed purge, Tiered Cache routing, and production observability. You will design, implement, test, roll out, and operate these systems, ensuring performance, correctness, and resilience across Cloudflare's global network. This position is ideal for individuals who enjoy tackling complex problems related to latency, correctness, and reliability in large-scale distributed systems.
Responsibilities
- Design, build, and operate the high-performance cache and proxy data plane at Cloudflare's edge.
- Contribute to Pingora and production services, improving asynchronous I/O, connection handling, memory/disk efficiency, and tail latency.
- Enhance cache correctness and performance through object placement, retention, admission/eviction policies, and range-request handling.
- Build globally distributed purge and Tiered Cache systems for quick content invalidation and efficient cache paths.
- Deliver shared platform capabilities, APIs, and rules for cache keys, TTLs, response handling, and programmatic access.
- Own features end-to-end, including design, implementation, testing, rollout, and production operation.
- Participate in on-call rotation, lead incident response, and improve reliability and operational tooling.
- Reason about failure modes and blast radius using traffic cohorts and rollback mechanisms.
- Partner with other engineering teams on cache and proxy platform extensions.
- Raise the engineering bar through code review, mentorship, and improving team standards.
Requirements
- Minimum 4 years of professional experience designing, building, and operating production software systems.
- Strong proficiency in at least one systems or backend language such as Rust, Go, C, or C++, with a willingness to work in Rust.
- Strong understanding of HTTP semantics and transport protocols (TCP, TLS, QUIC).
- Experience designing and implementing secure, resilient, high-performance distributed systems.
- Experience with Linux systems, networking, and concurrency.