JobCloudflare (Workers AI)Cloudflare (Workers AI)published Sep 9, 2026seen 1h

Senior Distributed Systems Engineer - Cache

Hybrid

Open original ↗

Captured source

source ↗
published Sep 9, 2026seen 1hcaptured 57mhttp 200method plain

Job Application for Senior Distributed Systems Engineer - Cache at Cloudflare

Back to jobs New Senior Distributed Systems Engineer - Cache Hybrid

About Us

At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company.

At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in.

Available Locations: Atlanta, GA, Austin, TX, Champaign, IL, Denver, CO, New York, NY, San Francisco, CA, Seattle, WA, Washington, DC

About the Department

Cloudflare's Engineering Team builds and runs the software that helps power millions of Internet properties. Product development covers both new features and functionality as well as scaling our existing software to meet the challenges of a massively growing customer base.

The Content Delivery group helps make the Internet faster, safer, and more reliable by building and optimizing one of the world's largest CDNs.

Role Summary

The Cache team builds and operates the high-performance reverse proxy and caching data plane at Cloudflare's edge. Built on Pingora, Cloudflare's open-source Rust framework for fast, reliable network services, this software serves cached content, routes traffic across cache tiers, and connects Cloudflare to customer origins. Its performance, correctness, and resilience directly affect request latency, origin load, and availability across Cloudflare's global network.

You will work primarily on the reverse proxy and the services that support it: backend routing and load balancing, cache storage and indexing, globally distributed purge, Tiered Cache routing, and production observability. These systems are also foundational building blocks for teams across Cloudflare, which use the proxy, caching, routing, and purge capabilities to power their own products. The team carries these capabilities through shared internal interfaces, distributed control planes, and customer-facing products including Cache Rules, Tiered Cache, Cache Reserve, Workers Cache API, and Always Online. Features ship end to end: you will help design, implement, test, roll out, and operate the systems you build.

This is a good fit if you enjoy problems where latency, correctness, and reliability are hard requirements; where changes must be introduced safely across a global fleet; and where improvements can benefit a significant portion of Internet traffic.

Role Responsibilities

Design, build, and operate the high-performance cache and proxy data plane that serves content at Cloudflare's edge and connects to customer origins

Contribute to Pingora and the production services built on it, improving asynchronous I/O, connection handling, memory and disk efficiency, and tail latency

Improve cache correctness and performance through object placement, retention, admission and eviction policies, range-request handling, and resilience to corrupt or incomplete data

Build globally distributed purge and Tiered Cache systems that invalidate content quickly, select efficient cache paths, improve hit rates, and reduce origin load

Deliver shared platform capabilities and customer-facing APIs and rules for cache keys, TTLs, response handling, purge behavior, and programmatic access from Workers

Own features end to end: write technical designs, implement services and libraries, add automated tests and observability, drive staged rollouts, and operate the resulting systems in production

Participate in the team's on-call rotation, lead incident response and post-mortems, and continually improve reliability, capacity, and operational tooling

Reason carefully about failure modes and blast radius, using traffic cohorts, measurable guardrails, and rollback mechanisms to introduce high-impact changes safely

Partner with product and engineering teams across Cloudflare that depend on or extend the cache and proxy platform

Raise the engineering bar through code review, mentorship, clear written communication, and improvements to team standards and tooling

Role Requirements (Must-Have Skills)

Minimum 4 years of professional experience designing, building, and operating production software systems

Strong proficiency in at least one systems or backend language such as Rust, Go, C, or C++, with a willingness to work in Rust

Strong understanding of HTTP semantics and transport protocols such as TCP, TLS, and QUIC

Experience designing and implementing secure, resilient, high-performance distributed systems

Experience with Linux systems, networking, concurrency, and performance analysis

Track record of production ownership, including monitoring, debugging, on-call, incident response, and ongoing reliability improvements

Experience using metrics, logs, traces, profiling, and experiments to understand production behavior and validate improvements

Strong written and verbal communication skills, including the ability to write clear technical designs and lead projects across team boundaries

Willingness to use AI-assisted engineering tools to accelerate development and investigation while remaining accountable for correctness, security, and design quality

Nice-to-Have Skills

Professional...

Excerpt shown — open the source for the full document.

Notability

notability 2.0/10

Routine job opening, not AI-specific or notable.