Thundering Herd Problem (Cache Stampede) Explained
The thundering herd problem hits when a cached value expires and every waiting request floods the backend at once. Here's why it happens and how to stop it.
The thundering herd problem, also called a cache stampede, happens when a cached value expires or a shared resource becomes momentarily unavailable, and a large number of clients that were all relying on it try to recompute or reacquire it at the exact same moment — overwhelming the backend that was supposed to be protected by the cache in the first place. It’s a common failure mode for any system where caching sits in front of an expensive operation, and it tends to show up precisely when a system is under the most load, which is the worst possible time.
How it actually happens
Picture a popular page or API response cached with a five-minute expiry. For those five minutes, thousands of requests per second are served instantly from cache, and the origin database or upstream service barely notices any traffic at all. Then the cache entry expires.
The next request after expiry finds nothing in the cache, so it falls through to the origin to recompute the value — reasonable so far. The problem is that dozens, hundreds, or thousands of other requests arrive in the same narrow window before that first request finishes and repopulates the cache. Each of them also finds an empty cache entry, and each of them also falls through to the origin. A load that the cache was comfortably absorbing at near-zero cost suddenly turns into a spike of duplicate, simultaneous work hitting a backend that was sized to handle the cached, low-traffic case — not the full uncached load all at once.
If the origin operation is expensive — a complex database query, a call to a slow third-party API, an ML inference call — this spike can be severe enough to slow the origin down, which makes each of those duplicate requests take longer, which keeps the cache empty longer, which lets even more requests pile in. In the worst case the origin falls over entirely, and even after it recovers, the same stampede repeats the moment it comes back up, because all the waiting requests retry at once.
Why “just cache it” doesn’t fully solve this
Caching is the standard fix for expensive, frequently-requested operations, but a naive cache — one that simply checks “is there a value, and has it expired” — only smooths out the steady-state traffic. It does nothing to coordinate what happens at the exact moment of expiry, when every request behaves identically: “no value, go compute it.” Without explicit coordination, the cache’s presence during normal operation actively sets up the stampede, by letting a huge amount of demand build up behind a single expiring entry instead of spreading that demand out.
Mitigation techniques that actually work
Request coalescing (single-flight). When a cache miss occurs, have the first request compute the value while every subsequent concurrent request for the same key waits on that first request’s result instead of independently hitting the origin. This is often implemented with an in-memory lock or a distributed lock keyed to the cache entry, ensuring only one computation happens per expired key regardless of how many requests arrive for it.
Stale-while-revalidate. Rather than deleting a cache entry the instant it expires, keep serving the stale value to incoming requests while a single background request refreshes it. Clients get a fast, if slightly outdated, response instead of blocking on a recompute, and only one request ever does the actual work. This pattern is covered in more detail in stale-while-revalidate explained, and it’s one of the most effective fixes because it eliminates the stampede without making any client wait longer than usual.
Jittered expiry. Instead of setting every cache entry to expire in exactly the same TTL, add a small random offset to each entry’s expiry time. This spreads what would have been one large, synchronized expiry event across a wider window, so the origin sees a trickle of recompute requests instead of a spike.
Early recomputation. Refresh a cache entry proactively, slightly before it actually expires — often triggered probabilistically, with the odds of an early refresh increasing as the entry approaches its expiry. Done well, the cache almost never actually goes empty, because it’s refreshed just ahead of the deadline by whichever request happens to trigger the probabilistic check.
Backend-side protections. Even with good caching discipline, it’s worth protecting the origin itself with rate limiting and backpressure, so that a stampede that does get through — from a bug, a misconfigured TTL, or a cold cache after a deploy — degrades gracefully rather than taking the whole system down.
| Technique | What it prevents | Tradeoff |
|---|---|---|
| Request coalescing | Duplicate concurrent recomputes for the same key | Requires a lock or shared in-flight tracker |
| Stale-while-revalidate | Clients blocking on a recompute at all | Clients briefly see stale data |
| Jittered expiry | Synchronized mass expiry across many keys | Doesn’t help a single hot key expiring alone |
| Early recomputation | The cache ever actually going empty | Slightly more background work under steady state |
| Rate limiting / backpressure | Backend collapse if a stampede does occur | Doesn’t prevent the stampede, only contains the damage |
Beyond caches: the same pattern elsewhere
The thundering herd problem isn’t exclusive to caching layers. The same shape of failure shows up whenever many clients are waiting on the same event and all react to it simultaneously — a service coming back online after downtime and immediately getting hit by every client’s simultaneous reconnect attempt, or a scheduled job firing across a large fleet of workers at the exact same second. Exponential backoff with jitter exists largely to address this broader version of the problem: spreading out retries so a fleet of clients doesn’t hammer a recovering service in perfect unison.
The takeaway
A thundering herd or cache stampede happens when a cache entry expires and every request that was relying on it falls through to the origin at the same moment, turning a load the cache was comfortably absorbing into a sudden spike against a backend sized for the cached case. Request coalescing and stale-while-revalidate address the problem directly by ensuring only one request ever does the actual recompute; jittered and early expiry spread the risk out before it concentrates; and rate limiting and backpressure are the backstop for when a stampede gets through anyway. The pattern is general enough to watch for anywhere a large number of clients can end up reacting to the same event at once, not just in caching layers.
Tagged
Keep reading
Takina · · 6 min read How to Break Up Long Tasks in JavaScript
Long tasks block the main thread for 50ms or more and make pages feel frozen. How to split them with yielding, scheduler.yield(), postTask, and workers.
The Lycoris Team · · 5 min read Head-of-Line Blocking Explained: HTTP/1.1, HTTP/2, QUIC
Head-of-line blocking is when one stalled item holds up everything queued behind it. How it affects HTTP/1.1, HTTP/2 over TCP, and how QUIC fixes it.
Takina · · 5 min read Microtasks vs Macrotasks in JavaScript, Explained
Microtasks (promise callbacks) run before the next macrotask (timers, events). How the two queues are ordered, and why it matters for your code.