Defending Against the Thundering Herd
Learn how cache stampedes happen when popular keys expire, and the techniques to prevent them: request coalescing, locks, early recomputation, and jittered TTLs.
Defending Against the Thundering Herd is a free Caching Strategies: Redis + CDN + Edge Computing lesson on CoddyKit — lesson 4 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Caching Strategies: Redis + CDN + Edge Computing learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
The Thundering Herd Problem
When a hot cache key expires, every concurrent request misses at once and rushes the origin together. This cache stampede (or thundering herd) can overwhelm the database in an instant.
Why It Is Dangerous
A single popular item serving 10,000 requests/second normally hits the cache. The moment it expires, those 10,000 requests all hit the database simultaneously, often causing a spike that takes the origin down.
Cause: Synchronized Expiry
The root cause is many keys (or many requests on one key) expiring at the same moment. The fix strategies all aim to spread or serialize the resulting recomputation.
Fix 1: Request Coalescing
Let only the first request recompute the value; everyone else waits for that result. This is also called single-flight.
in_flight = {}
def get(key, compute):
if key in in_flight:
return 'waiting for in-flight result'
in_flight[key] = True
return compute()
print(get('hot', lambda: 'computed once'))Fix 2: Mutex Lock
Use a distributed lock (e.g. a Redis key with NX) so only one process recomputes. Others briefly serve stale data or retry after a short wait.
lock = None
def acquire_lock(holder):
global lock
if lock is None:
lock = holder
return True
return False
print(acquire_lock('worker-1'))
print(acquire_lock('worker-2'))Fix 3: Jittered TTL
Add randomness to each entry's TTL so they do not all expire together. A base TTL plus random jitter spreads recomputation over time.
import random
base_ttl = 300
jitter = random.randint(0, 60)
print('TTL for this entry:', base_ttl + jitter, 'seconds')Fix 4: Early Recomputation
Refresh a value before it expires. When an entry is close to its TTL, a background task (or a probabilistic check) recomputes it so it never actually goes cold for users.
Probabilistic Early Expiration
A clever trick: as a key nears expiry, give each request a small, growing probability of recomputing early. One lucky request refreshes the value while others still serve the cached copy.
import random
time_left = 5
beta = 1.0
should_refresh = random.random() < (1 / max(time_left, 1)) * beta
print('Refresh early?', should_refresh)Fix 5: Serve Stale While Revalidating
Return the expired value immediately while a background job fetches fresh data. Users get a fast (slightly stale) response and the origin sees only one refresh request.
Combining Defenses
Real systems layer these: jittered TTLs to avoid synchronized expiry, plus coalescing or a lock to serialize the inevitable misses, plus stale-while-revalidate for the best user experience.
Watch for Cache Penetration Too
A related issue is penetration: requests for keys that never exist always miss and hit the origin. Cache negative results (or use a bloom filter) so missing keys are also absorbed.
Quick Check
Which technique prevents a cache stampede by ensuring only one request recomputes the value while the rest wait for that result?
Recap
You learned to defend against the thundering herd:
- Stampedes happen when hot keys expire and many requests miss at once.
- Coalescing and locks serialize recomputation.
- Jittered TTLs and early recomputation spread the load.
- Stale-while-revalidate keeps responses fast.
Combine these to keep your origin safe under load.
Frequently asked questions
Is the “Defending Against the Thundering Herd” lesson free?
Yes — the full text of “Defending Against the Thundering Herd” is free to read here on the web, and the Caching Strategies: Redis + CDN + Edge Computing course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Caching Strategies: Redis + CDN + Edge Computing course, upgrade to CoddyKit PRO.
What will I learn in “Defending Against the Thundering Herd”?
Learn how cache stampedes happen when popular keys expire, and the techniques to prevent them: request coalescing, locks, early recomputation, and jittered TTLs. You practise Caching Strategies: Redis + CDN + Edge Computing with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start Caching Strategies: Redis + CDN + Edge Computing?
No prior experience is required. Caching Strategies: Redis + CDN + Edge Computing on CoddyKit is structured for beginners through advanced learners; this is — lesson 4 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Defending Against the Thundering Herd” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this Caching Strategies: Redis + CDN + Edge Computing lesson?
Yes. Every Caching Strategies: Redis + CDN + Edge Computing lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Common Caching Patterns
- Cache Invalidation Strategies
- Cache Eviction Policies
- Defending Against the Thundering Herd