System designCore3 min

Cache failure modes

A cache hides how much load your database really gets. The dangerous moments are when it stops hiding it: a hot key expires, keys that don't exist, or the whole cache goes away.

1 · The problem

The database is sized for the misses

With a 99% hit ratio, the database sees 1% of the reads, and it's usually sized for that. Anything that drops the hit ratio multiplies its load at once.

99% hits
1,000 reads/s
95% hits
5,000 reads/s
80% hits
20,000 reads/s
Cache down
100,000 reads/s

Database reads at 100,000 requests/s. Losing the cache is a 100× jump, not a slowdown.

2 · Failure 1

Stampede: a hot key expires

A key read 5,000 times a second expires. Every request in the next few milliseconds misses, and all of them query the database for the same row, and all of them write it back.

FixHow
Request coalescingOnly one request per key loads it; the others wait for that result ("single flight")
A lock on refillThe first miss takes a short lock and refills; others serve the old value or wait briefly
Refresh earlyReload a hot key a little before it expires, with a chance that grows as expiry nears
Serve stale while refreshingKeep the expired value, return it, and refresh in the background

3 · Failure 2

Penetration: keys that don't exist

A request for an id that isn't there misses the cache, misses the database, and caches nothing, so the next one does the same. A bug or an attacker cycling through random ids sends every request to the database.

GET /event/999999Bloom filterall known idsCacheDatabasemaybe existsmissGET /event/999999Bloom filterall known idsCacheDatabasemaybe existsmiss

Fixes: cache the not found too, with a short TTL, and put a Bloom filter in front that says for sure when an id doesn't exist.

4 · Failure 3

Avalanche: many keys go at once

Load a million keys at startup with a one-hour TTL and they all expire in the same second an hour later. Or a cache node dies and its share of keys goes with it. Either way the database gets a wave it was never sized for.

FixHow
Jitter the TTLsOne hour ± 10%, so expiries spread out
Replicate the cacheA replica takes over a dead node's keys, warm
Warm before trafficFill a new or restarted cache before sending it requests
Protect the databaseRate-limit or shed load in front of it, and let a circuit breaker fail fast

5 · In a real system

Watch the cache die at peak

URL Shortener System Design

Redis goes down; Cassandra takes every read

Run it live: the cache dies at peak, every redirect falls through to Cassandra, clients retry, and the retries add to the load. Then try the fixes and see which ones keep redirects working.

Check yourself

3 questions

1. Hit ratio falls from 99% to 90% at a steady 50,000 reads/s. What happens to database reads?
2. Which fix is aimed at a stampede on one hot key?
3. Why add random jitter to TTLs?

Takeaways

Remember this

  • The database is sized for the misses; losing hits multiplies its load.
  • Stampede: coalesce refills of hot keys, refresh early, or serve stale.
  • Penetration: cache not-found results and filter unknown ids.
  • Avalanche: jitter TTLs, replicate and warm the cache, protect the database.