System designCore3 min
Cache failure modes
A cache hides how much load your database really gets. The dangerous moments are when it stops hiding it: a hot key expires, keys that don't exist, or the whole cache goes away.
1 · The problem
The database is sized for the misses
With a 99% hit ratio, the database sees 1% of the reads, and it's usually sized for that. Anything that drops the hit ratio multiplies its load at once.
Database reads at 100,000 requests/s. Losing the cache is a 100× jump, not a slowdown.
2 · Failure 1
Stampede: a hot key expires
A key read 5,000 times a second expires. Every request in the next few milliseconds misses, and all of them query the database for the same row, and all of them write it back.
| Fix | How |
|---|---|
| Request coalescing | Only one request per key loads it; the others wait for that result ("single flight") |
| A lock on refill | The first miss takes a short lock and refills; others serve the old value or wait briefly |
| Refresh early | Reload a hot key a little before it expires, with a chance that grows as expiry nears |
| Serve stale while refreshing | Keep the expired value, return it, and refresh in the background |
3 · Failure 2
Penetration: keys that don't exist
A request for an id that isn't there misses the cache, misses the database, and caches nothing, so the next one does the same. A bug or an attacker cycling through random ids sends every request to the database.
Fixes: cache the not found too, with a short TTL, and put a Bloom filter in front that says for sure when an id doesn't exist.
4 · Failure 3
Avalanche: many keys go at once
Load a million keys at startup with a one-hour TTL and they all expire in the same second an hour later. Or a cache node dies and its share of keys goes with it. Either way the database gets a wave it was never sized for.
| Fix | How |
|---|---|
| Jitter the TTLs | One hour ± 10%, so expiries spread out |
| Replicate the cache | A replica takes over a dead node's keys, warm |
| Warm before traffic | Fill a new or restarted cache before sending it requests |
| Protect the database | Rate-limit or shed load in front of it, and let a circuit breaker fail fast |
5 · In a real system
Watch the cache die at peak
URL Shortener System Design
Redis goes down; Cassandra takes every read
Run it live: the cache dies at peak, every redirect falls through to Cassandra, clients retry, and the retries add to the load. Then try the fixes and see which ones keep redirects working.
Check yourself
3 questions
Takeaways
Remember this
- The database is sized for the misses; losing hits multiplies its load.
- Stampede: coalesce refills of hot keys, refresh early, or serve stale.
- Penetration: cache not-found results and filter unknown ids.
- Avalanche: jitter TTLs, replicate and warm the cache, protect the database.