System designCore4 min
Caching strategies
Keep a copy of hot data somewhere faster, so most reads never reach the database. The hit ratio decides how much that buys you.
1 · The idea
A small, fast copy in front of a big, slow store
Most traffic goes to a small share of the data: today's popular events, a celebrity's profile, the home page. Keep those in memory (a local map, or a shared cache such as Redis) and answer from there. A read that finds its data in the cache is a hit; one that doesn't is a miss.
What you get depends almost entirely on the hit ratio. With a 1 ms cache in front of a 20 ms database, the average read costs:
hit × 1 ms + miss × (1 + 20) ms. Going from 90% to 99% hits also cuts the database's read load ten times.
2 · The default pattern
Cache-aside: the application fills the cache
- Read the cache.
A hit returns straight away.
- On a miss, read the database.
The application, not the cache, knows how to load the data.
- Store it with a TTL.
The next reader gets a hit. The expiry time bounds how stale a copy can get.
- On a write, update the database, then delete the key.
The next read loads the fresh value. Deleting is safer than writing the new value into the cache: two writers can't leave the older one behind.
Cache-aside is the usual choice because the cache stays optional: if it's down, reads still work, only slower.
3 · The alternatives
Who writes to the cache, and when
| Pattern | How it works | Good for | Watch out for |
|---|---|---|---|
| Cache-aside | The app reads the cache, loads misses from the database and fills the cache | Most read-heavy services | The first read of every key is a miss |
| Read-through | The cache loads misses itself, through a loader you give it | Keeping load logic in one place | The cache library becomes part of the read path |
| Write-through | Writes go to the cache, which writes the database before acknowledging | Data that's read right after it's written | Every write pays both latencies; cold data fills the cache |
| Write-behind | Writes go to the cache and reach the database later, in batches | Very high write rates, such as counters | Writes are lost if the cache dies before flushing |
| Write-around | Writes go only to the database; reads fill the cache | Data written once and rarely read again | A read right after a write is a miss |
4 · Where caches live
Every layer can cache
| Layer | Example | Hit saves |
|---|---|---|
| Browser | Cache-Control headers on images, scripts | The whole network trip |
| CDN edge | Static files and video near the user | The trip to your region |
| Reverse proxy | Whole responses for anonymous pages | Your application servers |
| In the application | A small in-process map of hot keys | A network hop to the shared cache |
| Shared cache | Redis or Memcached | A database query |
| Database | Its own buffer pool of hot pages | A disk read |
The nearer the user, the bigger the saving, and the harder it is to take a stale copy back.
5 · In a real system
A redirect that misses the cache
URL Shortener System Design
Cache-aside on the URL shortener's read path
A short link that isn't in Redis is read from Cassandra and then written into Redis, so the next click on it is a hit. Step through the miss to see each hop.
Check yourself
3 questions
Takeaways
Remember this
- The hit ratio is everything: each extra nine cuts both latency and database load.
- Cache-aside is the default: read the cache, load misses, set with a TTL, delete on write.
- Write-through, write-behind and write-around trade write latency, durability and freshness.
- Cache at every layer that helps, closest to the user first.