System designCore4 min

Caching strategies

Keep a copy of hot data somewhere faster, so most reads never reach the database. The hit ratio decides how much that buys you.

1 · The idea

A small, fast copy in front of a big, slow store

Most traffic goes to a small share of the data: today's popular events, a celebrity's profile, the home page. Keep those in memory (a local map, or a shared cache such as Redis) and answer from there. A read that finds its data in the cache is a hit; one that doesn't is a miss.

What you get depends almost entirely on the hit ratio. With a 1 ms cache in front of a 20 ms database, the average read costs:

No cache
20 ms
50% hits
11 ms
90% hits
3 ms
99% hits
1.2 ms

hit × 1 ms + miss × (1 + 20) ms. Going from 90% to 99% hits also cuts the database's read load ten times.

2 · The default pattern

Cache-aside: the application fills the cache

ApplicationCacheRedisDatabase1 · get key2 · on a miss, read3 · set key, TTLApplicationCacheRedisDatabase1 · get key2 · on a miss, read3 · set key, TTL
  1. Read the cache.

    A hit returns straight away.

  2. On a miss, read the database.

    The application, not the cache, knows how to load the data.

  3. Store it with a TTL.

    The next reader gets a hit. The expiry time bounds how stale a copy can get.

  4. On a write, update the database, then delete the key.

    The next read loads the fresh value. Deleting is safer than writing the new value into the cache: two writers can't leave the older one behind.

Cache-aside is the usual choice because the cache stays optional: if it's down, reads still work, only slower.

3 · The alternatives

Who writes to the cache, and when

PatternHow it worksGood forWatch out for
Cache-asideThe app reads the cache, loads misses from the database and fills the cacheMost read-heavy servicesThe first read of every key is a miss
Read-throughThe cache loads misses itself, through a loader you give itKeeping load logic in one placeThe cache library becomes part of the read path
Write-throughWrites go to the cache, which writes the database before acknowledgingData that's read right after it's writtenEvery write pays both latencies; cold data fills the cache
Write-behindWrites go to the cache and reach the database later, in batchesVery high write rates, such as countersWrites are lost if the cache dies before flushing
Write-aroundWrites go only to the database; reads fill the cacheData written once and rarely read againA read right after a write is a miss

4 · Where caches live

Every layer can cache

LayerExampleHit saves
BrowserCache-Control headers on images, scriptsThe whole network trip
CDN edgeStatic files and video near the userThe trip to your region
Reverse proxyWhole responses for anonymous pagesYour application servers
In the applicationA small in-process map of hot keysA network hop to the shared cache
Shared cacheRedis or MemcachedA database query
DatabaseIts own buffer pool of hot pagesA disk read

The nearer the user, the bigger the saving, and the harder it is to take a stale copy back.

5 · In a real system

A redirect that misses the cache

URL Shortener System Design

Cache-aside on the URL shortener's read path

A short link that isn't in Redis is read from Cassandra and then written into Redis, so the next click on it is a hit. Step through the miss to see each hop.

Check yourself

3 questions

1. A cache takes 1 ms and the database 20 ms. At a 90% hit ratio, what's the average read time?
2. With cache-aside, what should a write do to the cached copy?
3. Which pattern risks losing acknowledged writes if the cache crashes?

Takeaways

Remember this

  • The hit ratio is everything: each extra nine cuts both latency and database load.
  • Cache-aside is the default: read the cache, load misses, set with a TTL, delete on write.
  • Write-through, write-behind and write-around trade write latency, durability and freshness.
  • Cache at every layer that helps, closest to the user first.