Caching Strategies
Cache-aside, read-through, write-through and write-behind, how a cache stops being wrong (TTL, invalidation, versioned keys), and how to survive stampedes and hot keys.
A cache is a copy of data that is faster to reach than the source and may be wrong. Every caching strategy is an answer to three questions: who writes the copy (the application on a miss, the cache itself, or every write), how it stops being wrong (a time-to-live, an explicit invalidation, a versioned key) and what happens when many requests miss at once. Naming the strategy names its failure mode, which is the whole point of knowing the names.
Context
Caching predates the web by decades: the IBM System/360 Model 85 shipped the first CPU cache in 1968, and the vocabulary (hit, miss, eviction, write-through, write-back) comes from there. The web version arrived in 2003 when Brad Fitzpatrick wrote Memcached to keep LiveJournal's database alive, Redis followed in 2009, and Facebook's 2013 paper "Scaling Memcache" described most of the techniques still in use: leases against stampedes, regional pools, invalidation pipelines. Phil Karlton's line that there are only two hard things in computer science, cache invalidation and naming things, is old enough that nobody knows when he said it.
You have written a cache every time you typed cache.get(key) ?? (await db.query(); cache.set(...)). You have met the strategies as Rails' Rails.cache.fetch, Spring's @Cacheable, Next.js's revalidateTag, React Query's staleTime, a Cloudflare purge button, and an incident where users saw someone else's account because a per-user response was cached under a shared key.
The simplest strategy, cache-aside, is five lines:
async function getUser(id: string) {
const hit = await redis.get(`user:${id}`)
if (hit) return JSON.parse(hit) // hit: ~0.3 ms
const user = await db.user.findUnique({where: {id}}) // miss: ~5 ms
await redis.set(`user:${id}`, JSON.stringify(user), 'EX', 300) // 5 min TTL
return user
}- Hit / miss / hit ratio
- Whether the copy was there. A cache with a 50% hit ratio halves load; one at 99% removes it. Measure it or you are guessing.
- TTL
- Time to live: how long a copy is served before it is considered expired. The upper bound on staleness.
- Invalidation
- Removing or replacing a copy because the source changed, before its TTL would have.
- Eviction
- Removing copies to make room when memory is full: LRU (least recently used), LFU (least frequently used), or random.
- Stampede / thundering herd
- Many requests missing the same key at the same moment and all hitting the source together, typically right after an expiry.
- Hot key
- One key that receives a disproportionate share of traffic, so one cache node or one source row becomes the bottleneck.
Why it matters
A cache is the cheapest order-of-magnitude performance win available, which is why every system has several. It is also where the least understood bugs live: data that is stale for exactly the TTL after every edit, a "delete then read" race that leaves one key stale forever, a deploy that empties the cache and takes the database down with the resulting miss storm, and a Redis failover that turns a 10 ms endpoint into a 2 s one. Each of these has a name and a known fix.
Who writes the copy: the four patterns
| Pattern | Who fills / writes | Staleness | Failure mode |
|---|---|---|---|
| Cache-aside | App reads cache; on miss reads DB and sets cache. Writes go to DB; cache is deleted or left to expire. | Up to TTL after a write unless you invalidate | The default. Every miss is a full DB read; the delete-then-read race (below). |
| Read-through | App reads the cache only; the cache loads from DB on miss (a loader function, or a proxy like a CDN). | Same as cache-aside | Cleaner code, same semantics. The cache library owns stampede protection, or nobody does. |
| Write-through | Every write goes to the cache, which writes to DB synchronously before acking. | None for keys written through this path; reads always hit | Write latency doubles; cache holds every written key whether read or not. |
| Write-behind (write-back) | Writes go to the cache and are flushed to DB asynchronously, batched. | DB is behind the cache | Highest write throughput; a cache node loss loses unflushed writes. Only for data you can afford to lose or reconstruct. |
| Write-around | Writes go to DB only; cache fills on the next read. | Up to TTL | Avoids polluting the cache with write-once data (logs, uploads). |
How the copy stops being wrong
A TTL bounds staleness; invalidation ends it early. The obvious implementation of invalidation, delete the key when you write, has a race that can leave a key stale until its TTL, and if there is no TTL, forever.
| Technique | How | Trade-off |
|---|---|---|
| TTL | Every key expires. Short for volatile data, long for static. | Bounded staleness, zero coordination. The only fix that also covers the race above, which is why every key needs one. |
| Delete on write | After committing the DB write, delete the key. (Delete after, not before: deleting before leaves a longer window for a stale refill.) | Immediate freshness in the common case; the race in the rare case. Pair with a TTL. |
| Versioned keys | Include a version in the key: user:42:v7, or the row's updated_at. A write bumps the version; old copies are never read again and expire on their own. | No race, no delete; needs a cheap way to know the current version (often a second, tiny cached lookup). |
| Tags / dependencies | Record which keys depend on which entity; invalidate by tag (Next.js revalidateTag, Rails cache keys built from records). | Handles derived data (a page that shows five users); the dependency graph must be right. |
| Event-driven / CDC | A change-data-capture stream or the outbox publishes row changes; a subscriber invalidates or refreshes keys. | Decouples invalidation from every code path that writes; adds a pipeline with its own lag. |
Stampedes, hot keys and a cache that is down
A popular key expires. In the next 50 ms, 2,000 requests miss, and all 2,000 run the same expensive query. The database, which had been idle behind a 99% hit ratio, now receives more load than it was ever sized for, slows, and the misses pile up further. This is the stampede, and it also happens on cold start after a deploy or a cache restart, for every key at once.
async function getReport(id: string) {
const hit = await redis.get(`report:${id}`)
if (hit) return JSON.parse(hit)
// 2,000 concurrent misses → 2,000 of these
const report = await buildReport(id) // 800 ms, heavy on the DB
await redis.set(`report:${id}`, JSON.stringify(report), 'EX', 60)
return report
}async function getReport(id: string) {
const key = `report:${id}`
const hit = await redis.get(key)
if (hit) return JSON.parse(hit)
// one process rebuilds; others wait briefly or serve a stale copy
const lock = await redis.set(key + ':lock', '1', 'NX', 'PX', 5000)
if (!lock) {
const stale = await redis.get(key + ':stale') // kept for 10 min
if (stale) return JSON.parse(stale) // stale-while-revalidate
await sleep(50); return getReport(id) // or wait for the leader
}
const report = await buildReport(id)
const json = JSON.stringify(report)
await redis.multi()
.set(key, json, 'EX', 60 + Math.floor(Math.random() * 15)) // jitter
.set(key + ':stale', json, 'EX', 600)
.del(key + ':lock')
.exec()
return report
}| Problem | Fix | Note |
|---|---|---|
| Expiry stampede | Single-flight lock (one filler per key), serve stale while revalidating, jittered TTLs, probabilistic early refresh | Facebook uses "leases"; Nginx has proxy_cache_lock; most cache libraries have a coalescing option |
| Cold-start stampede | Warm the cache before taking traffic; roll deploys so nodes with warm local caches stay up; rate-limit misses to the DB | A cache flush on deploy is an outage plan |
| Hot key | Add a small in-process (L1) cache in front of Redis; replicate the key across N shards with a random suffix; move the hottest data into the app binary | Hashing spreads keys evenly, not load; one key on one node is a ceiling |
| Cache penetration | Cache negative results (the key does not exist) with a short TTL; a Bloom filter in front for sparse key spaces | Scrapers guessing IDs bypass the cache entirely otherwise |
| Cache unavailable | Short timeouts, circuit breaker, fall back to the source with a concurrency limit; degrade features rather than fail | A Redis timeout of 2 s makes every request 2 s; 50 ms and fail open |
Where caches live
| Layer | Scope | Typical TTL | Invalidation |
|---|---|---|---|
| Browser / HTTP cache | One user | Seconds to a year (hashed assets) | Cache-Control, ETag; see the HTTP caching topic |
| CDN edge | Everyone near that edge | Seconds to hours | Purge by URL or tag; stale-while-revalidate |
| In-process (LRU map) | One server instance | Seconds | TTL only, or a pub/sub message to every instance |
| Distributed (Redis, Memcached) | All instances of a service | Minutes | Delete on write, versioned keys, CDC |
| Database buffer pool / query cache | The database | Automatic | Automatic; the reason "the second query is fast" |
Pitfalls
- Keys without a TTL
Invalidation will miss a code path eventually, and a key with no TTL then stays wrong forever. A TTL is the safety net under every other mechanism; even "static" data gets one, measured in hours rather than seconds.
- Per-user data under a shared key
dashboard:summaryfilled from a request that carried user A's session is served to user B. Put the tenant or user id in the key for anything personalised, and never cache a response that varies on a cookie or Authorization header at a shared layer. - Caching the whole result set
A 5 MB JSON blob per key evicts thousands of small entries, saturates the network on every hit and is serialised on every miss. Cache the small, frequently read pieces (a user, a product), assemble from them, and keep an eye on average value size.
- Treating the cache as the source of truth
Write-behind without durable queuing, counters that only live in Redis, or a session store with no persistence: the first failover loses data and nobody can say what was lost. If the cache is the only copy, it is a database with the wrong guarantees.
- Synchronised expiry
Ten thousand keys warmed at deploy time with the same 300 s TTL expire in the same second five minutes later. Add random jitter to every TTL, and prefer refresh-ahead for the keys you know are hot.
- Not measuring the hit ratio
A cache with a 30% hit ratio is adding a round trip to 70% of requests for little benefit, and one dropping from 99% to 90% has just multiplied database load by ten. Hit ratio, evictions and p99 latency per cache are the three numbers to graph.
Interview questions
Q1Compare cache-aside and write-through. When would you pick each?
Cache-aside fills the cache lazily on a read miss and lets writes go to the database, so only data that is actually read gets cached and a write costs nothing extra, at the price of a miss and possible staleness after each write. Write-through updates the cache synchronously on every write, so reads after a write always hit and are fresh, at the price of higher write latency and caching data nobody reads. Cache-aside is the default; write-through when read-after- write freshness matters and the write rate is modest.
Q2How do you keep the cache consistent with the database when data changes?
Every key gets a TTL as the upper bound on staleness. On top of that, delete the key after the database commit, or better, use versioned keys so a write simply stops old copies being read. For derived data, tag keys with the entities they depend on and invalidate by tag. To avoid depending on every code path remembering, drive invalidation from the database change stream. And accept that plain delete-on-write has a race that only the TTL covers.
Q3What is a cache stampede and how do you prevent it?
A hot key expires and every concurrent request misses at once, all running the expensive source query together, which can take the database down that the cache was protecting. Prevention: let one request rebuild while the others wait or are served a slightly stale copy (a lock or lease per key), add jitter to TTLs so keys do not expire together, refresh hot keys shortly before expiry, and warm the cache before a fresh deployment takes traffic.
Q4Where would you put caches in a product catalogue service, and with what TTLs?
A CDN in front of product pages and images with a TTL of minutes and tag-based purge on price changes; a Redis cache-aside layer for product and inventory reads with a TTL around a minute and delete-on-write from the admin service; a small in-process LRU for the few hundred hottest products to absorb hot-key traffic; and HTTP caching headers so browsers keep static assets for a year under hashed URLs. Inventory counts stay short-lived or uncached at checkout.
Q5One product key gets 40% of all traffic. What breaks and what do you do?
Consistent hashing puts that key on one Redis node, so that node's CPU and network become the ceiling for the whole service. Add an in-process cache in every app instance for the hottest keys so most reads never reach Redis; replicate the key under several suffixed names spread across nodes and read a random one; and if it is truly static for minutes, ship it in the deploy or a local file.
Q6Redis goes down. What should the application do?
Fail open, quickly: a short client timeout, a circuit breaker that stops calling Redis after repeated failures, and a fallback to the database guarded by a concurrency limit so the miss storm cannot take it down too. Degrade features that depend on cached data rather than returning errors, and page on hit ratio dropping, not only on Redis being unreachable.
Q7LRU or LFU eviction?
LRU evicts what was not touched recently and adapts fast to changing hot sets, but a single scan of many cold keys can flush the useful ones. LFU keeps what is used often and resists scans, but is slow to forget keys that were hot last week. Redis defaults to an approximate LRU; switch to LFU when access patterns are stable and scans are common, and pick a policy that only evicts keys with a TTL if some keys must never be evicted.
- Cache-aside (app fills on miss) is the default; read-through moves the loader into the cache; write-through trades write latency for fresh hits; write-behind trades durability for write throughput.
- Every key gets a TTL, always. It is the only mechanism that bounds staleness when invalidation misses or races.
- Delete after the commit, not before, and know that the delete-then-read race exists; versioned keys and change-stream invalidation avoid it.
- Stampedes come from synchronised expiry and cold starts: single-flight per key, stale-while-revalidate, jittered TTLs, warm before traffic.
- Hashing spreads keys, not load: hot keys need an in-process layer or replication. Personalised data needs the user in the key.
- Fail open when the cache is down, with a short timeout and a concurrency cap on the source. Graph hit ratio, evictions and latency.