What's a cache really worth?

A cache hit is cheap; a miss pays full price. Your effective latency β€” or cost β€” is just the blend of the two, weighted by how often you hit. The catch is Amdahl's Law in disguise: the misses you can't eliminate play the role of the serial fraction, so they cap the win however fast your hits get. A 90% hit rate can't beat 10Γ—, no matter how quick the cache.

A cachehitA hit is served straight from the cache β€” the fast, cheap path. This is the latency a cached request returns in.returns inms, amissA miss isn't in the cache, so the request falls through to the slow source β€” a database, an API, a recomputation β€” and pays full latency.inms.
The cachehitsThe share of requests served from the cache. The single most important number here β€” and the one that's hardest to push toward 100%.% of the time.
5.9 ms effective latency Β· down from 50 ms uncached Β· 8.5Γ— faster
50Γ—
0%20%40%60%80%90%95%
perfect cache (1 Γ· miss rate)this cache
5.9 mseffective latency

At 90% hits a request averages 5.9 ms, down from 50 ms uncached β€” 8.5Γ— faster, against a floor of 1 ms.

8.5Γ—faster now
10Γ—miss-rate ceilingThe hard cap your miss rate sets: 1 Γ· (1βˆ’hit rate). Even a free, instant cache can't beat it β€” the requests you can't cache still cost full price. This is the 'serial fraction' of Amdahl's Law.
50Γ—cache ceilingThe most this cache could ever give, at a 100% hit rate: miss Γ· hit. Set by how much faster (or cheaper) the hit path is than a miss.
At 90% hits the misses are the ceiling: even an instant cache caps the win at 10Γ—. The 10% you can't cache dominate the average β€” to win more you have to cut the miss rate, not speed up hits.
Hit rateEffective latencyFaster
50%
25.5 ms
2Γ—
80%
10.8 ms
4.6Γ—
90%you
5.9 ms
8.5Γ—
95%
3.45 ms
14Γ—
99%
1.49 ms
34Γ—
99.9%
1.05 ms
48Γ—

The early gains are enormous β€” the first jump off a cold cache buys the most β€” then they flatten as the curve nears the cache ceiling. And each extra nine of hit rate is exponentially harder to win: more memory, more careful keys, longer TTLs. The misses you can never kill set the floor on effective latency, which is why past some point a faster miss path beats chasing a higher hit rate.

effective = hit Γ— hit-cost + miss Γ— miss-cost Β· Amdahl's Law for caches Β· see Amdahl's Law