A cache hit is cheap; a miss pays full price. Your effective latency β or cost β is just the blend of the two, weighted by how often you hit. The catch is Amdahl's Law in disguise: the misses you can't eliminate play the role of the serial fraction, so they cap the win however fast your hits get. A 90% hit rate can't beat 10Γ, no matter how quick the cache.
At 90% hits a request averages 5.9 ms, down from 50 ms uncached β 8.5Γ faster, against a floor of 1 ms.
| Hit rate | Effective latency | Faster |
|---|---|---|
| 50% | 25.5 ms | 2Γ |
| 80% | 10.8 ms | 4.6Γ |
| 90%you | 5.9 ms | 8.5Γ |
| 95% | 3.45 ms | 14Γ |
| 99% | 1.49 ms | 34Γ |
| 99.9% | 1.05 ms | 48Γ |
The early gains are enormous β the first jump off a cold cache buys the most β then they flatten as the curve nears the cache ceiling. And each extra nine of hit rate is exponentially harder to win: more memory, more careful keys, longer TTLs. The misses you can never kill set the floor on effective latency, which is why past some point a faster miss path beats chasing a higher hit rate.