A single call that's slow only 1% of the time sounds harmless. But a modern request fans out to dozens or hundreds of services and waits for them all โ and it's only fast if every one of them is. Those 1% chances compound, so a rare per-call tail quietly becomes the typical per-request experience. It's the same multiply-the-gaps math as availability in series, applied to latency.
Across 100 calls, about 63% of requests hit at least one slow one โ roughly 1 in 2.
| Fan-out | A request is slow |
|---|---|
| 1 call | 1% |
| 2 calls | 2% |
| 5 calls | 4.9% |
| 10 calls | 9.6% |
| 25 calls | 22% |
| 50 calls | 39% |
| 100 callsyou | 63% |
| 250 calls | 92% |
| 500 calls | 99% |
The fix isn't a faster average โ it's taming the tail itself: hedged or backup requests, tighter timeouts, fewer blocking dependencies. Cut the per-call slow rate or the fan-out and the whole curve drops, because it's their product that bites.