Why is a request always slow?

A single call that's slow only 1% of the time sounds harmless. But a modern request fans out to dozens or hundreds of services and waits for them all โ€” and it's only fast if every one of them is. Those 1% chances compound, so a rare per-call tail quietly becomes the typical per-request experience. It's the same multiply-the-gaps math as availability in series, applied to latency.

Each call isslowA 'slow' call is one above your latency target. If you care about the 99th percentile (p99), then 1% of calls are slow by definition.% of the time, and a request waits oncallsFan-out: how many independent backend calls one request makes โ€” microservices, shards, replicas โ€” and blocks on until the last returns..
100%50% โ€” a coin flip
1 call200400100 calls512
63%of requests slow

Across 100 calls, about 63% of requests hit at least one slow one โ€” roughly 1 in 2.

1 in 2requests slow
69coin-flip at
p37per-requestYour per-call tail, re-expressed per request: only this share of whole requests comes back fast. A per-call p99 can become a per-request p37.
Past the tipping point: a slow call is now the typical request, not a rare one. What looks like a clean 1% per call is a 63% chance per request โ€” your per-call p99 tail has become a per-request p37.
Fan-outA request is slow
1 call
1%
2 calls
2%
5 calls
4.9%
10 calls
9.6%
25 calls
22%
50 calls
39%
100 callsyou
63%
250 calls
92%
500 calls
99%

The fix isn't a faster average โ€” it's taming the tail itself: hedged or backup requests, tighter timeouts, fewer blocking dependencies. Cut the per-call slow rate or the fan-out and the whole curve drops, because it's their product that bites.