What does your timeout cost?

A timeout cuts the tail — but set it too tight and you abandon requests that would have finished, then retryA retry re-sends a request that timed out or errored. It rescues the occasional blip, but every retry is a second call: load the downstream feels on top of the originals. them, multiplying the load. Anchor it on your latency curve (median and p99), pick a timeout and a retry count, and read off how many attempts each request really costs — and the failure rate that survives.

With ams median andms p99The 99th-percentile latency: 1 in 100 requests is slower than this. Together with the median it fixes a log-normal tail., ams timeout and up toretries.
1.31×retry amplification

A 80 ms timeout abandons 24.8% of attempts. Retrying up to 2 times sends 1.31× the load downstream and cuts the failure rate to 1.5% — but a doomed request can take up to 240 ms before it gives up.

24.8%timed out / attemptP(latency > timeout) for one attempt, from the fitted log-normal. This is the share you abandon and then retry.
1.5%fails after retriesf^(R+1): every attempt, original and retries, times out. The SLO the timeout actually delivers.
240 msworst-case wait
p50p99
16.5 ms80 ms490 ms

Past p99 the curve is flat at — almost nothing times out, so almost nothing retries. Pull the timeout below the tail and amplification climbs fast toward , the cap when every request needs all 3 attempts. A timeout near the median doubles or triples your load for little gain.

TimeoutTimed out / attemptAmplification
p5050 ms
50%
fails 12.5%
1.75×
p90121 ms
10%
fails 0.1%
1.11×
p99250 ms
1%
fails 1 in 1000k
1.01×
p99.9424 ms
0.1%
fails 1 in 1000M

Setting the timeout at p99 abandons just 1% and barely amplifies, while still cutting the tail. Tighten it toward the median and you trade a sliver of latency for a flood of retries — the spark that lights a retry storm. Pair a generous timeout with capped retries and backoff.

Timeout & retry budget · amplification = (1 − f^(R+1)) / (1 − f), f = P(latency > timeout)