A timeout cuts the tail — but set it too tight and you abandon requests that would have finished, then retryA retry re-sends a request that timed out or errored. It rescues the occasional blip, but every retry is a second call: load the downstream feels on top of the originals. them, multiplying the load. Anchor it on your latency curve (median and p99), pick a timeout and a retry count, and read off how many attempts each request really costs — and the failure rate that survives.
A 80 ms timeout abandons 24.8% of attempts. Retrying up to 2 times sends 1.31× the load downstream and cuts the failure rate to 1.5% — but a doomed request can take up to 240 ms before it gives up.
Past p99 the curve is flat at 1× — almost nothing times out, so almost nothing retries. Pull the timeout below the tail and amplification climbs fast toward 3×, the cap when every request needs all 3 attempts. A timeout near the median doubles or triples your load for little gain.
| Timeout | Timed out / attempt | Amplification |
|---|---|---|
| p5050 ms | 50% fails 12.5% | 1.75× |
| p90121 ms | 10% fails 0.1% | 1.11× |
| p99250 ms | 1% fails 1 in 1000k | 1.01× |
| p99.9424 ms | 0.1% fails 1 in 1000M | 1× |
Setting the timeout at p99 abandons just 1% and barely amplifies, while still cutting the tail. Tighten it toward the median and you trade a sliver of latency for a flood of retries — the spark that lights a retry storm. Pair a generous timeout with capped retries and backoff.