You can't keep every trace or log line, so you sampleHead-based sampling: decide to keep or drop a trace at the start, independently and at random, with a fixed probability. Simple and cheap β but blind to whether the trace turns out to be interesting. a fraction. The catch: uniform sampling keeps the same slice of everything β so a one-off rare trace survives with exactly the sampling probability, no matter how much traffic you have. Catch rare events reliably and the storage bill follows. Find your balance.
You keep 1% of any single trace β so of one rare incident's traces, you'd hold onto just that. The event strikes about 173 times a day; you expect to capture 1.7, catching at least one on 82.4% of days β for 3.3 GiB of storage a day.
With 173 strikes a day, the curve climbs then flattens β past the knee, extra sampling buys little capture but keeps costing linearly in storage. A rarer event (or less traffic) drags the whole curve right: there's no cheap sample rate that catches a truly rare trace.
| Sample | Caught in a day | Storage / month |
|---|---|---|
| 0.1% | 15.9% β 0.17 captured | 9.9 GiB |
| 1%now | 82.4% β 1.7 captured | 98.9 GiB |
| 10% | ~100% β 17 captured | 989 GiB |
| 100% | ~100% β 173 captured | 9.7 TiB |
Capture saturates while storage keeps doubling with the rate β the two never line up. That gap is the case for tail-based samplingTail-based (or dynamic) sampling decides after a trace finishes: keep all errors and slow traces, sample the boring successes hard. It breaks the link between catching rare events and total volume.: keep 100% of errors and slow traces, sample the rest hard β so you catch the rare ones without paying to keep the routine ones.