Will your sample catch it?

You can't keep every trace or log line, so you sampleHead-based sampling: decide to keep or drop a trace at the start, independently and at random, with a fixed probability. Simple and cheap β€” but blind to whether the trace turns out to be interesting. a fraction. The catch: uniform sampling keeps the same slice of everything β€” so a one-off rare trace survives with exactly the sampling probability, no matter how much traffic you have. Catch rare events reliably and the storage bill follows. Find your balance.

Atrequests/sec, keeping% of traces atbytes each,
chasing an event that strikes 1 inrequests.
98.9 GiBkept per month

You keep 1% of any single trace β€” so of one rare incident's traces, you'd hold onto just that. The event strikes about 173 times a day; you expect to capture 1.7, catching at least one on 82.4% of days β€” for 3.3 GiB of storage a day.

82.4%caught in a dayP(catch β‰₯1) = 1 βˆ’ (1βˆ’p)^E over a day, where E is how many times the event strikes. Below ~90% you'll miss it on a meaningful share of days.
1.7captured / day
1.7Mtraces kept / day
100%0%
0.01%1%100%

With 173 strikes a day, the curve climbs then flattens β€” past the knee, extra sampling buys little capture but keeps costing linearly in storage. A rarer event (or less traffic) drags the whole curve right: there's no cheap sample rate that catches a truly rare trace.

SampleCaught in a dayStorage / month
0.1%
15.9%
β‰ˆ 0.17 captured
9.9 GiB
1%now
82.4%
β‰ˆ 1.7 captured
98.9 GiB
10%
~100%
β‰ˆ 17 captured
989 GiB
100%
~100%
β‰ˆ 173 captured
9.7 TiB

Capture saturates while storage keeps doubling with the rate β€” the two never line up. That gap is the case for tail-based samplingTail-based (or dynamic) sampling decides after a trace finishes: keep all errors and slow traces, sample the boring successes hard. It breaks the link between catching rare events and total volume.: keep 100% of errors and slow traces, sample the rest hard β€” so you catch the rare ones without paying to keep the routine ones.

Trace/log sampling Β· P(catch) = 1 βˆ’ (1 βˆ’ p)E, storage = rate Γ— p Γ— size