Sizing a test asks "how many users to spot this lift?" β this asks it backwards. Pin down the traffic you really have, and see the smallest effect it can detect, plus how likely you are to miss the smaller wins that are actually worth chasing.
Budget
We getvisitors per, run for.
They convert at a% baselineYour current conversion rate β the share of visitors who convert before you change anything. Every effect below is measured as a move away from this..
Call it at% powerThe chance the test detects a real effect when one genuinely exists. At 80% power, a true lift still goes unnoticed one run in five.,% confidenceHow sure you want to be that a win isn't a fluke. 95% confidence means a 5% chance of crowning a winner that's really just noise β a false positive.,.
28,000 visitors Β· 14,000 per arm at 80% power, Ξ± = 5% (2-sided)
Smallest detectable effectThe minimum detectable effect β the smallest true lift this much traffic can reliably catch at your chosen power. Anything smaller will usually slip by as a non-result.
+15.1%Relative liftA change measured against the baseline, not in raw points. A 5% rate rising to 5.25% is a +5% lift β not +0.25.
+0.76 pt5% β 5.76%
14,000Per arm
CoarseOnly a +15.1% lift or bigger will register. Smaller real gains will read as noise, and a null result won't rule them out.
If the true lift isβ¦
Chance you'd catch it
+2%
5.73%
likely miss
+5%
15.59%
likely miss
+10%
46.64%
likely miss
+15%
79.46%
marginal
+25%covered
99.51%
reliable
+40%covered
100%
reliable
two-proportion z-test, run in reverse Β· the companion to A/B Test