Say you're ranking items by how people rated them — every rating is a thumbs up (liked it) or down (didn't). Sort by raw percentage and a single 5-of-5 item outranks one with hundreds of mostly-good ratings. The Wilson score's lower boundA lower-confidence bound on the true approval rate. Given p̂ positives out of n ratings, it estimates the worst-case true rate you can defend at your chosen confidence — so an item needs more ratings before it can claim a high rank. fixes it: it scores each item by the worst-case true rating its sample can defend, so a tiny lucky tally is penalised for its uncertainty and climbs only as the ratings pile up.
Ranked by each item's worst-case true rating — the most its likes can defend at 95% confidenceIf you ran the same vote many times, the true rate would land above this bound that often. 95% is the common default; a higher bar demands more ratings before it promotes an item.. Each bar fills solid to that defensible rate, then fades to the raw rate it can't yet prove — so the faded tail is the uncertainty penalty a small sample pays.
Top pick: Bestseller at 84.2% lower bound — not New gadget, whose 100% rests on just 5 ratings.