You've sampled a population — bugs from a fuzzer, words in a text, users on a page — and counted the unique bugs that turned up. How many haven't appeared yet, and what fraction of the next draw will be brand new? Two estimators answer that from the same two crumbs of evidence — the singletonsA type that showed up exactly once across your whole sample. The fewer of these you have relative to n, the better-explored the population is. and doubletonsA type that showed up exactly twice. Together with the singletons they're enough to estimate how big the iceberg is under the water. .
On the next draw there's a 3.5% chance it's a brand-new bug — your singletons-to-n ratio is the whole answer.
The estimator points at about 34 more bugs hiding beyond the 120 you've already counted — a total of 154.
| f₁, f₂The frequency-of-frequencies counts that drive both estimators. Everything else is built from these and n. | 35 · 18 | singletons · doubletons of 1,000 draws |
| Good–Turing | 96.5% | C = 1 − f₁ ⁄ n = 1 − 35 ⁄ 1,000 |
| P(new on next draw) | 3.5% | = 1 − C = f₁ ⁄ n. Sometimes called the missing mass. |
| Chao1 unseen | 34 | Û = f₁² ⁄ (2·f₂) = 35² ⁄ (2·18) |
| Chao1 total Ŝ | 154 | Ŝ = S + Û = 120 + 34. A lower bound — the true total can be higher. |