Part of our Statistical paradoxes series · The law of small numbers
In the 1990s, school reformers found a pattern that looked like a gift. When you lined up schools by test scores and looked at the very top of the list, the best performers were disproportionately small. Small schools meant teachers who knew every student by name, fewer kids slipping through the cracks, a sense of community a 3,000-student high school could never offer. The story wrote itself, and some of the largest foundations in the country acted on it. The Bill & Melinda Gates Foundation, the Annenberg Foundation, the Pew Charitable Trusts and others put grants totaling in the billions behind creating smaller schools and breaking big ones apart.
The results were disappointing. In his first annual letter, in January 2009, Bill Gates wrote that "many of the small schools that we invested in did not improve students' achievement in any significant way." By then the statistician Howard Wainer had already pointed out the problem, and it had nothing to do with education. Working with Harris Zwerling, he looked at 1,662 Pennsylvania schools reporting fifth-grade test scores. Of the 50 top-scoring schools, six were among the smallest 3 percent, about four times as many as you'd expect. Then he checked the other end. Nine of the 50 lowest-scoring schools were among the 50 smallest too. Small schools weren't better. They were just more likely to land at either extreme.
A school's average test score is an average over its students, and averages over few people bounce around more than averages over many. That's not a theory about schools. It's arithmetic that Abraham de Moivre worked out in the 1730s: the spread of an average shrinks with the square root of the number of things you averaged. A school with 25 fifth-graders can post a spectacular year because a handful of strong students happened to land in the same cohort. A school with 500 fifth-graders can't; its luck averages out. Wainer called de Moivre's result "the most dangerous equation," because ignoring it has fooled people for centuries.
Psychologists Amos Tversky and Daniel Kahneman gave the mistake a name: the law of small numbers, the gut belief that a small sample should look like the population it came from. It doesn't. Small groups produce extreme results in both directions, and if you only look at one end of a ranking, you'll find small groups waiting there with a story attached.
Wainer's other favorite example makes the point without any policy at stake. Map the US counties with the lowest kidney-cancer rates and they're mostly rural, sparsely populated, in the Midwest, South and West. Clean air, fresh food, less stress? Now map the counties with the highest rates. Also mostly rural and sparsely populated. Rural life isn't protecting people and harming them at once. Small populations just produce noisy rates.
The tell is always the same: someone looked at the top of a list, never the bottom, and never asked how big each entry was.
Whenever a ranking, a leaderboard or a "top performers" list lands on your desk, ask one question: how big was each group, and what does the bottom of the list look like? If the same kind of unit, small schools, rural counties, tiny eval slices, crowds both ends, you're probably looking at noise with a narrative on top. Ask for the sample sizes, a plot of results against size, and an interval around each estimate before anyone spends money on the pattern.
If you're comparing sites, programs or model variants that differ a lot in size and want to know which differences are real before you act on them, talk to us.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}