← Back to blog
Data LiteracyProgram Evaluation

Part of our Statistical paradoxes series · The law of small numbers

The Best Schools Were Small. So Were the Worst.

4 min read

In the 1990s, school reformers found a pattern that looked like a gift. When you lined up schools by test scores and looked at the very top of the list, the best performers were disproportionately small. Small schools meant teachers who knew every student by name, fewer kids slipping through the cracks, a sense of community a 3,000-student high school could never offer. The story wrote itself, and some of the largest foundations in the country acted on it. The Bill & Melinda Gates Foundation, the Annenberg Foundation, the Pew Charitable Trusts and others put grants totaling in the billions behind creating smaller schools and breaking big ones apart.

The results were disappointing. In his first annual letter, in January 2009, Bill Gates wrote that "many of the small schools that we invested in did not improve students' achievement in any significant way." By then the statistician Howard Wainer had already pointed out the problem, and it had nothing to do with education. Working with Harris Zwerling, he looked at 1,662 Pennsylvania schools reporting fifth-grade test scores. Of the 50 top-scoring schools, six were among the smallest 3 percent, about four times as many as you'd expect. Then he checked the other end. Nine of the 50 lowest-scoring schools were among the 50 smallest too. Small schools weren't better. They were just more likely to land at either extreme.

The mechanism: small samples swing harder

A school's average test score is an average over its students, and averages over few people bounce around more than averages over many. That's not a theory about schools. It's arithmetic that Abraham de Moivre worked out in the 1730s: the spread of an average shrinks with the square root of the number of things you averaged. A school with 25 fifth-graders can post a spectacular year because a handful of strong students happened to land in the same cohort. A school with 500 fifth-graders can't; its luck averages out. Wainer called de Moivre's result "the most dangerous equation," because ignoring it has fooled people for centuries.

Psychologists Amos Tversky and Daniel Kahneman gave the mistake a name: the law of small numbers, the gut belief that a small sample should look like the population it came from. It doesn't. Small groups produce extreme results in both directions, and if you only look at one end of a ranking, you'll find small groups waiting there with a story attached.

Wainer's other favorite example makes the point without any policy at stake. Map the US counties with the lowest kidney-cancer rates and they're mostly rural, sparsely populated, in the Midwest, South and West. Clean air, fresh food, less stress? Now map the counties with the highest rates. Also mostly rural and sparsely populated. Rural life isn't protecting people and harming them at once. Small populations just produce noisy rates.

The tell is always the same: someone looked at the top of a list, never the bottom, and never asked how big each entry was.

Where else this shows up

  • Academic and grant-funded research. Rankings of hospitals, surgeons, labs or classrooms by outcome rate are routinely led and trailed by the units with the fewest cases. If your study compares sites, clinics or teachers, plot each unit's estimate against its sample size before you name winners. The classic display is a funnel plot: the expected spread narrows as n grows, and only points outside the funnel deserve a second look. Multilevel models, which pull small units' estimates toward the overall mean, exist precisely because raw small-group averages overstate how different units really are.
  • Non-profit program evaluation. Multi-site programs face this constantly. The site with 12 participants and a 92% completion rate gets featured in the annual report; the site with 11 participants and 40% gets a performance-improvement plan. Both may be running the same program equally well. Before you expand the "model site" or cut the struggling one, check whether the difference is bigger than chance would produce at those sizes, and ideally look at more than one year. A site that's top of the list two years running is far more informative than one spectacular cohort.
  • AI and LLM evaluation. Break an eval down by subgroup, language, customer or prompt category, and the slices with 15 items will hold both your best and worst scores. A model that gets 14 of 15 right on one category and 9 of 15 on another may not differ at all once you account for that much noise. Report the item count and a confidence interval next to every slice, and be especially skeptical of "the new model wins on these three categories" when those categories are the smallest ones.

What to ask for instead

Whenever a ranking, a leaderboard or a "top performers" list lands on your desk, ask one question: how big was each group, and what does the bottom of the list look like? If the same kind of unit, small schools, rural counties, tiny eval slices, crowds both ends, you're probably looking at noise with a narrative on top. Ask for the sample sizes, a plot of results against size, and an interval around each estimate before anyone spends money on the pattern.

If you're comparing sites, programs or model variants that differ a lot in size and want to know which differences are real before you act on them, talk to us.

Related reading

Sources and further reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}