← Back to blog
Data Literacy

The Tallest Parents Never Had the Tallest Children

4 min read

In 1875, Francis Galton mailed packets of sweet pea seeds to seven friends around England, sorted by the diameter of the seed that had produced them, and asked each friend to grow them out and mail back the offspring. He'd picked sweet peas on purpose — they self-fertilize, so there was only one parent's influence to track instead of two. When the harvests came back, he plotted parent seed size against offspring seed size and found something odd: the biggest parent seeds hadn't produced the biggest offspring, and the smallest hadn't produced the smallest. Every extreme had drifted back toward the middle. He worked out the slope of that drift, presented it to the Royal Institution in 1877, and started looking for the same pattern everywhere else.

He found it in people. In an 1886 paper compiling height records from 205 sets of parents and 928 of their adult children, Galton compared each child's height to the average of their parents' heights. The tallest parents in the sample had children who were taller than average — but reliably shorter than they were. The shortest parents had children shorter than average — but reliably taller than they were. Nobody in the data was obligated to end up average. It just kept happening anyway, at a fairly consistent rate: for every unit a set of parents stood above or below the population mean, their children landed roughly two-thirds as far from it. Galton called this “regression towards mediocrity.” The name got shortened to regression to the mean, and it describes something that shows up whenever a measurement is part real signal and part noise — which is nearly always.

The mechanism: an extreme score is usually part luck

Nothing is pulling anyone toward average. There's no force, no correction, no punishment for being unusual. What's actually happening is simpler and less mystical: any single measurement of height, test performance, or program outcome reflects a real underlying trait plus a pile of smaller, mostly random contributing factors — measurement error, timing, an unusually good or bad day, quirks of who happened to be in the sample. People who land at the extreme end of a measurement usually got there through some combination of real signal and unusually favorable (or unfavorable) noise. That noise doesn't repeat the same way twice. So the next measurement of the same person or group reflects the same real signal, but a fresh, more typical draw of noise — and the result lands closer to the middle, not because anything changed, but because the lucky (or unlucky) part of the first score was never going to happen twice in a row.

Where else this shows up

  • Academic and grant-funded research. A study that recruits participants specifically because they scored unusually high or low on a baseline measure — the most depressed, the most at-risk, the lowest-performing — will see many of them move toward the average at the very next measurement, treatment or no treatment, simply because they were selected on an extreme score in the first place.
  • Non-profit program evaluation. Choosing your worst-performing schools, clinics, or sites for a new intervention feels like sound targeting, but it also guarantees you selected the sites most likely to have had an unusually bad year — and unusually bad years tend to be followed by more ordinary ones, intervention or not.
  • AI and LLM evaluation. Flagging a model's worst-scoring runs on a benchmark, retraining or re-prompting, and then re-testing on a fresh sample will often show improvement even if nothing about the model actually got better — the worst runs were partly bad luck on that particular batch of prompts, and a new batch won't reproduce the same bad luck.

What to ask for instead

Before crediting an intervention for a group that improved, ask one question: was this group selected because it scored at an extreme, and if so, what would have happened to a similarly extreme group that got no intervention at all? That second group is what tells you how much of the “improvement” was regression to the mean and how much was real — and without it, an intervention aimed at the worst performers will look like it worked almost no matter what it actually did.

If you're designing an evaluation around a group selected for scoring unusually high or low on something, talk to us before you lock in a design that can't tell the difference between your program and ordinary statistical drift.

Related reading

Sources and further reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}