In 1875, Francis Galton mailed packets of sweet pea seeds to seven friends around England, sorted by the diameter of the seed that had produced them, and asked each friend to grow them out and mail back the offspring. He'd picked sweet peas on purpose — they self-fertilize, so there was only one parent's influence to track instead of two. When the harvests came back, he plotted parent seed size against offspring seed size and found something odd: the biggest parent seeds hadn't produced the biggest offspring, and the smallest hadn't produced the smallest. Every extreme had drifted back toward the middle. He worked out the slope of that drift, presented it to the Royal Institution in 1877, and started looking for the same pattern everywhere else.
He found it in people. In an 1886 paper compiling height records from 205 sets of parents and 928 of their adult children, Galton compared each child's height to the average of their parents' heights. The tallest parents in the sample had children who were taller than average — but reliably shorter than they were. The shortest parents had children shorter than average — but reliably taller than they were. Nobody in the data was obligated to end up average. It just kept happening anyway, at a fairly consistent rate: for every unit a set of parents stood above or below the population mean, their children landed roughly two-thirds as far from it. Galton called this “regression towards mediocrity.” The name got shortened to regression to the mean, and it describes something that shows up whenever a measurement is part real signal and part noise — which is nearly always.
Nothing is pulling anyone toward average. There's no force, no correction, no punishment for being unusual. What's actually happening is simpler and less mystical: any single measurement of height, test performance, or program outcome reflects a real underlying trait plus a pile of smaller, mostly random contributing factors — measurement error, timing, an unusually good or bad day, quirks of who happened to be in the sample. People who land at the extreme end of a measurement usually got there through some combination of real signal and unusually favorable (or unfavorable) noise. That noise doesn't repeat the same way twice. So the next measurement of the same person or group reflects the same real signal, but a fresh, more typical draw of noise — and the result lands closer to the middle, not because anything changed, but because the lucky (or unlucky) part of the first score was never going to happen twice in a row.
Before crediting an intervention for a group that improved, ask one question: was this group selected because it scored at an extreme, and if so, what would have happened to a similarly extreme group that got no intervention at all? That second group is what tells you how much of the “improvement” was regression to the mean and how much was real — and without it, an intervention aimed at the worst performers will look like it worked almost no matter what it actually did.
If you're designing an evaluation around a group selected for scoring unusually high or low on something, talk to us before you lock in a design that can't tell the difference between your program and ordinary statistical drift.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}