← Back to blog
Data Literacy

The Survival Rates That Improved While No One Got Better

3 min read

In 1985, three researchers at Yale — Alvan Feinstein, Daniel Sosin, and Carolyn Wells — published a comparison of lung cancer patients treated at the same institutions in two different eras: one cohort from 1953 to 1964, another from 1977. By every conventional measure, medicine had gotten better at treating the disease in the interim. The six-month survival statistics backed that up in the most striking way possible: patients in the 1977 cohort survived longer than patients in the earlier cohort, and that held true within every one of the three main disease stages taken separately. Stage-by-stage, patients were doing better. The total group was doing better. It looked like unambiguous progress.

It wasn't, in the sense anyone assumed. Between the two eras, imaging technology had improved — CT scans and other techniques could now find small metastases that used to stay invisible until much later. That didn't change what was happening inside any patient's body. It changed what doctors could see. Patients who would previously have been classified as early-stage, because their metastases hadn't yet been detected, were now correctly reassigned to a later stage the moment those same metastases showed up on a scan. Nobody's cancer had spread any further than it would have anyway. The only thing that moved was the label.

The mechanism: reshuffling the groups can move the average without moving anyone

Feinstein and his colleagues named this the Will Rogers phenomenon, after the humorist's line that when the Okies left Oklahoma and moved to California during the Depression, they raised the average intelligence of both states. The joke works because it's arithmetically true: if you move a below-average person out of a group, that group's average goes up, even though nobody in it got smarter. Move that same person into a group where they're above average, and that group's average goes up too. Both numbers improve. Nobody changed.

That's exactly what the reclassified lung cancer patients did to the statistics. A patient whose newly-detected metastasis bumped them from an early stage to a later one was, by definition, sicker than the patients who stayed behind in the early-stage group — so removing them raised the early-stage group's average survival. But that same patient was also healthier than the patients who'd already been in the later-stage group, because their disease had only just been detected rather than being advanced enough to have caused symptoms already. Adding them to the late-stage group raised its average too. Every stage looked like it was doing better. The overall population's actual prognosis, patient by patient, hadn't shifted at all — only the criteria for who got sorted into which bucket had.

Where else this shows up

  • Academic and grant-funded research. Any study comparing outcomes across eras, sites, or diagnostic criteria is vulnerable the moment the classification tool changes between measurements — a more sensitive instrument doesn't just measure the same thing more accurately, it can quietly move people between the groups being compared.
  • Non-profit program evaluation. A program that tightens its intake screening will often see its "success rate" rise the very next reporting period, not because services improved, but because the hardest cases that would have dragged the average down are now the ones being screened out before they're counted at all.
  • AI and LLM evaluation. Swap in a stricter or more sensitive content filter, a revised difficulty label for benchmark questions, or a new way of routing edge cases to a fallback model, and every remaining bucket can show an accuracy bump — without the underlying model getting better at anything. The improvement is in who got moved, not in what the system can do.

What to ask for instead

A rising average, measured stage by stage, group by group, or bucket by bucket, is not by itself evidence that anything got better for anyone. Before trusting it, ask one question: did the criteria for sorting people or cases into these groups change between the measurements being compared? If the answer is yes — a new diagnostic tool, a revised eligibility rule, a stricter or looser threshold — the comparison needs to hold the classification constant before it can tell you anything about real improvement, not just about reshuffling.

If you're comparing outcomes across a period where your measurement, screening, or classification method changed, talk to us before you report the improvement as real.

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}