← Back to blog
Data Literacy

The Protective Effect That Wasn't There

4 min read

In 1946, a biostatistician named Joseph Berkson pulled hospital admission records from the Mayo Clinic and found something that should have been a real medical finding. Among patients who'd been admitted to the hospital, having diabetes appeared to lower the odds of also having gallbladder disease — a negative correlation clean enough that, read at face value, it looked like diabetes was doing something protective. This wasn't a fringe result. Fourfold-table analysis — cross-tabulating two conditions and checking whether they moved together — was standard practice in clinical research at the time, and hospital records were the obvious place to run it.

Berkson didn't buy it, and instead of hunting for a biological mechanism, he worked out the statistics of how the sample itself had been built. He showed, with plain arithmetic, that a negative association could appear exactly as strong in hospital data as the one he was looking at, even if the two conditions had absolutely nothing to do with each other in the general population. The "finding" wasn't about diabetes or gallbladders at all. It was a byproduct of the fact that every single patient in the dataset had already cleared one bar: getting admitted to a hospital.

The mechanism: whatever got you into the sample can fake a relationship between traits that got you there

Picture hospital admission as a gate that either condition, on its own, is enough to open. A patient with severe diabetes doesn't need gallbladder disease to be admitted — diabetes alone is sufficient. A patient with severe gallbladder disease doesn't need diabetes, for the same reason. Now look only at the people who made it through the gate. Patients admitted because of diabetes are, on average, people whose diabetes alone was serious enough to get them there — they didn't also need a second condition to qualify. The same is true in reverse for the gallbladder patients. Because either condition alone is enough to explain a person's presence in the sample, having one makes it statistically less necessary — and therefore less likely — to also have the other, purely as an artifact of who ends up in the room. Outside the hospital, in the general population, diabetes and gallbladder disease can be entirely unrelated. Restrict your view to people who were selected by a process either condition could trigger on its own, and a negative correlation appears out of nothing.

Statisticians now call the admission variable in this setup a collider — a downstream outcome caused by both things you're studying — and the rule Berkson's arithmetic anticipated decades before the term existed is that conditioning on a collider can manufacture an association between its causes even when none exists. It's a different failure than the more famous Simpson's Paradox, where a real relationship flips or vanishes depending on how subgroups get pooled. Berkson's version invents a relationship out of two variables that were never related at all, just by restricting attention to a sample that both of them happened to help select.

Where else this shows up

  • Academic and grant-funded research. A clinical sample built from referrals, a survey sample built from volunteers, or any dataset built from people who "showed up" for some other reason inherits whatever quiet gate got them into it — and any two traits that could each independently explain a person's presence at that gate can end up looking correlated with each other, with no real relationship behind it.
  • Non-profit program evaluation. If either of two unrelated barriers — say, having reliable transportation or having paperwork ready — is enough on its own to keep someone from enrolling, looking only at who actually enrolled can make transportation access and paperwork readiness look like they trade off against each other, when in the eligible population they have nothing to do with one another.
  • AI and LLM evaluation. If a transcript gets escalated for human review whenever the model's confidence is low or whenever it trips a safety filter, studying only the escalated set can turn up a spurious negative relationship between confidence and safety compliance — an artifact of the escalation gate, not a real property of the model.

What to ask for instead

Before trusting a correlation found inside any selected group — hospitalized patients, program enrollees, escalated transcripts — ask one question: could more than one thing, on its own, have gotten someone into this sample? If the answer is yes, the correlation needs to be checked against the full, unfiltered population before it's treated as real. The math doesn't announce that it's been distorted by the selection; it just returns a clean, confident-looking number.

If you're looking at a correlation that only shows up in a referred, enrolled, or flagged subset of a larger population, talk to us before you treat it as a real relationship.

Related reading

Sources and further reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}