In 1973, UC Berkeley's graduate admissions numbers looked like exactly what a discrimination complaint is made of: men were admitted at about 44%, women at about 35%. Nine points, thousands of applicants, and a headline that wrote itself. The university asked statisticians Peter Bickel and Kenneth O'Connell to find the departments responsible.
They couldn't find one. Broken down department by department, women were admitted at about the same rate as men in most departments, and at a noticeably higher rate in several of them. The university-wide gap was real, but it wasn't coming from bias in any admissions decision. It was coming from where people applied: women applied disproportionately to the university's most competitive departments — the ones that rejected almost everyone, of any gender — while men applied more often to departments that admitted most of their applicants. Combine a lot of applications to hard-to-get-into departments with a lot of applications to easy ones, and the overall averages pull apart even when every department treats men and women identically. Bickel, Hammel, and O'Connell published the analysis in Science in 1975; it's been a standard teaching example ever since.
Statisticians call it Simpson's Paradox: a pattern that's clearly present in the combined data can shrink, vanish, or flip entirely once you split the data by the right subgroup. It isn't a trick of arithmetic and it isn't an edge case — it shows up whenever a variable you didn't control for (here, which department someone applied to) affects both the outcome and how the groups you're comparing are sized. The paradox was never in the math. It was in deciding, usually without noticing the decision was being made, what to lump together before running the comparison.
The Berkeley case is famous because it's clean and well documented. The pattern behind it is not rare at all:
Before trusting an aggregate comparison, ask what's being pooled to produce it, and whether the groups being compared are actually alike on anything that bears on the outcome. The fix is rarely a fancier statistic — it's reporting the breakdown next to the total, and being explicit about which one actually answers the question on the table.
If a top-line number and the subgroup numbers underneath it are telling you different stories, send us both — figuring out which one to trust is usually a short conversation, not a new study.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}