← Back to blog
Data Literacy

The Bias That Disappeared When You Looked Closer

2 min read

In 1973, UC Berkeley's graduate admissions numbers looked like exactly what a discrimination complaint is made of: men were admitted at about 44%, women at about 35%. Nine points, thousands of applicants, and a headline that wrote itself. The university asked statisticians Peter Bickel and Kenneth O'Connell to find the departments responsible.

They couldn't find one. Broken down department by department, women were admitted at about the same rate as men in most departments, and at a noticeably higher rate in several of them. The university-wide gap was real, but it wasn't coming from bias in any admissions decision. It was coming from where people applied: women applied disproportionately to the university's most competitive departments — the ones that rejected almost everyone, of any gender — while men applied more often to departments that admitted most of their applicants. Combine a lot of applications to hard-to-get-into departments with a lot of applications to easy ones, and the overall averages pull apart even when every department treats men and women identically. Bickel, Hammel, and O'Connell published the analysis in Science in 1975; it's been a standard teaching example ever since.

This has a name, and it's not rare

Statisticians call it Simpson's Paradox: a pattern that's clearly present in the combined data can shrink, vanish, or flip entirely once you split the data by the right subgroup. It isn't a trick of arithmetic and it isn't an edge case — it shows up whenever a variable you didn't control for (here, which department someone applied to) affects both the outcome and how the groups you're comparing are sized. The paradox was never in the math. It was in deciding, usually without noticing the decision was being made, what to lump together before running the comparison.

Where else this shows up

The Berkeley case is famous because it's clean and well documented. The pattern behind it is not rare at all:

  • Multi-site program evaluation. A program can look like it's underperforming overall while outperforming at every single site, if the sites serving the hardest-to-reach populations also happen to be the largest ones.
  • AI evaluation. A model can score worse "overall" than a competitor while beating it on every individual query type, if the two test runs happened to sample different mixes of easy and hard queries.
  • Pay and hiring equity analyses. An aggregate gap by gender or race can appear or disappear depending on whether you control for role, level, or department — and which of those is the right thing to control for is a judgment call about the question you're asking, not something a p-value can settle for you.

What to ask for instead

Before trusting an aggregate comparison, ask what's being pooled to produce it, and whether the groups being compared are actually alike on anything that bears on the outcome. The fix is rarely a fancier statistic — it's reporting the breakdown next to the total, and being explicit about which one actually answers the question on the table.

If a top-line number and the subgroup numbers underneath it are telling you different stories, send us both — figuring out which one to trust is usually a short conversation, not a new study.

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}