← Back to blog
Data Literacy

The Correlation That Was Positive for States and Negative for People

3 min read

In 1950, the sociologist William S. Robinson pulled two numbers out of the 1930 U.S. Census that shouldn't have disagreed with each other, and did. Looking across the 48 states, he calculated the correlation between the percentage of a state's population that was foreign-born and the state's overall literacy rate. It came out strongly positive: +0.53. States with more immigrants had, on average, higher literacy. Read at face value, the state-level number looked like reassuring evidence about immigrant populations — more immigrants, more literacy, right there in the numbers.

Then Robinson ran the comparison a different way: not state against state, but person against person. Using individual-level census records, he calculated the same correlation — foreign-born status against literacy — this time treating each person, not each state, as one data point. The sign flipped. At the individual level, the correlation was −0.11: immigrants were, on average, somewhat less likely to be literate than native-born citizens. Same two variables. Same census. Same year. One version of the analysis said immigrants and literacy moved together. The other said the opposite. Both were computed correctly. Neither was wrong about what it measured — they simply weren't measuring the same question.

The mechanism: an aggregate answers a question you didn't ask

Robinson's explanation was almost anticlimactic once he laid it out. Immigrants in 1930 were settling disproportionately in states that already had established schools, industry, and higher native-born literacy — the Northeast and the industrial Midwest, not the rural South, where native-born literacy itself ran lower. A state's overall literacy rate is a blend of two populations living side by side: a large native-born group and a smaller immigrant group. When immigrants moved to states where the native-born population was already comparatively literate, they raised both the state's foreign-born share and its overall literacy rate at once, for two reasons that had nothing to do with each other. The state-level correlation was capturing where immigrants chose to settle. The individual-level correlation was capturing something about immigrants themselves. Averaging away the individual and looking only at the group produced a correct answer to a question about geography that got mistaken for an answer to a question about people.

Robinson called this the ecological fallacy — using group-level data to draw conclusions about the individuals inside the group — and the name stuck because the error never went away. It's simply too easy to compute a clean aggregate correlation and forget that a group-level number and a person-level number are, mathematically, different claims that happen to share the same variable names.

Where else this shows up

  • Academic and grant-funded research. A county-level or school-level analysis showing that funding and test scores move together can't, on its own, tell you whether funding helped the students who needed it — the correlation can be fully explained by which kinds of districts get more funding in the first place, with no individual student's outcome actually moving.
  • Non-profit program evaluation. A report comparing outcomes across sites or regions — "counties with our program show X% lower unemployment" — is a claim about counties, not about the people the program served. The two can point in opposite directions if the counties that adopted the program differed from the ones that didn't in ways unrelated to it.
  • AI and LLM evaluation. A model that scores well on aggregate metrics across a benchmark's categories can still be performing badly for individual users or query types within each category — the category-level average is a different quantity than the per-case outcome, and a passing aggregate score doesn't guarantee any particular case behaves the way the average suggests.

What to ask for instead

Before trusting a correlation or a group-level comparison, ask one question: is this a claim about the group, or a claim about the individuals inside it — and which one do I actually need? If a group-level number is being used to justify a decision about individual people, cases, or users, it needs individual-level data to back it up, not a more careful re-analysis of the aggregate. An ecological correlation isn't wrong; it's just answering a narrower question than the one usually being asked of it.

If you're looking at a group-level number and using it to make a claim about individuals, talk to us before the two get treated as the same evidence.

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}