← Back to blog
Data LiteracyProgram Evaluation

Part of our Statistical paradoxes series · Nonresponse bias

The Poll of 2.4 Million That Picked the Loser

5 min read

In the autumn of 1936, the Literary Digest ran the biggest poll anyone had ever seen. The magazine mailed about 10 million mock ballots to Americans whose names came largely from telephone directories and automobile registration lists, and roughly 2.4 million of them came back. Nobody before or since has asked that many people who they planned to vote for. The verdict was clear: Alf Landon, the Republican governor of Kansas, would take about 57% of the vote and unseat Franklin Roosevelt.

Roosevelt won about 61% of the popular vote and carried every state but two. It was one of the most lopsided elections in American history, and the largest poll in history had called it for the other guy. Meanwhile a young pollster named George Gallup, working with a sample of roughly 50,000 people, got the winner right. Within two years the Literary Digest was out of business. Two and a half million answers had turned out to be worth less than fifty thousand.

The problem wasn't who got a ballot. It was who sent one back.

The textbook explanation for decades was the mailing list: in the Depression, people with phones and cars skewed wealthy, wealthy voters skewed Republican, so the sample was rigged before a single stamp was licked. It's a tidy story, and it's mostly wrong.

In 1988, political scientist Peverill Squire went back to a survey Gallup ran in May 1937 that asked people whether they'd received a Digest ballot and whether they'd returned it. Among people who got a ballot and mailed it back, Landon led. Among people who got a ballot and tossed it, Roosevelt led. Put the two groups together and the original list of 10 million favored Roosevelt. If everyone who received a ballot had answered, Squire concluded, the poll would at least have picked the right winner. A 2012 re-analysis by Dominic Lusinchi pushed further: phone and car owners backed Roosevelt too, and it was the nonrespondents, overwhelmingly Roosevelt voters, who did most of the damage.

That's nonresponse bias: when the people who answer differ from the people who don't in a way that's tied to what you're measuring. The plausible story is that people fed up with the New Deal had a reason to mail back a protest, while contented Roosevelt voters had better things to do. A 24% response rate isn't fatal on its own. A 24% response rate where the answer predicts whether you respond is.

And here's the part that should sting: the size of the sample did nothing to help. A bigger sample shrinks random error. It does nothing for systematic error. Two million biased answers just give you a very precise estimate of the wrong number. Ironically, the Digest had collected a clue it never used: it asked respondents how they voted in 1932. Statisticians Sharon Lohr and J. Michael Brick showed in 2017 that weighting the returned ballots by that question would have pointed to a Roosevelt majority in the Electoral College.

Where else this shows up

  • Grant-funded and academic surveys. A faculty climate survey with a 30% response rate, or a patient-satisfaction study mailed after discharge, is only as good as its answer to one question: are the people who replied different in ways that matter? Unhappy people and delighted people both write back. The quietly indifferent middle rarely does.
  • Non-profit program evaluation. Follow-up surveys are where this bites hardest. The participants who finished the program, kept their phone number, and feel good about it are the ones who answer six months later. The ones who dropped out, moved, or had a bad experience go silent. A glowing follow-up rate can be a measure of who's reachable, not of what the program did.
  • AI and LLM evaluation. Thumbs-up/thumbs-down feedback is a voluntary-response poll. Users who rate are not a random slice of users, and their reasons for rating shift as the product changes. A satisfaction score built from opt-in feedback can climb while the silent majority quietly churns. The same goes for eval sets built from logged conversations: only the queries people bothered to finish, or bothered to flag, make it in.
  • Big data generally. Millions of rows from an app, a portal, or a sensor network feel authoritative for the same reason 2.4 million ballots did. Volume is not coverage. If the process that generates the data is linked to the outcome, more data just locks in the bias with tighter error bars.

What to ask for instead

The next time someone hands you survey results, feedback scores, or a follow-up outcome, don't start with the sample size. Start with: who didn't answer, and what do we know about them? A credible analysis can tell you the response rate, compare responders to nonresponders on whatever was known about everyone beforehand (demographics, baseline scores, prior behavior, a 1932 vote), and show what happens to the result under a reasonable weighting or sensitivity check. If nobody can say anything about the people who stayed silent, the result tells you about the people who spoke up, and nothing more.

If you're designing a survey, planning follow-up data collection, or trying to work out how much a low response rate undermines numbers you already have, get in touch. Finding out who's missing is usually cheaper before the ballots go out than after.

Related reading

  • The Plane That Never Made It Back — The same blind spot from the other side: when the data that never reaches you is exactly the data that would change the answer.
  • How Much Attrition Is Too Much? — Nonresponse at follow-up, made practical: how to judge whether the participants you lost have quietly rewritten your evaluation's result.

Sources and further reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}