Part of our Statistical paradoxes series · Nonresponse bias
In the autumn of 1936, the Literary Digest ran the biggest poll anyone had ever seen. The magazine mailed about 10 million mock ballots to Americans whose names came largely from telephone directories and automobile registration lists, and roughly 2.4 million of them came back. Nobody before or since has asked that many people who they planned to vote for. The verdict was clear: Alf Landon, the Republican governor of Kansas, would take about 57% of the vote and unseat Franklin Roosevelt.
Roosevelt won about 61% of the popular vote and carried every state but two. It was one of the most lopsided elections in American history, and the largest poll in history had called it for the other guy. Meanwhile a young pollster named George Gallup, working with a sample of roughly 50,000 people, got the winner right. Within two years the Literary Digest was out of business. Two and a half million answers had turned out to be worth less than fifty thousand.
The textbook explanation for decades was the mailing list: in the Depression, people with phones and cars skewed wealthy, wealthy voters skewed Republican, so the sample was rigged before a single stamp was licked. It's a tidy story, and it's mostly wrong.
In 1988, political scientist Peverill Squire went back to a survey Gallup ran in May 1937 that asked people whether they'd received a Digest ballot and whether they'd returned it. Among people who got a ballot and mailed it back, Landon led. Among people who got a ballot and tossed it, Roosevelt led. Put the two groups together and the original list of 10 million favored Roosevelt. If everyone who received a ballot had answered, Squire concluded, the poll would at least have picked the right winner. A 2012 re-analysis by Dominic Lusinchi pushed further: phone and car owners backed Roosevelt too, and it was the nonrespondents, overwhelmingly Roosevelt voters, who did most of the damage.
That's nonresponse bias: when the people who answer differ from the people who don't in a way that's tied to what you're measuring. The plausible story is that people fed up with the New Deal had a reason to mail back a protest, while contented Roosevelt voters had better things to do. A 24% response rate isn't fatal on its own. A 24% response rate where the answer predicts whether you respond is.
And here's the part that should sting: the size of the sample did nothing to help. A bigger sample shrinks random error. It does nothing for systematic error. Two million biased answers just give you a very precise estimate of the wrong number. Ironically, the Digest had collected a clue it never used: it asked respondents how they voted in 1932. Statisticians Sharon Lohr and J. Michael Brick showed in 2017 that weighting the returned ballots by that question would have pointed to a Roosevelt majority in the Electoral College.
The next time someone hands you survey results, feedback scores, or a follow-up outcome, don't start with the sample size. Start with: who didn't answer, and what do we know about them? A credible analysis can tell you the response rate, compare responders to nonresponders on whatever was known about everyone beforehand (demographics, baseline scores, prior behavior, a 1932 vote), and show what happens to the result under a reasonable weighting or sensitivity check. If nobody can say anything about the people who stayed silent, the result tells you about the people who spoke up, and nothing more.
If you're designing a survey, planning follow-up data collection, or trying to work out how much a low response rate undermines numbers you already have, get in touch. Finding out who's missing is usually cheaper before the ballots go out than after.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}