In 2011, three researchers at Penn and Berkeley set out to prove something false and see if the numbers would back them up. They recruited 20 undergraduates and split them into two groups: one listened to the Beatles' "When I'm Sixty-Four," the other to an unrelated song called "Kalimba." Afterward, in what looked like a separate, unconnected survey, each student reported their birth date. The researchers ran the analysis controlling for each participant's father's age as a covariate, and got a real, statistically significant result: the students who'd heard the Beatles song were, by their own reported birth dates, a year and a half younger than the ones who hadn't — p = .04.
Nobody got younger from a pop song. That was the entire point. Simmons, Nelson, and Simonsohn built the study to demonstrate how easily a "real," publishable, statistically significant finding can be produced for something that's obviously false — using nothing but ordinary, defensible-looking choices made along the way. Which covariate to control for. When to stop collecting data. Which of several outcomes to report. Each decision looked reasonable in isolation, the kind of call any careful analyst might make without a second thought. Stacked together, they turned pure noise into a headline result, complete with a p-value under the conventional .05 threshold and a plausible-sounding write-up.
Statisticians call this p-hacking, or more precisely, the garden of forking paths — a term coined by Andrew Gelman and Eric Loken to describe what happens even without any intent to cheat. Every analysis involves small decisions: which covariates to include, whether to drop an outlier, which subgroup to look at first, whether to check the data after 20 participants or 40. Each decision is a fork in the path. A p-value only means "there's a 5% chance this result is due to chance" if the path to it was fixed before anyone looked at the data. The moment the path gets chosen after peeking — even innocently, even by a careful researcher with no agenda — the real false-positive rate stops being 5% and starts climbing.
Simmons and colleagues quantified exactly how far it climbs. Combining a handful of common, individually reasonable research habits — trying two similar outcome measures, adding one optional covariate, checking results periodically and stopping once something looks significant — pushed the true false-positive rate from a nominal 5% to well over 60% in their simulations. None of those habits requires bad faith. A researcher can make every individual call in good conscience, report the analysis honestly, and still end up with a result that has nothing to do with reality, simply because the choices were made after seeing how the data responded to each one. The paper wasn't reporting a scandal. It was reporting what normal, well-intentioned analysis looks like when the plan gets written after the results come in, rather than before.
Ask one question before trusting a p-value: was the analysis plan written down before anyone looked at the outcome data? Not whether the hypothesis sounds reasonable in hindsight — whether there's a plan, dated and specific, that predates the results, naming the outcome, the covariates, and the stopping rule in advance. If producing one means reconstructing what happened after the fact, the number you're looking at isn't measuring what it claims to. It's measuring how many forks got taken before something looked significant enough to report.
If you want a second set of eyes on an analysis plan before the data comes in — while it's still cheap to fix — reach out and let's talk it through.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}