← Back to blog
Study Design

The Beatles Song That Made People Younger

4 min read

In 2011, three researchers at Penn and Berkeley set out to prove something false and see if the numbers would back them up. They recruited 20 undergraduates and split them into two groups: one listened to the Beatles' "When I'm Sixty-Four," the other to an unrelated song called "Kalimba." Afterward, in what looked like a separate, unconnected survey, each student reported their birth date. The researchers ran the analysis controlling for each participant's father's age as a covariate, and got a real, statistically significant result: the students who'd heard the Beatles song were, by their own reported birth dates, a year and a half younger than the ones who hadn't — p = .04.

Nobody got younger from a pop song. That was the entire point. Simmons, Nelson, and Simonsohn built the study to demonstrate how easily a "real," publishable, statistically significant finding can be produced for something that's obviously false — using nothing but ordinary, defensible-looking choices made along the way. Which covariate to control for. When to stop collecting data. Which of several outcomes to report. Each decision looked reasonable in isolation, the kind of call any careful analyst might make without a second thought. Stacked together, they turned pure noise into a headline result, complete with a p-value under the conventional .05 threshold and a plausible-sounding write-up.

The mechanism: the garden of forking paths

Statisticians call this p-hacking, or more precisely, the garden of forking paths — a term coined by Andrew Gelman and Eric Loken to describe what happens even without any intent to cheat. Every analysis involves small decisions: which covariates to include, whether to drop an outlier, which subgroup to look at first, whether to check the data after 20 participants or 40. Each decision is a fork in the path. A p-value only means "there's a 5% chance this result is due to chance" if the path to it was fixed before anyone looked at the data. The moment the path gets chosen after peeking — even innocently, even by a careful researcher with no agenda — the real false-positive rate stops being 5% and starts climbing.

Simmons and colleagues quantified exactly how far it climbs. Combining a handful of common, individually reasonable research habits — trying two similar outcome measures, adding one optional covariate, checking results periodically and stopping once something looks significant — pushed the true false-positive rate from a nominal 5% to well over 60% in their simulations. None of those habits requires bad faith. A researcher can make every individual call in good conscience, report the analysis honestly, and still end up with a result that has nothing to do with reality, simply because the choices were made after seeing how the data responded to each one. The paper wasn't reporting a scandal. It was reporting what normal, well-intentioned analysis looks like when the plan gets written after the results come in, rather than before.

Where else this shows up

  • A grant-funded study chasing a stubborn null result home. A study funded to find an effect that comes back flat gets reanalyzed — a covariate added here, a subgroup split there — until something clears p < .05, and only that path makes it into the paper.
  • A program evaluation tracking a dozen outcome measures. When an evaluation has that many chances for something to move, checking all of them until one shows improvement reproduces this exact problem, with honest people and real data.
  • An AI eval re-run until the number looks right. Re-scoring a benchmark with a different prompt, seed, or rubric until the reported accuracy looks favorable is the same forking-paths problem in a different outfit — and it's just as invisible from the final number alone.

What to ask for instead

Ask one question before trusting a p-value: was the analysis plan written down before anyone looked at the outcome data? Not whether the hypothesis sounds reasonable in hindsight — whether there's a plan, dated and specific, that predates the results, naming the outcome, the covariates, and the stopping rule in advance. If producing one means reconstructing what happened after the fact, the number you're looking at isn't measuring what it claims to. It's measuring how many forks got taken before something looked significant enough to report.

If you want a second set of eyes on an analysis plan before the data comes in — while it's still cheap to fix — reach out and let's talk it through.

Related reading

Sources and further reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}