By the early 2020s, preregistration had become the thing serious researchers did. It grew out of a decade of replication failures across psychology, medicine, and economics that were eventually traced back to the same quiet habit: researchers trying several outcomes, models, or subgroups after seeing the data, then reporting only the version that came out significant. Preregistration was the fix — write down what you're going to test before you see the results, and you can't quietly go looking for whichever result turns out to be significant. Journals started asking for it. Funders started rewarding it. A growing share of published studies came with a public timestamp proving the hypothesis existed before the data did.
In 2024, four economists — Abel Brodeur, Nikolai Cook, Jonathan Hartley, and Anthony Heyes — decided to check whether it was actually working. They pulled test statistics from randomized controlled trials published in fifteen leading economics journals, 15,992 of them, and ran what's called a caliper test: if researchers are honestly reporting whatever the analysis produces, results should be smoothly distributed on either side of the conventional significance threshold. If researchers are nudging borderline results over the line — trying one more control variable, one more subgroup, one more model until something clears it — you see an unnatural pileup of results bunched just past the threshold and a corresponding gap just before it. Preregistered studies, on their own, showed the same pileup as studies that were never preregistered at all. The checkbox hadn't changed the underlying behavior.
The result wasn't that preregistration is worthless — it's that most of it wasn't doing the one job it's supposed to do. A preregistration that says "we will test whether the treatment affects the outcome" is true and timestamped and almost completely unconstraining. It doesn't say which outcome, out of several plausible ones. It doesn't say which model, which subgroups, which covariates, or what happens if the data looks messier than expected. Every one of those decisions is still sitting there, waiting to be made after the researcher has already seen the results — which is exactly the discretion preregistration was supposed to remove.
The one thing that did show up as effective in the same data: preregistration paired with a genuinely detailed pre-analysis plan — one that pins down the specific outcome variable, the specific model, and the specific subgroups in advance, before anyone has looked at how the numbers came out. Studies with that level of detail showed measurably less evidence of p-hacking and publication bias. The difference wasn't whether you preregistered. It was how much of your own future discretion you actually gave up when you did it.
The gap between "we preregistered" and "we removed our own wiggle room" turns up well outside economics journals:
The useful question isn't "did you preregister?" — on its own, that answer doesn't tell you much. The useful question is: what specifically did the plan pin down before anyone saw the data? A real pre-analysis plan names the exact outcome variable, the exact model specification, the exact subgroups, and the exact significance threshold, all before data collection is complete. If a preregistration can't answer "what would this analysis have looked like if the result had come out badly?", it isn't doing the job yet.
If you're writing a pre-analysis plan and want it to actually hold up — or reviewing one someone else wrote — talk to us before you're locked into it.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}