← Back to blog
Study Design

The Preregistration That Didn't Change a Thing

3 min read

By the early 2020s, preregistration had become the thing serious researchers did. It grew out of a decade of replication failures across psychology, medicine, and economics that were eventually traced back to the same quiet habit: researchers trying several outcomes, models, or subgroups after seeing the data, then reporting only the version that came out significant. Preregistration was the fix — write down what you're going to test before you see the results, and you can't quietly go looking for whichever result turns out to be significant. Journals started asking for it. Funders started rewarding it. A growing share of published studies came with a public timestamp proving the hypothesis existed before the data did.

In 2024, four economists — Abel Brodeur, Nikolai Cook, Jonathan Hartley, and Anthony Heyes — decided to check whether it was actually working. They pulled test statistics from randomized controlled trials published in fifteen leading economics journals, 15,992 of them, and ran what's called a caliper test: if researchers are honestly reporting whatever the analysis produces, results should be smoothly distributed on either side of the conventional significance threshold. If researchers are nudging borderline results over the line — trying one more control variable, one more subgroup, one more model until something clears it — you see an unnatural pileup of results bunched just past the threshold and a corresponding gap just before it. Preregistered studies, on their own, showed the same pileup as studies that were never preregistered at all. The checkbox hadn't changed the underlying behavior.

The box got checked. The discretion didn't go away.

The result wasn't that preregistration is worthless — it's that most of it wasn't doing the one job it's supposed to do. A preregistration that says "we will test whether the treatment affects the outcome" is true and timestamped and almost completely unconstraining. It doesn't say which outcome, out of several plausible ones. It doesn't say which model, which subgroups, which covariates, or what happens if the data looks messier than expected. Every one of those decisions is still sitting there, waiting to be made after the researcher has already seen the results — which is exactly the discretion preregistration was supposed to remove.

The one thing that did show up as effective in the same data: preregistration paired with a genuinely detailed pre-analysis plan — one that pins down the specific outcome variable, the specific model, and the specific subgroups in advance, before anyone has looked at how the numbers came out. Studies with that level of detail showed measurably less evidence of p-hacking and publication bias. The difference wasn't whether you preregistered. It was how much of your own future discretion you actually gave up when you did it.

Where else this shows up

The gap between "we preregistered" and "we removed our own wiggle room" turns up well outside economics journals:

  • Grant-funded research. A funder or IRB asks for a preregistered hypothesis, and a one-sentence statement of intent satisfies the requirement — while leaving the outcome measure, the analytic model, and the subgroup breakdowns entirely open for later.
  • Program evaluation. An evaluator commits publicly to "measuring impact" before the program runs, which sounds rigorous, without specifying which outcome counts as impact, over what time window, compared to what. Every one of those is a place a disappointing result can quietly become a better-looking one.
  • AI/LLM evaluation. A team commits to "evaluating before deployment," then picks the metric, the test set, and the passing threshold after they've already seen how the model performs on a few candidates — which is p-hacking with a different vocabulary.

What to ask for instead

The useful question isn't "did you preregister?" — on its own, that answer doesn't tell you much. The useful question is: what specifically did the plan pin down before anyone saw the data? A real pre-analysis plan names the exact outcome variable, the exact model specification, the exact subgroups, and the exact significance threshold, all before data collection is complete. If a preregistration can't answer "what would this analysis have looked like if the result had come out badly?", it isn't doing the job yet.

If you're writing a pre-analysis plan and want it to actually hold up — or reviewing one someone else wrote — talk to us before you're locked into it.

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}