← Back to blog
Study Design

The Antidepressant That Worked 94% of the Time — On Paper

4 min read

In 2008, a psychiatrist and former FDA reviewer named Erick Turner did something almost nobody had done before: he compared what got published about a class of drugs to what the FDA actually had on file about the same drugs. The FDA requires every company to register a trial before it starts and to report the result, published or not, as a condition of drug approval. That gave Turner and his colleagues a rare thing in medicine — a complete list of every trial that had actually been run, with no way for a disappointing result to just quietly disappear.

They pulled the FDA's records on 74 trials of twelve antidepressants, covering more than 12,000 patients, and lined them up against what had been published in medical journals. Of the 38 trials the FDA classified as positive, 37 were published as positive — almost perfect coverage. Of the 36 trials the FDA classified as negative or questionable, only 3 were published as negative. Twenty-two were published in a way that made them read as positive anyway, and eleven never appeared in a journal at all. Read the published literature alone, and the drugs looked like they worked in roughly 94% of trials. Read the FDA's complete file, the one nobody outside a regulatory agency normally sees, and the true figure was 51% — barely better than a coin flip. Turner published the comparison in the New England Journal of Medicine, and it remains one of the cleanest demonstrations on record of how far a published literature can drift from the evidence that actually exists.

The mechanism: the published record isn't a random sample of the studies that were run

This is publication bias, sometimes called the file-drawer problem: positive, exciting, statistically significant results are far more likely to get written up, submitted, accepted, and read than negative or null ones. No single person has to do anything dishonest for this to happen at scale. A researcher with a null result reasonably decides it's not worth the months of work to write it up. A journal editor, reviewing two submissions of equal rigor, picks the one with a striking finding over the one that found nothing. Each of those choices is defensible alone. Stacked across an entire field, they quietly filter the literature so that what survives to be read is systematically more optimistic than what was actually tried.

The reason this matters more than an ordinary blind spot is that the standard way of building confidence in a result — a systematic review or meta-analysis pooling "all the evidence" — is only as good as the sample of studies it can find. A meta-analysis run on the antidepressant literature before Turner's comparison was, unknowingly, averaging over a stacked deck: nearly every positive trial and only a handful of the negative ones. The math of the meta-analysis was fine. The input to it was missing exactly the studies that would have changed the answer, and there was no way to tell from inside the published record that anything was missing at all.

Where else this shows up

  • Academic and grant-funded research. A grant-funded study that finds no effect is far less likely to get written up than one that finds a striking one, even though the null result is often just as informative to the field and to the next researcher deciding whether to run a similar study.
  • Non-profit program evaluation. Pilot programs that don't move the needle tend to get quietly discontinued rather than written up and shared, while the pilot that happened to show promise gets the case study, the annual-report writeup, and the funding renewal — leaving anyone scanning "what's been tried" with a badly skewed picture of what actually works.
  • AI and LLM evaluation. Prompting strategies, fine-tuning tricks, and model comparisons that failed to beat the baseline rarely get written up, blogged about, or added to a leaderboard writeup, while the configuration that happened to win gets shared widely — which means the visible evidence for "what works" is built almost entirely from techniques that already looked good on the data they were tried on.

What to ask for instead

Don't stop at "what does the published evidence show." Ask the question the published record can't answer about itself: how many studies were run on this question, and how many of them aren't in front of me? Trial registries, pre-registration, and requiring a written analysis plan before data collection starts exist specifically to make that question answerable — Turner's comparison only worked because the FDA's registry made the unpublished trials findable at all.

If you're building a case on "the literature says," or deciding whether a pilot's good result would hold up if the less flattering runs were included, talk to us before you treat what's published as the whole picture.

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}