← Back to blog
Study DesignProgram Evaluation

The Program That Looked Like It Worked — Until Someone Checked

3 min read

In the late 1970s, a New Jersey prison began inviting teenagers with early brushes with the law to spend an afternoon inside, listening to inmates serving life sentences describe exactly what prison costs a person. The program, later known as Scared Straight, was raw and effective-looking on television and in the paperwork alike: kids came out visibly shaken, parents reported a real change at home, and the reoffending numbers the organizers tracked looked good. Within a few years, versions of it had spread to dozens of states.

Then researchers did what program staff hadn't: they found kids similar to the ones in the program who never went through it, and compared what happened to each group. When the Campbell Collaboration — a research clearinghouse that pools every rigorous trial of a given intervention — did this systematically in 2002 and updated it again in 2013, the result held across studies: kids who went through Scared Straight were more likely to reoffend than kids who got no intervention at all, by as much as 28 percentage points in some trials. A program that "obviously worked" was, on the best available evidence, making things worse.

The story was real. The comparison wasn't.

Nothing about the early success stories was fabricated. Kids really were shaken. Parents really did see a difference. What was missing was a control group: a set of similar kids, tracked the same way, for the same length of time, who didn't go through the program. Without one, there's no way to separate the program's effect from everything else that was already going to happen — kids referred to this kind of program are usually picked up at a moment of crisis, and crises tend to resolve on their own regardless of what happens next. That's regression to the mean, and it will hand you a convincing "before and after" story every time, whether or not the intervention did anything at all.

Where else this shows up

We see a version of this most months, usually from people who are careful and honest and still wrong for the same structural reason:

  • Program evaluation. A board sees pre/post numbers improve and calls it impact. Without a comparable group that didn't get the program, "improved" and "caused by us" are two different claims, and only one of them is on the table.
  • AI evaluation. A model looks great on the ten examples in the demo — often because those ten examples were chosen, consciously or not, because the model looked great on them. A held-out set the model has never seen, fixed before you look at results, is the control group here.
  • Grant reporting. Outcomes that were already trending in the right direction get attributed to the intervention that happened to start around the same time. Trend lines need a baseline period, not just an end point.

What to ask for instead

Not every project can afford a randomized controlled trial, and not every question needs one. But almost every claim of "this worked" can be pressure-tested with one question: compared to what? If the honest answer is "compared to how things were before, for the same people, with nothing else in the design to rule out other explanations," treat the result as a hypothesis worth testing properly — not as evidence you'd want to put in front of a funder, a reviewer, or a customer.

If you're staring at a result that looks too good, or trying to design a study so the result will actually mean something, talk to us before you write it up.

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}