← Back to blog
Program Evaluation

Difference-in-Differences vs. Pre/Post: Which Comparison Does Your Program Evaluation Need?

3 min read

A simple pre/post comparison tells you what changed while your program was running. It can't tell you whether the program caused that change, because it has no way to account for everything else that changed over the same window — the economy, the season, a policy shift, or just the passage of time. If there's any plausible reason the outcome could have moved without you, you need difference-in-differences, not a before-and-after number.

What a pre/post comparison actually measures

A pre/post comparison measures the total change in your outcome from before the program to after it — nothing more specific than that. If youth employment rose 8 points in the year after your workforce program launched, that 8 points includes your program's effect, plus whatever the local labor market did on its own, plus any seasonal pattern, plus the effect of any other initiative that happened to launch around the same time. A pre/post number can't separate those out, no matter how carefully you measured it. The design, not the measurement, is what's missing the information.

What difference-in-differences adds

Difference-in-differences (DiD) compares the change in your treated group to the change in a comparison group that didn't get the program over the same period — a similar population, region, or set of sites that was exposed to the same broader trends but not the intervention. The program's effect is estimated as the difference between those two changes, which cancels out whatever moved both groups regardless of the program. You don't need a randomized control group for this; a reasonably comparable group you can track over the same window is usually enough to start.

The design does rest on one assumption that's worth stating out loud to a funder or reviewer: parallel trends — that, absent the program, the treated and comparison groups would have moved together. You can't prove this outright, but you can support it: check whether the two groups were trending similarly before the program started. If their pre-program trends already diverged, DiD will misattribute that pre-existing gap to the program.

When a pre/post comparison is actually enough

  • The window is short and the outcome is stable. A two-week pre/post comparison of a process measure that never moves much on its own is a much safer bet than a year-long comparison of an outcome the economy also moves.
  • You're tracking implementation, not impact. "Did staff start following the new protocol" is a fair pre/post question. "Did client outcomes improve because of the new protocol" usually is not.
  • No comparison group of any kind exists or can be constructed. This doesn't make pre/post rigorous — it means you should report the result as descriptive, not causal, and say so plainly rather than implying a comparison you didn't run.

A quick gut-check

Ask one question before you pick a design: is there anything else, besides your program, that could plausibly have moved this outcome over the same time period? If the honest answer is "maybe," a pre/post number will get challenged the moment someone asks that same question back — and in our experience, a funder's own reviewer usually does. A comparison group, even an imperfect one, is what lets you answer it instead of hoping nobody asks.

If you're scoping an evaluation and aren't sure whether a comparison group is realistic for your program, this is where that conversation should start.

Related reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}