← Back to blog
Program Evaluation

How to Choose Outcomes for a Program Evaluation

3 min read

Pick outcomes by working backward from the one claim you want to be able to make about your program, then choose the smallest set of measures that could actually prove that claim wrong. Most evaluations go the other way: they start from whatever data is easy to collect, track a dozen things, and end up with a report that measures a lot and demonstrates very little.

Start with the claim, not the survey

Write down the sentence you'd most like to put in front of a funder a year from now: "Families who completed the program were more likely to stay stably housed," or "students in the tutoring program read at grade level sooner." That sentence tells you what your primary outcome is. Everything else is secondary. If you can't write the sentence, the problem isn't measurement yet; it's that the program's theory of change hasn't been pinned down, and no survey will fix that.

Then name one primary outcome. Not five. With five co-equal outcomes, the odds that at least one improves by chance alone are uncomfortably high, and an experienced funder knows it. Secondary outcomes are fine and often useful. Just decide which is which before the data comes in, and write it down.

What makes an outcome worth measuring

  • It's something the program could plausibly change in the time you have. A six-week job-readiness workshop probably won't move five-year earnings in a measurable way. It might move interview callbacks within three months. Pick the outcome your program's timeline can actually reach.
  • It's an outcome, not an output. Sessions delivered, attendance, and satisfaction scores tell you the program ran. They don't tell you whether anyone ended up better off. Track them, but don't let them stand in for the result.
  • It's measured the same way for everyone. If program participants are assessed by staff who know them and a comparison group fills out an online form, any difference you find may just be the difference between the two methods.
  • It has room to move. If 90% of participants already score near the top of a scale at intake, there's nowhere for the program to show an effect. Check baseline distributions before committing to a measure.
  • It's hard to game. Once a number is tied to funding, people start optimizing for the number. Prefer measures that are hard to improve without improving the underlying thing, like verified records over self-reports and observed behavior over stated intentions.

Validated instrument or homegrown survey?

If a published, validated instrument measures what you care about, use it. It has known reliability and makes your results comparable to other programs. And don't edit it: rewording items or dropping half the questions quietly throws away the validation you chose it for. If nothing fits, a homegrown measure is fine, but pilot it first. Run it with a handful of people like your participants and ask them what they thought each question meant. You'll find at least one question being read in a way you didn't intend.

Administrative data (school records, case files, public records) is often the most underused option. It's collected anyway, it doesn't depend on who answers a survey, and it usually has a pre-program history you can use as a baseline.

Lock it in before you look

Once you've chosen, put the primary outcome, how and when it's measured, and what counts as a meaningful change into a short written plan before any outcome data is analyzed. That one page is the difference between "we found an effect" and "we found an effect on the thing we said we'd look for", and funders give the second one a great deal more credit.

If you're setting up an evaluation and aren't sure which outcome can carry the weight of your program's claim, that's the design decision we help nonprofits get right before data collection starts.

Related reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}