← Back to blog
Study Design

When Does a Pilot Need to Become a Full RCT?

2 min read

A pilot needs to become a full randomized controlled trial the moment you start using its results to answer a causal question instead of a feasibility one. If the pilot is answering "can we recruit, does the measure hold up, roughly how much does this outcome vary," it's still doing its job. The moment anyone — you, a funder, a reviewer — starts pointing at the pilot's numbers to answer "did the intervention work," the question has quietly outgrown the design that's supposed to answer it.

What a pilot can honestly tell you

A well-run pilot is built to answer operational questions, not effect questions: whether you can recruit and retain the sample you'll need, whether the outcome measure behaves sensibly and captures what you think it does, what the likely variance in that outcome looks like (the number a real power analysis needs), and whether the protocol runs the way it does on paper once real people are involved. All of that is genuinely useful. None of it is evidence that the intervention caused a change, because a pilot usually isn't powered to detect a realistic effect size and often doesn't have a comparison group to rule out everything else that could explain an improvement.

The signs you've outgrown it

  • The pilot has no comparison group. A before/after number with nothing to compare it to can't distinguish the intervention's effect from regression to the mean, seasonal change, or anything else already in motion.
  • The sample was sized for feasibility, not power. A pilot of 20 or 30 people is usually plenty to check recruitment and measurement. It's rarely enough to detect a realistic effect size with any confidence — and running a significance test on it anyway invites a false negative, a false positive, or both across a few different cuts of the data.
  • Someone is treating the pilot's effect size as the real one. Effect estimates from small, unpowered, often uncontrolled pilots tend to run larger than the true effect — a form of selection bias in which pilots get talked about in the first place. Powering a full trial off a pilot's own inflated estimate is a common way to end up underpowered again.
  • A funder, reviewer, or board has started asking "does this work" instead of "does this run." That's a category change in the question being asked, and it needs a design built to answer it.

What the next stage actually needs

Moving from pilot to full trial isn't just "more people." It typically means adding a genuine comparison or control group, randomizing assignment if a causal claim is the goal, pre-specifying the primary outcome and analysis model before data collection starts, and running a real power analysis — using the pilot's variance estimate, not its effect estimate, as the input. Skipping any one of these and simply scaling up the pilot's original design tends to produce a bigger version of the same inconclusive study, just a more expensive one to run.

If you're deciding whether your pilot data is strong enough to build a case for a full trial, or you're scoping what that trial needs to look like, this is exactly the transition we help plan.

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}