Can I Power My Grant on My Pilot's Effect Size?
Small pilots give effect sizes too noisy to plan on, and the pilots that look most promising are the most misleading.
Symptoms
- Your preliminary data come from a pilot with 10 to 30 participants per group, and the effect size looked encouraging.
- The power analysis in your proposal plugs that effect size in directly, and the required sample is pleasantly small.
- A reviewer or study section asks you to justify the effect size, or notes that it seems optimistic.
What this usually means
Effect sizes from small pilots are too imprecise to anchor a power analysis. Kraemer and colleagues (2006) showed that when pilots are used this way, the most likely outcomes are that studies worth doing are abandoned and the studies that go ahead are underpowered.
The second problem has a name. Albers and Lakens (2018) call it follow-up bias: researchers go ahead only when the pilot implies a feasible sample, which happens when the pilot overestimated the effect. Even without publication bias, main studies designed this way end up with less power than planned.
A simulation in the R script below shows how large the problem is. The true effect is d = 0.40, which needs 100 participants per group for 80% power. Each pilot has 15 per group, and the main study goes ahead only if the pilot-based plan needs 150 per group or fewer:
- The middle 95% of pilot estimates ran from d = −0.33 to 1.25.
- Only 59% of pilots led to a main study. The other 41% would have abandoned a study with a real effect.
- The pilots that led to a main study estimated d = 0.67 on average, against a true 0.40.
- Those main studies had a median power of .46, and 86% had power below .80.
Common causes
1. Small samples give noisy estimates
With 15 per group, the standard error of d is around 0.37, so the estimate can easily land anywhere from no effect to a very large one.
2. Selection on promising pilots
The pilots that make a full study look affordable are disproportionately the ones that overestimated the effect.
3. Biased effect size measures
Albers and Lakens recommend against using η² from a pilot ANOVA in power analyses, because it overestimates the population effect more than ω² or ε².
4. The pilot answered a different question
Pilots often use a convenience sample, an early version of the protocol, or a short follow-up, so even an accurate estimate may not carry over to the main study.
Run these checks
- Put a confidence interval around the pilot effect size. If it spans small to large effects, the pilot cannot anchor the power analysis.
- Name the smallest effect worth detecting. Base it on practical or clinical importance, cost, or what would change a decision.
- Check the literature for estimates from similar interventions and outcomes, allowing for the fact that published effects tend to be inflated.
- Use the pilot for what it measures well: recruitment, retention, adherence, and whether the procedures work (Leon et al., 2011).
- Check the budget ceiling: compute the power your maximum affordable sample gives across a plausible range of effect sizes.
What not to do
- Don't plug the pilot's effect size straight into a power calculator.
- Don't present the pilot's p-value as evidence that the intervention works.
- Don't use η² from the pilot.
- Don't run a pilot only to estimate the effect size.
Treatment options
Power on the smallest effect of interest
Recommended by Albers and Lakens (2018). The study is then powered for the smallest effect that would matter, whatever the pilot showed.
Safeguard power
Perugini, Gallucci, and Costantini (2014) suggest powering on the lower limit of a 60% confidence interval around the pilot estimate, which leaves an 80% chance that the true effect is at least that large. It is more conservative, but with a very small pilot it still can't fully fix the problem.
Show power across a range of effects
A short table of power at several plausible effect sizes shows reviewers that you understand the uncertainty.
Sequential or adaptive designs
Planned interim analyses with controlled error rates, or an internal pilot that becomes part of the main sample, can use the early data without wasting it (Albers & Lakens, 2018).
Explain the justification
Whatever you choose, state the effect size, where it came from, and why it is the right basis for the sample size (Lakens, 2022).
Worked example
Using the simulation above (true d = 0.40, pilots of 15 per group, a budget ceiling of 150 per group):
| Basis for the power analysis | Result |
|---|---|
| Pilot effect size as observed | 59% of pilots lead to a main study; median power .46 |
| Safeguard power (lower 60% CI limit) | 26% of pilots fit the budget; median power .57 |
| Smallest effect of interest, d = 0.35 | 130 per group, with 80% power for any true effect of 0.35 or more |
The safeguard approach improves power but still falls short of .80 with a pilot this small, and it rules out most studies on budget grounds. Planning on the smallest effect of interest is the only option here that guarantees the intended power for the effects that matter. The R script below reproduces the simulation.
What to tell the reviewers
Our sample size is based on the smallest effect we consider practically important, d = 0.35, rather than on the effect estimated in our pilot, because estimates from pilots of this size are imprecise and tend to overstate the effects of studies that go forward (Kraemer et al., 2006; Albers & Lakens, 2018). With 130 participants per group, the study has 80% power at α = .05 for any effect of d = 0.35 or larger. The pilot established the feasibility of recruitment, retention, and the intervention protocol, which we report in the preliminary data section.
Every number above comes from one base-R script, with no packages to install.
Download case-003-pilot-effect-size.R →Sources
- Albers, C., & Lakens, D. (2018). When power analyses based on pilot data are biased: Inaccurate effect size estimators and follow-up bias. Journal of Experimental Social Psychology, 74, 187–195.
- Kraemer, H. C., Mintz, J., Noda, A., Tinklenberg, J., & Yesavage, J. A. (2006). Caution regarding the use of pilot studies to guide power calculations for study proposals. Archives of General Psychiatry, 63(5), 484–489.
- Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1), 33267.
- Leon, A. C., Davis, L. L., & Kraemer, H. C. (2011). The role and interpretation of pilot studies in clinical research. Journal of Psychiatric Research, 45(5), 626–629.
- Perugini, M., Gallucci, M., & Costantini, G. (2014). Safeguard power as a protection against imprecise power estimates. Perspectives on Psychological Science, 9(3), 319–332.