← The Analysis Clinic
Analysis Clinic · Case 003 · Diagnosed
My grant needs a defensible power analysis

Can I Power My Grant on My Pilot's Effect Size?

Small pilots give effect sizes too noisy to plan on, and the pilots that look most promising are the most misleading.

Power analysisPilot studiesGrant writingR

Symptoms

What this usually means

Effect sizes from small pilots are too imprecise to anchor a power analysis. Kraemer and colleagues (2006) showed that when pilots are used this way, the most likely outcomes are that studies worth doing are abandoned and the studies that go ahead are underpowered.

The second problem has a name. Albers and Lakens (2018) call it follow-up bias: researchers go ahead only when the pilot implies a feasible sample, which happens when the pilot overestimated the effect. Even without publication bias, main studies designed this way end up with less power than planned.

A simulation in the R script below shows how large the problem is. The true effect is d = 0.40, which needs 100 participants per group for 80% power. Each pilot has 15 per group, and the main study goes ahead only if the pilot-based plan needs 150 per group or fewer:

Common causes

1. Small samples give noisy estimates

With 15 per group, the standard error of d is around 0.37, so the estimate can easily land anywhere from no effect to a very large one.

2. Selection on promising pilots

The pilots that make a full study look affordable are disproportionately the ones that overestimated the effect.

3. Biased effect size measures

Albers and Lakens recommend against using η² from a pilot ANOVA in power analyses, because it overestimates the population effect more than ω² or ε².

4. The pilot answered a different question

Pilots often use a convenience sample, an early version of the protocol, or a short follow-up, so even an accurate estimate may not carry over to the main study.

Run these checks

  1. Put a confidence interval around the pilot effect size. If it spans small to large effects, the pilot cannot anchor the power analysis.
  2. Name the smallest effect worth detecting. Base it on practical or clinical importance, cost, or what would change a decision.
  3. Check the literature for estimates from similar interventions and outcomes, allowing for the fact that published effects tend to be inflated.
  4. Use the pilot for what it measures well: recruitment, retention, adherence, and whether the procedures work (Leon et al., 2011).
  5. Check the budget ceiling: compute the power your maximum affordable sample gives across a plausible range of effect sizes.

What not to do

Treatment options

Power on the smallest effect of interest

Recommended by Albers and Lakens (2018). The study is then powered for the smallest effect that would matter, whatever the pilot showed.

Safeguard power

Perugini, Gallucci, and Costantini (2014) suggest powering on the lower limit of a 60% confidence interval around the pilot estimate, which leaves an 80% chance that the true effect is at least that large. It is more conservative, but with a very small pilot it still can't fully fix the problem.

Show power across a range of effects

A short table of power at several plausible effect sizes shows reviewers that you understand the uncertainty.

Sequential or adaptive designs

Planned interim analyses with controlled error rates, or an internal pilot that becomes part of the main sample, can use the early data without wasting it (Albers & Lakens, 2018).

Explain the justification

Whatever you choose, state the effect size, where it came from, and why it is the right basis for the sample size (Lakens, 2022).

Worked example

Using the simulation above (true d = 0.40, pilots of 15 per group, a budget ceiling of 150 per group):

Basis for the power analysisResult
Pilot effect size as observed59% of pilots lead to a main study; median power .46
Safeguard power (lower 60% CI limit)26% of pilots fit the budget; median power .57
Smallest effect of interest, d = 0.35130 per group, with 80% power for any true effect of 0.35 or more

The safeguard approach improves power but still falls short of .80 with a pilot this small, and it rules out most studies on budget grounds. Planning on the smallest effect of interest is the only option here that guarantees the intended power for the effects that matter. The R script below reproduces the simulation.

What to tell the reviewers

Our sample size is based on the smallest effect we consider practically important, d = 0.35, rather than on the effect estimated in our pilot, because estimates from pilots of this size are imprecise and tend to overstate the effects of studies that go forward (Kraemer et al., 2006; Albers & Lakens, 2018). With 130 participants per group, the study has 80% power at α = .05 for any effect of d = 0.35 or larger. The pilot established the feasibility of recruitment, retention, and the intervention protocol, which we report in the preliminary data section.

Reproduce this Case

Every number above comes from one base-R script, with no packages to install.

Download case-003-pilot-effect-size.R →

Sources

← Case 002: The Reviewer Wants a Multiple-Comparisons Correction
All Cases →

Still stuck after the first checks?

Some problems turn on the details of your design, data, or the exact reviewer comment. A free 30-minute consult can identify the next defensible step and what it would take.

Book a free consult