← Back to blog
Study Design

How to Choose an Effect Size for a Power Analysis

3 min read

A power analysis needs a target effect, but choosing one is a substantive decision as well as a statistical one. Begin with the comparison your study will estimate and the difference that would matter to the people using the result. Then examine whether that difference is plausible in your setting.

Do not start by selecting “medium” from a software menu. A standardized convention may be useful for exploring scenarios, but it does not explain why your study should be designed around that value.

Define the target comparison first

Specify the outcome, population, groups, assessment time, and effect scale. A difference in final scores, a difference in change, and a treatment-by-time interaction are distinct targets. The calculation should match the planned analysis and its assumptions.

For observational research, also specify whether the target is an association or a causal contrast. An adequate sample size does not resolve confounding or make an inappropriate comparison valid.

Separate importance from expectation

The smallest difference worth detecting and the difference you expect are not necessarily the same. A prior study may suggest a large effect, while a smaller effect would still justify adopting a program. Designing only for the large estimate could leave the new study poorly equipped to detect the meaningful smaller difference.

Discuss importance in the outcome’s original units where possible. Consider practical consequences, costs, burden, and stakeholder priorities. If there is an established meaningful-difference threshold, check whether it applies to your population and whether it concerns an individual’s change or a between-group difference. Those uses are not interchangeable.

Translate to a standardized effect carefully

In a fictional study, a team decides that a four-point between-group difference on a skill score would matter. If the relevant standard deviation is ten points, the standardized mean difference is 4/10 = .40. If the standard deviation is fourteen points, the same four-point difference corresponds to approximately .286.

The substantive target has stayed the same, but the standardized effect has changed. Document where the variability estimate comes from. Paired, repeated-measures, baseline-adjusted, and clustered designs require additional inputs; a simple independent-groups effect size does not capture their entire structure.

Use prior evidence as evidence, not a default

Assess whether earlier studies used comparable outcomes, populations, intervention intensity, follow-up periods, and analyses. Examine uncertainty intervals and study quality rather than extracting one point estimate. Published estimates can be affected by selective reporting and differences in study conditions.

A small pilot’s treatment effect is often too imprecise to serve as the sole target for a definitive study. A pilot may be more informative about recruitment, measurement procedures, or variability, although those estimates also carry uncertainty.

Run a sensitivity grid

  • Vary the target difference across justified smaller and larger values.
  • Vary uncertain nuisance parameters, such as the standard deviation, baseline correlation, or intraclass correlation.
  • Include plausible missingness and distinguish recruited participants from usable observations.
  • For clustered studies, vary the number and size of clusters, not just the total participants.

Explain which scenario drives the proposed sample and what would change if assumptions are less favorable. If the required sample exceeds what is feasible, do not inflate the target effect to make the calculation fit. Reconsider the design, narrow the claim, or present the precision achievable with the available sample.

Report a reproducible justification

Record the target contrast, meaningful difference, supporting evidence, variability assumptions, significance level, power, allocation, planned analysis, and software. State how recruitment allowances were applied. Noninferiority, equivalence, and precision-based designs require their own justification; a superiority-study target difference cannot simply be reused as a margin.

Use the analysis-planning worksheet to connect the calculation to your question. See the grant power-analysis guide and the cluster-randomized trial guide. DASS can help review your design assumptions before you commit to recruitment.

Further reading

DELTA2 guidance on choosing the target difference and reporting sample-size calculations provides a framework for randomized trials. Its distinction between an important and realistic target is useful when examining your assumptions.

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}