Reviewer 2 Asked for a Post Hoc Power Analysis
Why "observed power" can't answer the reviewer's real question, and the three analyses that can.
Symptoms
- Your main result is not statistically significant, and a reviewer asks whether the study was "adequately powered" or asks you to "report post hoc power."
- Your software printed an "observed power" value next to the test, and it is low.
- You are tempted to write either "the study was underpowered to detect this effect" or "power was adequate, so there is no effect."
What this usually means
The reviewer's concern is fair. A nonsignificant result from a small study is weak evidence that there is no effect, and readers deserve to know how much the study can tell them. Observed power cannot answer that concern.
Observed power is the power you get by plugging the effect you observed back in as if it were the true effect. Hoenig and Heisey (2001) showed that, for any test, observed power is a one-to-one function of the p-value. It is the p-value expressed in a different unit, so it adds no information. For a two-sided test at α = .05:
| p-value | .01 | .03 | .05 | .10 | .20 | .40 | .80 |
|---|---|---|---|---|---|---|---|
| Observed power | .73 | .58 | .50 | .38 | .25 | .13 | .06 |
A result right at p = .05 always has observed power of about .50, and every nonsignificant result has about .50 or less. "Low observed power" after a nonsignificant result is guaranteed, so reporting it says nothing the p-value didn't. It also can't be read the other way: because higher observed power always goes with a smaller p-value, "high power and no significant effect" can never support the null. Hoenig and Heisey call this the power approach paradox.
Why the request comes up
1. The reviewer wants to know whether the null result is informative
This is the real question, and it has good answers. They just aren't power calculations.
2. A checklist asks for a power analysis
Reporting standards and journal checklists ask how the sample size was determined. That means the power analysis done before data collection, not one computed afterward.
3. The software offers it
Some packages print observed power next to the test, which makes it look like a standard result.
4. "Underpowered design" and "low observed power" get mixed up
Whether a design could detect a meaningful effect is a property of the planned sample and the effect you care about. It doesn't depend on the effect you happened to observe.
Run these checks
- Find your a priori power analysis. Look in the grant, preregistration, IRB protocol, or analysis plan. If it exists, report it as planned: the effect size assumed and why, α, target power, planned sample, and achieved sample.
- Compute the confidence interval for the effect, in the outcome's own units and as a standardized effect. The interval shows which effects are still compatible with your data.
- Name the smallest effect that would matter. Use the value from your planning if you have one. Decide it from theory, practice, or cost, never from the observed result.
- Run an equivalence test against that effect (two one-sided tests, or TOST; Lakens, 2017). It asks directly whether effects as large as the one you care about can be ruled out.
- If there was no a priori analysis, report a sensitivity analysis: the smallest effect your achieved sample could detect with, say, 80% power at your α (Lakens, 2022). Present it as a description of the design. Hoenig and Heisey warn against turning a "detectable effect size" computed from the observed variability into an argument for the null, so let the confidence interval and the equivalence test carry the inference.
What not to do
- Don't report observed power, from your software or by hand.
- Don't write that the null hypothesis is supported because power was adequate.
- Don't explain a null result as "underpowered to detect the observed effect." It is true of every nonsignificant result.
- Don't choose the equivalence bound after looking at the confidence interval.
Treatment options
Report the planned power analysis
If you had one, this answers the checklist version of the request completely.
Interpret the confidence interval
Say which effects the data rule out and which they don't. A wide interval means the study is inconclusive, not negative, and saying so plainly is more convincing than any power figure.
Add an equivalence test
If the interval sits inside the range of effects too small to matter, the equivalence test lets you say so formally.
Add a design sensitivity analysis
Useful when there was no planned power analysis, as long as it's framed as what the design could detect.
Worked example
Two groups of 40 participants are compared on an outcome scored 0 to 100. The difference is 0.94 points, t(78) = 0.37, p = .71. A post hoc power calculation would report observed power of .07, which only restates p = .71.
The analyses that answer the reviewer:
- Confidence interval. The 95% CI for the difference is −4.15 to 6.04 points.
- Equivalence test. The team had decided, before the study, that differences smaller than 6 points (half a standard deviation) would not matter in practice. TOST against ±6 points gives p = .026, and the 90% CI (−3.32 to 5.20) sits inside the bounds, so differences of 6 points or more can be rejected.
- Design sensitivity. With 40 per group and α = .05, the design had 80% power for standardized effects of d = 0.63 or larger.
The R script below reproduces every number here.
What to tell the reviewer
We thank the reviewer for raising the question of whether our null result is informative. We have not added a post hoc (observed) power analysis because observed power is a direct transformation of the p-value and cannot add information about a nonsignificant result (Hoenig & Heisey, 2001). Instead, we now report the 95% confidence interval for the group difference (−4.15 to 6.04 points) and an equivalence test against the smallest difference we specified in advance as meaningful (±6 points; Lakens, 2017). The equivalence test was significant (p = .026), indicating that differences of 6 points or more are unlikely. We also report that the achieved sample provided 80% power to detect effects of d = 0.63 or larger, and we note this as a limitation for smaller effects.
Every number above comes from one base-R script, with no packages to install.
Download case-001-post-hoc-power.R →Sources
- Hoenig, J. M., & Heisey, D. M. (2001). The abuse of power: The pervasive fallacy of power calculations for data analysis. The American Statistician, 55(1), 19–24.
- Lakens, D. (2017). Equivalence tests: A practical primer for t tests, correlations, and meta-analyses. Social Psychological and Personality Science, 8(4), 355–362.
- Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1), 33267.