← The Analysis Clinic
Analysis Clinic · Case 001 · Diagnosed
A reviewer or editor challenged my analysis

Reviewer 2 Asked for a Post Hoc Power Analysis

Why "observed power" can't answer the reviewer's real question, and the three analyses that can.

Power analysisEquivalence testingR

Symptoms

What this usually means

The reviewer's concern is fair. A nonsignificant result from a small study is weak evidence that there is no effect, and readers deserve to know how much the study can tell them. Observed power cannot answer that concern.

Observed power is the power you get by plugging the effect you observed back in as if it were the true effect. Hoenig and Heisey (2001) showed that, for any test, observed power is a one-to-one function of the p-value. It is the p-value expressed in a different unit, so it adds no information. For a two-sided test at α = .05:

p-value.01.03.05.10.20.40.80
Observed power.73.58.50.38.25.13.06

A result right at p = .05 always has observed power of about .50, and every nonsignificant result has about .50 or less. "Low observed power" after a nonsignificant result is guaranteed, so reporting it says nothing the p-value didn't. It also can't be read the other way: because higher observed power always goes with a smaller p-value, "high power and no significant effect" can never support the null. Hoenig and Heisey call this the power approach paradox.

Why the request comes up

1. The reviewer wants to know whether the null result is informative

This is the real question, and it has good answers. They just aren't power calculations.

2. A checklist asks for a power analysis

Reporting standards and journal checklists ask how the sample size was determined. That means the power analysis done before data collection, not one computed afterward.

3. The software offers it

Some packages print observed power next to the test, which makes it look like a standard result.

4. "Underpowered design" and "low observed power" get mixed up

Whether a design could detect a meaningful effect is a property of the planned sample and the effect you care about. It doesn't depend on the effect you happened to observe.

Run these checks

  1. Find your a priori power analysis. Look in the grant, preregistration, IRB protocol, or analysis plan. If it exists, report it as planned: the effect size assumed and why, α, target power, planned sample, and achieved sample.
  2. Compute the confidence interval for the effect, in the outcome's own units and as a standardized effect. The interval shows which effects are still compatible with your data.
  3. Name the smallest effect that would matter. Use the value from your planning if you have one. Decide it from theory, practice, or cost, never from the observed result.
  4. Run an equivalence test against that effect (two one-sided tests, or TOST; Lakens, 2017). It asks directly whether effects as large as the one you care about can be ruled out.
  5. If there was no a priori analysis, report a sensitivity analysis: the smallest effect your achieved sample could detect with, say, 80% power at your α (Lakens, 2022). Present it as a description of the design. Hoenig and Heisey warn against turning a "detectable effect size" computed from the observed variability into an argument for the null, so let the confidence interval and the equivalence test carry the inference.

What not to do

Treatment options

Report the planned power analysis

If you had one, this answers the checklist version of the request completely.

Interpret the confidence interval

Say which effects the data rule out and which they don't. A wide interval means the study is inconclusive, not negative, and saying so plainly is more convincing than any power figure.

Add an equivalence test

If the interval sits inside the range of effects too small to matter, the equivalence test lets you say so formally.

Add a design sensitivity analysis

Useful when there was no planned power analysis, as long as it's framed as what the design could detect.

Worked example

Two groups of 40 participants are compared on an outcome scored 0 to 100. The difference is 0.94 points, t(78) = 0.37, p = .71. A post hoc power calculation would report observed power of .07, which only restates p = .71.

The analyses that answer the reviewer:

The R script below reproduces every number here.

What to tell the reviewer

We thank the reviewer for raising the question of whether our null result is informative. We have not added a post hoc (observed) power analysis because observed power is a direct transformation of the p-value and cannot add information about a nonsignificant result (Hoenig & Heisey, 2001). Instead, we now report the 95% confidence interval for the group difference (−4.15 to 6.04 points) and an equivalence test against the smallest difference we specified in advance as meaningful (±6 points; Lakens, 2017). The equivalence test was significant (p = .026), indicating that differences of 6 points or more are unlikely. We also report that the achieved sample provided 80% power to detect effects of d = 0.63 or larger, and we note this as a limitation for smaller effects.

Reproduce this Case

Every number above comes from one base-R script, with no packages to install.

Download case-001-post-hoc-power.R →

Sources

← All Cases
Case 002: The Reviewer Wants a Multiple-Comparisons Correction →

Still stuck after the first checks?

Some problems turn on the details of your design, data, or the exact reviewer comment. A free 30-minute consult can identify the next defensible step and what it would take.

Book a free consult