The Reviewer Wants a Multiple-Comparisons Correction
Whether to correct depends on the claim you're making, not on how many p-values are in the paper.
Symptoms
- A reviewer writes that "the authors conducted many tests without correcting for multiple comparisons" and asks for Bonferroni.
- The paper has one primary outcome and several secondary outcomes, subgroup analyses, or pairwise comparisons after an ANOVA.
- Several findings that are significant now would not survive a strict correction.
What this usually means
Every test run at α = .05 has a 5% chance of a false positive when there is no effect. Across several independent tests of true null hypotheses, the chance that at least one comes out significant grows quickly: 14% with 3 tests, 40% with 10, and 64% with 20. A simulation of 10 independent null tests in the R script below gives the same 40%.
Whether that matters depends on how the results are used. Rubin (2021) separates three situations:
- Disjunction testing: the claim is supported if any test in a set is significant ("the intervention improved outcomes" when it improved at least one of five). Here the chance of a false positive across the set is the relevant error, and an adjustment is needed.
- Conjunction testing: the claim needs all tests to be significant. No adjustment is needed.
- Individual testing: each test answers its own question and is interpreted on its own. No adjustment is needed for that test's claim.
Bender and Lange (2001) make the same point for clinical research: adjustment is required in confirmatory studies whenever several tests feed one final conclusion, while exploratory analyses need not be adjusted but must be presented as exploratory. Rothman (1990) argued against routine adjustment altogether. The practical question is therefore not "how many tests did you run?" but "which tests support which claim?"
Common causes
1. Several primary outcomes without an order
If any one of them counts as success, they form one family.
2. Secondary outcomes reported as findings
The abstract highlights the secondary outcomes that were significant, as if they were confirmatory.
3. Pairwise comparisons after an omnibus test
All pairs of groups are compared and the significant pairs are reported.
4. Subgroup analyses
The effect is tested separately by sex, age, site, or baseline severity.
5. Exploratory analyses written up as confirmatory
Analyses chosen after seeing the data are described in the language of hypothesis tests.
Run these checks
- List every test in the paper and mark each one as prespecified primary, prespecified secondary, or exploratory. Use the preregistration, grant, or analysis plan, not memory.
- Write down the claim each test supports. Tests that feed the same conclusion form one family.
- Choose the error rate for each family. Control the familywise error rate (any false positive) for confirmatory claims. Control the false discovery rate (the expected share of false positives among significant results) when screening many outcomes where a few false leads are acceptable.
- Pick the method: Holm (1979) for familywise control, Benjamini and Hochberg (1995) for false discovery rate control.
- Plan what you will report: the family definition, the method, and both adjusted and unadjusted p-values, with effect sizes and confidence intervals.
What not to do
- Don't apply Bonferroni to every p-value in the paper. It treats unrelated questions as one claim.
- Don't define the families after seeing which results survive.
- Don't quietly drop outcomes that didn't survive the correction.
- Don't use plain Bonferroni when you want familywise control. Holm's procedure controls the same error rate and rejects at least as many hypotheses.
Treatment options
Holm for a confirmatory family
Use it when any single false positive would undermine the claim. It makes no assumptions about how the tests are related.
Benjamini–Hochberg for an exploratory family
Use it when you are screening many outcomes and want to limit the proportion of false leads rather than rule out every one.
A fixed testing order
Prespecify an order and test each hypothesis at the full α only if the one before it was significant. This protects the primary claim and gives secondary outcomes a principled place. It has to be planned before the results are known.
Relabel instead of correcting
Results that were never part of a confirmatory claim can be reported as exploratory, with effect sizes and intervals and without significance language, as hypotheses for future work.
Worked example
A study reports one prespecified primary outcome (p = .004) and eight secondary outcomes (p = .011, .019, .032, .041, .090, .210, .480, .730). Without any correction, five results are below .05.
If the reviewer treats all nine as one family:
| Method | Results below .05 |
|---|---|
| None | 5 |
| Bonferroni | 1 (primary) |
| Holm | 1 (primary) |
| Benjamini–Hochberg | 2 (primary and the first secondary, adjusted p = .049) |
If the analysis plan named the primary outcome in advance, it is tested on its own at α = .05 and stands (p = .004). The eight secondary outcomes form their own family. None survives Holm (smallest adjusted p = .088) or Benjamini–Hochberg (smallest adjusted p = .076).
The honest conclusion holds under every reasonable choice: the primary result is supported, and the secondary results are suggestive at most and should be presented as exploratory. The R script below reproduces both tables.
What to tell the reviewer
We agree that the secondary analyses require attention to multiplicity. As specified in our analysis plan, the primary outcome was a single confirmatory test, which we report at α = .05 (p = .004). We now treat the eight secondary outcomes as one family and report Holm-adjusted p-values alongside the unadjusted values (Holm, 1979). No secondary outcome remains significant after adjustment, and we have revised the abstract and discussion to describe these results as exploratory. We report effect sizes and confidence intervals for all outcomes so readers can judge their magnitude.
Every number above comes from one base-R script, with no packages to install.
Download case-002-multiple-comparisons.R →Sources
- Bender, R., & Lange, S. (2001). Adjusting for multiple testing: When and how? Journal of Clinical Epidemiology, 54(4), 343–349.
- Benjamini, Y., & Hochberg, Y. (1995). Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B, 57(1), 289–300.
- Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70.
- Rothman, K. J. (1990). No adjustments are needed for multiple comparisons. Epidemiology, 1(1), 43–46.
- Rubin, M. (2021). When to adjust alpha during multiple testing: A consideration of disjunction, conjunction, and individual testing. Synthese, 199, 10969–11000.