← Back to blog
Study Design

Statistical Power for Cluster-Randomized Trials: Why More Participants Isn't Enough

2 min read

In a lot of applied research — education, health systems, workplace interventions — you can't randomize individuals without contaminating the comparison. You randomize classrooms, clinics, or teams instead, and everyone within a cluster gets the same treatment. That design choice, cluster randomization, has a direct and often underestimated cost: it reduces your effective sample size below your total headcount, sometimes drastically.

Why clustering costs you power

People within the same cluster tend to resemble each other more than people in different clusters — same teacher, same clinic culture, same team norms. That similarity is measured by the intraclass correlation (ICC). The bigger the ICC and the bigger the cluster, the more your 500 students really behave, statistically, like a much smaller number of independent observations. The standard way to quantify this is the design effect: 1 + (m − 1) × ICC, where m is the average cluster size. A design effect of 3 means you effectively need three times the individually-randomized sample size to detect the same effect.

A concrete illustration

Say you have 20 schools of 30 students each (600 students total) and an ICC of 0.15 — a fairly typical value for achievement outcomes. The design effect is 1 + (29 × 0.15) ≈ 5.35. Your 600 students provide roughly the statistical power of about 112 independently randomized students. A power analysis that ignores clustering and treats this as 600 independent observations will be badly overconfident about what the study can detect.

What this means for planning a study

  • Power analysis for a cluster-randomized design needs the ICC (from prior literature or a pilot), the planned cluster size, and the number of clusters — not just a target headcount.
  • Adding participants within existing clusters has rapidly diminishing returns. Adding more clusters almost always buys more power than adding more people per cluster.
  • The analysis model has to match the design: a standard regression that ignores clustering will understate standard errors and can produce false positives. Multilevel (mixed-effects) models that explicitly model the cluster structure are the standard approach.

If you're planning a cluster or multi-site trial, the sample-size math is the single most consequential thing to get right before you start recruiting sites — this is exactly where we help.

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}