Powering a Cluster-Randomized Grant When the ICC Is Unknown
Show how the required number of clusters changes across plausible ICCs, and separate participant attrition from cluster loss.
Symptoms
- You know approximately how many participants each site can recruit, but lack a reliable ICC.
- A proposal uses a single small ICC without a sensitivity table or allowance for lost sites.
What this usually means
The sample size depends on similarity within clusters as well as the effect size. For equal cluster sizes in a simple parallel design, the design effect is 1 + (m − 1) × ICC. An uncertain ICC creates an uncertain sample requirement; a single optimistic value hides that uncertainty.
Common causes
- An ICC from another outcome, population, or adjustment model has been borrowed without checking comparability.
- A large participant target is mistaken for a large number of independent clusters.
- Attrition is represented by inflating participants only, even though entire sites might be lost.
Run these checks
- Define the outcome scale, target effect, design, and intended analysis.
- Find ICC evidence from comparable settings and show a plausible range.
- Specify retained participants per cluster and realistic variation in cluster size.
- Include separate scenarios for individual dropout and cluster loss.
- Validate the final plan with a method that handles the actual cluster count, covariates, and allocation.
What not to do
Do not report this design-effect screen as an exact guarantee of 80% power. It uses a normal approximation and omits small-cluster degrees-of-freedom penalties. Do not use it unchanged for binary outcomes, stepped-wedge trials, repeated cross sections, or highly unequal clusters.
Treatment options
Use the range to test feasibility before committing to a design. If a plausible ICC makes the plan infeasible, evaluate additional clusters, a different target effect justified by practical importance, or design improvements. Predefine baseline adjustment and use a dedicated calculation or simulation aligned with the final analysis.
Worked example
For a continuous outcome with standardized difference 0.35, two-sided α = .05, target power .80, and 25 retained participants per cluster, the individually randomized normal approximation needs 128.14 participants per arm. Applying the design effect gives this screening table:
| ICC | Design effect | Clusters per arm | Participants per arm |
|---|---|---|---|
| .01 | 1.24 | 7 | 175 |
| .03 | 1.72 | 9 | 225 |
| .05 | 2.20 | 12 | 300 |
| .10 | 3.40 | 18 | 450 |
At ICC .05, retaining 20 rather than 25 participants requires 13 analyzable clusters per arm by the same approximation. To retain 12 clusters with 10% cluster loss, the simple recruitment allowance is 14 per arm. These are separate scenarios, not a combined final recommendation.
What to tell the reviewers
We evaluated recruitment requirements across ICCs .01–.10 and separately considered participant attrition and cluster loss. The design-effect calculation was used for feasibility screening. The final sample justification will use the planned analysis and a method that accounts for the number and sizes of clusters.
See it in R and Python
Both languages use the same normal-approximation feasibility screen. This is not a final small-cluster power calculation.
Python dependencies: NumPy and SciPy. Install with python -m pip install numpy scipy.
# Continuous outcome, parallel CRT, equal clusters; base R only.
# Normal approximation SCREEN, not a final small-cluster power calculation.
d <- 0.35; m <- 25; alpha <- 0.05; target <- 0.80
n_ind <- 2*(qnorm(1-alpha/2)+qnorm(target))^2/d^2
icc <- c(0.01,0.03,0.05,0.10)
de <- 1+(m-1)*icc
clusters <- ceiling(n_ind*de/m)
print(data.frame(ICC=icc,design_effect=de,clusters_per_arm=clusters,people_per_arm=clusters*m),row.names=FALSE,quote=FALSE)
cat(sprintf('Individual-randomization normal approximation: %.2f per arm\n',n_ind))
# Participant attrition changes retained cluster size, not just total sample.
retained <- 20
cat(sprintf('At ICC=.05 and 20 retained per cluster: %d clusters per arm\n',ceiling(n_ind*(1+(retained-1)*.05)/retained)))
# Cluster loss must be allowed for separately.
cat(sprintf('For %d analyzable clusters/arm and 10%% cluster loss: recruit %d/arm\n',clusters[3],ceiling(clusters[3]/.90)))
stopifnot(all(diff(clusters)>=0))
# DASS Analysis Clinic Case 005; Python dependencies: numpy, scipy.
from pathlib import Path
import numpy as np
from scipy import stats, optimize
HERE = Path(__file__).resolve().parent
def ols(X, y):
b = np.linalg.lstsq(X, y, rcond=None)[0]
residual = y - X @ b
df = len(y) - X.shape[1]
cov = (residual @ residual / df) * np.linalg.inv(X.T @ X)
se = np.sqrt(np.diag(cov))
ci = np.column_stack((b-stats.t.ppf(.975, df)*se,b+stats.t.ppf(.975,df)*se))
return b, cov, ci
def power(n, d):
# Matches R power.t.test(strict=FALSE): rejection tail in effect direction.
return stats.nct.sf(stats.t.ppf(.975,2*n-2),2*n-2,d*np.sqrt(n/2))
effect=.35; m=25
n=2*(stats.norm.ppf(.975)+stats.norm.ppf(.8))**2/effect**2
print('Normal approximation only; individual n per arm:',n)
for icc in [.01,.03,.05,.10]:
de=1+(m-1)*icc; k=int(np.ceil(n*de/m)); print(f'ICC={icc:.2f}; DE={de:.2f}; clusters/arm={k}; people/arm={k*m}')
print('20 retained per cluster, ICC .05:',int(np.ceil(n*(1+19*.05)/20)))
print('12 analyzable clusters and 10% cluster loss:',int(np.ceil(12/.9)))
assert [int(np.ceil(n*(1+24*r)/25)) for r in [.01,.03,.05,.10]]==[7,9,12,18]
Every number above comes from one base-R script, with no packages to install.
Download case-005-unknown-icc.R →