← The Analysis Clinic
Analysis Clinic · Case 008 · Diagnosed
My data aren't what we planned

Missing Data Across Waves: What Will Reviewers Accept?

Make the missing-data assumptions visible, choose a method that matches the study, and show sensitivity to departures from MAR.

Missing dataLongitudinal studiesmiceR

Symptoms

What this usually means

The analysis needs an assumption about unobserved outcomes. MAR means missingness can depend on observed information, but not on the missing outcome after conditioning on that information. It cannot be verified from observed data alone. The missing-data method and the sensitivity analysis should address the target outcome and treatment question.

Common causes

Run these checks

  1. Describe missingness by wave and arm and report known reasons.
  2. Define the estimand and distinguish missing assessments from events that alter the outcome question.
  3. Include relevant observed predictors of outcomes and missingness. Match the imputation model to interactions and dependence.
  4. Inspect imputation diagnostics and Monte Carlo stability; use enough imputations for the amount of missing information.
  5. Specify plausible departures from MAR with subject-matter input.

What not to do

Do not present MAR as a tested fact, replace missing values with a single guess and ignore uncertainty, or promise that imputation repairs every form of dropout. Complete cases can be valid under some conditions; validity must be argued for the chosen analysis.

Treatment options

Likelihood-based longitudinal models or multiple imputation can use incomplete records under their assumptions. For multiple waves or clustered studies, preserve that structure. Add an explicit sensitivity analysis, such as controlled shifts in imputed outcomes; its scenarios are assumptions rather than estimates of the missingness mechanism.

Worked example

A synthetic two-wave study enrolls 600 participants; 321 have observed follow-up. Missing outcomes number 110 in control and 169 in intervention. Missingness is generated from observed baseline and arm, so MAR holds by construction. Thirty normal-regression imputations include those predictors. The complete-data adjusted effect is 0.482 and the complete-case adjusted estimate is 0.474.

Assumption for missing intervention outcomesPooled effectApproximate 95% CI
MAR; no shift0.4490.226 to 0.672
Shift down by 0.5 outcome units0.169−0.056 to 0.395
Shift down by 1.0 outcome unit−0.110−0.340 to 0.120

Sensitivity intervals use Rubin variance pooling and a normal approximation. The script also reports the MAR finite-df mice::pool interval, 0.223 to 0.675. The shifts are illustrative choices, applied only to originally missing intervention outcomes. This baseline/follow-up example does not demonstrate a multilevel imputation model.

What to tell the reviewers

We report missing follow-up by arm and use multiple imputation conditional on treatment and baseline outcome. We state the MAR assumption and compare prespecified downward shifts for missing intervention outcomes. The sensitivity results show how departures from that assumption affect the estimated contrast.

See it in R and Python

Python performs Bayesian normal-regression multiple imputation for one missing continuous follow-up outcome, then pools the effect and outcome-shift sensitivity analyses. It uses the shared incomplete R-generated data but independent posterior draws, so its pooled estimates differ slightly from mice. The worked-example table above reports R results. This compact implementation is not a general chained-equations or multilevel imputation package.

Python dependencies: NumPy and SciPy. Install with python -m pip install numpy scipy.

Download the shared synthetic CSV beside the Python script before running it. Identical seeds in R and Python do not generally generate identical samples.

# Synthetic two-wave study: baseline fully observed, follow-up incomplete.
# Requires mice; no automatic installation. Multilevel/multiwave extensions
# need an imputation model that respects their dependence structure.
if (!requireNamespace('mice',quietly=TRUE)) stop('Install mice before running this example.')
set.seed(8008)
n <- 600
baseline <- rnorm(n)
arm <- rep(0:1,each=n/2)
followup <- .5*arm + .8*baseline + rnorm(n)
missing <- rbinom(n,1,plogis(-.5+.9*baseline+.6*arm))==1
full <- data.frame(arm,baseline,followup)
d <- full; d$followup[missing] <- NA
cc <- lm(followup~arm+baseline,d)
methods <- c(arm='',baseline='',followup='norm')
imp <- mice::mice(d,m=30,maxit=5,method=methods,seed=8009,printFlag=FALSE)
# Rubin pooling for a scalar arm coefficient, with normal-approximation CIs.
# The mice::pool result also supplies finite-df inference for the MAR fit.
pooled <- summary(mice::pool(with(imp,lm(followup~arm+baseline))),conf.int=TRUE)
print(pooled)
scalar <- function(delta) {
 vals <- vapply(seq_len(imp$m),function(i) {
  complete <- mice::complete(imp,i)
  # Controlled shift only for missing outcomes in the intervention arm.
  complete$followup[missing & arm==1] <- complete$followup[missing & arm==1]+delta
  f <- lm(followup~arm+baseline,complete)
  c(q=coef(f)['arm'],u=vcov(f)['arm','arm'])
 },numeric(2))
 q <- mean(vals[1,]); se <- sqrt(mean(vals[2,])+(1+1/imp$m)*var(vals[1,]))
 c(delta=delta,estimate=q,lower=q-qnorm(.975)*se,upper=q+qnorm(.975)*se)
}
cat(sprintf('Enrolled=%d; follow-up observed=%d; missing control=%d; missing intervention=%d\n',n,sum(!missing),sum(missing & arm==0),sum(missing & arm==1)))
cat(sprintf('Complete-data arm=%.3f; complete-case adjusted arm=%.3f\n',coef(lm(followup~arm+baseline,full))['arm'],coef(cc)['arm']))
print(t(vapply(c(0,-.5,-1),scalar,numeric(4))))
cat('mice version:',as.character(packageVersion('mice')),'\n')
stopifnot(imp$m==30,all(is.finite(pooled$estimate)))

# Maintainer-only export; run from repository root. Synthetic data only.
if (Sys.getenv("CLINIC_EXPORT_DATA") == "1") {
write.csv(d,'analysis-clinic/examples/case-008-shared-data.csv',row.names=FALSE,quote=FALSE,na='')
}
# DASS Analysis Clinic Case 008; Python dependencies: numpy, scipy.
from pathlib import Path
import numpy as np
from scipy import stats, optimize
HERE = Path(__file__).resolve().parent

def ols(X, y):
    b = np.linalg.lstsq(X, y, rcond=None)[0]
    residual = y - X @ b
    df = len(y) - X.shape[1]
    cov = (residual @ residual / df) * np.linalg.inv(X.T @ X)
    se = np.sqrt(np.diag(cov))
    ci = np.column_stack((b-stats.t.ppf(.975, df)*se,b+stats.t.ppf(.975,df)*se))
    return b, cov, ci

def power(n, d):
    # Matches R power.t.test(strict=FALSE): rejection tail in effect direction.
    return stats.nct.sf(stats.t.ppf(.975,2*n-2),2*n-2,d*np.sqrt(n/2))

# Bayesian normal-regression MI for one incomplete continuous outcome.
# Fully observed arm/baseline, flat coefficient prior and p(sigma^2) proportional
# to 1/sigma^2. Not a general MICE or multilevel imputation implementation.
d=np.genfromtxt(HERE/'case-008-shared-data.csv',delimiter=',',names=True)
arm=d['arm']; baseline=d['baseline']; y=d['followup']; missing=np.isnan(y)
X=np.column_stack((np.ones(len(y)),arm,baseline)); observed=~missing
Xo=X[observed]; yo=y[observed]; bhat=np.linalg.lstsq(Xo,yo,rcond=None)[0]
df=len(yo)-X.shape[1]; sse=np.sum((yo-Xo@bhat)**2); inv=np.linalg.inv(Xo.T@Xo)
rng=np.random.default_rng(8009); completed=[]
for _ in range(30):
 variance=sse/rng.chisquare(df)
 beta=rng.multivariate_normal(bhat,variance*inv)
 out=y.copy(); out[missing]=X[missing]@beta+rng.normal(0,np.sqrt(variance),missing.sum())
 completed.append(out)
def pooled(delta):
 estimates=[]; variances=[]
 for out in completed:
  shifted=out.copy(); shifted[missing & (arm==1)]+=delta
  b,cov,_=ols(X,shifted); estimates.append(b[1]);variances.append(cov[1,1])
 q=np.mean(estimates); se=np.sqrt(np.mean(variances)+(1+1/30)*np.var(estimates,ddof=1))
 return q,q-stats.norm.ppf(.975)*se,q+stats.norm.ppf(.975)*se
print('Observed:',observed.sum(),'missing by arm:',[np.sum(missing & (arm==g)) for g in [0,1]])
print('Complete-case adjusted effect:',bhat[1])
for delta in [0,-.5,-1]:print('delta / pooled effect / approximate CI:',delta,pooled(delta))
print('Python posterior draws differ from mice; the page table reports the R run.')
assert observed.sum()==321 and np.isclose(bhat[1],.474,atol=.0005)
assert pooled(0)[0]>pooled(-.5)[0]>pooled(-1)[0]

Download Python script · Download shared synthetic CSV

Reproduce this Case

Requires the mice package. The example was checked with mice 3.19.0 on R 4.6.0.

Download case-008-missing-waves.R →

Sources

← Case 007: The Effect Flipped Sign When I Added a Covariate
Case 009: My Coauthor Reran the Code and Got Different Results →

Still stuck after the first checks?

Some problems turn on the details of your design, data, or the exact reviewer comment. A free 30-minute consult can identify the next defensible step and what it would take.

Book a free consult