← The Analysis Clinic
Analysis Clinic · Case 007 · Diagnosed
My results don't make sense

The Effect Flipped Sign When I Added a Covariate

Check the analysis rows and the adjustment question before treating a sign reversal as evidence of a mistake.

RegressionCovariate adjustmentR

Symptoms

What this usually means

The unadjusted association and the association conditional on a covariate answer different questions. A reversal may reflect confounding or suppression, but the regression output alone cannot decide which causal interpretation is justified. Adjustment can also introduce bias when a variable is a mediator or collider.

Common causes

Run these checks

  1. Compare both models on the same rows and verify reference levels and units.
  2. State whether you want a marginal association, a conditional association, or a causal effect.
  3. Use substantive knowledge and a causal diagram to decide whether adjustment is appropriate.
  4. Inspect scatterplots, uncertainty, predictor correlations, and influence. Assess functional form and interactions.

What not to do

Do not select covariates to restore a preferred direction. A tolerable VIF does not establish correct adjustment, and a large VIF does not by itself explain a reversal. Do not call the adjusted coefficient causal without the necessary design and assumptions.

Treatment options

Report unadjusted and prespecified adjusted estimates with their distinct interpretations. Separate the effects of changing the sample and changing the model. If an adjustment decision is uncertain, present justified sensitivity models rather than optimizing the sign.

Worked example

Eight deliberately constructed observations satisfy z = x + w and y = −x + 2z + e, where x, w, and e are orthogonal. Both models use all eight rows.

Modelx coefficient95% CI
y ~ x+1.000−1.234 to 3.234
y ~ x + z−1.000−2.626 to 0.626

The adjusted z coefficient is +2.000; correlation between x and z is .707 and the x VIF is 2.000. Thus a reversal is possible with modest VIF and no sample change. Both focal intervals include zero. This algebraic illustration supplies no real-world causal evidence.

What to tell the reviewers

We compared the unadjusted and adjusted models on identical observations. The sign change therefore reflects adjustment rather than sample exclusion. We describe the different questions the models address, justify the covariate using the study design, and report confidence intervals for both estimates.

See it in R and Python

Both languages construct the same deterministic observations and reproduce the sign reversal, t-based confidence intervals, and VIF.

Python dependencies: NumPy and SciPy. Install with python -m pip install numpy scipy.

# Exact synthetic illustration; base R only. Orthogonal residuals.
x <- rep(c(-1,1),each=4)
w <- rep(c(-1,-1,1,1),2)
e <- rep(c(-1,1),4)
z <- x+w
y <- -x+2*z+e
d <- data.frame(x,z,y)
a <- lm(y~x,d); b <- lm(y~x+z,d)
cat(sprintf('Same rows: unadjusted n=%d; adjusted n=%d\n',nobs(a),nobs(b)))
cat(sprintf('Unadjusted x=%.3f; adjusted x=%.3f; adjusted z=%.3f\n',coef(a)['x'],coef(b)['x'],coef(b)['z']))
vif <- 1/(1-summary(lm(x~z,d))$r.squared)
cat(sprintf('Correlation x,z=%.3f; x VIF=%.3f\n',cor(x,z),vif))
cat(sprintf('Unadjusted x CI=[%.3f,%.3f]; adjusted x CI=[%.3f,%.3f]\n',confint(a)['x',1],confint(a)['x',2],confint(b)['x',1],confint(b)['x',2]))
stopifnot(abs(coef(a)['x']-1)<1e-8,abs(coef(b)['x']+1)<1e-8)
# DASS Analysis Clinic Case 007; Python dependencies: numpy, scipy.
from pathlib import Path
import numpy as np
from scipy import stats, optimize
HERE = Path(__file__).resolve().parent

def ols(X, y):
    b = np.linalg.lstsq(X, y, rcond=None)[0]
    residual = y - X @ b
    df = len(y) - X.shape[1]
    cov = (residual @ residual / df) * np.linalg.inv(X.T @ X)
    se = np.sqrt(np.diag(cov))
    ci = np.column_stack((b-stats.t.ppf(.975, df)*se,b+stats.t.ppf(.975,df)*se))
    return b, cov, ci

def power(n, d):
    # Matches R power.t.test(strict=FALSE): rejection tail in effect direction.
    return stats.nct.sf(stats.t.ppf(.975,2*n-2),2*n-2,d*np.sqrt(n/2))

x=np.repeat([-1.,1.],4); w=np.tile([-1.,-1.,1.,1.],2); e=np.tile([-1.,1.],4)
z=x+w; y=-x+2*z+e
naive=ols(np.column_stack((np.ones(8),x)),y)
adjusted=ols(np.column_stack((np.ones(8),x,z)),y)
vif=1/(1-np.corrcoef(x,z)[0,1]**2)
print('Unadjusted x / CI:',naive[0][1],naive[2][1]); print('Adjusted x / CI:',adjusted[0][1],adjusted[2][1]);print('Adjusted z:',adjusted[0][2],'VIF:',vif)
assert np.isclose(naive[0][1],1) and np.isclose(adjusted[0][1],-1) and np.isclose(vif,2)

Download Python script

Reproduce this Case

Every number above comes from one base-R script, with no packages to install.

Download case-007-coefficient-sign.R →

Sources

← Case 006: My Mixed Model Won’t Converge with Crossed Random Effects
Case 008: Missing Data Across Waves: What Will Reviewers Accept? →

Still stuck after the first checks?

Some problems turn on the details of your design, data, or the exact reviewer comment. A free 30-minute consult can identify the next defensible step and what it would take.

Book a free consult