← The Analysis Clinic
Analysis Clinic · Case 010 · Diagnosed
My model won't run

My Logistic Regression Has Separation

Investigate perfect prediction before interpreting huge odds ratios or treating an optimizer’s last iterate as an estimate.

Logistic regressionSeparationRPython

By Data Analysis & Statistical Solutions (DASS) · Published · Presentation updated

About this example: The calculations use synthetic data, simulations, or specified illustrative inputs rather than client records. The example demonstrates particular assumptions and does not diagnose your study. Review your design, data, and model assumptions before applying a suggested approach. This update adds provenance and navigation information; it does not represent a new technical review of every calculation.

Symptoms

Your logistic regression returns huge coefficients, enormous standard errors, fitted probabilities near zero or one, or a convergence warning. A reported odds ratio may be implausibly large while its Wald test is uninformative.

What this means

Complete separation occurs when a linear combination of predictors perfectly separates the outcome categories. Along a separating direction, the likelihood keeps improving as coefficients grow; the usual finite maximum-likelihood estimate does not exist. Quasi-complete separation allows boundary ties while retaining a separating direction. Sparse data can contribute, but separation is not synonymous with every sparse-data or convergence problem.

Run these checks

  1. Confirm which outcome is modeled and inspect coding errors, leakage, and predictors derived from the outcome.
  2. Cross-tabulate categorical predictors against the outcome and inspect continuous predictors by outcome.
  3. Remember that separation may involve combinations of variables; no empty cell in a single table does not rule it out.
  4. Review fitting warnings and use a suitable separation diagnostic for the specified model. Investigate instability rather than trusting a convergence flag alone.

What not to do

Do not report the final numerical iterate as a stable odds ratio, increase iteration limits as the sole remedy, or remove a scientifically required adjustment variable simply to eliminate the warning. Do not turn arbitrary category merging into a routine fix.

Defensible options

Correct genuine coding or leakage problems first. If separation remains, consider a bias-reduced approach such as Firth logistic regression, or a Bayesian model with justified proper priors. Penalization yields finite estimates under appropriate conditions, but the choice of penalty or prior matters. Ridge regression is another regularization approach, particularly for prediction, and is not interchangeable with Firth’s bias-reduction method. Exact approaches may suit some small, simple designs but have computational and inferential limitations.

Choose the approach around the estimand and purpose, report the method, and examine sensitivity. A finite estimate does not create information absent from the data or establish out-of-sample predictive performance.

Worked example

These six synthetic observations have predictor values −3, −2, −1 for outcome 0 and 1, 2, 3 for outcome 1. With intercept fixed at zero for this illustration, the log likelihood increases from −0.977554 at slope 1 to −0.295107 at slope 2, −0.013522 at slope 5, and −0.000091 at slope 10. Its supremum is approached as the slope increases without bound. There is no finite maximizing slope to report.

The downloadable scripts reproduce these likelihood values using stable probability calculations. They demonstrate the failure of ordinary maximum likelihood; they do not implement or compare treatment methods. A real analysis needs a justified remedy and uncertainty estimates.

What to tell the reviewer

Adapt after completing the investigation: We identified separation in the specified logistic model. Ordinary maximum-likelihood estimates for the affected coefficients were not finite. We therefore used [justified method], documented its settings and assumptions, and examined sensitivity to [relevant alternatives]. We report uncertainty and the limitations imposed by the sparse outcome patterns.

See interpreting logistic regression and use the troubleshooting worksheet to record the diagnosis and decision.

See it in R and Python

R and Python evaluate the same likelihood path for six specified synthetic observations. No finite ordinary maximum-likelihood slope exists.

Python dependencies: NumPy and SciPy. Install with python -m pip install numpy scipy.

# Synthetic complete separation; base R. No finite ML slope exists.
x <- c(-3,-2,-1,1,2,3)
y <- as.integer(x > 0)
print(table(x > 0, y))
for (b in c(1,2,5,10)) {
  p <- plogis(b*x)
  loglik <- sum(ifelse(y == 1, log(p), log1p(-p)))
  cat(sprintf("slope=%g loglik=%.6f\n", b, loglik))
}
# Deliberately do not report the last optimizer iterate as an ML estimate.
# Same specified synthetic example as R; no finite ML slope exists.
import numpy as np
from scipy.special import log_expit
x=np.array([-3.,-2.,-1.,1.,2.,3.])
y=(x>0).astype(int)
for b in [1,2,5,10]:
    ll=np.sum(np.where(y==1,log_expit(b*x),log_expit(-b*x)))
    print(f"slope={b} loglik={ll:.6f}")
assert np.all(x[y==0]<0) and np.all(x[y==1]>0)

Download Python script

Reproduce this Case

Base R only. Python requires NumPy and SciPy. The example evaluates likelihoods; it does not fit a penalized or Bayesian remedy.

Download case-010-separation.R →

Sources

← Case 009: My Coauthor Reran the Code and Got Different Results
Case 011: My Interaction Is Significant—What Do I Report? →

Still stuck after the first checks?

Some problems turn on the details of your design, data, or the exact reviewer comment. A free 30-minute consult can identify the next defensible step and what it would take.

Book a free consult