Source-linked AI summary
The phase transition for the existence of the maximum likelihood estimate in high-dimensional logistic regression
Emmanuel J. Candes, Pragya Sur
TL;DR
The paper asks when the MLE exists in high-dimensional logistic regression, where geometric separation criteria do not provide an a priori predictive boundary. Using convex-geometric analysis of Gaussian-covariate models, it derives an explicit h_MLE boundary: above it the MLE existence probability tends to 0, while below it the probability tends to 1.
Problem
Geometric separation characterizations do not concretely predict when the logistic-regression MLE will exist for covariates drawn from a distribution.
Method
The paper uses convex geometry to derive an explicit phase-transition boundary for high-dimensional logistic regression with Gaussian covariates.
Results
κ > h_MLE(β0, γ0) yields P{MLE exists} → 0, whereas κ < h_MLE(β0, γ0) yields P{MLE exists} → 1.
Takeaways & Limitations
MLE existence undergoes a sharp transition governed by dimensionality, the intercept, and the limiting variance of the linear predictor.
Takeaways & Limitations
The analysis assumes Gaussian covariates and leaves whether the formulas extend to more general covariate distributions for future research.
Abstract
from arXiv · showhide
This paper rigorously establishes that the existence of the maximum likelihood estimate (MLE) in high-dimensional logistic regression models with Gaussian covariates undergoes a sharp `phase transition'. We introduce an explicit boundary curve $h_{\text{MLE}}$, parameterized by two scalars measuring the overall magnitude of the unknown sequence of regression coefficients, with the following property: in the limit of large sample sizes $n$ and number of features $p$ proportioned in such a way that $p/n \rightarrow κ$, we show that if the problem is sufficiently high dimensional in the sense that $κ> h_{\text{MLE}}$, then the MLE does not exist with probability one. Conversely, if $κ< h_{\text{MLE}}$, the MLE asymptotically exists with probability one.
1 Introduction
The paper addresses when the logistic-regression MLE exists by moving from abstract separation conditions to an explicit high-dimensional phase-transition boundary for Gaussian covariates. It shows that this boundary depends on the intercept and the limiting variance of the linear predictor, with existence probability switching sharply across it.
- 1.1 Data geometry and the existence of the MLE: The MLE exists exactly when the sample points overlap; complete or quasi-complete separation prevents its existence.This geometric characterization motivates the search for a predictive criterion based on the covariate distribution.
- 1.2 Limitations: Existing separation characterizations do not tell analysts in advance when the MLE will exist for covariates sampled from a distribution.The paper frames this as the practical limitation of replacing one abstract condition with another.
- 1.3 Cover’s result: For independent labels with equal marginal probabilities, Cover’s result gives a phase transition at κ = p/n = 1/2 between overlap and separation.Below 1/2 the data asymptotically overlap, whereas above 1/2 they are asymptotically separated.
- 1.3 Cover’s result: The paper asks whether analogous phase transitions occur when class labels depend on covariates, a question motivated by the importance of likelihood-based inference.The paper explicitly identifies this as its central question.
- 1.4 Phase transitions: For Gaussian covariates, the paper rigorously establishes and explicitly computes the MLE phase-transition boundary h_MLE.The analysis considers Gaussian models with non-singular covariance and high-dimensional asymptotics p/n → κ.
- 1.4 Phase transitions: The boundary depends on β0 and the limiting predictor variance, with κ > h_MLE implying asymptotic nonexistence and κ < h_MLE implying asymptotic existence.The transition is sharp, and h_MLE decreases with both β0 and γ0; the paper derives the formula using convex-geometry ideas.
2 Main Result
The paper derives a sharp phase-transition boundary for MLE existence in high-dimensional logistic regression, analyzes special signal regimes and intercept specifications, and validates the predictions empirically.
- 2.1 Model with intercept: For κ > hMLE(β0, γ0), limn,p→∞ P{MLE exists} = 0, while for κ < hMLE(β0, γ0), limn,p→∞ P{MLE exists} = 1.
- 2.1 Model with intercept: The boundary h_MLE(β0, γ0) is symmetric in β0 and decreases as either |β0| or γ0 increases.
- 2.2 Special cases: When signal vanishes, the phase transition is at κ = 1/2, recovering and extending Cover’s result for symmetric independent responses.
- 2.2 Special cases: With infinite signal strength, MLE existence requires p/n → 0.
- 2.3 Model without intercept: For models without an intercept, the same phase-transition location holds, with a boundary given by equation (8).
- 2.4 Comparison with empirical results: Across simulated Gaussian designs, empirical MLE-existence probabilities show sharp transitions closely aligned with the theoretical curves; at P(yi = 1) = 0.9, the boundary is theoretically κ = 0.255.In all replications, the MLE existed for κ < 0.24 and did not exist for κ > 0.28.
3 Conic Geometry
The section recasts MLE nonexistence as a separating-hyperplane problem and then as intersection of a random subspace with a convex cone. Convex-geometric bounds yield the phase transition at κ = h_MLE(β0, γ0).
- 3 Conic Geometry: No MLE occurs exactly when a nonzero affine combination classifies every observation with nonnegative signed margin, equivalently when a separating hyperplane exists.
- 3.1 Gaussian covariates: A nonsingular covariance transformation preserves separating hyperplanes, so the analysis reduces to independent standard-normal predictors without loss of generality.
- 3.1 Gaussian covariates: After rotational invariance concentrates the signal in the first coordinate, the response and first predictor are modeled through (Y, V), while the remaining predictors are independent standard normals.
- 3.2 Convex cones: The no-MLE event decomposes into univariate separation or a nontrivial intersection between L = span(X2, …, Xp) and the cone C(W), with W = span(Y, V).
- 3.2 Convex cones: Univariate separation has exponentially small probability, so asymptotic MLE existence is governed by whether the random subspace L intersects C(W).
- 3.3 Proof of Theorem 1: The approximate kinematic formula and statistical dimension quantify this intersection, with Q_n = h_MLE(β0, γ0) + O_P(n^-1/2).
- 3.3 Proof of Theorem 1: If κ > h_MLE(β0, γ0), no MLE occurs with probability tending to one; if κ < h_MLE(β0, γ0), the complementary existence result follows similarly.
4 Proof of Theorem 3
The proof establishes concentration of a random convex objective around its expectation and controls the minimizer using convexity, smoothness, and local approximation. It shows that the relevant minimizer converges at rate O_P(n^-1/4).
- The proof defines J(x) = ||x_+||^2/2, forms the matrix A from y and V, and represents the random objective F and its expectation f.
- Strict convexity of f gives a unique population minimizer λ0, while λ⋆ denotes a minimizer of the random objective F.
- Sub-exponential concentration controls both F(λ) and its gradient around their expectations, supplying the uniform bounds used in the minimizer analysis.
- Convexity and Lipschitz-gradient bounds reduce control of λ⋆ to behavior on a circle around λ0, and a finite net extends pointwise bounds over that circle.
- The minimizer satisfies ||λ⋆ − λ0|| = O_P(n^-1/4), establishing the rate needed in Theorem 3.
5 Conclusion
The paper establishes a Gaussian-covariate MLE phase transition and derives an explicit boundary for models with or without an intercept. Its convex-geometric approach may extend to broader covariate distributions, which remains future work.
- The paper establishes a phase transition for MLE existence in high-dimensional logistic regression with Gaussian covariates.
- It derives a simple phase-transition boundary for models fitted with or without an intercept using convex geometry and a kinematic formula.
- The authors suggest that the phenomena and formulas may hold for more general covariate distributions, leaving that extension to future research.