Source-linked AI summary
On the Distribution of Penalized Maximum Likelihood Estimators: The LASSO, SCAD, and Thresholding
Benedikt M. Potscher, Hannes Leeb
TL;DR
The paper asks how LASSO, SCAD, and thresholding estimators are distributed in finite samples and asymptotically under conservative versus consistent model-selection tuning. It derives these distributions and related uniform rates, finding highly nonnormal behavior that can persist or intensify in large samples, slower-than-n^-1/2 uniform convergence under consistent selection, and an intrinsic impossibility of distribution estimation.
Problem
Distributional behavior of penalized estimators remains incompletely understood, especially across finite samples, large samples, and model-selection tuning regimes.
Method
The paper derives finite-sample and moving-parameter asymptotic distributions for LASSO, SCAD, and hard-thresholding estimators, together with uniform convergence rates and distribution-estimation results.
Results
The estimators’ distributions are highly nonnormal in finite samples and can remain highly nonnormal, or become more pronounced, in large samples under consistent model selection.
Takeaways & Limitations
Nonnormal distributional behavior is not merely a small-sample effect, and the consistent-selection tuning regime does not deliver uniform n^1/2-consistency.
Takeaways & Limitations
Restricting the parameter space to avoid nonuniformity can exclude parameters larger than n^-1/2 that the same procedure classifies as nonzero with probability tending to one.
Abstract
from arXiv · showhide
We study the distributions of the LASSO, SCAD, and thresholding estimators, in finite samples and in the large-sample limit. The asymptotic distributions are derived for both the case where the estimators are tuned to perform consistent model selection and for the case where the estimators are tuned to perform conservative model selection. Our findings complement those of Knight and Fu (2000) and Fan and Li (2001). We show that the distributions are typically highly nonnormal regardless of how the estimator is tuned, and that this property persists in large samples. The uniform convergence rate of these estimators is also obtained, and is shown to be slower than 1/root(n) in case the estimator is tuned to perform consistent model selection. An impossibility result regarding estimation of the estimators' distribution function is also provided.
1 Introduction
The paper addresses incomplete distributional understanding of penalized maximum likelihood estimators by studying finite-sample and asymptotic behavior for thresholding, LASSO, and SCAD estimators under different tuning regimes.
- Distributional properties of penalized maximum likelihood estimators, including finite-sample and large-sample limits, remain incompletely understood.
- The paper studies hard-thresholding, LASSO, and SCAD estimators in a tractable model designed to reveal strengths and weaknesses relevant to more complex settings.
- The analysis covers both conservative and consistent model-selection tuning regimes.
- The paper derives uniform convergence rates and shows that consistent model-selection tuning yields a rate slower than n^-1/2.
- It also shows that the finite-sample distribution of these estimators cannot be estimated in any reasonable sense.
- The paper connects penalized maximum likelihood estimators to classical post-model-selection estimators whose distributional properties have also been studied.
2 The Model and the Estimators
The paper analyzes three penalized estimators in an orthogonal Gaussian model: hard-thresholding, soft-thresholding/LASSO, and SCAD, with tuning parameters governing their shrinkage and selection behavior.
- The analysis begins with an orthogonal linear regression model, where separable penalties make component estimators mutually independent.
- In the univariate Gaussian location formulation, observations are independent N(θ, σ2), with known variance normalized to one.
- Hard-thresholding sets the estimator to zero unless |ȳ| exceeds the positive tuning threshold ηn, otherwise returning ȳ.
- For ηn=n^-1/4, hard-thresholding is a simple instance of Hodges’ estimator.
- Soft-thresholding applies sign(ȳ)(|ȳ|−ηn)+ and coincides with the LASSO in this setting.
- SCAD combines soft-thresholding for small |ȳ| with hard-thresholding for large |ȳ|, joined by piecewise linear interpolation and controlled by a>2.
3 Model Selection Probabilities
Model-selection behavior depends on the tuning sequence: conservative and consistent regimes have distinct asymptotic selection probabilities, while moving-parameter analysis reveals nonuniform behavior near the restricted model.
- All three estimators select the restricted model when their estimate equals zero and the unrestricted model otherwise.
- The condition ηn→0 ensures that incorrectly selecting the restricted model for nonzero θ has probability vanishing asymptotically.
- When n^1/2ηn→e<∞, the procedures are conservative and retain a positive limiting probability of selecting the unrestricted model at θ=0.
- When n^1/2ηn→∞, the procedures are consistent model selectors and the unrestricted-model selection probability vanishes at θ=0.
- Moving-parameter asymptotics are introduced because pointwise asymptotics can miss essential finite-sample behavior in model-selection problems.
- At the boundary |ζ|=1, if n^1/2(ηn−ζθn)→r, the restricted-model selection probability converges to Φ(r).
- Under consistent selection, deviations with θn/ηn→ζ and |ζ|<1 can remain asymptotically undetected, including deviations larger than n^-1/2.
- The convergence speed of selection probabilities depends on the relevant rates of ηn, θn, and ηn−ζθn.
4 Consistency, uniform consistency, and uniform convergence rate of ˆθH, ˆθS, and ˆθSCAD
All three estimators are uniformly consistent when ηn→0, but their uniform convergence rate depends on tuning: conservative selection permits n^1/2-consistency, whereas consistent selection is slower.
- The condition ηn→0 is equivalent to consistency of the hard-thresholding estimator and carries over to soft-thresholding and SCAD.
- All three estimators are uniformly consistent when ηn→0, with the relevant supremum probability converging to zero.
- The uniform convergence bound uses an=min{n^1/2,ηn^-1}, yielding a rate controlled by the smaller of n^1/2 and ηn^-1.
- Under conservative model-selection tuning, the estimators are uniformly n^1/2-consistent.
- Under consistent model-selection tuning, the theorem guarantees only uniform ηn^-1-consistency, and the estimators do not converge faster than ηn uniformly.
- When n^1/2ηn→0, each estimator is uniformly asymptotically equivalent to the unrestricted estimator ȳ.
5 The distributions of ˆθH, ˆθS, and ˆθSCAD
The estimators have finite-sample and asymptotic distributions that are often nonnormal, multimodal, and sensitive to local parameter sequences. Consistent model-selection tuning can also undermine uniform n^1/2-consistency, while pointwise asymptotics may miss important finite-sample behavior.
- Finite-sample distributions: Hard-thresholding has a singular point-mass component and an absolutely continuous component formed by an excised unrestricted-normal distribution.The two components correspond to selecting the restricted and unrestricted models, respectively.
- Finite-sample distributions: Finite-sample distributions are typically highly nonnormal and can be multimodal for hard-thresholding, soft-thresholding, and SCAD.Hard-thresholding combines a point mass with an excised normal component; SCAD similarly combines a singular component with a multimodal continuous part.
- Asymptotic distributions: Under local sequences, hard-thresholding, soft-thresholding, and SCAD converge to distributions that can include point masses, truncated or shifted normal pieces, and multimodality.The limiting forms depend on the behavior of n^1/2θ_n and the tuning parameter, including boundary cases and divergent sequences.
- Asymptotic distributions: Pointwise asymptotic distributions can misrepresent finite-sample behavior when θ is close to, but not equal to, zero, whereas local-sequence limits track it more closely.This discrepancy is documented for hard-thresholding, soft-thresholding, and SCAD.
- Consistent model selection: Consistent model-selection tuning produces nonuniform behavior: hard-thresholding and SCAD are not uniformly n^1/2-consistent, while soft-thresholding is not even pointwise n^1/2-consistent.The paper characterizes additional limiting regimes, including point-mass limits and escape of mass.
6 Impossibility results for estimating the distribution of ˆθH, ˆθS, and ˆθSCAD
The paper shows that estimating the finite-sample cdfs of hard-thresholding, soft-thresholding, and SCAD estimators is intrinsically difficult, including for consistent procedures. The unavoidable error depends on the tuning sequence and remains substantial except when the estimator approaches the unrestricted maximum likelihood estimator.
- Problem and setup: The cdfs of the centered and scaled estimators depend on the unknown parameter in a complicated manner, making their estimation intrinsically difficult.The paper treats hard-thresholding, soft-thresholding, and SCAD within a unified framework.
- Consistent cdf estimators: No uniformly consistent estimator exists for the cdf when the estimator of that cdf is itself consistent.This impossibility applies to each of the hard-thresholding, soft-thresholding, and SCAD estimators under the stated tuning sequence.
- Consistent cdf estimators: For consistent cdf estimators, the unavoidable worst-case error is at least any ε < (Φ(t + e) −Φ(t −e))/2.Here e is the limit of n^1/2η_n, and the bound applies over parameter neighborhoods shrinking at rate n^-1/2.
- Dependence on tuning: When estimator tuning is consistent, e = ∞ and the error range equals 1/2; under conservative tuning, e < ∞ and the range is smaller but can remain substantial.Only e = 0 makes the bound zero, corresponding to uniform asymptotic equivalence with the unrestricted maximum likelihood estimator.
- Arbitrary cdf estimators: The same non-uniformity affects arbitrary cdf estimators, although their lower bound is 1/2 rather than 1.The result includes both a large-sample limit statement and a finite-sample statement.
- Dependence on tuning: A larger η_n produces a more sparse estimator and directly enlarges the range of errors in which any cdf estimator performs poorly.In large samples, the corresponding role is played by e = lim_n n^1/2η_n.
7 Conclusion
The estimators have highly nonnormal distributions in finite and large samples, with behavior depending strongly on tuning and parameter asymptotics. Consistent-selection tuning also produces slower uniform convergence and makes distribution-function estimation impossible in a uniform sense.
- Finite-sample distributions are highly nonnormal because they combine a singular normal component with an absolutely continuous component that may be multimodal.
- Moving-parameter asymptotics show that nonnormal behavior can persist at any sample size, including when parameters approach a lower-dimensional submodel.
- Under consistent model-selection tuning, large-sample distributions can remain highly nonnormal and become more pronounced, with SCAD mass potentially escaping to ±∞.
- Uniform convergence is always obtained, but its rate is uniformly n^1/2 under conservative tuning and slower than n^1/2 under consistent-selection tuning.
- The centered and scaled finite-sample cdf cannot be estimated uniformly consistently, although pointwise-consistent estimators exist.
- Risk is favorable near the lower-dimensional model but reverses outside that neighborhood, while consistent-selection tuning makes worst-case risk increase indefinitely with sample size.
A Appendix
The appendix characterizes moving-parameter limits for hard-, soft-, and SCAD thresholding under consistent-selection tuning. It also establishes sharp impossibility results for uniformly estimating their finite-sample distribution functions.
- Under consistent-selection tuning, the relevant asymptotic scale is η_n^-1 rather than n^1/2.
- For hard-thresholding, the limit is a pointmass at zero when 1<|ζ|≤∞.
- Soft-thresholding converges to the pointmass δ_{−sign(ζ) min(1,|ζ|)} when θ_n/η_n converges to ζ.
- SCAD converges to pointmasses at −sign(ζ)min(1,|ζ|) for |ζ|≤2 and at −sign(ζ)(a−|ζ|)/(a−2) for 2<|ζ|<a.
- The appendix proves that no uniformly consistent estimator exists for the distribution cdf at |t|≤1, even over compact sets containing zero.
- For |t|>1, trivial constant estimators are uniformly consistent as the cdf converges uniformly to one for t>1 and zero for t<−1.