Source-linked AI summary

Valid Post-Selection and Post-Regularization Inference: An Elementary, General Approach

Victor Chernozhukov, Christian Hansen, Martin Spindler

arXiv:1501.03430v3math.STecon.EM

TL;DR

The paper addresses how to conduct valid inference on a low-dimensional parameter when a high-dimensional nuisance parameter is estimated after selection or regularization. It develops a general orthogonal-score framework with sufficient conditions and shows that inference can remain regular despite irregular nuisance estimation, while noting stronger requirements for one-step estimation.

  • Problem

    Regularization and variable selection complicate inference about low-dimensional parameters when nuisance parameters are high-dimensional.

  • Method

    The paper uses empirical estimating equations made orthogonal to nuisance estimation errors, together with high-quality nuisance estimators, and develops sufficient conditions for affine-quadratic models.

  • Results

    Inference for α can remain valid, including through generalized Neyman C(α) statistics, even when η estimators are non-regular and not asymptotically linear.

  • Takeaways & Limitations

    Orthogonalized estimating equations provide a general route to post-selection and post-regularization inference across broad model classes without requiring perfect variable selection.

  • Takeaways & Limitations

    The one-step estimator requires stronger regularity conditions and can suffer from higher-order biases relative to the argmin estimator.

Abstract

from arXiv · show

Here we present an expository, general analysis of valid post-selection or post-regularization inference about a low-dimensional target parameter, $α$, in the presence of a very high-dimensional nuisance parameter, $η$, which is estimated using modern selection or regularization methods. Our analysis relies on high-level, easy-to-interpret conditions that allow one to clearly see the structures needed for achieving valid post-regularization inference. Simple, readily verifiable sufficient conditions are provided for a class of affine-quadratic models. We focus our discussion on estimation and inference procedures based on using the empirical analog of theoretical equations $$M(α, η)=0$$ which identify $α$. Within this structure, we show that setting up such equations in a manner such that the orthogonality/immunization condition $$\partial_ηM(α, η) = 0$$ at the true parameter values is satisfied, coupled with plausible conditions on the smoothness of $M$ and the quality of the estimator $\hat η$, guarantees that inference on for the main parameter $α$ based on testing or point estimation methods discussed below will be regular despite selection or regularization biases occurring in estimation of $η$. In particular, the estimator of $α$ will often be uniformly consistent at the root-$n$ rate and uniformly asymptotically normal even though estimators $\hat η$ will generally not be asymptotically linear and regular. The uniformity holds over large classes of models that do not impose highly implausible "beta-min" conditions. We also show that inference can be carried out by inverting tests formed from Neyman's $C(α)$ (orthogonal score) statistics.

1. Introduction

The paper develops a general framework for valid inference on a low-dimensional parameter with high-dimensional nuisance parameters estimated by selection or regularization. Its approach combines orthogonalized estimating equations with high-quality nuisance estimators and applies the framework to affine-quadratic and instrumental-variables models.

  • Motivation: High-dimensional models require regularization, but inference must account explicitly for selection or shrinkage to characterize estimator behavior accurately.The paper also notes that specification searches in conventional low-dimensional models can invalidate usual inference when variable selection is ignored.
  • Applications: The paper supplies sufficient conditions for affine-quadratic models and studies post-regularization inference with many instruments and controls in linear instrumental variables models.It also reports a simulation example and an empirical illustration involving logit demand estimation.
  • Framework: The framework targets valid inferential statements for a low-dimensional parameter α in the presence of a high-dimensional nuisance parameter η.It is designed to encompass existing results and to apply across broad classes of high-dimensional models.
  • Framework: Orthogonality or immunization of the estimating equations makes inference about α less sensitive to regularization-induced errors in estimating η.The condition can be established using Neyman’s orthogonalized score in likelihood settings and extended to GMM.
  • Nuisance estimation: Approximate sparsity provides one example of the structure that high-quality estimators of η can exploit for informative inference.The paper emphasizes that its general results do not require sparsity-based estimation strategies.
  • Paper organization: The presentation develops general results, orthogonality constructions, approximately sparse estimation results, and applications to endogenous-variable inference.The paper’s notation and matrix-norm definitions support the formal development of these results.

2. A Testing and Estimation Approach to Valid Post-Selection and Post-Regularization Inference

The paper develops valid testing and estimation procedures for a low-dimensional target with a high-dimensional nuisance parameter. Orthogonal estimating equations, structured nuisance estimation, and regularity conditions support uniform inference after selection or regularization.

  • Framework: The framework identifies α through empirical equations M(α, η)=0 while estimating the high-dimensional nuisance parameter η under structured assumptions.The nuisance dimension may greatly exceed the sample size, motivating selection or regularization methods.
  • Testing: The generalized C(α)-statistic is a quadratic form in normalized, orthogonalized scores and can be inverted to construct confidence sets.Under the stated conditions, the score is asymptotically normal and its quadratic form is asymptotically χ2.
  • Adaptivity and orthogonality: Adaptivity requires that replacing η0 with η̂ changes the empirical equations by o_Pn(n^-1/2), so nuisance estimation is first-order negligible.This can hold even when η̂ is non-regular and not asymptotically linear.
  • Adaptivity and orthogonality: Orthogonality makes the equations locally insensitive to nuisance perturbations, while additional estimation-quality and complexity conditions control the remaining empirical-process terms.The paper notes that η̂ often needs to converge faster than n^-1/4, but that rate alone is insufficient.
  • Testing: Proposition 1 establishes uniformly valid confidence sets after selection or regularization when adaptivity, normality, and variance consistency hold.The results apply uniformly over sequences of probability laws satisfying the stated conditions.
  • Estimation: Adaptive argmin estimators and weighted equations inherit uniformly valid inference, while one-step estimators are first-order equivalent but require stronger regularity conditions.Finite-sample evidence cited by the paper indicates that argmin estimators can work better because one-step estimators may have higher-order biases.

3. Achieving Orthogonality Using Neyman’s Orthogonalization

Neyman orthogonalization constructs estimating equations whose nuisance derivative vanishes at the true parameter, protecting inference for α from first-order nuisance-estimation effects. In likelihood and GMM settings, the construction also yields optimal scores or instruments and supports valid testing and estimation under structured nuisance estimation.

  • Classical likelihood case: Neyman’s construction projects the score for α onto the orthocomplement of the nuisance tangent space, producing orthogonal equations.The paper relates this construction to semiparametric efficiency theory and notes its applicability to high-dimensional finite-dimensional nuisance parameters.
  • Classical likelihood case: The likelihood score is modified as ψ(w_i, α, β) = ∂_αℓ(w_i, α, β) − µ∂_βℓ(w_i, α, β), with η combining β and vec(µ).The matrix µ is the orthogonalization parameter, and its true value solves an auxiliary equation.
  • Orthogonality property: The resulting moment function satisfies the required orthogonality property: its nuisance derivative at the true parameter values is zero.This immunizes the target equation against first-order perturbations from estimating nuisance quantities.
  • Classical likelihood case: Under suitable regularity conditions, Neyman’s orthogonalization holds for quasi-likelihood scores, including misspecified likelihoods and dependent data.The projection-based alternative, however, requires the information matrix equality and therefore does not generally provide valid orthogonalization under misspecification.
  • Inference: C(α) statistics and empirical moment equations using regularized η estimators provide the basis for testing and estimating α.Under the stated high-level conditions, the resulting variance can attain the optimal GMM variance, and the same conclusions apply to the one-step estimator.
  • GMM problems: In GMM, µ premultiplies the original moments to create an orthogonal moment that is also interpretable as an optimal instrument or optimal score.The paper defines η to include the original nuisance parameter and the vectorized orthogonalization matrix.

4. Achieving Adaptivity In Affine-Quadratic Models via Approximate Sparsity

In affine-quadratic models, orthogonality removes the first-order nuisance-estimation term, while approximate sparsity and estimator-quality conditions control the remaining terms. Under these conditions, testing and estimation are adaptive despite high-dimensional regularization.

  • Affine-quadratic model: The affine-quadratic framework is useful because it includes widely used linear models while making the key adaptivity arguments transparent.The paper states that the derivations extend readily to more complicated models.
  • Testing: Orthogonality makes the leading nuisance-estimation term T1,j vanish, but the remainder terms T2,j and T3,j require additional structure.These terms arise from empirical-process variation and the quadratic remainder in the nuisance expansion.
  • Exact sparsity: Under exact sparsity, moderate-deviation, sparse-norm, and estimation-quality conditions, the adaptivity condition holds for affine-quadratic testing.The elementary result specifically combines affine-quadratic structure with orthogonality and sparsity-based bounds.
  • Scope and conditions: The sparsity requirement can be relaxed in special cases using sample splitting, but it appears unavoidable in general.The stated relaxation is s log(p_n)^c/n → 0 for some constant c.
  • Approximate sparsity: Approximate sparsity permits η0 to contain no zero components by decomposing it into a sparse component and a small non-sparse remainder.This structure is presented as more realistic and richer than exact sparsity.
  • Approximate sparsity: With approximate sparsity, the condition s^2 log(p_n)^2/n → 0 and bounds on moderate deviations and second-derivative norms imply adaptive testing.The result requires both sparse and pointwise control of the second-derivative matrix.
  • Estimation: Additional deviation and second-derivative bounds extend adaptivity from testing to estimation in the affine-quadratic model.The paper states that these conditions establish the adaptivity condition for both testing and estimation.

5. Analysis of the IV Model with Very Many Control and Instrumental Variables

The IV analysis develops valid post-selection inference when high-dimensional controls and instruments are handled through orthogonal moments and structured nuisance estimation. Simulations and a logit-demand application show that the resulting procedures retain good inferential behavior and produce more plausible estimates than naive alternatives.

  • Model and assumptions: The model allows many controls and instruments relative to the sample size while keeping the dimension of the endogenous parameter fixed.Informative estimation and inference require restrictions such as approximate sparsity on the high-dimensional nuisance parameter.
  • Orthogonal inference: Orthogonality makes the moment condition relatively insensitive to small errors in the nuisance estimate and immunizes inference against small selection mistakes.This permits valid inference despite non-regular estimation of the nuisance parameter.
  • Validity results: Under the stated approximate-sparsity, smoothness, and regularity conditions, C(α)-statistics and associated confidence sets are uniformly valid.The result applies uniformly over the specified class of probability laws and remains valid with an estimated variance matrix.
  • Simulation illustration: In simulation, Oracle and Double-Selection estimators are centered correctly and approximately normal, whereas naive estimators are centered far from zero.The naive procedures have poor finite-sample approximations because selection errors create many-instrument or omitted-variable bias.
  • Empirical illustration: In the demand application, selection-based price estimates become more elastic and more consistent with the theoretical prediction than baseline estimates.Using the larger variable set reduces the number of products with estimated inelastic demand from 139 to 12.
  • Empirical illustration: The flexible selection methods yield more sensible structural estimates at most a modest cost in increased estimation uncertainty.The application uses augmented controls and instruments because theory does not clearly determine the relevant characteristics, instruments, or functional form.

6. Overview of Related Literature

The literature develops uniformly valid inference after model selection or regularization, especially for low-dimensional targets with high-dimensional nuisance parameters. Related approaches include sparsity-based estimation, orthogonal scores, one-step debiasing, and inference conditional on selected models.

  • Regularization methods: Sparsity-based estimators, including Lasso-type methods, dominate much of the literature on high-dimensional nuisance parameters.Alternative regularization schemes include shrinkage and ridge regression in many-instrument settings.
  • Uniform post-selection inference: Recent work emphasizes uniformly valid inference over large model classes where perfect model selection is impossible.This contrasts with earlier approaches that rely on perfect recovery and then ignore model-selection effects.
  • Orthogonal scores: Orthogonal-score methods support valid confidence regions by inverting Neyman’s C(α) tests and can also yield point estimators through score minimization.These methods were developed for high-dimensional approximately sparse models.
  • Debiasing: One-step debiasing is approximately equivalent to solving orthogonal estimating equations through a Gauss–Newton step.The general framework suggests exact and one-step solutions are first-order asymptotically equivalent, although higher-order differences may remain.
  • Selective inference: Conditional selective-inference methods instead target parameters of a data-dependent pseudo-true model after conditioning on the selection event.This approach is logically distinct from inference on low-dimensional parameters in the presence of high-dimensional nuisance parameters.

Appendix A. The Lasso and Post-Lasso Estimators in the Linear Model

The appendix defines Lasso and Post-Lasso estimators for linear prediction. Lasso uses penalized convex optimization with covariate-specific loadings, while Post-Lasso applies OLS to the variables selected by Lasso.

  • Setup: The linear-model setup considers outcomes y_i and predictor vectors x_i in a high-dimensional prediction model.The predictors are represented as a p-vector for each observation.
  • Lasso: Lasso is defined through a convex penalized optimization program with covariate-specific penalty loadings.The loadings accommodate non-Gaussian, heteroscedastic, and dependent data and preserve rescaling equivariance.
  • Post-Lasso: Post-Lasso estimates coefficients by applying ordinary least squares while constraining coefficients outside the selected support to zero.Equivalently, it is OLS using only regressors selected by Lasso.
  • Selection: The selected support contains exactly the covariates whose estimated Lasso coefficients are nonzero.This support determines the model used by the Post-Lasso estimator.
  • Properties: Lasso targets prediction without overfitting and can attain near-optimal regression-function rates under suitable conditions, whereas ℓ1 regularization creates shrinkage bias.Post-Lasso is intended to reduce some shrinkage bias while retaining the same convergence rate under sensible conditions.
  • Implementation: Practical Lasso performance depends on penalty parameters and loadings chosen so the penalty dominates the score with high probability.The loading and penalty choices require additional arguments when p is large or observations are not i.i.d. Gaussian.

Appendix B. Proofs

The proof begins by fixing an arbitrary sequence of models or data-generating processes indexed by n.

  • Proof setup: The proof considers any sequence {P_n} in the model class.
  • Proof setup: The asymptotic argument is formulated along sequences indexed by sample size n.
  • Proof setup: The opening setup places the subsequent probability statements within this sequence-based framework.

B.1. Proof of Proposition 2.

The proof establishes the estimator’s rate through successive bounds, linearization, and a one-step correction. Orthogonality controls the nuisance-estimation contribution, leading to a root-n approximation under the proposition’s conditions.

  • Rate proof: The proof first decomposes estimation error into nuisance, linearization, and empirical-process terms.The terms I1, I2, and I3 isolate these components in the bound for M(α̂, η0).
  • One-step correction: A one-step estimator ᾱ is shown to remain close to α0 at the root-n scale and to approximate the linearized map.The argument establishes ∥√n(ᾱ − α0)∥ ≲ P_n n^-1/2 under the proposition’s conditions.
  • Linearization: The proof derives a root-n rate for α̂ and then defines a linearization map around α0.The map is bL(α) := M̂(α0, η0) + Γ1(α − α0).
  • Rate proof: Orthogonality ∂η′M(α0, η0) = 0 makes the nuisance contribution sufficiently small under the stated conditions.This is the key step used to obtain the n^-1/2 rate.
  • Conclusion: The final claims follow from the preceding approximation together with the Continuous Mapping Theorem and Lemma 8.

B.2. Proof of Proposition 3.

The proof begins by defining feasible and infeasible one-steps, while omitting analogous verification of the remaining conditions.

  • Step 1 defines the feasible and infeasible “one-steps.”

B.3. Proof of Proposition 4.

The proof establishes successive bounds on auxiliary quantities, derives a vanishing remainder, and transfers root-n equivalence among estimators. It then invokes Proposition 1 and Proposition 2 for uniform inference validity.

  • The proof bounds ∥ˆF∥, ∥ˆFΓ1 − I∥, and ∥ˆF − F∥ using equations (20) and (11).
  • The next step combines the preceding bounds with condition (21) to show that n^1/2D converges in probability to zero.
  • The triangle inequality and earlier steps yield √n∥ˇα − ¯α∥ →P 0 and √n∥ˇα − ˆα∥ →P 0.
  • The conditions of Proposition 1 are declared satisfied, after which its conclusions follow immediately.

B.4. Proof of Lemma 2.

The lemma’s proof expands the relevant expression and uses previously established bounds to conclude the stated result.

  • Previously established bounds are used to complete the proof of Lemma 2.
  • The proof expands the expression using the same approach as the proof of Lemma 3.

B.5. Proof of Lemma 3 and 4.

The proofs control expansion terms through orthogonality, sparsity growth conditions, and norm inequalities, while auxiliary lemmas provide probabilistic tools for the high-dimensional analysis.

  • The expansion separates terms T1,j through T4,j, with T1,j equal to zero by orthogonality.
  • Under s^2 log(p_n)^2/n → 0, T3,j and T4,j vanish in probability at the bound √n s log(p_n)/n.
  • For bounded d and k, the claim follows from the assumed growth conditions after bounding the elements of the relevant derivative terms.
  • The appendix invokes normal quantile notation and cites prior results for moderate deviations and related high-dimensional bounds.
  • Lemma 7 supplies laws of large numbers for large matrices in sparse norms under sub-Gaussian or bounded-coordinate assumptions and corresponding growth restrictions.
  • Lemma 8 transfers convergence to a multivariate standard normal into probability comparisons over convex sets.

Appendix D. Proof of Proposition 5

The appendix verifies the conditions needed for Proposition 5 through a sequence of preparation, nuisance-estimation, and stability arguments. These steps establish root-n control of the target estimator and consistency of variance-related quantities.

  • Proof structure: The proof verifies the assumptions of Lemmas 4 and 5, from which the desired result follows via Propositions 1 and 2 and Lemma 2.The verification is organized into explicit steps covering preparation and conditions for both lemmas.
  • Nuisance estimation: The estimator of the nuisance parameter η satisfies performance bounds with probability tending to one under the stated assumptions.These bounds are obtained by modifying arguments from Belloni et al. (2012).
  • Nuisance estimation: The required modification allows errors to be uncorrelated, rather than mean independent, with control regressors and instruments.The extension also accommodates regressing an estimated response variable on control regressors in the third step of Algorithm 1.
  • Target estimation: |α̂ − α0| ≲P_n n^-1/2, providing the root-n bound needed for the final stability argument.The bound follows after verifying the conditions of Lemma 5 under the regularity conditions.
  • Variance consistency: The estimated variance-related quantity V̂_n converges to V_n because Γ̂_1(η̂) converges to Γ_1 and Ω̂ converges to Ω.The proof separately establishes consistency of Ω̂(α0) and then controls the difference between Ω̂ and its version evaluated at the true target.
  • Remainder control: The proof concludes that the remainder term D converges in probability to zero after applying the stated norm bounds, Markov’s inequality, and the regularity conditions.The argument uses bounds for estimator components and empirical moments under Conditions AS.1, SM, and RF.
Loading 1501.03430v3…