Source-linked AI summary
Locally Robust Semiparametric Estimation
Victor Chernozhukov, Juan Carlos Escanciano, Hidehiko Ichimura, Whitney K. Newey, James M. Robins
TL;DR
Economic and causal parameters often depend on nonparametric or high-dimensional first steps, creating bias concerns for GMM. The paper constructs orthogonal moments by adding nonparametric influence functions and uses cross-fitting, automatic auxiliary-function estimation, and new doubly robust moments. It shows reduced model-selection and regularization bias, robustness to auxiliary-function misspecification, and general asymptotic inference for high-dimensional applications.
Problem
Economic and causal parameters can depend on nonparametric or high-dimensional first steps, while plug-in GMM can be biased by first-step model selection and regularization.
Method
The paper constructs orthogonal GMM moments by adding the nonparametric influence function to identifying moments, with cross-fitting and automatic estimation of additional unknown functions.
Results
The paper shows that debiased GMM reduces model-selection and regularization bias, remains robust when additional influence-function functions are misspecified, and supports asymptotic inference in high-dimensional quantile and dynamic discrete-choice applications.
Takeaways & Limitations
Orthogonal moments provide a general route to debiased GMM for machine-learning first steps and yield new doubly robust moment equations.
Takeaways & Limitations
Debiased GMM is computationally more complicated because it requires estimating additional unknown functions, and the paper leaves extensions beyond conditional location first steps for future work.
Abstract
from arXiv · showhide
Many economic and causal parameters depend on nonparametric or high dimensional first steps. We give a general construction of locally robust/orthogonal moment functions for GMM, where moment conditions have zero derivative with respect to first steps. We show that orthogonal moment functions can be constructed by adding to identifying moments the nonparametric influence function for the effect of the first step on identifying moments. Orthogonal moments reduce model selection and regularization bias, as is very important in many applications, especially for machine learning first steps. We give debiased machine learning estimators of functionals of high dimensional conditional quantiles and of dynamic discrete choice parameters with high dimensional state variables. We show that adding to identifying moments the nonparametric influence function provides a general construction of orthogonal moments, including regularity conditions, and show that the nonparametric influence function is robust to additional unknown functions on which it depends. We give a general approach to estimating the unknown functions in the nonparametric influence function and use it to automatically debias estimators of functionals of high dimensional conditional location learners. We give a variety of new doubly robust moment equations and characterize double robustness. We give general and simple regularity conditions and apply these for asymptotic inference on functionals of high dimensional regression quantiles and dynamic discrete choice parameters with high dimensional state variables.
MIT and NBER
The paper is classified under JEL codes C13, C14, C21, and D24 and concerns semiparametric estimation with locally robust moments.
- The paper’s keywords include local robustness, orthogonal moments, double robustness, semiparametric estimation, bias, and GMM.
- Its JEL classifications are C13, C14, C21, and D24.
1 Introduction
The paper develops orthogonal, debiased GMM for economic and causal parameters with nonparametric or high-dimensional first steps. It establishes robustness and inference results, develops applications and doubly robust moments, and compares debiased GMM with plug-in GMM.
- Core contribution: The paper constructs orthogonal GMM moments by adding the nonparametric influence function to identifying moments.Cross-fitting evaluates each observation’s moment using first-step estimates trained on other observations.
- Core contribution: Debiased GMM reduces model-selection and regularization bias and is especially relevant when machine-learning first steps use many regressors or state variables.Cross-fitting also avoids the need for Donsker conditions for many machine-learning first steps.
- Comparison with plug-in GMM: Compared with plug-in GMM, debiased GMM supports valid confidence intervals under local model-selection alternatives and root-n consistency under some regularized first steps.The paper states that plug-in GMM conditions are less general and less simple because of an additional first-step-specific remainder.
- Applications: The paper gives debiased estimators for functionals of high-dimensional conditional quantiles and dynamic discrete-choice parameters with high-dimensional state variables.The dynamic discrete-choice estimator uses machine learners for conditional choice probabilities and a novel Lasso estimator of conditional value-function differences.
- Robustness and theory: The nonparametric influence function remains mean-zero when additional functions are misspecified, so estimating those functions need not achieve an n^-1/4 rate.The paper also provides regularity conditions for orthogonality using a Gateaux derivative characterization.
- Robustness and theory: The paper derives new classes of doubly robust moment functions and characterizes double robustness, alongside partial robustness results.The classes include affine functionals of nonparametric regressions, other conditional moment restrictions, and density estimators.
2 Debiased GMM
Debiased GMM augments identifying moments with a nonparametric influence function that removes first-order effects from estimating nonparametric first steps. Cross-fitting and automatic estimation of auxiliary functions support implementation and general asymptotic theory.
- Estimator setup: The estimator targets a finite-dimensional parameter θ whose identifying moments depend on an unknown function γ.Identification requires θ0 to be the unique solution of E[g(W, γ0, θ)] = 0 over Θ.
- Orthogonal moments: The nonparametric influence function φ characterizes the Gateaux derivative of the expected identifying moment with respect to the first-step function.It is defined through local distributional perturbations and can be calculated from the derivative characterization.
- Orthogonal moments: The orthogonal moment is ψ(W, γ, α, θ) = g(W, γ, θ) + φ(W, γ, α, θ), combining identifying moments with an influence-function adjustment.The adjustment removes the first-order effect of estimating γ and α on expected moments.
- Cross-fitting: Cross-fitting evaluates moments on observations excluded from the samples used to estimate γ, α, and initial θ, reducing own-observation bias and avoiding Donsker conditions.The construction partitions observations into groups and trains first-step estimators using observations outside each group.
- Automatic estimation: The paper provides automatic estimators of auxiliary functions in the influence function without requiring the form of α0 to be known.The method uses the orthogonal moments and first step and generalizes an earlier automatic approach beyond least-squares projection functionals.
- Inference and efficiency: Debiased GMM’s efficiency depends on the moment functions, first step, and weighting matrix, with Ψ-hat^-1 identified as the optimal GMM weighting choice.The adjustment does not affect identification and is included to remove the first-order effect of γ-hat.
- Example: A conditional covariance example uses γ0(X) = E[Y|X] and α0(X) = E[Z|X], with θ0 = E[Zγ0(X)] = E[α0(X)γ0(X)].The example motivates the estimator for the component of E[Cov(Z,Y|X)] that depends on unknown functions.
3 Example 3: Dynamic Discrete Choice
The paper develops a debiased GMM estimator for dynamic discrete choice models with high dimensional state variables, using machine learners for conditional choice probabilities. A Monte Carlo study compares debiased and plug-in estimators under increasingly rich Lasso dictionaries.
- Model and estimator: The estimator targets dynamic discrete choice structural parameters through learners of conditional choice probabilities with high dimensional state variables.The model assumes binary choices, stationary first-order Markov states, and a renewal choice structure.
- Model and estimator: The conditional choice probability is modeled as Λ(a(X_t, θ, γ_2, γ_3)), with a combining utility differences and continuation-value terms.The first steps include γ_1(X_t), γ_2(X_t), and γ_3.
- First steps and orthogonalization: The implementation uses cross-fitted Lasso learners, nested sample splitting, and additional estimated functions for the nonparametric influence function.The construction estimates γ_1, γ_2, γ_3 and three corresponding influence-function components.
- Monte Carlo design: The Monte Carlo study uses 500 replications, T = 10, sample sizes n = 100, 300, 1000, and 10,000, and three dictionaries ranging from linear terms to interactions.The richest dictionary includes linear terms, squares, and all pairwise products.
- Monte Carlo results: Across the cases, debiased GMM has much smaller bias than plug-in GMM, while its coverage is close to nominal for the richest dictionary.Plug-in confidence-interval coverage is far from nominal; debiased GMM is no more variable in larger samples or with smaller dictionaries, and is sometimes less variable.
- Monte Carlo results: The bias correction can partially out the effect of estimated first steps in identifying moments, helping explain the estimator’s low variance.The cancellation occurs between the effect of changing γ in identifying moments and its effect in the estimated influence-function term.
4 Neyman Orthogonality
Neyman orthogonality is established by combining identifying moments with a nonparametric influence function whose first-order effect offsets first-step estimation. The results also establish robustness to additional nuisance functions and conditions for small-bias behavior.
- Definition and construction: Neyman orthogonality means that the unknown functions γ and α have no first-order effect on the expected moment.The paper derives this through distributional paths and the zero-mean property of the nonparametric influence function.
- Robustness: The nonparametric influence function can retain mean zero when additional functions on which it depends, and potentially the parameter, are not at their true values.This robustness follows when an alternative distribution generates the selected nuisance function.
- Regularity conditions: Under sufficiently rich pathwise derivatives, the total pathwise condition implies a zero Hadamard derivative with respect to γ.Theorem 3 formalizes this result using a normed linear set and tangential Hadamard differentiability.
- Small-bias property: If the expected orthogonal moment is twice continuously Frechet differentiable, its departure from zero is bounded quadratically in the first-step error.This supports regularity conditions for root-n consistency when the moment is nonlinear in γ.
- Definition and construction: Adding the nonparametric influence function to identifying moments makes the resulting moment function Neyman orthogonal under stated regularity conditions.The influence-function term partials out the effect of varying the first step in the identifying moments.
- Small-bias property: Orthogonal moments shrink faster than the nonparametric first-step bias, transferring the small-bias property to the resulting GMM estimator.The paper notes that mean-square norms make this property broadly applicable to machine-learning first steps.
- Interpretation and novelty: The paper presents the pathwise condition and regularity results as novel, including robustness of the influence function to additional unknown functions.It contrasts this construction with approaches based on semiparametric efficient influence functions or score specifications.
- Interpretation and novelty: The construction is nonparametric and estimator-based, relying on identifying moments and the nonparametric limit of the first-step estimator rather than model specification.Consequently, the orthogonality results do not depend on correct specification of a model.
5 Comparing Debiased and Plug-in GMM
The paper contrasts debiased GMM, which uses orthogonal moments to reduce first-step bias, with plug-in GMM, whose validity depends on restrictive conditions tied to the first-step learner.
- Model selection: Debiased GMM remains valid under first-step model selection, whereas plug-in GMM can have invalid confidence intervals under local alternatives.Model selection can create first-order bias for plug-in GMM, while debiased GMM's bias is second-order.
- Regularity conditions: Plug-in GMM's key sample-orthogonality condition can fail when model selection omits variables or regularization leaves residuals correlated with functions of covariates.Although α0 is not explicitly estimated in plug-in GMM, its form affects whether the required condition holds.
- Special cases: Plug-in GMM is asymptotically equivalent to the target sample average when α0(X) is a linear combination of the same variables as γ0(X), but this need not hold generally.The paper identifies Z = Y as a case where α0(X) = γ0(X) holds a priori.
- Regularization: Regularized first steps can prevent plug-in GMM from being root-n consistent, while debiased GMM can retain root-n consistency under sufficient regularity conditions.For Lasso, the plug-in bias can be of order ln(p), whereas debiased GMM has second-order bias of size ln(p)/n.
- Asymptotic theory: The debiased estimator's regularity conditions are more general and simpler, while plug-in conditions are specific to the first-step estimator and more complicated.For plug-in GMM, verifying the additional condition depends heavily on the nature of the first-step learner.
6 Automatic Estimation of α0
The paper estimates the unknown function α0 using moment conditions derived from orthogonality, then regularizes a sieve approximation for high-dimensional settings. This construction supports debiased GMM for functionals of conditional location learners while accommodating conditional quantiles.
- Method scope: The approach focuses on estimating α0 from a known form of the nonparametric influence function, which is suited to high-dimensional settings where that function may itself be high dimensional.Estimating the entire influence function numerically is discussed as an alternative, but kernel methods are not considered suitable for high-dimensional machine learning settings.
- Moment construction: Orthogonality supplies population moment conditions for α0 by taking the Gateaux derivative of orthogonal moments with respect to first-step perturbations.The resulting conditions are converted into sample moments by replacing expectations with averages and unknown first steps with estimates.
- Regularized estimation: A sieve approximation and Lasso or Dantzig regularization estimate α0 from sample moments formed over choices of δ.The sample moments use observations outside the relevant estimation fold, as required for debiased GMM.
- Location-function applications: For conditional location learners, α0 is the Riesz representer associated with the linear functional E[m(W, γ)].The framework covers conditional means and conditional quantiles through different convex loss functions.
- Regularized estimation: The proposed estimator generalizes earlier minimum-distance constructions to twice differentiable convex losses without requiring an explicit estimator of the loss curvature in a denominator.Curvature is incorporated through a weighted second-moment estimator Q-hat.
- Location-function applications: For conditional quantiles, kernel terms estimate the conditional density at zero, and nested sample splitting supports asymptotic inference using mean-square convergence of the first-step estimator.Here Q = E[f(0|X)b(X)b(X)′].
- Method scope: Constructing α0 for first steps other than conditional location functions is left for future work, including identification and asymptotic theory.The stated scope boundary concerns the automatic construction of α0, not the general orthogonal-moment framework.
7 Double Robustness
The paper characterizes when orthogonal moments are doubly robust and derives new classes by adding nonparametric influence functions to identifying moments. Affineness in the first step yields a central characterization, with applications to conditional moment restrictions and density first steps.
- Definition and characterization: Double robustness means the moment expectation remains zero when one first-step component is incorrect, providing two routes for the moment condition to hold.The paper also notes that doubly robust moments have simpler asymptotic-normality conditions than general debiased GMM.
- Definition and characterization: When the first step space is linear, a moment is doubly robust if and only if its population expectation is affine in the first step.This converts the local orthogonality condition into a global property.
- New constructions: If both the identifying moment and its nonparametric influence function are affine in the first step, their sum is doubly robust.The construction produces new doubly robust moments for conditional moment restrictions and density-valued first steps.
- Conditional moment restrictions: Theorem 8 characterizes double robustness for moments g(W,γ,θ) + α(X)λ(W,γ), including an expected outer-product representation when E[g(W,γ,θ0)] is affine.This form encompasses the familiar doubly robust treatment-effect moment as a special case.
- Conditional moment restrictions: For conditional moment restrictions, the paper derives doubly robust moments of the form m(w,γ) − θ + α(x)[y − γ(z)] when m is linear and α0 represents the associated functional.The result allows Z to differ from X, covering a nonparametric instrumental-variables setting.
- Density first steps: For a density-valued first step, adding its influence function to an affine identifying moment yields a doubly robust moment, including a novel class of such conditions.The paper also reports identification results for functionals even when the underlying first step is not identified.
8 Asymptotic Theory
The asymptotic theory establishes general, simple conditions for debiased GMM using orthogonal moments and cross-fitting. It supports asymptotic normality, variance consistency, double robustness, and applications to conditional quantile and dynamic discrete-choice estimators.
- General theory: The framework accommodates doubly robust moments through alternative small-bias conditions imposed on the first-step estimator.One condition requires a faster than n^-1/4 rate for the first step, while other alternatives exploit affine moment structure.
- General theory: Cross-fitting removes the need for Donsker conditions, which is important for machine-learning first steps.The result differs from earlier asymptotic theory by requiring no Donsker conditions.
- Inference: The asymptotic theory delivers consistency of the estimated variance matrix and asymptotic normality of semiparametric GMM.The variance estimator converges to the asymptotic variance under the stated lemma conditions.
- General theory: The theory uses orthogonal moments and cross-fitting to obtain general asymptotic results under simple mean-square convergence conditions.The conditions separately treat identifying moments and the nonparametric influence function.
- Applications: For conditional quantile functionals, the influence-function component equals the Riesz representer divided by the conditional density at the quantile.Kernel weighting estimates this component without directly inverting an estimated conditional density.
- Applications: The dynamic discrete-choice application derives convergence results for estimated value functions and asymptotic inference under high-dimensional sparse-approximation conditions.The theorem imposes rate and regularity restrictions involving the value-function learner and additional influence-function learners.
9 Appendix A: Proofs of Theorems
The appendix proves the paper’s orthogonality, double-robustness, and asymptotic results by decomposing estimation errors into first-step, influence-function, interaction, and parameter components. It then verifies the required conditions for the paper’s applications.
- Proof strategy: The proofs establish orthogonality by differentiating the influence-function contribution along distributional paths and matching it to the first-step effect.The resulting identity yields the key derivative relation used in the main theory.
- Inference: The proofs establish consistency of the variance estimator and asymptotic normality by applying laws of large numbers, Slutsky’s theorem, and the preceding lemmas.The same framework yields Jacobian convergence and the final semiparametric GMM result.
- Double robustness: Double robustness is characterized through affine dependence of the expected identifying moment on the first step when the additional function is correct.The appendix derives the corresponding moment equation and its converse.
- Proof strategy: The main remainder decomposition separates identifying-moment error, influence-function error, parameter error, and their interaction.This decomposition is used to show that the sample moment has a negligible remainder under the stated assumptions.
- Proof strategy: Cross-fitting makes the estimated first steps conditionally independent of observations used in the moment averages.Conditional inequalities and mean-square bounds then control the remainder terms.
- Applications: The conditional quantile proof verifies the general theorem using the quantile definition, density regularity, and the influence-function representation.The application therefore inherits the general debiased-GMM inference result.
- Applications: The dynamic discrete-choice proof derives value-function rates using maximal inequalities, sparse approximation, and standard plug-in arguments.These rates feed into the theorem for asymptotic inference on structural parameters.
10 Appendix B: Convergence Rate for ˆαℓin Example 2.
Appendix B develops convergence-rate conditions for estimating the additional influence-function components in the conditional quantile application. The analysis uses bounded regressors, approximate sparsity, kernel weighting, and weighted expectations.
- Rate conditions: The appendix imposes bounded-basis and approximate-sparsity conditions for estimating the influence-function component with Lasso-type learners.The sparse approximation error and coefficient size are controlled through Assumptions B1 and B2.
- Kernel estimation: Weighted design matrices and moment vectors are controlled using bounded basis functions and the conditional density weighting.These quantities support the convergence analysis for the automatic estimator.
- Kernel estimation: Kernel weighting targets the weighted expectation involving the conditional density at zero, which determines the relevant convergence rate.The proof changes variables around the residual and controls the resulting terms using mean-value expansions and concentration inequalities.