Source-linked AI summary

Automatic Debiased Machine Learning of Causal and Structural Effects

Victor Chernozhukov, Whitney K Newey, Rahul Singh

arXiv:1809.05224v5math.STecon.EM

TL;DR

High-dimensional regressions create useful prediction tools but can yield biased and poorly calibrated inference when plugged into causal or structural effects. The paper develops automatic debiasing based on the identifying function and Lasso-estimated Riesz representers, while allowing flexible regression learners. It establishes root-n inference and applies the method to job-training treatment effects and demand elasticities with correlated preferences.

  • Problem

    Causal and structural parameters often depend on high-dimensional regressions, but regularization and model selection can make plug-in inference biased, poorly covered, and not root-n consistent.

  • Method

    The paper constructs automatic debiased machine-learning estimators using orthogonal moments and a Lasso minimum distance learner of the Riesz representer, while permitting flexible regression learners.

  • Results

    The paper establishes root-n consistency, asymptotic normality, and consistent variance estimation for varied causal and structural estimators, then applies them to treatment effects and demand elasticities.

  • Takeaways & Limitations

    Automatic debiasing extends inference to causal and structural effects built from high-dimensional or nonparametric regressions without requiring an explicit bias-correction formula.

  • Takeaways & Limitations

    The method requires a sufficiently fast convergence rate for the Riesz-representer estimator, supported by approximation conditions such as Besov or Hölder structure.

Abstract

from arXiv · show

Many causal and structural effects depend on regressions. Examples include policy effects, average derivatives, regression decompositions, average treatment effects, causal mediation, and parameters of economic structural models. The regressions may be high dimensional, making machine learning useful. Plugging machine learners into identifying equations can lead to poor inference due to bias from regularization and/or model selection. This paper gives automatic debiasing for linear and nonlinear functions of regressions. The debiasing is automatic in using Lasso and the function of interest without the full form of the bias correction. The debiasing can be applied to any regression learner, including neural nets, random forests, Lasso, boosting, and other high dimensional methods. In addition to providing the bias correction we give standard errors that are robust to misspecification, convergence rates for the bias correction, and primitive conditions for asymptotic inference for estimators of a variety of estimators of structural and causal effects. The automatic debiased machine learning is used to estimate the average treatment effect on the treated for the NSW job training data and to estimate demand elasticities from Nielsen scanner data while allowing preferences to be correlated with prices and income.

1. Introduction

The paper develops automatic debiasing for causal and structural parameters that depend on high-dimensional regressions, addressing inference problems caused by regularization and model selection. Its framework extends debiased machine learning across many effects, regression learners, and targeted estimation settings.

  • Motivation: High-dimensional regressions make machine learning useful for estimating policy effects, derivatives, decompositions, treatment effects, and structural parameters.The paper considers settings with many covariates, prices, or other regressors.
  • Inference problem: Regularization and model selection can create bias that makes plug-in confidence intervals poorly covered and estimators non-root-n consistent.The problem arises because prediction-oriented bias-variance trade-offs are unsuitable for inference on causal and structural effects.
  • Automatic debiasing: Neyman orthogonal moments remove the first-order effect of regression estimation, while a Lasso minimum distance learner automatically estimates the required Riesz representer.The construction depends on the identifying moment function rather than a prespecified formula for the representer.
  • Scope: The framework covers linear and nonlinear regression functionals, including policy effects, average derivatives, treatment effects, causal mediation, and regression decompositions.It provides debiased machine learning estimators for effects where such estimators were previously unavailable.
  • Scope: Automatic debiasing avoids requiring an explicit bias-correction formula and supports regression learners including neural nets, random forests, Lasso, and other high-dimensional methods.The paper also develops a targeted version called Auto-TML.
  • Novelty: The paper directly estimates Riesz representers for broad linear and nonlinear functionals in high-dimensional settings without strong Donsker class assumptions.Its sparse linear approximation is restrictive parametrically but general in high-dimensional or nonparametric settings.

2. Average Linear Effects and Orthogonal Moment Functions

The paper formulates linear effects of regression functions through orthogonal moments and a Riesz representer, covering policy effects, derivatives, treatment effects, and economic applications. Identification and inference remain robust to regression approximation when the representer is well approximated, while the framework also extends to nonlinear functionals.

  • The parameter θ0 is an expectation of a known functional m(W, γ) of an observation W and regression γ.
  • The paper extends the framework beyond single linear regression functionals to nonlinear functionals of multiple regressions, including causal mediation and regression decompositions.
  • The framework covers average policy effects, weighted average derivatives, average treatment effects, and bounds on average equivalent variation.The examples use known distribution shifts, treatment propensity scores, or price and income information to define the relevant functionals.
  • The Riesz representer α0 satisfies E[m(W, γ)] = E[α0(X)γ(X)] for square-integrable γ, linking the functional to a regression-weighting function.Its existence corresponds to a finite semiparametric variance bound.
  • Orthogonal moments identify θ0 when either the regression equals the conditional expectation or the Riesz representer is contained in the learner’s function space.This permits identification even when a high-dimensional regression learner converges to a projection rather than E[Y|X].
  • For average treatment effects, the orthogonal moment reduces to the doubly robust moment under nonparametric regression, while high-dimensional settings require approximating inverse propensity components.

3. Estimation

The estimation procedure combines cross-fitting with an automatic Lasso-based learner for the auxiliary function, allowing Auto-DML to use broad classes of regression learners. The section also states key dictionary conditions, describes related targeted estimation, and notes scope boundaries for variance estimation and linear approximation.

  • Cross-fitting: Cross-fitting estimates the regression and auxiliary functions on complementary observations before solving the orthogonal moment equation for the target parameter.In practice, 5-fold or 10-fold cross-fitting is often used.
  • Regression learners: Any regression learner with a mean-square convergence rate that is a power of 1/n can be used to construct Auto-DML.Examples include neural nets, random forests, Lasso, boosting, and other high-dimensional methods.
  • Dictionary conditions: The dictionary must contain possible probability limits of the regression learner and have linear combinations that approximate every function in that class.These conditions link the regression learner and the dictionary used to estimate the auxiliary function, preserving orthogonality.
  • Dictionary conditions: The dictionary needs to approximate the projected auxiliary function rather than the original auxiliary function itself.For the average treatment effect, the relevant projection concerns the difference of inverse propensity scores, not the inverse propensity score directly.
  • Scope and limitations: The approach can use nonlinear dictionary functions for neural nets and random forests, while its variance estimator requires consistency of the regression learner.The paper also notes that linearity of the auxiliary learner may be less appealing in some structural models and that practical comparisons remain open.
  • Auxiliary-function learning: The automatic auxiliary learner uses a penalized minimum-distance construction based on known moments and the dictionary, without requiring the form of the auxiliary function.Its linear representation can also avoid inverting a learned conditional probability or density, potentially reducing instability in high-dimensional settings.

4. Large Sample Inference

The paper establishes convergence rates and asymptotic inference for Auto-DML under approximate sparsity, sparse-eigenvalue or alternative approximation conditions, and rate requirements linking the two learners. Under these conditions, Auto-DML and Auto-TML support root-n inference with consistent variance estimation across several causal and structural examples.

  • General results: The analysis derives mean-square convergence rates for the Lasso learner of ᾱ and uses them to establish root-n consistency and asymptotic normality for Auto-DML.The results also address consistency of the asymptotic variance estimator.
  • Conditions: Approximate sparsity requires a sparse linear combination of dictionary functions to approximate ᾱ, without requiring strict sparsity.The dictionary is assumed to support the relevant approximation, with sparse-eigenvalue and boundedness conditions used in the main rate analysis.
  • Conditions: Theorem 1 provides convergence-rate results for the Lasso minimum-distance learner under the paper’s approximation and regularity assumptions.The theorem is built on assumptions involving the dictionary, approximation rate, regularization, and matrix convergence.
  • Alternative rate result: Theorem 2 shows that, under an alternative condition, the Lasso learner of ᾱ converges faster than n^-1/4+c for every c > 0 without a sparse-eigenvalue condition.This alternative relies on the stated coefficient approximation condition rather than the main sparse-eigenvalue requirement.
  • Inference: Theorem 3 yields asymptotic normality and permits usual asymptotic tests and confidence intervals, while Auto-TML is also consistent and asymptotically normal under similar conditions.The result combines the convergence rates for the auxiliary learner with conditions on the regression learner and variance estimation.
  • Examples: Corollaries establish asymptotic normality for policy effects, average derivatives, and treatment-effect examples under primitive conditions such as boundedness and propensity-score overlap.For the treatment-effect example, the propensity score must be bounded away from 0 and 1.

5. Nonlinear Effects of Multiple Regressions

The section develops Auto-DML for effects that are nonlinear in multiple regressions, including causal mediation. It constructs automatic bias corrections from the identifying function and regression learners, with conditions yielding root-n inference.

  • Scope: Auto-DML covers expectations of nonlinear functions of multiple regressions, including causal mediation and regression decompositions.These effects take the form θ0 = E[m(W, γ)] when m is nonlinear in multiple regression functions.
  • Construction: The estimator uses a separate Lasso learner for each Riesz-representation component and combines their bias corrections across regressions.Each correction uses a Lasso learner multiplied by the corresponding regression residual.
  • Construction: For nonlinear functionals, cross-fitting evaluates derivative-based correction terms independently of the regression learner used to construct them.This supports uniform consistency over a potentially large dictionary using mean-square convergence rates.
  • Construction: The method is automatic because it requires the functional m(w, γ) and regression residuals rather than a manually derived full bias correction.For linear functionals, the derivative construction reduces to evaluating m(Wi, γ) at dictionary elements.
  • Rates and inference: Nonlinear functionals require larger penalty choices than ln(p)/n because derivative estimates depend on machine-learned regressions.A penalty proportional to n^-1/4 generally suffices when the regressions converge faster than n^-1/4.
  • Rates and inference: Although the nonlinear estimator is not doubly robust, it has zero first-order bias and can be root-n consistent and asymptotically normal under regularity conditions.The required correction learner consistently estimates the influence-function component for the nonlinear functional.

6. Regression Decomposition and the Average Treatment Effect on the Treated

This section applies Auto-DML to regression decompositions and the average treatment effect on the treated. The NSW application compares estimates across covariate specifications, comparison samples, and regression learners.

  • ATET: The ATET is the average effect of changing treatment on outcomes conditional on covariates, averaged over the treated subpopulation.Under conditional mean independence of potential outcomes and treatment, the same parameter is the ATET.
  • ATET: Auto-DML estimates the ATET using a Lasso learner of the Riesz representation and a regression learner with cross-fitted residual corrections.The learned weights approximately balance treated and untreated covariate averages.
  • Application: The NSW application studies job-training effects using experimental and quasi-experimental comparison data under common-support restrictions.The outcome is 1978 earnings, and treatment indicates participation in job training.
  • Application: Auto-DML produces comparable NSW estimates across neural-network, random-forest, and Lasso regressions for large covariate sets.The paper reports similar estimates after automatic bias correction for each learner.
  • Application: 1794 (633) is the experimental NSW difference-in-means benchmark, while the corresponding reported estimate for the PSID comparison is 1763 (1026).The paper states that the other results are broadly consistent with these comparisons.

7. Panel Average Derivative and Demand Elasticities

The section applies Auto-DML to panel demand models with household preferences correlated with prices and expenditure. It estimates own-price elasticities for milk and soda using Nielsen scanner data and debiases the resulting plug-in-sensitive estimates.

  • Model: The demand application estimates own-price elasticities while allowing individual preferences to correlate with prices and total expenditure.The model uses a panel structure with correlated random slopes.
  • Model: The panel model represents budget shares through dictionary functions of prices, expenditure, and covariates, with household-specific preference coefficients.Average marginal effects of these functions yield income, own-price, and cross-price elasticities.
  • Estimation: Auto-DML estimates own-price elasticity by applying a smooth transformation to an Auto-DML estimate of the average derivative.Consistency and asymptotic normality follow through the continuous mapping and delta methods.
  • Data: The Nielsen application uses 1483 Houston-area households observed monthly from 2004–2006, with 609 households included throughout the period.The data permit an unbalanced panel, with 12 to 36 monthly observations per household.
  • Results: Allowing correlated random coefficients lowers the milk and soda elasticity estimates relative to cross-sectional estimates, especially for milk.The paper compares Auto-DML estimates with cross-sectional and fixed-effects estimates.
  • Results: Debiased elasticity estimates differ from plug-in estimates by much more than their associated standard errors.The paper presents this comparison as evidence of the importance of debiasing in the application.

8. Conclusions

The paper presents automatic debiasing for parameters depending on high-dimensional or nonparametric regressions and establishes inference under sufficiently fast mean-square convergence.

  • Conclusions: Auto-DML requires only the form of the object of interest, while allowing any regression learner that converges in mean square sufficiently quickly.The paper establishes root-n consistency, asymptotic normality, and consistent asymptotic variance estimation for varied causal and structural parameters.

A.1. Tuning.

The tuning procedure extends equation (3.7) by selecting regularization through analyst-specified hyperparameters and cross-validation. It outputs fold-specific minimum-distance Lasso estimates using a low-dimensional sub-dictionary.

  • Theoretical Procedure: The procedure takes observations excluding fold ℓ, a p-dimensional dictionary b, and a low-dimensional sub-dictionary b_low as inputs.The dictionary includes an intercept, and the output is the estimator ˆρ_ℓ trained on observations outside fold ℓ.
  • Theoretical Procedure: The tuning objective minimizes a generalized Lasso criterion with regularization r_L and coordinate-specific loadings ˆD_ℓ,j.The penalty includes a special intercept term and the loadings are diagonal entries of ˆD_ℓ.
  • Cross-Validation Procedure: The hyperparameters (c1,c2,c3) are analyst-chosen, with practice values (1, 0.1, 0.1), and c1 scales r_L.Cross-validation selects c1 from {5/4, 1, 3/4, 1/2}.
  • Cross-Validation Procedure: Cross-validation chooses c1 by minimizing CV(c1) over {5/4, 1, 3/4, 1/2}.The procedure defines a cross-validated loss and solves the corresponding discrete optimization problem.
  • Cross-Validation Procedure: The iterative tuning procedure is justified by an argument analogous to Algorithm A.1 and Theorem 1 of Belloni et al. (2012).The normalization and regularization formulas are described as analogues of those earlier procedures.

A.2. Optimization.

The optimization section derives coordinate-wise soft-thresholding for the minimum-distance Lasso objective. It abstracts from sample splitting and normalization, then uses convexity to establish convergence to the minimizer.

  • Optimization: The minimum-distance Lasso objective is minimized using a generalized coordinate-descent approach with coordinate-wise soft-thresholding.The method extends standard Lasso coordinate descent to the minimum-distance Lasso objective.
  • Optimization: The derivation abstracts from sample splitting, moment estimation, normalization, and special intercept treatment, and scales the objective by 1/2.This abstraction simplifies the notation used for the update derivation.
  • Optimization: The coordinate update follows by combining the subgradient of the penalty with the subgradient of the full objective.The derivation proceeds from the penalty subgradient, summarizes the objective subgradient, and rearranges it into a component-wise update.
  • Optimization: Coordinate descent converges to the objective minimizer because the differentiable component is convex and the coordinate penalties are convex.The convergence argument invokes the stated convexity conditions.

A.3. Minimum Distance Lasso Using Simulated Data.

A known-truth simulation validates the minimum-distance Lasso implementation by comparing successive methodological choices against the hdm LassoShooting.fit implementation. The final estimator is then used in the paper’s empirical examples.

  • Validation: The minimum-distance Lasso implementation is compared with hdm’s LassoShooting.fit at each methodological point of departure.The comparison covers the formulation, theoretical r_L, normalization ˆD, iteration, and stabilization.
  • Validation: The validation exercise confirms the validity of the techniques introduced in the tuning procedure.The evaluated techniques include the minimum-distance formulation, theoretical regularization, normalization, iteration, and stabilization.
  • Simulation Design: The simulation uses a 101-dimensional design with ground truth ρ0=(1,1,1,0,0,...).The design is used to validate recovery when the true coefficient vector is known.
  • Implementation: The final-row estimator is used in the empirical examples of Sections 6 and 7 and is the estimator defined by the tuning procedure.Before theoretical regularization is introduced, the implementation uses r_L=0.5.

Appendix B. Proofs of Results

The appendix proves estimation-rate and asymptotic-inference results for the debiased procedure under stated approximation, sparsity, moment, and regularity assumptions. The results support asymptotic normality and conventional tests and confidence intervals.

  • Rate Results: The proofs allow m(w,γ) to be nonlinear in γ by permitting ε_n to exceed √(ln(p)/n).These lemmas establish conditions used in proving Theorem 1.
  • Rate Results: The Lasso proof derives bounds for the estimation error using the quadratic form, ℓ1 penalty, first-order conditions, and approximation assumptions.The argument repeatedly applies sparsity restrictions, norm inequalities, and bounds on the empirical moment error.
  • Rate Results: Under Assumption 7, the appendix establishes approximation properties corresponding to ξ=1/2 and uses them to control projection and coefficient errors.The results construct sparse approximations and relate them to the projection of α0(X) on the dictionary.

B.1. Panel Average Derivative and Demand Elasticities.

The paper extends elasticity and derivative estimation through smooth transformations and delta-method variance calculations, with clustered covariance estimators for household-level panel observations.

  • The asymptotic variance of the transformed parameter is derived from the variance of the underlying estimator using the delta method.
  • The covariance estimator recognizes that repeated household observations form clusters.
  • The framework extends with light modification to own-price, income, and cross-price elasticities.
  • Elasticities are obtained as smooth transformations of concatenated derivative components, including scalar and vector components.

C.1. Regression Decomposition and ATET.

Using cross-validation rather than theoretical iteration to tune regularization produces ATET estimates that are broadly similar, with larger standard errors.

  • Cross-validation-based Auto-DML ATET estimates are broadly similar to the main results, with larger standard errors.

C.2. Panel Average Derivative and Demand Elasticities.

The OLS elasticity specification uses expenditure and product prices, with time-averaged regressors and clustered delta-method standard errors for milk elasticity results.

  • The OLS specification concatenates log expenditure and log price for each good as regressors.
  • Time averages of the regressors are used for the household-level covariates.
  • The specification has K = 16 and p = 288 dimensions.
  • Clustered standard errors are calculated using the delta method for the milk elasticity results.
Loading 1809.05224v5…