Source-linked AI summary

Double/Debiased/Neyman Machine Learning of Treatment Effects

Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey

arXiv:1701.08687v1stat.MLstat.ME

TL;DR

The note applies a double machine-learning procedure to ATE and ATTE estimation. Its theorem establishes concentration, approximate normality, and uniformly valid confidence regions, subject to nuisance-rate conditions.

  • Problem

    The note considers inference for the ATE and ATTE as target parameters.

  • Method

    The DML approach uses sample-splitting and a double procedure to estimate treatment-effect parameters.

  • Results

    The estimator concentrates around the target, is approximately unbiased and normally distributed, and supports uniformly asymptotically valid confidence regions.

  • Takeaways & Limitations

    The procedure provides asymptotically valid estimation and inference for the treatment-effect targets considered.

  • Takeaways & Limitations

    The result requires a nuisance-parameter estimation-rate condition, and the conditions needed to attain those rates are not restated.

Abstract

from arXiv · show

Chernozhukov, Chetverikov, Demirer, Duflo, Hansen, and Newey (2016) provide a generic double/de-biased machine learning (DML) approach for obtaining valid inferential statements about focal parameters, using Neyman-orthogonal scores and cross-fitting, in settings where nuisance parameters are estimated using a new generation of nonparametric fitting methods for high-dimensional data, called machine learning methods. In this note, we illustrate the application of this method in the context of estimating average treatment effects (ATE) and average treatment effects on the treated (ATTE) using observational data. A more general discussion and references to the existing literature are available in Chernozhukov, Chetverikov, Demirer, Duflo, Hansen, and Newey (2016).

1. Scores for Average Treatment Effects

The paper formulates ATE and ATTE estimation with Neyman-orthogonal scores that remain sufficiently insensitive to nuisance-estimation errors, while using machine learning for complicated nuisance functions. The approach applies to binary treatment settings with heterogeneous effects under unconfoundedness.

  • ATE and ATTE are estimated under the unconfoundedness assumption using observational data with binary treatment D.
  • The model permits very general heterogeneity in treatment effects because treatment is not additively separable.
  • The propensity score m0(Z) and outcome function g0(D, Z) are unknown potentially complicated nuisance functions estimated with machine learning methods.
  • Neyman-orthogonal scores satisfy an identification condition and orthogonality with respect to nuisance functions, supporting inference after machine-learning estimation.
  • Data splitting is the second critical ingredient enabling use of a wide array of modern machine-learning estimators.
  • For ATE and ATTE, doubly robust or efficient scores are available, and the respective scores obey identification and Neyman orthogonality properties.

2. Algorithm and Result

The DML estimator combines Neyman-orthogonal scores with K-fold cross-fitting to estimate ATE and ATTE using machine-learning nuisance estimators. Under stated moment, overlap, and nuisance-rate conditions, it is approximately unbiased, asymptotically normal, uniformly inferentially valid, and efficient.

  • Algorithm: DML uses K-fold cross-fitting: nuisance estimators are trained on observations outside each fold, then fold-specific estimators are averaged.Sample splitting helps eliminate overfitting bias, while orthogonality reduces regularization and modeling biases.
  • Algorithm: The procedure applies machine-learning estimators of nuisance parameters to orthogonal scores for both the ATE and ATTE.The paper discusses boosted linear and tree models, random forests, ensemble methods, and hybrid machine-learning methods.
  • Conditions: The formal result assumes bounded moments, propensity-score overlap, nondegenerate variation, and rate conditions for the machine-learning nuisance estimators.The nuisance-rate requirement is non-primitive, and the paper does not restate method-specific conditions needed to obtain those rates.
  • Conditions: Cross-fitting weakens the sparsity requirement to sgsm/n, ignoring logarithmic factors, rather than (sg)^2 + (sm)^2 ≪ n without sample splitting.If the propensity score is known, only consistency of the regression estimator is needed; analogous comparisons extend to approximately sparse models.
  • Result: Because the scores are efficient scores, both ATE and ATTE estimators attain the semiparametric efficiency bound of Hahn (1998).This efficiency statement accompanies the asymptotic validity results under the specified class of distributions.

3. Accounting for Uncertainty Due to Sample-Splitting

The note modifies an asymptotically valid procedure to account for finite-sample uncertainty caused by dependence on particular sample splits. It repeats estimation across repartitioned data, aggregates estimates, and incorporates their spread into standard errors.

  • Sample splitting can affect finite-sample estimates even though the particular partition has no asymptotic impact.
  • The proposed modification repeats the main estimation procedure S times, repartitioning the data in each replication.Each partition uses auxiliary samples for nuisance-function estimation and main samples for estimating the parameter of interest.
  • Point estimates can be aggregated across replications using either the sample average or sample median.Aggregation reduces sensitivity to particular random splits; the median is more robust to extreme estimates.
  • The median estimator is more robust to outliers than the mean estimator across random partitions.
  • Standard errors add a component capturing the spread across replications to usual sampling uncertainty.This produces more conservative inference than relying on a single replication's standard error.

Appendix A. Practical Implementation and Empirical Examples

The appendix applies the DML method to observational and randomized empirical examples, examining flexible nuisance estimation, cross-fitting, and uncertainty from repeated sample splits. Results are broadly consistent across machine-learning methods, while precision varies with cross-fitting folds and confounding adjustment.

  • Empirical examples: The appendix considers an observational 401(k) eligibility study and a randomized Pennsylvania Reemployment Bonus experiment.The examples illustrate flexible confounding control in observational data and empirical properties in a randomized setting.
  • Robustness across splits: Mean and median ATE estimates are broadly consistent across flexible methods and similar across repeated sample splits.Repeated estimation uses 100 random repartitions, reporting mean and median estimates with uncertainty measures incorporating sampling and split variability.
  • 401(k) example: 401(k) eligibility is treated as exogenous only after conditioning on income and other variables related to job choice.Eligibility is not randomly assigned, motivating flexible control for potential confounding.
  • Method implementation: Flexible machine-learning methods address the tension between controlling for many confounds and retaining power to learn treatment effects.The appendix uses tree-based methods, Lasso, and neural networks to estimate nuisance functions.
  • 401(k) results: $19,559 is the estimated ATE without controls, whereas estimates flexibly accounting for confounding are substantially attenuated.The uncontrolled estimate has an estimated standard error of 1413 and is not a valid causal estimate if confounding variables are neglected.
  • Cross-fitting and uncertainty: 5-fold cross-fitting produces considerably lower standard errors than 2-fold cross-fitting for all methods, but no general fold-precision relationship is established.The appendix suggests that larger auxiliary samples may improve nuisance-function learning, while explicitly limiting the generality of that interpretation.
  • Pennsylvania Bonus results: The Pennsylvania Bonus Experiment finds a negative and significant ATE on unemployment duration across methods, except random forests in the interactive model, significant at 10%.All other reported methods are significant at the 5% level.

Appendix B. Proofs

Appendix B proves the asymptotic results for the ATE estimator, with the ATTE result following similarly. The proof verifies nuisance-estimation conditions, variance consistency, and Gaussian convergence using empirical-process and limit-theorem arguments.

  • Proof strategy: The appendix proves the result for the ATE estimator and states that the ATTE result follows similarly.The proof begins by choosing an arbitrary sequence of probability measures in the model class.
  • Nuisance conditions: Equations (B.1) and (B.2) provide minimal conditions on nuisance-parameter estimators that can replace more primitive textual conditions.The proof then establishes these assertions step by step.
  • Convergence arguments: Conditional convergence in probability implies unconditional convergence in probability through a simple lemma based on expectation bounds.The lemma also notes that conditional moment convergence yields the result via Markov’s inequality.
  • Asymptotic normality: The proof uses independence across data blocks, the Lindeberg-Feller theorem, and the Cramer-Wold device to obtain a Gaussian limit.The normal limit remains valid when the population standard deviation is replaced by its estimate under the stated conditions.
  • Variance conditions: The proof verifies variance consistency by decomposing the empirical second-moment difference and applying concentration and inequality arguments.It also establishes the required upper and lower variance bounds, completing the proof of the theorem.
Loading 1701.08687v1…