Source-linked AI summary
Double/Debiased Machine Learning for Treatment and Causal Parameters
Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, James Robins
TL;DR
The paper addresses the difficulty of using prediction-focused ML to estimate low-dimensional causal parameters when regularization biases nuisance estimates. It combines orthogonal scores with auxiliary ML predictions and cross-fitting, yielding root-n inference that is approximately unbiased and normal under supported conditions. The method supports flexible ML estimators, while efficiency may be partly sacrificed by regularizing the score.
Problem
Prediction-oriented ML can yield poorly behaved causal-parameter estimators because regularization biases nuisance-function estimates, while classical Donsker conditions are unsuitable for high-dimensional settings.
Method
The paper combines Neyman-orthogonal scores, auxiliary ML prediction problems, and cross-fitting to construct debiased estimators using flexible nuisance-function learners.
Results
The resulting estimators support root-N-consistent estimation and valid inference for low-dimensional parameters, with approximately unbiased and normal behavior under the paper’s conditions.
Takeaways & Limitations
Double machine learning provides a general procedure for causal and treatment-effect inference with highly complex nuisance functions and broad classes of ML methods.
Takeaways & Limitations
Using a regularized score can reduce efficiency, although it may provide additional robustness by requiring weaker conditions than full-efficiency estimation.
Abstract
from arXiv · showhide
Most modern supervised statistical/machine learning (ML) methods are explicitly designed to solve prediction problems very well. Achieving this goal does not imply that these methods automatically deliver good estimators of causal parameters. Examples of such parameters include individual regression coefficients, average treatment effects, average lifts, and demand or supply elasticities. In fact, estimates of such causal parameters obtained via naively plugging ML estimators into estimating equations for such parameters can behave very poorly due to the regularization bias. Fortunately, this regularization bias can be removed by solving auxiliary prediction problems via ML tools. Specifically, we can form an orthogonal score for the target low-dimensional parameter by combining auxiliary and main ML predictions. The score is then used to build a de-biased estimator of the target parameter which typically will converge at the fastest possible 1/root(n) rate and be approximately unbiased and normal, and from which valid confidence intervals for these parameters of interest may be constructed. The resulting method thus could be called a "double ML" method because it relies on estimating primary and auxiliary predictive models. In order to avoid overfitting, our construction also makes use of the K-fold sample splitting, which we call cross-fitting. This allows us to use a very broad set of ML predictive methods in solving the auxiliary and main prediction problems, such as random forest, lasso, ridge, deep neural nets, boosted trees, as well as various hybrids and aggregators of these methods.
1. INTRODUCTION AND MOTIVATION
The paper develops double/debiased machine learning for root-N estimation and valid inference on low-dimensional causal parameters with highly complex nuisance functions. Orthogonal scores and cross-fitting address regularization bias and overfitting while accommodating flexible ML methods.
- Motivation: The paper targets root-N-consistent estimation and valid inference for low-dimensional causal or treatment-effect parameters with high-dimensional nuisance functions.The nuisance functions may be estimated with random forests, lasso, neural nets, boosted trees, hybrids, and ensembles.
- Motivation: In the partially linear model, the main coefficient θ0 represents a treatment-effect or lift parameter, while g0 and m0 capture outcome structure and treatment confounding.The controls X may be high-dimensional relative to the sample size.
- Regularization Bias: Naively plugging regularized ML estimates into the main regression can produce substantive bias because regularization keeps variance controlled while slowing nuisance-function convergence below n^-1/2.The resulting scaled bias can diverge because the relevant terms are not centered at zero.
- Double Machine Learning: Double machine learning adds an auxiliary prediction problem, approximately partialling out treatment confounding and removing regularization bias from the target estimator.The construction links orthogonalization to instrumental-variables and debiased-lasso interpretations.
- Cross-Fitting: Cross-fitting swaps main and auxiliary samples across splits and averages the resulting estimates, reducing overfitting bias while recovering efficiency otherwise lost through sample splitting.The paper presents cross-fitting as a practical way to handle highly adaptive ML methods without strong entropy conditions.
- Orthogonalization: Neyman orthogonality makes the score locally insensitive to nuisance estimates, so the leading regularization effect depends on products of nuisance-estimation errors rather than their individual first-order errors.This permits useful inference even when nuisance functions converge relatively slowly.
- High-Complexity Settings: Classical Donsker conditions are unsuitable for settings whose model complexity or covariate dimension grows with sample size, motivating the paper’s high-complexity analysis.The paper notes that Donsker conditions can exclude even high-dimensional linear models.
2. CONSTRUCTION OF NEYMAN ORTHOGONAL SCORE/MOMENT FUNCTIONS
This section develops Neyman orthogonal scores for finite- and infinite-dimensional nuisance parameters across likelihood, GMM, and conditional-moment settings. These scores reduce sensitivity to nuisance estimation, while efficiency may require stronger conditions or additional variance-function estimation.
- The framework constructs orthogonal scores for low-dimensional parameters with nuisance parameters estimated in broad likelihood, GMM, and functional settings.The construction applies beyond likelihood formulations and includes infinite-dimensional nuisance parameters.
- Neyman orthogonality requires the pathwise derivative of the expected score with respect to nuisance deviations to vanish at the truth.The condition is imposed over a nuisance realization set containing likely values of the nuisance estimator.
- Starting from a nonorthogonal score, the construction transforms it so moment conditions are locally insensitive to nuisance-parameter errors.A quasi-likelihood lemma establishes orthogonality, and analogous results are given for GMM scores.
- Efficiency and orthogonality need not coincide: regularized constructions can deliver near-orthogonality and robustness while sacrificing some efficiency.In a conditional-moment example, an efficient orthogonal score requires estimating the heteroscedasticity function and imposing additional smoothness assumptions.
- Concentrating out nuisance functions yields Neyman orthogonal scores without requiring the criterion function to be a likelihood.When the criterion is the true log-likelihood, this approach also yields an efficient score under the stated regularity conditions.
- The resulting scores can connect to influence-function adjustments and attain semiparametric efficiency under suitable model conditions.The partially linear example illustrates the influence-function route, while the likelihood setting links the score to the semiparametric bound.
3. DML: POST-REGULARIZED INFERENCE BASED ON NEYMAN-ORTHOGONAL ESTIMATING EQUATIONS
DML combines Neyman-orthogonal scores with cross-fitting to estimate low-dimensional parameters while accommodating complex machine-learning nuisance estimators. Under stated regularity and identification conditions, the resulting estimators achieve root-N concentration, approximate Gaussianity, and uniformly valid confidence regions.
- Nuisance estimation: DML estimates nuisance parameters with machine-learning methods suited to approximate sparsity, trees, neural networks, or ensembles.Ensemble methods have guarantees approximately no worse than the best component method.
- Cross-fitting: K-fold cross-fitting estimates nuisance functions on complement folds and evaluates Neyman-orthogonal estimating equations on held-out folds.DML1 and DML2 differ in how fold-specific estimates are aggregated; DML2 often behaves better in finite samples.
- Finite-sample guidance: The choice of K has no asymptotic impact under the conditions, while moderate values such as 4 or 5 have worked better than K = 2 in reported examples and simulations.Larger K provides more observations for estimating high-dimensional nuisance functions.
- Regularity conditions: Neyman orthogonality, smoothness, and identification conditions control sensitivity to nuisance-estimation errors and support the DML theory.The identification condition bounds the singular values of the relevant matrix between c0 and c1.
- Asymptotic properties: DML1 and DML2 concentrate in a 1/√N neighborhood of θ0 and are approximately linear and centered Gaussian under the theorem's conditions.The result applies uniformly over expanding classes of probability distributions, unlike methods not based on orthogonal scores.
- Inference: The asymptotic theory supports confidence intervals and regions that are uniformly valid over the specified model classes.The result extends to consistent variance estimators and to nonlinear-score settings under their additional conditions.
4. INFERENCE IN PARTIALLY LINEAR MODELS
The paper applies DML to partially linear regression and instrumental-variable models, using orthogonal scores to estimate regression coefficients with complex nuisance functions. Under the stated conditions, both versions are first-order equivalent and support uniformly valid inference.
- Partially linear regression: In partially linear regression, θ0 is the coefficient in Y = Dθ0 + g0(X) + U, with conditional exogeneity making it a causal or treatment effect parameter.The treatment interpretation requires D to be conditionally exogenous given covariates.
- Partially linear regression: DML uses orthogonal scores that residualize treatment and outcome-related components through nuisance functions such as g, m, or ℓ.The displayed scores satisfy both the moment condition and the orthogonality condition at the true nuisance parameters.
- Partially linear regression: Cross-fitting broadens the usable ML methods and relaxes sparsity requirements relative to approaches relying on lasso-type estimators without cross-fitting.The paper describes this as a consequence of using held-out data for nuisance estimation and score evaluation.
- Partially linear regression: The partially linear regression DML1 and DML2 estimators are first-order equivalent and have uniformly valid asymptotic confidence regions under the regularity conditions.The theorem permits replacing the asymptotic variance with a consistent estimator.
- Partially linear regression: Under conditional homoscedasticity, the asymptotic variance reduces to the semiparametric efficiency bound for θ.This is a special-case efficiency result rather than a general efficiency guarantee.
- Rate conditions: Cross-fitting weakens the required joint sparsity conditions: if the propensity function is very sparse, the regression function may be comparatively dense, and vice versa.If the propensity function is known or estimated at the N^-1/2 rate, only consistency of the regression-function estimator is needed.
- Partially linear IV models: For partially linear IV models, DML uses orthogonal scores based on residualized instruments and obtains first-order equivalent estimators with uniformly valid inference.The IV model replaces conditional exogeneity with an instrument Z satisfying the stated conditional moment restrictions.
5. INFERENCE ON TREATMENT EFFECTS IN THE INTERACTIVE MODEL
The interactive model uses orthogonal scores with machine-learned outcome and propensity functions to estimate ATE and ATTE, under unconfoundedness and regularity conditions. The resulting DML estimators are first-order equivalent, uniformly valid for inference, and asymptotically efficient.
- Model and target parameters: The model estimates ATE and ATTE with binary treatment using outcome and propensity functions that may be unknown and complicated.The nuisance functions are learned using machine-learning methods.
- Scope: Without unconfoundedness or conditional exogeneity, the corresponding quantities measure association rather than causal treatment effects.They are then called average predictive effect and average predictive effect for the exposed.
- Orthogonal scores: Estimating ATTE requires the treated-group outcome function, propensity function, and a constant treatment probability, but not the untreated-group outcome function.The constant nuisance component simplifies the variance formula without affecting the DML estimators.
- Orthogonal scores: Under unconfoundedness, the ATE and ATTE scores satisfy both the identifying moment condition and Neyman orthogonality.Orthogonality limits the first-order impact of nuisance-function estimation errors.
- Inference: DML1 and DML2 estimators for ATE and ATTE are first-order equivalent and have uniformly valid confidence regions under Assumption 5.1.The variance may be estimated using the estimator defined in the general DML theory.
- Inference: The ATE and ATTE scores are efficient, so both estimators attain the semiparametric efficiency bound.This efficiency result is attributed to the scores used for the two treatment-effect parameters.
- Regularity and rate conditions: Sample splitting weakens the required sparsity condition relative to estimation without sample splitting, allowing one nuisance function to be relatively dense when the other is very sparse.If the propensity score is known, consistency of the outcome-regression estimator is sufficient.
- Inference on LATE: The same DML framework is extended to LATE with binary treatment and binary instrument, yielding first-order-equivalent estimators and uniformly valid inference under Assumption 5.2.The LATE result uses a score specialized to the instrumental-variable setting.
6. EMPIRICAL EXAMPLES
The empirical examples apply DML to randomized and observational treatment settings, using multiple machine-learning methods and cross-fitting schemes. Across applications, estimates are broadly consistent across nuisance estimators, while fold choice affects standard-error precision differently across examples.
- Empirical design: The empirical section studies unemployment insurance bonuses, 401(k) eligibility and participation, and institutions using DML.The applications include randomized, self-selected, and instrumented treatment settings.
- Machine-learning methods: Five nuisance-function estimators include Random Forest, Reg. Tree, Boosting, Lasso, and Neural Net.The study also evaluates Ensemble and Best combinations of these methods.
- Unemployment insurance bonus: The unemployment-bonus estimates are negative and significant at the 5% level across all methods, with no practical difference between standard-error estimators.The result holds for the reported estimation methods and standard-error calculations.
- 401(k) eligibility: The flexible 401(k) eligibility estimates are substantially below the $19,559 no-controls baseline, while results across flexible methods are broadly consistent.The no-controls estimate has an estimated standard error of 1413, but the passage cautions that it is not valid under neglected confounding.
- 401(k) participation: 401(k) participation effects are uniformly positive and statistically significant, and flexible ML estimates are broadly consistent with slightly attenuated linear-IV estimates.The linear-IV specification reports an estimated effect of $13,102 with estimated standard error (1922).
- Cross-fitting and robustness: Five-fold cross-fitting yields lower standard errors than two-fold cross-fitting in the 401(k) examples, but the opposite pattern appears in the institutions example.Adjusting for sample-split variability increases standard errors without changing qualitative conclusions, and nuisance-method choice does not substantively alter conclusions.
A.6. Proof of Lemma 2.1
The proof establishes the lemma by using invertibility of the relevant Jacobian block to identify the unique nuisance value satisfying the required moment condition.
- Proof: Invertibility of Jββ gives the unique solution μ0 in equation (2.10).The proof then invokes equation (2.6) to establish the moment condition.
- Proof: The asserted claim follows because the constructed nuisance value satisfies E[ψ(W; θ0, η0)] = 0.The proof concludes after applying the stated definition or remark.
- Proof: The identity matrix and Kronecker-product notation specify the dimensions and structure of the derivative expression used in the argument.The notation is used to verify the asserted claim after solving for μ0.
A.7. Proof of Lemma 2.2
The proof verifies the required condition by selecting a neighborhood for β and an arbitrary matrix μ, then representing the combined nuisance parameter η accordingly.
- Proof setup: The argument chooses β satisfying the stated norm bound and lets μ be any dθ × dβ matrix.These objects define η as the stacked vector containing β and vec(μ).
- Conclusion: The proof concludes after verifying the required condition for this parameterization.The passage states that the argument completes the lemma.
- Proof setup: The combined nuisance parameter is written as η = (β′, vec(μ)′)′ for the subsequent verification.This parameterization organizes the components needed in the lemma's condition.
A.8. Proof of Lemma 2.3
The proof verifies orthogonality by showing that the derivative of the expected score with respect to the nuisance parameter equals zero.
- Orthogonality condition: The nuisance derivative is zero: ∂η′EP ψ(W; θ0, η0) = [μ0Gβ, EP m(W, θ0, β0)′ ⊗ Idθ×dθ] = 0.This establishes the required orthogonality relation.
- Notation: Idθ×dθ denotes the dθ × dθ identity matrix, while ⊗ denotes the Kronecker product.These definitions specify the matrix structure in the derivative expression.
A.9. Proof of Lemma 2.4
The proof follows the argument of Lemma 2.2 under the stated parameter restrictions and concludes the lemma.
- The proof adopts the strategy of Lemma 2.2 for β satisfying the stated ℓ1-neighborhood condition.It also allows any dθ × k matrix µ and defines η from β and vec(µ).
- The lemma's proof is completed after this analogous argument.
A.10. Proof of Lemma 2.5
The proof introduces a path-indexed score and relies on regularity conditions to justify interchanging expectation and derivatives.
- The proof defines Q(W; θ, r) as the score evaluated along a path from η0(θ) toward η(θ).The path parameter r ranges over [0, 1].
- Regularity conditions justify interchanging EP with ∂θ and ∂θ with ∂r in equation (A.2).
- The argument concludes by completing the lemma's proof.
A.11. Proof of Lemma 2.6
The proof verifies the lemma's required conditions by representing nuisance elements through µ and h, using identities, and showing key terms vanish.
- The proof represents each η ∈ T as a pair (µ, h) with µ integrable and h ∈ H.
- It establishes one required condition using the stated assumptions and another through a preceding identity.The proof specifically invokes equation (2.22) for the latter equality.
- For equation (2.3), the proof takes arbitrary η = (µ, h) ∈ T_N = T and analyzes the resulting terms.
- The decomposition sets I1 = 0 by the argument used in (A.3), while I2 = 0 for the displayed conditional-expectation relation.
- The proof concludes after these verifications.
A.12. Proof of Theorem 3.1 (DML2 case)
The theorem proof establishes uniform bounds through five steps, combining sample splitting, empirical-process arguments, near-orthogonality, and matrix conditioning to obtain the stated result.
- Step 1: The proof reduces uniformity over P ∈ P_N to verifying equation (3.10) for arbitrary sequences P_N ∈ P_N.It first notes that equation (3.11) follows immediately from the assumptions.
- Step 1: The event E_N that all fold-specific estimators lie in T_N has probability at least 1 − KΔ_n = 1 − o(1).This follows from Assumption 3.2 and the union bound because K is fixed and Δ_n = o(1).
- Step 1: The proof combines the stepwise bounds with ρ_N = o(1), then applies the Lindeberg-Feller CLT and Cramer-Wold device to obtain equation (A.4).It subsequently establishes bounds (A.5)–(A.8) in four further steps.
- Step 2: Step 2 combines I1,k = O_PN(N^-1/2) and I2,k = O_PN(r_N) to establish the bound in equation (A.13).The argument uses conditional non-stochasticity of the fold-specific estimator and Assumption 3.2.
- Step 3: Step 3 invokes the Neyman near-orthogonality condition to control the relevant empirical-process expansion.The proof uses Taylor expansion and the imposed λ_N near-orthogonality condition.
A.13. Proof of Theorem 3.1 (DML1 case)
The proof establishes the uniform result by reducing it to (3.10), then applies fixed-K arguments and matrix stability to obtain the asymptotic conclusion.
- Uniform validity over P ∈PN is reduced to proving (3.10), because (3.11) follows immediately from the assumptions.
- Because K is fixed and independent of n, the proof reuses the arguments from Steps 2–5 in Section A.12.
- N −1/2 + rN ⩽ρN = o(1) and Assumption 3.1 keep the singular values of bJ0,k bounded away from zero with probability 1 −o(1).
- The bounds (A.18)–(A.21) are then used uniformly over folds to complete the intermediate estimates.
- The final bound combined with the Lindeberg-Feller CLT and Cramer-Wold device yields (A.17) and completes the proof.
A.14. Proof of Theorem 3.2.
The proof derives uniform asymptotic results for DML estimators by controlling score and nuisance-estimation errors, establishing linearization, and verifying the required assumptions for nonlinear and instrumental-variable scores.
- Theorem 3.2’s proof is uniform over P ∈PN, and its second claim follows immediately from the first and Theorem 3.1.
- Theorem 3.3: Theorem 3.3 reduces the nonlinear case to the linear proof once subsample DML estimators are shown to be approximately linear.
- Lemma 6.3: The nonlinear proof proceeds through preliminary rates, linearization, empirical-process bounds, and a final remainder bound.
- Linearization: Neyman near-orthogonality controls nuisance-estimation effects, while bounded singular values of J0 and Σ0 support inversion and normalization.
- Theorems 4.1 and 4.2: For the instrumental-variable score, Theorem 4.2 follows from earlier DML theorems after verifying orthogonality, Jacobian bounds, moment conditions, and variance nondegeneracy.
- Theorems 4.1 and 4.2: The score ψ(W; θ, η) = (Y −Dθ −g(X))(Z −m(X)) is decomposed into a coefficient term and a remainder term for the verification.