Source-linked AI summary
Causal mediation analysis with double machine learning
Helmut Farbmacher, Martin Huber, Lukáš Lafférs, Henrika Langen, Martin Spindler
TL;DR
Causal mediation analysis must address confounder selection and identification assumptions when estimating treatment pathways in high-dimensional data. This paper combines efficient-score causal mediation estimators with double machine learning and sample splitting, establishing root-n-consistent, asymptotically normal effects under regularity conditions. In NLSY97 data, health insurance has a moderate short-term effect on general health, but routine checkups do not mediate it.
Problem
Causal mediation analysis requires controlling observed confounders, but researcher-based covariate preselection creates model-selection uncertainty in high-dimensional settings.
Method
The paper combines efficient score functions with double machine learning, sample splitting, and machine-learning estimates for direct and indirect effects.
Results
The proposed effect estimators are root-n consistent and asymptotically normal under specific regularity conditions.
Takeaways & Limitations
In NLSY97 data, health insurance has a moderate short-term effect on general health, but routine checkups do not mediate that effect.
Takeaways & Limitations
The approach relies on selection-on-observables assumptions and regularity conditions, including a constraint on important confounders relative to sample size.
Abstract
from arXiv · showhide
This paper combines causal mediation analysis with double machine learning to control for observed confounders in a data-driven way under a selection-on-observables assumption in a high-dimensional setting. We consider the average indirect effect of a binary treatment operating through an intermediate variable (or mediator) on the causal path between the treatment and the outcome, as well as the unmediated direct effect. Estimation is based on efficient score functions, which possess a multiple robustness property w.r.t. misspecifications of the outcome, mediator, and treatment models. This property is key for selecting these models by double machine learning, which is combined with data splitting to prevent overfitting in the estimation of the effects of interest. We demonstrate that the direct and indirect effect estimators are asymptotically normal and root-n consistent under specific regularity conditions and investigate the finite sample properties of the suggested methods in a simulation study when considering lasso as machine learner. We also provide an empirical application to the U.S. National Longitudinal Survey of Youth, assessing the indirect effect of health insurance coverage on general health operating via routine checkups as mediator, as well as the direct effect. We find a moderate short term effect of health insurance coverage on general health which is, however, not mediated by routine checkups.
1 Introduction
The paper addresses uncertainty in selecting confounders for causal mediation analysis by combining efficient-score methods with double machine learning. It establishes robust, root-n-consistent inference under specific regularity conditions and applies the approach to health insurance and routine checkups.
- Motivation: Researcher-driven covariate preselection creates model-selection uncertainty and can yield incorrect inference when covariates are refined by predictive power.The issue is especially consequential when potential confounders are numerous or exceed the sample size.
- Method: The paper combines causal mediation analysis based on efficient score functions with double machine learning for data-driven control of observed confounders.The framework targets valid inference under selection-on-observables assumptions and specific regularity conditions.
- Method: Neyman orthogonality makes direct and indirect effect estimation less sensitive to local errors in plug-in estimates.The relevant plug-in models include the conditional mean outcome, mediator density, and treatment probability.
- Method: A Bayes-transformed score avoids estimating the conditional mediator density and remains Neyman orthogonal.The alternative is particularly useful for continuous or multidimensional mediators.
- Application: In the NLSY97 application, health insurance has a moderate health-improving direct effect, while the indirect effect through routine checkups is very close to zero.The paper reports no evidence that health insurance affects general health through routine checkups in this application.
2 Definition of direct and indirect effects
The paper decomposes a binary treatment’s average effect into natural direct and indirect effects, distinguishing treatment pathways through a mediator from other mechanisms. It also defines the controlled direct effect, which fixes the mediator at a prescribed value and generally differs from the natural direct effect when treatment–mediator interaction exists.
- Core decomposition: The average treatment effect is decomposed into a direct effect and an indirect effect operating through the mediator.The treatment is binary, and the mediator and outcome are represented through potential variables.
- Natural effects: The natural direct effect switches treatment while holding the potential mediator fixed at the value naturally realized under treatment state d.This blocks the causal mechanism through the mediator.
- Natural effects: The natural indirect effect switches potential mediator values while holding treatment fixed at d, thereby blocking the direct effect.The paper denotes it by δ(d) and defines it using potential outcomes under mediator values induced by treatment states.
- Core decomposition: The average treatment effect equals θ(1)+δ(0) and also θ(0)+δ(1), using natural direct and indirect effects defined across opposite treatment states.These alternative decompositions reflect the treatment state used to define the mediator.
- Effect heterogeneity: Differences between θ(1) and θ(0), or between δ(1) and δ(0), allow heterogeneous effects arising from interaction between treatment and mediator.The paper illustrates this possibility using health insurance, routine checkups, and general health.
- Controlled direct effect: The controlled direct effect γ(m) switches treatment while fixing the mediator at a prescribed value m for the entire population.It equals the natural direct effect only when treatment and mediator do not interact; the practical relevance depends on whether intervening on the mediator is feasible and desirable.
3 Assumptions and identification
Identification relies on conditional independence and common support assumptions given pre-treatment covariates, ruling out relevant treatment, mediator, and outcome confounding. Efficient scores identify counterfactual outcomes and remain consistent when one of three nuisance models is misspecified.
- Assumptions: Conditional independence of treatment requires potential mediators and outcomes to be independent of treatment given covariates X.This rules out confounders jointly affecting treatment and mediator or outcome conditional on X.
- Assumptions: Conditional independence of the mediator requires potential outcomes to be independent of the mediator given treatment D and covariates X.With pre-treatment X, this excludes post-treatment confounders of the mediator–outcome relation.
- Assumptions: The mediator–outcome restriction requires careful scrutiny and becomes less plausible when the treatment-to-mediator interval is large amid time-varying variables.The paper emphasizes this as a practical limitation of the identification strategy.
- Assumptions: Common support requires Pr(D = d|M = m, X = x) > 0 for both treatment states across the support of M and X.This ensures comparable units across treatment and mediator states rather than deterministic treatment or mediation.
- Identification: The efficient score uses the conditional mediator density, treatment probability, and conditional outcome mean to identify E[Y (d, M(1 −d))].The nuisance components are f(M|D, X), pd(X), and µ(D, M, X).
- Identification: The score is multiply robust: estimation remains consistent if one of the three nuisance models is misspecified.This property supports data-driven estimation of mediation effects with flexible nuisance-model specifications.
4 Estimation of the counterfactual with K-fold Cross-Fitting
The paper estimates mediation counterfactuals with efficient scores, K-fold sample splitting, and cross-fitting to control overfitting from machine-learned nuisance models. Under stated regularity and prediction conditions, the resulting estimators are root-n consistent and asymptotically normal.
- Cross-fitting: The estimation strategy combines efficient scores with sample splitting to estimate E[Y (d, M(1 −d))].Nuisance models are fitted on complementary subsamples and evaluated on held-out observations.
- Cross-fitting: K-fold cross-fitting averages estimated efficient scores across held-out observations to obtain the counterfactual estimate in the full sample.The procedure splits W into K subsamples, predicts nuisance models within each held-out fold, and averages the resulting scores.
- Nuisance estimation: The nuisance models estimate the treatment probability, mediator density, and conditional outcome mean using complementary training data.These correspond to pd(X), f(M|D, X), and µ(D, M, X).
- Theory: Neyman orthogonality and linearity of the score support double machine learning when plug-in nuisance estimates converge at rate n^-1/4.The paper notes that lasso, random forests, boosting, and neural networks can achieve this rate under suitable conditions.
- Effects: Total, direct, and indirect effects are obtained as differences between estimated potential outcomes.For example, ˆ∆=ˆΛ1−ˆΛ0, while direct and indirect effects use combinations of ˆΛd and ˆΨd.
5 Simulation study
The simulation study evaluates two effect-estimation approaches using high-dimensional data-generating processes and post-lasso nuisance estimation. Both approaches generally approach the true effects at a root-n rate, with differences depending on sample size and setting.
- Design: The simulations use p = 200 covariates, sample sizes n = 1000 and 4000, and 1000 repetitions per data-generating process.The designs vary covariate importance and confounding strength.
- Design: The study compares estimators based on Theorems 1 and 2, with the latter using a modified score function that avoids conditional mediator densities.Nuisance parameters are estimated by post-lasso regression with 3-fold cross-fitting.
- Results: Asymptotic standard errors decently estimate the actual standard deviation of the point estimators.This assessment is based on simulation results for absolute bias, standard deviation, and root mean squared error of the standard-error estimates.
6 Application
The application uses NLSY97 data to decompose the effect of health insurance coverage on young adults’ general health into direct and routine-checkup-mediated pathways. Insurance has a moderate short-term health-improving effect, but routine checkups do not mediate it.
- Data and design: The study decomposes health insurance coverage’s effect on general health into an indirect pathway through routine checkups and a direct effect through other mechanisms.The application focuses on relatively young individuals and short-term health effects.
- Data and design: The analysis includes 770 control variables, including 601 dummy variables and 252 missing-value dummies, drawn from pre-treatment information.Controls cover demographic, socioeconomic, family, labor-market, and health-related characteristics.
- Results: The total effects are statistically significant at the 5% level, and negative ATEs indicate a short-term health-improving effect of insurance coverage.Direct effects are similar to the ATEs and statistically significant at least at the 10% level in 3 out of 4 cases.
- Results: Indirect effects through routine checkups are generally close to zero and not statistically significant at the 10% level in 3 out of 4 cases.The results suggest that insurance coverage does not importantly affect young adults’ general health through routine checkups in the short run.
7 Conclusion
The paper combines causal mediation analysis with double machine learning for high-dimensional adjustment under selection-on-observables assumptions. It establishes root-n consistency and asymptotic normality under specific conditions, finds favorable finite-sample behavior, and reports a moderate insurance effect not mediated by routine checkups.
- Contribution: The paper proposes estimators for natural direct and indirect effects and the controlled direct effect using efficient score functions, sample splitting, and machine-learning plug-in estimates.The approach avoids ad hoc pre-selection of control variables under selection-on-observables assumptions.
- Theory: The proposed effect estimators are root-n consistent and asymptotically normal under specific regularity conditions.These properties are established for the estimators developed in the paper.
- Evidence: The simulation study finds that the proposed estimators perform well in samples with several thousand observations.The detailed simulation results report convergence toward the true effects at a root-n rate.
- Application: In NLSY97, health insurance coverage has a moderate short-term effect on general health, but routine checkups do not mediate that effect.The application therefore points to other mechanisms for the observed short-term effect.
A Simulation results for standard errors
Appendix A reports simulation results for standard errors, using absolute bias, standard deviation, and root mean squared error as evaluation metrics.
- The simulation evaluates standard-error estimates using absolute bias, standard deviation, and root mean squared error.The reported metrics are abbreviated as abias, sd, and rmse.
B Proofs
The proofs establish the theorem conditions by verifying the cited regularity assumptions for the relevant causal parameters.
- The proofs verify Assumptions 3.1 and 3.2 under Theorems 3.1 and 3.2 from Chernozhukov et al. (2018).
B.1 Proof of Theorem 1
The proof of Theorem 1 verifies score regularity, moment conditions, linearity, continuity, and Neyman orthogonality for the counterfactual parameters. It also establishes nuisance-estimation rate conditions through boundedness, overlap, norm inequalities, and product-rate restrictions.
- Parameter-specific score verification: The proof treats Ψd0 = E[Y(d, M(1−d))] and Ψdm0 = E[Y(d, m)] using score-based arguments.
- Nuisance parameters: The nuisance vector η contains the outcome regression, mediator density, and treatment propensity models.
- Rate conditions: The nuisance estimators must satisfy an L2 product-rate condition of at most δn n^-1/2 for the outcome and mediator components.
- Score properties: Neyman orthogonality follows because derivative terms cancel by Bayes’ Law for the counterfactual score.
- Regularity bounds: The proof uses bounded outcomes and overlap conditions to control norms of the outcome and auxiliary nuisance functions.
B.2 Proof of Theorem 2
The proof of Theorem 2 verifies analogous score properties for the fixed-mediator parameter while incorporating treatment probabilities conditional on the mediator. It establishes the required product-rate and regularity conditions through bounds on the nuisance functions.
- Score construction: The fixed-mediator score is constructed for estimating E[Y(d, m)].
- Nuisance parameters: The nuisance vector includes the outcome regression, auxiliary function, treatment probabilities conditional on mediator and covariates, and marginal propensity.
- Rate conditions: The proof imposes product-rate restrictions linking errors in the outcome, mediator-conditional treatment, auxiliary, and marginal propensity models.
- Score properties: The score is shown to be Neyman orthogonal through derivative cancellations justified by iterated expectations, probability identities, and Bayes’ Law.
- Regularity verification: Bounds for the second-order derivative terms establish score regularity under the stated regularity conditions.