Source-linked AI summary
The State of Applied Econometrics - Causality and Policy Evaluation
Susan Athey, Guido Imbens
TL;DR
Policy evaluation often relies on observational data because randomized experiments can be expensive or infeasible, making causal inference difficult. The paper reviews identification strategies, supplementary analyses, and machine-learning methods for causal inference. It emphasizes tools for assessing credibility, handling many covariates, and improving policy analysis, while excluding policies never implemented in an available dataset and outcomes not observed there.
Problem
Many policy questions lack feasible randomized experiments, while observational data make it difficult to identify causal effects credibly.
Method
The paper synthesizes developments in identification strategies, supplementary analyses, and machine learning for causal inference in policy evaluation.
Results
The review shows how these methods can support more credible policy analysis, including causal analyses with many covariates and flexible models.
Takeaways & Limitations
Applied researchers should use supplementary analyses to assess identification and modeling assumptions alongside primary causal analyses.
Takeaways & Limitations
The review excludes policies never implemented in an available dataset and outcomes not observed in that dataset.
Abstract
from arXiv · showhide
In this paper we discuss recent developments in econometrics that we view as important for empirical researchers working on policy evaluation questions. We focus on three main areas, where in each case we highlight recommendations for applied work. First, we discuss new research on identification strategies in program evaluation, with particular focus on synthetic control methods, regression discontinuity, external validity, and the causal interpretation of regression methods. Second, we discuss various forms of supplementary analyses to make the identification strategies more credible. These include placebo analyses as well as sensitivity and robustness analyses. Third, we discuss recent advances in machine learning methods for causal effects. These advances include methods to adjust for differences between treated and control units in high-dimensional settings, and methods for identifying and estimating heterogeneous treatment effects.
1 Introduction
The review addresses how applied researchers can draw credible causal conclusions about policy effects when randomized experiments are unavailable. It synthesizes identification strategies, supplementary credibility checks, and machine-learning approaches for causal analysis.
- Motivation: Observational policy evaluation is difficult because naive comparisons can confuse correlation with causality.For example, comparing employment in high- and low-minimum-wage states need not estimate the change caused by raising the minimum wage.
- Identification strategies: The review surveys identification strategies for learning causal effects from observational data, including regression discontinuity and synthetic control methods.Regression discontinuity uses treatment assignment thresholds and attributes boundary deviations to treatment under a continuity assumption.
- Identification strategies: It also considers methods for network settings and approaches that combine experimental and observational data.
- Supplementary analyses: Supplementary analyses assess the credibility of identification strategies, modeling choices, and robustness to assumptions.Their purpose is to support and strengthen confidence in the primary analyses rather than to focus on goodness-of-fit measures.
- Machine learning: The review gives special emphasis to combining machine learning with causal analysis for datasets with many covariates and more flexible models.It discusses using these methods to improve policy-evaluation credibility and approach supplementary analyses more systematically.
- Scope: The paper focuses narrowly on causal or design-based methods relevant to policy analysis and emphasizes recommendations for applied work.It also identifies areas of ongoing and open research.
2 New Developments in Program Evaluation
The section reviews major developments in causal inference and program evaluation, spanning potential-outcome foundations, identification strategies, supplementary approaches, and machine-learning methods.
- Foundations: Causal effects compare potential outcomes for the same unit, but only one treatment level and corresponding outcome can be observed.This is the fundamental problem of causal inference, so estimates rely on comparisons across units receiving different treatment levels.
- Foundations: Under unconfoundedness, associational relations conditional on covariates can receive a causal interpretation as conditional average treatment effects.The literature includes matching, weighting, and propensity-score estimators, where e(x) is the conditional probability of treatment given covariates.
- Scope: The article presents a selective review rather than a comprehensive treatment of the econometric literature, excluding instrumental variables and detailed bounds and partial-identification analyses.It also notes that the perspective is subjective and discusses areas with ongoing work and open research questions.
- Machine learning and treatment heterogeneity: The review connects causal-effect estimation with machine learning when there are many covariates, including settings where covariates may outnumber units.It also covers heterogeneous treatment effects and multiple treatment levels under unconfoundedness.
- Identification strategies: Recent identification developments include regression discontinuity, synthetic control methods, causal methods for networks, external validity, and causal interpretations of regression.The review describes synthetic control methods as a particularly important development in program evaluation during the last decade.
2.1 Regression Discontinuity Designs
Regression discontinuity designs identify treatment effects from discontinuities in treatment incentives at a threshold, under comparability of observations just to either side. Applied guidance emphasizes local polynomial estimation, bandwidth selection, credibility checks, and analysis of heterogeneity.
- Design and identification: Regression discontinuity designs exploit threshold-based discontinuities in treatment incentives to evaluate binary-treatment effects.Sharp designs have a treatment-probability jump from zero to one; fuzzy designs have a smaller jump.
- Design and identification: The estimand compares conditional outcome expectations at the threshold, scaled by the discontinuity in treatment probability.In sharp designs the denominator equals one; in fuzzy designs the interpretation is the average effect for compliers at the threshold.
- Design and identification: Identification interprets the estimand locally when individuals just right and left of the threshold are comparable.The resulting effect applies to individuals close to the threshold, or to threshold compliers in the fuzzy case.
- Estimation and inference: Local linear estimation has substantially better finite-sample properties than nonparametric methods that ignore threshold effects and is now standard.Bandwidth selection should target the conditional expectation at the threshold rather than the entire regression function.
- Extensions and illustration: Multiple exogenous variables defining the threshold can permit estimates at different margins and improve analysis of heterogeneous causal effects.The Jacob-Lefgren application finds relatively little evidence of heterogeneity in program estimates.
2.2 Synthetic Control Methods and Difference-In-Differences
Difference-in-differences compares post-treatment differences between treated and control groups after adjusting for their pre-treatment difference. Synthetic control methods extend this design by choosing weighted controls that resemble the treated unit before treatment, while changes-in-changes avoids functional-form assumptions.
- Difference-in-differences estimates treatment effects by subtracting the pre-treatment outcome difference from the post-treatment difference.
- Synthetic control methods replace a single control or simple control average with a weighted average matched to the treated unit’s lagged outcomes and covariates.Weights are chosen to make the weighted controls resemble the treated unit during pre-treatment periods.
- The synthetic control approach has become widely used because it offers a simple improvement over standard control comparisons.
- Synthetic control implementation commonly restricts weights to be non-negative and to sum to one, but these restrictions can hurt fit for units at distributional extremes.Allowing negative weights or weights that sum to another value, or using best-subset regression or LASSO, may improve fit.
- Changes-in-changes provides a nonlinear difference-in-differences model that does not rely on functional-form assumptions.The method identifies the untreated outcome distribution for the treated group and has an efficient estimator based on empirical outcome distributions.
2.3 Estimating Average Treatment Effects under Unconfoundedness in Settings with Multivalued Treatments
Multivalued-treatment evaluation requires estimands and adjustment methods that differ from binary-treatment analyses. Generalized propensity scores and weak unconfoundedness extend dimension reduction and matching-based approaches to multiple treatment levels.
- Research on multivalued treatments addresses a setting that has received substantially less attention than binary-treatment evaluation.
- The natural uniform-policy estimand compares average outcomes if all units were switched from treatment level w1 to treatment level w2.
- Applying binary-treatment methods only to units receiving w1 or w2 generally estimates a conditioned effect rather than the unconditional policy effect.
- Weak unconfoundedness preserves propensity-score dimension reduction for multivalued treatments, even though no single scalar covariate function generally maintains full conditional independence.
- Matching and propensity-score subclassification methods developed for binary treatments can be extended to multivalued treatments without increasing estimator complexity.
2.4 Causal Effects in Networks and Social Interactions
Causal analysis of networks must account for interference, peer effects, network formation, and dependence across units. The literature develops models and inference methods for these settings, but remains broad, fragmented, and dependent on strong identification or asymptotic assumptions.
- Network settings violate the no-interference premise because one unit’s treatment can affect other units’ outcomes.
- The network literature covers many settings and is therefore fragmented; this review discusses only a subset and is explicitly brief and incomplete.
- Peer effects include shared-environment correlations, effects of peers’ characteristics, and effects of peers’ outcomes.
- Identification of peer effects can require very strong assumptions and is unrealistic in many settings.Empirical work has often ruled out some effects to identify others.
- Network formation models support asymptotic approximations by specifying how new units and links relate to the observed network as the sample expands.
- Randomization inference tests network hypotheses by recalculating statistics under alternative treatment assignments rather than relying on large-sample normality.This approach can test hypotheses without invoking large-sample properties of test statistics.
2.5 Randomization Inference and Causal Regressions
Randomization-based analysis provides a causal interpretation of uncertainty and regression beyond standard sampling-based reasoning. The section connects regression coefficients and standard errors to treatment assignment mechanisms while emphasizing the importance of design and population context.
- Randomization inference: Randomization inference derives estimator distributions from treatment assignments and avoids relying on large-sample approximations.It clarifies testing for treatment effects and unbiased estimation through the act of randomization.
- Randomization inference: The randomization perspective informs observational-study interpretation and complications from finite-population inference and clustering.
- Causal regressions: An explicitly causal perspective is especially useful when samples are convenience samples or contain all units in a population rather than random draws.
- Causal regressions: Under complete random assignment of a binary cause, conventional Eicker-Huber-White standard errors have a causal assignment-based interpretation.
- Causal regressions: With a randomly assigned binary cause, additional regressors need not satisfy assumptions about their relation to the outcome.
- Causal regressions: Clustering adjustments can reflect clustered treatment assignment rather than only clustered outcomes or regressors.
2.6 External Validity
External validity asks whether causal estimates apply beyond the studied sample. The paper reviews tests and extrapolation methods for instrumental variables and regression discontinuity designs.
- Causal studies can have strong internal validity yet provide little guarantee that their effects apply to other populations or settings.
- Instrumental variables: Recent approaches assess external validity by testing whether local effects can be generalized across complier, always-taker, and never-taker groups.
- Instrumental-variables estimates may have only local validity for compliers rather than validity for the entire sample.The local average treatment effect is the average effect for individuals whose treatment status is affected by the instrument.
- Regression discontinuity: Regression-discontinuity estimates are principally valid near the forcing-variable threshold, making extrapolation to other population values a central concern.
- Regression discontinuity: Higher-order derivatives can support extrapolation in sharp and fuzzy regression-discontinuity designs without additional covariates, under smoothness assumptions.
- Regression discontinuity: Conditioning on additional exogenous covariates and finding that forcing-variable correlations with outcomes vanish can justify extrapolation away from the threshold.
2.7 Leveraging Experiments
Experimental and observational data can be combined when neither design alone answers the causal question. The paper describes approaches using intermediate outcomes, surrogates, and multiple experiments to extend inference.
- Experiments provide internal validity, while large representative observational studies can provide external validity and precision.
- Intermediate outcomes and surrogates: When an experiment observes only an intermediate outcome, surrogate variables and observational data can help estimate the effect on an unobserved primary outcome.
- Intermediate outcomes and surrogates: Surrogate-based estimation requires surrogacy and comparability of the experimental and observational samples.
- Intermediate outcomes and surrogates: Two methods estimate the average effect by imputing missing experimental outcomes from observational surrogate relationships or imputing observational treatment status from experimental relationships.
- Selection on unobservables: Intermediate outcomes can also help adjust for selection on unobservables by comparing observational estimates with experimental estimates.
- Multiple experiments: Combining multiple experiments can improve efficiency, support prediction in another population, or estimate effects for treatments with different characteristics, though such inferences are not design-validated.
3 Supplementary Analyses
Supplementary analyses probe the identification strategy behind primary causal estimates rather than replace them. They include placebo tests, sensitivity analyses, robustness checks, and design-specific diagnostics.
- Supplementary analyses use implications of identification assumptions in the data to assess the credibility of primary analyses.
- Placebo analyses: Placebo analyses replace the primary outcome with a pseudo outcome known not to be affected by treatment, whose true estimand is zero.
- Sensitivity and robustness: Sensitivity and robustness analyses examine how primary estimands change when critical identifying assumptions are weakened or specifications vary.
- Placebo analyses: In the lottery application, the actual-outcome estimate is −$5,740, while the lagged pseudo-outcome estimate is −$530.
- Placebo analyses: A joint lottery placebo test produced a p-value of 0.135, supporting unconfoundedness in that study.
- Placebo analyses: In the Lalonde application, substantial adjusted differences remained and cast doubt on unconfoundedness.
- Design-specific diagnostics: For regression discontinuity, a density difference of 0.10 with standard error 0.08 provided little evidence of a threshold discontinuity.
4 Machine Learning and Econometrics
Machine-learning methods differ from traditional econometric approaches because they prioritize prediction and data-driven tuning. The paper explains how these tools can support causal analysis, especially with high-dimensional covariates.
- Unsupervised learning: Unsupervised learning reduces covariate dimensionality by replacing many indicators, such as words in documents, with a smaller number of groups.
- Prediction and causal inference: Supervised machine learning predicts outcomes from covariates in new data, whereas causal inference targets counterfactual outcomes under treatment assignments.
- The review focuses on supervised machine-learning methods that can improve causal analysis, particularly in high-dimensional settings.
- Prediction and causal inference: Machine-learning methods commonly trade bias for lower variance and emphasize prediction risk rather than asymptotic normality centered on the estimand.
- Cross-validation: Cross-validation selects tuning parameters by evaluating predictions on held-out subsamples, balancing model complexity, bias, and variance.
4.1 Prediction Problems
This section reviews nonparametric and penalized regression methods for prediction, emphasizing their tuning, regularization, and limitations in high-dimensional settings. It also distinguishes predictive performance from interpretability and causal estimation.
- Nonparametric regression: Nonparametric regression methods include nearest neighbors, kernel regression, and series regression, but their performance deteriorates when covariate dimension is high.Nearest-neighbor methods average observations close in Euclidean distance, while kernel methods weight nearby observations more heavily; both face high-dimensional limitations.
- Penalized regression: Penalized regression adds a parameter penalty to the least-squares objective, producing well-defined estimates even when K > N.The penalty is selected through cross-validation and can yield sparse models under norms such as L0 or L1.
- Penalized regression: LASSO sets some coefficients exactly to zero and shrinks the remainder, whereas ridge regression smoothly shrinks all coefficients toward zero.The two methods correspond to different penalty norms and Bayesian prior interpretations.
- Interpretation and prediction: LASSO can make models easier to interpret even when it does not predict better than ridge regression.Whether sparsity matters depends on the application; it may be less important when the model is used only for prediction.
- Model selection: Cross-validation balances in-sample fit against out-of-sample prediction by selecting the penalty that minimizes mean-squared error on an independent dataset.The in-sample versus out-of-sample fit gap is unobserved when the model is estimated, motivating cross-validation.
- Inference and interpretation: LASSO has formal asymptotic results, but standard confidence intervals require conditions including many regressors having exactly zero true coefficients.The section also notes that causal interpretation of individual LASSO coefficients requires caution because the objective targets prediction rather than unbiased estimation.
4.2 Machine Learning Methods for Average Causal Effects
Machine learning methods extend average-treatment-effect estimation to settings with many covariates. The section contrasts predictive variable selection with procedures designed to preserve valid causal estimation and control imbalance.
- High-dimensional adjustment: Many pretreatment covariates can make unconfoundedness more plausible, motivating machine learning methods that flexibly adjust for high-dimensional confounding.These methods often resemble fixed-dimensional estimators while accommodating substantially more covariates.
- Propensity scores: Propensity-score weighting and matching can be sensitive to specification, especially for units with propensity scores near zero or one.Small changes such as using logit rather than probit models can substantially change weights and produce non-robust estimators.
- Propensity scores: High-dimensional propensity scores can be estimated with random forests, boosting, or LASSO, but weight variability may remain problematic and motivate trimming.Trimming changes the estimand by eliminating observations with extreme estimated propensity scores.
- Predictive versus causal estimation: Ordinary LASSO can omit weak outcome predictors that are important confounders, so its coefficients generally should not receive a causal interpretation.The prediction objective prioritizes outcome correlation and may neglect variables correlated with treatment but only weakly correlated with outcomes.
- Predictive versus causal estimation: Double selection uses LASSO to select covariates related to outcomes and treatment, then includes their union in a final ordinary least-squares regression.This procedure addresses omitted-variable bias and improves estimator properties for average treatment effects.
- Balancing methods: Covariate-balancing estimators reduce extrapolation by reweighting controls to resemble treated units and then using regularized regression for remaining differences.Unlike propensity-score methods, the approach controls bias even when treatment assignment cannot be represented by a sparse model.
4.3 Heterogenous Causal Effects
The section examines methods for discovering and estimating heterogeneous treatment effects while addressing false discoveries and inference after subgroup selection. It emphasizes causal trees and forests with sample splitting and valid confidence intervals.
- Motivation and multiple testing: Heterogeneous treatment effects are relevant for assigning units to optimal treatments, but searching across many covariates and subgroups can generate spurious differences.The problem becomes more severe with many covariates and is linked to multiple testing.
- Motivation and multiple testing: Multiple testing creates false discoveries because some hypotheses are expected to be rejected even when their null hypotheses are true.List et al. propose testing treatment-effect differences across discretized low-versus-high covariate values and correcting for multiplicity.
- Multiple testing: Bootstrap-based testing can improve on standard multiple-testing approaches when covariates are highly correlated.The method accounts for correlation among test statistics, where correlated covariates can induce nearly identical sample divisions.
- Causal trees: Causal trees partition covariate space by treatment-effect heterogeneity, estimate effects within subgroups, and use sample splitting to support nominally valid confidence intervals.One sample builds the partition and the other estimates subgroup treatment effects.
- Causal trees: Causal-tree criteria target treatment-effect prediction rather than outcome prediction and penalize partitions with high treatment-effect variance.Treatment effects are unobserved individually, so the infeasible criterion is estimated using averages within leaves.
- Causal trees: Sample splitting sacrifices half the data for treatment-effect estimation but keeps confidence intervals valid regardless of the number of covariates.It also permits more complex modeling in the second sample after the partition is formed.
- Causal forests: Causal forests provide smooth estimates of τ(x), and their predictions are asymptotically normal and centered on the true conditional average treatment effect.They combine causal trees rather than standard prediction-focused random-forest trees and include variance estimation for confidence intervals.
4.4 Machine Learning Methods with Instrumental Variables
The section considers high-dimensional predictive methods in instrumental-variables settings, where many instruments can make standard procedures unreliable. It discusses regularized first-stage estimation and the associated conditions for valid inference.
- Regularized instrumental variables: LASSO methods can estimate the first and second stages in high-dimensional instrumental-variables settings while providing conditions for valid confidence intervals.The approach applies shrinkage to predictive stages rather than treating the large instrument set with standard methods.
- High-dimensional instruments: In instrumental variables, the first stage predicts endogenous variables using exogenous variables and excluded instruments.This is ordinarily a predictive exercise, but standard methods can perform poorly when the instrument set is large.
- High-dimensional instruments: Many instruments may arise from interactions or flexible transformations, creating settings where conventional instrumental-variables methods have poor properties.Alternative approaches include many-instrument asymptotics, hierarchical Bayes, and random-effects methods.
- Network instruments: In network settings, randomized encouragements can generate many instruments that each weakly affect a particular individual.This creates a distinct many-instrument setting for instrumental-variables analysis.
5 Conclusion
The review highlights newer approaches for estimating policy impacts, stronger supplementary checks on identification credibility, and machine-learning tools for causal inference. It argues these developments can reduce reliance on unnecessary modeling assumptions and increase the credibility of policy analysis.
- The review highlights recently developed approaches for estimating the impact of policies.
- Supplementary analyses help analysts assess the credibility of estimation and identification strategies.The review places greater emphasis on these analyses relative to previous literature.
- The review covers recent developments in machine learning for causal inference, including new estimation methods.
- Machine learning can buttress policy-evaluation credibility when causal strategies require flexible control for many covariates in observational data.
- The literature may help researchers avoid unnecessary functional-form and other modeling assumptions, increasing the credibility of policy analysis.