Source-linked AI summary

Doubly Robust Difference-in-Differences Estimators

Pedro H. C. Sant'Anna, Jun B. Zhao

arXiv:1812.01723v3econ.EM

TL;DR

The paper studies robustness and efficiency properties of DID estimators for the ATT. It derives doubly robust estimands and estimators, characterizes efficiency, and develops estimators that can also be doubly robust for inference.

  • Problem

    The paper studies the robustness and efficiency properties of DID estimators for the ATT.

  • Method

    It derives doubly robust ATT estimands and proposes doubly robust DID estimators, including settings with repeated cross-section data.

  • Results

    The proposed estimators attain the semiparametric efficiency bound when all working models are correctly specified, and some are also doubly robust for inference.

  • Takeaways & Limitations

    The paper connects doubly robust DID estimation with semiparametric efficiency and inference robustness.

  • Takeaways & Limitations

    The doubly robust property concerns consistency rather than inference, although some further estimators have doubly robust asymptotic linear representations.

Abstract

from arXiv · show

This article proposes doubly robust estimators for the average treatment effect on the treated (ATT) in difference-in-differences (DID) research designs. In contrast to alternative DID estimators, the proposed estimators are consistent if either (but not necessarily both) a propensity score or outcome regression working models are correctly specified. We also derive the semiparametric efficiency bound for the ATT in DID designs when either panel or repeated cross-section data are available, and show that our proposed estimators attain the semiparametric efficiency bound when the working models are correctly specified. Furthermore, we quantify the potential efficiency gains of having access to panel data instead of repeated cross-section data. Finally, by paying articular attention to the estimation method used to estimate the nuisance parameters, we show that one can sometimes construct doubly robust DID estimators for the ATT that are also doubly robust for inference. Simulation studies and an empirical application illustrate the desirable finite-sample performance of the proposed estimators. Open-source software for implementing the proposed policy evaluation tools is available.

1 Introduction

The paper develops doubly robust DID estimators for the ATT under conditional parallel trends, covering both panel and repeated cross-section data. It derives efficiency bounds, studies panel-data efficiency gains, and proposes estimators that can also be doubly robust for inference.

  • The analysis covers both panel data and repeated cross-section data, whose efficiency bounds incorporate different restrictions implied by the identifying assumptions.
  • DID estimators for the ATT remain consistent when either the propensity score model or the comparison-group outcome-evolution model is correctly specified.
  • Panel data provide efficiency gains over repeated cross-section data, with larger gains when pre- and post-treatment cross-section sample sizes are more imbalanced.
  • When working models are correctly specified, the panel-data DR DID estimator is locally efficient, while a repeated-cross-section estimator attains the efficiency bound when all specified models are correct.
  • The paper cautions that DR consistency does not automatically imply DR inference, because asymptotic variance depends on which nuisance models are correctly specified.
  • Attention to nuisance-parameter estimation can yield DR DID estimators that are computationally simple, locally efficient, and doubly robust for inference.

2 Difference-in-differences

The paper develops doubly robust DID estimands for the ATT under panel or repeated cross-section data. These estimands remain valid when either the propensity score or outcome-regression working model is correctly specified, while efficiency depends on the data structure and model conditions.

  • Data requirements: The repeated-cross-section framework rules out compositional changes in treatment status and covariates across periods.The sampling framework accommodates several sampling schemes but excludes changing composition in (D,X).
  • Identification: The main identification challenge is recovering the treated group’s untreated post-treatment outcome from observed data.The ATT is expressed as the treated group’s observed post-treatment outcome minus its counterfactual untreated outcome.
  • Identification: DID identifies the ATT under conditional parallel trends and overlap assumptions.Conditional parallel trends permits covariate-specific time trends but rules out unit-specific trends; overlap requires treated and untreated units across covariate values.
  • Doubly robust estimands: The proposed DR DID estimands identify the ATT if either, but not necessarily both, the propensity score or outcome regression model is correctly specified.This property applies with either panel or repeated cross-section data and is less demanding than relying exclusively on OR or IPW.
  • Efficiency: Panel data can yield more efficient ATT estimators than repeated cross-section data.The panel-data efficient influence function does not depend on treated-group outcome regressions, unlike the repeated-cross-section case; efficiency loss from cross-sections is convex in λ.

3 Estimation and inference

The paper develops estimators and inference procedures built from the doubly robust DID moments, using first-step nuisance estimates for propensity scores and outcome regressions. Under suitable working models, some estimators are doubly robust for both consistency and inference, with local semiparametric efficiency for selected panel and repeated-cross-section estimators.

  • Estimation strategy: The proposed estimators use a two-step strategy: estimate nuisance functions first, then plug fitted values into sample analogues of the DR DID moments.The nuisance functions include propensity scores and outcome-regression components for panel or repeated-cross-section data.
  • Generic estimators: Generic first-step estimators yield doubly robust and locally semiparametrically efficient panel estimators when working nuisance models are correctly specified.The corresponding asymptotic variance achieves the semiparametric efficiency bound under correct specification.
  • Doubly robust inference: The improved estimators target double robustness for inference, so their asymptotic variance does not depend on which nuisance working model is correctly specified.This removes an estimation effect from first-step nuisance estimation and can produce simpler, more stable inference.
  • Working-model conditions: The improved procedures impose linear outcome-regression and logistic propensity-score working models with symmetric covariate entry.These assumptions are more stringent than those for generic DR estimators but are described as computationally tractable and easier to implement.
  • Panel-data estimator: The improved panel estimator is asymptotically normal and equals the semiparametric efficiency bound when the working models are correctly specified.The construction uses inverse probability tilting for the propensity score and weighted least squares for outcome-regression parameters.
  • Repeated-cross-section estimators: With repeated cross-sections, both proposed estimators are doubly robust for consistency and inference, but only the second is locally semiparametrically efficient.The second estimator attains the efficiency bound, whereas the first does not in general; simulations show this efficiency loss can be large.

4 Monte Carlo simulation study

The simulations evaluate ATT estimators under panel and repeated-cross-section designs, comparing bias, accuracy, inference, and efficiency across alternative model-specification settings. The proposed DR DID estimators generally retain low bias when either the propensity score or outcome regression is correctly specified, while panel data are substantially more efficient than repeated cross-sections.

  • Simulation design: The Monte Carlo study compares DID estimators using bias, RMSE, confidence-interval coverage and length, asymptotic variance, and semiparametric efficiency bounds.The designs use n = 1,000 observations and 10,000 simulations.
  • Panel data: TWFE is severely biased and has almost zero confidence-interval coverage when covariate-specific trends matter, because its estimand is not the ATT.The simulations therefore show that TWFE-based policy evaluations can be misleading in these designs.
  • Estimator stability: Normalized IPW estimators are more stable than unnormalized IPW estimators, while the latter can have substantially larger RMSE and asymptotic variance.The repeated-cross-section simulations also find that standardized IPW estimators are more stable and efficient across designs.
  • Repeated cross-section data: Repeated-cross-section designs have much larger efficiency bounds, RMSEs, asymptotic variances, and confidence-interval lengths than panel designs.The efficiency loss can be striking, and becomes more pronounced when the pre- and post-treatment repeated-cross-section sample sizes are imbalanced.
  • Robustness to misspecification: DR DID estimators show little to no bias when either working model is correctly specified, whereas outcome-regression and IPW estimators become biased when their respective models are misspecified.When both models are correctly specified, semiparametric estimators show little to no Monte Carlo bias and DR estimators can be close to the efficiency bound.
  • Efficiency: DR DID estimators can deliver important efficiency gains over competing estimators, especially when outcome-regression models are correctly specified.Their estimated asymptotic variances are close to the semiparametric efficiency bound when both propensity-score and outcome-regression models are correctly specified.

5 Empirical illustration: the effect of job training on earnings

The empirical illustration uses NSW job-training data with CPS comparison observations to assess evaluation bias against experimental ATT benchmarks. Across samples and specifications, the proposed DR DID estimators combine relatively favorable bias properties with smaller standard errors than IPW estimators.

  • Data and design: The experimental ATT benchmarks are $886, $1794, and $2748 for the LaLonde, DW, and early RA samples, respectively.Their standard errors are $488, $671, and $1005, respectively.
  • Evaluation criterion: Evaluation bias is measured by applying estimators to randomized-out individuals and CPS comparison observations, where consistency implies an ATT of zero.Deviations from zero are labeled evaluation bias.
  • Empirical findings: TWFE estimates are stable across specifications but usually exhibit large positive and statistically significant evaluation biases.Regression-based DID estimators are generally most precise but can be severely biased downward for the LaLonde sample.
  • Empirical findings: Abadie’s IPW estimators tend to have the largest standard errors but relatively small evaluation biases, while normalized weights can improve IPW stability.The proposed DR DID estimators share favorable bias properties with Abadie’s IPW estimator and have smaller standard errors than IPW estimators.
  • Empirical findings: The proposed DR DID estimators are presented as an attractive alternative for evaluating the NSW job-training program.The comparison is made across linear, DW, and augmented DW covariate specifications.

6 Concluding remarks

The paper concludes that its DR DID estimators remain valid under misspecification of one nuisance model and can attain semiparametric efficiency under correct specification. It also identifies extensions to conditional effects and data-adaptive first-step estimation, while leaving their detailed analysis for future work.

  • Main conclusions: The proposed estimators are consistent for the ATT when either the propensity-score or outcome-regression model is correctly specified.They achieve the semiparametric efficiency bound when the nuisance-function working models are correctly specified.
  • Main conclusions: The results cover both panel and repeated-cross-section data and show that nuisance-parameter estimation choices can yield DR estimators that are also DR for inference.The paper illustrates these tools through simulations and an empirical application.
  • Extensions and scope: The proposed estimation and inference tools are not directly applicable to the infinite-dimensional conditional ATT CATT(X1).Combining the DR formulation with Chen and Christensen’s methodology may provide uniformly valid inference for CATT and nonlinear functionals.
  • Extensions and scope: Data-adaptive or machine-learning first-step estimators create technical challenges because they generally belong to non-Donsker function classes.The detailed analysis of these extensions is left to future work.

generic first-step estimators

The paper’s first-step analysis formalizes parametric nuisance models and their estimation for panel and repeated-cross-section data. The assumptions allow pseudo-true parameters under possible misspecification while requiring smoothness, consistency, linear expansions, overlap, and integrability conditions.

  • Data structures: The observed data are W = (Y0,Y1,D,X) for panel data and W = (Y,T,D,X) for repeated cross-sections.The two data structures generate corresponding first-step conditions and parameter vectors.
  • Generic first-step framework: The generic first-step framework represents nuisance functions with parametric models g(x;θ), where θ lies in a compact parameter space.The framework permits misspecification through pseudo-true parameters.
  • Regularity conditions: The assumptions require continuity, twice continuous differentiability near the pseudo-true parameter, a unique interior pseudo-true value, and strong consistency of the estimator.They also require an asymptotically linear expansion for the first-step estimator.
  • Regularity conditions: The framework imposes an overlap condition requiring propensity scores to stay between ε and 1−ε almost surely.This condition holds uniformly over interior propensity-score parameters.
  • Implementation: Under mild moment conditions, the requirements accommodate linear or nonlinear outcome regressions and logit or probit propensity-score models estimated by standard methods.Examples include least squares and quasi-maximum likelihood.

A.1 Panel data case

The panel-data estimator is consistent for the ATT when either the propensity score or comparison-group outcome-evolution model is correctly specified. Under correct nuisance-model specification, it is semiparametrically efficient, while generic variance estimation is not generally doubly robust for inference.

  • Either the propensity score model or the comparison-group outcome-evolution model suffices for consistency of the proposed estimator for the ATT.
  • The estimator has an asymptotically linear representation, is √n-consistent, and is asymptotically normal.
  • When nuisance-function models are correctly specified, the proposed doubly robust DID estimator attains the semiparametric efficiency bound.
  • The generic estimator is doubly robust for consistency but generally not doubly robust for inference because the asymptotic variance depends on which nuisance models are correctly specified.
  • Including all correction terms when estimating the variance is recommended because omitting them may produce asymptotically invalid inference.

A.2 Repeated cross-section data case

For repeated cross-section data, both proposed ATT estimators are consistent under alternative correct working models, but their efficiency and inference properties differ. One estimator attains the efficiency bound under stronger specification requirements than in the panel-data case.

  • Either of the stated working-model conditions is sufficient for consistency of the repeated-cross-section estimators.
  • The second estimator attains the semiparametric efficiency bound, whereas the first does not when the specified propensity-score and outcome conditions hold.
  • Both proposed repeated-cross-section estimators are √n-consistent and asymptotically normal.
  • The repeated-cross-section estimators are doubly robust for consistency but not generally for inference.
  • The second estimator is semiparametrically efficient when the propensity-score model and all treated- and comparison-unit outcome-regression models are correctly specified.
  • These efficiency requirements are stronger than those required when panel data are available.

B Appendix B: Influence function of the DR DID estimators with repeated

Appendix B develops the influence-function components underlying the large-sample properties of the repeated-cross-section doubly robust DID estimators. It also notes that estimating treated-group outcome-regression coefficients creates no estimation effect.

  • The influence functions play a major role in analyzing the large-sample properties of the repeated-cross-section estimators.
  • Appendix B states the precise definitions of the influence functions and the estimation effects associated with nuisance parameters.
  • The repeated-cross-section weighting functions depend on treatment and time variables, with the comparison-group weight also depending on covariates and propensity-score parameters.
  • Estimating outcome-regression coefficients for the treated group does not produce an estimation effect.
Loading 1812.01723v3…