Source-linked AI summary

Design-based Analysis in Difference-In-Differences Settings with Staggered Adoption

Susan Athey, Guido Imbens

arXiv:1808.05293v3econ.EMcs.LGmath.ST

TL;DR

The paper asks how to estimate and conduct inference for average treatment effects when panel units adopt treatment at staggered dates and remain treated. It uses a design-based framework centered on adoption-date assignment and shows that standard DID is unbiased for a weighted average of causal effects under random adoption-date assignment, while standard variance estimators are conservative.

  • Problem

    The paper studies estimation and inference for average treatment effects when units adopt treatment at different dates and remain exposed thereafter.

  • Method

    The paper develops a design-based Difference-In-Differences analysis that models adoption dates as assigned treatments under exclusion restrictions.

  • Results

    Under random adoption-date assignment, the standard DID estimand is a weighted average of different causal effects, and standard variance estimators are conservative.

  • Takeaways & Limitations

    The standard DID estimator has a design-based causal interpretation in staggered-adoption settings, but inference should account for its conservative conventional variance estimates.

  • Takeaways & Limitations

    The analysis relies on strong exclusion restrictions that have no testable implications without additional information.

Abstract

from arXiv · show

In this paper we study estimation of and inference for average treatment effects in a setting with panel data. We focus on the setting where units, e.g., individuals, firms, or states, adopt the policy or treatment of interest at a particular point in time, and then remain exposed to this treatment at all times afterwards. We take a design perspective where we investigate the properties of estimators and procedures given assumptions on the assignment process. We show that under random assignment of the adoption date the standard Difference-In-Differences estimator is is an unbiased estimator of a particular weighted average causal effect. We characterize the proeperties of this estimand, and show that the standard variance estimator is conservative.

1 Introduction

The paper develops a design-based analysis of Difference-In-Differences with staggered, absorbing adoption dates. Under random adoption-date assignment, it characterizes the standard DID estimand and its variance, showing that conventional variance estimators are conservative.

  • The setting has units that adopt treatment at different dates and remain exposed thereafter.
  • The paper studies identification, estimation, and inference under assignment-process restrictions and exclusion restrictions without generally imposing functional-form assumptions.
  • The analysis treats adoption dates as the source of estimator uncertainty rather than sampling units from a larger population.
  • The standard DID estimator can be interpreted as a weighted average of causal effects under a random adoption-date assumption.
  • The paper derives the exact variance of the DID estimator and shows that the Liang-Zeger and clustered-bootstrap variance estimators are conservative.

2 Set Up

The paper formulates staggered adoption through adoption dates as discrete treatment levels, while observed treatment exposure is represented by a binary indicator over time. It defines causal effects by comparing potential outcomes under alternative adoption dates.

  • The observed panel records adoption dates and realized outcomes, with potential outcomes indexed by adoption date.
  • Each unit’s discrete treatment is its adoption date, which may occur in any observation period or never occur.
  • Once adopted, a unit remains exposed to treatment in all subsequent periods.
  • The binary indicator W_it records whether treatment has been adopted by unit i by time t.
  • Average causal effects compare population-average potential outcomes under one adoption date with those under another.
  • The benchmark effect compares never adopting with adopting in the first period, because it contrasts untreated and already-treated potential outcomes.

3 Assumptions

The assumptions separate restrictions on adoption-date assignment from exclusion restrictions on potential outcomes and auxiliary homogeneity or population assumptions. Random adoption enables exact finite-sample results but is strong and not testable under exchangeability without additional information.

  • Assumption sets: The paper distinguishes design assumptions, substantive exclusion restrictions, and auxiliary assumptions about effects, sampling, and outcome models.
  • Design assumption: The design assumption concerns assignment of adoption dates conditional on potential outcomes and possibly pretreatment variables.
  • Limitations: The random adoption assumption is strong and has no testable implications for the joint distribution of outcomes and adoption dates when units are exchangeable.
  • Relaxing random adoption: Additional pretreatment variables can relax complete randomization by requiring random adoption within groups sharing those variables.
  • Random adoption: Under random adoption, the marginal adoption-date distribution and adoption-date shares are fixed, enabling exact finite-sample results for the standard DID estimator.

3.2 Exclusion Restrictions

The section formalizes exclusion restrictions that determine when adoption timing can be treated as a binary exposure and when those restrictions have testable implications.

  • No Anticipation: Assumption 2 rules out effects of future adoption dates on outcomes before treatment begins.It does not restrict the distribution of adoption dates, but may fail when the policy is anticipated before implementation.
  • Invariance to History: Assumption 3 requires outcomes to depend on current exposure rather than how long a unit has been treated.This is described as a strong restriction, though it may be more plausible for repeated cross-sections within clusters such as states.
  • Binary Treatment: Together, Assumptions 2 and 3 imply that treatment can be represented by the binary indicator W(a,t)=1_{a≤t}.The potential outcomes can then be simplified to treated and untreated states.
  • Assumption Status: Assumptions 2 and 3 are substantive exclusion restrictions that cannot be guaranteed by randomizing adoption dates.They are often imposed implicitly when realized outcomes are modeled using only contemporaneous treatment exposure.
  • Testability: Without additional information, the exclusion restrictions have no testable implications because they constrain jointly unobservable potential outcomes.Combined with random assignment, however, they impose testable restrictions when T≥2 and there is variation in adoption timing.

3.3 Auxiliary Assumptions

This section introduces auxiliary assumptions used to simplify treatment-effect structure, variance analysis, and interpretation of the standard DID estimator.

  • Treatment-Effect Homogeneity: Assumption 4 requires treatment effects comparing adoption dates to be constant across units.It restricts heterogeneity across units while allowing the section to impose separate restrictions on time variation.
  • Treatment-Effect Stability: Assumption 5 restricts treatment-effect variation over time and, with Assumptions 2 and 3, yields a constant binary treatment-effect setup.The resulting implication is summarized in Lemma 4.
  • Random Sampling: Assumption 6 treats the sample as a random draw from an infinitely large population over adoption dates and potential outcomes.This supplies additional structure for viewing potential outcomes as random and for analyzing average potential outcomes.
  • Standard DID Setup: The standard DID specification uses additive unit and time effects and implicitly assumes treatment effects are additive and constant across units and periods.The section studies the estimator’s expectation and finite-sample variance under these and related assumptions.
  • Estimator Analysis: The paper interprets DID under randomized adoption dates and proposes a variance estimator smaller than Liang-Zeger and clustered-bootstrap estimators.It also states that the usual random-sampling-based variance is generally larger than the relevant variance.

4.2 The Interpretation of Difference-In-Differences Estimators

Under random adoption dates, the standard DID estimator has a causal interpretation as a weighted combination of adoption-date comparisons, with the interpretation refined by exclusion assumptions.

  • Weighted Representation: The DID estimator can be written as a weighted average of simple estimators for causal effects of changing adoption dates.The representation is mechanical until assumptions on the assignment mechanism support a causal interpretation.
  • Component Comparisons: One component compares never adopting with adoption in the first period, switching exposure from untreated to treated at time t.Its weights sum to one, although some weights may be negative.
  • Component Comparisons: A second component compares never adopting with adoption after time t, so neither potential outcome is treated at time t.The weights for this component also sum to one.
  • Component Comparisons: A third component compares adoption by time t with adoption at the initial time, so both potential outcomes are treated at time t.Its weights are described as summing to one.
  • Causal Interpretation: With no anticipation, DID becomes a weighted average of effects comparing never adopting with adoption at or before t.These effects represent switching from no exposure to exposure, with weights summing to one.

4.3 The Randomization Variance of the Difference-In-Differences Estimators

The section derives the randomization variance of DID under randomized adoption dates and explains why standard variance procedures can be conservative.

  • Variance Derivation: The randomization variance is derived for the DID estimator under randomized adoption dates.The analysis begins from a representation with fixed weights under the assignment mechanism.
  • Variance Components: The exact variance includes covariance terms that are difficult, and generally impossible, to estimate from the available data.Estimating the component variances is straightforward, but inference on the covariance terms is not generally possible.
  • Randomization Versus Sampling: Under sampling-based variance calculations, adoption-date shares are stochastic across samples, generally producing a larger variance.This differs from the randomization distribution, where the weights are fixed.
  • Special Case: In a two-period example, the special-case variance agrees with the Neyman variance for a completely randomized experiment.The example has some units adopting in the second period and others never adopting, with first-period potential outcomes equal to zero.

4.4 Estimating the Randomization Variance of the Difference-In-Differences Estimators

The section develops a conservative randomization-based variance estimator for the DID estimator and compares it with standard clustered approaches. It shows that standard estimators incorporate additional adoption-date-weight variation and generally offer limited room for improvement.

  • Limits of variance estimation: There is no unbiased estimator for the DID estimator’s variance in general.The section notes that the relevant covariance terms are not directly estimable.
  • Source of conservativeness: The estimator’s conservativeness arises partly because nonnegative covariance terms enter the exact variance with a negative sign and are ignored or bounded.Ignoring these terms produces an upwardly biased variance estimator.
  • Proposed variance estimator: The proposed variance estimator is conservative for the DID estimator under the maintained assignment assumption.The theorem establishes that the estimator overstates or equals the randomization variance.
  • Comparison with standard estimators: The Liang-Zeger and clustered-bootstrap estimators are generally more conservative because they also account for variation in adoption-date weights.The randomization scheme fixes adoption-date fractions, whereas these procedures allow those fractions to vary, introducing additional uncertainty.
  • Special cases: In the two-period case, the proposed variance reduces to the Neyman variance from randomized experiments.Small improvements may exploit heteroskedasticity, but the section states that such gains are generally modest.

5 Some Simulations

The simulations compare exact and estimated DID variances across designs that vary potential-outcome dependence and treatment-effect heterogeneity. They evaluate variance estimates through average variance and 95% confidence-interval coverage.

  • Simulation design: The designs vary potential-outcome dependence, treatment-effect patterns, and adoption-date distributions.Designs include constant effects, adoption-date-specific effects, and positive or negative correlations across potential outcomes.
  • Simulation design: The simulations keep potential-outcome sets fixed within a design while repeatedly drawing adoption dates from the design-specific distribution.This preserves the design’s adoption-date fractions across simulations.
  • Evaluation: 95% confidence intervals are formed using Normal-based point estimates plus or minus 1.96 standard errors.The study reports average variance estimates and coverage rates for these intervals.
  • Simulation design: The simulation compares five variances: the exact randomization variance and four estimators.The estimators include the proposed conservative estimator, Liang-Zeger clustering, and two clustered bootstraps.
  • Results: In Design B, Liang-Zeger and the ordinary clustered bootstrap substantially overestimate variance, while the fixed-adoption-date bootstrap and proposed estimator achieve appropriate coverage.The comparison isolates the effect of allowing adoption-date fractions to vary across resamples.

6 Conclusion

The paper develops a design-based analysis of DID with staggered adoption and characterizes both the estimator’s causal estimand and its variance. Under random adoption dates, standard DID averages multiple causal effects, while common variance estimators are unnecessarily conservative.

  • Contribution: The paper develops a design-based approach to DID estimation with staggered adoption.It studies what the standard DID estimator estimates and how its variance behaves under assignment assumptions.
  • Causal interpretation: Under random adoption dates, the standard DID estimand is a weighted average of different causal effects.Examples include switching from never adopting to adopting in the first period or to adopting later.
  • Inference: The standard Liang-Zeger and clustered-bootstrap variance estimators are unnecessarily conservative.The paper proposes an improved variance estimator.

Appendix

The appendix supplies proofs for the paper’s lemmas and theorems, deriving the DID estimand and variance results from the stated assumptions. It also establishes bounds for variance terms that are not directly estimable.

  • Main proofs: Theorem 1’s proof derives the DID estimator’s causal interpretation by applying the paper’s assignment and exclusion assumptions.The proof also obtains special cases when particular causal effects are zero.
  • Variance lemmas: The appendix proves preliminary lemmas characterizing variances, covariances, and components of the DID estimator.These results use finite-population randomization arguments and support the main variance theorems.
  • Variance bounds: The appendix bounds a variance component that is not directly estimable because it enters the variance expression with a negative sign.A lower bound is constructed rather than estimating the component directly.
Loading 1808.05293v3…