Source-linked AI summary
Difference-in-Differences with Multiple Time Periods
Brantly Callaway, Pedro H. C. Sant'Anna
TL;DR
Staggered DiD applications often involve multiple periods, varying treatment timing, and covariate-dependent trends. This paper develops group-time treatment-effect methods with simultaneous inference and finds evidence that higher minimum wages reduced teen employment.
Problem
Many DiD applications involve multiple periods, varying treatment timing, and non-parallel outcome dynamics associated with observed characteristics.
Method
The paper identifies and estimates group-time average treatment effects, supports regression, weighting, and doubly robust estimands, and aggregates effects across dimensions.
Results
The proposed estimators are asymptotically normal and support computationally convenient, valid simultaneous bootstrap inference; the application finds evidence of reduced teen employment.
Takeaways & Limitations
The framework provides interpretable causal-effect analysis for staggered DiD settings and indicates that increasing the minimum wage may reduce teen employment.
Takeaways & Limitations
The empirical application has important limitations, including pseudo group-time average treatment-effect estimates in pre-treatment periods.
Abstract
from arXiv · showhide
In this article, we consider identification, estimation, and inference procedures for treatment effect parameters using Difference-in-Differences (DiD) with (i) multiple time periods, (ii) variation in treatment timing, and (iii) when the "parallel trends assumption" holds potentially only after conditioning on observed covariates. We show that a family of causal effect parameters are identified in staggered DiD setups, even if differences in observed characteristics create non-parallel outcome dynamics between groups. Our identification results allow one to use outcome regression, inverse probability weighting, or doubly-robust estimands. We also propose different aggregation schemes that can be used to highlight treatment effect heterogeneity across different dimensions as well as to summarize the overall effect of participating in the treatment. We establish the asymptotic properties of the proposed estimators and prove the validity of a computationally convenient bootstrap procedure to conduct asymptotically valid simultaneous (instead of pointwise) inference. Finally, we illustrate the relevance of our proposed tools by analyzing the effect of the minimum wage on teen employment from 2001--2007. Open-source software is available for implementing the proposed methods.
1 Introduction
The paper develops a unified staggered Difference-in-Differences framework for multiple periods, varying treatment timing, and conditional parallel trends. It identifies interpretable group-time effects, supports flexible aggregation and estimation, and provides simultaneous inference robust to treatment-effect heterogeneity and dynamics.
- Motivation and scope: The framework targets DiD applications with multiple periods, variation in treatment timing, and parallel trends that may hold only conditional on observed covariates.It focuses on staggered adoption, where treated units remain treated thereafter.
- Framework: The analysis separates identification, aggregation, and estimation/inference, allowing interpretable causal parameters under arbitrary treatment-effect heterogeneity and dynamic effects.This structure avoids interpreting standard two-way fixed-effects regressions as causal effects in staggered DiD settings.
- Identification: Group-time average treatment effects measure effects for treatment-timing group g at time t without directly restricting heterogeneity across covariates, adoption timing, or post-treatment dynamics.These parameters support learning about heterogeneity and constructing more aggregated causal measures.
- Identification and estimation: The paper provides nonparametric point-identification conditions and three estimand families based on outcome regression, inverse probability weighting, and doubly robust methods.The estimands accommodate covariate-specific trends and are developed for both panel data and repeated cross sections.
- Aggregation: Aggregation schemes summarize many group-time effects while highlighting heterogeneity across groups and periods or producing an overall effect comparable to the two-period, two-group ATT.The framework forms families of aggregate parameters and partial aggregations for selected heterogeneity dimensions.
- Inference: Simultaneous confidence bands asymptotically cover the entire group-time treatment-effect path while accounting for dependence across estimators, improving uncertainty visualization over pointwise intervals.The procedure is computationally convenient and supports asymptotically valid inference.
2 Identification
The section defines staggered treatment adoption and develops group-time average treatment effects as the central causal parameters. Under the paper’s assumptions, these effects are nonparametrically point-identified using outcome regression, inverse probability weighting, or doubly robust methods.
- Treatment setup: Treatment is irreversible: no unit is treated initially, and once treated, it remains treated in subsequent periods.The treatment process is represented by each unit’s first-treatment period G; never-treated units are assigned G = ∞.
- Causal parameters: The framework uses ATT(g, t), the average treatment effect for units first treated in group g, measured at time t.The family of group-time effects accommodates multiple treatment groups and periods without imposing treatment-effect homogeneity across groups or time.
- Causal parameters: The family of ATT(g, t) parameters can reveal how treatment effects vary across groups and evolve over time.Fixing g while varying t studies dynamics for a group, while comparing groups examines heterogeneity across treatment cohorts.
- Identification results: The identifying assumptions allow covariate-specific trends without restricting the relationship between treatment timing and potential outcomes.The paper characterizes this as a weaker alternative to a randomization-based assumption, while noting a trade-off between assumption strength and identification.
- Identification results: Under the paper’s assumptions, group-time average treatment effects are nonparametrically point-identified in the multiple-period, multiple-group design.The identification results support outcome regression, inverse probability weighting, and doubly robust estimands.
3 Summarizing Group-Time Average Treatment Effects
This section develops weighted aggregation schemes for group-time average treatment effects to summarize policy impacts and treatment-effect heterogeneity. It emphasizes interpretable dynamic and overall-effect summaries that avoid problematic TWFE weighting.
- Aggregation framework: Researchers aggregate ATT(g, t) using known or estimable weighting functions chosen to answer policy questions and highlight heterogeneity across groups, time, and exposure length.The framework supports many possible summary parameters, including overall, dynamic, group-specific, and calendar-time aggregations.
- TWFE limitations: TWFE aggregate coefficients can place negative or design-driven weights on underlying effects, potentially producing negative estimates even when treatment effects are positive for every unit.The weights depend on group sizes, treatment timing, and the number of periods, complicating causal interpretation.
- Overall treatment effect: The general-purpose overall-effect parameter averages treatment effects experienced by all units that ever participate, matching the canonical two-period, two-group ATT interpretation.Aggregating ATT(g, t) with nonnegative weights rules out negative-weight problems, although positive weights alone are not sufficient for a reasonable overall summary.
- Partial aggregations: The proposed partial aggregations address exposure dynamics, differences between earlier- and later-treated groups, and cumulative policy effects across groups through a specified time.The discussion assumes no treatment anticipation and availability of a never-treated group.
- Dynamic effects: The proposed dynamic aggregation highlights heterogeneity by exposure length without compositional changes and bypasses the pitfalls of dynamic TWFE specifications.Its comparisons across exposure lengths are not based on dynamic TWFE regressions.
4 Estimation and Inference
The section develops doubly robust estimators for group-time treatment effects and establishes their asymptotic properties. It also proposes multiplier-bootstrap inference yielding asymptotically valid simultaneous confidence bands.
- Estimation: The proposed two-step strategy estimates nuisance functions by group and period, then plugs their fitted values into sample analogues of the ATT(g, t) estimands.The nuisance functions include generalized propensity scores and comparison-group outcome regressions.
- Estimation: The doubly robust approach combines outcome regression and inverse probability weighting, requiring correct specification of either the comparison-group outcome evolution or the propensity score.This provides robustness against model misspecification relative to using outcome regression or inverse probability weighting alone.
- Estimation: The estimators extend earlier doubly robust DiD methods to multiple groups and periods, allow treatment anticipation, and use Hájek-type weights summing to one in finite samples.The finite-sample normalization is described as potentially improving finite-sample properties.
- Inference: The asymptotic analysis focuses on a never-treated comparison group under a large n, fixed T framework, with not-yet-treated comparisons following by symmetric arguments.The theoretical results rely on parametric nuisance-function models and associated regularity assumptions.
- Inference: A multiplier bootstrap procedure provides asymptotically valid inference and relatively straightforward confidence bands that are simultaneously valid across groups and periods.The procedure is stated to have correct asymptotic coverage and to be robust against multiple-testing problems.
5 The Effect of Minimum Wage Policy on Teen Employment
The application compares TWFE with the proposed staggered DiD approach using state-level minimum-wage timing and county-level teen employment data from 2001–2007. Both approaches indicate employment reductions, but the proposed method yields different estimates and reveals dynamic effects, with results requiring caution because of pre-treatment evidence against parallel trends and cross-state treatment heterogeneity.
- Data and design: The study defines treatment groups by the year states first increased minimum wages, while states that did not increase them form the untreated group.The analysis uses county-level teen employment from 2001–2007, when the federal minimum wage was flat at $5.15 per hour.
- Dynamic effects: Dynamic effects become more negative with exposure length: employment falls 2.7% in the first year, 7.1% in the second, 12.5% in the third, and 13.6% in the fourth.These dynamic estimates are robust to balancing group composition across exposure lengths.
- Proposed-method results: Using the proposed approach, the average effect across states that increased the minimum wage is a 3.1% reduction in teen employment.The estimated group-time effects range from 0.9% lower employment in 2006 for states first treated in 2006 to 7.1% lower employment in 2007 for states first treated in 2004.
- Interpretation and limitations: The application suggests that estimation-method choice can produce qualitatively different conclusions, but pre-treatment violations and cross-state treatment-size heterogeneity warrant caution.Some pseudo group-time effects in pre-treatment periods differ significantly from zero, providing suggestive evidence against parallel trends.
6 Conclusion
The paper proposes well-defined group-time treatment effects for staggered Difference-in-Differences and flexible identification, estimation, aggregation, and inference procedures. An application to minimum-wage increases shows evidence of reduced teen employment and differences from two-way fixed-effects results.
- 6 Conclusion: The proposed ATT(g, t) measures the average treatment effect in period t for units first treated in period g.Unlike a post-treatment dummy in a two-way fixed-effects regression, ATT(g, t) corresponds to a well-defined treatment effect parameter.
- 6 Conclusion: ATT(g, t) can be aggregated to summarize heterogeneity by dimensions such as treatment exposure or to obtain an overall treatment effect.
- 6 Conclusion: The methodology accommodates covariate-conditional parallel trends, never-treated or not-yet-treated comparison groups, and anticipated treatment participation.These features allow units to adjust behavior before treatment implementation.
- 6 Conclusion: Nonparametric identification supports outcome regression, inverse probability weighting, and doubly robust estimands, while estimators are consistent and asymptotically normal.A multiplier bootstrap provides simultaneous confidence bands for ATT(g, t), with generally low computational costs and implementation available in the R did package.
- 6 Conclusion: The minimum-wage application found some evidence of reduced teen employment and notable differences from the conventional two-way fixed-effects approach.The authors therefore recommend that applied researchers consider methods robust to treatment-effect heterogeneity and dynamics.
Appendix A: Proofs of Main Results
Appendix A proves the paper’s main identification and asymptotic results using auxiliary lemmas. The proofs establish alternative control-group representations of ATT(g, t), point identification, and central-limit-theorem-based inference results.
- Auxiliary lemmas: Auxiliary Lemma A.1 represents ATT(g, t) as a conditional outcome change difference between group g and never-treated units.This holds under Assumptions 1, 2, 3, 4, and 6 for g ∈Gδ and t ∈{2, . . . T −δ} with t ≥g −δ.
- Auxiliary lemmas: Auxiliary Lemma A.2 instead represents ATT(g, t) using units untreated at t + δ from groups not yet treated.The result relies on Assumptions 1, 2, 3, 5, and 6 and applies when g −δ ≤t < ¯g.
- Main theorem proofs: Theorem 2 establishes point identification for all admissible group-time pairs and derives asymptotic linear representations for the proposed estimators.The proof invokes Theorem A.1(a) of Sant’Anna and Zhao (2020) and the Lindeberg–Lévy central limit theorem.
- Main theorem proofs: Theorem 3’s proof applies a conditional multiplier central limit theorem and continuous mapping arguments to establish the limiting distribution of the estimator vector.The covariance is expressed through the expectation of products of influence-function terms.
Appendix B: Additional Results for Repeated Cross Sections
This appendix extends the paper’s identification results from panel data to repeated cross sections. Under time-specific random sampling and compositional stability, it establishes identification results that support analogous estimation and aggregation procedures.
- Data and setup: The appendix extends the identification results to repeated cross-section data, where each pooled-sample unit is observed in one time period with outcomes, treatment timing, covariates, and group indicators.The observed data are represented as (Y, G2, . . . , GT , C, T, X).
- Data and setup: The sampling assumption requires independent random samples for each period and makes (G2, . . . , GT , C, X) invariant to observation time, ruling out compositional changes.Conditional on T = t, observations are independently and identically distributed from the corresponding time-specific distribution.
- Identification and estimation: Stationarity allows mixture-distribution draws to estimate generalized propensity scores because covariates, group indicators, and treatment status are observed for all units.The appendix defines outcome regressions and weights under the mixture distribution for subsequent estimands.
- Identification and estimation: The identification results imply a two-step estimator for ATT(g, t) with repeated cross sections, with analogous asymptotic arguments and aggregation into summary causal-effect measures.The proposed estimators parallel the panel-data procedure and can be aggregated as in the main text.