Source-linked AI summary
The role of parallel trends in event study settings: An application to environmental economics
Michelle Marcus, Pedro H. C. Sant'Anna
TL;DR
Staggered DID studies must choose among parallel trends assumptions that differ in strength and identify different causal parameters. The paper compares these assumptions, develops GMM-based estimators, and shows that stronger assumptions trade robustness for efficiency. In the Clean Water Act application, the transition to state management has little to no effect on violation rates, while conclusions about corruption-related heterogeneity depend on the chosen assumption.
Problem
Staggered DID procedures use different parallel trends assumptions and can identify different causal parameters, while TWFE estimates may be weighted averages with negative weights.
Method
The paper compares staggered-DID assumptions and parameters, separates identification from estimation, and develops GMM estimators that exploit restrictions implied by stronger assumptions.
Results
The paper documents a robustness–efficiency trade-off across parallel trends assumptions and finds little to no Clean Water Act effect on violation rates, while corruption-related effects depend on the assumption.
Takeaways & Limitations
Researchers should make the parallel trends assumption explicit because it influences estimator choice, causal parameters, and the interpretation of DID results.
Takeaways & Limitations
TWFE coefficients can have negative weights and may lack a sensible causal interpretation when treatment effects vary across groups and time.
Abstract
from arXiv · showhide
Difference-in-Differences (DID) research designs usually rely on variation of treatment timing such that, after making an appropriate parallel trends assumption, one can identify, estimate, and make inference about causal effects. In practice, however, different DID procedures rely on different parallel trends assumptions (PTA), and recover different causal parameters. In this paper, we focus on staggered DID (also referred as event-studies) and discuss the role played by the PTA in terms of identification and estimation of causal parameters. We document a ``robustness'' vs. ``efficiency'' trade-off in terms of the strength of the underlying PTA, and argue that practitioners should be explicit about these trade-offs whenever using DID procedures. We propose new DID estimators that reflect these trade-offs and derived their large sample properties. We illustrate the practical relevance of these results by assessing whether the transition from federal to state management of the Clean Water Act affects compliance rates.
1 Introduction
In staggered DID settings, different parallel trends assumptions identify different causal parameters and create a robustness–efficiency trade-off. The paper compares these assumptions, proposes corresponding estimators, and illustrates their practical importance using Clean Water Act management transitions.
- Motivation: Staggered DID is more challenging than canonical 2 × 2 DID because treatment occurs across multiple periods and groups, with procedures relying on different parallel trends assumptions and causal parameters.The paper focuses on staggered adoption designs with binary treatments to compare these assumptions and parameters directly.
- Parallel trends assumptions: C&S considers weaker assumptions based on never-treated or not-yet-treated comparison groups, yet all these approaches can recover the same variety of average treatment effects.The assumptions differ in strength because they impose different restrictions on pre-treatment trends and comparison groups.
- Implications: The paper argues that researchers should state the parallel trends assumption explicitly because it helps determine both the appropriate estimator and the interpretation of the resulting treatment parameter.This links research-design assumptions to estimator choice rather than treating TWFE specifications as self-interpreting.
- Robustness and efficiency: Researchers may favor weaker parallel trends assumptions when stronger versions impose many restrictions relative to the available observations, although stronger assumptions can support more efficient estimation and specification tests.The paper uses GMM to form more efficient estimators under stronger assumptions and connects overidentification to Hansen–Sargan J-tests.
- Application: Clean Water Act transitions from federal to state management have little to no effect on violation rates, robust across parallel trends assumptions and causal parameters.The broader corruption comparison changes with the assumed PTA: corruption-specific trends yield essentially no evidence of heterogeneous effects, while pooled comparison groups yield a larger decrease for more corrupt states.
2 Difference-in-differences with multiple time periods
In staggered DID designs, treatment timing creates multiple groups and periods, making parallel-trends assumptions and causal parameters more complex than in canonical 2 × 2 DID. The section contrasts stronger and weaker assumptions, identifies TWFE interpretation problems, and discusses alternative estimators and inference.
- Framework: Staggered adoption involves multiple treatment groups and time periods, making DID substantially more challenging than the canonical two-group, two-period design.The paper focuses on panel-data settings with binary treatment, staggered adoption, no anticipation, and overlap across treatment cohorts.
- The different parallel trends assumptions: PTA 2.6 is weaker than PTA 2.5 and PTA 2.7 because it imposes no parallel pre-trends, though it requires a never-treated group.PTA 2.7 is weaker than PTA 2.5 because it does not restrict pre-trends before the first treatment or for the earliest treatment group.
- Pitfalls of TWFE regression specifications: TWFE coefficients generally represent weighted averages of heterogeneous treatment effects, and some weights can be negative, undermining their causal interpretation as policy-effect summaries.The paper therefore cautions against interpreting standard TWFE event-study coefficients as sensible causal parameters in general.
- Treatment effect parameters and estimators: The paper develops interpretable treatment-effect estimators whose estimates are unbiased, consistent, and asymptotically normal under the relevant assumptions.The discussed estimators target group-time or aggregated treatment effects rather than relying on the problematic causal interpretation of TWFE coefficients.
- Inference: Simultaneous confidence bands are important when estimating multiple treatment-effect parameters because ignoring multiple testing can produce misleading inference.The paper also notes that event-study estimators avoid pitfalls associated with using dynamic TWFE to assess parallel-trends credibility.
3 Not all parallel trends assumptions are made equal
Different PTAs imply different identification and testing properties in staggered DID. Their choice creates a robustness–efficiency trade-off that affects suitable estimators and interpretation of pre-treatment tests.
- Pre-treatment event-study tests provide direct or placebo-type evidence depending on the invoked PTA.Under PTA 2.5 and PTA 2.7, overidentified restrictions can support direct tests; under PTA 2.6, tests are only placebo-type evidence.
- Failing to reject pre-treatment tests does not establish identifying-assumption validity because tests may lack power or offsetting violations may occur.TWFE pre-treatment coefficients can also be contaminated by post-treatment effects.
- PTA 2.6 is weakest because it imposes no pre-treatment trend restrictions, but it requires a never-treated group.Its just-identified system cannot directly test the assumption.
- Researchers should explicitly state the PTA because it determines transparency, pre-test interpretation, and estimator efficiency.When a never-treated group is sufficiently large or pre-trend restrictions are undesirable, PTA 2.6 and its estimator may favor robustness over efficiency.
- PTA 2.5 imposes parallel pre-treatment trends across all groups and periods, whereas PTA 2.7 restricts them only from the first treatment period onward.The distinction matters when early pre-treatment periods may reflect a different economic environment.
4 Using GMM to estimate DID parameters
The paper constructs efficient DID estimators by combining PTA-implied and observational moment restrictions in a GMM framework. Under PTA 2.7, the resulting estimator is semiparametrically efficient and can support specification testing, though implementation may become burdensome.
- The efficient GMM estimator combines all linearly independent PTA and observational moment restrictions to estimate the unknown parameters.A preliminary consistent estimator and sample-average augmented moments are used in the construction.
- Under PTA 2.7, the GMM estimator is semiparametrically efficient.Its asymptotic result assumes finite second moments, positive-definite covariance, and Assumptions 2.1–2.4.
- Under PTA 2.7, GMM is generally more efficient than the simpler not-yet-treated, never-treated, and interaction-weighted estimators.Using the semiparametric efficiency bound generally yields tighter confidence intervals.
- The overidentified GMM system permits a Sargan–Hansen J-test for violations of PTA 2.7.Under the null, deviations of J from zero should remain within sampling error; violations should produce a large J.
- Implementation can be difficult when treatment groups or periods are numerous: the application has 780 moments, 195 overidentification restrictions, and 759 state-year observations.In such cases, the paper expects researchers to favor the simpler inefficient estimator.
5 A simpler and more robust DID estimator
The section develops a simpler DID estimator under a weaker parallel trends assumption that uses not-yet-treated units, avoids explicit pre-trend restrictions, and remains easy to compute. It establishes identification and asymptotic properties, then compares when this estimator may be preferred to alternatives.
- A weaker parallel trends assumption: The weaker parallel trends assumption uses the outcome evolution among units not yet treated by time t to identify ATT(g, t).It does not require every individual not-yet-treated group to serve as a comparison group, so the effects are nonparametrically just-identified.
- Identification: Under this assumption, ATT(g, t) equals ATTny+ (g, t) for post-treatment periods t ≥ g.The result applies for 2 ≤ g ≤ t ≤ T.
- Estimator construction: ATTny+ (3, 4) uses data from all groups, whereas ATTny (3, 4) uses only the early-treated and never-treated groups.The broader data use can make ATTny+ estimators more precise, while the weaker assumption does not impose equal pre-treatment trends across comparison groups.
- Estimation and inference: The sample analogue is easy to compute from combinations of sample means, and the resulting estimators are √n-consistent with a joint asymptotic distribution.The section also provides influence functions and inference procedures, including standard errors, pointwise intervals, and bootstrapped simultaneous confidence intervals.
- Scope and caveat: The estimator is intended for post-treatment periods; pre-treatment analysis requires a separate local-deviation estimand and a stronger parallel trends assumption.The paper notes that this pre-treatment estimand can provide indirect evidence because the main assumption cannot be directly tested.
- Event-study extension: The estimator can also support event-study estimates that are consistent and asymptotically normal.These estimates are constructed by weighting the group-time estimators across treatment cohorts and event times.
- Practical comparison: Researchers may favor this estimator when they want to avoid pre-trend restrictions, use all groups, avoid a never-treated group, or prioritize easy implementation over GMM efficiency.The paper also identifies settings with few clusters or a relatively small never-treated group as potentially relevant.
6 The effect of enforcing the Clean Water Act
The Clean Water Act application examines staggered state authorization using event-study procedures tied to alternative parallel trends assumptions. Overall authorization has little to no effect on violation rates, while corruption-specific findings change substantially depending on whether differential counterfactual trends are allowed.
- 6.1 Data: The analysis follows Grooms (2015), constructing state-year measures of facilities with inspections, violations, or enforcement actions and excluding 27 always-treated states.The remaining sample contains 23 states with authorization timing observed during the sample period.
- 6.2 Specifications and assumptions: The authors estimate event-study treatment-effect dynamics using four procedures, including dynamic TWFE and estimators based on explicitly stated parallel trends assumptions.The event-study results report 20 treatment leads and 20 treatment lags with point-wise and simultaneous 90% confidence intervals.
- 6.3.1 Baseline results: Across parallel trends assumptions and estimators, state authorization has little to no effect on violation rates.Summary measures likewise indicate a close-to-zero effect, consistent with Grooms (2015).
- 6.3.2 Corrupt vs. Non-corrupt Results: The contrasting corruption results demonstrate why empirical DID analyses should make their underlying parallel trends assumptions explicit.The application shows that conclusions about heterogeneous effects depend on the selected assumption.
- 6.3.2 Corrupt vs. Non-corrupt Results: Allowing corruption-specific counterfactual trends yields no evidence that authorization effects differ between corrupt and non-corrupt states.This result contrasts with the TWFE specification in the corruption comparison.
- 6.3.2 Corrupt vs. Non-corrupt Results: Imposing the stronger assumption of common counterfactual trends produces a large post-authorization violation-rate decrease in corrupt states relative to non-corrupt states.The relative drop appears to increase with elapsed treatment time.
7 Conclusion
The paper shows that parallel trends assumptions shape identification, estimation, and treatment-effect summaries in event-study settings, creating a robustness–efficiency trade-off that should be made explicit.
- Different parallel trends assumptions can identify and estimate different treatment-effect parameters when treatment timing varies.
- The strength of the underlying parallel trends assumption creates a trade-off between robustness and efficiency in DID analysis.
- Estimators can use all restrictions implied by the underlying parallel trends assumption to achieve semiparametric efficiency.
- Practitioners should state the parallel trends assumption they invoke because doing so promotes more transparent and objective analysis.