Source-linked AI summary

A Practical Guide to Counterfactual Estimators for Causal Inference with Time-Series Cross-Sectional Data

Licheng Liu, Ye Wang, Yiqing Xu

arXiv:2107.00856v3stat.MEstat.AP

TL;DR

Conventional fixed-effects models face important drawbacks, including unrealistic assumptions about time-varying confounders. The paper introduces counterfactual estimation that fits controls and imputes counterfactuals for treated units, improving current TWFE practices and dynamic treatment-effect analysis.

  • Problem

    Fixed-effects models face important drawbacks, including assumptions about time-varying confounders that may be unrealistic.

  • Method

    The framework fits data in controls and imputes counterfactual outcomes for treated observations, with diagnostic tools for assessing identifying assumptions.

  • Results

    The framework provides a simple but powerful way to improve current practices with TWFE models and offers advantages over existing approaches.

  • Takeaways & Limitations

    The paper improves existing practice for estimating and plotting dynamic treatment effects.

  • Takeaways & Limitations

    The identifying assumptions can be unrealistic because they require the absence of time-varying confounders.

Abstract

from arXiv · show

This paper introduces a simple framework of counterfactual estimation for causal inference with time-series cross-sectional data, in which we estimate the average treatment effect on the treated by directly imputing counterfactual outcomes for treated observations. We discuss several novel estimators under this framework, including the fixed effects counterfactual estimator, interactive fixed effects counterfactual estimator, and matrix completion estimator. They provide more reliable causal estimates than conventional twoway fixed effects models when treatment effects are heterogeneous or unobserved time-varying confounders exist. Moreover, we propose a new dynamic treatment effects plot, along with several diagnostic tests, to help researchers gauge the validity of the identifying assumptions. We illustrate these methods with two political economy examples and develop an open-source package, fect, in both R and Stata to facilitate implementation.

1. Introduction

The paper introduces a counterfactual-imputation framework for TSCS data that addresses TWFE weighting problems and can accommodate treatment reversal and decomposable time-varying confounders. It also develops multiple estimators plus visualization and diagnostic tools for assessing identifying assumptions and model choice.

  • Motivation: TWFE models can be biased when strict exogeneity, constant treatment effects, or no carryover effects fail, while existing remedies may be limited or inefficient.Strict exogeneity excludes time-varying confounders and feedback; heterogeneous effects can produce negative weights, and some alternatives require staggered adoption or discard observations.
  • Counterfactual framework: The framework treats treated observations as missing, models outcomes using control-condition data, and imputes treated observations’ counterfactual outcomes.It focuses on dichotomous treatments in general panel structures where treatment may switch back and forth.
  • Framework benefits: By excluding treated observations from model fitting and imposing uniform weights on treated effects, counterfactual estimators avoid negative weights and address bias from treatment-effect heterogeneity.The framework also accommodates models that can potentially relax conventional strict exogeneity assumptions and facilitates diagnostics and visualization.
  • Estimators: The paper presents FEct, IFEct, and MC estimators, using untreated outcomes to construct lower-rank approximations that account for potential time-varying confounders.FEct is a special case of IFEct and provides a simple solution to the TWFE weighting problem.
  • Diagnostics and scope: The paper’s second main contribution is a set of visualization and diagnostic tools for assessing identifying assumptions and selecting suitable models.The approach is designed to accommodate complex TSCS structures, including treatment reversal, and can serve as a building block for doubly robust estimators.

2. Counterfactual Estimators

The section develops a counterfactual-estimation framework for TSCS data that targets treatment effects by imputing treated units’ untreated potential outcomes. It defines the identifying assumptions and introduces fixed-effects, interactive-fixed-effects, and matrix-completion estimators, with diagnostic and implementation support.

  • Setup and assumptions: The untreated potential outcome combines observed covariates, unobserved attributes, and idiosyncratic error through an additive functional-form model.The setup also imposes strict exogeneity and a low-dimensional decomposition of unobserved attributes.
  • Estimands: The framework targets the average treatment effect on units whose treatment status changes during the observed period.Effects for never-treated and always-treated units are difficult to identify or estimate and are therefore excluded from preprocessing.
  • Estimation strategy: The estimation strategy separates observations under control and treatment conditions to impute counterfactual outcomes for treated observations.The framework nests TWFE and IFE models while allowing treatment effects to vary across units and periods.
  • Estimator comparison: IFEct outperforms matrix completion in some settings, whereas matrix completion performs better in others, so estimator choice depends on the setting.Simulation results indicate that both inferential methods work well with reasonable sample sizes, such as T = 20 and N = 50.

3. Diagnostics

The section introduces diagnostic tools for assessing identifying assumptions, including a counterfactual dynamic-treatment-effects plot, placebo tests, and extensions testing pretrends and carryover effects. The plot summarizes residualized treatment effects without a researcher-chosen base category, while the tests help identify assumption failures.

  • Diagnostics: The diagnostic toolkit combines a counterfactual dynamic-treatment-effects plot with placebo, no-pretrend, and no-carryover tests for evaluating identifying assumptions.The latter two tests extend the placebo-test framework.
  • Dynamic treatment-effects plot: The proposed plot averages Yit − ˆYit(0) for treated units by time relative to treatment onset, with pretreatment residual averages expected to converge to zero under valid assumptions.Because untreated averages are already subtracted, the plot does not require researchers to choose a base category.
  • Dynamic treatment-effects plot: The plot relaxes the constant-treatment-effect assumption and provides an intuitive visual check for data or modeling problems, but it cannot identify the specific source of assumption failure.Potential failures include anticipation, time-varying confounding, and feedback from past outcomes.
  • Simulation evidence: In the simulated example, FEct shows strong pretrends and sizable positive posttreatment bias, MC shows smaller posttreatment biases, and IFEct estimates are close to the truth.These patterns illustrate that failing to adjust for relevant factors can bias causal estimates.
  • Statistical tests: The placebo test can detect failures such as feedback effects and is robust to model misspecification, while the extended tests assess pretrends and carryover effects.In the reported example, IFEct has a placebo effect statistically indistinguishable from zero (p = 0.534), whereas FEct and MC fail the placebo assessment.
  • Statistical tests: The carryover-effects results indicate no carryover effects regardless of estimator or test, consistent with the data-generating process.The no-pretrend extension uses an equivalence test to address limitations of testing joint zero means with an F test.

4. Empirical Examples

The empirical examples apply counterfactual estimators and diagnostics to staggered-adoption and treatment-switching designs. FEct reproduces conventional estimates in the first example, while IFEct and MC address diagnostic concerns in the second.

  • Direct democracy and naturalization rates: The Swiss municipality dataset contains 1,211 municipalities observed for 19 years from 1991 to 2009.The outcome is minority immigrants’ naturalization rate, and treatment indicates whether decisions are made by popular referendums.
  • Direct democracy and naturalization rates: 1.339%: naturalization rates increased on average after municipalities shifted decision-making from popular referendums to elected officials under twoway fixed effects.The reported standard error was 0.161.
  • Direct democracy and naturalization rates: IFEct selected zero factors and MC selected maximum regularization, so both reduced to FEct and produced identical estimates.FEct results were substantively the same as conventional TWFE estimates, while counterfactual estimators made identifying-assumption checks more transparent.
  • Partisan alignment and grant allocation: In the England grant-allocation example, FEct pretreatment residual averages consistently deviated from zero, suggesting potential violations of identifying assumptions.IFEct and MC produced pretreatment residual averages very close to zero.
  • Partisan alignment and grant allocation: IFEct and MC yielded placebo effects close to zero, while IFEct approximated the data better than MC with an almost completely flat pretrend.With FEct, the null hypothesis of a non-zero placebo effect could not be rejected at the 5% level.
  • Partisan alignment and grant allocation: Positive carryover effects persisted at least three years after partisan alignment ended, so removing three post-treatment periods made IFEct the most suitable model.After re-estimation, IFEct passed all diagnostic tests.

5. Conclusion

The paper frames counterfactual estimation as a practical way to improve TWFE-based causal analysis of TSCS data by imputing treated units’ counterfactual outcomes. It unifies estimators, diagnostics, and implementation guidance to help researchers assess identifying assumptions and analyze treatment effects.

  • Counterfactual estimation: The framework fits data on controls and imputes counterfactual outcomes for treated observations, with diagnostic tests probing identifying assumptions.Its central principle is to “fit data in the controls and impute counterfactuals to the treated.”
  • Estimators: FEct, IFEct, and MC address heterogeneous-treatment-effect negative weights, general panel treatment structures, time-varying covariates, and decomposable time-varying confounders.IFEct and MC are existing methods placed in the same framework to enable diagnostics and assumption evaluation.
  • Diagnostics: The paper improves dynamic-treatment-effect plots and develops statistical tests based on out-of-sample predictions of untreated potential outcomes.Researchers are encouraged to use visual and statistical tests together to gauge identifying assumptions.
  • Practical workflow: The recommended workflow starts with FEct and diagnostics, escalates to IFEct or MC when placebo or pretrend tests fail, and removes post-treatment periods when carryover tests fail.Optional subgroup analysis can identify which units drive detected effects.
  • Implementation: The authors provide panelView and fect packages in R and Stata to support analysis of TSCS data.The paper presents these tools as assistance for improving research practice.

Replication Materials

The data and materials needed to verify the article’s computational reproducibility are publicly available through the Harvard Dataverse Network.

  • Replication Materials: Data and materials for verifying the article’s computational reproducibility are available through the American Journal of Political Science Dataverse within the Harvard Dataverse Network.The materials cover the article’s procedures and analyses.

Supplementary Information for … A.2. The MC Algorithm

The supplementary material describes iterative procedures for the IFEct and matrix completion algorithms. The MC method uses singular-value shrinkage, while hard imputation produces estimates almost identical to IFEct.

  • A.1. The IFEct Algorithm: The IFEct algorithm updates coefficients using untreated data, imputes treated outcomes through conditional expectations, and iterates over completed data.The procedure initializes parameters with a twoway fixed effects model and updates estimates across successive steps.
  • A.1. The IFEct Algorithm: The IFEct procedure updates factor estimates and factor loadings by minimizing a least-squares objective over completed data.It also updates the grand mean and twoway fixed effects before estimating treated counterfactuals.
  • A.1. The IFEct Algorithm: The algorithm estimates treated counterfactuals and then obtains the ATT and ATTs as in FEct.These are the final two steps of the IFEct algorithm.
  • A.2. The MC Algorithm: The MC method defines observed and unobserved matrix entries and begins with L0(θ) = PO(Y).The method uses PO(A) and O(A) to distinguish matrix entries based on whether (i, t) belongs to O.
  • A.2. The MC Algorithm: MC applies singular-value decomposition and soft imputation, replacing each σi(A) with max(σi(A) −θ, 0).The shrinkage operator reconstructs the matrix using the modified singular values.
  • A.2. The MC Algorithm: The MC algorithm iteratively calculates Lh+1(θ), while hard imputation yields estimates almost identical to the IFEct algorithm.Hard imputation retains σi(A) when σi(A) ≥θ and replaces it with zero otherwise.

A.3. The Difference-in-Means Tests and Equivalence Tests … B. Proofs

The appendix defines equivalence tests based on reversed null hypotheses and discusses how carryover effects affect assumptions and diagnostics. It also outlines placebo and carryover testing procedures for assessing pretrends and temporal interference.

  • A.3. The Difference-in-Means Tests and Equivalence Tests: The primary equivalence test uses two one-sided tests (TOST) for each pretreatment period and declares equivalence only when every period rejects the null hypothesis.The TOST is preferred because its threshold is more intuitive and easier to interpret than the alternative equivalence F test.
  • A.3. The Difference-in-Means Tests and Equivalence Tests: The equivalence F test reverses the null hypothesis and rejects it when the statistic is below the 100αth percentile of a non-central F-distribution.The distribution is F(m + 1, Ntr −m−1, Ntrκ2), with Ntrκ2 as the centrality parameter.
  • A.3. The Difference-in-Means Tests and Equivalence Tests: With an equivalence threshold of 0.6, the equivalence test is more lenient than the F test and can declare equivalence where the F test suggests inequivalence.The equivalence test shares the TOST’s advantages, but its threshold is less intuitive.
  • A.4. Discussion on the No Carryover Effects Assumption: Carryover effects violate SUTVA temporally because a unit’s outcome can depend on its own treatment status in earlier periods.This is temporal interference and does not imply failure of the strict assumption.
  • A.4. Discussion on the No Carryover Effects Assumption: Carryover effects are mainly a concern when treatment switches on and off; placebo-test prediction errors deviating from zero provide evidence that the assumption may be invalid.Such deviations may also result from temporal shocks unrelated to carryover effects.
  • A.4. Discussion on the No Carryover Effects Assumption: Under staggered adoption, δit can be interpreted as the current treatment effect plus cumulative carryover effects relative to the never-treated potential-outcome history.A possible remedy is to recode several periods after treatment ends as treated so carryover effects can fully appear.
  • A.5. Procedures for the Diagnostic Tests: The placebo test removes l observations immediately before treatment for treated units, estimates placebo-period ATT, and uses bootstrap or jackknife confidence intervals.The no-pretrend test is the special case l = 1, repeated across successive pre-treatment periods.
  • A.5. Procedures for the Diagnostic Tests: The no-carryover test removes the first l observations after treatment ends, estimates carryover-period ATT, and tests it using DIM or equivalence methods.Researchers can allow limited carryover by removing specified post-treatment periods, then re-estimating ATT and repeating the diagnostics.

B.1. Unbiasedness and Consistency of FEct and IFEct

Under the stated model specifications and regularity conditions, FEct and IFEct estimators are unbiased, with consistency established as the relevant panel dimensions increase. FEct unit fixed effects can remain inconsistent when only N grows, whereas the ATT estimators are consistent.

  • FEct: Under model specification (A1) and regularity conditions, FEct estimates of β, α_i, and ξ_t are unbiased, while β and ξ_t are consistent.The consistency result is stated for the coefficient and time fixed-effect estimates.
  • FEct: α_i is inconsistent when only N →∞ because the number of estimated parameters changes.This limitation applies to the FEct unit fixed-effect estimates, not to the stated ATT consistency result.
  • IFEct: Under model specification (A2) and regularity conditions, IFEct estimates of β, α_i, ξ_t, λ_i, and f_t are unbiased and consistent as N and T increase.The result covers the observed coefficients, additive effects, interactive-effect loadings, and factors.
  • FEct: The FEct ATT estimators are unbiased and consistent as N →∞ under model specification (A1) and the stated regularity conditions.The proof concludes that the estimation error for the ATT converges to zero as N increases.
  • IFEct: The IFEct ATT estimators are unbiased and consistent as N, T →∞ under model specification (A2) and the stated regularity conditions.This follows from the unbiasedness and consistency of the IFEct parameter estimates.

B.2. FEct as a Weighting Estimator

Under model specification (A1), FEct can be represented as a weighting estimator for untreated observations. Its weights place greater emphasis on untreated observations when unit or period comparisons are sparser, and in the example all four FEct weights are equal and positive.

  • Under model specification (A1), Proposition 3 characterizes FEct as a weighting estimator.
  • FEct weights untreated observation (j, s) more heavily when fewer untreated observations exist in unit j or period s.
  • In the three-unit, four-period example, FEct assigns four equal and positive weights to the treated-observation estimates.The displayed decomposition gives each component a coefficient of 1/4.

C. Inferential Methods · D. Additional Monte Carlo Evidence · D.1. Describing the DGP of the Simulated Example

The paper uses unit-clustered block bootstrap and jackknife procedures for inference, finding that both accurately estimate FEct variance in simulations. Additional Monte Carlo exercises compare inferential methods and estimators using a simulated DGP with heterogeneous, gradually increasing treatment effects and specified covariate, factor, and error structures.

  • C. Inferential Methods: Block bootstrap resamples an equal number of units with replacement, replicating each selected unit’s full time series of outcomes, treatment status, and covariates.Standard errors and confidence intervals use conventional standard-deviation and percentile methods, respectively.
  • C. Inferential Methods: Jackknife drops one unit’s entire time series per run and re-estimates treatment effects, providing an alternative when the treated-unit count is small but exceeds one.The variance is computed from the resulting treatment-effect estimates.
  • C. Inferential Methods: Both FEct bootstrap and jackknife QQ plots lie almost exactly on the 45-degree lines, while the inconsistent twoway fixed effects plot does not pass through (0, 0).This supports extending block bootstrap and jackknife variance estimation to counterfactual estimators including FEct, IFEct, and MC.
  • D. Additional Monte Carlo Evidence: The additional Monte Carlo exercises examine finite-sample inferential properties, differences between IFEct and MC, and the equivalence test versus the F test.The simulated DGP is described before these exercises.
  • D.1. Describing the DGP of the Simulated Example: The outcome DGP combines covariates, unit and time fixed effects, two latent factors, and i.i.d. disturbances, with f1t following a linear trend plus white noise.The covariates, factor loadings, unit effects, and factor noise are specified as normal or i.i.d. normal variables, while time effects follow a stochastic drift.

D.2. IFEct versus MC

In simulations with 200 units and 30 periods, IFEct performs better with few strong factors, while MC catches up to and eventually outperforms IFEct as factors become numerous and weak. MC’s robustness reflects that it does not directly estimate factors and loadings, whereas cross-validation has more difficulty detecting weaker factors.

  • Simulation design: The simulations use 200 units and 30 time periods, with all treated units receiving treatment in period 21 (T0 = 20).The number of factors varies from 1 to 9 while total factor contributions to outcome variance remain constant.
  • Estimator comparison: IFEct performs better than MC when only a few factors are present and each has a relatively strong signal.This comparison reflects the hard-impute IFEct estimator versus the soft-impute MC estimator.
  • Estimator comparison: MC catches up with and eventually beats correctly specified IFEct as the number of factors grows and each factor produces weaker signals.The results compare mean squared prediction errors for treated counterfactuals across 500 simulations.
  • Estimator comparison: When factors become weaker, cross-validation has more difficulty selecting them, resulting in worse predictive performance.The figure compares IFEct using the correct or cross-validated number of factors with MC using a cross-validated tuning parameter.
  • Estimator comparison: MC remains robust to many factors because it does not directly estimate factors and loadings.This supports MC’s advantage when factors are sparsely distributed or numerous and weak.

D.3. F Test versus the Equivalence test. · E. Additional Information on the Empirical Examples · E.1. Replicating Hainmueller and Hangartner (2015)

The section compares the F and equivalence tests under unobserved confounding and adds empirical-example materials, including treatment-assignment patterns from a replication. The simulations show that test performance depends on sample size and bias magnitude.

  • D.3. F Test versus the Equivalence test.: The F test has low power in small samples, whereas TOST does not.This comparison concerns testing the no-pre-trend assumption.
  • D.3. F Test versus the Equivalence test.: When sample size is relatively large, TOST are more liberal than the F test toward small biases.The paper frames this as a trade-off between the two tests.
  • D.3. F Test versus the Equivalence test.: 600 simulations compare the tests using N = 100 units, comprising 50 treated and 50 controls over 40 periods.The simulations estimate a FEct model that omits the time-varying confounder and then expand the sample to N = 300.
  • D.3. F Test versus the Equivalence test.: With N = 100, the F test often fails to reject zero residual average even when confounder-induced biases are large.The equivalence test’s probability of declaring equivalence falls quickly as bias increases, making it more powerful for detecting imbalances.
  • D.3. F Test versus the Equivalence test.: With N = 300, the F test’s non-rejection rate declines quickly as confounder influence grows, while the equivalence test signals increasingly influential confounding.The equivalence test declares equivalence when confounding is relatively inconsequential and raises an alarm as influence increases.
  • D.3. F Test versus the Equivalence test.: Figure A10 compares F and equivalence tests under unobserved confounding, with N = 100 in plot (a) and N = 300 in plot (b).Each plotted dot is based on 600 simulations.
  • E.1. Replicating Hainmueller and Hangartner (2015): Figure A11 plots treatment status for the first 50 units in the Hainmueller and Hangartner replication, showing staggered adoption across municipalities.Municipalities are ordered by when they began adopting indirect democracy for naturalization decisions, and the plot uses panelView.

E.2. Replicating Fouirnaies and Mutlu-Eren (2015)

This replication examines partisan alignment’s effects on UK local-council grants using treatment-status, pre-trend, carryover, placebo, and cohort-specific analyses. It reports cohort-specific estimates from FEct and IFEct based on councils’ alignment timing.

  • Treatment definition: Treatment status is defined by when English local councils become politically aligned with the government party.Councils are ordered by alignment timing, and the treatment-status plot is produced with panelView.
  • Diagnostic tests: The analysis evaluates carryover effects three years after treatment ends and identifies periods used for placebo and no-carryover tests.Blue dots indicate placebo-test periods, while blue and red dots identify periods removed during Step 1 and periods used to test no carryover effects, respectively.
  • Cohort-specific effects: Cohort-specific effects of partisan alignment on specific grants are estimated with FEct and IFEct.Cohorts are defined by the timing of a council’s first alignment with the government party and broadly correspond to three treatment-status blocks.
Loading 2107.00856v3…