Source-linked AI summary
When Should You Adjust Standard Errors for Clustering?
Alberto Abadie, Susan Athey, Guido Imbens, Jeffrey Wooldridge
TL;DR
The paper asks when and how standard errors should be adjusted for clustering, beyond conventional assumptions about correlated outcomes. It develops a sampling-and-design framework and new variance estimators, concluding that robust errors can be too small while conventional cluster errors can be unnecessarily large.
Problem
Conventional clustering frameworks do not fully incorporate treatment-assignment design, leaving unclear when and how clustering should be used for treatment-effect inference.
Method
The paper develops a framework that shifts attention from outcome-process features to the population treatment effect and derives new variance estimators and procedures.
Results
The framework shows that robust standard errors can be too small, whereas conventional cluster standard errors can be unnecessarily large.
Takeaways & Limitations
The decision on when and how to cluster standard errors depends on the framework's sampling and assignment processes rather than outcome residual correlation alone.
Takeaways & Limitations
CCV and TSCB are designed for settings where within-cluster quantities can be precisely estimated, while broader framework principles may apply elsewhere.
Abstract
from arXiv · showhide
In empirical work it is common to estimate parameters of models and report associated standard errors that account for "clustering" of units, where clusters are defined by factors such as geography. Clustering adjustments are typically motivated by the concern that unobserved components of outcomes for units within clusters are correlated. However, this motivation does not provide guidance about questions such as: (i) Why should we adjust standard errors for clustering in some situations but not others? How can we justify the common practice of clustering in observational studies but not randomized experiments, or clustering by state but not by gender? (ii) Why is conventional clustering a potentially conservative "all-or-nothing" adjustment, and are there alternative methods that respond to data and are less conservative? (iii) In what settings does the choice of whether and how to cluster make a difference? We address these questions using a framework of sampling and design inference. We argue that clustering can be needed to address sampling issues if sampling follows a two stage process where in the first stage, a subset of clusters are sampled from a population of clusters, and in the second stage, units are sampled from the sampled clusters. Then, clustered standard errors account for the existence of clusters in the population that we do not see in the sample. Clustering can be needed to account for design issues if treatment assignment is correlated with membership in a cluster. We propose new variance estimators to deal with intermediate settings where conventional cluster standard errors are unnecessarily conservative and robust standard errors are too small.
1. Introduction
The paper reframes clustering through sampling and treatment-assignment design, showing when conventional adjustments are needed and proposing intermediate variance estimators. Its framework separates sampling from assignment roles and clarifies why robust and conventional cluster standard errors can respectively be too small or conservative.
- Implications: Residual correlation within clusters neither proves that clustering is needed nor rules it out when absent.The relevant decision depends on the sampling and assignment processes rather than residual correlation alone.
- Implications: Clustered standard errors can be unnecessarily conservative under random sampling, while robust standard errors can be too small.The paper also shows that treatment-effect heterogeneity introduces additional variance components affecting clustering adjustments.
- Framework: The framework allows clustering in both sampling and treatment assignment, nesting clustered sampling and clustered treatment assignment while also covering intermediate cases.Treatment may depend on cluster membership without being perfectly determined by clusters, leaving within-cluster treatment variation.
- Framework: The framework separates sampling and assignment roles: the data do not reveal sampling-related clustering needs but do reveal assignment-related needs.This design perspective avoids relying solely on an assumed error-component structure.
- New estimators: The paper proposes CCV analytic variance formulas and a bootstrap procedure, TSCB, to improve on robust and conventional cluster standard errors.These procedures target settings where conventional cluster errors are too large and robust errors are too small.
- Empirical illustration: In the Census example, state-clustered standard errors are approximately twenty-six times larger than robust errors, while CCV and TSCB lie between them.For the OLS estimate, CCV and TSCB are 0.0035 and 0.0036, compared with robust 0.0012 and cluster 0.0269.
2. A Framework for Clustering
The framework treats uncertainty as arising from sampling and treatment assignment, separating variation across units, clusters, and assignments. It models two-stage clustered sampling and assignment mechanisms that may depend on cluster membership.
- The framework identifies three sources of sampling variation: which units, which clusters, and which treatments are observed.
- It distinguishes sampling uncertainty from treatment-assignment uncertainty, allowing all three components to affect estimator variance.
- 2.2. The Sampling Process: Clustered sampling first selects clusters with probability q_k, then samples units from selected clusters with probability p_k.
- 2.3. The Assignment Process: The assignment process can be random, partially clustered, or clustered depending on whether treatment probabilities vary across clusters and within clusters.
3. The Least Squares Estimator and its Variance
The least squares estimator’s variance depends on sampling, assignment, cluster sizes, and treatment-effect heterogeneity. Robust variance can underestimate uncertainty, whereas cluster variance is generally conservative and can be extremely conservative.
- When clustered sampling is present, the variance increases through a term that vanishes when average treatment effects are homogeneous across clusters.
- The need to adjust for clustered sampling is not identifiable from the sample, whereas the need to account for clustered assignment may be sample-informative.
- The asymptotic variance depends on the sampling fraction, clustered sampling, clustered assignment, and heterogeneity in potential outcomes.
- Cluster variance cannot underestimate the true variance asymptotically, but it can be extremely conservative outside special cases.
- Robust variance can severely underestimate the true variance when clusters explain substantial heterogeneity in potential outcomes.
4. Two New Variance Estimators
The paper proposes analytic and resampling-based variance estimators that correct cluster-variance bias in intermediate settings. Their performance improves when cluster-specific treatment-effect heterogeneity can be estimated, but they require sufficient treated and control observations per cluster.
- The authors propose two variance estimators: one analytic correction and one based on resampling methods.
- The estimators rely on estimating treatment-effect variation across clusters and therefore require substantial treated and control observations per cluster.
- The proposed estimators substantially improve on cluster variance when cluster variance has a large upward bias.
- The proposed estimators can be conservative when treatment effects are homogeneous or when within-cluster treated and control samples are too small.
- The proposed estimators perform very well in simulations, while formal derivations of their properties are left for future work.
- The two-stage-cluster-bootstrap extends resampling to settings where a large fraction of clusters is observed.
5. The Fixed Effect Estimator
The fixed-effect analysis shows that robust standard errors can underestimate variance while conventional cluster standard errors can be too large. The authors propose CCV and TSCB estimators that combine information from both conventional approaches.
- CCV and TSCB are proposed as alternative variance estimators for settings where robust errors are too small and cluster errors unnecessarily large.
- The fixed-effect estimator is obtained by regressing treatment and outcomes after removing cluster means.
- The fixed-effect estimator has a large-sample normal distribution under additional regularity conditions.
- Robust standard errors can underestimate the true variance, while cluster variance estimates are generally too large.
- The proposed estimator is a convex combination of robust and cluster variance estimates, with weights chosen to correct cluster-estimator bias.
- The fixed-effect variance estimator is undefined when treatment has no within-cluster variation.
6. Simulations
Simulations compare robust, cluster, CCV, and TSCB standard errors with simulated estimator variability across sampling and assignment designs. The proposed methods generally outperform conventional alternatives, with TSCB especially accurate in the baseline design.
- In the baseline design, the normalized standard deviation of the least-squares estimator is 5.91, closely matching its asymptotic standard error of 5.90.
- The robust standard error averages 1.90, while the cluster standard error averages 44.86 against the normalized estimator variability.
- CCV averages 6.32, about 7 percent above the normalized standard deviation, while TSCB averages 5.80.
- For the fixed-effect estimator, robust standard errors are about 16 percent too small, whereas cluster errors are too large by a factor of 20.
- Across designs, CCV and especially TSCB outperform robust and cluster standard errors.
- Reducing the sampled-unit fraction worsens analytic CCV performance, whereas bootstrap-based TSCB continues to perform well.
7. Implications for Practice
The appropriate clustering adjustment depends on the sampling and treatment-assignment processes rather than merely on within-cluster outcome correlation. Conventional clustering can be unnecessarily conservative, while partially clustered assignment permits smaller alternatives.
- Clustered variance is asymptotically correct when there is no cluster sampling, and also when large clusters are substantially sampled but few units per cluster are observed.
- In randomized unit-level assignment, clustering can produce unnecessarily wide confidence intervals even when outcomes are correlated within clusters.
- For partially clustered assignment with large clusters, CCV and TSCB can produce standard errors considerably smaller than conventional clustered errors.
- When cluster sizes are large and treatment varies within clusters, CCV and TSCB can substantially reduce standard errors.
- When treatment is assigned by cluster, standard errors should be clustered at the assignment level, such as villages when all farmers share treatment status.
8. Conclusion
The paper develops a sampling-and-assignment framework for deciding when and how to cluster standard errors. It shows that conventional methods can be too small or too large and proposes alternatives for intermediate settings.
- The decision to cluster depends on sampling and assignment processes, not on within-cluster error components in outcomes.
- Robust standard errors can be too small, while conventional cluster standard errors can be unnecessarily large.
- CCV and TSCB provide more precise standard errors when clusters are large and treatment assignment varies sufficiently within clusters.
- If sampling is not clustered, standard errors should be clustered at the treatment-assignment level because potential-outcome sampling is determined by assignment.
- When the sampled-cluster fraction is non-negligible and cluster-average treatment effects vary, conventional clustered standard errors may be inaccurate.
- The analysis is restricted to linear estimators, while deriving formulas for other sampling and assignment processes remains future work.
On-Line Appendix
The paper is titled “When Should You Adjust Standard Errors for Clustering?” and is authored by Alberto Abadie, Susan Athey, Guido W. Imbens, and Jeffrey M. Wooldridge.
- The paper examines when standard errors should be adjusted for clustering.
- The authors are Alberto Abadie, Susan Athey, Guido W. Imbens, and Jeffrey M. Wooldridge.
A.1. Setting and notation
The setting defines clustered populations, potential outcomes, sampling, and treatment assignment processes. Sampling and assignment can each operate in two stages and can generate within-cluster dependence.
- Each population contains units partitioned into strata or clusters, with treatment and no-treatment potential outcomes for every unit.
- The population treatment effect is defined using unit-level potential outcomes and cluster membership.
- Potential outcomes are assumed to be uniformly bounded in absolute value across populations and units.
- Sampling is independent of potential outcomes and assignments and proceeds by sampling clusters first, then units within sampled clusters.
- Treatment assignment first draws cluster-specific probabilities, then assigns treatment independently to units conditional on those probabilities.
- The assignment-probability variance determines whether treatment is random across clusters or correlated within clusters.
A.2. Base case: Difference in means
The base case studies the simple difference in means between treated and untreated sampled units under assumptions controlling sampled-cluster counts, within-cluster sample sizes, and cluster-size imbalance.
- The estimator is the coefficient on treatment in a regression of outcome on a constant and the treatment indicator.
- The analysis compares treated and untreated units using the simple difference of means.
- The assumptions require the expected number of sampled clusters to diverge, sampled units per cluster not to vanish, and cluster sizes not to become excessively imbalanced.
A.2.1. Large k distribution
The large-k analysis derives asymptotic behavior for the difference-in-means estimator under clustered sampling and assignment, and separately treats the case where no clustering is required.
- A.2.1. Large k distribution: The derivation decomposes outcomes into cluster-level effects and residual components to analyze the estimator’s variance.
- A.2.1. Large k distribution: The sample treatment proportions consistently estimate their population counterparts, with analogous results for treated and untreated groups.
- A.2.1. Large k distribution: Positive within-cluster assignment variance or incomplete cluster sampling creates within-cluster dependence in the relevant terms.
- A.2.1. Large k distribution: The variance remains bounded away from zero when the stated cross-cluster variation conditions hold together with the within-cluster sampling-size assumption.
- A.2.1. Large k distribution: A Lyapunov-condition argument establishes the central limit theorem under bounded potential outcomes and the sampling and cluster-size assumptions.
- A.2.1. Large k distribution: When all clusters are sampled and assignment variance is zero, no clustering is required; the relevant asymptotic condition is nkpk → ∞.
A.2.2. Estimation of the variance
This section establishes convergence properties for variance estimators based on regression residuals. The proof decomposes estimation errors into bounded-in-probability and vanishing terms, then applies array laws of large numbers.
- Residuals are defined from regressions of Yk,i on a constant and Wk,i, using estimated coefficients pαk and pτk.
- The cluster variance estimator is shown to have bounded asymptotic order under the stated normalization.
- The proof shows that residual-estimation errors converge to zero in probability by factorizing each term into bounded and L1-vanishing components.
- Analogous calculations establish that the robust variance estimator's discrepancy terms vanish in probability.
A.3. Fixed effects
This section analyzes fixed-effects variance components by separating within-cluster errors from between-cluster treatment-effect variation. It derives conditions under which the resulting estimator is well-defined and identifies when a bias term disappears or becomes negligible.
- Estimator conditions: The estimator requires the limiting expected treatment-assignment variance within clusters to remain bounded away from zero for large-sample definition.
- Bias term: If cluster sizes do not vary across clusters, the bias term Bk equals zero; otherwise, cluster-size variation can generate bias.
- Estimator conditions: The asymptotic analysis assumes bounded potential outcomes, growing expected sampled clusters, diverging observations per sampled cluster, and a bounded maximum-to-minimum cluster-size ratio.
- Fixed-effects variance decomposition: The variance expression separates contributions from intra-cluster heterogeneity in potential outcomes and treatment effects from inter-cluster variation in average treatment effects.
- Fixed-effects variance decomposition: The inter-cluster treatment-effect term can dominate the variance in large samples when the stated growth conditions hold.