Source-linked AI summary
The Econometrics of Randomized Experiments
Susan Athey, Guido Imbens
TL;DR
Randomized experiments avoid many challenges of observational studies, but causal inference observes at most one potential outcome for each unit. This review collects important statistical methods, emphasizing randomization-based approaches and covering complex designs, heterogeneity, non-compliance, and interactions; these methods support valid confidence intervals and exact p-values for some network tests.
Problem
Randomized experiments avoid many challenges of observational studies, yet causal inference observes at most one potential outcome for each unit.
Method
The review collects important statistical methods, focusing primarily on randomization-based approaches from classic methods through recent work on non-compliance.
Results
The reviewed methods allow researchers to construct valid confidence intervals and calculate exact p-values for tests through general networks.
Takeaways & Limitations
The review provides methods for analyzing randomized experiments across designs and complications including non-compliance, heterogeneous effects, covariates, and network interactions.
Takeaways & Limitations
Treatment-effect prediction methods can suffer in mean-squared error, while the range of conditions in which one method improves on sample splitting remains uncertain.
Abstract
from arXiv · showhide
In this review, we present econometric and statistical methods for analyzing randomized experiments. For basic experiments we stress randomization-based inference as opposed to sampling-based inference. In randomization-based inference, uncertainty in estimates arises naturally from the random assignment of the treatments, rather than from hypothesized sampling from a large population. We show how this perspective relates to regression analyses for randomized experiments. We discuss the analyses of stratified, paired, and clustered randomized experiments, and we stress the general efficiency gains from stratification. We also discuss complications in randomized experiments such as non-compliance. In the presence of non-compliance we contrast intention-to-treat analyses with instrumental variables analyses allowing for general treatment effect heterogeneity. We consider in detail estimation and inference for heterogeneous treatment effects in settings with (possibly many) covariates. These methods allow researchers to explore heterogeneity by identifying subpopulations with different treatment effects while maintaining the ability to construct valid confidence intervals. We also discuss optimal assignment to treatment based on covariates in such settings. Finally, we discuss estimation and inference in experiments in settings with interactions between units, both in general network settings and in settings where the population is partitioned into groups with all interactions contained within these groups.
1 Introduction
The chapter develops statistical methods for designing and analyzing randomized experiments, emphasizing inference justified by treatment assignment rather than hypothetical sampling. It addresses covariate imbalance, stratification, non-compliance, heterogeneous effects, clustering, and interference.
- Randomization-based inference: Randomization-based inference treats potential outcomes as fixed and treatment assignment as random, unlike sampling-based inference.The chapter presents this perspective as its central methodological theme and relates it to regression analysis.
- Covariate adjustment and design: Regression adjustments for covariates require additional assumptions, whereas stratifying on covariates and averaging within strata is directly justified by randomization.The chapter recommends using experimental design to address covariate differences rather than relying solely on post-randomization model-based adjustment.
- Scope and coverage: The chapter surveys formal methods for experiment design and analysis, focusing largely on single binary treatments and experimental rather than observational settings.Its scope includes power analysis, pairwise randomization, re-randomization, regression, and clustered designs.
- Treatment-effect heterogeneity: The chapter covers estimation and inference for average, quantile, conditional, and heterogeneous treatment effects, including valid confidence intervals for identified subpopulations.It focuses on methods that identify groups with different average treatment effects and estimate conditional average treatment effects.
- Covariate adjustment and design: Stratification yields variances no higher than those under completely randomized assignment, but pairing units can complicate variance estimation.The authors recommend small strata while retaining at least two treated and two control units where within-stratum variances can be estimated.
- Complications and dependence: It extends experimental analysis to non-compliance, clustered randomization, and interactions or spillovers in clusters and general networks.The chapter also discusses exact p-values for tests involving interactions between units.
2 Randomized Experiments and Validity
Randomized experiments control treatment assignment, supporting causal inference and internal validity, but their suitability and generalizability remain limited by the causal question, ethics, interference, and differences across settings.
- Randomized Experiments versus Observational Studies: Randomized experiments control assignment independently of observed and unobserved unit characteristics, unlike observational studies.This control can eliminate selection bias in comparisons between treated and control units.
- Randomized Experiments versus Observational Studies: Randomized experiments cannot answer every causal question, including effects of interventions on a single unit and many macroeconomic questions.Repeated interventions may instead permit experiments or quasi-experiments in some settings.
- Randomized Experiments versus Observational Studies: Ethical constraints can make experiments infeasible, such as when educational services cannot be withheld from individuals.Researchers may then use observational studies, possibly with randomized inducements to participate.
- Internal Validity: Well-executed randomized experiments have internal validity, but interference between units can complicate that conclusion.Internal validity concerns whether observed treatment–outcome covariance reflects a causal relationship within the study population.
- External Validity: External validity cannot be guaranteed because consent and treatment-effect heterogeneity may prevent results from generalizing across populations, settings, treatments, or outcomes.RCT results cannot automatically be extrapolated outside the context in which they were obtained.
- External Validity: Seven randomized experiments on microfinance programs found remarkable consistency across studies, illustrating one way to assess generalizability.Comparisons across multiple settings can reveal whether findings persist when unit characteristics, treatments, or treatment rates vary.
3 The Potential Outcome / Rubin Causal Model Framework for Causal Inference
The potential-outcomes framework defines causal effects through comparisons of outcomes under alternative treatments, while randomized assignment supplies a known basis for causal inference. Its analyses require multiple units and assumptions about interference and assignment mechanisms.
- Potential Outcomes: The fundamental problem of causal inference is that each unit reveals at most one of its potential outcomes.Credible and precise causal inference therefore requires additional assumptions or information.
- Multiple Units and Interference: With N units and binary treatments, the full treatment vector has 2^N possible values, so cross-unit comparisons require restrictions on interactions.The no-interference restriction makes unit i’s potential outcomes depend only on its own treatment, whereas interference can arise among classmates or labor-market participants.
- Assignment Mechanisms: Randomized experiments use a known assignment mechanism that does not depend on potential outcomes, distinguishing them from observational studies.The framework treats the assignment mechanism as central because it explains why units received their observed treatments.
- Potential Outcomes: Causal effects compare a unit’s potential outcomes under different treatments, although only one outcome is observed after assignment.The framework represents treatment effects using contrasts such as Y(1) − Y(0) or ratios such as Y(1)/Y(0).
- Experimental Designs: Completely randomized experiments assign a fixed number of units to treatment and the remainder to control.Stratified designs instead fix treated counts within covariate-defined strata, while related designs aim to exclude assignments likely to be uninformative.
4 The Analysis of Completely Randomized Experiments
For completely randomized experiments, inference conditions on the realized sample and treatment counts, using the randomization distribution rather than hypothetical repeated sampling. The section develops exact tests, average-effect estimation, regression connections, and quantile-effect analyses.
- Foundations: Randomization-based inference treats treatment assignment as the source of uncertainty while holding the sample and treatment-control counts fixed.The analysis may regard the sample as the population of interest or as a random sample from an infinite population, but its primary estimand is finite-sample based.
- Scope: The section covers exact p-values for sharp hypotheses, average treatment effects, regression analyses, and quantile treatment effects.The regression discussion explains how randomization justifies conventional regression analyses.
- Exact Tests: 1.79 thousand dollars is the observed difference in post-treatment earnings between treatment and control groups in the Lalonde application.Reassigning treatment while fixing group sizes produces an exact p-value of 0.0044, leading to rejection of the no-effect sharp null.
- Exact Tests: Rank-based randomization tests can improve power with outliers and thick-tailed distributions, but may perform poorly with many zeros and very thick tails among nonzero outcomes.In the Lalonde application, rank and mean statistics differ because many outcomes are zero, limiting the practical robustness gain.
- Average Treatment Effects: The conventional variance estimator is generally upwardly biased, producing conservative confidence intervals.The bias vanishes in two important cases discussed by the authors, including settings where treatment effects are constant.
- Average Treatment Effects: 0.0076 is the normal-approximation p-value, compared with an exact Fisher p-value of 0.0044 for the same application.The section contrasts randomization-based exact inference with approximation-based inference for estimated effects.
5 Randomization Inference and Regression Estimators
The section connects randomization-based inference with regression estimators, showing when regression preserves design-based properties and warning that model-based analyses can rely on difficult assumptions.
- Randomization and regression: Regression analyses can combine randomization assumptions, modeling assumptions, and large-sample approximations, especially with nonlinear methods.The authors therefore recommend care when using regression rather than randomization-based methods.
- Randomization and regression: Randomization fixes the treatment assignment mechanism, whereas conventional regression treats treatment-error independence as a population assumption that is difficult to assess.The residual’s interpretation is often unclear because it is used to capture unobserved outcome factors.
- Regression under randomization: Random assignment implies zero average residuals within both treatment arms, but not homoskedasticity or independence of errors from treatment.Valid confidence intervals therefore require Eicker-Huber-White robust standard errors.
- Regression under randomization: In the simple no-covariate case, least squares estimates the same effect as the difference in means and is unbiased for the average causal effect.The equivalence concerns estimation; inference can differ because robust variance estimators are not generally unbiased.
- Covariate adjustment: Covariate adjustment can reduce asymptotic variance by a factor related to 1 − R2 and support subgroup-effect estimation, but stratification can capture these gains by design.Indicator-based covariates preserve finite-sample properties and yield clearer interpretations than multivalued covariates in the regression function.
6 The Analysis of Stratified and Paired Randomized Experiments
Stratified and paired randomization use covariates in the design to improve precision, with paired designs as the limiting case of strata containing two units.
- Designs: Stratified randomization partitions the covariate space and conducts a completely randomized experiment within each subset.A paired experiment is the extreme case with exactly two units per subset and one treated unit per pair.
- Stratified analysis: The overall stratified effect averages within-stratum estimates using stratum-share weights, while equal treatment proportions make it equal to the overall difference in means.The completely randomized variance can be conservative because it ignores the precision gain from stratification.
- Paired analysis: Paired effects are estimated from within-pair treated-control differences and combined by averaging across pairs.The paired variance estimator reflects the design, whereas treating the experiment as completely randomized can substantially overstate uncertainty.
- Variance estimation: The usual within-stratum variance estimator requires at least two treated and two control units per stratum, so it is infeasible for one-treated/one-control pairs.A separate variance estimator based on variation across pair effects is used in the paired case.
- Illustration: In the Children’s Television Workshop experiment, the completely randomized standard error was 7.8, almost twice the paired-design estimate.The authors describe this as a substantial gain from paired randomization in that application.
7 The Design of Randomized Experiments and the Benefits of Stratification
The section develops power calculations and argues that ex ante stratification improves expected precision without a finite-sample tradeoff relative to complete randomization.
- Design recommendation: The authors recommend stratifying as much as possible until each stratum contains at least two treated and two control units.They generally prefer stratified over paired designs because paired analyses can impose additional costs despite possible precision benefits.
- Power calculations: Power calculations determine the minimum sample size needed to detect a prespecified treatment effect with prespecified significance level and power.The calculation depends on α, β, τ, σ2, and the treatment allocation share γ.
- Benefits of stratification: Stratification improves precision by preventing chance covariate imbalances while preserving the validity of randomization-based inference.Without stratification, the experiment remains valid and can still yield exact p-values, but inference may be less precise.
- Benefits of stratification: Ex ante stratification with common treatment probabilities cannot be worse than complete randomization, even in small samples with weak or zero covariate-outcome correlation.The authors conclude that committing to stratification can only improve precision, not lower it.
8 The Analysis of Clustered Randomized Experiments
Clustered randomization assigns treatment at the cluster level, requiring choices about estimands and analysis units when cluster sizes or within-cluster interactions matter.
- Efficiency and setting: Clustered designs are generally less efficient than completely randomized or stratified designs for a fixed sample size, but their motivations differ.The design may be chosen because assignment or sampling is naturally organized around schools, villages, states, or other clusters.
- Motivation and design: Cluster randomization can accommodate within-cluster interference and simplify sampling, potentially collecting data on more units for the same effort.When there is no interference across clusters, cluster-level assignment supports analyses that avoid modeling within-cluster interactions.
- Choice of analysis unit: Cluster-level analysis is transparent and usually preferred because unit-level covariates often add modest precision relative to cluster averages.When cluster sizes vary substantially, analyses should consider both cluster-average and unit-average targets.
- Estimands: Clustered experiments involve distinct population-average and cluster-average treatment effects, weighted respectively by cluster size and equally across clusters.The choice depends on substantive interest and the informativeness or precision of the available analysis.
- Estimands: Cluster-average effects can be estimated more precisely than population-average effects when a few large clusters coexist with many small clusters.In an extreme case, inference for the population effect is difficult because a mega-cluster remains entirely in one treatment group, while cluster-average inference may be precise.
- Reporting: The authors recommend reporting population-average and cluster-average analyses together when the population effect is substantively primary.The more precise cluster-average analysis can complement noisier inference for the population-average effect.
9 Noncompliance in Randomized Experiments
Non-compliance breaks the validity of comparing outcomes by treatment received because receipt can be systematically related to outcomes. The section presents three weak-assumption approaches and contrasts them with stronger-assumption analyses.
- Types of non-compliance: One-sided non-compliance occurs when control units cannot access the active treatment; two-sided non-compliance allows some control assignees to receive it.These patterns depend on whether violations occur only among treatment assignees or in both assignment groups.
- Motivation: Non-compliance makes treatment receipt endogenous: assignment is exogenous, but receipt may differ systematically across units and relate to outcomes.Therefore, randomization validates comparisons by assignment, not necessarily comparisons by post-treatment receipt.
- Approaches: The three valid approaches are intention-to-treat analysis, instrumental variables estimation of the complier effect, and partial-identification bounds for the full-population receipt effect.The instrumental-variables estimand is the local average treatment effect for compliers, while bounds provide a range for the population average effect.
- Stronger-assumption analyses: As-treated and per-protocol analyses require stronger assumptions than the three primary approaches, including unconfoundedness or dropping non-adhering units.The as-treated comparison relies on selection on observables, while per-protocol analysis excludes units who do not receive their assigned treatment.
- Intention-to-treat: Intention-to-treat estimates the causal effect of assignment by comparing realized outcomes across assignment groups while ignoring treatment receipt.Its confidence intervals are valid in large samples under randomization and SUTVA without additional assumptions.
- Instrumental variables and bounds: Instrumental variables identifies the average receipt effect for compliers under monotonicity and the exclusion restriction.Under these assumptions, the local average treatment effect is identified, and the corresponding bounds can be sharp.
10 Heterogenous Treatment Effects and Pretreatment Variables
The section frames treatment-effect heterogeneity and external validity as distinct uses of pretreatment variables. Covariate-specific effects can be combined with a target population’s covariate distribution to transport average effects.
- Motivation: Researchers study heterogeneity in treatment effects in addition to estimating average effects for the full sample or population.This differs from using pretreatment variables only to improve precision for an overall average treatment effect.
- External validity: External validity asks whether applying a treatment in a different setting produces the same effect.Differences between populations can be addressed when they are captured by observable pretreatment variables.
- Covariate-specific effects: Estimating τ(x) = E[Y_i(1) − Y_i(0)|X_i = x] allows average treatment effects to be estimated for populations with known covariate distributions.The target average is obtained by accounting for differences in the distribution of X_i.
10.1 Randomized Experiments with Pretreatment Variables
Subpopulation analyses can estimate and compare treatment effects, but post hoc subgroup selection and many covariate tests create multiple-testing concerns. The reviewed alternatives accommodate many covariates and complex heterogeneity while supporting valid inference.
- Subpopulation analysis: Researchers can estimate treatment effects separately across prespecified subpopulations and test equality between them.For example, an educational program’s effects may be analyzed separately for girls and boys.
- Multiple testing: Post hoc subgroup selection can invalidate p-values because of multiple testing.This concern arises when researchers examine many potential subgroups after seeing the data.
- Multiple testing: With 100 independent binary pretreatment variables unrelated to treatment effects, some large absolute t-statistics can arise by chance.The example motivates pre-analysis plans and multiple-testing corrections.
- Flexible heterogeneity: Recently developed methods address heterogeneity when covariates are numerous relative to sample size or the treatment-effect model is complex.These methods are presented as alternatives to simple subgroup analyses and aim to preserve valid confidence intervals.
10.2 Testing for Treatment Effect Heterogeneity
Testing treatment-effect heterogeneity requires accounting for many correlated hypotheses and can also use flexible tests of whether τ(x) varies with covariates. The section highlights both computational and specification limits.
- Flexible tests: Nonparametric tests can assess whether the conditional treatment effect τ(x) is constant across covariates under unconfoundedness.The procedure uses increasingly rich basis functions to approximate conditional expectations and test equality.
- Assumptions: Randomized assignment implies the unconfoundedness assumption used for these conditional heterogeneity tests.This connects the testing framework to randomized-experiment identification.
- Multiple testing: Testing each covariate separately for treatment-effect differences creates a multiple-testing problem requiring adjusted confidence intervals.Bonferroni correction may be overly conservative when covariates and their test statistics are correlated.
- Multiple testing: Bootstrap-based procedures can address multiple testing while accounting for correlation among test statistics.List, Shaikh, and Xu propose a computationally feasible approach for this setting.
- Limitations: A limitation of prespecified testing is that researchers cannot easily explore every covariate interaction or discretization.The approach requires fixing the hypothesis set before analysis.
10.3 Estimating Treatment Effect Heterogeneity
The section reviews parametric, nonparametric, and subgroup-based approaches for estimating heterogeneous treatment effects, emphasizing valid inference with covariates. It highlights honest sample splitting, data-driven partitions, and methods whose coverage can remain stable as covariates increase.
- Approaches: Researchers can model treatment-effect heterogeneity parametrically, estimate τ(x) nonparametrically, or construct covariate-space partitions that maximize heterogeneity.The section also considers regression interactions and regularized regression for systematic covariate selection.
- Recursive partitioning: Data-driven partitions identify subgroups with differing treatment effects without prespecifying the form of heterogeneity.The method uses a criterion to choose covariates and thresholds, and can explore interaction effects through a meaningful partition.
- Honest estimation: Sample splitting separates partition selection from subgroup estimation, reducing bias and supporting valid confidence intervals.An independent sample estimates treatment effects and standard errors after the partition is chosen; the resulting estimates are unbiased on the two subsamples.
- Recursive partitioning: The partitioning method outputs subgroups selected for treatment-effect heterogeneity together with subgroup treatment-effect estimates and standard errors.Its criterion is designed to minimize expected mean-squared error of treatment effects.
- Nonparametric methods: Modified random forests can produce asymptotically normal treatment-effect estimates with consistent asymptotic-variance estimation.Compared with K-nearest-neighbor matching or kernel methods, they can retain coverage with more covariates while producing more accurate treatment-effect estimates.
- Limitations: Confidence intervals need not deteriorate as covariate counts grow, but prediction mean-squared error can worsen instead.With too many covariates relative to sample size, existing methods may not provide valid confidence intervals; a single partition can also miss other heterogeneity.
11 Experiments in Settings with Interactions
The section examines randomized experiments when units affect one another through spillovers, peer effects, or equilibrium interactions. It reviews estimands and designs for clustered and general-network settings, while emphasizing unresolved methodological challenges.
- Forms of interference: Interference can arise through treatment spillovers, peer characteristics, peer behavior, or deliberate interactions between individuals.Examples include fertilizer leaching across agricultural plots and educational programs affecting untreated students through treated peers.
- Motivation: Interactions may complicate inference about an overall average effect or be the researcher's primary object of interest.The literature contains many approaches, and which methods will be most useful for empirical work remains unclear.
- Empirical examples: Randomized studies find externalities and causal peer effects in education, labor markets, dormitories, and military squadrons.Evidence includes substantial deworming externalities, labor-market redistribution, roommate effects, and outcomes varying with squadron composition.
- Empirical examples: In labor-market experiments, average treatment effects that increase when marginal treatment rates are lower suggest job redistribution from controls to trained individuals.Researchers varied assignment rates across labor markets and compared within-market treatment differences across markets.
- Clustered settings: For clustered interference, direct, indirect, total, and overall effects can be studied using two-stage randomization of clusters and then units.A common assumption is that outcomes depend on the fraction of treated peers, not their identities.
- Limitations: Without the fraction-treated assumption, proliferating indirect effects make unbiased estimation difficult.The assumption is often made, sometimes implicitly, in empirical work.
- General networks: General-network experiments ask what can be learned about interaction effects from randomizing a binary treatment on one network.The network is represented by a symmetric N × N adjacency matrix with zero diagonal.
12 Conclusion
The review presents statistical methods for analyzing randomized experiments, primarily from a randomization-based perspective. It spans classical inference, non-compliance, clustering, treatment-effect heterogeneity, and experiments with interactions.
- Scope: The review focuses primarily on randomization-based rather than model-based methods for analyzing randomized-experiment data.Its coverage runs from classic Fisher and Neyman methods through recent developments.
- Scope: The topics include non-compliance, clustered experiments, treatment-effect heterogeneity, and experiments in settings with interactions.These topics organize the review's modern methodological coverage.