Source-linked AI summary
External Validity: From Do-Calculus to Transportability Across Populations
Judea Pearl, Elias Bareinboim
TL;DR
The paper addresses how causal effects learned experimentally can be generalized to a target population where only observational data are available. It introduces selection diagrams and uses do-calculus to decide transportability and derive the required bias-free combination of evidence. The framework also clarifies measurement needs, theoretical boundaries, and assumptions underlying such transfers.
Problem
The paper addresses how to infer causal effects in a target population from experimental findings in a study population when only observational data are available in the target.
Method
The paper uses selection diagrams to encode population differences and reduces transportability questions to symbolic derivations in do-calculus.
Results
The procedures decide whether transport is possible and identify how experimental and observational findings can be combined to estimate target-population causal effects without bias.
Takeaways & Limitations
Transport formulae specify essential measurements in both populations and provide principled calibration for differences between them.
Takeaways & Limitations
The results are worst-case under nonparametric assumptions allowing every variable to be an effect modifier, and real applications may involve measurement error, selection bias, finite samples, uncertain graphs, or unmeasured confounding.
Abstract
from arXiv · showhide
The generalizability of empirical findings to new environments, settings or populations, often called "external validity," is essential in most scientific explorations. This paper treats a particular problem of generalizability, called "transportability," defined as a license to transfer causal effects learned in experimental studies to a new population, in which only observational studies can be conducted. We introduce a formal representation called "selection diagrams" for expressing knowledge about differences and commonalities between populations of interest and, using this representation, we reduce questions of transportability to symbolic derivations in the do-calculus. This reduction yields graph-based procedures for deciding, prior to observing any data, whether causal effects in the target population can be inferred from experimental findings in the study population. When the answer is affirmative, the procedures identify what experimental and observational findings need be obtained from the two populations, and how they can be combined to ensure bias-free transport.
1. INTRODUCTION: THREATS VS. ASSUMPTIONS
The paper reframes external validity as transportability across differing environments and develops formal licenses for when causal findings may be transferred. It uses causal diagrams and do-calculus to decide feasibility and derive unbiased combinations of experimental and observational evidence.
- Scope of contribution: The paper offers theoretical limits on generalization, including which population differences can be addressed by design and which prohibit transfer.These limits supplement existing statistical methodologies for combining evidence from diverse studies.
- Threats versus assumptions: Prior literature primarily describes threats to generalization rather than formal assumptions that license transport across differing environments.Meta-analysis and hierarchical models commonly pool studies without explicitly distinguishing experimental from observational evidence.
- Motivation: External validity concerns generalizing empirical findings across populations, settings, treatments, or measurement variables.The paper focuses on transferring causal effects from experimental studies to new populations.
- Formal language: Causal diagrams provide a language for precisely characterizing environments and encoding differences among populations.Intervention and counterfactual models make formalization of transportability possible.
- Threats versus assumptions: The paper formulates licenses to transport as assumptions that, if true, permit causal results to be transferred across studies.This shifts the focus from communicating what may go wrong to specifying conditions under which transport is justified.
- Formal language: Do-calculus yields algorithms for deciding whether transportability is feasible and combining experimental and observational findings into unbiased target-population estimates.The procedures also identify the evidence required from the study and target populations.
2. PRELIMINARIES: THE LOGICAL FOUNDATIONS OF CAUSAL INFERENCE
The preliminaries represent causal analysis as inference from defended assumptions, causal queries, and data. Structural equation models, interventions, d-separation, and do-calculus provide the machinery for identifying causal effects and adapting it to transportability.
- Causal models as inference engines: Causal analysis combines qualitative assumptions, causal queries, and experimental or nonexperimental data.The model encodes assumptions mathematically, while queries concern causal or counterfactual relationships among variables.
- Structural equation models: A structural equation model consists of exogenous variables, observable endogenous variables, functions determining endogenous values, and a distribution over exogenous variables.The functions represent causal processes and remain invariant unless explicitly intervened on.
- Causal models as inference engines: The inference engine produces logical implications, conditional claims, and testable data-fitness measures from assumptions, queries, and data.Testable implications can include conditional independencies or equality constraints used to assess model compatibility.
- Interventions and causal effects: The do(x) operator simulates a physical intervention by replacing the equation for X with the constant X = x while leaving other model components unchanged.The resulting postintervention distribution of Y supports comparisons of treatment efficacy across intervention levels.
- Identification and causal calculus: Identification asks whether a causal query can be estimated from data together with a partially specified model such as a causal graph.In nonparametric models, identification requires expressing the intervention quantity in terms of observable probabilities.
- Identification and causal calculus: D-separation encodes conditional independencies implied by causal assumptions, while do-calculus removes intervention operators when deriving identifiable expressions.For transportability, the goal changes from eliminating do-operators to separating them from variables representing population disparities.
3. INFERENCE ACROSS POPULATIONS: MOTIVATING EXAMPLES
The examples show that transporting causal effects from a randomized Los Angeles study to New York City requires formulas tailored to how population differences relate causally to treatment, outcome, and intermediate variables. Age-specific effects can be reweighted by the target age distribution, but proxy variables and treatment-dependent biomarkers require different reasoning and, in some cases, target-population observational data.
- Example 1: Example 1 asks how age-specific causal effects from a Los Angeles randomized trial can identify the overall treatment effect in older New York City.The setup combines experimental effects by age with the target population’s differing age distribution.
- Example 1: The initial transport formula combines Los Angeles age-specific experimental effects with New York City’s observational age distribution.Its validity depends on assumptions about invariance across cities that the paper seeks to make explicit.
- Example 2: With linguistic skills as an unmeasured-age proxy, differences in skill distributions may reflect age differences or changed skill–age relations, so the appropriate formula depends on causal context.The example asks whether the Los Angeles skill-specific effect can identify the New York City effect when age is unavailable.
- Example 2: If skill–age relations are shared and population differences reflect genuine age differences, age may modify treatment responses and the simple proxy-based transport formula becomes invalid.The paper concludes that skill-specific causal effects need not remain invariant across populations.
- Example 3: For a treatment-dependent biomarker on the pathway from treatment to disease, the overall effect is not a simple average of biomarker-specific effects.The original weighting rule applies only when the biomarker is unaffected by treatment; the correct rule instead reweights by the target conditional distribution P*(z|x).
- Example 3: In the surrogate-endpoint setting, target-population observational data on P*(z,x) can support reassessment when treatment affects the outcome directly and through the biomarker, under stated assumptions.The required transport formula differs from both earlier formulas and uses P*(z|x) estimated in the target population.
4. FORMALIZING TRANSPORTABILITY
Transportability is formalized as a causal problem: population differences must be localized to mechanisms, represented with selection diagrams, and analyzed through do-calculus to determine when transport is possible.
- 4.1 Selection Diagrams and Selection Variables: Transportability depends on causal relations and mechanisms, not merely on statistical differences between populations.Observed disparities such as P(z) ≠ P*(z) can arise from different causal mechanisms and therefore lead to different transport formulas.
- 4.1 Selection Diagrams and Selection Variables: A second randomized experiment in the target population can permit transport formulas under more relaxed assumptions, but at higher cost.The paper specifically notes that this can allow X and Z to be confounded.
- 4.1 Selection Diagrams and Selection Variables: Transportability assumptions are constrained when causal mechanisms are unmeasured or when structural changes between domains violate the diagram's representation.The paper excludes nature-induced disparities from design differences and notes that differing causal directionality lies beyond its scope.
- 4.1 Selection Diagrams and Selection Variables: Selection diagrams distinguish population differences in age distributions, in how Z depends on age, and in how Z depends on X.Dashed arcs additionally represent latent variables affecting pairs of variables.
- 4.1 Selection Diagrams and Selection Variables: The absence of a selection node pointing to an outcome mechanism represents the assumption that age-specific causal effects are invariant across populations.In Figure 4(a), this absence supports direct transportability of the z-specific effects.
- 4.1 Selection Diagrams and Selection Variables: Selection diagrams encode population differences by adding selection variables that point to mechanisms suspected to differ across domains.Switching populations is represented by conditioning on different values of the selection variables.
- 4.2 Transportability: Definitions and Examples: Theorem 1 turns the declarative definition of transportability into an effective procedure based on sequential do-calculus derivations.The procedure determines whether transportability can be demonstrated from the available causal representation.
- 4.2 Transportability: Definitions and Examples: When population differences affect only the conditional mechanism for Z, pa(Z)-specific effects and the overall target effect may be directly transportable even when z-specific effects are not.In Example 6, confounding arcs make these quantities nonidentifiable despite their direct transportability.
5. TRANSPORTABILITY OF CAUSAL EFFECTS—A GRAPHICAL CRITERION
The paper gives graphical criteria for deciding whether causal effects transport between populations and for deriving the required transport formulas. The criteria use S-admissibility and recursive applications of Theorem 3 to determine which experimental and target-population observations suffice.
- Theorems 2 and 3 decide algorithmically whether a causal relation is transportable and specify its transport formula.
- S-admissibility: A set Z is S-admissible when it d-separates Y from the selection variables S after intervening on X.
- S-admissibility: If observed pretreatment covariates Z are S-admissible, the average causal effect is transportable using the weighting formula of Corollary 1.
- Recursive criterion: Theorem 3 extends transportability to treatment-dependent covariates through trivial transportability, S-admissibility, or d-separation conditions applied recursively.
- Transport formulas: In the Figure 6(d) example, experimental factors and target-population observational factors combine without requiring estimation of the joint intervention effect in the experiment.
- Graphical cases: Theorem 3 establishes transportability in several graphical examples, but Figure 6(f) is nontransportable because no S-admissible set exists and condition 3 fails.
6. CONCLUSIONS
The paper frames transportability as a formal problem of licensing causal generalization across populations rather than merely warning about threats. Its criteria guide measurement choices and clarify both theoretical limits and assumptions under which additional transport licenses may be available.
- Selection diagrams and transport formulas let investigators determine whether target-population causal effects can be inferred and calibrate results for population differences.
- The framework identifies essential measurements in experimental and observational studies, reducing measurement costs and sampling variability.
- Scope and assumptions: Theorems 2 and 3 provide worst-case inferences because every variable may potentially modify effects.
- Scope and assumptions: Assuming noninteractive or monotonic relationships can yield additional transport licenses beyond those sanctioned by the theorems.
- Extensions: The method can also support transporting statistical findings between observational studies by avoiding repeated measurements and improving precision through pooled data.
- Scope and assumptions: Immediate use requires sufficient qualitative background knowledge about how the study and target populations differ.
- Scope and assumptions: Real-world applications also face measurement error, selection bias, finite-sample variability, graph uncertainty, and possible unmeasured confounding.
APPENDIX
The appendix derives transport formulas for causal effects in the models of Figures 6(d) and 7, using conditions of Theorem 3 and S-admissibility criteria.
- The appendix derives the transport formula for the causal effect in the model of Figure 6(d), equation (5.9).
- For the Figure 6(d) derivation, the conditions reference S-admissibility of Z on CE(X,Y) and of T on CE(X,W).
- The Figure 6(d) derivation also uses the empty set {} of CE(X,W) and multiple conditions of Theorem 3.
- The appendix derives the transport formula for the causal effect in the model of Figure 7, equation (5.10), using S-admissibility of Z on CE(X,Z).