Source-linked AI summary

Causal inference methods for combining randomized trials and observational studies: a review

Bénédicte Colnet, Imke Mayer, Guanhua Chen, Awa Dieng, Ruohong Li, Gaël Varoquaux, Jean-Philippe Vert, Julie Josse, Shu Yang

arXiv:2011.08047v4stat.ME

TL;DR

RCTs control confounding but may not represent target populations, whereas observational data can be representative but confounded. This review synthesizes methods for combining both sources, connects potential-outcomes and structural-causal-model approaches, and evaluates them in simulations and trauma data, while highlighting assumptions and overlap concerns.

  • Problem

    RCTs may lack external validity because of unrepresentativeness, while observational studies may conflate treatment effects with confounding.

  • Method

    The paper reviews identification and estimation methods for combining RCTs and observational data across generalizability, confounding assessment, efficiency, and causal frameworks.

  • Results

    The review finds that combining RCTs and observational data can improve statistical power and external validity, with simulations and tranexamic-acid analyses illustrating the methods.

  • Takeaways & Limitations

    Choosing between integrated analyses depends on assumptions about transportability or unconfoundedness, the selected variables, the target population, and covariate overlap.

  • Takeaways & Limitations

    Methods require that conditional treatment effects agree between observational data and the RCT given measured covariates, an assumption threatened by unobserved treatment modifiers.

Abstract

from arXiv · show

With increasing data availability, causal effects can be evaluated across different data sets, both randomized controlled trials (RCTs) and observational studies. RCTs isolate the effect of the treatment from that of unwanted (confounding) co-occurring effects but they may suffer from unrepresentativeness, and thus lack external validity. On the other hand, large observational samples are often more representative of the target population but can conflate confounding effects with the treatment of interest. In this paper, we review the growing literature on methods for causal inference on combined RCTs and observational studies, striving for the best of both worlds. We first discuss identification and estimation methods that improve generalizability of RCTs using the representativeness of observational data. Classical estimators include weighting, difference between conditional outcome models, and doubly robust estimators. We then discuss methods that combine RCTs and observational data to either ensure uncounfoundedness of the observational analysis or to improve (conditional) average treatment effect estimation. We also connect and contrast works developed in both the potential outcomes literature and the structural causal model literature. Finally, we compare the main methods using a simulation study and real world data to analyze the effect of tranexamic acid on the mortality rate in major trauma patients. A review of available codes and new implementations is also provided.

1 Introduction

RCTs provide strong control of confounding but may not represent the target population, while observational data can be representative yet confounded. This review organizes methods for combining both sources across generalizability, observational credibility, efficiency, and related causal-inference frameworks.

  • RCTs isolate treatment effects through randomized allocation but may lack representativeness and external validity.
  • Observational data can improve representativeness, while RCTs can help assess confounding in observational analyses.
  • Combining both sources can improve estimation of heterogeneous treatment effects when RCTs are under-powered.
  • The review covers generalizability, transportability, recoverability, and data fusion, connecting potential-outcomes and structural-causal-model literatures.
  • The paper reviews identification, estimation, software, simulations, and an application to treatment effects in trauma patients.

2 Problem setting

The paper formalizes how RCT and observational samples are represented and combined, focusing on separately sampled non-nested designs. It defines treatment effects, propensity and sampling quantities, and distinguishes generalizability from transportability.

  • Notations in the PO framework: Each individual is represented by covariates, potential outcomes, treatment assignment, and trial eligibility or willingness to participate.
  • Data structure: The review considers observational data with covariates alone or with observed treatment and outcome, alongside RCT observations.
  • Treatment effects: The CATE is defined conditionally on covariates, with separate population and trial versions based on potential outcomes.
  • Treatment effects: The population ATE and trial ATE can differ because they average treatment effects over different populations.
  • Study designs: In non-nested designs, RCT and observational samples are obtained separately, and the trial-selection score cannot be identified from the data.
  • Study designs: Generalizability concerns targets within or among trial-eligible individuals, whereas transportability includes individuals who are not trial-eligible.

3 When observational data have no treatment and outcome information

This section reviews how observational covariate data can generalize randomized-trial findings to a target population, emphasizing identification assumptions and estimators that adjust for population shifts. It covers weighting, outcome modeling, calibration, and doubly robust approaches, alongside practical limits from overlap and unmeasured treatment-effect modifiers.

  • Scope: The section considers generalizing trial findings to a target population when observational data provide covariates but no treatment or outcome information.The observational sample is treated as a random sample from the target population.
  • Identification assumptions: Identification requires consistency, conditional randomization in the trial, assumptions connecting trial participation to potential outcomes or treatment effects, and positivity of trial participation.Positivity requires adequate covariate overlap between the trial and target populations.
  • Identification assumptions: Ignorability over trial participation requires controlling for shifted baseline covariates that predict outcomes, whereas weaker assumptions need only shifted treatment-effect modifiers.Outcome-predictive covariates are not necessarily treatment-effect modifiers.
  • Estimation strategies: IPSW generalizes trial effects by weighting RCT observations according to inverse trial-participation odds, while stratification aggregates within-stratum effects using target-population proportions.Stratification is proposed to mitigate extreme weights.
  • Estimation strategies: Plug-in g-formula estimators model conditional outcome means among trial participants and marginalize predictions over the target population’s empirical covariate distribution.Under the stated assumptions, the plug-in g-formula converges toward the target treatment effect.
  • Robust estimation: Calibration and augmented estimators can provide balance or double robustness; augmented IPSW is consistent if either the participation-odds or outcome model is correctly specified.Calibration weighting also supports local efficiency when both calibration weights and outcome models are correctly specified.
  • Practical issues: Machine-learning nuisance estimation requires caution, including cross-fitting, and rate conditions can be needed for asymptotic normality or root-n consistency.The cited results require consistently estimated nuisance functions at rates such as n^1/4.
  • Practical issues: Generalization is constrained by limited overlap and by shifted treatment-effect modifiers absent from the covariates shared across datasets.When overlap fails, generalization may target an eligible subpopulation; missing modifiers can violate identifiability.

4 When observational data contain treatment and outcome information

Section 4 reviews how treatment-and-outcome information in observational data can complement RCTs, either by addressing unmeasured confounding or improving treatment-effect estimation. It covers bias-function, confounding-function, surrogate-outcome, and integrative approaches, including a reliability-based decision about pooling data.

  • 4.1 Dealing with unmeasured confounders in observational data: RCTs can ground observational analyses when unmeasured confounding threatens standard observational treatment-effect estimators.Former RCTs may serve as negative controls, while sensitivity analysis addresses remaining confounding.
  • 4.1 Dealing with unmeasured confounders in observational data: Secondary outcomes can identify primary-outcome effects when hidden confounding is shared, allowing observational treatment-effect differences on the surrogate to adjust RCT information.This is framed as latent unconfoundedness and a missing-data problem.
  • 4.1 Dealing with unmeasured confounders in observational data: Bias-function correction combines an observational CATE estimate with an RCT-estimated confounding bias function on their common covariate support.Under a low-complexity parametric bias model, the corrected CATE is consistent under the stated identification conditions.
  • 4.1 Dealing with unmeasured confounders in observational data: A confounding-function model jointly identifies heterogeneous treatment effects and confounding effects by coupling RCT and observational data under parametric assumptions.The framework can generalize RCT ATEs without requiring covariate-distribution overlap, and its integrative CATE estimator is strictly more efficient than the RCT estimator when predictors are linearly independent.
  • 4.2 Toward more efficient estimation: When both RCT transportability and observational unconfoundedness hold, pooling can improve CATE efficiency; a preliminary test can exclude observational data when comparability is doubtful.The elastic integrative estimator uses a data-adaptive switch intended to achieve mean squared error no worse than the RCT-only estimator.

5 Structural causal models (SCM) and transportability

Section 5 connects structural causal models and do-calculus to transportability and data fusion. It presents selection-diagram assumptions and transport formulas, while noting that practical SCM estimators and detailed implementation guarantees remain limited.

  • 5.1 Formulating transportability in the SCM framework: SCM and potential-outcomes formulations are formally equivalent when potential outcomes are identified with intervention distributions.The review uses this correspondence to connect causal-effect expressions across the two frameworks.
  • 5.1 Formulating transportability in the SCM framework: SCM transportability generalizes RCT intervention distributions to a target population by combining trial conditional causal distributions with target-population covariate information.The RCT identifies P(Y | do(A = a), X, S = 1), which is combined with target-population information through a transport formula.
  • 5.1 Formulating transportability in the SCM framework: Selection diagrams encode population differences through arrows from selection status to covariates, treatment, or outcomes.An arrow from selection to outcomes can prevent transportability, whereas the corresponding absence of an arrow supports the transportability assumption.
  • 5.1 Formulating transportability in the SCM framework: Post-treatment covariate adjustment can identify a transport formula when X is S-admissible but not S-ignorable.The displayed selection diagram represents different effects of A on X across populations.
  • 5.1 Formulating transportability in the SCM framework: SCM methods offer versatile identification tools, but practical estimators with public implementations and detailed consistency, convergence-rate, or robustness results remain scarce.Available work extends weighting and double/debiased machine-learning methods, while partial-identification methods provide bounds when effects are not identifiable.

6 Software for combining RCT and observational data

Section 6 reviews software resources and evaluates generalization estimators in simulations. The simulations show unbiased estimation under well-specified models and distinct robustness patterns under model misspecification.

  • 6.1 Software: The review catalogs implementations for both causal-effect identifiability and estimation, emphasizing software as a bridge between causal theory and practice.It also provides reproducible R implementations for several estimators introduced in the review.
  • 6.2 Simulation study of generalization estimators: The well-specified simulation compares IPSW, normalized IPSW, stratification, plug-in g-formula, calibration weighting, AIPSW, and ACW over 100 simulations.Figure 5 reports estimated ATEs relative to the true effect and an RCT-only baseline.
  • 6.2 Simulation study of generalization estimators: The true target-population ATE was 27.4, while the RCT-only estimate averaged 14.24 because treatment-effect modifiers differed between trial and population.All estimators were unbiased in this well-specified scenario, although IPSW had greater variability than the other estimators.
  • 6.2 Simulation study of generalization estimators: When the sampling propensity model is misspecified, IPSW is biased, whereas outcome-model misspecification biases the plug-in g-estimator.Figure 6 evaluates these cases across the listed estimators over 100 simulations.
  • 6.2 Simulation study of generalization estimators: AIPSW remains unbiased when either the sampling propensity or outcome model is misspecified, while CW and ACW remain robust when both models are misspecified.The latter result demonstrates robustness of calibration against slight model misspecification.

7 Application: Effect of Tranexamic Acid

The application combines CRASH-3 trial data with the Traumabase registry to assess TXA mortality effects and generalize trial findings to a target TBI population. Results broadly align with the RCT but vary across estimators, with missingness and distribution differences complicating interpretation.

  • 7 Application: Effect of Tranexamic Acid: The analysis assesses TXA’s potential effect on mortality among TBI patients using CRASH-3 and the observational Traumabase registry.The application targets intracranial bleeding and uses data from both an RCT and an observational study.
  • 7.1.2 Purely-observational results from two different estimation strategies: The observational AIPW analyses found no evidence of a TXA mortality effect, whereas IPW results indicated a possible deleterious effect.The observational analysis adjusted for 17 confounding variables and included 21 additional outcome-predictive variables unrelated to treatment.
  • 7.2.1 Context: CRASH-3 found head-injury-related death in 18.5% of TXA-treated patients versus 19.8% with placebo, with a nonsignificant RR of 0.94 [95% CI 0.86 - 1.02].A positive effect was reported only for mild and moderate cases.
  • 7.3 Analyses: Covariate comparisons reveal treatment-assignment bias in the observational study and balanced treatment groups in the RCT.The combined analysis therefore requires assessment of common baseline covariates, treatment, and outcome before comparing datasets.
  • 7.3.2 Analyses: Generalization estimators split between supporting CRASH-3’s nonsignificant effect and indicating a deleterious effect, while missing values may affect estimator performance.Large confidence intervals for GRF weight estimators are likely related to imbalanced proportions of missing values.

8 Conclusion

The review presents combining RCTs and observational data as a way to improve causal inference, including external validity and statistical power. It emphasizes that trustworthy analyses depend on identification choices, appropriate covariates, and careful handling of missing data.

  • 8 Conclusion: Combining observational data and RCTs can improve statistical power and the external validity of causal inference.The review focuses substantially on generalizing and transporting RCT findings across populations.
  • 8 Conclusion: The literature uses varying terminology and scattered implementations, while proposed methods still lack sufficient real-world benchmarks.Relevant terms include generalizability, representativeness, external validity, transportability, and data fusion.
  • 8 Conclusion: Choosing between analyses depends on assumptions concerning transportability or unconfoundedness, and the application found generalized estimates concordant with the RCT for at least half of estimators.The purely observational AIPW analysis also supported the RCT findings.
  • 8 Conclusion: Domain expertise and causal graphs can guide covariate selection and help assess whether a causal question is identifiable.The structural causal model framework is presented as a principled way to avoid biased causal-effect estimates.
  • 8 Conclusion: Missing-data methods require care because strategies such as weighting and multiple imputation rely on untestable assumptions about the missingness mechanism.Missing values are typically more frequent in observational data, while RCTs can also experience missed visits and dropout.
  • 8 Conclusion: In a single RCT, difference-in-means estimation is unbiased for the population ATE only when the trial is a random sample of the target population.Otherwise, the estimator can be biased for the population average treatment effect despite random treatment assignment within the trial.

B Estimation of ATE in observational data

Observational-data estimation of treatment effects requires unconfoundedness and overlap, after which weighting, regression, and doubly robust approaches can identify and estimate ATEs and CATEs.

  • Identification assumptions: Unconfoundedness requires all confounding factors affecting treatment and outcome to be measured in covariates X.This assumption makes treatment assignment conditionally as good as random.
  • Identification assumptions: Overlap requires the propensity score e(x) to remain bounded away from 0 and 1.Formally, some η > 0 must satisfy η < e(x) < 1 − η almost surely.
  • Identification: Under unconfoundedness and overlap, observational data identify the ATE through reweighting and regression formulations.The review introduces these identification formulas before discussing corresponding estimators.
  • Estimators: Inverse Propensity Weighting reweights observations using the propensity score to balance treated and untreated groups on covariates.Treated observations with small propensity scores receive larger weights, with the reverse applying to untreated observations.
  • Theoretical properties: The review surveys theoretical results for IPSW, stratification, calibration weighting, and augmented calibration weighting, including consistency, asymptotic normality, bias, and variance.Several results depend on correctly specified or consistently estimated propensity or selection models.
  • Nested designs: Nested trial designs identify both the overall treatment effect and the effect among nonparticipants, because nonrandomized sampling probabilities are observed.The latter quantity can characterize heterogeneity within the cohort and variables related to sampling or treatment effects.

E.1 When observational data have no outcome and treatment information

Nested trial designs embed an RCT within a cohort, allowing observed sampling probabilities and target-population distributions to support IPSW, g-formula, and doubly robust estimation.

  • Estimators: IPSW, the g-formula, and doubly robust estimators are adapted to nested designs using observed participation and population-specific covariate distributions.The IPSW weights differ from the non-nested case because πS can be estimated directly.
  • Nested trial design: In a nested design, trial participation is observed, so the sampling probability of nonrandomized individuals is identifiable.The design distinguishes participants with S = 1 from nonparticipants with S = 0.
  • G-formula: The nested-design g-formula estimates E[Y(a) | S = 0] by averaging trial conditional outcome means over the nonrandomized population.The identifying expression uses E[E[Y | X, S = 1, A = a] | S = 0].
  • Implementations: Available implementations include causal-effect identification tools, causal-fusion interfaces, IPSW code, and nested-design g-formula code.The cited software supports graphical identification, transportability queries, weighting, regression, and sandwich variance estimation.
  • SCM framework: Structural causal models represent variables, structural functions, exogenous-variable distributions, and interventions through graphs and do-operators.The do(A = a0) operation replaces A’s mechanism with a constant and deletes incoming arrows into A.

F.1.2 Sample selection bias

The review frames sample selection bias as estimating a causal effect from selected data and shows how graphical assumptions can make the effect identifiable or combine selected and population data.

  • Selection bias: Sample selection bias arises when available data follow P(A, Y | S = 1) but the target is P(Y | do(A = a)).S indicates whether a unit belongs to the observed sample.
  • Graphical identification: When selection is d-separated from the outcome by treatment and treatment is unconfounded, the experimental distribution can recover the causal effect.In that case, P(y | do(a)) = P(y | a) = P(y | a, S = 1).
  • Covariate selection: The SCM framework can select covariates that both control confounding and remain usable under biased sampling.In the illustrated case, X is the available adjustment set because it is separated from S.
  • Data fusion: The S-backdoor admissible criterion permits combining biased outcome data with unbiased population covariate data when X blocks paths between selection and outcome and is measured in both sources.The resulting post-stratification formula recalibrates selected-sample information using P(X = x) from the target population.
  • Caveat: Post-treatment covariates generally invalidate the post-stratification formula because S-ignorability is rarely satisfied in that setting.The review therefore restricts the formula’s interpretation to appropriate pre-treatment adjustment variables.

G Additional simulation results

This section provides additional simulation results following the earlier simulation analysis.

  • Additional simulation results: The section presents additional results for the simulations described in Section 6.2.

G.1 Distributional shift between RCT and observational samples

The simulations examine how distributional shifts between RCT and observational samples affect ATE estimation. Stronger shifts increase variance for weighted and calibration-weighting estimators, while stratification improves as the number of strata rises.

  • Distributional shift: The simulated RCT covariates are generally lower than those in the observational sample, while overlap remains valid.Each target-sample observation retains a non-zero probability of inclusion in the experimental sample.
  • Distributional shift: A stronger shift in X1 is created by changing its sampling-model coefficient from −0.5X1 to −1.5X1.
  • Estimator behavior: Stronger covariate shifts increase the variance of weighted and calibration-weighting estimators.
  • Estimator behavior: Increasing the number of stratification strata from 3 to 15 improves prediction in the simulations.

G.3 Impact of a hidden treatment effect modifier

The analysis shows that a treatment-effect modifier affected by sampling must be included for unbiased transport, while missing covariate information and observational confounding complicate causal analysis.

  • Hidden treatment-effect modifier: Using only X1 in the sampling propensity score leaves IPSW unbiased because X1 is the shifted treatment-effect modifier.
  • Hidden treatment-effect modifier: When X1 both affects RCT sampling and modifies treatment effects, omitting X1 from IPSW produces strong bias.
  • Homogeneous treatment effects: With homogeneous treatment effects, the RCT alone consistently estimates the population ATE, so transport methods are unnecessary on the absolute-difference scale.
  • Missing values: Traumabase covariates have 0 to nearly 60% missing values, with codes that may indicate MCAR or informative MNAR mechanisms.
  • Missing values: Missing-data causal analysis requires assumptions about both complete-case effect identification and the mechanism generating missingness.
  • Covariate adjustment: Expert elicitation and causal graphs are used to identify plausible confounders because direct treatment-effect estimation from Traumabase is confounded.
  • Study alignment: CRASH-3 restricts treatment administration to within 3 hours, whereas Traumabase lacks exact timing and administration-type information.
  • Study alignment: The studies differ in outcome definitions, since Traumabase includes several causes of 28-day in-hospital death beyond TBI-related death.

H.3 Additional analysis

Additional analyses visualize distributional differences between CRASH-3 and Traumabase across key clinical covariates. These comparisons include age, Glasgow score, systolic blood pressure, sex, and pupil reactivity.

  • Analysis overview: The additional analysis also examines principal components, propensity-score histograms and scatter plots, and patient strata by injury severity.
  • Distributional shift: Histograms compare empirical distributions between Traumabase and CRASH-3 for age, Glasgow score, systolic blood pressure, sex, and pupil reactivity.
  • Distributional shift: CRASH-3 contains more young patients, whereas Traumabase contains more moderate cases associated with higher Glasgow scores.

H.3.2 Principal component analysis

The principal-component and propensity-score analyses characterize relationships and differences between CRASH-3 and Traumabase. Stratified treatment-effect analyses show discrepant evidence, while clinically relevant transport to some subgroups remains unsupported by the stated assumptions.

  • Principal component analysis: PCA of the combined data relates Glasgow coma scale and pupil reactivity, reflecting their shared association with intracranial injury and consciousness.
  • Propensity-score analysis: Conditional-odds estimates from logistic regression and forests include extreme values, with grf strengthening this trend.
  • Propensity-score analysis: The forest method uses Traumabase missing values when learning the propensity-score model.
  • Treatment effects by stratum: Across severity strata, Traumabase shows no average treatment effect, whereas CRASH-3 finds a beneficial effect for mild TBI.
  • Treatment effects by stratum: Treatment effects may be biased toward the null among patients with very severe brain injury and poor prognosis regardless of treatment.
  • Transport limitations: For mild-to-moderate TBI with major extracranial bleeding, the presented methods cannot satisfy the assumptions needed for transport.
  • Software and data: The review provides an inventory of publicly available identification and estimation software for generalization.
  • Software and data: Tables report study sample sizes and ATE estimates for Traumabase and reproduced CRASH-3 analyses across injury-severity strata.
Loading 2011.08047v4…