Source-linked AI summary

Outcome-wide longitudinal designs for causal inference: a new template for empirical studies

Tyler J. VanderWeele, Maya B. Mathur, Ying Chen

arXiv:1810.10164v1stat.ME

TL;DR

The paper addresses limitations of single exposure–outcome observational analyses, including unmeasured confounding, investigator discretion, and fragmented evidence across outcomes. It proposes an outcome-wide longitudinal design that analyzes multiple subsequent outcomes with coordinated confounder control, sensitivity metrics, and multiple-testing procedures. The authors argue that this design can reduce analytic discretion, improve efficiency, and broaden the assessment of exposure effects.

  • Problem

    Single exposure–outcome analyses remain vulnerable to unmeasured confounding, investigator degrees of freedom, selective reporting, and inefficiently fragmented evidence.

  • Method

    The outcome-wide longitudinal design applies one exposure analysis across multiple subsequent outcomes using coordinated covariate and modeling decisions, E-values, and multiple-testing methods.

  • Results

    In the data analysis example, 17 of 24 tests remained significant after Bonferroni correction, while 15 of 17 continuous-outcome tests remained significant under Romano–Wolf correction.

  • Takeaways & Limitations

    The design can convey effects across many outcomes in one study while reducing outcome-specific analytic discretion and improving efficiency for the research community.

  • Takeaways & Limitations

    Outcome-wide analyses do not eliminate selective outcome reporting or all researcher degrees of freedom, because analysts can still vary covariates, modeling approaches, or reported outcomes across studies.

Abstract

from arXiv · show

In this paper we propose a new template for empirical studies intended to assess causal effects: the outcome-wide longitudinal design. The approach is an extension of what is often done to assess the causal effects of a treatment or exposure using confounding control, but now, over numerous outcomes. We discuss the temporal and confounding control principles for such outcome-wide studies, metrics to evaluate robustness or sensitivity to potential unmeasured confounding for each outcome, and approaches to handle multiple testing. We argue that the outcome-wide longitudinal design has numerous advantages over more traditional studies of single exposure-outcome relationships including results that are less subject to investigator bias, greater potential to report null effects, greater capacity to compare effect sizes, a tremendous gain in the efficiency for the research community, a greater policy relevance, and a more rapid advancement of knowledge. We discuss both the practical and theoretical justification for the outcome-wide longitudinal design and also the pragmatic details of its implementation.

1. INTRODUCTION

The paper proposes the outcome-wide longitudinal design, which applies confounding-control analyses for one exposure across multiple subsequent outcomes. It supplements this structure with E-values, simultaneous analytic decisions, and multiple-testing methods to address unmeasured confounding and investigator discretion.

  • Addressing criticisms: Covariate-control and basic modeling decisions are made simultaneously across outcomes to reduce outcome-specific investigator discretion.The conventional template permits choices about models and covariates that can be made separately across analyses, potentially favoring expected results.
  • Proposed template: The outcome-wide longitudinal design applies a single-exposure causal analysis simultaneously to multiple outcomes measured after the exposure.The proposed template extends the existing single exposure–outcome structure rather than replacing confounding control.
  • Proposed template: The design reports E-values to characterize how sensitive each estimate is to potential unmeasured confounding.E-values are presented as a metric related to the robustness of estimates to one or more potential unmeasured confounders.
  • Motivation and scope: The approach supports broader assessment of an exposure’s effects across outcomes relevant to human flourishing, including happiness, health, purpose, relationships, and financial security.The authors frame broad outcome coverage as relevant to policy, public health, and knowledge advancement.
  • Motivation and scope: Applying one exposure analysis across numerous outcomes can convey more information in a single study and improve efficiency for researchers, readers, and the research community.The authors contrast this with searching across many separate single exposure–outcome studies.

2. LONGITUDINAL DESIGNS FOR CAUSAL INFERENCE

The longitudinal framework uses temporally ordered exposure, covariate, and outcome data to support causal interpretation under conditional no-confounding assumptions. It emphasizes rich, outcome-aware confounder control while coordinating covariates across outcomes and accounting for timing, mediation, reverse causation, and prior exposure.

  • Temporal ordering and assumptions: Causal interpretation requires the exposure to precede subsequent outcomes and measured covariates to support comparability between exposed and unexposed groups.Cross-sectional data generally cannot establish the temporal ordering needed for this interpretation.
  • Temporal ordering and assumptions: Under conditional no confounding, observed exposure–outcome contrasts can identify causal effects on difference and ratio scales.The paper expresses this using counterfactual outcomes and observed conditional associations.
  • Timing and prior measures: Baseline outcome adjustment can help address reverse causation, but covariate timing requires balancing residual confounding against adjustment for mediators.A prior wave that is too distant may be less useful for confounding control, whereas contemporaneous adjustment can block part of a total effect.
  • Confounder selection: Confounder selection should include common causes of exposure and outcome while excluding variables on the exposure-to-outcome pathway when estimating total effects.The paper also discusses the disjunctive cause criterion, which includes pre-exposure causes of the exposure, outcome, or both.
  • Outcome-wide confounding control: Outcome-wide analyses favor a common covariate set across outcomes because outcome-specific covariate choices increase discretion and complicate analysis and reporting.The authors also note that outcomes may affect one another, making outcome-specific causal assessment difficult.
  • Timing and prior measures: Controlling for prior exposure can further address reverse causation and constrain explanations based on unmeasured confounding.The paper describes prior-exposure adjustment through covariate inclusion or stratification.

3. E-VALUES FOR UNMEAURED CONFOUNDING AND OTHER BIASES

Outcome-wide analyses use sensitivity metrics to assess how strongly unmeasured confounding would need to be to explain observed associations, while also addressing measurement error and missing-data biases. The E-value is recommended for each outcome, but its interpretation depends on measured confounding control and other sources of bias.

  • Sensitivity analysis for unmeasured confounding: For an observed risk ratio of 1.3, an unmeasured confounder associated 1.92-fold with both exposure and outcome would suffice to reduce the association to the null, whereas weaker confounding would not.
  • Sensitivity analysis for unmeasured confounding: Outcome-wide studies should report E-values for both each effect estimate and the confidence-interval limit closest to the null.The confidence-interval E-value indicates the minimum confounding needed to shift the interval to include the null and allows robustness to vary across outcomes.
  • Sensitivity analysis for unmeasured confounding: The E-value quantifies the minimum confounding strength, conditional on measured covariates, that could explain away an observed risk-ratio association.It allows multiple unmeasured confounders by interpreting the parameters as maximum associations across confounder levels.
  • Sensitivity analysis for unmeasured confounding: The E-value is conservative: its confounding threshold can be sufficient under some scenarios but insufficient under others, because the maximum-bias relation is an inequality.Actual bias may be smaller, particularly when an unmeasured confounder is rare, because the E-value assumes an unfavorable distribution of the confounder.
  • Sensitivity analysis for other types of bias: Causal interpretation also requires attention to measurement error, missing data, and exposure- or outcome-related selection or censoring.Nondifferential measurement error often biases estimates toward the null, whereas differential measurement error often biases them away from the null.
  • Sensitivity analysis for other types of bias: For missing data, the authors generally recommend outcome-wide multiple imputation alongside complete-case comparisons, with additional sensitivity analyses when appropriate.If missing-data concerns require more careful outcome-specific handling, abandoning the outcome-wide approach may be preferable.

4. MULTIPLE TESTING METRICS

Outcome-wide analyses require methods that address false positives across many exposure–outcome tests while preserving interpretable evidence below strict adjusted thresholds. The paper discusses Bonferroni, correlation-aware metrics, and contextual interpretation of p-values.

  • Bonferroni correction: Bonferroni correction supports the stronger conclusion that at least J associations are true when J associations reject at the a/K threshold, with error rate at most a.This extends its usual global-null interpretation, which only concerns whether at least one true association exists.
  • Bonferroni correction: In large cohorts with moderate numbers of tests, Bonferroni correction often changes the detectable effect-size range only slightly.The penalty can nevertheless be substantial in smaller studies, studies with many outcomes, or trials powered for one primary outcome.
  • Bonferroni correction: Researchers can report actual p-values alongside the number of tests and the Bonferroni threshold, allowing readers to assess both nominal and corrected evidence.The paper presents this as an alternative to choosing definitively between corrected and uncorrected reporting.
  • Correlation-aware metrics: Correlation-aware multiple-testing metrics can preserve familywise error while being less conservative and can compare observed rejections with a null expectation accounting for outcome correlation.The authors describe a confidence-interval approach and an R package for continuous outcomes.
  • Interpreting evidence: P-values that miss adjusted thresholds should not automatically be discarded because p-values provide continuous evidence and differences such as 0.04 versus 0.06 may be small.The paper recommends considering nominal results, effect sizes, and evidence across studies rather than imposing a single magical cutoff.

5. DATA ANALYSIS EXAMPLE

The paper illustrates the outcome-wide design using parental warmth in childhood and 24 later flourishing, mental-health, and health-behavior outcomes. Most associations remained significant under stringent corrections, and several estimates appeared robust to moderate unmeasured confounding.

  • Study design: N = 2,984 MIDUS participants were analyzed longitudinally to assess childhood parental warmth against multiple later outcomes.The sample was drawn to include siblings and twin pairs, with one sibling randomly selected for simplicity.
  • Study design: 24 outcomes had median correlation magnitude 0.25, ranging from 0.0007 to 0.88, and one additional standard deviation of parental warmth predicted 0.20 standard deviations greater mid-life flourishing.The flourishing estimate was adjusted for demographics and childhood family factors, with 95% CI [0.16, 0.24].
  • Study results: 18 of 24 outcomes were significantly associated with parental warmth at a = 0.05, including 17 significant at a = 0.01, with all effect directions indicating improved flourishing outcomes.The associations were reported from the outcome-wide analysis in Table 2.
  • Study results: E-values for several flourishing outcomes and depression exceeded 1.5, indicating that confounding associations of 1.5-fold each could shift confidence intervals to the null, whereas weaker confounding could not.The paper interprets at least some flourishing estimates as reasonably robust to moderate unmeasured confounding.
  • Multiple-testing results: After correction, 17 of 24 tests remained significant by Bonferroni, while 15 of 17 continuous outcomes remained significant under Romano–Wolf correction at both a = 0.05 and a = 0.01.Relative to the global-null expectation, the analysis found 10 excess hits at a = 0.05 and 13 at a = 0.01.

6. REPORTING OF OUTCOME WIDE-ANALYSES

The paper recommends compact, standardized reporting that preserves information across many outcomes while making exposure timing, measurement, effect-size comparisons, and confounding sensitivity transparent.

  • Reporting structure: Outcome-wide reports should organize extensive information efficiently, including overall sample demographics and standardized outcome scales where appropriate.Standardization is especially recommended when outcome scales are unfamiliar, and population standard deviations should be reported.
  • Reporting structure: E-values should be reported for each outcome for both the effect estimate and its confidence interval.This supports assessment of sensitivity to potential unmeasured confounding across the outcome set.
  • Measurement and timing: Because the exposure is fixed across outcomes, its measurement details should be discussed in the main text alongside the timing of exposure, outcomes, and covariates.Some detailed outcome-measure information may be placed in an online supplement.
  • Effect-size reporting: For continuous exposures, primary analyses should use a meaningful nominal comparison, a per-standard-deviation scale, or exposure tertiles or a median split.The preferred option depends on whether the exposure scale is well understood and substantively interpretable.
  • Effect-size reporting: Approximate risk ratios can facilitate comparisons across continuous outcomes, but the paper recommends using them as supplementary rather than primary analyses.They can be obtained by meaningful dichotomization, median split, or approximate conversion from standardized effect sizes.

7. ADVANTAGES OF OUTCOME-WIDE LONGITUDINAL DESIGNS

Outcome-wide longitudinal analyses present effects of one exposure across many subsequent outcomes, increasing information, comparability, and efficiency while reducing some investigator-driven modeling choices.

  • 7.1 Conveys More Information: A single outcome-wide publication conveys effects of one exposure across a broad range of outcomes, reducing the need to search across separate studies.The authors argue this can improve efficiency for readers, researchers, editors, and peer reviewers.
  • 7.5 Policy Relevance: Presenting beneficial and harmful effects across outcomes at once may provide more useful evidence for public-health and policy recommendations.The authors highlight exposures that may have beneficial effects on some outcomes and harmful effects on others.
  • 7.2 Greater Potential to Report Null Results: Outcome-wide analyses may facilitate reporting null results, which can be as or more informative than findings suggesting an effect.Null outcomes might also serve as negative controls and provide evidence that positive findings are not solely due to unmeasured confounding.
  • 7.3 Less Temptation to Choose Models: Using common covariates and modeling across outcomes reduces opportunities to optimize separate results in line with investigator expectations.The approach does not eliminate selective analysis or reporting, and researchers may still run multiple outcome-wide analyses.
  • 7.4 The Comparison of Effect Sizes: Outcome-wide studies allow more direct comparison of an exposure’s effect sizes across outcomes within the same sample.Comparisons across separate studies can instead reflect differences between populations rather than differences between outcomes.

8. DESIGN VARIATIONS

The paper considers alternatives and extensions of outcome-wide designs, including exposure-wide, lagged exposure-wide, interaction, time-varying, instrumental-variable, and mediator-wide applications. These variations differ in feasibility and confounding risks.

  • 8.1 Challenges of Exposure-Wide Designs: Exposure-wide studies assess associations between one outcome and many exposures simultaneously, but may face substantial confounding and collider stratification biases.The authors emphasize that each exposure may require a separate confounder set, making bias assessment difficult as the number of exposures grows.
  • 8.2 Lagged Exposure-Wide Designs: A lagged exposure-wide design examines one later outcome while modeling contemporaneous exposures separately and controlling for earlier exposures and covariates.The proposed setup uses exposures at wave 2, an outcome at wave 3, and covariates at wave 1, with all confounding-control decisions held constant across regressions.
  • 8.2 Lagged Exposure-Wide Designs: Under the stated confounding-control assumption, each exposure coefficient provides a consistent causal-effect estimate on the regression model’s relevant scale.For linear regression, the coefficient corresponds to the contrast between potential outcomes under exposure levels 1 and 0.
  • 8.3 Interaction Outcome-Wide Studies: With two relatively contemporaneous exposures, outcome-wide regressions can report their main effects and interactions across multiple outcomes.The framework can also quantify proportions attributable to each exposure and their interaction, with comparative conversion to a difference scale.
  • 8.4 Time-Varying Exposures: Outcome-wide designs could in principle address exposure trajectories, but time-varying exposures require more complex confounding assumptions and causal models beyond this paper’s scope.The authors refer readers to specialized causal-inference treatments for these analyses.
  • 8.5 Quasi-Experimental Designs: Outcome-wide analyses may be plausible with instrumental variables when the instrument for treatment or exposure is itself randomized, depending on context.The paper gives treatment-compliance assignment and draft-lottery number as examples of potentially relevant settings.
  • 8.6 Mediation: A mediator-wide design can be biased when mediators affect one another but are evaluated singly, one at a time.The proposed variation fixes the exposure and outcome while examining numerous potential mediators individually.

9. CONCLUSION

The paper proposes the outcome-wide longitudinal design as a template for assessing causal effects across multiple outcomes. It presents the design as potentially preferable when numerous outcomes matter, while emphasizing the need for broad outcome data and future refinement.

  • 9. CONCLUSION: The outcome-wide longitudinal design extends causal-effect studies across outcomes and incorporates confounding-control, unmeasured-confounding, and multiple-testing considerations.The authors describe the approach as a template for empirical studies rather than a general theory of causal inference.
  • 9. CONCLUSION: The design may become preferable to single exposure-outcome studies when many outcomes are relevant and longitudinal data with confounding control are available.The authors retain a role for careful evaluation of individual exposure-outcome relationships.
  • 9. CONCLUSION: Using numerous outcomes can support more objective inference, reporting of null results, consistent evaluation of unmeasured confounding, and comparison of effect sizes.These advantages are presented as potential benefits of outcome-wide designs in relevant longitudinal or panel-data settings.
  • 9. CONCLUSION: Outcome-wide studies depend critically on having a broad range of outcomes available, including outcomes spanning human flourishing.The paper links such data collection to policy, public-health, and knowledge-advancement aims.

E-value for point

The supplied material identifies supplementary longitudinal analyses of parental warmth and health, well-being, and flourishing outcomes. These analyses use alternative outcome and exposure representations, with some procedures addressing missing data and multiple testing.

  • E-value for point: Complete-case analyses examined the longitudinal association between standardized parental warmth and mid-life health and well-being outcomes, with sample sizes ranging from 2,618 to 2,742.The analysis used the Midlife in the United States Study from 1994-1995 to 2004-2006.
  • E-value for point: A supplementary analysis examined parental warmth tertiles and mid-life health and well-being outcomes in a sample of 2,948 participants.The table title identifies this as a longitudinal analysis.
  • E-value for point: Another analysis converted continuous flourishing-outcome effect estimates to the risk-ratio scale for parental warmth tertiles in a sample of 2,948 participants.The analysis used the 2004-2006 questionnaire wave.
  • E-value for point: The reported significance notation distinguishes p<0.05 and p<0.01 before Bonferroni correction from p<0.05 after correction, using a cutoff of 0.002 for 24 outcomes.The correction threshold is defined as 0.05/24 outcomes.
  • E-value for point: A further analysis dichotomized continuous outcomes at the median split when examining parental warmth tertiles and flourishing in mid-life.This analysis also used a sample of 2,948 participants.
Loading 1810.10164v1…