Source-linked AI summary

Pleiotropy robust methods for multivariable Mendelian randomization

Andrew J. Grant, Stephen Burgess

arXiv:2008.11997v1stat.ME

TL;DR

Pleiotropy threatens valid causal estimation in Mendelian randomization, and few robust methods exist for multivariable analyses. The paper introduces three summary-data methods and evaluates them across pleiotropy settings. MVMR-Robust performs best with few invalid instruments, MVMR-Lasso has the best overall mean squared error, and MVMR-Median performs nearly as well as MVMR-Lasso.

  • Problem

    Pleiotropy can bias Mendelian-randomization estimates, while relatively few pleiotropy-robust methods are available for multivariable analyses.

  • Method

    The paper presents three pleiotropy-robust multivariable Mendelian-randomization methods for summary-level data and evaluates them in simulations and an applied analysis.

  • Results

    MVMR-Robust outperformed MVMR-PRESSO with relatively few invalid instruments, while MVMR-Lasso performed best overall in mean squared error, including when half the variants were invalid.

  • Takeaways & Limitations

    The methods provide complementary options for multivariable Mendelian randomization across different pleiotropy settings, alongside MVMR-Egger.

  • Takeaways & Limitations

    Regularization techniques generally cannot compute accurate standard errors, although inflated type I error for MVMR-Lasso can be mitigated with a three-sample approach.

Abstract

from arXiv · show

Mendelian randomization is a powerful tool for inferring the presence, or otherwise, of causal effects from observational data. However, the nature of genetic variants is such that pleiotropy remains a barrier to valid causal effect estimation. There are many options in the literature for pleiotropy robust methods when studying the effects of a single risk factor on an outcome. However, there are few pleiotropy robust methods in the multivariable setting, that is, when there are multiple risk factors of interest. In this paper we introduce three methods which build on common approaches in the univariable setting: MVMR-Robust; MVMR-Median; and MVMR-Lasso. We discuss the properties of each of these methods and examine their performance in comparison to existing approaches in a simulation study. MVMR-Robust is shown to outperform existing outlier robust approaches when there are low levels of pleiotropy. MVMR-Lasso provides the best estimation in terms of mean squared error for moderate to high levels of pleiotropy, and can provide valid inference in a three sample setting. MVMR-Median performs well in terms of estimation across all scenarios considered, and provides valid inference up to a moderate level of pleiotropy. We demonstrate the methods in an applied example looking at the effects of intelligence, education and household income on the risk of Alzheimer's disease.

1 Introduction

Mendelian randomization uses genetic variants to estimate causal effects, but pleiotropy can invalidate instruments and bias estimates. This paper addresses the limited availability of pleiotropy-robust methods for multivariable analyses by proposing and evaluating new approaches.

  • Mendelian randomization uses genetic variants as instrumental variables to estimate causal effects from observational data while addressing unmeasured confounding.
  • Pleiotropy occurs when variants affect traits beyond the risk factor, creating alternative pathways that can invalidate instruments and bias causal estimates.
  • Multivariable Mendelian randomization fits multiple risk factors jointly, helping account for measured pleiotropy and distinguish direct effects from total effects involving mediators.
  • Few pleiotropy-robust methods exist for multivariable Mendelian randomization, whereas many methods address the single-risk-factor setting under different assumptions.
  • Valid multivariable causal estimation typically requires that all genetic-variant pathways to the outcome are accounted for through measured risk factors.
  • The paper proposes new multivariable approaches robust to different pleiotropy forms, evaluates them in simulations, and demonstrates them using Alzheimer’s disease data.

2 Modelling assumptions

The model represents multiple risk factors, genetic variants, confounders, and outcome relationships while allowing correlated risk factors and pleiotropic direct effects. It defines balanced, directional, and outlier pleiotropy and describes assumptions underlying robust estimation.

  • Data-generating model: The data-generating model includes K risk factors, p genetic variants, confounders, and an outcome, with genetic effects represented by a p × K matrix.
  • Data-generating model: Risk factors may be correlated through correlated error terms and their shared association with confounders, even when genetic variants are mutually independent.
  • Instrument validity and pleiotropy: A genetic variant is valid when it predicts at least one risk factor, has no confounder association, and has no direct or unmeasured pathway to the outcome.
  • Pleiotropy patterns: Balanced pleiotropy has direct effects with mean zero, whereas directional pleiotropy has direct effects with a nonzero mean.
  • Pleiotropy patterns: Outliers are a relatively small number of nonzero pleiotropic effects that may be large in magnitude.
  • Trait associations and orientation: Categorical traits can be analyzed through logistic or ordinal logistic regression, with associations interpreted as changes in log odds per additional effect allele.
  • Trait associations and orientation: Genetic-variant orientation can affect methods that model pleiotropic effects, so analyses may be repeated under alternative orientations when no primary risk factor exists.
  • The InSIDE assumption: Pleiotropy-robust methods often require InSIDE, meaning instrument strength is independent of direct effects; its finite-sample version will rarely hold exactly because of random variation.

3 Methods

The paper reviews multivariable inverse-variance weighting, Egger regression, outlier-robust procedures, robust regression, median-based estimation, and regularization approaches for pleiotropy.

  • Existing methods: MVMR-IVW uses weighted least squares and is consistent with valid instruments, or with balanced pleiotropy when the InSIDE assumption holds.
  • Existing methods: MVMR-IVW is sensitive to outliers and directional pleiotropy.
  • Existing methods: MVMR-Egger adds an intercept to account for pleiotropy but relies on InSIDE, has lower precision, and can depend on variant orientation.
  • Existing methods: The outlier procedure repeatedly compares observed and simulated residual sums of squares, then applies Bonferroni-corrected empirical p-values to identify variants for removal.
  • Robust regression: Robust regression caps sufficiently large residuals and is expected to improve efficiency with few invalid instruments, but may perform poorly when many instruments are invalid.
  • Median-based estimation: MVMR-Median extends least absolute deviations regression to multiple risk factors, offering reduced sensitivity to outliers and robustness to directional pleiotropy.

4 Simulations

The simulations compare multivariable Mendelian randomization methods across balanced and directional pleiotropy, varying invalid-variant proportions and causal-effect settings. Results show distinct trade-offs among robust, lasso, median, IVW, Egger, and PRESSO approaches, with patterns broadly retained under fewer instruments, correlated risk factors, and one-sample designs.

  • Simulation design: The study simulated 1,000 replications across three pleiotropy scenarios, 10%, 30%, or 50% invalid variants, and two sets of causal effects.The first set included θ1 = 0.2; the second set had θ1 = 0, with the first risk factor’s causal effect reported.
  • Comparative performance: All methods performed well for bias under balanced pleiotropy, whereas MVMR-IVW and MVMR-Egger became increasingly biased as directional pleiotropy increased.MVMR-IVW and MVMR-Egger were also less precise and very low powered in these settings.
  • Comparative performance: MVMR-Robust outperformed MVMR-PRESSO in every considered scenario, with lower bias, greater precision, and correct type I error rates.MVMR-PRESSO had low bias at low pleiotropy but performed poorly at moderate or high levels.
  • Comparative performance: MVMR-Lasso was generally the most precise, matching MVMR-Robust in mean squared error at 10% pleiotropy while retaining performance at higher levels.It had inflated type I error rates, despite comparable power to MVMR-Robust at 10% pleiotropy and better power as pleiotropy increased.
  • Comparative performance: MVMR-Median had bias comparable to MVMR-Lasso across pleiotropy levels, but was less precise and lower powered while maintaining type I error rates closer to significance level.MVMR-Median was uniformly outperformed across scenarios only by MVMR-Lasso in mean squared error.
  • Sensitivity analyses: With fewer instruments, mean squared error increased slightly across methods, while comparative performance remained similar under fewer instruments, correlated risk factors, and one-sample Mendelian randomization.The parametric bootstrap, which ignores risk-factor correlation, still performed well.

5 Applied example: the causal effect of intelligence, education and household income on Alzheimer’s disease

The applied analysis examines intelligence, education, and household income as multivariable risk factors for Alzheimer’s disease using 213 shared genetic instruments. Across methods, intelligence retained evidence of a protective effect, while education and income effects attenuated toward the null after multivariable adjustment.

  • Data and instruments: The analysis used 213 genetic variants available in both the household-income and Alzheimer’s disease datasets as instruments.Associations for intelligence and education came from separate GWAS, while Alzheimer’s disease associations came from Lambert et al.
  • Data and instruments: Genetic-variant associations with years of education and household income appeared reasonably strongly correlated, while correlations involving intelligence were low to moderate.Figure 3 showed little evidence of directional pleiotropy but suggested possible outliers.
  • Univariable results: Univariable analyses estimated protective effects for intelligence, education, and household income on Alzheimer’s disease.Estimated log causal odds ratios were −4.20 for intelligence, −0.59 for education, and −0.60 for household income, with the reported confidence intervals excluding zero.
  • Multivariable results: In the MVMR-IVW model, education and household income estimates attenuated to the null, whereas intelligence remained protective with an estimated odds ratio of −0.47 (95% CI −0.86 to −0.07).The multivariable model therefore retained evidence for an intelligence effect independent of education and household income.
  • Robustness analysis: Pleiotropy-robust methods broadly agreed with MVMR-IVW, although MVMR-Egger suggested a null intelligence effect and MVMR-Lasso identified 15 pleiotropic variants for removal.The MR-PRESSO outlier test detected no outliers.
  • Interpretation: The consistency of findings supports the assertion that intelligence has a causally protective effect on Alzheimer’s disease independent of education and household income.Education and income may influence Alzheimer’s disease through intelligence, pleiotropic pathways, or both.

6 Discussion

The paper presents pleiotropy-robust methods for multivariable Mendelian randomization and positions them as complementary sensitivity analyses. Their performance depends on the extent and form of pleiotropy, while modelling and measurement-error assumptions constrain interpretation.

  • Simulation findings: When relatively few instruments are invalid, MVMR-Robust outperformed MVMR-PRESSO in all scenarios considered.This advantage was reported under low levels of invalidity.
  • Simulation findings: MVMR-Lasso performed best overall by mean squared error, including when half the genetic variants were invalid and pleiotropy was directional.Its type I error rates were inflated, but this could be mitigated when a three-sample approach was possible.
  • Simulation findings: MVMR-Median performed almost as well as MVMR-Lasso by mean squared error and retained correct type I error rates at higher levels of pleiotropy.The methods were also demonstrated in an analysis of intelligence, education, household income and Alzheimer’s disease.
  • Applied analysis: Residuals-versus-fitted values from MVMR-IVW can visualise heterogeneity, potential outliers and pleiotropy, helping determine an appropriate robust method.The applied example showed little evidence of directional pleiotropy but possible outliers.
  • Limitations: Interpretation is constrained by assumptions of linearity, homogeneous effects, no measurement error in genetic variant–risk-factor associations, and sufficiently strong conditional instruments.For non-continuous traits, logistic-regression summary statistics may introduce bias through odds-ratio noncollapsibility; weak instruments may cause weak-instrument bias.
  • Methods and contribution: The methods extend pleiotropy-robust Mendelian randomization to multiple risk factors and provide a suite of sensitivity analyses alongside MVMR-Egger.They are designed to handle different forms of pleiotropy using summary-level data.

7 Software

The paper provides R code for its methods and simulation results, alongside existing R packages used to implement related methods.

  • Software: R code is available for performing the paper’s methods and reproducing its simulation results.The code is hosted at the stated GitHub repository.
  • Software: MendelianRandomization and MR-PRESSO are among the existing R packages used to implement the various methods.
  • Software: robustbase, quantreg, and glmnet are also used to implement the various methods.

Funding

The authors report fellowship support from the Wellcome Trust and Royal Society, along with funding from the NIHR Cambridge Biomedical Research Centre.

  • Funding: Andrew J. Grant and Stephen Burgess are supported by a Sir Henry Dale Fellowship jointly funded by the Wellcome Trust and the Royal Society.The grant number is 204623/Z/16/Z.
  • Funding: The research was funded by the NIHR Cambridge Biomedical Research Centre.
  • Funding: The stated views belong to the authors and are not necessarily those of the NHS, NIHR, or Department of Health and Social Care.

S.1 Algorithm for performing the regularization approach

The regularization approach represents genetic associations and their uncertainty in matrices, projects onto a weighted design space, and solves the penalized problem for a chosen λ.

  • Definitions: β̂X is defined as a p × K matrix of genetic variant–risk factor associations, while β̂Y is a length-p outcome-association vector.
  • Definitions: S is a p × p diagonal matrix whose diagonal entries are σ^-2, and θ0 is represented as a length-p vector.
  • Algorithm: The procedure projects onto the column space of S^1/2β̂X and solves the regularized problem for a given λ.
  • Algorithm: The algorithm uses S^1/2 in its computational procedure.
  • Algorithm: The kth element of θ̂λ estimates θk for the selected λ, while Step 1 is a standard lasso.

S.2 Supplementary simulation results

The supplementary simulations examine additional settings involving more risk factors, correlated risk factors, and one-sample estimation, using tables and figures that report estimation and inference performance.

  • Overview: The supplementary simulation section reports results from the additional studies described in Section 4.
  • Additional risk factors: When p = 20, Tables S1–S2 and Figure S2 examine βXjk drawn from Uniform (0, 0.22).
  • Correlated risk factors: For correlated risk factors, Tables S3–S4 and Figure S3 use cor(vXik, vXil) = 0.5 for all k ≠ l.
  • One-sample setting: For one-sample Mendelian randomization, Tables S5–S6 and Figure S4 examine associations estimated from the same sample.
  • Performance measures: Figures S1–S4 report logarithms of mean squared errors across scenarios S1–S3 and invalid-variant proportions of 10, 30, or 50%.Figure S1 focuses on causal-effect estimates for θ1 and θ4; Figures S2–S4 vary the simulation setting.
Loading 2008.11997v1…