Source-linked AI summary

Incorporating biological information into linear models: A Bayesian approach to the selection of pathways and genes

Francesco C. Stingo, Yian A. Chen, Mahlet G. Tadesse, Marina Vannucci

arXiv:1111.5419v1stat.AP

TL;DR

The paper addresses how to identify biologically relevant pathways and genes when gene-level selection alone may overlook pathway structure and gene dependence. It proposes a Bayesian microarray model using pathway summaries, network-informed priors, and structured MCMC moves, and reports improved pathway separation, fewer false positives, identification of otherwise missed predictors, and improved prediction accuracy. The approach is constrained when pathway databases are incomplete or gene-network information is unavailable.

  • Problem

    Gene selection alone may miss pathway-level biology, motivating methods that identify both relevant genes and pathways from microarray data.

  • Method

    A Bayesian model uses pathway memberships for PLS-based pathway summaries, gene networks for a Markov random field prior, and biological structure to guide MCMC moves.

  • Results

    The method improves separation of relevant from nonrelevant pathways, reduces false positives, can identify predictors otherwise missed, and improves prediction accuracy.

  • Takeaways & Limitations

    Integrating pathway and network knowledge can aid interpretation of selected genes and pathways while supporting prediction in gene-expression analyses.

  • Takeaways & Limitations

    The approach cannot use the dependence structure and MRF prior for all genes when pathway databases are incomplete or gene-network information is unavailable.

Abstract

from arXiv · show

The vast amount of biological knowledge accumulated over the years has allowed researchers to identify various biochemical interactions and define different families of pathways. There is an increased interest in identifying pathways and pathway elements involved in particular biological processes. Drug discovery efforts, for example, are focused on identifying biomarkers as well as pathways related to a disease. We propose a Bayesian model that addresses this question by incorporating information on pathways and gene networks in the analysis of DNA microarray data. Such information is used to define pathway summaries, specify prior distributions, and structure the MCMC moves to fit the model. We illustrate the method with an application to gene expression data with censored survival outcomes. In addition to identifying markers that would have been missed otherwise and improving prediction accuracy, the integration of existing biological knowledge into the analysis provides a better understanding of underlying molecular processes.

1. Introduction.

The paper argues that selecting genes alone can miss pathway-level biology, so it proposes a Bayesian model integrating pathway memberships and gene networks to select both pathways and genes. The approach uses pathway summaries, structured priors, and MCMC moves, with simulations indicating improved separation and fewer false positives.

  • Motivation: Gene selection alone may be insufficient because disease-related processes involve interacting genes, pathways, and alternative signaling routes.The introduction notes that blocking one pathway or a few pathways may not prevent signaling through downstream genes, alternative branches, or different pathways.
  • Contribution: The proposed Bayesian model incorporates pathway memberships and gene networks into DNA microarray analysis to identify relevant pathways and genes.It uses biological information to define pathway summaries, specify priors, and structure MCMC moves.
  • Related work: The model is designed to select both pathways and genes, whereas earlier approaches often targeted either important pathways or important genes.The paper positions its contribution as a more comprehensive use of pathway-gene relationships and gene dependence structures.
  • Method: The method combines latent indicators, a Markov random field prior for gene dependence, and pathway expression measures based on PLS components.Pathway summaries are derived from the first latent components of PLS regressions on selected genes within each pathway.
  • Results: The simulation studies report better separation between relevant and nonrelevant pathways and fewer false positives when using the MRF prior.The paper evaluates the approach on simulated data and applies it to gene-expression data with survival outcomes.

2. Model specification.

The model jointly selects phenotype-related pathways and genes by combining pathway summaries from selected genes with pathway and gene-network information in Bayesian priors. It supports continuous, categorical, and censored survival outcomes.

  • Data and model structure: The model uses pathway membership and gene-network matrices to represent biological relationships among genes and pathways.S records gene membership in pathways, while R records direct gene-network links; both are constructed from pathway databases.
  • Data and model structure: Binary indicators θ and γ select pathways and genes, respectively, allowing inference to identify both relevant groups and genes within selected groups.A gene is selected when γj = 1, and pathway scores use selected genes belonging to selected pathways.
  • Pathway activity summaries: Each pathway is summarized by a first latent PLS component computed from the expression levels of its selected genes.The component is formed from Xk(γ) using the eigenvector associated with the largest eigenvalue of the covariance-based PLS criterion.
  • Outcome models: The framework extends from continuous responses to categorical outcomes through probit models and to censored survival outcomes through an accelerated failure time model with latent variables.For censored survival, observed times and censoring indicators constrain latent Zi, with Zi = log(Yi) for uncensored observations and Zi > log(Yi) for censored observations.
  • Prior specification: A Markov random field prior uses gene-network dependence so that larger η favors selecting genes whose neighbors are already selected.The prior also includes a sparsity parameter µ; for isolated genes it reduces to an independent Bernoulli distribution.
  • Prior specification: A hyperprior on η restricts attention to positive dependence and is scaled using the detected phase-transition value ηPT.The restriction avoids negative dependence and is intended to minimize the phase-transition problem in which selected-model size increases sharply.

3. Model fitting.

Model fitting uses a marginalized MCMC scheme that updates pathway and gene inclusion indicators, the network-smoothing parameter, and outcome-specific latent variables. Posterior inference then estimates pathway and gene inclusion from visited models, while prediction constructs pathway scores without future outcomes.

  • MCMC sampling: The MCMC algorithm updates pathway and gene indicators with Metropolis–Hastings moves that add or delete genes and pathways.Moves can change both indicators together, only gene inclusion, or only pathway inclusion.
  • MCMC sampling: The sampling scheme separately updates the MRF parameter η and outcome-specific latent variables for classification or survival models.The MCMC steps sample inclusion indicators, η, and additional latent variables introduced by probit or accelerated failure time models.
  • MCMC sampling: An auxiliary-variable proposal handles the unavailable MRF normalizing-constant ratio when updating η.The auxiliary variable w shares γ’s state space, and its proposal uses the proposed η to construct the Metropolis–Hastings update.
  • Posterior inference: Posterior pathway selection uses marginal inclusion probabilities, while gene selection can condition on representation of at least one pathway containing the gene.The procedure obtains these probabilities from the relative frequencies of inclusion in visited models; Rao–Blackwellized estimates are not straightforward here and may be computationally expensive.
  • Prediction: For new observations, prediction uses least-squares estimates with pathway scores generated by PCA because future outcomes are unavailable for fitting PLS.For categorical or censored survival outcomes, latent predictions are converted to observed-outcome predictions using the corresponding outcome relationship.

4. Application.

The model was assessed on simulated pathway-structured data and applied to breast-cancer microarray data, using KEGG information and MCMC posterior probabilities to select pathways and genes. Simulations supported pathway and gene recovery, while the application showed improved predictive error and biologically interpretable selected networks.

  • Simulation studies: Simulations used 70 KEGG pathways, 1,098 genes, four selected pathways, 15 relevant genes, and 100 samples across three coefficient settings.The settings were β = ±0.5, β = ±1, and β = ±1.5.
  • Simulation studies: All four relevant pathways were selected at a marginal posterior probability cutoff of 0.8 in all three simulated scenarios.Lowering the cutoff to 0.5 introduced two false-positive pathways in each scenario.
  • Simulation studies: At β = ±1.5, all 15 relevant genes were selected without false positives using a marginal posterior probability threshold of 0.12.At lower signal levels, fewer relevant genes were recovered or more false positives were included.
  • Simulation studies: With the MRF prior, the average separation between the least-probable relevant pathway and most-probable nonrelevant pathway was 0.28, versus 0.18 without it.The authors also report increased sensitivity for selecting true variables with the MRF prior.
  • Application to microarray data: The microarray analysis included 76 lymph-node-negative breast-cancer patients, 196 pathways, and 3,592 probes mapped through KEGG gene networks.Missing expressions were imputed using 10-nearest-neighbor averaging.
  • Application to microarray data: The pathway-informed model achieved validation MSE 1.4497 using 12 pathways and 41 probe sets, compared with MSE 1.9317 from a gene-only analysis.Using 7 pathways and 14 genes, the model achieved MSE 1.7614 for a more size-matched comparison.
  • Application to microarray data: Posterior probabilities prioritized pathways and genes for follow-up, while MRF islands helped interpret connected branches of selected pathways.Selected pathways included MAPK signaling and KEGG pathways in cancers, and selected genes included PKCA.

5. Discussion.

The proposed Bayesian model integrates pathway membership and gene-network dependence to select pathways and genes, using biological information in summaries, priors, and MCMC moves. Simulations show that the MRF prior can improve pathway separation and gene selection, while incomplete databases limit its applicability.

  • Method: Biological information defines pathway summaries, dependence-aware priors, and the MCMC moves used to fit the model.The MRF prior captures gene dependence, while pathway information structures the sampling scheme.
  • Simulation findings: The MRF prior improved separation between relevant and nonrelevant pathways in simulations.The reported separation was measured by the difference between the lowest posterior probability among relevant pathways and the highest among nonrelevant pathways.
  • Simulation findings: In a small-coefficient simulation, the MRF model selected all correct genes without false positives, whereas the model without MRF included 3 false positives.The effect of the prior depended on the concordance between the prior network and the data.
  • Limitations: Pathway databases are incomplete, and gene-network information is unavailable for many genes.When networks are unavailable, the proposed MRF dependence structure cannot be used for all genes.
  • Method: The model uses PLS pathway summaries, Bayesian pathway selection, and gene selection within selected pathways.It incorporates biological information into both model construction and variable selection.
  • Comparison: The discussion contrasts this pathway-aware Bayesian approach with group lasso methods that select entire groups or none of their variables.The authors note that genes within pathways may not be highly correlated, while genes across pathways can be correlated or belong to multiple pathways.

SUPPLEMENTARY MATERIAL

The supplementary material documents the MCMC implementation and theoretical discussion supporting the model's sampling procedure.

  • Supplementary methods: The supplement describes the MCMC steps for updating (θ,γ).
  • Supplementary methods: The supplement discusses ergodicity of the Markov chain on the restricted space.
  • Supplementary methods: The supplementary material focuses on implementation and theoretical properties of the parameter-update chain.
Loading 1111.5419v1…