Source-linked AI summary

CoMEt: A Statistical Approach to Identify Combinations of Mutually Exclusive Alterations in Cancer

Mark D. M. Leiserson, Hsin-Ta Wu, Fabio Vandin, Benjamin J. Raphael

arXiv:1503.08224v1q-bio.QMstat.CO

TL;DR

Cancer genomics needs to identify driver alterations despite heterogeneous tumors and incomplete knowledge of cancer pathways. CoMEt performs de novo mutual-exclusivity analysis using exact frequency-conditional statistics and ensemble sampling of multiple alteration sets. It outperforms earlier approaches and finds pathway-associated and potentially novel cancer genes across four TCGA cancer types.

  • Problem

    Cancer genomics must distinguish driver from passenger alterations even though tumors of the same type contain heterogeneous driver combinations and known pathways are incomplete.

  • Method

    CoMEt uses an exact mutual-exclusivity test conditioned on alteration frequencies and MCMC sampling to identify collections of multiple, including overlapping, alteration sets while analyzing subtypes.

  • Results

    CoMEt outperforms earlier de novo methods on simulated and real data and identifies mutually exclusive collections overlapping known pathways while implicating novel cancer genes across multiple TCGA cancer types.

  • Takeaways & Limitations

    CoMEt provides testable hypotheses about pathway alterations, novel cancer genes, and relationships between subtype- and pathway-related mutual exclusivity.

  • Takeaways & Limitations

    Computing the exact score requires enumerating fixed-margin contingency tables, a problem that is generally #P-complete and intractable even for small alteration sets.

Abstract

from arXiv · show

Cancer is a heterogeneous disease with different combinations of genetic and epigenetic alterations driving the development of cancer in different individuals. While these alterations are believed to converge on genes in key cellular signaling and regulatory pathways, our knowledge of these pathways remains incomplete, making it difficult to identify driver alterations by their recurrence across genes or known pathways. We introduce Combinations of Mutually Exclusive Alterations (CoMEt), an algorithm to identify combinations of alterations de novo, without any prior biological knowledge (e.g. pathways or protein interactions). CoMEt searches for combinations of mutations that exhibit mutual exclusivity, a pattern expected for mutations in pathways. CoMEt has several important feature that distinguish it from existing approaches to analyze mutual exclusivity among alterations. These include: an exact statistical test for mutual exclusivity that is more sensitive in detecting combinations containing rare alterations; simultaneous identification of collections of one or more combinations of mutually exclusive alterations; simultaneous analysis of subtype-specific mutations; and summarization over an ensemble of collections of mutually exclusive alterations. These features enable CoMEt to robustly identify alterations affecting multiple pathways, or hallmarks of cancer. We show that CoMEt outperforms existing approaches on simulated and real data. Application of CoMEt to hundreds of samples from four different cancer types from TCGA reveals multiple mutually exclusive sets within each cancer type. Many of these overlap known pathways, but others reveal novel putative cancer genes. *Equal contribution.

1 Introduction

Cancer genomics must identify driver alterations despite heterogeneous combinations of mutations and incomplete pathway knowledge. CoMEt addresses this challenge de novo by statistically modeling mutual exclusivity and identifies multiple pathway- and subtype-related alteration sets in TCGA data.

  • Motivation: Mutational heterogeneity makes it difficult to distinguish driver alterations from passengers because tumors of the same type carry different driver combinations.These combinations can arise when alterations perturb different genes within the same biological pathway.
  • Motivation: De novo analysis is valuable because pathway databases and interaction networks are incomplete, lack tissue specificity, and may not represent a particular cancer cell.Exhaustive testing of all possible combinations is nevertheless computationally infeasible.
  • Prior limitations: Existing mutual-exclusivity methods can be biased toward high-frequency genes and often cannot jointly identify overlapping sets or account for subtype-specific mutations.These limitations can obscure lower-frequency alterations and confound exclusivity signals.
  • CoMEt contributions: CoMEt uses an exact frequency-conditional exclusivity test, jointly identifies multiple and overlapping alteration sets, and analyzes cancer subtypes simultaneously.Its statistical procedure is designed to improve sensitivity to rare alterations while avoiding spurious subtype-specific sets.
  • Results: CoMEt outperforms earlier approaches on simulated and real data and identifies pathway-overlapping combinations plus potentially novel genes across four TCGA cancer types.Reported candidates include IL7R and EPHB3 in gastric cancer and SRCRB4D in glioblastoma.

2 Methods

CoMEt represents diverse alterations in a binary matrix, scores mutual exclusivity conditional on observed alteration frequencies, and samples collections of alteration sets when exhaustive enumeration is infeasible. It summarizes significant sampled collections with a marginal probability graph and uses exact or approximate procedures to evaluate exclusivity.

  • Data representation: CoMEt converts alteration measurements into an m × n binary alteration matrix representing presence or absence across samples.Alterations may include gene mutations, specific nucleotide mutations, epigenetic changes, or other binary events.
  • Mutual-exclusivity scoring: The score Φ(M) is an exact mutual-exclusivity P-value conditional on each alteration’s observed frequency.It uses a 2^k contingency table with fixed margins, so the test targets surprisingly high exclusivity rather than general dependence.
  • Collection scoring: CoMEt combines set scores into collection scores under an independence-across-sets null hypothesis and uses a mid P-value to reduce conservativeness from discrete exact distributions.The collection score is formed by multiplying the scores of its alteration sets.
  • Collection sampling and summarization: MCMC samples collections in proportion to Φ(M)^−1, and a marginal probability graph summarizes pairwise support across significant collections.Edges below threshold δ are removed, and connected components become output modules that can differ in number and size from the input parameters.
  • Mutual-exclusivity scoring: The test statistic T(XM) sums samples containing exactly one alteration, distinguishing exclusivity from other forms of non-independence.For k = 2, this is equivalent to a one-sided Fisher’s exact test; for larger sets, it specifically measures exclusivity.

2.4 Sampling collections of mutually exclusive alterations with MCMC

CoMEt uses Metropolis–Hastings MCMC to sample collections of alteration sets because the collection space is generally too large for exhaustive enumeration.

  • 2.4 Sampling collections of mutually exclusive alterations with MCMC: CoMEt uses Metropolis–Hastings MCMC to sample collections of t alteration sets in proportion to Φ(M)^−1.The method targets low-score, highly significant collections when enumerating all possible collections is infeasible.

2.5 Marginal probability graph

CoMEt summarizes MCMC-sampled collections of mutually exclusive alteration sets with a marginal probability graph. This graph supports extraction of overlapping or otherwise connected alteration sets while accounting for subtype-specific confounding and collection-level significance.

  • Marginal probability graph: CoMEt computes pairwise posterior probabilities that alterations belong to the same set from MCMC samples, then uses them to construct a marginal probability graph.The graph summarizes the posterior distribution over collections rather than selecting only one collection.
  • Marginal probability graph: Thresholding graph edges at δ and taking connected components yields highly exclusive alteration sets, including overlapping pathways linked by cut nodes.Connected components are preferred over cliques because they can represent overlapping alteration sets.
  • Statistical significance: Collection significance is evaluated against permuted alteration matrices that preserve sample and alteration frequencies.This controls for the large number of possible collections when assessing unusually high scores.
  • Subtype confounding: Subtype-specific alterations can appear mutually exclusive in mixed samples even when they are not in the same biological pathway.The paper addresses this confounding by analyzing subtype information through an additional CoMEt run and combining the resulting collection frequencies.
  • Comparison with MEMo: CoMEt differs from MEMo because MEMo tests coverage, whereas CoMEt tests exclusivity directly; for three or more genes, maximum coverage need not equal maximum exclusivity.Coverage-based testing can therefore deflate or inflate significance relative to the exclusivity statistic.

3 Results

CoMEt was evaluated against existing methods on simulated and real cancer datasets. It ranked implanted pathways more effectively, identified collections more accurately in simulations, and recovered significant, biologically relevant alteration modules in TCGA data.

  • Simulated individual gene sets: CoMEt ranked the implanted pathway first for every simulated coverage γP ≥0.3, while Dendrix and muex did so only for γP ≥0.7 and γP ≥0.9, respectively.At γP = 0.1 and 0.2, CoMEt was over an order of magnitude better than the other approaches.
  • Simulated individual gene sets: CoMEt ran in under 4 minutes, compared with under 1 minute for Dendrix and often an hour or more for muex across simulated datasets.The comparison covered average runtimes of each weight function across all gene sets.
  • Simulated collections: CoMEt outperformed Multi-Dendrix in 11/12 simulated datasets and achieved ARI > 0.5 on all 12 datasets.The ARI difference exceeded 0.2 for 8/12 datasets, and CoMEt achieved ARI > 0.8 for 7/12 datasets.
  • Simulated collections: CoMEt remained superior to Multi-Dendrix in 11/12 datasets when both methods used the true pathway numbers and sizes.The authors attribute this result to CoMEt’s statistical score and MCMC sampling.
  • Real cancer data: On GBM datasets, CoMEt identified more significant exclusive sets and more genes in the COSMIC Cancer Census than muex.On the Multi-Dendrix GBM dataset, its collections overlapped the Rb, p53, and PI(3)K pathways, with Φ(M) ranging from 10^-8 to 10^-19.
  • Real cancer data: CoMEt produced almost identical GBM and BRCA results with or without the MutSigCV filter, unlike Multi-Dendrix.Without filtering, Multi-Dendrix included highly altered genes such as TTN and MUC16.
  • TCGA AML: In TCGA AML, CoMEt identified four modules spanning driver alterations, receptor tyrosine kinase and RAS signaling, chromatin regulation, and DNA methylation.The modules included known AML driver alterations and recovered previously reported TET2–IDH2 mutual exclusivity with a proposed mechanistic explanation.

(a) TCGA GBM

CoMEt identifies mutually exclusive alteration modules across TCGA gastric and breast cancer subtype analyses, recovering pathway-related and subtype-associated patterns. The results also include candidate genes whose roles remain uncertain or potentially novel.

  • TCGA STAD: CoMEt identified five mutually exclusive modules in TCGA STAD, including known cancer genes and novel candidate genes.Two modules indicate subtype-specific altered genes and pathways.
  • TCGA STAD: The second STAD module recapitulated mutual exclusivity between CDH1 mutations and ARHGAP6-CLDN18 fusions enriched in the genomically stable subtype.It also included PCDHA11 and EPHB3-region amplification.
  • TCGA STAD: The fourth STAD module contained CCNE1 amplification, SMAD4 mutations, and MET splice-site mutations in 115/217 samples.These genes participate in cell-cycle, TGF-β, and RTK/RAS signaling, respectively.
  • TCGA BRCA: CoMEt identified three subtype-specific modules and three mutated-gene modules in TCGA BRCA.CCND1 amplification was associated with luminal B, and ERBB2 amplification with the HER2-enriched subtype.
  • TCGA BRCA: A BRCA module combined subtype-associated alterations with pathway-related mutual exclusivity, including CDH1, AKT1, and PIK3CA in luminal A.Another set included TP53 and chromosome-region 4q13.3 amplification associated with basal-like tumors.

4 Discussion

The discussion presents CoMEt as a de novo method for analyzing diverse alteration types and collections of mutually exclusive events. Across TCGA datasets, it recovers pathway overlaps, subtype relationships, and candidate cancer genes that support further investigation.

  • 4 Discussion: CoMEt identifies collections of mutually exclusive alterations de novo without prior biological knowledge.Its score conditions on alteration frequencies to detect exclusivity involving rare mutations.
  • 4 Discussion: CoMEt was reported to outperform Dendrix, muex, and Multi-Dendrix on simulated and real data.The comparison is stated at the level of earlier de novo methods.
  • 4 Discussion: Across multiple TCGA cancer types, CoMEt identified significantly exclusive collections overlapping known pathways and implicating novel cancer genes.The results also distinguish subtype-driven exclusivity from pathway- or interaction-related exclusivity.
  • 4 Discussion: CoMEt supports binary data covering point mutations, indels, copy-number aberrations, rearrangements, splice-site mutations, fusions, and subtype annotations.The authors also suggest possible use with germline variants.
  • 4 Discussion: The authors propose pan-cancer analysis as an application because CoMEt can analyze cancer-type-specific and other exclusive alterations simultaneously.They also suggest broader uses for the tail-enumeration strategy underlying the exact statistics.

Exclusive Alterations in Cancer

The supplied section identifies the paper’s authors and their institutional affiliations. Mark D.M. Leiserson, Hsin-Ta Wu, Fabio Vandin, and Benjamin J. Raphael are listed as authors.

  • Exclusive Alterations in Cancer: Mark D.M. Leiserson, Hsin-Ta Wu, Fabio Vandin, and Benjamin J. Raphael are listed as authors.Leiserson and Wu are marked as equal contributors.
  • Exclusive Alterations in Cancer: The authors are affiliated with Brown University’s Department of Computer Science and Center for Computational Molecular Biology.The listed affiliations also include the University of Southern Denmark’s Department of Mathematics and Computer Science.
  • Exclusive Alterations in Cancer: Fabio Vandin is additionally affiliated with the University of Southern Denmark.

S1 Methods

CoMEt uses an MCMC procedure to sample collections of mutually exclusive gene sets, assess convergence across chains, and summarize the resulting marginal probability graph. Practical settings balance biologically meaningful collection sizes against runtime and convergence.

  • MCMC sampling: The MCMC state space consists of possible collections M, with transitions designed so the chain is ergodic and converges to a unique stationary distribution.The target distribution weights collections in proportion to Φ(M)^-1.
  • MCMC sampling: More exclusive collections receive higher sampling weights because CoMEt uses Φ(M)^-1.This weighting is implemented through a Metropolis-Hastings procedure.
  • MCMC sampling: Each iteration proposes a collection by replacing or swapping a randomly selected gene, depending on whether that gene is already present.The initialization selects tk genes and assigns k genes to each of t sets.
  • Convergence: CoMEt assesses convergence by comparing sampling distributions from multiple chains with different initial gene sets.The pipeline uses total variation distance and increases iterations when convergence metrics remain away from zero.
  • Marginal probability graph: The marginal probability threshold δ is selected by locating an “L-corner” in the edge-count versus edge-weight distribution.The procedure uses a minimum edge-weight threshold and piecewise linear fits before and after candidate values.
  • Experimental settings: Experiments generally used k = 4, t = 4, 100 million iterations, and 5 to 10 random initializations.Subtype analyses set t to the number of predefined subtypes; larger k and t can increase runtime and slow convergence.

S2 Data

The supplementary data section describes cancer datasets from TCGA and the simulation framework used to evaluate CoMEt.

  • Cancer datasets: AML data include whole-exome and copy-number measurements from 200 TCGA patients, organized into nine expert-defined gene categories.The categories include spliceosome, cohesin complex, MLL-X fusions, transcription factors, epigenetic modifiers, kinases, phosphatases, and RAS proteins.
  • Cancer datasets: GBM analyses used three datasets spanning 236–291 samples, with alterations drawn from mutations, copy-number aberrations, and expression-supported events.The datasets included Multi-Dendrix, muex, and pan-cancer GBM collections with different gene and alteration filters.
  • Cancer datasets: The STAD dataset was filtered for hypermutators and mutation-frequency criteria, yielding 217 gastric-cancer patients.The analysis combined mutations with focal copy-number aberrations, fusions, rearrangements, and splicing events.
  • Cancer datasets: The BRCA dataset contains whole-exome and copy-number data from 507 TCGA patients across basal-like, HER2-enriched, luminal A, and luminal B subtypes.Each analyzed subtype contained at least 10% of the total samples.
  • Simulation design: Simulated datasets implanted a k-gene mutually exclusive pathway, added background alterations, and assigned mutation frequencies across n samples.The reported configuration used m = 100, n = 500, k = 3, three pathway frequencies, five background genes, and 100 million CoMEt iterations from three starts.

S3 Supplementary Tables

The supplementary tables document simulated pathway coverages, comparative recovery results, and breast-cancer alteration collections.

  • Supplementary tables: Table S1 compares alteration collections reported by Multi-Dendrix and CoMEt on TCGA pan-cancer breast-cancer data with k = 3 and t = 4.It also marks differences between datasets with and without the MutSigCV filter.
  • Supplementary tables: The simulation parameter choices for background genes and mutation probability were derived from TCGA glioblastoma and breast-cancer data.Background frequencies matched five highly mutated glioblastoma genes, while q was estimated from a breast-cancer mutation matrix.
  • Supplementary tables: Table S2 reports pathway coverages γP for simulated datasets containing non-overlapping pathways.
  • Supplementary tables: Table S3 reports mean adjusted Rand index across 25 simulated datasets for CoMEt and Multi-Dendrix under consensus and true-parameter settings.The experiments vary the number of implanted pathways and genes per pathway.

S1 Supplementary Figures

The supplementary figures examine convergence, test-statistic behavior, computational performance, simulated mutation distributions, and CoMEt outputs across datasets.

  • Supplementary figures: Figure S1 plots total variation distance across iterations for MCMC runs of 1M and 10M iterations.
  • Supplementary figures: Figure S2 contrasts the MEMo coverage statistic Γ with CoMEt’s exclusivity statistic T in cases where Γ deflates or inflates the P-value.
  • Supplementary figures: Figure S3 compares exact and binomial P-values for all k = 3 alteration sets in GBM and BRCA, colored by co-occurring alterations.Exact P-values are much smaller than binomial P-values mainly when co-occurrences are relatively low; those cases are fastest for tail enumeration.
  • Supplementary figures: Figures S4–S6 show the CoMEt visualization application, average pathway ranks against Dendrix and muex, and runtime comparisons across 25 simulated datasets.
  • Supplementary figures: Figures S7–S9 display simulated mutation-frequency distributions, GBM edge-weight distributions, and CoMEt results for AML with t = 4 and k = 4.
  • Supplementary figures: Figure S10 presents mutation matrices for TCGA GBM, AML, STAD, and BRCA, with alterations as rows and samples as columns.Orange marks co-occurring alterations, blue marks exclusive alterations, and grey marks unaltered cells.
Loading 1503.08224v1…