Source-linked AI summary
Fast and Scalable Inference of Multi-Sample Cancer Lineages
Victoria Popic, Raheleh Salari, Iman Hajirasouliha, Dorna Kashef-Haghighi, Robert B. West, Serafim Batzoglou
TL;DR
Heterogeneous tumor samples and the scale of multi-sample datasets make specialized cancer phylogeny methods necessary. LICHeE infers lineage trees and subclonal mixtures from SSNV presence patterns and VAFs, and evaluations found effective reconstruction across simulations and published datasets. Its main scope boundaries are reliance on suitable deep-sequencing-derived measurements, incomplete explicit CNV modeling, and unresolved ranking of equally valid trees.
Problem
Sample heterogeneity and large numbers of samples make manual and traditional phylogeny methods inadequate for scalable, fine-grained tumor lineage reconstruction.
Method
LICHeE partitions SSNVs by presence patterns, clusters them by VAFs, and searches a constraint network for valid lineage trees while inferring sample subclonal mixtures.
Results
LICHeE effectively reconstructed lineage trees and sample heterogeneity in simulations and three published ultra-deep multi-sample cancer datasets, often finding unique valid trees.
Takeaways & Limitations
LICHeE provides automated multi-sample cancer lineage inference that can uncover heterogeneity and improve on traditional tree-building analyses within its supported data scope.
Takeaways & Limitations
LICHeE does not explicitly detect or incorporate CNVs, and it does not define criteria for ranking multiple trees that equally satisfy the available ordering constraints.
Abstract
from arXiv · showhide
Somatic variants can be used as lineage markers for the phylogenetic reconstruction of cancer evolution. Since somatic phylogenetics is complicated by sample heterogeneity, novel specialized tree-building methods are required for cancer phylogeny reconstruction. We present LICHeE (Lineage Inference for Cancer Heterogeneity and Evolution), a novel method that automates the phylogenetic inference of cancer progression from multiple somatic samples. LICHeE uses variant allele frequencies of SSNVs obtained by deep sequencing to reconstruct multi-sample cell lineage trees and infer the subclonal composition of the samples. LICHeE is open-sourced and available at http://viq854.github.io/lichee.
Background
Cancer phylogeny reconstruction must account for heterogeneous samples and scale beyond manual or traditional tree-building approaches. LICHeE uses SSNV presence patterns and VAFs to infer lineage trees and subclonal mixtures efficiently across multiple samples.
- Background: Tumors contain mixtures of cell subpopulations with distinct somatic variants, complicating phylogenetic reconstruction from multiple samples.Existing approaches include manual analysis and traditional phylogeny methods, which are not designed to handle sample heterogeneity and large multi-sample datasets.
- Background: LICHeE reconstructs multi-sample tumor phylogenies and decomposes tumor subclones from targeted deep-sequencing SSNV datasets.Given SSNV VAFs from multiple samples, it finds lineage trees consistent with presence patterns, VAFs, and the cell division process while estimating sample heterogeneity.
- Background: LICHeE reduces search complexity through VAF constraints, runs in a few seconds on hundreds of SSNVs, and does not require preprocessing.Its SSNV-calling heuristic uses present and absent thresholds, then assigns greyzone variants using the inferred tree; CNVs are not automatically incorporated, but CP values can be used instead.
- Background: LICHeE partitions SSNVs by sample occurrence, clusters them by VAF similarity, and encodes possible evolutionary orderings in an acyclic constraint network.The network's nodes are SSNV clusters, and its edges represent possible predecessor relationships satisfying ordering constraints.
- Background: The method searches the constraint network for valid lineage trees whose parent-daughter edges preserve somatic VAF consistency, then reconstructs sample composition from the tree.The individual samples are added as leaves after the tree search, with their subclone composition inferred by tracing corresponding cell lineages.
- Background: LICHeE was evaluated on simulations and ultra-deep multi-sample datasets from ccRCC, HGSC, and breast-cancer xenoengraftment studies.It produced near-identical trees to published ccRCC trees, better-supported HGSC trees, and additional heterogeneity findings; unique trees were found for each ccRCC patient and all HGSC patients except Case5.
Lineage Tree Reconstruction on ccRCC Data
LICHeE reconstructs ccRCC lineage trees using SSNV presence patterns and VAF constraints, revealing heterogeneity and generally matching published trees while resolving conflicting or weakly supported branches.
- 602 validated variants from eight ccRCC individuals supported comparisons between LICHeE trees and published maximum-parsimony trees.The original study clustered VAFs within each sample before reconstructing trees from presence patterns.
- LICHeE-generated trees highly match published results for several patients, while omitting branches lacking sufficient SSNV evidence.For EV006, one partially shared group had no supporting mutations, and another was excluded because it was supported by only one SSNV.
- 0.34 parent VAF versus group averages of 0.32, 0.27, and 0.21 constrained RMH004’s three overlapping groups to at most one descendant branch.LICHeE removed the two least populated conflicting branches to produce a valid tree.
- LICHeE suggests two subclones in RMH004 sample R4, although the evidence is weak because the supporting shared group contains only three mutations and has low frequency.The VAF constraint rules out assigning the R4-private group beneath the R4/R2 shared group.
- Applying VAF constraints to tree topologies and clustering SSNVs across samples reveals heterogeneity that presence-pattern reconstruction or single-sample clustering can miss.Cross-sample clustering can distinguish mixed lineages even when subclones have similar VAFs in individual samples.
Lineage Tree Reconstruction on HGSC Data
LICHeE was evaluated on six HGSC cases and identified tree structures and mixed lineages that the published neighbor-joining analyses did not fully support from the mutation data.
- 340 validated somatic mutations from 19 tumor samples across six HGSC patients were used to compare LICHeE with published neighbor-joining trees.The published trees used Pearson correlation distances computed from binary sample presence patterns.
- In Case1, VAFs and binomial-test results supported mutations shared by samples a, b, and d, rather than the published early divergence of d followed by c.LICHeE’s proposed shared group was supported both by the study’s test and by its hard-threshold caller.
- Case5 contained mutations in samples a, b, c, e, and f but not d, contradicting the reported tree’s implied profile excluding c.The authors attribute the published result to private mutations in c lowering its Pearson correlation with other samples.
- Traditional neighbor joining could not reveal sample heterogeneity, whereas LICHeE detected mixed lineages in samples b, e, and f.The authors state that Pearson correlation was unsuitable for these data.
- Case5 had two mutation clusters with mean VAFs [0.29, 0.34, 0.3, 0.19, 0.4] and [0.17, 0.14, 0.13, 0.1, 0.13] across samples a, b, c, e, and f.The group (b, f) was supported by 14 mutations, while group (e, f, b) had weak support from three mutations.
Lineage Tree Reconstruction on Xenoengraftment Data
LICHeE was compared with single-cell phylogenies and clonal genotypes in breast-cancer xenoengraftment data, closely reproducing validated lineage structures and sample compositions.
- LICHeE was applied to deep-sequencing VAF data from breast-cancer xenoengraftment cases SA494 and SA501 and compared with single-cell phylogenies and clonal genotypes.The source study used deep-genome and single-cell sequencing, PyClone, and Bayesian phylogenetic inference.
- For SA501, LICHeE mirrored the single-cell ancestral genotype splitting into sibling clades A and B, with samples X1 and X2 mixing A-D and X4 dominated by E.The single-cell study found genotype A absent from X4 and genotype E private to X4.
- LICHeE collapsed genotypes B and C because their bulk CP values were highly similar across X1, X2, and X4.Additional samples in the PyClone analysis showed larger CP differences, but were not part of this LICHeE clustering input.
- For SA494, LICHeE reconstructed the exact topology of two mutually exclusive clades emerging from an ancestral clone shared by samples T and X4.It also identified two X4-private mutation clusters, with the descendant cluster present in about 20% of the sample.
Simulations
LICHeE was evaluated on simulated heterogeneous lineage trees generated under randomized and localized sampling, with sequencing and sampling noise added to simulated VAFs. Performance depended on coverage, sample number, and CNVs, while mutation ordering remained highly accurate in most settings.
- Simulation design: The simulator generated branching cancer lineages by introducing SSNV or CNV events into expanding cell populations.Samples were formed by selecting subclones through randomized or spatially localized schemes.
- Simulation design: Localized sampling selected subclones mostly from distinct branches to mimic biopsies from spatially distinct sites.Each localized sample could include up to five subclones from one subtree, one from a neighboring subtree, and up to 20% normal contamination.
- Simulation design: Simulated VAFs incorporated multinomial sampling, Binomial sequencing noise at 100X, 1,000X, or 10,000X coverage, and Q30 base-call accuracy.The Binomial model used the true VAF as its mean parameter.
- Simulation results: 81% of SSNVs were preserved at 100X coverage with 15 samples, whereas 1,000X coverage preserved 91-94% of SSNVs and 83-88% of mutation pairs.With CNVs, the 1,000X results declined to 91-92% SSNVs and 66-70% SSNV pairs.
- Simulation results: 99-100% of ancestor-descendant relationships were correctly ordered without CNVs, compared with 92-96% when 80% of SSNVs lay in CNV regions at 1,000X coverage.Reverse sibling-to-ancestor-descendant violations remained rare, reaching up to 7%.
- Scope: LICHeE is best applied to deep-sequencing data with low VAF variance, while lower-coverage whole-genome sequencing and direct aneuploidy or large-CNV modeling remain future work.The method can also use copy-number values from existing tools that correct for SCNAs, LOH, and sample purity.
Grouping and Clustering SSNVs
LICHeE groups SSNVs by their presence patterns across samples, assigns uncertain variants using VAF similarity and set cover, and then clusters each group by VAF profiles. The resulting clusters carry VAF centroids used in downstream lineage inference.
- Presence-profile grouping: Each SSNV receives a binary profile indicating presence or absence in every sample, and SSNVs with identical profiles form one group.An all-ones profile represents germline variants when a normal control is included.
- Presence-profile grouping: VAF thresholds classify SSNVs as robustly present or absent when their profiles meet a minimum same-profile support requirement.The default minimum number of other robust SSNVs sharing a profile is one.
- Presence-profile grouping: Non-robust SSNVs are assigned to candidate groups using VAF similarity, and unassigned variants are organized through a set cover formulation to minimize new groups.Similarity compares VAFs across samples using the minimum-to-maximum VAF ratio.
- VAF clustering: Within each SSNV group, EM clustering of a n×s VAF matrix produces clusters with associated VAF centroid vectors.Small or noisy clusters may be removed, or neighboring clusters may be collapsed based on centroid distance.
- VAF clustering: For patient EV007, the evolutionary constraint network summarizes each SSNV group using its binary profile, VAF centroid, and number of assigned SSNVs.The network edges encode potential precedence relationships between groups.
Evolutionary Constraint Network Construction
LICHeE constructs a directed acyclic constraint network whose nodes are SSNV clusters and whose edges encode evolutionary precedence allowed by VAF and sample-presence constraints. It then organizes candidate relationships using profile levels and VAF error minimization.
- Network representation: Each network node represents an SSNV cluster, while an edge u → v indicates that u could precede v evolutionarily; the root represents germline.The network is a DAG, and every node can connect to the root.
- Edge constraints: A candidate edge requires the parent VAF to be no smaller than the child VAF beyond an error margin and disallows a zero-VAF parent with a nonzero-VAF child.The margin combines cluster standard errors with a configurable parameter.
- Profile-level organization: Nodes are assigned levels by the Hamming weight of their binary profiles, and conflicting profiles at the same level cannot be connected.Between levels, only higher-level nodes can parent lower-level nodes.
- Edge resolution: For nodes sharing both a level and binary profile, LICHeE directs the edge to minimize VAFERR.The indicator function identifies samples where the child VAF exceeds the parent VAF.
- Example: Figure 6 illustrates the resulting constraint network for ccRCC patient EV007.The reported spanning tree is highlighted within the network.
Phylogenetic Tree Search
LICHeE searches constraint-network spanning trees that satisfy local and global VAF consistency, ranks valid alternatives, and exposes ambiguity or search failure to users.
- Tree search: Valid lineage trees are spanning trees of the constraint network that satisfy the VAF-based requirement at every node.The method permits child VAF sums to exceed parent values within error margin ϵ because not all branches need be observed.
- Sample decomposition: Each valid lineage tree supports sample decomposition by enumerating paths from the germline root to the last mutation-containing node in each sample.These paths define the inferred subpopulations within each sample.
- Tree search: The extended Gabow–Myers search prunes trees immediately when adding an edge violates the VAF constraint.Its runtime is O(|V| + |E| + |E|N), where N is the number of spanning trees, and is typically a few seconds in practice.
- Tree ranking: A quadratic-programming step checks global consistency and ranks valid trees by their fit to the VAF data.Because evaluating every tree can be impractical, LICHeE first ranks trees locally and applies quadratic programming to the top k, with k = 5 by default.
- Ambiguity handling: When several trees fit equally well because unobserved branches create ambiguity, LICHeE reports ranked alternatives and provides a GUI for topology exploration.The authors do not define more sophisticated optimality criteria and suggest suboptimal trees may contain biologically significant signals.
- Failure handling: No valid lineage tree may be found when the noise margin is too narrow or an SSNV group has a misclassified binary profile.The current recovery procedure removes non-robust SSNV-group nodes one by one, beginning with the smallest nodes.
Visualization
LICHeE provides an interactive GUI for examining constraint networks and lineage trees, tracing sample lineages, and inspecting SSNV and sample-composition information.
- Tree visualization: The GUI visualizes constraint networks and phylogenetic trees, with each input sample shown as a leaf connected to nodes containing its SSNVs.Selecting a sample leaf displays that sample’s lineage using a DFS traversal from the germline root.
- Interactive exploration: Users can click nodes for additional information, rearrange or remove nodes, and collapse nodes belonging to the same group.The program also reports the number of trees found and the highest-ranking tree’s score, while allowing users to choose how many trees to display.
Implementation
LICHeE was implemented in Java, released as open-source software, and developed and evaluated by the listed authors.
- Software: LICHeE was implemented in Java and made freely available online as open-source software.The paper identifies the project website as the distribution location.
- Disclosures: SB disclosed advisory and scientific-board roles at DNAnexus, 23andMe, and Eve Biomedical.These affiliations are reported in the paper’s disclosure information.
- Contributions: VP, RS, and SB developed the method, while VP implemented the LICHeE program and simulator and drafted the manuscript.The authors jointly performed the evaluation and analysis, with additional manuscript contributions from RS, SB, DK, and IH.
Appendix A — Evaluation of Phylosub
The appendix evaluates PhyloSub on the Gerlinger et al. dataset, finding that shared mutations generally cluster correctly while partially shared mutations are often grouped together despite differing presence patterns.
- Evaluation setup: PhyloSub was run with default parameters on the Gerlinger et al. dataset, comparing each patient’s top three tree structures with the published tree.The authors found it difficult to identify similarities between the reported and published topologies.
- Clustering results: Shared mutations were generally clustered correctly, whereas partially shared mutations were often combined into one group that violated their presence pattern.For patient EV003, the reported trunk included all shared mutations plus one private and two partially shared mutations.
- Clustering results: For EV003, the second branch contained several partially shared mutations and remaining private mutations with very different presence patterns.The example is illustrated in Figure 8, which presents the PhyloSub result for patient EV003.
Appendix B — Experimental Details of the ccRCC, HGSC, and Xenoengraftment Comparison
The comparison evaluates LICHeE on ultra-deep multi-sample cancer datasets, using dataset-specific validation criteria and noise thresholds. HGSC and xenoengraftment analyses apply explicit filtering and threshold choices to accommodate data quality and validation uncertainty.
- ccRCC: The ccRCC dataset validated 602 nonsynonymous substitutions and indels across multiple samples from 8 individuals using ultra-deep amplicon sequencing.Sequencing averaged >400x coverage, and 587 mutations met the analysis criterion of >100x minimum coverage.
- HGSC: The HGSC dataset included 19 tumor samples from 6 patients and validated 340 somatic mutations using deep amplicon sequencing.Median sequencing coverage exceeded 5000x.
- HGSC: HGSC analyses used Tabsent = 0.005 and Tpresent = 0.01, except Case5, which used 0.01 and 0.04 because of higher noise.At least three mutations were required as evidence for a tree node, and samples g and h were excluded from Case5.
- Xenoengraftment: The xenoengraftment study used deep-genome and single-cell sequencing to assess clonal dynamics in breast cancer tissue transplanted into immunodeficient mice.Single-cell analyses covered passages SA501 and SA494, with LICHeE run using dataset-specific thresholds and a minimum cluster size of 6.