Source-linked AI summary
The inference of gene trees with species trees
Gergely J. Szöllosi, Eric Tannier, Vincent Daubin, Bastien Boussau
TL;DR
Gene trees often differ from species trees, motivating models that describe their relationship and integrate sequence evolution with gene-family evolution. This review synthesizes such models and reports improved inference accuracy, while noting that current approaches remain incomplete and face scaling challenges for large genomic datasets.
Problem
Gene and species histories are linked but often differ, while most existing phylogenetic models focus on sequence-based gene-tree reconstruction rather than jointly modeling both histories.
Method
The review examines probabilistic gene-tree–species-tree models combined with sequence-evolution models, covering duplication, loss, lateral transfer, and incomplete lineage sorting.
Results
Methods modeling dependence between gene trees and species trees have shown improved accuracy for inferring both gene trees and species trees.
Takeaways & Limitations
Better gene trees provide a stronger basis for studying genome evolution and reconstructing ancestral chromosomes and ancestral gene sequences.
Abstract
from arXiv · showhide
Molecular phylogeny has focused mainly on improving models for the reconstruction of gene trees based on sequence alignments. Yet, most phylogeneticists seek to reveal the history of species. Although the histories of genes and species are tightly linked, they are seldom identical, because genes duplicate, are lost or horizontally transferred, and because alleles can co-exist in populations for periods that may span several speciation events. Building models describing the relationship between gene and species trees can thus improve the reconstruction of gene trees when a species tree is known, and vice-versa. Several approaches have been proposed to solve the problem in one direction or the other, but in general neither gene trees nor species trees are known. Only a few studies have attempted to jointly infer gene trees and species trees. In this article we review the various models that have been used to describe the relationship between gene trees and species trees. These models account for gene duplication and loss, transfer or incomplete lineage sorting. Some of them consider several types of events together, but none exists currently that considers the full repertoire of processes that generate gene trees along the species tree. Simulations as well as empirical studies on genomic data show that combining gene tree-species tree models with models of sequence evolution improves gene tree reconstruction. In turn, these better gene trees provide a better basis for studying genome evolution or reconstructing ancestral chromosomes and ancestral gene sequences. We predict that gene tree-species tree methods that can deal with genomic data sets will be instrumental to advancing our understanding of genomic evolution.
Introduction
Gene trees capture genomic histories that are linked to, but often differ from, species histories because multiple evolutionary processes shape them. This review examines models that integrate these processes with sequence evolution to improve inference of gene and species trees.
- Introduction: Gene trees and species trees are linked but often differ because duplication, loss, lateral transfer, recombination, and incomplete lineage sorting generate distinct histories.Incomplete lineage sorting can leave topological signatures when segregating mutations cross speciation events.
- Introduction: Up to 30% of the human genome sequence is estimated to be more closely related to Gorilla than to Chimpanzee because of incomplete lineage sorting.
- Introduction: Integrating sequence alignments with gene-family evolution and species-tree information can support inference of gene trees from known species trees and species histories from gene alignments.The proposed probabilistic framework can account for duplication, loss, lateral transfer, and/or incomplete lineage sorting.
- Introduction: Parsimony is conceptually straightforward but requires difficult relative weighting of events and can produce equivalent solutions that hinder efficient exploration or integration.
- Introduction: Probabilistic models combine sequence-evolution and gene-family-evolution probabilities, allowing integration or sampling across many evolutionary scenarios rather than selecting a single parsimonious scenario.Conditional independence permits multiplication of the model probabilities when the species tree is independent of sequence given the gene tree.
- Introduction: Methods modeling dependence between gene trees and species trees have shown improved accuracy for inferring both types of trees.The review presents their assumptions, operation, and results, emphasizing probabilistic models while also discussing parsimony where probabilistic methods are unavailable.
species tree
Gene family evolution involves incomplete lineage sorting, duplication and loss, transfer, and hybridization-related processes. Although models exist for individual processes and some combinations, no coherent model yet handles all processes together.
- Gene family evolution includes incomplete lineage sorting, gene duplication and loss, and gene transfer.
- Hybridization can affect large genome portions and replace genes in the receiving species, while allopolyploidization allows both genomes to cohabit across generations.
- Published models account for individual processes, and recent approaches integrate several of them.
- No published model currently handles all processes together in a coherent statistical framework.
Gene birth-death generates gene trees along the species tree
Gene family evolution models generate gene trees within the constraints of a species tree, often by applying birth-death processes along species-tree branches. Simplifying these models to gene-tree topologies reduces complexity but discards branch-length information.
- Gene family evolution can be viewed as generating a gene tree inside a species tree, with speciation events constraining the process.
- A birth-death process begins above the species-tree root and splits into processes in child lineages at each speciation node.
- For a species tree with n branches, including the branch above the root, the model contains n birth-death processes whose parameters may be dependent.
- Dependence can be modeled through parameter evolution along the species tree or through lateral gene transfer between lineages.
- Topology-only models simplify inference by replacing dated event probabilities with probabilities of event sequences or lineage counts, but discard potentially useful gene-tree branch-length information.
species tree.
Coalescent models explain gene–species tree discordance and use it to infer population parameters and species trees. Extensions address genomic recombination and uncertain species assignments, but rely on assumptions about sampling and hybridization.
- Allele histories can differ from species histories in bifurcation timing and topology, especially with short speciation intervals or large ancestral populations.
- Coalescent models use discordance across loci to estimate effective population sizes along the species tree.
- The multispecies coalescent is widely used to infer species trees from multiple loci with one or several alleles, using branch lengths or only topologies.
- Hidden Markov models along chromosomes use changes in local gene trees to provide information about recombination rates from species trees and genomic alignments.
Models of gene duplication and loss
Gene duplication-and-loss models range from probabilistic reconciliation and birth-death formulations to branch-specific models for joint gene-tree and species-tree inference. Speed improvements often involve simplifying event or branch-length modeling, with corresponding realism or parameterization trade-offs.
- Duplication-and-loss models commonly use fixed duplication and loss rates and may ignore population-level processes.
- Birth-death models can combine species-tree dates with gene-family-specific sequence-evolution rates in a hierarchical framework.
- Poisson-process models compute parsimonious reconciliation probabilities for gene-tree topologies against species trees, gaining speed by omitting a loss parameter.
- Branch-specific duplication and loss parameters enable models on non-dated species trees, while topology-only reconciliation speeds computation.
- Joint inference still requires gene-tree branch lengths for sequence evolution, and independent gene-family rates reduce global parameters but greatly increase family-specific parameters.
Models of lateral gene transfer
Lateral gene transfer models represent transfers and, in some cases, gene replacement alongside gene-tree and species-tree evolution. Their assumptions and computational costs vary, and existing approaches have important limits.
- Lateral gene transfer incorporates a gene from a different species into a genome.
- Transfer-only models usually represent gene replacement, depending on whether a homologous recipient gene is conserved or lost.
- Birth-death models differ from replacement models because replacement breaks the independence assumption between lineages.
- Suchard’s model used random walks between gene-tree and unrooted species-tree topologies, with path probabilities summed under a Poisson event model.
- Suchard’s approach explained a forest of over 140 gene trees but was computationally limited to trees with only 6 to 8 species.
- The approach did not distinguish time-consistent transfers from “back-to-the-future” transfers because its species-tree topology was not anchored in time.
Models that combine the above
Combined models extend duplication-loss frameworks to include transfer, incomplete lineage sorting, hybridization, and population-level processes. The proposed combinations broaden process coverage but retain simplifying assumptions or omissions.
- DLCoal: DLCoal requires a dated species tree, effective population sizes, duplication and loss rates, a dated locus tree, and a gene tree.
- DLCoal: The DLCoal model generates a locus tree through duplication and loss, then generates a gene tree through a coalescent process along that locus tree.
- DTLSR and ODT: DTLSR and ODT modify birth-death processes to include gene duplications and transfers without requiring the additional locus-tree object used by DLCoal.
- ODT extensions: The ex-ODT extension permits transfers involving extinct or otherwise unrepresented species by modeling a broader evolving species population.
- Model boundaries: ODT and DTLSR do not model population-level processes or allele fixation, whereas the two model types could be combined hierarchically.
- Hybridization: Hybridization models use rooted networks with two parental branches and a probability γ assigning genes to parent A versus parent B.
Simulation and inference
The reviewed pipeline treats phylogenetic analysis as linked statistical inferences, but many early steps remain sequential and isolated. Gene-tree/species-tree models and simulation procedures provide more integrated alternatives while facing transfer-related computational challenges.
- Pipeline: The phylogenetic pipeline includes sequencing correction, assembly, annotation, gene-family clustering, alignment, tree reconstruction, and species-tree inference.
- Pipeline: Most pipeline steps are sequential, so later analyses disregard uncertainty from earlier estimates and provide no feedback.
- Integrated inference: Gene-tree/species-tree models connect gene-tree construction with species-tree construction, enabling communication between these pipeline stages.
- Simulation: Transfer simulation couples contemporaneous species-tree branches and prevents algorithms relying on independent gene lineages.
- Simulation: No simulation method described accounts for transfers through complete phylogenies containing extinct and unsampled branches without potentially expensive computation.
- Simulation: A parametric bootstrap-like ALE+ex-ODT procedure simulated alignments while retaining empirical gene-family-size and evolutionary-rate properties.
- Likelihood computation: Likelihood calculations use dynamic programming over gene trees and species trees, with sequence data and gene-family data supplying leaf information.
Impact on systematics and genome evolution
Models linking gene and species trees improve species-tree estimation, gene-tree reconstruction, and interpretation of genome evolution. They also expose biological history encoded in gene-tree conflict, though current methods remain sensitive to input quality and scale.
- Gene-tree/species-tree methods repeatedly improve species-tree estimation, gene-tree estimation, and genomic-evolution studies over methods that do not model gene-family evolution.
- Coalescent-based species-tree methods can outperform concatenation, but erroneous or unresolved input gene trees can produce unresolved species trees.
- Gene duplication, loss, and transfer models have reconstructed species phylogenies consistent with the literature, including analyses using 18,896 gene trees.
- Gene transfer supplies information about ancestral species history, including event timing, species-tree orientation, and rooting.
- In Cyanobacteria, transfer information supported a root differing from outgroup-based rooting, with the unusual root supported by more than 200 transfer events.
- PHYLDOG jointly reconstructed species and gene trees with good accuracy on simulated and real data, outperforming sequence-alignment-only methods.
- Using a species tree improves downstream ancestral reconstruction: ancestral sequences were more accurate, and one reconstructed protein was more stable with better enzymatic capabilities.
Next challenges
Future work must make gene-tree/species-tree models more computationally efficient while integrating additional evolutionary processes and data-analysis stages. Existing results support broader biological applications, but a unified model remains unavailable.
- Joint inference is computationally difficult because likelihood calculations require summing over large spaces of gene trees or allele histories.
- The main challenge is scaling increasingly complex evolutionary models to large datasets containing many species and gene families.
- Analytical approaches such as SNAPP bypass gene-tree sampling by integrating over possible allele histories, but their improvement over multispecies-coalescent models remains open.
- No model currently unifies coalescent, duplication, loss, transfer, substitution, and rearrangement processes.
- Integrating gene-tree/species-tree models with alignment methods could improve inference of alignments, gene trees, and possibly species trees.
- Phylogenetically aware methods have shown promise across pipeline stages, including genome assembly, despite the absence of a fully integrated pipeline.
- For larger datasets, efficiency improvements beyond parallelization, staged family handling, and importance sampling will be needed.
Conclusion
The relationship between gene and species trees has become clearer through new evolutionary models and inference algorithms. Future methods must simultaneously increase biological realism and scale to growing datasets.
- Conceptual and methodological advances have clarified the relationship between gene trees and species trees and enabled statistical inference based on birth-death processes and dynamic programming.
- Future developments face a tension between modeling genome evolution more accurately and scaling to increasingly large datasets.
Y Wu. COALESCENT-BASED SPECIES TREE INFERENCE FROM GENE TREE TOPOLOGIES UNDER INCOMPLETE LINEAGE SORTING BY MAXIMUM
The paper presents gene tree–species tree models as representations of evolutionary processes operating from sites and genes to whole species genomes. These models include birth–death events, coalescent processes, and probabilistic reconciliation within phylogenomic inference.
- Evolutionary processes form a hierarchy from point mutations at sites to gene duplication, loss, and transfer, and finally speciation and extinction among genomes.
- Birth–death processes model speciation and extinction at the species level and gene-family evolution inside a species tree.
- The multispecies coalescent models duplication and loss as population-level processes in which new loci or deletions drift until fixation or extinction.
- Models differ in the evolutionary processes they represent, including substitutions, fixed duplication and loss, incomplete lineage sorting, and transfer events.
- The phylogenomics pipeline includes sequencing, assembly, gene-tree construction, and species-tree construction, with models specifying likelihoods for inferred trees.
- Amalgamated likelihood estimation combines clades sampled from gene-tree posteriors to explore reconciled gene trees, including trees absent from the original sample.