Source-linked AI summary

Efficient Exploration of the Space of Reconciled Gene Trees

Gergely J. Szöllősi, Wojciech Rosikiewicz, Bastien Boussau, Eric Tannier, Vincent Daubin

arXiv:1306.2167v1q-bio.PEq-bio.BMq-bio.GN

TL;DR

Existing clade-probability approaches can approximate many gene-tree topologies from small samples but ignore dependencies between nonoverlapping clades. ALE extends dynamic programming to sum over amalgamated reconciled gene trees, improving reconstruction accuracy and reducing apparent phylogenetic discord.

  • Problem

    Conditional clade probabilities approximate posterior probabilities from small tree samples but ignore that nonoverlapping clades are not necessarily independent.

  • Method

    ALE extends dynamic programming to simultaneously sum over all reconciled gene trees that can be amalgamated from sampled clades, under a duplication-transfer-loss reconciliation model.

  • Results

    24.3%, 59.1% and 45.8% reductions in duplications, transfers and losses, respectively, were observed for dataset II; simulations also showed substantially improved reconstruction accuracy.

  • Takeaways & Limitations

    Joint-likelihood inference reduces apparent phylogenetic discord, with most joint-tree bipartitions supported by sequence while some receive strong support only jointly.

  • Takeaways & Limitations

    Simulations did not rule out overfitting of the species tree for sample sizes larger than those tested.

Abstract

from arXiv · show

Gene trees record the combination of gene level events, such as duplication, transfer and loss, and species level events, such as speciation and extinction. Gene tree-species tree reconciliation methods model these processes by drawing gene trees into the species tree using a series of gene and species level events. The reconstruction of gene trees based on sequence alone almost always involves choosing between statistically equivalent or weakly distinguishable relationships that could be much better resolved based on a putative species tree. To exploit this potential for accurate reconstruction of gene trees the space of reconciled gene trees must be explored according to a joint model of sequence evolution and gene tree-species tree reconciliation. Here we present amalgamated likelihood estimation (ALE), a probabilistic approach to exhaustively explore all reconciled gene trees that can be amalgamated as a combination of clades observed in a sample of trees. We implement ALE in the context of a reconciliation model, which allows for the duplication, transfer and loss of genes. We use ALE to efficiently approximate the sum of the joint likelihood over amalgamations and to find the reconciled gene tree that maximizes the joint likelihood. We demonstrate using simulations that gene trees reconstructed using the joint likelihood are substantially more accurate than those reconstructed using sequence alone. Using realistic topologies, branch lengths and alignment sizes, we demonstrate that ALE produces more accurate gene trees even if the model of sequence evolution is greatly simplified. Finally, examining 1099 gene families from 36 cyanobacterial genomes we find that joint likelihood-based inference results in a striking reduction in apparent phylogenetic discord, with 24%, 59% and 46% percent reductions in the mean numbers of duplications, transfers and losses.

MATERIALS AND METHODS

The method combines conditional clade probabilities with probabilistic reconciliation to sum over, sample, and optimize reconciled gene trees assembled from observed clades.

  • Gene Tree Reconciliation using Conditional Clade Probabilities: Conditional clade probabilities estimate gene-tree topology probabilities from clade and split frequencies in an MCMC tree sample.The recursion terminates at single-leaf clades and assigns zero probability to trees containing unobserved clades.
  • Gene Tree Reconciliation using Conditional Clade Probabilities: ALE extends reconciliation methods to iterate over all gene-tree topologies whose clades can be amalgamated from the sampled trees.It replaces bifurcating reconciliation events with alternative resolutions of the corresponding clades.
  • Gene Tree Reconciliation using Conditional Clade Probabilities: The joint likelihood sums alignment likelihood and reconciliation probability over amalgamated gene-tree topologies given a species tree and reconciliation model.The model is formulated for gene-lineage evolution through reconciliation events in the species tree.
  • Gene Tree Reconciliation using Conditional Clade Probabilities: Dynamic programming replaces gene-tree branches with clades and sums over observed daughter-clade splits weighted by conditional probabilities.The recursion applies this construction across reconciliation events to calculate the total joint likelihood.
  • Gene Tree Reconciliation using Conditional Clade Probabilities: Stochastic backtracking samples reconciled gene trees from the likelihood sum, while maximization identifies the most likely reconciled tree.For the manuscript’s data, likelihood calculation takes a few seconds.

Validation based on “real” gene trees

The validation simulates amino-acid sequences from reconciled gene trees based on realistic cyanobacterial gene-family data, then reconstructs them under a simpler sequence model.

  • Validation based on “real” gene trees: Simulated sequences use gene-tree topologies, branch lengths, and alignment sizes derived from 1099 cyanobacterial gene families across 36 genomes.The reconstructed reconciled gene trees first serve as the histories used to generate simulated sequences.
  • Validation based on “real” gene trees: Sequences are simulated with a complex LG model including across-site rate variation and invariant sites, but reconstructed with a Poisson model lacking rate variation.This setup tests reconstruction under a simplified sequence-evolution model.

Sequence data

The sequence-data workflow builds two cyanobacterial gene-family datasets, obtains MCMC tree samples, and generates reconciled reference trees and simulated alignments for evaluation.

  • Sequence data: Alignments are produced from extracted amino-acid sequences using MUSCLE and subsequently cleaned with GBLOCKS.The cleaned alignments are used for downstream tree sampling and simulation.
  • Sequence data: Dataset I contains 342 universal single-copy families, while dataset II contains 1099 families with at least ten genes in any of 36 cyanobacterial genomes.Families with more than 150 genes were excluded when constructing the simulated dataset.
  • Sequence data: PhyloBayes MCMC samples are obtained under LG+Γ4+I, with at least 3000 post-burn-in samples for the cleaned alignments.For dataset I, 10000 trees are sampled for the universal single-copy families.
  • Sequence data: ALEsample generates at least 5000 reconciled trees for sampling DTL rates, and ALEml identifies maximum-likelihood DTL rates and reconciled trees.Majority consensus and fully resolved reference trees are derived from the ALEsample trees.
  • Sequence data: Amino-acid sequences are simulated with bppseqgen under an LG model with Gamma-distributed site rates and 10% invariant sites.The complex simulation model uses α = 0.1 for the Gamma distribution.

Inference for simulated data

Across simulated and biological cyanobacterial datasets, joint-likelihood reconstruction improves gene-tree accuracy and substantially reduces inferred reconciliation events, with no observed simulation trend indicating overfitting.

  • Inference for simulated data: Joint-likelihood reconstruction substantially improves accuracy over sequence-only inference across both simulated datasets.The improvement remains significant when joint inference uses a simple sequence model and sequence-only inference uses the complex simulation model.
  • Inference for simulated data: Universal single-copy families imply nearly single-copy ancestral genomes under joint likelihood, whereas sequence-only trees infer 248 single-copy, 34 zero-copy, and 60 multi-copy roots.The single-copy assumption is used to assess reconstruction accuracy on biological data without a known correct tree.
  • Inference for simulated data: For dataset I, joint reconciled trees reduce required duplications by 81.6%, transfers by 70.9%, and losses by 70.2%.These reductions compare events needed to reconcile joint trees with those for sequence-only trees.
  • Inference for simulated data: Mean Robinson–Foulds distance to the species tree is 13.02 for simple-model joint ML trees, versus 17.77 and 21.80 for complex- and simple-model sequence-only trees.The corresponding mean distance for the real gene trees is 11.41.
  • Inference for simulated data: Increasing sample size shows decreasing distance to real and species trees, with no observed overfitting signal in the tested simulations.The authors caution that larger sample sizes could not be ruled out from the available analysis alone.

Analysis of the signal for the phylogenetic discord

Joint-likelihood inference largely agrees with sequence support while adding a minority of strongly supported bipartitions absent from sequence-only analyses. Conversely, sequence-supported bipartitions are rarely rejected by joint inference, and simulations show similar patterns.

  • 71% of bipartitions in joint trees have statistical support > 0.95 according to sequence alone.
  • 6.4% of bipartitions in joint trees have support > 0.95 according to joint likelihood but < 0.05 according to sequence alone.
  • Simulations show similar statistical-support patterns to the observed bipartitions.
  • 85.8% of bipartitions have an absolute difference in sequence-only and joint-likelihood support < 0.1.
  • 1.4% of bipartitions are strongly supported by joint likelihood but not sequence, compared with 0.18% strongly supported by sequence but not joint likelihood.

I. DISCUSSION

ALE exhaustively explores reconciled gene trees from sampled clades and improves reconstruction accuracy and support relative to sequence-only inference. In cyanobacterial data, joint inference reduces inferred evolutionary events and apparent phylogenetic discord, while its applicability remains constrained by species-tree availability and approximation assumptions.

  • Method: ALE exhaustively explores the joint likelihood of many reconciled gene trees by amalgamating clades from a small sampled tree set.It is implemented with a reconciliation model allowing duplication, transfer, and loss, while the computational scheme can extend to other models.
  • Accuracy: Joint inference reconstructs more accurate gene trees than sequence-only inference, even when ALE uses a simplified sequence-evolution model and sequence-only inference uses the correct model.This pattern is supported by simulations based on realistic cyanobacterial gene-family topologies, branch lengths, and alignment sizes.
  • Statistical support: 90.3% versus 80.0% of real-data consensus bipartitions have support > 0.95 in joint versus sequence-only trees.In simulations, the corresponding values are 92.4% versus 83.6%.
  • Reconciliation: 3.6 versus 8.7 transfers per family and 25.8 to 11.4 Robinson-Foulds distance show reduced inferred transfer and phylogenetic discord in joint trees.Transfers comprise 69% of birth events in the reported reconciliation analysis.
  • Limitations: Results beyond cyanobacteria are limited by the availability of well-supported dated species phylogenies, and ALE approximates posterior probabilities using conditional clade probabilities.The method also relies on finite tree samples; in simulations, 98% of bipartitions in real gene trees were present in sampled trees.
  • Interpretation: Joint inference reduces apparent discord, which the authors suggest largely reflects uncertainty in sequence-only reconstructions.Their broader conclusion is based on combining information aggregated across gene families through a putative species tree.

A MINIMAL MODEL OF SPECIATION AND GENE BIRTH AND DEATH

The reconciliation model represents speciation with a constant-size continuous-time Moran process and models gene evolution through duplication, transfer, and loss. Only a small fraction of extant species is sampled, while gene histories may traverse extinct or unsampled lineages.

  • Species history: Lateral gene transfer can place gene histories on extinct and unsampled species branches, not only on the phylogeny of sampled species.Because donors may later go extinct or remain unsampled, the complete historical species phylogeny cannot feasibly be reconstructed directly.
  • Speciation model: The minimal speciation model keeps N species constant and uses a continuous-time Moran process with speciation rate σ.Each event creates two descendants and causes a randomly chosen species to go extinct; only n ≪ N extant species are sampled.
  • Gene evolution: Genes evolve independently through a birth-and-death process allowing duplication at rate δ, transfer at rate τ/(N − 1), and loss at rate λ.Gene copies can also arise or disappear through the modeled speciation dynamics.

1. Amalgamated Sum Over Reconciled Trees

The method recursively sums probabilities over reconciled gene trees by mapping sampled clade splits onto time-discretized species-tree branches and accounting for extinction, speciation, duplication, transfer, and loss events.

  • Time discretization: Time is discretized along species-tree speciation intervals into D subintervals of height Δt_i = (t_i+1 − t_i)/D.The recursion uses species-tree time slices indexed by speciation times and further subdivides each slice.
  • Lineage probabilities: Extinction probabilities track whether lineages on represented or unrepresented species branches leave observed descendants.The calculation also uses single-gene propagators to describe lineage evolution through the species tree.
  • Amalgamated sum: Using these probabilities, the procedure recursively maps gene-tree branches and clades onto represented and unrepresented species-tree branches to sum amalgamated reconciliations.The accompanying figure materials evaluate reconstruction accuracy, DTL events, and statistical support for simulated and cyanobacterial data.
  • Reconciliation recursion: The dynamic program sums over all observed MCMC clade splits γ′,γ′′ while considering no event, duplication, unrepresented speciation, transfer, and loss outcomes.For each split, the recursion combines probabilities for the two descendant lineages and accounts for whether descendants remain represented.
  • Boundary conditions: At species-tree speciation times, represented speciations can be followed by loss, while terminal conditions assign probability 1 when an observed leaf belongs to the terminal branch.The terminal condition assigns zero otherwise.

2. ALE implementation

ALE provides sampling and maximum-likelihood implementations that explore reconciled trees from conditional clade probabilities on a dated binary species tree, with evaluations spanning simulated sample sizes and a cyanobacterial species tree.

  • Common inputs: Both ALE implementations take a dated binary species tree and conditional clade probabilities from an MCMC sample of gene-tree topologies.They set σ = 2N, assuming the species-tree height equals its expected value under the coalescent.
  • Evaluation: The reconstruction-accuracy experiment compares real trees, sequence-only models, and ALEml estimates from samples of 10 through 10000 gene trees.The comparison uses 342 simulated universal single-copy families and includes COMPLEX and SIMPLE sequence-evolution models.
  • ALEsample: ALEsample estimates duplication, transfer, and loss rates with Metropolis-Hastings sampling under the joint likelihood and samples reconciled trees by stochastic backtracking.The rate prior is implicitly flat, and zero-rate boundaries are absorbing.
  • ALEml: ALEml optimizes duplication, transfer, and loss rates by downhill simplex, then finds the maximum-likelihood reconciled gene tree by backtracking.The optimization maximizes L_joint(A|S, δ, τ, λ, σ = 2N).
  • Species-tree setting: The species-tree input is a maximum-likelihood chronologically ordered phylogeny based on 36 genomes and 8332 homologous gene families.The implementation is available from the cited ALE repository.

MAXIMUM ENTROPY DISTRIBUTION FOR MARGINAL SPLIT FREQUENCIES

The paper shows that conditional-clade probabilities derived from marginal split frequencies define a maximum-entropy distribution over rooted gene trees. This distribution is obtained by enforcing total-probability and split-frequency constraints while recursively resolving clades.

  • Conditional-clade probabilities define a distribution over rooted gene-tree topologies from marginal split frequencies.The construction indexes all rooted trees and uses indicators for clades and their resolving splits.
  • For each clade, splits divide it into complementary daughter clades, and sums over possible splits provide the recursive decomposition.The notation distinguishes a clade γ from a split ξ into complementary γ′ and γ′′, with summation over compatible splits.
  • The distribution matches observed split frequencies while satisfying total probability.The constraints require the probabilities to sum to one and reproduce the observed frequencies for each split.
  • The maximum-entropy distribution is obtained by maximizing entropy under these linear constraints.The derivation introduces a Lagrangian and sets derivatives with respect to tree probabilities to zero.
  • Recursive normalization of inner-clade sums establishes that the resulting factors collectively satisfy the required probability normalization.The recursion starts from single-split clades and extends normalization to ancestral clades.
Loading 1306.2167v1…