Source-linked AI summary

Inferring phylogenetic networks with maximum pseudolikelihood under incomplete lineage sorting

Claudia Solís-Lemus, Cécile Ané

arXiv:1509.06075v3q-bio.PEmath.STstat.APstat.CO

TL;DR

Statistical inference of phylogenetic networks remains limited by scalability, especially when modeling incomplete lineage sorting and reticulate inheritance. The paper introduces SNaQ, a quartet-based maximum pseudolikelihood method for multi-locus data, and reports faster inference than full likelihood while retaining accuracy. Applications to Xiphophorus illustrate its use for evolutionary relationships involving hybridization.

  • Problem

    Existing statistical methods for explicit phylogenetic-network inference have limited scalability, despite the need to model nonvertical inheritance and incomplete lineage sorting.

  • Method

    SNaQ estimates semi-directed phylogenetic networks by maximizing a pseudolikelihood built from quartet concordance factors under a coalescent-with-hybridization model.

  • Results

    The method was much faster than full likelihood without compromising accuracy in simulations and was applied to reconstruct hybridization-related relationships among Xiphophorus fishes.

  • Takeaways & Limitations

    Quartet-based pseudolikelihood provides a computationally scalable alternative for estimating biologically interpretable networks from multi-locus data.

  • Takeaways & Limitations

    Inference assumes level-1 networks, and some reticulations remain difficult to detect or orient when taxon sampling is limited or branches are long.

Abstract

from arXiv · show

Phylogenetic networks are necessary to represent the tree of life expanded by edges to represent events such as horizontal gene transfers, hybridizations or gene flow. Not all species follow the paradigm of vertical inheritance of their genetic material. While a great deal of research has flourished into the inference of phylogenetic trees, statistical methods to infer phylogenetic networks are still limited and under development. The main disadvantage of existing methods is a lack of scalability. Here, we present a statistical method to infer phylogenetic networks from multi-locus genetic data in a pseudolikelihood framework. Our model accounts for incomplete lineage sorting through the coalescent model, and for horizontal inheritance of genes through reticulation nodes in the network. Computation of the pseudolikelihood is fast and simple, and it avoids the burdensome calculation of the full likelihood which can be intractable with many species. Moreover, estimation at the quartet-level has the added computational benefit that it is easily parallelizable. Simulation studies comparing our method to a full likelihood approach show that our pseudolikelihood approach is much faster without compromising accuracy. We applied our method to reconstruct the evolutionary relationships among swordtails and platyfishes ($Xiphophorus$: Poeciliidae), which is characterized by widespread hybridizations.

Models

The paper models explicit phylogenetic networks with reticulation events, inheritance probabilities, and incomplete lineage sorting, then estimates them using quartet-based pseudolikelihood. Identifiability depends on network structure, taxon sampling, branch lengths, and the level-1 assumption.

  • Network models: Explicit networks represent ancestral species and reticulation events, unlike implicit split networks, whose internal nodes lack direct biological interpretation.The model treats hybridization, introgression, and horizontal gene transfer as network processes that are not distinguishable without additional biological information.
  • Network models: A rooted network is a directed acyclic graph with tree nodes having one parent and two children, and hybrid nodes having two parents and one child.Tree edges lead to tree nodes, whereas hybrid edges lead to hybrid nodes.
  • Network models: The model uses n taxa, h hybridization events, cycle sizes k_i, coalescent branch lengths t, and inheritance probabilities γ describing parental gene contributions.Only identifiable branch lengths are estimated; with one sequenced individual per taxon, external-edge lengths are not identifiable.
  • Quartet pseudolikelihood: Quartet concordance factors are genomic support proportions for the three possible unrooted quartets on each four-taxon subset, and network CFs are weighted averages over underlying species trees.Under the coalescent, quartet CFs are independent of root placement, allowing semi-directed network inference.
  • Identifiability: Identifiability improves with larger reticulation cycles and sufficient taxon sampling, but four taxa often cannot separate parameters or distinguish hybrid-node placement.For k ≥4, hybridizations can generally be detected with n ≥5, while direction can become identifiable when n ≥5; the method assumes level-1 networks.
  • Quartet pseudolikelihood: SNaQ maximizes a pseudolikelihood built from four-taxon subnetworks, estimating topology, branch lengths, and inheritance probabilities through heuristic network search.The method uses observed quartet CFs from multi-locus data and searches networks with bounded numbers of hybridizations.

Results

SNaQ provided accurate phylogenetic-network inference while substantially improving scalability over PhyloNet, and recovered major vertical patterns even when reticulations were difficult to resolve. Applied to Xiphophorus, it supported two reticulations and widespread uncertainty in some placements.

  • Simulated data: SNaQ's pseudolikelihood approach improved speed while retaining substantial accuracy relative to PhyloNet in simulations.PhyloNet was too slow for larger scenarios, preventing accuracy comparisons there.
  • Simulated data: For n = 10 and 300 loci, a single PhyloNet replicate required over 400 hours, and PhyloNet could not run for n = 15.
  • Simulated data: With h ≥2, SNaQ's directed-network accuracy decreased, but the unrooted topology was still correctly inferred in most replicates.For n ≥10, estimated gene trees particularly degraded hybrid-edge direction.
  • Simulated data: The major tree was correctly estimated from 300 or more genes in nearly all scenarios, even when minor or major hybrid-edge assignments were incorrect.The major tree captures the network's backbone, or major vertical inheritance pattern.
  • Simulated data: Hybridization detection depended on network geometry: good diamonds were recovered from 100 genes, whereas bad diamonds and directed k = 4 cycles required substantially more evidence.Undirected cycles could remain accurately recovered even when hybrid-edge direction was difficult to infer.
  • Xiphophorus fishes evolution: In 1183 Xiphophorus transcripts, h = 2 best fit the data, with reticulations involving X. xiphidium (γ = 0.17) and X. nezahuacoyotl (γ = 0.20).The latter was associated with high concordance-factor support for a clade uniting X. nezahuacoyotl and the nigrensis group.
  • Xiphophorus fishes evolution: Bootstrap analyses supported most of the inferred tree and the X. xiphidium reticulation, but gave split donor-lineage support for the X. nezahuacoyotl reticulation.One alternative donor lineage received 75% support.

Discussion

The discussion positions quartet-based pseudolikelihood as a scalable, biologically interpretable approach while identifying restrictive assumptions, identifiability limits, and unresolved model-selection and missing-data issues.

  • SNaQ extends statistical network inference to multi-locus data while retaining an evolutionary model and biological interpretability.
  • The method assumes level-1 networks, excluding intersecting cycles and requiring further work for more complex network searches.
  • Unrooted quartets lack identifiability when n < 5, while long branches can obscure gene-flow direction and hybrid-node placement even with more taxa.
  • The major tree drops hybrid edges with inheritance γ < 0.5 and can require fewer genes to recover than the horizontal signal.
  • Missing 4-taxon sets may be necessary for large networks because their number grows approximately as n^4/24.
  • Information criteria are inappropriate for selecting h under pseudolikelihood, and quartet dependence leaves comparison theory incomplete.

A Quartet CFs under the coalescent with hybridization

The section formulates quartet concordance factors for four-taxon networks with one hybridization and uses subnetwork equivalence to simplify their representation.

  • For a level-1 four-taxon network with one hybridization, quartet CFs are expressed as inheritance-weighted combinations of coalescent branch-length terms.
  • The transformation z_i = 1 − exp(−t_i) simplifies the quartet CF formulas for network subnetworks.
  • Subnetworks with type 1 produce two equal minor CFs and are therefore equivalent, in quartet CFs, to subnetworks without hybridization.

A.3 Four-taxon networks with more than one hybridization

Multiple hybridizations in level-1 four-taxon networks can be reduced to equivalent networks with zero or one hybridization by simplifying network blobs.

  • Hybridizations of type 1 can be removed after pruning taxa when the resulting subnetworks retain the relevant quartet-CF equivalence.
  • Figure S1 illustrates networks with multiple type-1 hybridizations that are equivalent to four-taxon trees in expected CFs.

B.1 5-taxon network: topology identifiability

The section analyzes whether quartet concordance factors distinguish a five-taxon network with one hybridization from its corresponding tree and establishes generic identifiability under stated conditions.

  • A five-taxon network yields 5 four-taxon subsets, each with 3 possible quartets, producing 15 quartet CF equations.
  • Topology identifiability is tested by whether one set of quartet CF values can simultaneously satisfy the network and tree equation systems.
  • Regardless of tree branch lengths u1 and u2, the example five-taxon network is identifiable when t0 > 0, t1 > 0, t11 < ∞, and γ ∈ (0, 1).
  • Across possible five-taxon networks with one hybridization, generic identifiability holds when every tree edge has ti > 0, γ ∈ (0, 1), and the hybridization cycle has k ≥ 3 nodes.

B.2 n-taxon network: topology identifiability

Topology identifiability depends on the hybridization cycle and the taxon sampling across its subnetworks. Quartet-based conditions can determine when hybridizations are detectable or indistinguishable from trees.

  • Identifiability conditions: For k = 2, the network is not identifiable from a tree, while larger cycles require conditions on branch lengths and inheritance probabilities.The displayed cases specify parameter constraints for k = 3, 4, and 5.
  • Quartet information: Identifiability can be studied using quartets containing at most two taxa from each subnetwork.With one sampled tip, hybridizations within that subnetwork do not affect quartet concordance factors.
  • Cycle-based characterization: The analysis characterizes n-taxon networks by the number of nodes k in their hybridization cycle.Figure S3 considers k = 2, 3, 4, and 5 from left to right.
  • Identifiability target: The study asks whether branch lengths and heritabilities are identifiable when the network topology is fixed.

C.1 Identifiability of parameters in a 5-taxon network

The 5-taxon analysis expresses quartet concordance factors as a constrained equation system. Network structure imposes invariants, and insufficient independent equations can prevent identification of all parameters.

  • Equation system: A 5-taxon network yields 15 quartet concordance-factor equations in branch lengths and inheritance parameter γ.The three concordance factors for each 4-taxon set must sum to 1.
  • Invariants: Network structure imposes consistency invariants, including equalities among concordance factors from different quartet sets.For example, the displayed equations require c1 = c4, c2 = c5, and c3 = c6.
  • Tree comparison: For a species tree, 13 independent invariants leave two algebraically independent equations for two branch-length parameters.These equations can therefore be solved for the two unknown parameters, although uniqueness is not established here.
  • Network identifiability: The Fig. S2 network has four parameters but only three algebraically independent equations, producing infinitely many solutions.The parameters γ, t0, t1, and t11 cannot all be solved from the quartet concordance factors.
  • Reparametrization: Reparametrization reduces the bad diamond I concordance factors to three identifiable parameters: x, y, and t11.Here x = γ(1 − exp(−t0)) and y = (1 − γ)(1 − exp(−t1)).

C.2 n-taxon network: parameter identifiability when h = 1

For a single-hybridization network, parameter identifiability varies with cycle size and taxon distribution. Some configurations are fully identifiable, while others require fixing or reparametrizing parameters.

  • Good triangle: For k = 3, parameters are not identifiable when n ≤5.When n ≥6 and every cycle subnetwork has at least two taxa, six independent equations identify only six of seven parameters.
  • Diamonds: For k = 4, all parameters are identifiable when n0 ≥2, n2 ≥2, or both n1 and n3 ≥2.The remaining configurations are bad diamonds I and II, where parameters are not all identifiable.
  • Larger cycles: For k = 5, all parameters are identifiable by combining information from quartets and previously analyzed k = 4 subnetworks.The analysis identifies cycle parameters and then uses good or bad diamond subnetworks to identify remaining branch lengths.
  • Larger cycles: For k >5, extracting k = 5 subnetworks identifies different parameter subsets that together span all parameters.
  • Scope of parameters: Branch lengths leading to one-taxon subnetworks are irrelevant to quartet concordance factors and are excluded from optimization.In bad diamond I, several external or hybrid branch lengths are therefore non-identifiable.

D Heuristic search in the space of networks

The heuristic search explores network topologies and parameters through randomized local moves while enforcing level-1 and taxon constraints. Multiple stopping criteria control the search.

  • Initialization: SNaQ can initialize searches from a user-specified network or a fast quartet-based tree estimate such as ASTRAL.Branch lengths without supplied values are initialized from average observed concordance factors spanning the corresponding branch.
  • Search strategy: The default procedure performs 10 independent searches with randomized starting perturbations before navigating network space.NNIs or hybrid-edge-origin moves are used with probability 0.7 by default.
  • Proposal moves: Search moves relocate hybrid-edge origins or targets, reverse hybrid-edge direction, apply NNIs, or add hybridizations.Moves modify topology while retaining or initializing branch lengths according to the move type.
  • Validity constraints: Every proposed network must remain level 1, respect the maximum hybridization count, and permit at least one valid root placement.Invalid proposals fail immediately and are replaced by new random moves.
  • Stopping criteria: The search stops when pseudolikelihood changes are below 0.001, pseudo-deviance is below tolerance, or failed-move limits are reached.Move-specific limits are used to avoid repeatedly proposing the same network.

E Identifiability from quartets versus triples

Unrooted quartets can distinguish networks that rooted triples cannot, because quartets include additional ingroup-only subsets. The paper demonstrates this analytically and numerically for two example networks.

  • Unrooted quartets provide more information for network identification than rooted triples.
  • Rooted triples augmented with an outgroup correspond to some unrooted quartets, while quartets also include ingroup-only subsets.
  • The example's quartet concordance factors are derived from formulas incorporating inheritance probability and branch-length parameters.
  • For Ψ1 and Ψ2, the first four quartet concordance factors are identical, but the fifth differs, making the networks distinguishable.
  • The shared concordance factors are 0.18584, 0.69153, 0.90743, and 0.914114, whereas CFAB|CD is 0.92956 for Ψ1 and 0.956292 for Ψ2.

F Simulated data

The simulations used gene trees generated on specified true networks, including examples with three hybridizations and deliberately challenging bad-diamond structures.

  • Gene trees were simulated on the true networks used in the study.
  • The simulation set included a bad diamond I with n = 6 and γ = 0.2, and a bad diamond II with n = 10 and γ = 0.3.

G Xiphophorus fish network analysis

The Xiphophorus analysis evaluates estimated networks across h = 0 to 5 hybridizations using pseudo-deviance and visual network structure. Hybrid edges indicate mixed parental ancestry, with inheritance probabilities shown for minor and major edges.

  • Pseudo-deviance scores were examined against the number of hybridizations for the Xiphophorus fish data.
  • The Xiphophorus labels include X. birchmanni and X. malinche.
  • Estimated networks were produced for h = 0 to 5, with h = 0 represented as a rooted tree and h ≥1 as semi-directed networks.
  • Hybrid edges point toward hybrid nodes representing ancestral species with mixed parental origins.
Loading 1509.06075v3…