Source-linked AI summary

Predicting evolution from the shape of genealogical trees

Richard A. Neher, Colin A. Russell, Boris I. Shraiman

arXiv:1406.0789v2q-bio.PE

TL;DR

Whether reconstructed genealogical-tree shape contains information about sampled individuals’ relative fitness and future population composition remains a central predictive question. The authors infer fitness from genealogical branching patterns and rank sequences using a local branching index that summarizes nearby tree length. Sequences ranked by LBI predict influenza A/H3N2 progenitor lineages with high accuracy.

  • Problem

    Whether reconstructed genealogical-tree shape contains information about sampled individuals’ relative fitness and future population composition remains a central predictive question.

  • Method

    The authors infer fitness from genealogical branching patterns and rank sequences using a local branching index that summarizes nearby tree length.

  • Results

    Sequences ranked by LBI predict influenza A/H3N2 progenitor lineages with high accuracy.

  • Takeaways & Limitations

    Meaningful predictions from influenza genealogies imply persistent fitness variation in circulating A/H3N2 populations.

  • Takeaways & Limitations

    Predictions are often suboptimal during antigenic-cluster transitions driven by large-effect mutations.

Abstract

from arXiv · show

Given a sample of genome sequences from an asexual population, can one predict its evolutionary future? Here we demonstrate that the branching patterns of reconstructed genealogical trees contains information about the relative fitness of the sampled sequences and that this information can be used to predict successful strains. Our approach is based on the assumption that evolution proceeds by accumulation of small effect mutations, does not require species specific input and can be applied to any asexual population under persistent selection pressure. We demonstrate its performance using historical data on seasonal influenza A/H3N2 virus. We predict the progenitor lineage of the upcoming influenza season with near optimal performance in 30% of cases and make informative predictions in 16 out of 19 years. Beyond providing a tool for prediction, our ability to make informative predictions implies persistent fitness variation among circulating influenza A/H3N2 viruses.

Results · The fitness distribution on a tree · Fitness inference is insensitive to model assumptions

The study develops a tree-based fitness inference method using selection-biased diffusion and message passing, then tests its robustness when evolution occurs through discrete mutations. In simulations, fitness rankings correlate with true fitness and improve as mutation rates increase, with only weak dependence on model parameter Γ.

  • The fitness distribution on a tree: The method infers node fitness by combining branchwise propagators with message passing across the reconstructed genealogical tree.It calculates marginal fitness distributions for internal and external nodes by propagating information up and down the tree.
  • The fitness distribution on a tree: The joint fitness distribution assigns relative fitness to ancestral and sampled nodes and factorizes into propagators along tree branches, with normalization by Z(T).Each node’s fitness is measured relative to the population mean at its sampling time, and the structure parallels phylogenetic likelihood calculations.
  • The fitness distribution on a tree: Branch propagators retain ancestral-fitness information over short intervals but approach the population distribution over long intervals.Backward-time inference uses Bayesian inversion to estimate ancestral fitness from descendants.
  • The fitness distribution on a tree: The model treats fitness evolution as a selection-biased random walk, approximated by diffusion when many mutations contribute to fitness.Selection biases lineage survival toward higher fitness, while the branch fitness diffusion constant captures stochastic fitness changes.
  • Results: Spearman’s correlation coefficients around 0.5 show that inferred fitness rankings are well predicted, and rankings improve with increasing mutation rates.This result is shown for a typical simulation and is consistent across adaptive-mutation rates, while depending only weakly on Γ.
  • Fitness inference is insensitive to model assumptions: Simulations tested the infinitesimal-mutation SBD model against discrete-mutation evolution with fixed fitness variance and beneficial mutation rates nA = 0.02, . . . , 0.16 per generation.The simulated genome otherwise contained predominantly deleterious mutations, providing a test of model robustness under adaptive evolution in a changing environment.
  • Fitness inference is insensitive to model assumptions: Using Γ = 0.2 and 0.5, the analysis found that larger Γ performs better at low mutation rates when fitness diversity is dominated by only a few mutations.Increasing mutation rates improves the SBD approximation, whereas the choice of Γ has only a weak overall effect; Γ represents the relative importance of stochastic diffusion and selection.

High inferred fitness predicts progenitor sequences · Local branching density as a heuristic ranking

Sequences assigned high inferred fitness tend to identify progenitor lineages of future populations, while a model-independent local branching index provides a nearly comparable ranking using surrounding tree length. The heuristic is relatively insensitive to inference parameters and captures branching patterns associated with gradually changing fitness.

  • High inferred fitness predicts progenitor sequences: High-fitness predictions identify future progenitor sequences in simulations, with Fig. 2D evaluating their normalized sequence distance against the post-hoc optimal pick.The comparison uses the Hamming distance between the highest-fitness prediction and the population 200 generations later, normalized by the average present-to-future distance.
  • High inferred fitness predicts progenitor sequences: Using sequences sampled across a 100-generation interval produces highly similar fitness-inference results to sampling 200 sequences from one generation.
  • Local branching density as a heuristic ranking: Fitness ranking and progenitor prediction depend little on Γ and ω/σ, even though faithful posterior-fitness inference requires numerical branch propagation and parameter knowledge.This parameter insensitivity suggests that ranking is driven mainly by a more universal tree quantity.
  • Local branching density as a heuristic ranking: For short time periods, internal-node fitness estimates increase with downstream branch length, because subtree length polarizes inferred fitness toward high values.For a fixed number of descendants, star-like subtrees maximize length and indicate rapid branching or multiple mergers.
  • Local branching density as a heuristic ranking: Because fitness changes gradually along lineages, high-fitness nodes are expected to show both upstream and downstream branching within a neighborhood whose size depends on fitness decorrelation.
  • Local branching density as a heuristic ranking: The local branching index λi(τ) ranks internal and terminal nodes by exponentially discounted surrounding tree length, with τ setting the neighborhood scale.Within the SBD model, τ corresponds to the equilibration timescale of high-fitness lineages, of order Tc/√log N.
  • Local branching density as a heuristic ranking: LBI rankings are almost as accurate as complex SBD fitness inference, as Fig. 3 compares Spearman correlation with true fitness across pairwise diversity and memory timescales.The LBI can be computed efficiently using the same message-passing techniques used for posterior fitness inference.

Prediction of seasonal influenza A/H3N2 progenitor lineages

The authors used genealogical trees of seasonal influenza A/H3N2 HA1 sequences from 1995–2013 to predict progenitors of the following northern-hemisphere winter season. Their rankings identified future circulating lineages with variable accuracy and performed comparably to the influenza-specific predictor of Łuksza and L¨assig (2014).

  • Prediction setup: The method analyzed up to 100 HA1 sequences per region from May–February to predict the closest relative circulating the following October–March season.Samples covered Asia and North America from 1995 to 2013, using publicly available Influenza Research Database sequences and maximum-likelihood trees built with FastTree.
  • Prediction performance: The highest-ranked internal nodes predicted 1997–1999, 2003, 2006–2009, and 2013 reasonably well, but failed in 1995, 1996, and 2002.Predictions had intermediate accuracy in the remaining years; highest-ranked external nodes performed similarly except in 1997.
  • Comparison with existing predictors: Across years, internal-node ranking was comparable to Łuksza and L¨assig (2014), whereas external-node ranking was slightly worse.The two approaches nevertheless produced very similar year-to-year predictions, potentially because their model uses downstream synonymous mutations to capture epistatic interactions.
  • Evaluation metric: Prediction quality was quantified by normalized genetic distance, where d = 0 denotes an optimal prediction and d = 1 denotes a random pick.The authors averaged d across years and bootstrapped years to compare internal- and external-node rankings with Łuksza and L¨assig (2014) and naive predictors.

Inferred fitness increases are associated with epitope mutations … Derivation of the fitness inference algorithm

The method infers relative fitness from genealogical branching patterns under a small-effect mutation model and uses this information to predict future populations. High inferred fitness is associated with epitope substitutions, while predictive performance persists even when genetic diversity is low.

  • Inferred fitness increases are associated with epitope mutations: Branches in the top quartile of inferred fitness increases are enriched for nonsynonymous substitutions, with restriction to epitopes A–D increasing enrichment to approximately 2-fold.Further restriction to the seven Koel loci increases enrichment slightly, although their small number limits power to detect additional enrichment.
  • Discussion: The association between high inferred fitness and epitope substitutions is consistent with antigenic novelty driving influenza evolution, despite the model being agnostic to sequence and protein structure.The relevant epitopes historically have high d_n/d_s, suggesting positive selection.
  • Discussion: The algorithm models adaptive evolution as fitness dynamics on genealogical trees and infers individual-node fitness probabilistically; simulations show that the highest-fitness sequence tends to match the future progenitor.The framework uses a selection-biased diffusion model in which evolution proceeds through many small-effect mutations.
  • Discussion: Predictive power increases with nonneutral genetic diversity but remains at rather low pairwise distances, even where the diffusion model is a poor approximation.This persistence supports a relationship between fitness and genealogical structure beyond the model’s best-fitting regime.
  • Discussion: Prediction is limited in years involving antigenic-cluster transitions, when specific large-effect mutations cause drastic antigenic changes; sampling, migration, and demographic structure may also perturb genealogical patterns.These factors can hamper prediction even though reconstructed influenza genealogies remain informative.
  • Derivation of the fitness inference algorithm: The branching patterns of reconstructed genealogies contain information about sampled individuals’ relative fitness that can predict future population composition when fitness differences depend on multiple mutations.The algorithm requires only a reconstructed genealogy as input and is intended for applications ranging from RNA viruses to cancer cell populations.
  • Derivation of the fitness inference algorithm: The inference derives from a branching-process approximation: offspring sampling probabilities yield numerically solved branch propagators, which combine into the posterior fitness distribution.The method operates on a static sequence set from one time point and uses influenza history only for validation.

Offspring number distributions

The paper models offspring counts through a backward master equation for P(n|x,t), incorporating birth, death, mutation, and environmental deterioration. Under short-tailed mutational effects and high mutation rate, the generating function and reproductive value quantify lineage growth before coalescence.

  • Offspring number distributions: The offspring distribution P(n|x,t) is derived from a backward master equation incorporating birth at rate 1+x, death at rate one, mutation, and environmental deterioration at velocity v.The mutation contribution averages over fitness effects μ(s) with total rate u, while fitness increases backward by Δtv because the environment deteriorates forward in time.
  • Offspring number distributions: Assuming short-tailed fitness effects and mutation rate u large relative to typical effects, mutation dynamics reduce to directional and diffusive fitness changes characterized by D = u⟨s^2⟩/2 and σ^2 = v − u⟨s⟩.The generating-function equation is solved numerically to approximate fitness distributions on genealogical trees.
  • Offspring number distributions: The reproductive value R(x,t), defined as the expected offspring number after t generations, is obtained from the generating function and satisfies a linear equation.This approximation is valid only for times short compared with the coalescence time Tc.
  • Offspring number distributions: Initially, a lineage grows clonally at rate x, growth slows as the rest of the population adapts at rate σ^2, and descendant mutations alter offspring fitness.These effects determine lineage dynamics before coalescence.

Lineage sampling probability

The generating function φ_ω(x,t) gives the probability that a lineage is represented in a sample, with asymptotic forms that are accurate when φ_ω is small or saturated. The approximation satisfies the initial condition, long-time behavior, and neutral limit.

  • Lineage sampling probability: φ_ω(x,t) represents the probability that a lineage is included in a sample of fraction ω = M/N.Here, M is the sample size and N is the population size.
  • Lineage sampling probability: Sampling probability is obtained by summing the probability that none of n offspring are sampled, then subtracting that sum from 1.Each term (1 − ω)^n is the probability that none of the n offspring enter the sample.
  • Lineage sampling probability: The approximation is accurate when φ_ω is small or when sufficiently large x causes saturation, φ_ω(x,t) ≈ x.It also satisfies φ_ω(x,0) = ω, approaches x for x > 0 at long times, and recovers φ_ω(0,t) = ω/(1 + ωt) in the neutral limit.

Branch propagator

The branch propagator describes the fitness distribution of descendants conditioned on an unbranched lineage contributing to the sample. Its dynamics incorporate mutation-driven fitness diffusion, lineage reproduction, and sampling, yielding characteristic ancestral and terminal-progenitor distributions.

  • Governing equation: The propagator equation combines reproductive growth, sampling-dependent non-branching, mutation-driven fitness drift, and fitness diffusion, with a delta-function initial condition.The derivation treats birth events whose unsampled branch does not contribute to the present sample and shifts mean fitness through time.
  • Assumption: The small-fitness-difference assumption is appropriate for populations with small fitness variance, and violating it changes the propagator’s quantitative rather than qualitative behavior.The derivation assumes y ≪1, corresponding to small σ and small fitness differences between generations.
  • Terminal branch propagator: For positive-fitness ancestors, terminal progenitor probability initially increases with ancestral age but eventually decreases as lineage survival becomes unlikely.For short times and moderate parental fitness, the terminal propagator simplifies to the reproductive value.
  • Numerical behavior: Numerical solutions show that descendant fitness distributions broaden over time, while ancestral fitness distributions converge far in the past to a common curve favoring fit ancestors.Fit ancestors are more likely to be sampled because they leave more offspring, whereas excessively fit ancestors are opposed by lineage non-survival.

Tree-based inference

The method infers ancestral fitness by factorizing fitness probabilities over a genealogical tree and marginalizing them through iterative message passing. Recent nodes with greater downstream tree length are inferred toward higher fitness, with star-like topologies indicating rapid expansion of exceptionally fit clones.

  • Tree-based inference: The joint ancestral-fitness distribution factorizes over tree nodes and branches, allowing polytomies under unconstrained population size and noninteracting branches.These approximations are considered appropriate when selection dominates because coalescent properties depend only weakly on the stated constraints.
  • Tree-based inference: Iterative message passing integrates out leaf and ancestral fitness variables, combining upstream and downstream branch messages to obtain each node’s marginal fitness distribution.The method computes marginal distributions by multiplying all messages entering a node and normalizing the result.
  • Tree-based inference: Mean marginal fitness is used to rank internal and external nodes, while downstream total tree length polarizes recent nodes toward the high-fitness edge.For a fixed number of descendants, downstream tree length is maximized by a star topology.
  • Tree-based inference: Star-like genealogies are associated with rapid expansion of clones founded by exceptionally fit individuals.The inference links multiple mergers and expanded downstream tree length to exceptionally high founder fitness.

Calculating the local branching index (LBI)

The local branching index (LBI) is computed as an exponentially discounted tree length surrounding each node. Its calculation uses message passing, combining upward messages from children with downward messages from parents before evaluating the node’s discounted tree length.

  • Calculating the local branching index (LBI): LBI measures the integrated exponentially discounted tree length surrounding a node.The calculation parallels the message-passing framework used to evaluate fitness distributions.
  • Calculating the local branching index (LBI): Up-messages propagate information from each node’s children toward its parent using branch lengths.For node i, the expression depends on branch length b_i and a sum over its children ij.
  • Calculating the local branching index (LBI): Down-messages propagate information from each parent to its child, completing the message-passing calculation.After all up and down messages are calculated, they determine the exponentially discounted tree length.

Implementation of the inference algorithm

The inference algorithm uses message passing on a discretized fitness grid to estimate node-fitness marginals from reconstructed trees. Predictions rank internal and external nodes by expected fitness, with branch lengths converted to model time units and clade expansion rates estimated separately from longitudinal frequencies.

  • Implementation: Message passing computes marginal fitness distributions for every external and internal tree node on a discrete fitness grid implemented in Python with SciPy and NumPy.The implementation uses the `survival_gen_func` and `fitness_inference` classes.
  • Prediction procedure: The algorithm builds a maximum-likelihood FastTree tree, passes it to fitness inference, and predicts the highest-ranked node by expected fitness.FastTree was modified to resolve short branches better.
  • Model parameterization: Branch propagation is parameterized by fitness diffusion D, fitness standard deviation σ, and sampling fraction ω, with time measured in σ−1 units and selection strength in σ units.The dimensional diffusion constant is Γ = Dσ−3, and the generating-function initial condition is φω(x, 0) = ω/σ.
  • Time calibration: Branch lengths are converted from nucleotide distance to σ−1 time units using π ≈2µ⟨T2⟩ and, for the SBD model, ⟨T2⟩σ ≈Γ−1.The conversion factor β depends on the chosen Γ, mutation rate µ, and average pair coalescent time.
  • Frequency analysis: Clade expansion rates are estimated by fitting a line to log frequencies measured across three equal May–February intervals, using a pseudocount of 5.The frequencies quantify the fraction of sequences below each internal node in each interval.

Simulations · Influenza data

The study combines individual-based simulations with processed influenza A/H3N2 HA1 sequences to investigate genealogical evolution under mutation and selection. Simulations vary genomic mutation rates, while influenza data span human-host viruses collected from 1968 to 2014.

  • Simulations: Simulations use FFPopSim for individual-based populations with fixed fitness variance σ = 0.03.The simulations introduce mutations at random sites in random individuals at rate µ.
  • Simulations: The total genomic mutation rate u = Lµ is varied from 0.016 to 0.256 across L = 2000 simulated sites.This parameterization varies mutation input while keeping the simulated genome length fixed.
  • Simulations: Mutations are defaulted to deleterious effects drawn from an exponential distribution.The simulation passage also describes an environment-changing setup, but the supplied text ends before its details.
  • Influenza data: The influenza dataset contains human-host A/H3N2 sequences covering the entire HA1 domain from 1968 to 2014.Sequences were downloaded from IRD and aligned using its default alignment feature.
  • Influenza data: The HA1 alignment was manually inspected and trimmed to the HA1 domain.This preprocessing step followed alignment with IRD’s default settings.
  • Influenza data: The dataset excludes obvious outliers, laboratory strains, sequences with indels, and sequences containing more than 4 ambiguous nucleotides.These exclusions were performed manually after alignment inspection.

Appendix A: Figure 2 – supplements · Appendix B: Figure 3 – supplements · Appendix D: Figure 4 – supplements

The supplements show that prediction quality depends on genetic diversity and memory scale, while high LBI identifies sequences and clades associated with future population success. LBI predictions are often similar to, and sometimes better than, an alternative forecasting method.

  • Appendix A: Figure 2 – supplements: Prediction rank correlation increases with pairwise diversity, with larger Γ favored at small distances and smaller Γ at large distances.The small-distance regime corresponds to few large-effect mutations, whereas large distances spread fitness variation across many loci.
  • Appendix A: Figure 2 – supplements: Continuous sampling does not reduce rank correlation at moderate or large mutation rates, and predicted-strain distance to the future population behaves similarly.This supplement uses 200 simulated sequences sampled over 100 generations rather than from one time point.
  • Appendix B: Figure 3 – supplements: Sequences with the highest LBI tend to lie close to the progenitor of populations 200 generations later.The supplement measures predicted-sequence distance relative to the average distance between the sampled and future populations.
  • Appendix B: Figure 3 – supplements: LBI-based predictions vary as the memory time scale τ changes from 2^-6 to 4, with separate trajectories for internal and external nodes.The supplement examines prediction variation across years and memory scales.
  • Appendix B: Figure 3 – supplements: In many years, the highest-LBI sequence is very similar to the sequence predicted by Łuksza and L¨assig (2014), while either method is closer to the future in other years.Łuksza and L¨assig aimed to minimize amino-acid distance at epitope positions, whereas this comparison concerns nucleotide distance.
  • Appendix B: Figure 3 – supplements: High LBI predicts clade expansion, as clades ranked highly by LBI show excess representation among clades that grow into the following year.The analysis covers clades below 75% frequency in samples collected from May to February.
  • Appendix D: Figure 4 – supplements: Influenza A/H3N2 prediction accuracy based on LBI improves as the memory time scale τ increases.Accuracy is measured by nucleotide distance to the future sample, scaled so the optimal pick has d = 0 and a random pick has d = 1.
Loading 1406.0789v2…