Source-linked AI summary
Prediction, dynamics, and visualization of antigenic phenotypes of seasonal influenza viruses
Richard A. Neher, Trevor Bedford, Rodney S. Daniels, Colin A. Russell, Boris I. Shraiman
TL;DR
Seasonal influenza antigenic evolution must be tracked as viruses evade immunity, but the relationship between serological measurements and viral sequence evolution can be modeled directly. The paper maps HI antigenic changes onto phylogenetic paths and shows that HA-sequence-based models predict titers with accuracy comparable to traditional approaches, while exposing important scope limits.
Problem
Seasonal influenza viruses rapidly evade immunity, motivating models that relate antigenic differences measured by HI assays to HA sequence evolution.
Method
The paper models HI titers using titer drops along phylogenetic branches or amino acid substitutions, with sparse genetic effects and virus- and antiserum-specific terms.
Results
The models accurately predict measured HI titers and achieve similar or better accuracy than traditional cartographic approaches.
Takeaways & Limitations
Integrating HI data into sequence-derived phylogenies can support antigenic visualization, vaccine strain selection, and efforts to predict successful influenza lineages.
Takeaways & Limitations
Prediction requires relevant branches or substitutions to be constrained by training data, and tree additivity can fail when amino acid positions mutate repeatedly.
Abstract
from arXiv · showhide
Human seasonal influenza viruses evolve rapidly, enabling the virus population to evade immunity and re-infect previously infected individuals. Antigenic properties are largely determined by the surface glycoprotein hemagglutinin (HA) and amino acid substitutions at exposed epitope sites in HA mediate loss of recognition by antibodies. Here, we show that antigenic differences measured through serological assay data are well described by a sum of antigenic changes along the path connecting viruses in a phylogenetic tree. This mapping onto the tree allows prediction of antigenicity from HA sequence data alone. The mapping can further be used to make predictions about the makeup of the future seasonal influenza virus population, and we compare predictions between models with serological and sequence data. To make timely model output readily available, we developed a web browser based application that visualizes antigenic data on a continuously updated phylogeny.
Results
The paper models antigenic distance as titer drops along phylogenetic paths, while accounting for antiserum potency, virus avidity, and sparse genetic effects.
- Models: The tree model expresses logarithmic HI titers as titer drops along branches connecting test viruses with reference viruses.The substitution model instead sums effects of amino acid substitutions separating the sequences.
- Models: Antiserum potency and virus avidity absorb systematic variation across antisera and viruses, respectively.These effects raise or lower expected titers for all measurements involving the corresponding antiserum or virus.
- Inference: The model parameters are fit with a sparse cost function that favors a few large genetic effects and many zero effects.The genetic component uses ℓ1 regularization and non-negativity constraints, while virus and antiserum variation use ℓ2 regularization.
- Inference: Decomposing HI data into potency, avidity, and genetic components smooths distance relationships and assigns shared titer drops to branches or substitutions.Branch or substitution effects capture drops associated with multiple antiserum-virus pairs.
The tree and substitution models accurately predict HI titers
The tree and substitution models predict held-out HI titers from phylogenetic or sequence information, including for viruses entirely absent from training data, while revealing limits for novel clades.
- Held-out prediction: Approximately 0.5 log2 titer levels is the prediction accuracy for H3N2 antiserum-virus combinations held out from training.Accuracy is somewhat lower for influenza B lineages, and little consistent antigenic evolution is observed in A(H1N1pdm09).
- Held-out prediction: Approximately 0.75 log2 titer levels is the prediction accuracy for viruses completely absent from training data.Both models achieve this accuracy, although virus avidity cannot be estimated for these viruses.
- Held-out prediction: The increased error for unseen viruses is largely due to virus-to-virus variability not captured by the HA phylogeny.
- Prediction scope: Genetic effects require relevant branches or substitutions to be constrained by measurements in the training set.For a completely novel clade, the model predicts titers equal to the clade base for all subtending viruses.
- Prediction scope: The models provide smoothed predictions for every antiserum-virus combination in a phylogenetic tree, with confidence varying by the quality and amount of antigenic data.The tree model also predicts titers exceeding homologous titers, while unavailable avidities explain the absence of strongly negative predictions for unseen viruses.
Amino acid substitutions associated with titer drops
Antigenic changes concentrate at prominent HA sites, but inferred substitution effects vary with genetic background, amino acid direction, and substitution combinations.
- Sites and effects: Seven HA positions near the receptor-binding site, known as Koel 7, play a particularly prominent role in antigenic evolution.
- Context dependence: The number of substitutions at prominent sites predicts antigenic distance poorly, with r2 values of about 0.25.
- Sites and effects: The top inferred A(H3N2) substitution contributions are often associated with Koel 7 sites and are similar across overlapping intervals when genetic context is shared.
- Cluster transitions: N145K contributes 1.95 antigenic units in the Sichuan/87-to-Beijing/89 transition, with minor substitutions bringing the estimate close to a 3.9-unit map distance.Each antigenic unit equals a two-fold decrease in HI titer.
- Context dependence: Substitution effects depend strongly on genetic background and direction: N145K and K145N have large inferred effects in different periods, whereas reverse changes can be small.K158R is associated with a 2 unit titer drop without reaching high frequencies.
Cumulative antigenic evolution
Cumulative antigenic change is estimated by summing branch contributions from the tree root to its leaves, revealing faster drift in H3N2 than in influenza B lineages.
- Rates: Approximately 0.7 log2 titer units per year is the antigenic advance estimated for H3N2 viruses.
- Rates: Influenza B advances 0.16 units per year in the Vic lineage and 0.09 units per year in the Yam lineage.
- Effect distribution: Approximately half of cumulative A(H3N2) antigenic evolution comes from many substitutions with effects smaller than one unit.
- Effect distribution: 20% of cumulative A(H3N2) antigenic evolution is accounted for by a few substitutions with effects above two units.Some large effects cannot be individually resolved because substitutions are colinear.
Limitations of the models
The models perform similarly overall, but each has failure modes tied to how antigenic effects are represented and constrained by available data.
- The tree and substitution models have similar accuracy, but one can outperform the other in particular circumstances.
- The tree model assumes antigenic titers are additive along phylogenetic paths, which can fail when the same HA site mutates independently in different branches.
- The substitution model can be more accurate for independent mutations at the same site, but it can fail when the same substitution occurs in different genetic backgrounds.
- Back-in-time measurements are underrepresented, limiting support for reverse substitutions in the substitution model; the tree model avoids this issue by assuming symmetric effects.
- Model deviations usually affect isolated clades with few constraining measurements, and the visualization helps identify these inaccuracies.
Antigenic change and the success of clades
Antigenic advancement is associated with clade and substitution success, but it does not determine which lineages spread because false positives and other factors remain.
- Large antigenic changes appear to fix more often than small changes, although several antigenically evolved clades fail to spread.
- cHI-based antigenic advancement predicts a following season’s dominant lineage better than the population average, with quality comparable to LBI.
- Maximal cHI can select clades far from the future population, indicating false positives that worsen when smaller clades are included.
- Successful clades tend to be antigenically advanced, suggesting that cHI correlates with viral fitness but is not the only determinant of clade success.
- Fixation probability increases with inferred antigenic effect, but substitutions with very large effects can still fail to spread.
Visualization of antigenic evolution
The authors map antigenic measurements and predictions onto a continuously updated phylogenetic tree, enabling interactive comparison of sequence evolution, titers, and antigenic distances.
- The approach maps HI titer data onto the phylogenetic tree instead of projecting sequence evolution onto a two-dimensional antigenic map.
- nextflu tracks seasonal influenza evolution in near real time and integrates tree and sequence titer models into the augur processing pipeline.
- Tree tips are colored by predicted log2 titer distance from a selectable focal reference virus, while crosses identify vaccine strains.
- Users can inspect all available measurements and predictions from both models in tooltips and recolor the tree by selecting reference viruses.
- Measured titers and predictions can be toggled to reveal titer noise and model inaccuracies, while predictions provide de-noised antigenic distances.
- The tree can also be colored by antigenic evolution accumulated along branches from the root, analogous to one dimension of antigenic cartography.
Discussion
The tree-based and substitution-based models predict influenza antigenicity from HA sequence context, with accuracy similar to or better than traditional cartography. HI data integrates naturally with phylogeny, but false positives limit its predictive use.
- Tree-based and substitution-based models predict HI titers with similar or better accuracy than traditional antigenic cartography.
- At most n −2 internal branches define tree-like pairwise distances among n viruses, while polytomies typically reduce the number of relevant branches.For A(H3N2), HI data covered 1796 of 2473 viruses and 622 internal branches constrained by HI data.
- H3N2 shows fast antigenic drift, whereas both influenza B lineages show slower drift; the method generally estimates lower drift, especially for influenza B.
- HI-based prediction produces substantial false positives at 1% clade frequency, although fewer occur at 5%.Some substitutions with large antigenic effects fail to spread, and antigenically distant clades can die out.
- Genetic background dependence of large-effect HA substitutions is identified as a route toward reducing false positives and improving prediction.The discussion links this dependence to possible fitness costs and compensatory stabilizing mutations.
- HI data integrates naturally onto sequence-derived phylogeny, while current HI-based prediction does not outperform sequence-based methods.
- Placing HI data in genealogical context may help optimize targeted acquisition of HI measurements.
Data
The study combines human influenza HA sequences with HI measurements collected from publications and WHO reports spanning 2002–2015. Despite changes in assay modalities, the model describes the multiyear data with reasonable accuracy.
- HA sequences from human influenza A and B viruses were downloaded from GISAID, while HI data came from publications and WHO CC London reports from 2002 to 2015.HI data collected before 2011 was curated in Bedford et al. (2014).
- Reasonable accuracy was maintained across years despite changes in red blood cells, neuraminidase inhibitors, and other HI assay modalities.Differences in assay methodology can largely be absorbed by antiserum potency and virus avidity terms.
Data processing and model fit
The analysis fits tree and substitution models to HI measurements after sequence alignment and phylogenetic reconstruction. A constrained, regularized quadratic optimization estimates branch effects, virus avidities, and antiserum potencies.
- The pipeline subsamples viruses, aligns sequences, builds a phylogenetic tree, retains all antigenically measured strains, and then fits tree and substitution models to HI data.
- Antigenic distance is the log2 difference between homologous and test-virus HI titers, and tree distance sums branch effects along the virus–antiserum path.
- ℓ1 regularization makes branch effects sparse, while ℓ2 regularization penalizes large antiserum and virus avidity values without enforcing sparsity.
- The model has S antiserum potencies, V virus avidities, and internal tree-branch parameters, with typically only a fraction of branches receiving nonzero effects.
- The optimization is formulated as a canonical quadratic program with inequality constraints on the unknown parameter vector.The matrix Q and vector q specify the cost function, while G and h encode the constraints.
- The model vector contains branch titer drops, antiserum potencies, and virus avidities, with each measurement represented by a row of a binary matrix A.A row marks path branches, virus avidity, and antiserum potency for the corresponding titer prediction.
- Non-negativity constraints force branch titer drops to be positive, and the fitted quadratic programming problem is solved with cvxopt.
Visualization
The browser-based visualization extends nextflu and auspice to display HI titers and inferred antigenic effects on continuously updated influenza phylogenies.
- The visualization adds HI-titer coloring and tooltips to the standard nextflu tree display.
- Virus-lineage pages show structural positions where the substitution model inferred large contributions to antigenic change.Structures are visualized with JSmol, using 5HMG for H3N2 and 4LXV for H1N1pdm.