Source-linked AI summary

Quantification of the effect of mutations using a global probability model of natural sequence variation

Thomas A. Hopf, John B. Ingraham, Frank J. Poelwijk, Michael Springer, Chris Sander, Debora S. Marks

arXiv:1510.04612v1q-bio.BM

TL;DR

The paper addresses the challenge of predicting genetic-variation effects when systematic mutation assays leave much protein sequence space unexplored. It models correlated amino-acid variation in protein families with a statistical energy that includes residue interactions, and finds reasonably accurate predictions for several protein mutational datasets. The approach is particularly informative when experimental selection resembles the evolutionary pressure represented in the sequences.

  • Problem

    Systematic functional assays provide useful mutation data, but they cover only a small fraction of potential mutational space, motivating use of natural sequence variation.

  • Method

    The paper uses sequence-only probabilistic models of protein families, including residue interactions, to quantify mutation effects through statistical-energy differences.

  • Results

    The models reasonably accurately predict experimentally determined effects across several proteins, including higher-order mutations and influenza nucleoprotein and hemagglutinin variants.

  • Takeaways & Limitations

    Natural sequence variation can support quantitative prediction of mutation effects and analysis of residue interactions in protein function.

  • Takeaways & Limitations

    Inferring a model with all pairwise residue interactions presents a limitation, and correlations decrease when a large fraction of mutants becomes non-viable.

Abstract

from arXiv · show

Modern biomedicine is challenged to predict the effects of genetic variation. Systematic functional assays of point mutants of proteins have provided valuable empirical information, but vast regions of sequence space remain unexplored. Fortunately, the mutation-selection process of natural evolution has recorded rich information in the diversity of natural protein sequences. Here, building on probabilistic models for correlated amino-acid substitutions that have been successfully applied to determine the three-dimensional structures of proteins, we present a statistical approach for quantifying the contribution of residues and their interactions to protein function, using a statistical energy, the evolutionary Hamiltonian. We find that these probability models predict the experimental effects of mutations with reasonable accuracy for a number of proteins, especially where the selective pressure is similar to the evolutionary pressure on the protein, such as antibiotics.

Introduction

The paper develops a sequence-only statistical approach for predicting mutation effects by modeling residue interactions in protein families. It addresses limited experimental coverage and reports reasonably accurate predictions across several protein mutational datasets.

  • Motivation: Systematic mutation assays cover only a small fraction of potential mutational space and depend strongly on the assayed functional property.Natural evolutionary variation provides a broader record of sequences that retained sufficient function.
  • Approach: The authors present a general approach that uses only sequence variation among evolutionarily related protein families to predict mutation effects through residue interactions.The model builds on statistical models of epistatic constraints previously used for protein-structure prediction.
  • Limitations of prior approaches: Previous computational approaches often assume independent positions and do not systematically model transitive inter-residue dependencies or context-dependent mutation effects.These limitations persist despite evidence that mutation effects depend on context throughout protein evolution.
  • Evaluation: The analysis compares model-derived statistical energies with experimental assays from systematic mutation scans across 13 protein families.The tests examine fitness effects, thermostability, and higher-order mutations, including influenza proteins.
  • Approach: The model estimates statistical energies for protein sequences and uses energy differences between wild-type and mutant sequences to quantify mutation effects.The evolutionary sequence model includes site-specific biases and pairwise epistatic constraints.
  • Results: The approach reasonably accurately predicts experimentally determined higher-order mutation effects and effects of thousands of influenza nucleoprotein and hemagglutinin variants.The results also indicate that across-species variation can inform understanding of human influenza protein evolution.

Probability model for sequences in a protein family

The EVH models protein-family sequence probability distributions using position-specific amino-acid preferences and pairwise residue interactions. Statistical energy differences between mutant and wild-type sequences quantify mutation effects while incorporating sequence context and epistasis.

  • The model represents each protein family as a sequence-space distribution constrained by single-position amino-acid preferences and pairwise amino-acid combinations.
  • The statistical energy combines single-residue fields and pairwise couplings across residue positions.
  • The evolutionary Hamiltonian, or EVH, is the model’s statistical energy, and its parameters are estimated from protein-family sequence data using penalized pseudo-likelihood.
  • Mutation effects are quantified as log-odds or statistical-energy differences between mutant and wild-type sequences.
  • Explicit interaction energies incorporate the sequence context of mutations and thereby capture epistatic and cooperative effects.

Predictions of effects of mutations

EVH predictions correlate with experimentally measured effects across several proteins and assay types. The model also evaluates unobserved combinations of mutations and evolutionary paths through sequence space.

  • 34,745 single and double PABP mutations correlate positively with experimental effects, with r=0.61 for individual mutations and r=0.68 for average effects per site.
  • 1685 M.HaeIII mutations correlate positively with selection-experiment effects, with r=0.68 for all mutations and r=0.79 for average effects per site.
  • EVH statistical energies capture melting-temperature variation with r=0.72 for 47 Fyn SH3 mutations and r=0.71 for 22 rat trypsin-2 mutations.
  • Because EVH assigns energy to any sequence, it can predict effects for higher-order mutations and explore regions of sequence space absent from natural variation.
  • EVH predictions can support experiment design, evolutionary-dynamics studies, and preliminary estimates of mutational effects before experiments.

Especially good model predictions for particular selection experiments

Prediction accuracy depends on how closely experimental selection pressure matches the evolutionary pressure encoded by sequence variation. Correlations strengthen under relevant antibiotic selection but can decline when effects saturate.

  • Evolutionary sequence variation reflects contributions to natural selection, whereas laboratory assays may target properties that are not crucial during evolution.
  • For more than 4000 β-lactamase mutations, increasing ampicillin pressure raises correlations from r=0.15 to r=0.68 for specific effects and from r=0.26 to r=0.78 for average site effects.
  • For bacterial-kinase assays, the lowest kanamycin dose gives the best average-per-site correlation, r=0.74, while higher pressure causes saturation and weaker correlation.
  • Combining bacterial-kinase data across eight antibiotics improves correlations, possibly reflecting selection pressures experienced by the enzyme.
  • Across assays including thermostability, binding, and kinetics, predictions correlate best when experimental selection pressure is most relevant to the protein’s in vivo function.

Interactions in epistatic models essential for good predictions

Epistatic EVH models generally outperform independent models because pairwise interactions encode context-dependent sequence constraints. The advantage is especially pronounced for binding, active-site, specificity, and stability-related mutations.

  • Across the 13 analyzed protein datasets, epistatic EVH predictions generally agree more closely with experimental measurements than independent-model predictions.
  • For binding or active-site residues, the epistatic model achieves r=0.61 versus r=0.28 for the independent model.
  • Pairwise epistatic constraints increase explained melting-temperature variation by 20% for SH3 and 45% for trypsin.
  • In SH3, co-dependent substitutions form a contiguous predicted interaction network in the protein core.
  • Context dependence explains why alanine substitutions that occur elsewhere in a family can be deleterious in the target sequence background.
  • Including interactions is crucial for capturing sequence constraints and deleterious effects at functional-specificity residues.

Discussion

The model uses natural sequence variation to predict mutation effects and investigate protein-function constraints, but its accuracy is bounded by data sparsity, model complexity, and experimental noise.

  • Discussion: Sequence scarcity creates a tradeoff between model complexity and fit, motivating additional sequence data and improved regularization.The authors highlight overfitting as a concern and recommend interpreting prediction accuracy alongside experimental noise.
  • Discussion: Experimental replicate correlations of ~r = 0.7 constrain the accuracy expected from computational mutation-effect predictions.The discussion explicitly links this replicate correlation to the interpretation of prediction accuracy.
  • Discussion: The approach can study specific variants and combinations across numerous protein families using only sequence variation.The authors propose global probability models as a way to turn evolutionary sequence records into quantitative resources for connecting genotype to phenotype.
  • Discussion: Future analyses may benefit from global sequence-family models that explicitly incorporate interactions between positions.The proposed scope extends beyond proteins as genomic sequencing becomes more accessible.
  • Discussion: Mutation datasets were assembled through a literature search of quantitative high-throughput mutagenesis experiments, excluding proteins with insufficient sequence diversity.The compilation also included low-throughput measurements for SH3 and trypsin.

Inference of epistatic statistical model of sequences

The paper infers a global probability model of protein sequences with site-specific fields and pairwise couplings, then converts sequence probabilities into mutation-effect predictions and context-dependence measures.

  • Inference of epistatic statistical model of sequences: The model is a maximum-entropy distribution constrained by empirical single-site and pairwise amino-acid frequencies.It is described as a Markov random field, or 20-state Potts model.
  • Inference of epistatic statistical model of sequences: The sequence probability uses site parameters hi(σi), pair parameters Jij(σi,σj), and a partition function Z over all sequences.The distribution assigns probabilities across sequence space for a protein family.
  • Inference of epistatic statistical model of sequences: Parameters are estimated by pseudolikelihood maximization because full likelihood inference is intractable over the 20^N sequence space.L2 regularization is added to improve generalization in the under-sampled regime.
  • Inference of epistatic statistical model of sequences: Evolutionary couplings summarize pairwise constraints using Frobenius norms, average-product correction, and a significance cutoff.These couplings are compared with structural contacts using residue-distance thresholds.
  • Inference of epistatic statistical model of sequences: The model’s statistical energy differences are not directly comparable between proteins, so effects are rescaled using the mean of the 5% most deleterious predicted single mutants.This rescaling reduces the influence of single outliers.
  • Inference of epistatic statistical model of sequences: An independent model uses only site-specific amino-acid preferences, allowing comparison with the pairwise epistatic model.The difference between the two predicted effects quantifies context dependence between sites.
  • Inference of epistatic statistical model of sequences: Experimental mutation effects are classified with two-component Gaussian mixture models assigning mutants to the component with higher posterior probability.The input effects are sequencing-read enrichment ratios before and after functional selection, transformed into log-space where applicable.

Extended Data Tables

The extended data tables catalog the mutagenesis experiments and report correlations between model predictions and experimental results.

  • Extended Data Table 1 lists the set of mutagenesis experiments.
  • Extended Data Table 2 reports correlations between prediction and experiment.
  • Extended Data Table 3 reports correlations between prediction and experiment for an additional category or analysis.The table caption is truncated in the supplied passage.

Extended Data Figure Legends

The extended data figures compare evolutionary-model predictions with mutational experiments, examine context-dependent effects and residue interactions, and relate model behavior to structural and selective features.

  • Scatter plots display the quantitative relationship between evolutionary predictions and experimental results.
  • The epistatic model predicts all tested single-mutation effects using context-dependent evolutionary information.Displayed positions were covered by the sequence alignments used to infer the statistical models.
  • The model predicts compensatory mutational paths in which G177E, G177D, or L202Q enable otherwise deleterious mutations at other sites.A179M and D136R are tolerated in the backgrounds of L202Q and G177E, respectively, and may require the enabling mutation first.
  • Correspondence between predicted and experimental effects for bacterial aminoglycoside kinase depends on antibiotic selection pressure.Some mutations become detectable as deleterious in vivo only when selective pressure increases, whereas overall correlations decrease when many mutants are non-viable.
  • Predicted epistatic residue pairs largely correspond to structurally proximate residues and define a global interaction network.
  • Epistatic predictions generally agree better with experimental data than independent predictions because they account for context-dependent mutation effects.Epistatic effects also differ more strongly from independent effects for high-effect mutations and specificity-determining residues.
Loading 1510.04612v1…