Source-linked AI summary

A three-state prediction of single point mutations on protein stability changes

Emidio Capriotti, Piero Fariselli, Ivan Rossi, Rita Casadio

arXiv:0705.1490v2q-bio.BMq-bio.QM

TL;DR

Predicting protein stability changes from single-point mutations is difficult because near-zero experimental ΔΔG values are uncertain and the mutation classes are imbalanced. The paper introduces SVM predictors that classify mutations as destabilizing, neutral, or stabilizing using sequence or structure information. The sequence model reaches Q3 0.52 and <C> 0.28, while the structure model reaches Q3 0.58 and <C> 0.39.

  • Problem

    Experimental uncertainty can obscure the sign of near-zero ΔΔG values, while destabilizing mutations outnumber stabilizing ones in training data.

  • Method

    Support vector machines classify single-point mutations into destabilizing, neutral, or stabilizing classes using protein sequence or tertiary-structure information.

  • Results

    The sequence-based predictor reaches Q3 0.52 and <C> 0.28, while the structure-based predictor reaches Q3 0.58 and <C> 0.39.

  • Takeaways & Limitations

    The three-state predictors provide more detailed mutation-effect predictions, and structure-based context performs better for highly exposed residues.

  • Takeaways & Limitations

    The method relies on a thermodynamic reversibility assumption that treats reverse mutations as having equal-magnitude free-energy changes with opposite signs.

Abstract

from arXiv · show

A basic question of protein structural studies is to which extent mutations affect the stability. This question may be addressed starting from sequence and/or from structure. In proteomics and genomics studies prediction of protein stability free energy change (DDG) upon single point mutation may also help the annotation process. The experimental SSG values are affected by uncertainty as measured by standard deviations. Most of the DDG values are nearly zero (about 32% of the DDG data set ranges from -0.5 to 0.5 Kcal/mol) and both the value and sign of DDG may be either positive or negative for the same mutation blurring the relationship among mutations and expected DDG value. In order to overcome this problem we describe a new predictor that discriminates between 3 mutation classes: destabilizing mutations (DDG<-0.5 Kcal/mol), stabilizing mutations (DDG>0.5 Kcal/mol) and neutral mutations (-0.5<=DDG<=0.5 Kcal/mol). In this paper a support vector machine starting from the protein sequence or structure discriminates between stabilizing, destabilizing and neutral mutations. We rank all the possible substitutions according to a three state classification system and show that the overall accuracy of our predictor is as high as 52% when performed starting from sequence information and 58% when the protein structure is available, with a mean value correlation coefficient of 0.30 and 0.39, respectively. These values are about 20 points per cent higher than those of a random predictor.

Introduction

Predicting stability free-energy changes after single-point mutations is a key Structural Bioinformatics problem complicated by experimental uncertainty and imbalanced data. The paper addresses these issues with a three-class predictor that distinguishes destabilizing, neutral, and stabilizing mutations.

  • Motivation: Experimental uncertainty can reverse the apparent sign of ΔΔG for measurements near zero, complicating mutation-effect prediction.The training data are also intrinsically asymmetric, with destabilizing mutations outnumbering stabilizing ones.
  • Contribution: The predictor classifies mutations as destabilizing, neutral, or stabilizing instead of using only two classes.This explicitly separates mutations with negligible effects on protein stability.
  • Contribution: The method predicts free-energy changes from either protein sequence or protein structure.

Material and Methods

The study builds three-state SVM predictors from curated ProTherm mutation data, using thermodynamic reversibility to balance classes and sequence or structural inputs to encode mutation environments. Evaluation uses cross-validation and comparisons with baseline and prior predictors.

  • Dataset: The curated dataset contains 1681 single-point mutations from 58 proteins selected from reversible experimental measurements in ProTherm.
  • Classification: Mutations are labeled destabilizing below -0.5 Kcal/mol, stabilizing above 0.5 Kcal/mol, or neutral between these thresholds.The threshold choice is intended to balance the dataset and reflects reported experimental standard-error limits.
  • Data balancing: Thermodynamic reversibility augments the data with reverse mutations, assuming equal-magnitude free-energy changes with opposite signs.This addresses the asymmetric abundance of mutation classes.
  • Predictor inputs: The SVM uses a 42-value input encoding temperature, pH, the residue substitution, and either structural or sequence-neighbourhood context.The mutation is encoded across 20 residue types, while the final 20 values represent the local environment.
  • Evaluation: Performance is assessed with Q3, coverage, specificity, correlation, ROC area, and 20-fold cross-validation after sequence-similarity clustering.The predictors are also compared with SVM-BASE and mapped I-Mutant classifications.

Results and Discussion

The three-state SVM predictor performs better with structural information and distinguishes mutation effects most effectively for mutations with relevant stability changes.

  • Sequence-based Predictor: 0.52 Q3 and 0.28 mean correlation are achieved by the sequence-based predictor using a 31-residue window.Sequence context improves accuracy by 3% and mean correlation by 2% over SVM-BASE.
  • Structure-based Predictor: 0.58 Q3 and 0.39 mean correlation are achieved by the structure-based predictor using a 12 Å radius.Q3 and mean correlation increase with the reliability index.
  • Prediction analysis: Correlation coefficients are higher for destabilizing and stabilizing mutations than for neutral mutations, while environmental information improves neutral-mutation prediction.The improvement for neutral mutations is reflected in the ROC analyses.
  • Prediction analysis: Structure-based predictions generally have higher ROC areas for all three classes, with the largest improvement for neutral mutations.This pattern is also reported for the sequence-based method.
  • Comparison with previous methods: The new three-class methods improve over the earlier two-class methods by 3% sequence-based and 6% structure-based accuracy.Correlation improves by 1% and 5%, respectively; structure-based information is especially advantageous for highly exposed residues.

Conclusions

The paper concludes that reversible thermodynamic data and a neutral mutation class improve the balance and interpretability of single-point mutation stability prediction.

  • Conclusions: Thermodynamic reversibility generates a balanced dataset and makes the predictive methods intrinsically symmetric.The authors relate this symmetry to energy-based methods.
  • Conclusions: The neutral class groups mutations with ΔΔG near zero and partially prevents learning wrong associations from experimental errors.The authors suggest applying this approach to identify more stabilizing or destabilizing mutations than neutral ones.

Authors' contributions

The authors divide responsibilities across data extraction, predictor implementation, thermodynamic analysis, results review, discussion, and paper writing.

  • Authors' contributions: EC extracted ProTherm data, implemented the predictors, and wrote the paper.PF, IR, and RC contributed to the thermodynamic hypothesis, results review, discussion, and writing.
  • Authors' contributions: PF, IR, and RC contributed to the thermodynamic hypothesis, reviewed results, and helped write the paper.

Figures

The figures document the database distribution, predictor performance, ROC analyses, mutation-class analyses, and sequence–structure comparisons.

  • Sequence-based predictor: The sequence-based predictor uses reliability-index analyses to relate prediction accuracy and correlation to confidence.The relevant figure reports Q3 and C against RI, with DB denoting the fraction of the dataset above each threshold.
  • Structure-based predictor: The structure-based predictor similarly evaluates Q3 and C as functions of reliability index for structure-derived predictions.The figure caption defines DB as the fraction of DB3D at or above a given RI threshold.
  • ROC analyses: ROC figures compare the best sequence- and structure-based methods with SVM-BASE for non-neutral and neutral mutation predictions.Panels distinguish |ΔΔG|>0.5 Kcal/mole from |ΔΔG|≤0.5 Kcal/mole.
  • Mutation-class analyses: Mutation-class analyses report residue-pair accuracy for neutral, destabilizing, and stabilizing mutations using sequence and structure predictors alongside database frequencies.The neutral and destabilizing analyses explicitly describe symmetrised experimental data.
  • Sequence–structure comparison: Across residue-accessibility groups, the figures compare overall accuracy and mean correlation for sequence- and structure-based methods.The groups are highly buried, intermediate-accessibility, and exposed residues.

Tables

The tables report cross-validation performance for sequence windows, structure-centered environments, and comparisons with an I-Mutant predictor.

  • Sequence-based method: Table 1 reports sequence-based SVM cross-validation performance for different window lengths centered on the mutated residue.
  • Structure-based method: Table 2 reports structure-based SVM cross-validation performance for different protein environments centered on the mutated residue’s C-α.
  • Predictor comparison: Table 3 compares the best sequence-based and structure-based SVM methods with the I-Mutant-based predictor.The comparison uses SVM-WIN31 and SVM-3D12; related notation is supplied in the accompanying material.
Loading 0705.1490v2…