Source-linked AI summary

Machine learning for protein folding and dynamics

Frank Noé, Gianni De Fabritiis, Cecilia Clementi

arXiv:1911.09811v1physics.bio-phphysics.chem-phq-bio.BMstat.ML

TL;DR

Protein-folding and dynamics research needs methods that can predict structures, model energy landscapes, analyze simulations, and enhance rare-event sampling. This review synthesizes machine-learning approaches across these tasks, highlighting advances alongside unresolved physical, computational, and interpretive challenges. It concludes that machine learning is becoming important while physical knowledge and trained scientific judgment remain essential.

  • Problem

    Protein folding and dynamics require improved approaches for structure prediction, energy-function design, simulation analysis, and rare-event sampling.

  • Method

    The paper reviews machine-learning methods for protein structure prediction, learned force fields, simulation analysis, and enhanced sampling.

  • Results

    Machine learning has advanced protein structure prediction and is being applied to force-field design, simulation analysis, and enhanced sampling of protein dynamics.

  • Takeaways & Limitations

    Machine learning can provide new tools and reveal patterns in protein folding and molecular science, especially when methods incorporate physical symmetries, invariances, and conservation laws.

  • Takeaways & Limitations

    Important challenges remain, including long-range interactions, computational speed, loss of equilibrium kinetics under enhanced sampling, and the continued need for physical knowledge and scientific interpretation.

Abstract

from arXiv · show

Many aspects of the study of protein folding and dynamics have been affected by the recent advances in machine learning. Methods for the prediction of protein structures from their sequences are now heavily based on machine learning tools. The way simulations are performed to explore the energy landscape of protein systems is also changing as force-fields are started to be designed by means of machine learning methods. These methods are also used to extract the essential information from large simulation datasets and to enhance the sampling of rare events such as folding/unfolding transitions. While significant challenges still need to be tackled, we expect these methods to play an important role on the study of protein folding and dynamics in the near future. We discuss here the recent advances on all these fronts and the questions that need to be addressed for machine learning approaches to become mainstream in protein simulation.

Introduction

Machine learning is increasingly being applied to protein folding and dynamics, including structure prediction, energy-function design, simulation analysis, and enhanced sampling. The paper reviews these advances and the challenges to broader adoption.

  • Machine learning extracts complex patterns and relationships from large datasets for evaluating new data.
  • In protein folding and dynamics, machine learning has been used for multiple purposes across fundamental-science problems.
  • Structure prediction uses machine learning to associate folded protein structures with sequence information, with recent benefits visible in CASP competitions.
  • Neural networks can represent energy functions that include multi-body terms difficult to model analytically.
  • Unsupervised learning can extract metastable states from high-dimensional simulations and connect them to measurable observables.
  • The review covers recent machine-learning contributions and anticipates their increasing significance for protein folding and dynamics.

Machine learning for protein structure prediction

Machine learning has substantially advanced protein structure prediction by using evolutionary information and deep-learning architectures. In CASP13, AlphaFold ranked first with a simplified workflow based heavily on machine learning.

  • Structure prediction infers a protein’s folded structure from its sequence information.
  • Recent machine-learning successes apply deep learning to evolutionary information, including patterns of co-evolution between contacting amino acids.
  • CASP evaluates structure-prediction methods every two years using blind predictions for sequences whose structures have not been released.
  • Historically, top CASP predictors differed little, indicating incremental workflow improvements rather than a clearly superior method.
  • In CASP13, AlphaFold ranked first with a simplified machine-learning-heavy workflow that predicted distance histograms from co-evolutionary data.
  • AlphaFold used an autoencoder and a knowledge-based potential derived from predicted distance histograms to generate or minimize structures.

Folding proteins with machine learned force fields

Machine-learned force fields represent energy functions directly from atomic coordinates and train on reference properties such as quantum-mechanical energies and forces. Their application remains constrained by long-range interactions and simulation speed.

  • Machine-learned force fields represent classical energy functions with neural networks or other models instead of specifying functional forms a priori.
  • These models can be trained to reproduce energies and forces obtained from quantum-mechanical calculations.
  • Neural networks can approximate many-body energy relationships that are difficult to encode analytically.
  • The CGnet example represents a protein’s effective energy with a neural network and visualizes Chignolin’s folding free-energy landscape using TICA coordinates.
  • Long-range electrostatic and van der Waals interactions may be missed when training uses quantum calculations on only small molecules.
  • Neural-network energy and force calculations are faster than ab initio calculations but slower than standard classical force fields.

Machine learning of coarse-grained protein folding models

Machine learning is used to build coarse-grained protein models that reduce computational cost while representing effective interactions. Neural networks are attractive because coarse-graining produces nonlinear multi-body terms that are difficult to specify with general functional forms.

  • Coarse-grained models map groups of atoms to effective beads and assign bead-level energy functions to reproduce protein properties.
  • Coarse-grained models can study larger systems and longer timescales with reduced computational resources.
  • Renormalizing local degrees of freedom creates multi-body terms in coarse-grained effective energies, even when the reference force field is pairwise.
  • Suitable general functional forms for these coarse-grained multi-body effects are challenging to define.
  • Neural networks naturally capture nonlinearities and multi-body terms without requiring a specific functional form.

Machine learning for analysis and enhanced simulation of protein dynamics

Machine learning supports both the extraction of slow collective variables and metastable states from protein simulations and the enhancement of rare-event sampling. Approaches range from MSM-based analysis and VAMPnets to adaptive, biased, and generative sampling methods, with trade-offs in kinetic information.

  • Protein-dynamics analysis targets slow collective variables, metastable states, and increased transitions between them.
  • Markov state models estimate transitions between metastable states from short, nonequilibrium trajectories and predict equilibrium and long-time behavior.
  • The conventional MSM construction pipeline remains error prone and dependent on substantial expert knowledge despite improvements in feature selection, dimensionality reduction, clustering, and transition estimation.
  • VAMPnets learn slow reaction coordinates from time-lagged simulation data and can hierarchically decompose the NTL9 state space into metastable states.
  • Adaptive sampling selects new simulation starting states using learned slow coordinates or metastable states, accelerating rare transitions such as protein folding and unfolding.
  • Enhanced-sampling methods use learned collective variables or generative models to improve rare-event sampling, including adaptive metadynamics, VES, and Boltzmann Generators.
  • Enhanced sampling can reconstruct target-state equilibrium distributions through reweighting but generally loses equilibrium kinetic information.

Conclusions

Machine learning offers new tools for protein folding and structure prediction, but physical knowledge, symmetries, conservation laws, and scientific interpretation remain important. Methods incorporating these physical principles perform better than black-box approaches across the reviewed areas.

  • Machine learning can advance molecular science, including protein folding and structure prediction, while physical and chemical knowledge remains essential.
  • Methods incorporating physical symmetries, invariances, and conservation laws perform better than black-box methods across the reviewed areas.
  • Scientists remain essential for interpreting learned patterns and using them to formulate general principles.
Loading 1911.09811v1…