Source-linked AI summary
Inverse Statistical Physics of Protein Sequences: A Key Issues Review
Simona Cocco, Christoph Feinauer, Matteo Figliuzzi, Remi Monasson, Martin Weigt
TL;DR
Protein sequences change extensively during evolution even as structure and function remain conserved, creating a need to infer the constraints governing homologous sequence families. The review applies inverse statistical-physics and Potts-model approaches to sequence data, showing that pairwise couplings support contact prediction while also highlighting sampling-bias challenges. These models therefore connect sequence variability with structural and functional constraints, within the assumptions of the modeling framework.
Problem
Homologous proteins can differ substantially in amino-acid sequence while conserving structure and function, raising the question of how sequence variability reflects evolutionary constraints.
Method
The review models MSAs of homologous proteins with Boltzmann and generalized Potts distributions, inferring parameters from empirical sequence statistics.
Results
Strong pairwise couplings provide accurate residue-contact predictions, whereas fields alone can produce non-functional sequences and may be insufficient for generative protein design.
Takeaways & Limitations
Inverse statistical-physics models can extract information about evolutionary constraints and relate sequence variability to conserved protein structure and function.
Takeaways & Limitations
Current models assume sequences are independent samples from an unknown distribution, although biological sequences retain phylogenetic and common-ancestry correlations.
Abstract
from arXiv · showhide
In the course of evolution, proteins undergo important changes in their amino acid sequences, while their three-dimensional folded structure and their biological function remain remarkably conserved. Thanks to modern sequencing techniques, sequence data accumulate at unprecedented pace. This provides large sets of so-called homologous, i.e.~evolutionarily related protein sequences, to which methods of inverse statistical physics can be applied. Using sequence data as the basis for the inference of Boltzmann distributions from samples of microscopic configurations or observables, it is possible to extract information about evolutionary constraints and thus protein function and structure. Here we give an overview over some biologically important questions, and how statistical-mechanics inspired modeling approaches can help to answer them. Finally, we discuss some open questions, which we expect to be addressed over the next years.
INTRODUCTION
Homologous proteins can vary substantially in sequence while conserving structure and function, motivating inverse statistical-physics models of evolutionary constraints. These models use aligned sequence families to connect conservation and coevolution with protein structure, function, and mutational effects.
- Homologous proteins may differ in more than 70 or 80% of amino acids while retaining highly similar three-dimensional structures and biological functions.
- Large families of homologous sequences enable statistical analysis of sequence variability to uncover constraints associated with conserved protein structure and function.
- An MSA represents protein-family sequences as rows and aligned residue positions as columns, with gaps treated here as a 21st amino acid.
- Inverse modeling reconstructs sequence distributions despite an unknown Hamiltonian and sparse coverage of the q^L sequence space, requiring parameter-reduced models.
- Model specification can use assumed Hamiltonian forms with parameter inference or maximum entropy to match selected empirical observables, introducing subjective choices.
- Strong pairwise couplings accurately predict residue contacts, while local fields alone can generate non-functional sequences and may be insufficient for generative protein design.
STATISTICAL PHYSICS OF THE INVERSE POTTS PROBLEM
The inverse Potts framework begins with an MSA and estimates sequence statistics while correcting for sampling biases caused by shared ancestry and uneven species coverage. Mutual information alone cannot distinguish direct couplings from indirect correlations, motivating inverse methods.
- The method starts from an MSA of aligned protein-family sequences and infers local fields and pairwise couplings for a generalized Potts model.
- Empirical modeling uses single-site amino-acid frequencies and pairwise co-occurrence frequencies extracted from MSA columns.
- Sequence-specific weights define an effective sequence number to reduce distortions from similar homologous sequences and unevenly sampled species.
- Weights based on Hamming-distance neighborhoods commonly use x ≃ 0.2−0.3, a range reported to produce accurate and robust contact-prediction results.
- Mutual information is positive for statistically dependent sites but is an inaccurate proxy for pairwise coupling because correlations can arise indirectly through coupling networks.
From Amino-Acid Frequencies to Potts models
Potts-model inference fits empirical sequence statistics by constraining model averages to match observed frequencies. Pairwise models capture correlations more realistically than independent-site models but require computationally expensive inference over an exponential sequence space.
- Pairwise constraints produce a more realistic model but exact inference requires a partition-function sum over q^L sequences.
- Independent-site models fit local fields to first-order frequencies but factorize positions and cannot reproduce second-order co-occurrence statistics.
- Pairwise Potts models infer local fields and couplings so Boltzmann-distribution first and second moments equal empirical single- and two-site frequencies.
- The coupled inference system contains Np = Lq + L(L−1)q^2/2 equations that must be solved simultaneously.
Overparametrization and gauge invariance
The inverse Potts problem is overparameterized because frequency constraints are dependent and gauge transformations leave the Boltzmann distribution unchanged. Regularization is therefore needed to address undersampling and overfitting in protein MSAs.
- Overparametrization and gauge invariance: Only L(q−1) one-point and L(L−1)/2 (q−1)^2 pairwise frequencies are independent because pairwise frequencies marginalize to single-site frequencies.These constraints total Nc = L(q−1) + L(L−1)/2 (q−1)^2.
- Overparametrization and gauge invariance: Gauge transformations leave P unchanged, reducing the fitted parameter count until it matches the number of independent constraints.Inference becomes well defined once a gauge is fixed.
- Overparametrization and gauge invariance: The lattice-gas and zero-sum gauges provide two equivalent ways to remove this parameter redundancy.The lattice-gas gauge uses a reference amino acid, whereas the zero-sum gauge preserves symmetry among Potts states.
- Cross-entropy minimization and Bayesian interpretation: Cross-entropy minimization is equivalent to minimizing the KL divergence between empirical and Boltzmann distributions, or maximizing Bayesian log-likelihood with a prior.The Potts fields and couplings are conjugate to one- and two-point frequencies.
- Regularization: Typical proteins require Nc ≃ 5 · 10^5−5 · 10^7 inferred parameters from M ≃ 10^2−10^5 sequences, making regularization necessary.Undersampling can create spurious correlations, while L1 and L2 penalties or pseudocounts reduce such effects.
Methods of Approximate Inference
Exact averages required by the inverse equations are generally intractable for proteins with hundreds of residues, so approximate inference schemes are used.
- Methods of Approximate Inference: Exact evaluation requires the partition function Z, whose sum over all q^N sequences is impractical for proteins with hundreds of residues.The properly regularized cross entropy remains convex, allowing local convex optimization once averages are approximated.
Boltzmann Machine Learning (BM)
Boltzmann machine learning iteratively adjusts Potts fields and couplings by comparing MCMC-estimated marginals with empirical frequencies. Its computational cost restricts direct application to relatively small protein families.
- Boltzmann Machine Learning (BM): Boltzmann machine learning starts from initial fields and couplings, estimates one- and two-point marginals by MCMC, and updates parameters against empirical frequencies.The update follows the cross-entropy gradient with step size ϵ.
- Boltzmann Machine Learning (BM): Small ϵ and sufficiently precise MCMC sampling guarantee convergence toward the fixed-point equations.The guarantee applies to the iterative procedure under those conditions.
- Boltzmann Machine Learning (BM): Efficient implementations make Boltzmann machine learning applicable to families smaller than about L = 200, while studies of hundreds or thousands of families remain out of reach.The limiting factor is the procedure’s high computational cost.
- Boltzmann Machine Learning (BM): Covariance inversion replaces an exponential-time inference problem of time q^L with approximately L^3 operations.This can be performed on a standard desktop even for proteins of L ≃ 1000 amino acids.
Mean-field Approximation (MF)
Mean-field and related approximations reduce the computational burden of inverse Potts inference, but mean-field inference is not exact and lacks rigorous error bounds. Pseudolikelihood methods avoid partition-function evaluation and are statistically consistent under the stated model assumptions.
- Mean-field Approximation (MF): Mean-field inference obtains couplings by inverting the covariance matrix and then resolves the fields using single-site frequencies and couplings.Covariances are calculated through linear response.
- Mean-field Approximation (MF): Mean-field inference is currently among the most computationally efficient approximations, but it does not converge to the exact solution even with infinite data.No rigorous bounds on its inference errors are known.
- Pseudolikelihood Maximization (PLM): Pseudolikelihood replaces the full log probability with site-dependent terms whose normalization requires L sums over q amino acids rather than one sum over q^L sequences.This avoids the exponential-time partition-function calculation.
- Pseudolikelihood Maximization (PLM): Pseudolikelihood is statistically consistent: with infinite data sampled from a pairwise Boltzmann distribution, it recovers the true fields and couplings.Independent minimization of site terms enables parallel implementation, though finite data can produce asymmetric couplings.
- Pseudolikelihood Maximization (PLM): Pseudolikelihood becomes slower for larger samples because its objective depends on the complete set of sampled configurations and scales linearly with M.This contrasts with methods based only on empirical single- and two-site frequencies.
Adaptive Cluster Expansion (ACE)
Adaptive Cluster Expansion (ACE) truncates a cross-entropy expansion to strongly interacting clusters, then derives Potts parameters from empirical frequencies. It reproduces sampled frequencies and correlations accurately while adapting model complexity to sampling quality, at greater computational cost.
- Adaptive Cluster Expansion (ACE): ACE represents minimal cross entropy as cluster contributions based on empirical one- and two-site frequencies.Each contribution captures information not deducible from the cluster’s subclusters.
- Adaptive Cluster Expansion (ACE): Clusters are recursively expanded from pairs and retained when their cross-entropy contributions exceed threshold θ.The expansion stops when newly created clusters have |∆S| below θ.
- Adaptive Cluster Expansion (ACE): Derivatives of the cross entropy with respect to empirical frequencies provide the Potts fields and couplings.The threshold is tuned so Monte Carlo samples reproduce the data statistics within finite-sampling errors.
- Adaptive Cluster Expansion (ACE): ACE accurately reproduces sampled frequencies and correlations, unlike Gaussian, mean-field, and pseudolikelihood approximations.Its convergence depends on interaction-graph structure rather than correlation magnitude.
- Adaptive Cluster Expansion (ACE): ACE uses threshold θ as a regularizer, adapting network sparsity to sampling quality.Its main disadvantage is computational burden from exact partition-function calculations for each cluster.
Other Approaches
The review compares independent-site, approximate, and more exact approaches for modeling protein sequences. Approximate methods can recover strong coevolutionary topology efficiently, whereas accurate energies, probabilities, and sequence sampling require more precise inference.
- Other Approaches: Independent-site profile models capture position-specific conservation, while profile hidden Markov models additionally handle unaligned sequences and insertions or deletions.These models assume statistical independence between positions.
- Other Approaches: Mean-field and pseudolikelihood inference reproduce empirical statistics poorly, whereas Boltzmann-machine learning and ACE are more precise.Mean-field inference also causes serious equilibration problems in MCMC simulations for the PF00014 example.
- Other Approaches: For contact prediction, inference quality matters less than correctly recovering the topology of the coevolutionary network.Approximate methods can therefore remain useful even when exact parameter values are inaccurate.
- Other Approaches: The average-product correction subtracts a null-model contribution based on the single-site properties of the two positions.It improves contact predictions when applied to coupling strengths in the zero-sum gauge.
- Other Approaches: 80% of BM’s top 50 coupled pairs also appear among PLM’s top 50, compared with 70% for MF and 62% for ACE.These overlaps rise to 100%, 88%, and 76%, respectively, when the comparison set expands to the first 100 pairings.
- Other Approaches: MF and PLM are suitable for identifying strongly coevolving pairs, but ACE or BM are required when model energies, probabilities, or sampling must be accurate.This distinction reflects topology-focused versus distribution-focused applications.
- Other Approaches: ACE and BM reproduce empirical Hamming-distance histograms, while IND and PLM show significant deviations.Hamming distance is not directly fitted by the Potts model.
Residue-residue contact prediction and tertiary
Direct-coupling methods use evolutionary sequence variation to infer residue contacts within and between proteins, supporting structure and complex prediction. Coupling-based rankings outperform raw correlations, but practical scope is limited by sequence separation, interface conservation, matching, and computational cost.
- Residue-residue contact prediction and tertiary: Potts-model couplings are ranked to predict residue contacts, providing structural information without detailed biophysical modeling.The strongest couplings are expected to identify contacting residue pairs.
- Residue-residue contact prediction and tertiary: Pairs with sequence separation |i −j| ≤4 are excluded because gap stretches often produce large, structurally less informative scores.Long-range contacts are more informative for tertiary-structure prediction.
- Residue-residue contact prediction and tertiary: In PF00014, coupling-based predictions produce significantly fewer false positives than mutual-information rankings.The comparison uses 4,915 sequences of length 53 mapped onto the 5PTI structure.
- Residue-residue contact prediction and tertiary: DCA consistently improves over correlations across tests involving thousands of protein families.Approximate inference can suffice because contact prediction mainly requires coevolutionary network topology.
- Residue-residue contact prediction and tertiary: Predicted contacts constrain the number of possible protein structures and have been used in thousands of novel structure predictions.They also provide prior information for structure-prediction methods.
- Residue-residue contact prediction and tertiary: The same Potts formalism infers interprotein contacts and can be extended to organism-wide protein-protein interaction networks.The model introduces LaLbq^2 cross-protein coupling parameters for two protein families.
- Residue-residue contact prediction and tertiary: Interprotein inference requires paired MSAs from the same organisms, with paralog matching handled using biological priors or additional probabilistic modeling.A protein-pair interaction score can average the n largest FAPC values; n = 4 is reported to perform well.
- Residue-residue contact prediction and tertiary: Reliable interprotein signals are most evident at large, widely conserved interfaces, whereas smaller or partially conserved interfaces are difficult to detect.Inference is also computationally expensive because one Potts model is needed per protein pair.
entire sequences
Potts models assign sequence energies that can assess compatibility with a protein family and discriminate experimentally folding from non-folding sequences. In WW-domain tests, preserving pairwise amino-acid correlations improved folding outcomes, and inference techniques performed comparably for sequence ranking.
- Mutation effects: Potts energies can predict mutation fitness effects, with applications ranging from pathogen virulence and drug resistance to disease-causing mutations.The approach has been tested across viral, bacterial, and human proteins.
- Scoring entire sequences: Potts-model energies quantify how compatible a sequence is with the natural sequence variability of a protein family.The energy is minus the sequence log-probability up to a normalization constant.
- Experimental sequence design: 31% of CC sequences and 67% of NAT sequences folded correctly, whereas none of the R or IC sequences folded under the experimental conditions.CC preserves pairwise frequencies, while IC preserves single-site frequencies but destroys covariation.
- Scoring entire sequences: SCA scores whole shuffled alignments, whereas Potts-model energies score individual sequences and were reported to predict which sequences fold.The distinction matters when evaluating or designing sequences one at a time.
- Energy and folding: Energy was an excellent discriminator between folding and non-folding sequences across the NAT, CC, IC, and R data sets.Most folding CC and natural sequences had lower energies than IC and independent-model sequences; higher-energy CC and natural sequences were mostly non-folding.
- Inference robustness: MF, PLM, and BM showed quantitatively different inference results but comparable performance in distinguishing folding from non-folding sequences.The authors attribute this to sequence ranking requiring parameters that are sufficiently accurate rather than maximally precise.
Generative aspects and entropy: from lattice
Potts models support generative studies of structurally valid protein sequences and provide entropy-based measures of sequence-family diversity. Applications extend from lattice proteins and WW domains to estimating conservation and mutational landscapes in HIV.
- Generative aspects: ACE-inferred Potts models generated lattice-protein sequences likely to fold into the same structure, whereas independent-model sequences rarely folded.This result held in the idealized 27-residue, 3 × 3 × 3 lattice setting, but not for less precise MF- and PLM-based models.
- Generative aspects: Pairwise statistical information was reported as necessary and sufficient for generating structurally valid proteins in the lattice-protein setting.The claim is explicitly framed as applying to this simple model and as confirming the WW-domain studies.
- Entropy: Approximately 1.2 nats per residue position, or e^1.2 ≃3.3 different amino acids per site, estimates WW-domain sequence diversity under the Potts model.Because the model includes only one- and two-point statistics, the true entropy may differ through higher-order statistical contributions.
- HIV applications: Cross-entropies of HIV proteins identify which viral proteins are more conserved and therefore less inclined to mutations.The analysis used alignments from the Los Alamos database and was illustrated for 14 HIV protein families.
- HIV applications: Potts-model mutation energy costs correlated with in-vitro viral replicative power at coefficient −0.76.The comparison used mutations relative to the wild-type sequence.
CONCLUSION AND OUTLOOK
Inverse statistical-physics methods use expanding sequence data to study evolutionary constraints and biological structure and function. The review highlights their promise while emphasizing unresolved issues in phylogeny, higher-order interactions, prior knowledge, and large-scale validation.
- CONCLUSION AND OUTLOOK: Large sequence datasets create both an opportunity for sophisticated modeling and a challenge for analyzing biological systems.The review places this work at the interface of statistical physics, bioinformatics, and biology.
- CONCLUSION AND OUTLOOK: Analyzing variability among evolutionarily related proteins can reveal which structural and functional characteristics constrain evolution over millions of years.This connects sequence statistics to biological principles governing protein evolution.
- Open questions: There is no fundamental reason higher-order statistical couplings should not contribute, so the success and limits of pairwise models remain unresolved.This issue is especially important for generative models that should produce sequences statistically indistinguishable from natural sequences.
- Open questions: Current models treat amino acids as abstract symbols and do not yet incorporate complementary physicochemical and structural knowledge systematically.The review calls for principled integration of heterogeneous information sources.
- Open questions: Current models assume sequences are independently and identically distributed, although biological sequences retain phylogenetic memory of common ancestors.Empirical reweighting provides corrections, but the review identifies the correct treatment of phylogenetic structure as an open question.
- Open questions: Large-scale predictions such as unknown protein-structure prediction remain rare because methods have generally been tested on sample cases with known answers.This is identified as a limitation of the current state of application.