Source-linked AI summary

Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling

Wengan He, Yongsheng Luo, Lihong Jiang, Wenhui Xu, Yu Li

arXiv:2608.16094v1cs.AIcs.LG

TL;DR

Protein structure prediction has expanded beyond MSA-driven monomer folding, but its methodological evolution across representations, architectures, learning strategies, confidence, and evaluation requires synthesis. This review organizes that evolution into four phases and three transitions, clarifying how recent models broaden structural modeling while retaining uneven capabilities across molecular classes and interactions.

  • Problem

    Monomer folding captures only part of biologically relevant structure, motivating methodological synthesis of models that handle interfaces and chemically distinct molecular entities.

  • Method

    The review analyzes methodological evolution across representations and data, architectures and learning strategies, and confidence and evaluation, organizing the field into four phases and three transitions.

  • Results

    The framework distinguishes complementary generalization profiles: MSA-based methods benefit from homologous coverage, whereas some MSA-free predictors help on orphan or low-homology targets.

  • Takeaways & Limitations

    Recent systems integrate previously separate tasks, but specialized inductive biases remain practically advantageous for challenging domains such as antibody modeling.

  • Takeaways & Limitations

    Current unified frameworks retain uneven accuracy, training coverage, and chemical treatment across molecular classes and interaction types.

Abstract

from arXiv · show

Accurate protein structure prediction is fundamental to structural biology because protein structure underlies molecular function and provides a basis for mechanistic interpretation. Recent advances in deep learning have transformed the field from multiple sequence alignment (MSA)-driven monomer folding into broader frameworks capable of modeling protein complexes and increasingly heterogeneous molecular systems. Existing reviews have summarized this progress from the perspectives of representative models, application domains, and protein design. Building on these efforts, this review focuses on the methodological evolution of the field itself. It examines recent developments through three closely related dimensions: representations and data, architectures and learning strategies, and confidence and evaluation. Within this perspective, the field is organized into four methodological phases and three cross-cutting transitions: from explicit evolutionary coupling features and early contact prediction to learned sequence representations in AlphaFold2, RoseTTAFold, and ESMFold; from protein-only monomer folding to increasingly integrated modeling of heterogeneous molecular systems in AlphaFold-Multimer, RoseTTAFoldNA, and AlphaFold3; and, more recently, from prediction-oriented structure inference to design-oriented generative modeling in RFdiffusion and related frameworks. This framework provides a clearer understanding of how methodological shifts have shaped the capabilities, limitations, and practical roles of recent models.

1. Introduction

This review traces protein structure prediction as a methodological evolution driven by deep learning, from evolutionary coupling and monomer folding toward heterogeneous-system modeling and generative design. It organizes this progression around representations and data, architectures and learning strategies, and confidence and evaluation, distinguishing four phases by modeling objective and system scope.

  • Motivation: Protein structure constrains molecular function and enables mechanistic interpretation, motivating computational methods for structure prediction.The Protein Data Bank is the principal global archive of experimentally determined macromolecular structures.
  • Methodological evolution: Deep learning shifted the field from co-evolution-assisted contact inference to end-to-end coordinate prediction with near-experimental accuracy for many monomeric proteins.Later developments expanded modeling to protein–protein complexes, protein–nucleic acid assemblies, and ligand-inclusive heterogeneous molecular systems.
  • Review perspective: The review emphasizes methodological evolution through representations and data, architectures and learning strategies, and confidence and evaluation.It differs from reviews organized primarily around representative model families, application scope, or protein design.
  • Methodological phases: The field is organized into four methodological phases spanning explicit evolutionary modeling, end-to-end deep folding, unified complex structure modeling, and generative structure modeling and design.The associated transitions proceed from explicit evolutionary features to learned sequence representations, from protein-only monomers to heterogeneous molecular systems, and from prediction to design-oriented generative modeling.
  • Phase assignment: Phases are assigned primarily by modeling objective and system scope rather than architecture alone.AlphaFold3 is placed in Phase III because its principal task is structure inference for specified heterogeneous biomolecular assemblies, despite using diffusion-based denoising in its structure module.

2. Explicit Evolutionary Modeling

Explicit evolutionary modeling used MSAs to derive co-evolutionary features for contact or distance prediction, followed by downstream three-dimensional reconstruction. Its modular, MSA-dependent design provided informative structural constraints but limited end-to-end correction, geometric consistency, and performance for proteins with shallow or biased homologous coverage.

  • Pipeline: The early pipeline constructed MSAs, extracted co-evolutionary features, predicted contacts or distances, and reconstructed structures downstream.The workflow used external folding or reconstruction procedures based on fragment assembly, distance geometry, or knowledge-based potentials.
  • Pipeline: Evolutionary constraints were treated as signals of three-dimensional structural relationships, with DCA and related statistics quantifying residue couplings.Residues that are spatially proximal tend to co-evolve to maintain structural stability and function.
  • Learning strategy: Deep learning improved long-range contact prediction, but engineered evolutionary signals, staged constraint prediction, and downstream reconstruction remained the dominant logic.RaptorX and related frameworks used MSA-derived features while preserving the fragmented methodological structure.
  • Limitations: The modular design prevented gradients from propagating through the full pipeline and allowed errors in MSA construction, coupling estimation, or contact prediction to reach downstream folding.Global geometric consistency was not directly optimized during contact prediction, so locally plausible constraints could remain incompatible with a coherent three-dimensional structure.
  • Limitations: Performance depended strongly on homolog depth, diversity, and alignment quality, while predicted contacts or distances remained underdetermined intermediate representations.Shallow or biased evolutionary coverage can produce sparse or noisy signals, and reconstruction quality also depends on downstream optimization and restraint consistency.

3. From Explicit Evolutionary Features to Learned Sequence Representations

Protein structure prediction shifted from multi-stage pipelines that explicitly extract evolutionary features from query-specific MSAs toward pretrained protein language models that learn sequence representations directly. This transition reduces dependence on external alignment preprocessing and enables more integrated prediction, while preserving model- and target-dependent trade-offs and motivating hybrid approaches.

  • Explicit evolutionary features: Earlier MSA-based pipelines retrieve homologs, construct an MSA, extract covariance or coupling features, and use them to predict contacts or distance distributions.Their core assumption is that correlated mutations across homologous sequences reveal structural constraints.
  • Learned sequence representations: PLM-based predictors encode query sequences with pretrained protein language models, replacing query-specific MSAs with contextual embeddings learned through self-supervised pretraining.The embeddings are passed to downstream folding or coordinate-prediction modules.
  • Methodological contrast: The methodological contrast is between alignment-dependent, multi-stage workflows and more integrated prediction pipelines based on learned sequence representations.MSA-based methods obtain target-specific evolutionary information, whereas PLM-based methods encode statistical regularities acquired during pretraining.
  • Generalization profiles: MSA-based methods remain effective with informative homologous coverage, while some MSA-free predictors can help on orphan or low-homology targets without target-specific alignments.This advantage does not generalize to all PLM architectures or difficult proteins.
  • Limitations and hybridization: PLM approaches trade less interpretable distributed embeddings and substantial pretraining requirements against reduced inference-time alignment preprocessing, with fine-grained constraint recovery remaining model- and target-dependent.Hybrid systems combining PLM embeddings with MSA-derived or pairwise geometric information may therefore be advantageous in some settings.

4. From Monomer Folding to Unified Modeling of Heterogeneous Molecular Assemblies

This section traces the methodological expansion from single-chain monomer folding to relatively unified modeling of heterogeneous molecular assemblies. The transition changes represented entities, learned geometric and chemical relationships, prediction targets, and evaluation criteria, while laying groundwork for generative structural modeling.

  • 4.1. Monomer Folding as the Foundational Paradigm: Monomer folding maps a single amino acid sequence to a dominant three-dimensional structure, commonly evaluated with TM-score, RMSD, and lDDT.These metrics are insufficient alone for complexes and ligand-containing systems, which require interface- and chemistry-aware measures.
  • 4.1. Monomer Folding as the Foundational Paradigm: Biologically relevant function often depends on binding, multimerization, nucleic-acid recognition, ligand association, and higher-order assembly, motivating intermolecular and chemically aware modeling.Monomer folding captures only a subset of these structural phenomena.
  • 4.2. From Multi-Chain Complexes to Heterogeneous Systems: The field progresses from multi-chain protein complexes through protein–DNA and protein–RNA systems toward models integrating proteins, nucleic acids, ligands, ions, and modified residues.AlphaFold-Multimer addresses inter-chain geometry, RoseTTAFoldNA targets nucleic-acid complexes, RoseTTAFold All-Atom broadens biomolecular scope, and AlphaFold3 supports joint inference across these entities.
  • 4.2. From Multi-Chain Complexes to Heterogeneous Systems: AlphaFold3 replaces the AlphaFold2 structure module with a diffusion module that denoises atom coordinates, yet remains prediction-oriented because molecular identities are specified.Repeated random seeds can produce multiple candidate predictions, but diffusion architecture alone does not make a model design-oriented.
  • 4.2. From Multi-Chain Complexes to Heterogeneous Systems: Relatively unified frameworks still show uneven accuracy, training coverage, and chemical treatment across molecular classes and interaction types.Evaluation must combine global or interface-level similarity with ligand pose quality and local stereochemical validation.
  • 4.2. From Multi-Chain Complexes to Heterogeneous Systems: Heterogeneous modeling narrows the conceptual boundary between prediction and generation and prepares approaches that model structure as a distribution of plausible configurations.This bridge follows from representing interaction-rich molecular systems in shared all-atom spaces with iterative refinement.

5. From Prediction-Oriented Structure Inference to Generative Protein Design

The section distinguishes prediction-oriented inference from generative protein design by conditioning and primary objective, rather than by sample count or diffusion use. It frames generative modeling as a complementary expansion toward controlled creation of novel molecular solutions.

  • Core distinction: Prediction-oriented models infer plausible structures for specified molecular inputs, whereas generative design creates novel molecular candidates under structural or functional constraints.AlphaFold2 returns a ranked structural hypothesis, while AlphaFold3 remains prediction-oriented despite diffusion-based sampling because sequence and composition are given.
  • Prediction limitations: Prediction-oriented inference becomes insufficient when targets occupy multiple biologically relevant states or when context-dependent dynamics are central to function.Confidence metrics and seed variability do not automatically represent a faithful physical ensemble; validation against experimental ensemble data, simulations, or state-specific benchmarks remains necessary.
  • Distributional prediction versus design: Distributional structure prediction samples alternative conformations for a specified sequence, whereas de novo design searches for new molecular solutions.EigenFold exemplifies sequence-conditioned distributional structure prediction, while RFdiffusion and La-Proteina are design-oriented generative systems.
  • Generative modeling: Generative design models can produce backbones, sequences, side-chain arrangements, or joint sequence–structure representations under constraints.Diffusion, flow-matching, and autoregressive methods provide alternative generative routes; RFdiffusion supports motif scaffolding, symmetric assemblies, and binder design.
  • Evaluation: Prediction and design require different evaluation criteria: predictive models emphasize structural correspondence and confidence calibration, while design models additionally require novelty, validity, diversity, designability, and experimental function.Generated candidates also face stereochemical, energetic, folding, expression, stability, binding, and biological-function constraints.
  • Practical integration: Prediction, generation, ranking, and experimental validation can be iterated in practical design loops, making generative modeling complementary to predictive structure models.Predictive models assess designed sequences against target folds or interaction geometries, while generative models propose candidates satisfying specified constraints.

6. Synthesis, Challenges, and Future Directions

Protein structure prediction has evolved from explicit physical and evolutionary modeling to learned representations, integrated heterogeneous-system modeling, and generative design. Future progress depends on addressing data imbalance, interpretability, evaluation, computational accessibility, and the integration of prediction, distributional inference, and design.

  • Methodological synthesis: The field has been methodologically reconfigured from explicit evolutionary constraints to learned representations, heterogeneous-system modeling, and generative design.This evolution changes how structural information is represented, learned, and utilized rather than merely improving benchmark accuracy.
  • Challenges: Training and evaluation remain biased toward well-represented biomolecular subsets because heterogeneous complexes, alternative conformations, and chemically diverse interactions are unevenly available.Monomer-focused structure and sequence resources are relatively rich, while examples of broader molecular systems are more unevenly distributed.
  • Challenges: Latent representations in protein language models, all-atom predictors, and generative models are harder to map onto biological mechanisms than explicit co-evolutionary features.Improved mechanistic interpretation would support validation, failure analysis, and appropriate scientific use.
  • Challenges: Broader modeling scope requires evaluation beyond TM-score, RMSD, and lDDT, including interface, ligand-pose and chemistry, ensemble, and generative-design assessments.Generative design specifically requires measures of validity, novelty, diversity, designability, and expected performance.
  • Future directions: Future systems should integrate dynamics, ensembles, context dependence, experimental evidence, and iterative prediction–design workflows grounded in structural and functional constraints.Progress will require ensemble-aware data and benchmarks, while multimodal inputs may include cryo-EM density, cross-linking, mutational scans, biochemical assays, and ligand-binding measurements.
  • Future directions: Hybrid frameworks should keep predictive evaluation, distributional exploration, and generative design distinct while improving efficiency, robustness, reproducibility, uncertainty calibration, and accessibility.Reporting should include training-data provenance, inference settings, sampling counts, uncertainty calibration, and hardware requirements.

CRediT authorship contribution statement

The authors contributed across conceptualization, methodology, investigation, validation, data curation, visualization, supervision, and manuscript writing. Wengan He also handled original-draft writing, while other authors contributed to original-draft or review-and-editing work.

  • Wengan He contributed conceptualization, methodology, investigation, data curation, visualization, and original-draft writing.
  • Yongsheng Luo contributed conceptualization, methodology, investigation, and original-draft writing; Lihong Jiang and Wenhui Xu contributed validation and review/editing.
  • Yu Li contributed conceptualization, supervision, and review/editing.
Loading 2608.16094v1…