Source-linked AI summary
Sequence co-evolution gives 3D contacts and structures of protein complexes
Thomas A. Hopf, Charlotta P. I. Schärfe, João P. G. L. M. Rodrigues, Anna G. Green, Chris Sander, Alexandre M. J. J. Bonvin, Debora S. Marks
TL;DR
Protein interactions are numerous, but most lack detailed 3D structural information. This paper uses evolutionary couplings across paired protein sequences to predict inter-protein contacts and model complexes, finding accurate contact predictions and 3D models across known complexes while extending predictions to unknown complexes. The approach could support residue-level and genome-wide interaction analysis as sequence databases expand.
Problem
Most known protein interactions lack 3D structural information, limiting detailed characterization of their interfaces.
Method
The method scores correlated evolutionary changes between paired protein sequences to predict inter-protein residue contacts and uses them to calculate complex structures.
Results
The method identified contacts with 74% precision within 10Å and 69% within 8Å for pairs scoring above 0.8, while 11/15 top-ranked docked models had interface RMSDs under 5Å.
Takeaways & Limitations
With sufficient sequence data, evolutionary couplings can provide residue-level information for protein interactions and help determine structures of complexes lacking known 3D models.
Takeaways & Limitations
The approach depends on large numbers of evolutionarily related sequences and assumes interactions are conserved across species and paralogs.
Abstract
from arXiv · showhide
Protein-protein interactions are fundamental to many biological processes. Experimental screens have identified tens of thousands of interactions and structural biology has provided detailed functional insight for select 3D protein complexes. An alternative rich source of information about protein interactions is the evolutionary sequence record. Building on earlier work, we show that analysis of correlated evolutionary sequence changes across proteins identifies residues that are close in space with sufficient accuracy to determine the three-dimensional structure of the protein complexes. We evaluate prediction performance in blinded tests on 76 complexes of known 3D structure, predict protein-protein contacts in 32 complexes of unknown structure, and demonstrate how evolutionary couplings can be used to distinguish between interacting and non-interacting protein pairs in a large complex. With the current growth of sequence databases, we expect that the method can be generalized to genome-wide elucidation of protein-protein interaction networks and used for interaction predictions at residue resolution.
Introduction
The paper extends evolutionary-coupling analysis from individual proteins to protein complexes, addressing the scarcity of 3D interaction information. It develops and evaluates a method that predicts inter-protein contacts and supports 3D modeling, including for complexes of unknown structure.
- ~80% of known protein interactions lack 3D information, leaving at least ~30,000 human and ~6000 E. coli interactions incompletely characterized.
- The study tests whether correlated evolutionary changes across interacting proteins can identify spatially close residue pairs.
- An inter-protein evolutionary-coupling score ranks residue pairs using their overall score distributions.
- The method produces accurate predictions for most top-ranked inter-protein couplings and supports accurate 3D models of docked complexes.
- The approach predicts evolutionary couplings for 32 complexes without known 3D structures, including the ATP synthase a-, b- and c-subunit interaction.
Results
EVcomplex uses evolutionary couplings between paired proteins to identify residue contacts, guide complex docking, and predict interactions or structural details for complexes lacking structures. Tests on known complexes, unknown pairs, and ATP synthase show accurate contact and model recovery, while performance depends on assumptions and score interpretation.
- Method: The method pairs homologous protein sequences and uses EVcouplings to calculate inter-protein evolutionary coupling scores for residue pairs.The approach requires assumptions about which homologous proteins interact and applies a global maximum-entropy model with pseudolikelihood maximization.
- Blinded prediction of known complexes: 74% (69%) of predicted residue pairs with EVcomplex scores greater than 0.8 were accurate within 10Å (8Å) of experimental complex structures.At least one predicted contact exceeded the 0.8 threshold in 53 of 76 complexes.
- Blinded prediction of known complexes: Three complexes yielded more than 20 predicted contacts with over 80% precision: histidine kinase–response regulator, tRNA synthetase, and vitamin B importer.They produced 78, 32, and 21 residue pairs, respectively.
- Limitations: The 0.8 score threshold can miss correct contacts, while false positives arise from assumptions about conserved interactions, paralogs, residue proximity, and conformational states.In ethanolamine ammonia-lyase, five additional correct contacts scored slightly below 0.8; flexible ATP synthase subunits can also produce conflicting evolutionary constraints.
- Docking known complexes: Over 70% of docking models generated with evolutionary-coupling restraints were within 4Å interface RMSD of experimental structures, versus less than 0.5% of controls.All 15 best models had interface RMSDs under 6Å, and 11 of 15 top-ranked models were under 5Å.
- Functional constraints: Evolutionary-coupling networks linked ATP-binding regions to transporter subunits and reproduced experimentally identified functional constraints in MetNI and BtuCD.For MetNI, the top 10 inter-protein pairs were within 8Å and produced an average 1.4Å interface RMSD across 100 models.
- De novo prediction of unknown complexes: 32 of 82 structurally unknown protein pairs had at least one high-scoring predicted contact, and 17 of 19 DinJ–YafQ predictions were within 8Å of the subsequently solved structure.In ATP synthase, 24 of 28 possible subunit pairs were correctly classified as interacting or non-interacting.
Discussion
The approach can infer protein interactions and approximate complex structures from evolutionary couplings, but its accuracy and applicability depend on sequence availability, pairing assumptions, and complex architecture. The authors identify genome-wide scaling as a future opportunity alongside technical challenges involving homomultimers, flexibility, and less obligate interactions.
- Limitations: A primary limitation is dependence on large numbers of evolutionarily related sequences, especially when protein pairs occur in limited taxonomic branches.The current method also imposes a genome-distance requirement to reduce complications from uncertain paralog pairing and divergent interactions.
- Future scope: Approximately 1/10th of the 3000 known E. coli protein interactions can currently be analyzed with EVcomplex.The authors project broader inference as genome sequence availability and diversity increase.
- Structural modeling: Complex modeling from predicted contacts succeeds in many tested cases, but homomultimeric inter-protein couplings must be separated from intra-protein signals.This deconvolution is identified as an important technical challenge for future work.
- Interaction specificity: ATP synthase analysis provides a proof of principle for identifying interacting proteins and specific cross-protein residue couplings simultaneously.The authors do not yet know how well the approach handles less obligate interactions, although two-component signaling results suggest optimism.
- Scoring: A uniform EVcomplex threshold of 0.8 accounts heuristically for raw coupling scores, alignment depth, and concatenated-sequence length.The authors recommend examining contacts below this cutoff when independent biological knowledge or higher-scoring contacts support them.
- Future scope: The method may accelerate genome-wide exploration of protein interactions and complex structures at residue-level resolution.The authors frame this as a future direction conditioned on successful de novo calculation of co-evolved residues.
Materials and methods
The study constructs paired sequence alignments, computes normalized evolutionary-coupling scores, and evaluates predicted contacts against known structures. It then uses high-confidence couplings as docking restraints and assesses the resulting models by interface RMSD.
- Selection of interacting protein pairs: Candidate protein pairs were drawn from E. coli interaction data, including experimental, literature-curated, and PDB-supported interactions, with three additional complexes added.Pairs were aligned and concatenated using genomic proximity rules for homologous proteins.
- Computation of evolutionary couplings: EVcomplex analyzes concatenated paired sequences with a global maximum-entropy model and PLM to generate intra- and inter-protein coupling scores.Columns containing more than 80% gaps were excluded, and sequence weights reflected cluster size.
- Definition of a scale-free score for the assessment of interactions: The normalized EVcomplex score adjusts raw coupling reliability for alignment sequence depth and inference-problem size.Sequence sufficiency required Neff/L > 0.3, where Neff is the effective sequence count and L is the concatenated alignment length.
- Prediction of interactions in a set of subunits: For de novo structural interpretation, EVfold generated monomer models for structurally unsolved ATP synthase subunit-a and UmuC.Coupling parameters used PLM, with sequences clustered and weighted at 90% identity.
- Prediction of interactions in a set of subunits: The highest-ranked inter-EC score was used as an interaction proxy: scores above 0.8 indicated likely interactions, 0.75–0.8 weak predictions, and below 0.75 rejection.The protocol was also applied to all 28 pairings among eight E. coli ATP synthase subunits.
- Docking and structural evaluation: Fifteen diverse complexes with at least five high-scoring inter-protein couplings were docked using HADDOCK with Cα distance restraints.Docking comprised rigid-body minimization, semi-flexible torsion-angle refinement, and explicit-solvent refinement, generating 500, 100, and 100 models respectively.
- Comparison of predicted to experimental structures: Docked models were compared with cognate crystal structures using interface backbone RMSD, defining interfaces as residues within 6 Å of the partner.Mobile terminal helices and flexible domains were excluded from specified complex comparisons.