Source-linked AI summary
Direct-coupling analysis of residue co-evolution captures native contacts across many protein families
Faruck Morcos, Andrea Pagnani, Bryan Lunt, Arianna Bertolino, Debora S. Marks, Chris Sander, Riccardo Zecchina, Jose' N. Onuchic, Terence Hwa, Martin Weigt
TL;DR
The paper addresses whether sequence correlations can reliably identify native residue contacts beyond TCS proteins while separating direct from indirect effects. It develops and evaluates a faster mean-field implementation of DCA across protein domains, finding many correctly predicted contacts and broader structural signals, subject to sufficient homologous sequences.
Problem
The study asks how well DCA identifies native residue contacts in proteins other than TCS proteins, where indirect correlations complicate simple covariance-based prediction.
Method
The authors introduce mfDCA, a mean-field implementation of DCA, and evaluate its sequence-based contact predictions across 131 large domain families.
Results
DCA identifies many true contacts, recapitulates global contact-map structure across the examined domains, and detects signals involving alternative conformations, ligands, and inter-domain interactions.
Takeaways & Limitations
Predicted contacts can guide computational studies of alternative conformations, protein complex formation, and de novo protein-domain structure prediction.
Takeaways & Limitations
The approach requires a large number of homologous sequences, and earlier studies were limited to a few proteins with such sequence depth.
Abstract
from arXiv · showhide
The similarity in the three-dimensional structures of homologous proteins imposes strong constraints on their sequence variability. It has long been suggested that the resulting correlations among amino acid compositions at different sequence positions can be exploited to infer spatial contacts within the tertiary protein structure. Crucial to this inference is the ability to disentangle direct and indirect correlations, as accomplished by the recently introduced Direct Coupling Analysis (DCA) (Weigt et al. (2009) Proc Natl Acad Sci 106:67). Here we develop a computationally efficient implementation of DCA, which allows us to evaluate the accuracy of contact prediction by DCA for a large number of protein domains, based purely on sequence information. DCA is shown to yield a large number of correctly predicted contacts, recapitulating the global structure of the contact map for the majority of the protein domains examined. Furthermore, our analysis captures clear signals beyond intra- domain residue contacts, arising, e.g., from alternative protein conformations, ligand- mediated residue couplings, and inter-domain interactions in protein oligomers. Our findings suggest that contacts predicted by DCA can be used as a reliable guide to facilitate computational predictions of alternative protein conformations, protein complex formation, and even the de novo prediction of protein domain structures, provided the existence of a large number of homologous sequences which are being rapidly made available due to advances in genome sequencing.
Introduction
The paper evaluates whether DCA can identify native residue contacts beyond previously studied TCS proteins and introduces a faster implementation for large-scale analysis. Across 131 domain families, mfDCA captures many intra-domain contacts and broader structural signals from sequence data.
- Introduction: DCA addresses indirect correlations that caused simple covariance analysis to identify residue pairs far apart in structure.Direct Coupling Analysis aims to disentangle direct from indirect correlations.
- Introduction: The study asks how well DCA identifies native residue contacts in proteins other than TCS proteins.
- Introduction: mfDCA uses a mean-field approximation and is 10^3–10^4-times faster than mpDCA, enabling rapid analysis of many long protein sequences.The speed advantage addresses the computational cost of the slowly converging message-passing implementation.
- Introduction: Across 131 large domain families, mfDCA captures a large number of intra-domain contacts and recapitulates the global structure of contact maps.The evaluation uses domain families with accurate structural information.
- Introduction: Strong correlations between distant residues can reflect inter-domain contacts, alternative conformations, and ligand-mediated couplings.
- Introduction: mfDCA outperforms simple covariance analysis and a recent approximate Bayesian analysis for contact prediction.
Results and Discussion
The study evaluates DCA-based residue-contact prediction across many protein-domain families and finds that direct-information rankings recover native contacts more effectively than mutual information. The approach also reveals biologically meaningful long-range, inter-domain, conformational, and ligand-mediated couplings.
- Approach: DCA estimates residue-position occupancy and correlation from multiple sequence alignments to predict spatial proximity between residues.The analysis uses domain sequences collected in MSAs and evaluates predicted pairs against structural contacts.
- Approach: mfDCA is 10^3-10^4-times faster than the message-passing approach, enabling systematic analysis of hundreds of protein families and alignments up to about 500 amino acids per row.The faster heuristic addresses the computational cost that previously limited large-scale DCA studies.
- Contact prediction: 84% of the top-20 mfDCA direct-information pairs are true contacts on average across 131 domain families.The predicted distance distribution shows characteristic peaks around 3-5Å and 7-8Å that are absent from the background distribution of all residue pairs.
- Beyond intra-domain contacts: High-DI signals can reflect alternative conformations or ligand-mediated coupling rather than direct residue-residue contact.In DosR, displaced contacts involve the C-terminal helix and may reflect inter-domain interaction; in a metalloenzyme, Glu110 and His7 share interaction with a Mn(II) ion.
Methods
The study builds sequence-based DCA analyses from Pfam alignments, fits a maximum-entropy model to empirical frequencies, and introduces an efficient mean-field implementation for estimating direct couplings. Reweighting and pseudo-count regularization address sampling bias and matrix invertibility, while direct information ranks residue pairs.
- Data extraction: Sequences are organized into local Pfam-HMM multiple alignments, without further global-alignment refinement.The alignment columns define protein-domain length, and the study uses an in-house mapping to connect predictions with PDB residues.
- Sequence statistics and reweighting: Sequence reweighting corrects sampling bias by reducing the influence of sequences sharing more than 80% identity, and it systematically improves results.Performance depends only weakly on the precise threshold across 70–90%.
- Maximum-entropy modeling: DCA fits a maximum-entropy pairwise model whose couplings and fields reproduce empirical single-site and pair frequencies.A gauge choice fixes equivalent parameterizations by measuring fields and couplings relative to the last amino acid.
- Efficient DCA implementation: The new algorithm is about 3–4 orders of magnitude faster for L = 70 and directly analyzes alignments with L ≈ 1000, limited by working memory.Its speedup comes from performing parameter inference in a single computational step rather than iterative message passing.
- Small-coupling expansion: The mean-field approximation estimates couplings by inverting a connected-correlation matrix, but pseudo-counts are required when the unregularized matrix is noninvertible.The study uses a pseudo-count proportional to Meff after sensitivity analysis.
- Direct information: Direct information ranks residue pairs by the mutual-information signal attributable to direct coupling alone.For each pair, an isolated two-site model uses the inferred coupling and auxiliary fields compatible with empirical single-residue counts.
Figure legends
The figure legends compare DCA-derived direct information with mutual information and approximate Bayesian predictions across protein domains and structural examples. They highlight accurate intra-domain contacts alongside signals from oligomerization, alternative conformations, and ligand coordination.
- Global performance: DI results clearly outperform MI and an approximate Bayesian approach across the 131 domains studied.The comparison concerns top-ranked contacts and predicted structural-distance distributions.
- Contact distances: The top 10, 20, and 30 mfDCA predictions show distance peaks around 3–5 Å and 7–8 Å.The legends use an 8 Å structural-contact cutoff and compare predicted pairs against native contacts.
- Inter-domain interactions: Oligomerization contacts occur in 21 of the 131 families and constitute a significant fraction of long-distance, high-DI predictions.The Sigma-54 interaction-domain example contains three inter-domain contacts separated by less than 5 Å in a ring-like heptamer.
- Alternative conformations: The top-20 DI contact rate is 100% for the NarL DNA-binding domain but 65% for the full-length DosR structure.The difference is associated with a major reorientation of the GerE domain’s C-terminal helix between structures.
- Ligand-mediated couplings: A high-DI Glu110–His7 pair coordinates an Mn(II) ion in the protein’s dimer configuration.The accompanying structural example also depicts K+ ions coordinated by individual monomers.
- Long-range contacts: mfDCA yields more accurate predictions at larger sequence separations than MI in the illustrated domains.The figure bins MI and mfDCA predictions by separation along the protein sequence and marks true positives separately.
P(NAPx > n)
The section consists solely of a reference to Figure 7, without describing its content.
- Figure 7 is referenced, but its plotted quantity and interpretation are not specified.