Source-linked AI summary

Identification of direct residue contacts in protein-protein interaction by message passing

M. Weigt, R. A. White, H. Szurmant, J. A. Hoch, T. Hwa

arXiv:0901.1248v1q-bio.BMcond-mat.stat-mechq-bio.QM

TL;DR

Protein-interaction specificity is difficult to resolve because covariance methods cannot separate direct from indirect residue correlations. The paper combines covariance analysis with global inference through message passing and applies it to bacterial two-component signaling proteins. It successfully identifies spatially proximal residue pairs for both SK/RR heterointeractions and RR/RR homointeractions without ad hoc tuning parameters.

  • Problem

    Covariance-based sequence analysis identifies correlated residues but cannot distinguish direct from indirect correlations, limiting inference of protein-interaction surfaces.

  • Method

    The method first identifies correlated residues by covariance analysis, then uses statistical message passing to infer direct coupling between residue positions.

  • Results

    The method distinguishes interacting from non-interacting residues for both SK/RR heterodimers and RR/RR homodimers with vastly improved accuracy and no ad hoc tuning parameters.

  • Takeaways & Limitations

    The approach demonstrates that global inference can identify direct protein interactions from sequence information alone and may extend to general interfaces with sufficient homologous sequences.

  • Takeaways & Limitations

    At present, analyses are limited to proteins that are highly amplified in sequenced genomes, while single-copy protein pairs require more sequence data.

Abstract

from arXiv · show

Understanding the molecular determinants of specificity in protein-protein interaction is an outstanding challenge of postgenome biology. The availability of large protein databases generated from sequences of hundreds of bacterial genomes enables various statistical approaches to this problem. In this context covariance-based methods have been used to identify correlation between amino acid positions in interacting proteins. However, these methods have an important shortcoming, in that they cannot distinguish between directly and indirectly correlated residues. We developed a method that combines covariance analysis with global inference analysis, adopted from use in statistical physics. Applied to a set of >2,500 representatives of the bacterial two-component signal transduction system, the combination of covariance with global inference successfully and robustly identified residue pairs that are proximal in space without resorting to ad hoc tuning parameters, both for heterointeractions between sensor kinase (SK) and response regulator (RR) proteins and for homointeractions between RR proteins. The spectacular success of this approach illustrates the effectiveness of the global inference approach in identifying direct interaction based on sequence information alone. We expect this method to be applicable soon to interaction surfaces between proteins present in only 1 copy per genome as the number of sequenced genomes continues to expand. Use of this method could significantly increase the potential targets for therapeutic intervention, shed light on the mechanism of protein-protein interaction, and establish the foundation for the accurate prediction of interacting protein partners.

Introduction

Protein-interaction surfaces are difficult to characterize experimentally, motivating sequence-based computation. This study combines covariance analysis with message passing to infer direct contacts in bacterial two-component signaling proteins.

  • Motivation: Experimental interaction-surface methods are arduous or serendipitous, and co-crystal structures are particularly challenging for transient partners.Structural evidence also requires independent validation that the interaction reflects a physiologically relevant complex.
  • Problem: Covariance methods detect correlated substitutions but cannot distinguish direct from indirect residue correlations.Indirect correlations can arise through conformational changes or networks of smaller interactions.
  • Method: Message passing addresses the global optimization challenge of disentangling direct interactions from covariance-derived correlated residue pairs.The method starts from correlated residues and infers the set of direct couplings using statistical-physics-inspired message passing.
  • Study system: The study applies a two-stage covariance and message-passing analysis to more than 2,500 coupled sensor-kinase and response-regulator pairs.The bacterial two-component system is highly amplified and provides extensive sequence, genetic, and structural information for evaluation.
  • Results: The method distinguishes interacting from non-interacting residues for both SK/RR heterodimers and RR/RR homodimers with vastly improved accuracy without ad hoc tuning parameters.The authors propose applicability to general protein interfaces when sufficient homologous sequences are available.

Results and Discussion

Global inference separates direct from indirect residue correlations, identifying spatially proximal contacts in SK/RR heterointeractions and RR homointeractions from sequence data. The resulting DI rankings are robust and outperform MI for predicting residue proximity.

  • Disentangling direct and indirect correlations: DI ranks the 32 MI-selected pairs by direct coupling, distinguishing nine strong-contact candidates from 22 pairs whose high MI reflects many weak links.Group I links connect eight SK positions with five RR positions and are expected to represent physical interface contacts; Group II links are not expected to contact directly.
  • SK/RR structural interpretation: Group I pairings define an SK/RR interaction mode between SK α1/α2-helices and the RR α1-helix, bringing catalytic-site residues together.Mapped Group I residues are exposed on the relevant helices, whereas Group II pairings cannot be spatially aligned consistently with direct interactions.
  • Structural validation: ≤ 6Å separates the mapped Group I residue pairs at the Spo0B–Spo0F interface, and 5 of 6 near-threshold pairs also fall within this distance.Strongly directly coupled residues are anti-correlated with structural distance.
  • Predictive performance: DI maintains specificity one for the top 2.5% of 408 SK/RR scoring pairs, whereas MI produces a false positive after one true positive and falls to 30–40% specificity.The comparison supports DI as a substantially better indicator of residue proximity than MI.

Concluding Remarks

The paper combines covariance analysis with global inference implemented by message passing to separate direct from indirect residue correlations. This approach enhances contact-pair specificity, but its current applicability depends on sufficient homologous sequences.

  • Concluding Remarks: A novel computational method combines covariance analysis with global inference to infer structural details of protein-protein interactions from primary sequence information.Message passing disentangles correlations arising from direct versus indirect interactions.
  • Concluding Remarks: The combined approach enhances the specificity of contact-pair prediction compared with traditional purely local covariance methods such as MI.The comparison is explicitly made against mutual-information-based approaches.
  • Concluding Remarks: Current applicability relies on approximately 10 structurally homologous protein sequences in a typical bacterial genome.The authors anticipate that expanding genomic databases may reduce this constraint for many proteins occurring in one genomic copy.
  • Concluding Remarks: A possible E86-R108 salt bridge is not realized in the ArcA structure because of a likely crystallographic artifact involving a neighboring ArcA dimer.The neighboring-dimer contact is unavailable in solution.
  • Concluding Remarks: The molecular interaction details revealed may provide potential targets for antibiotic drug design when precise structural information is unavailable.The paper frames this as a possible application of the interaction information obtained by the method.
  • Concluding Remarks: The method may aid interpretation of correlations in other large biological datasets, including mRNA and protein profiles and neuronal spike activities.The paper presents this as a broader potential application of disentangling direct and indirect interactions.

Material and Methods

The study constructs aligned, nonredundant datasets of interacting domains from bacterial genomic and Pfam information, then applies message passing to a reduced set of correlated positions. Computational cost motivates position selection and limits feasible sequence sizes.

  • Material and Methods: Domains were aligned and culled from the non-redundant RefSeq database, using HMM searches restricted to complete domains.Unique species were included to avoid oversampling organisms with multiple sequenced strains.
  • Material and Methods: The SK/RR study used Pfam domains PF00512 and PF00072, while RR/RR analysis considered proteins containing PF00072 with additional domain requirements.Functional association was inferred from chromosomal adjacency for the SK/RR interaction study.
  • Material and Methods: Global inference estimates model parameters from sequence-alignment marginal distributions, with direct couplings and local amino-acid biases as model parameters.Technical details are reported in the supplementary text.
  • Material and Methods: The computationally efficient but semi-heuristic message-passing approach estimates single- and two-variable marginal distributions using belief and susceptibility propagation.The approach is exact on tree-like coupling graphs and also works efficiently on loopy graphs.
  • Material and Methods: The reduced model selects up to approximately 60 positions with high MI values and considers all intra- and inter-protein pairs among them.This reduction is required because the full-sequence computational cost is O(212N4).
  • Material and Methods: Inference for the reduced model required about 4 days on a single CPU, whereas selecting 100 residues would require more than one month.Smaller sets of N=32,40,50 produced qualitative results that did not depend on the MI cutoff once all Fig. 3 network nodes were included.

Direct information

Direct information quantifies the coupling between two sequence positions after global inference. It is computed from a two-position distribution generated using the inferred direct coupling while preserving the observed single-position marginals.

  • Direct information: Direct information is calculated from the contribution of the direct coupling eij(Ai,Aj) to the inferred two-residue distribution.The calculation focuses on positions i and j in the statistical model.
  • Direct information: The relevant two-position contribution is obtained from a hypothetical system containing only positions i and j.These positions are coupled by eij(Ai,Aj).
  • Direct information: The hypothetical two-position system retains the correct single-variable marginals fi(Ai) and fj(Aj).This construction isolates the direct interaction while preserving the individual residue preferences.
  • Direct information: DI measures the direct coupling strength between positions i and j.The mathematical definition of the two-position distribution is provided in the supplementary text.

Figure Legends

The figure legends describe how MI and DI classify correlated residue pairs, map inferred interactions onto protein structures, and relate DI to structural proximity. They emphasize that DI distinguishes strong direct links from low direct correlations and predicts contacts more specifically than MI.

  • Figure 1: Figure 1 separates residue pairs into high-MI/high-DI Group I and low-direct-correlation Group II, with cutoff-borderline pairs highlighted in blue.Strong links in Group I are expected to represent physical contacts at the SK/RR dimer interface.
  • Figure 1: Figure 1 maps strong direct correlations in red and low direct correlations in green onto HisKA and response-regulator structures.Residues connected by red or green lines distinguish the direct-correlation categories.
  • Figure 2: Figure 2 plots minimal residue-pair distance against DI and MI for 408 pairings mapped to the Spo0B/Spo0F co-crystal structure.DI is shown with red symbols and MI with blue symbols.
  • Figure 2: Figure 2 compares contact-pair prediction specificity across rank percentiles for DI and MI.Specificity is defined as the fraction of ranked pairings within 6Å in the Spo0B/Spo0F co-crystal structure.
  • Figure 3: Figure 3 shows identified dimer contact pairs and maps them onto exemplary OmpR-class response-regulator structures.Contact pairings localize to the α4- and α5-helices; pairing 89:109 forms a salt bridge in all three examples.
Loading 0901.1248v1…