Source-linked AI summary

3D RNA and functional interactions from evolutionary couplings

Caleb Weinreb, Adam J. Riesselman, John B. Ingraham, Torsten Gross, Chris Sander, Debora S. Marks

arXiv:1510.01420v3q-bio.BM

TL;DR

RNA structure and functional interactions remain difficult to characterize as sequence discovery outpaces structural and interaction research. The paper uses evolutionary couplings from sequence co-variation to infer RNA and RNA-protein contacts, enabling accurate contact detection, blinded 3D structure prediction, and functional-interaction discovery.

  • Problem

    RNA sequence discovery outpaces knowledge of non-coding RNA structure and functional interactions, including RNA-RNA contacts in divergent RNase P systems.

  • Method

    The study applies global maximum entropy models of sequence co-variation to infer nucleotide-nucleotide and nucleotide-amino acid contacts, then uses them for RNA structure prediction and interaction analysis.

  • Results

    Evolutionary couplings predict RNA 3D contacts with good accuracy, exceed mutual information for long-range and non-Watson-Crick contacts, and provide proof of principle for RNA-protein contact detection.

  • Takeaways & Limitations

    Evolutionary coupling analysis can reveal structural and functional interactions across RNAs and RNA-protein complexes from sequence information, including contacts relevant to riboswitches and HIV Rev-RRE interactions.

  • Takeaways & Limitations

    Sequence abundance remains the rate-limiting component, and the folding pipeline is not optimized for large RNAs over 120 nt.

Abstract

from arXiv · show

Non-coding RNAs are ubiquitous, but the discovery of new RNA gene sequences far outpaces research on their structure and functional interactions. We mine the evolutionary sequence record to derive precise information about function and structure of RNAs and RNA-protein complexes. As in protein structure prediction, we use maximum entropy global probability models of sequence co-variation to infer evolutionarily constrained nucleotide-nucleotide interactions within RNA molecules, and nucleotide-amino acid interactions in RNA-protein complexes. The predicted contacts allow all-atom blinded 3D structure prediction at good accuracy for several known RNA structures and RNA-protein complexes. For unknown structures, we predict contacts in 160 non-coding RNA families. Beyond 3D structure prediction, evolutionary couplings help identify important functional interactions, e.g., at switch points in riboswitches and at a complex nucleation site in HIV. Aided by accelerating sequence accumulation, evolutionary coupling analysis can accelerate the discovery of functional interactions and 3D structures involving RNA.

Results

Evolutionary couplings accurately recover RNA contacts, including long-range and non-Watson-Crick interactions, and improve structure prediction. They also extend to RNA-protein complexes and reveal functional interactions in riboswitches, HIV RRE, and RNase P.

  • Method: The maximum-entropy model fits natural RNA sequence distributions and identifies strongly coupled nucleotide pairs as evolutionary couplings.The model incorporates single-site biases and pairwise coupling terms, fitted by penalized maximum likelihood.
  • RNA contact prediction: ECs contained 2.4 times more long-range contacts than enhanced mutual information and captured 16% of annotated non-Watson-Crick base pairs, 1.7 times more than MI.Long-range enrichment ranged from 0.8–10.8 times across the 22 examples, while non-Watson-Crick enrichment ranged from 0.5–4.0 times.
  • RNA contact prediction: ECs were over 90% accurate for the top 900 contacts in the eukaryotic ribosome and produced 2.8-fold fewer false positives than enhanced mutual information.They also predicted more Watson-Crick, non-Watson-Crick, and long-range contacts.
  • 3D structure prediction: Models using EC-derived tertiary contacts deviated less from experimental structures than secondary-structure-only controls in all five RNA families.The five comparisons were statistically significant, with p < 0.01 for each.
  • RNA-protein complexes: ECs predicted RNA-protein contacts and improved docking, with all six selected complexes showing lower i-RMSD than controls and four containing a top-three model below 3.5 Å.The selected complexes had at least 75% true positives among their top four predicted contacts.

Discussion

Evolutionary couplings provide structural and functional information for RNA and RNA–protein systems, while the approach remains limited by contact-selection, folding, and sequence-coverage constraints.

  • Contributions: Evolutionary couplings predict long-range and non-Watson–Crick RNA contacts and support blinded 3D structure prediction with good accuracy.They also identify functionally important interactions in riboswitches and the HIV RRE–Rev complex.
  • RNA–protein interactions: Evolutionary couplings can detect intermolecular contacts in RNA–protein complexes from sequence information alone.High-throughput application requires improved curation and phasing of RNA–protein alignments.
  • Limitations: The pipeline reports the top L/2 contacts, although choosing the number flexibly would better accommodate variation in RNA contact abundance.The folding pipeline is also not optimized for RNAs longer than 120 nt.
  • Limitations: Sequence abundance remains the rate-limiting component, motivating broader sequence coverage and inclusion of additional structured transcripts.The paper identifies expanding RFAM coverage as a future research direction.

Experimental Procedures

The pipeline calculates evolutionary couplings by fitting a global probability model to each sequence alignment using pseudo-maximum likelihood.

  • Experimental Procedures: Evolutionary couplings are calculated by fitting a global probability model to each alignment with pseudo-maximum likelihood.The complete pipeline is provided as a single script.

1. Sample re-weighting (and definition of Meff)

Sequence re-weighting reduces phylogenetic bias by weighting each sequence according to its number of neighbors and summing those weights into an effective alignment size.

  • 1. Sample re-weighting (and definition of Meff): Each sequence is down-weighted by the number of neighboring sequences it has in sequence space to reduce phylogenetic bias.The effective alignment size is defined as the sum of sequence weights.
  • 1. Sample re-weighting (and definition of Meff): The model is fit with a pseudo-maximum likelihood approximation and L2 regularization.These choices define the probability-model fitting procedure.

3. Post-processing

Post-processing converts fitted pairwise couplings into corrected evolutionary-coupling scores by applying a Frobenius norm and average-product correction.

  • 3. Post-processing: Frobenius norms of fitted pairwise couplings quantify co-variation strength before correction.Average-product correction removes characteristic under-sampling distortions, and the corrected scores are reported as final evolutionary couplings.

3D structure prediction

The study adapts a global maximum-entropy model to infer RNA evolutionary couplings and uses them to predict tertiary contacts and three-dimensional structures. The workflow addresses phylogenetic and sampling artifacts while combining coupling-derived restraints with coarse-grained modeling and all-atom refinement.

  • 3D structure prediction: Five 70-120-nt RNA families with known structures and at least one contact spanning d!! ≥ L/4 were modeled using NAST decoys, clustering, XPLOR refinement, and all-atom structures.For each family, 600-1000 coarse-grained decoys were generated, the lowest-energy-per-contact 20% were clustered with k=4, and four candidate models were refined.
  • Modeling evolutionary couplings: The model represents RNA-family sequence probabilities with single-site fields and pairwise nucleotide couplings, whose matrix slices describe co-variation between position pairs.The partition function normalizes the probability distribution, while coupling strength is assessed from the Frobenius norm of fitted coupling matrices.
  • Model fitting: Pseudolikelihood replaces direct maximum likelihood because computing the partition function is intractable, and parameters are optimized with L-BFGS.The objective balances fit to the data against regularization, which is needed because the parameter count can greatly exceed the effective sequence count.
  • Corrections and weighting: Sequence reweighting reduces phylogenetic redundancy by assigning sequences weights inversely proportional to the size of their sequence neighborhoods.The neighborhood threshold was set to θ=0.2 for RFAM alignments and θ=0.1 for most RNA-protein alignments, with θ=0.033 for HIV Rev-RRE alignments.
  • Corrections and weighting: Average product correction removes phylogeny- and undersampling-related bias from Frobenius-norm scores, producing the evolutionary-coupling scores used for contact inference.The correction subtracts normalized row and column averages from each matrix position.

Prediction of mutational effects

The inferred global probability model is also used to estimate how mutations alter sequence-model energy. This provides predictions of mutations likely to disrupt key interactions in the T box riboswitch and RNase P.

  • Prediction of mutational effects: Mutation effects are calculated as the energy difference between the mutated sequence and the original sequence.The sequence energy combines single-site fields and pairwise coupling terms from the fitted global model.
  • Prediction of mutational effects: The method predicts mutations likely to disrupt key interactions in the T box riboswitch and RNase P.These predictions use the inferred parameters h and J and follow a previously described methodology.

Computing MI

The study compares evolutionary couplings with mutual-information measures of co-evolution. It separates the global-model contribution from sequence reweighting and average-product correction by also computing an enhanced MI score.

  • Computing MI: Raw mutual information is computed from single-position and joint-position residue frequencies in the alignment.The frequency terms represent marginal and pairwise probabilities for states at positions i and j.
  • Computing MI: Evolutionary-coupling scores differ from raw MI through a global maximum-entropy model, phylogenetic sequence down-weighting, and APC correction.Enhanced MI incorporates the latter two features without replacing the pairwise MI framework with the global model.

Annotating interactions

Known RNA structures are used to annotate predicted contacts by structural distance and biochemical interaction type. The analysis focuses on the top L/2 contacts separated by more than four positions along the chain.

  • Annotating interactions: For each alignment, the analysis examines the top L/2 contacts with chain-distance > 4 and labels true positives when minimum atom distance is < 8 Å.Contacts are further classified by secondary-structure distance and biochemical interaction type using RFAM structures and FR3D annotations.
  • Annotating interactions: Long-range contacts are defined by secondary-structure distance dss > 4, where dss is the shortest path through secondary-structure contacts or chain adjacency.Biochemical interaction types are assigned from crystal-structure annotations downloaded from RNA3DHub.

Computing 3D structures from evolutionary couplings

The pipeline uses evolutionary-coupling-derived tertiary restraints to generate and refine blinded RNA 3D structure predictions. It produces four candidate structures per RNA family and evaluates them against known structures and an unrestrained control.

  • Blinded predictions target 70–120-nt RNA families with known structures and at least one highly long-range contact.Highly long-range contacts satisfy d_ij ≥ L/4.
  • Each family begins with 200 random unfolded structures satisfying secondary-structure constraints, which are folded using top evolutionary-coupling contacts as tertiary restraints.The number of top long-range contacts varies from 20 to the RNA length L.
  • Decoys are ranked by energy per contact, and the lowest-energy 20% are clustered with k = 4 before representative models are converted to all-atom structures.Energy per contact is E/N, where E is NAST energy and N is the number of restraint contacts.
  • The four representatives are converted to all-atom models, refined by simulated annealing in XPLOR, and compared with the crystal structure using all-atom RMSD and Molprobity.A no-tertiary-restraint pipeline provides the control comparison.

RNA-‐protein 3D structure prediction

RNA–protein contact prediction requires phased RNA and protein alignments, which are concatenated using taxonomy information before evolutionary couplings are calculated. The highest-confidence contacts are then used as restraints for rigid-body docking and interface-accuracy assessment.

  • RNA alignments come from RFAM, while protein alignments come from PFAM or UniProt searches when PFAM lacks full coverage.Protein sequences are filtered by gap content, length coverage, and alignment quality.
  • RNA–protein sequences are phased by NCBI taxonomy ID, retaining representatives with average Hamming distance no greater than 1%.One RNA and one protein representative are randomly selected for each retained taxonomy ID.
  • The RNA evolutionary-coupling model is extended to a full alphabet containing amino acids, with no other model changes.This produces couplings between RNA nucleotides and protein amino acids.
  • Six of 21 complexes with at least 75% true positives among their top four contacts are docked in HADDOCK using those contacts as 5±2 Å distance restraints.The remaining docking parameters use defaults, and interface accuracy is assessed with i-RMSD.
  • The HIV RRE analysis uses both an RFAM alignment and a custom LANL alignment built from HIV env sequences and a reference RRE covariance model.The SL4 and SL5 conformations are evaluated using structures from the pNL4-3 genome and RNAeval energy calculations.

Evolutionary couplings for T box riboswitch and RNase P

The supplementary analyses characterize evolutionary-coupling strength, structural prediction workflows, and contact patterns in RNA and RNA–protein systems. They also identify settings where low-ranking contacts may be false positives and compare coupling-based predictions with controls or alternatives.

  • Functional RNA families: For archaeal RNase P, low-ranking contacts meeting the long-range criteria appeared isolated at random map positions, a hallmark of false positives.The analysis defines long-range contacts using d_ss ≥ 12 for these large sequences.
  • Evolutionary-coupling strength: Secondary-structure contacts generally have higher normalized evolutionary-coupling ranks, whereas long-range contacts show a broader rank distribution.Some long-range contacts outrank some secondary-structure contacts.
  • Long-range contact prediction: For the eukaryotic ribosome, mutual information predicts dramatically fewer true-positive long-range contacts than evolutionary couplings and 2.8-fold more false positives.The comparison uses the top L/2 predicted contacts.
  • RNA structure prediction: For five known RNA families, four evolutionary-coupling-based candidate structures are compared with the experimental structure, with the most accurate candidate highlighted.The related decoy analysis contrasts evolutionary-coupling restraints with no restraints.
  • RNA–protein complexes: For 21 RNA–protein complexes, true-positive rates are examined across increasing contact numbers and effective sequence depth, followed by rigid-body docking for selected complexes.The docking analysis focuses on six complexes with sufficiently accurate top contacts.
  • Prediction workflow: The RNA 3D-structure workflow generates random secondary structures, folds them with evolutionary couplings, prunes violated contacts, clusters decoys, and performs all-atom conversion and refinement.These stages form the pipeline illustrated in Figure S7.
Loading 1510.01420v3…