Source-linked AI summary
CABS-dock web server for the flexible docking of peptides to proteins without prior knowledge of the binding site
Mateusz Kurcinski, Michal Jamroz, Maciej Blaszczyk, Andrzej Kolinski, Sebastian Kmiecik
TL;DR
Protein-peptide interactions are important but difficult to characterize because peptide flexibility and receptor flexibility complicate docking. CABS-dock performs site-agnostic flexible docking from random peptide conformations and positions, and over 80% of bound and unbound cases yielded high- or medium-accuracy models.
Problem
Structural characterization of protein-peptide interactions is difficult because current docking methods do not efficiently handle peptide conformational fluctuations and receptor flexibility.
Method
CABS-dock uses one efficient simulation to search for binding sites while allowing full peptide flexibility and small receptor-backbone fluctuations.
Results
Over 80% of bound and unbound benchmark cases produced models with high or medium accuracy.
Takeaways & Limitations
CABS-dock supports protein-peptide docking without prior binding-site knowledge and offers optional control over excluded binding modes or receptor-fragment flexibility.
Takeaways & Limitations
Prediction results may qualitatively differ between runs, and the tested docking protocol used default receptor-flexibility settings.
Abstract
from arXiv · showhide
Protein-peptide interactions play a key role in cell functions. Their structural characterization, though challenging, is important for the discovery of new drugs. The CABS-dock web server provides an interface for modeling protein-peptide interactions using a highly efficient protocol for the flexible docking of peptides to proteins. While other docking algorithms require pre-defined localization of the binding site, CABS-dock doesn't require such knowledge. Given a protein receptor structure and a peptide sequence (and starting from random conformations and positions of the peptide), CABS-dock performs simulation search for the binding site allowing for full flexibility of the peptide and small fluctuations of the receptor backbone. This protocol was extensively tested over the largest dataset of non-redundant protein-peptide interactions available to date (including bound and unbound docking cases). For over 80% of bound and unbound data set cases, we obtained models with high or medium accuracy (sufficient for practical applications). Additionally, as optional features, CABS-dock can exclude user-selected binding modes from docking search or to increase the level of flexibility for chosen receptor fragments. CABS-dock is freely available as a web server at http://biocomp.chem.uw.edu.pl/CABSdock
INTRODUCTION
CABS-dock addresses the difficulty of modeling flexible protein-peptide interactions by combining binding-site search, peptide modeling, and complex refinement in one simulation. It was tested on bound and unbound benchmark cases involving peptides of 5–15 amino acids.
- Protein-peptide interactions are important in cell functions, but their structural details remain relatively poorly understood and experimentally difficult to investigate.The dynamic and transient nature of peptide binding contributes to this difficulty.
- Current docking algorithms struggle with peptide conformational fluctuations and the computational cost of simultaneously treating receptor flexibility.Even small receptor fluctuations can be costly for many computational models.
- CABS-dock unifies binding-site prediction, flexible peptide modeling, and complex refinement into one efficient simulation of coupled folding and binding.The approach couples peptide folding and binding to a flexible receptor structure.
- Unlike approaches requiring prior binding-site knowledge, CABS-dock searches for docking sites without that information, extending previous success beyond very short peptides.Earlier site-agnostic docking had been successful only for peptides of 2–4 amino acids.
- The method was evaluated on 103 bound and 68 unbound protein receptor cases using peptides 5–15 amino acids long.The benchmark contains experimentally determined receptor structures with and without a peptide, respectively.
- Previous validation studies produced complex arrangements close to native structures while allowing fully flexible peptides without binding-site or peptide-conformation information.These studies included intrinsically disordered peptides, antibody-targeting peptides, and peptide co-activators of nuclear receptors.
- CABS-dock uses a coarse-grained multiscale CABS protein model designed to efficiently treat conformational changes while preserving local accuracy.The CABS model represents each amino acid with up to four interaction centers, uses Monte Carlo dynamics, and applies statistical potentials.
Protocol overview
The CABS-dock protocol generates random peptide structures and positions, simulates binding with replica-exchange Monte Carlo while restraining the receptor near native conformations, and selects and reconstructs final models.
- Protocol overview: Random peptide structures are generated and randomly placed on a sphere centered at the receptor’s geometrical center.The sphere radius is the receptor’s longest dimension plus 20 angstroms.
- Protocol overview: Binding and docking are simulated with Replica Exchange Monte Carlo using 10 replicas whose temperatures decrease during the simulation.The procedure produces 10 trajectories, each containing 1000 snapshots, for 10000 models in total.
- Protocol overview: Restraint penalties increase linearly after a 1 Å violation, with reduced or zero slope for user-selected semi-flexible or fully flexible receptor fragments.The penalty slope is halved for semi-flexible fragments and set to zero for fully flexible fragments.
- Protocol overview: Final-model selection first filters trajectories by binding state and energy, then clusters 1000 selected models using k-medoids.Ten consensus medoids from clustering with k=10 are selected as final models.
- Protocol overview: The selected final models are reconstructed from C-alpha traces to all-atom representations and optimized with Modeller using the DOPE statistical potential.
Docking without prior knowledge of the binding site
CABS-dock was validated for blind protein–peptide docking on bound and unbound cases without using bound peptide structures or binding-site information. Over 80% of cases produced high- or medium-accuracy models, with comparable performance for bound and unbound receptors.
- The protocol was validated against the largest available dataset of non-redundant protein–peptide interactions.
- The accuracy assessment uses peptide ligand RMSD after receptor superimposition, with high quality defined as RMSD < 3 Å and low quality as RMSD > 5.5 Å.
- CABS-dock used neither bound peptide structures nor binding-site information in the blind prediction test.
- Performance for bound cases was on the same level as for unbound cases because small interface differences were handled by default receptor flexibility.
- Over 80% of bound and unbound dataset cases yielded high- or medium-accuracy models sufficient for practical applications or further refinement.
- Prediction runs generally produced either consistent or qualitatively distinct predictions, so ambiguous cases should be examined across independent runs.
SERVER DESCRIPTION
The server accepts a receptor structure, peptide sequence, and optional supporting information for protein–peptide docking. Input constraints include standard amino-acid sequences, receptor backbone completeness, and a maximum peptide length of 30 amino acids.
- Required inputs include a protein receptor structure or PDB code with chain identifiers and a peptide sequence in one-letter code.
- Accepted receptor files may contain single- or multi-chain proteins, with chains up to 500 amino acids.
- Each receptor residue should have complete N, Cα, C and O backbone atoms, while side-chain atoms may be missing.
- Non-standard amino acids are automatically converted to standard counterparts.
- The peptide sequence must use standard amino acids and may contain at most 30 amino acids.
- If peptide secondary structure is not provided, PSI-PRED is used automatically; assignments use H, E or C codes.
- Overpredicting regular secondary structure is more harmful than underpredicting it, so ambiguous residues are better assigned coil.
Output interface
The output interface provides docking models, clustering information, and peptide–receptor contact maps. Users can inspect, download, and analyze representative and trajectory-derived structures through dedicated tabs.
- The interface is organized into Project information, Docking prediction results, Clustering details, and Contact maps tabs.
- Docking prediction results provide 10 final models representing structural clusters for 3D viewing and download.
- Users can download archives containing final models, cluster models, complete trajectories, receptor input structures, and log files.
- Clustering details show model composition by trajectory affiliation and provide interactive 3D viewing and PDB downloads.
- Cluster tables summarize properties including density and diversity.
- Contact maps display peptide–receptor residue contacts, allow users to set a distance cutoff, and provide a text contact list.
Advanced options
Advanced options let users allocate more simulation effort, exclude selected binding modes, or increase flexibility in chosen receptor fragments. These controls target difficult cases and enable focused exploration of receptor conformational alternatives.
- Users can exclude unlikely binding modes by marking receptor residues or resubmitting completed jobs with selected models excluded.
- Selected receptor fragments can receive higher flexibility by removing distance restraints that maintain near-native conformations.
- Users may choose moderate or full flexibility through the Mark flexible regions option.
- The simulation length can be increased from the default 50 to a maximum of 200 Monte Carlo macrocycles.
- Longer simulations may benefit difficult cases such as large receptors or peptides longer than 20 residues.
- Independent simulation runs provide an alternative strategy for demanding cases and can be analyzed together.
Online documentation
CABS-dock provides online documentation, free server access, persistent result links, and queued computational processing. The website supports visualization and job-status tracking for runs typically taking about three hours.
- Documentation: Documentation describes the method and provides a tutorial for accessing and interpreting results.It is updated regularly according to user needs or method improvements.
- Access and results: The server is free to all users and does not require login.Submitted jobs receive web links that can be bookmarked and revisited later.
- Access and results: Jobs are listed on a queue page unless users select the option to hide them from the results page.Results remain available only for a limited period.
- Implementation: The website uses Python with Flask and Jinja2, while molecular visualization uses 3Dmol.js and JSmol.The site runs with Apache2 and SQLite3 for user queue storage.
- Job processing: The queue is checked every 5 minutes, and jobs move through pending, running, and done or error states.Computations run on a Linux supercomputer cluster with about 100 CPU threads; a typical run takes about 3 hours.
SUMMARY
This work delivers an easy-to-use web server interface for the CABS-dock protein-peptide docking protocol. The authors anticipate applications to new systems and future modeling procedures, with extensions including refinement, experimental data, and increased receptor-fragment flexibility.
- SUMMARY: The CABS-dock protocol had already been used successfully in studies of protein-peptide interactions.The present work focuses on developing an accessible web-server interface for that protocol.
- SUMMARY: The authors expect the server to be applied to new systems and incorporated into new modeling procedures.
- SUMMARY: Potential extensions include refinement steps, incorporation of experimental data, and greater flexibility for appropriate receptor fragments.Predicted restraints are given as one example for increasing receptor-fragment flexibility.
TABLE AND FIGURES LEGENDS
The figures and tables depict the CABS-dock workflow, benchmark performance, web interface, and supplementary evaluation datasets. The illustrated workflow starts from random peptide conformations and positions, generates and filters models, and selects representative final models.
- Figure 1: The protocol illustration shows simulation starting from random peptide conformations and positions.The receptor is colored green, modeled peptide conformations magenta, and the reference experimental peptide structure yellow.
- Figure 1: The simulation produces 10,000 models, which are filtered and clustered by binding modes and peptide conformations before 10 representative final models are selected.
- Figure 1: In the benchmark illustration, 7 of 10 final models occupy the native binding site, and the best model is within 1.37 Å of the native structure.
- Figure 2: Figure 2 summarizes percentages of high-, medium-, and low-accuracy models for 103 bound and 68 unbound benchmark cases.The percentages are reported for all 10,000 models and the top 10 selected models.
- Supplementary tables: Supplementary Tables S1 and S2 report bound and unbound case results, respectively, across three prediction runs and progressively filtered model sets.Table S3 lists receptor-pair PDB codes in bound and unbound forms.
- Figure 3: Figure 3 presents example web-server output interfaces for docking prediction results and clustering details.