Source-linked AI summary

Structural biology in the clouds: The WeNMR-EOSC Ecosystem

Rodrigo Vargas Honorato, Panagiotis I. Koukos, Brian Jiménez-García, Andrei Tsaregorodtsev, Marco Verlato, Andrea Giachetti, Antonio Rosato, Alexandre M. J. J. Bonvin

arXiv:2107.01056v1q-bio.BMcs.CEcs.DC

TL;DR

Understanding macromolecular structures, dynamics, and interactions requires diverse computational approaches and tools. This review presents the WeNMR-EOSC ecosystem and reports its worldwide use at substantial scale.

  • Problem

    Researchers need diverse approaches and tools to understand macromolecular structure, dynamics, function, and interactions for biological and biotechnological applications.

  • Method

    The paper reviews WeNMR-EOSC structural biology services, technological solutions for distributed computing, and their use by a worldwide community.

  • Results

    >12 million jobs from over 23,000 users across 125 countries accounted for ~4,000 CPU-years in 2020.

  • Takeaways & Limitations

    WeNMR-EOSC provides user-friendly web services that let researchers run complex structural biology workflows using local and distributed computing resources.

Abstract

from arXiv · show

Structural biology aims at characterizing the structural and dynamic properties of biological macromolecules at atomic details. Gaining insight into three dimensional structures of biomolecules and their interactions is critical for understanding the vast majority of cellular processes, with direct applications in health and food sciences. Since 2010, the WeNMR project (www.wenmr.eu) has implemented numerous web-based services to facilitate the use of advanced computational tools by researchers in the field, using the high throughput computing infrastructure provided by EGI. These services have been further developed in subsequent initiatives under H2020 projects and are now operating as Thematic Services in the European Open Science Cloud (EOSC) portal (www.eosc-portal.eu), sending >12 millions of jobs and using around 4000 CPU-years per year. Here we review 10 years of successful e-infrastructure solutions serving a large worldwide community of over 23,000 users to date, providing them with user-friendly, web-based solutions that run complex workflows in structural biology. The current set of active WeNMR portals are described, together with the complex backend machinery that allows distributed computing resources to be harvested efficiently.

1 INTRODUCTION

Structural biology investigates the 3D structures, dynamics, and interactions of proteins and nucleic acids to understand biological processes and support biotechnology and health applications. The section introduces the WeNMR-EOSC ecosystem, its structural biology services, and solutions for efficient, user-friendly distributed computing for a worldwide community.

  • Motivation: Proteins and nucleic acids’ 3D structures, dynamics, and interactions are crucial for understanding biological processes and advancing biotechnology and health applications.Applications include developing new drugs.
  • WeNMR-EOSC ecosystem: WeNMR provides specialized structural biology software and user-friendly, efficient, cost-effective distributed execution through its EOSC Thematic service.The passage identifies DIRAC Work as part of this distributed execution approach.
  • Contribution: The section presents WeNMR-EOSC services, technological challenges and solutions for using distributed computing resources, and their use by an active worldwide community.The overview focuses on how the ecosystem supports efficient use of distributed resources.

2 SERVICES · 2.1 AMPS-NMR

Within WeNMR-EOSC, freely available web services address the multifaceted computational needs of macromolecular structural biology. AMPS-NMR provides a user-friendly portal for refining experimental NMR structures and running biomolecular molecular dynamics simulations.

  • 2 SERVICES: WeNMR-EOSC develops diverse tools as freely available web services for researchers studying macromolecular function, behavior, and dynamics.
  • 2 SERVICES: The EOSC portal hosts the currently active WeNMR-EOSC services described in the following overview.
  • 2.1 AMPS-NMR: AMPS-NMR offers a user-friendly interface for restrained molecular dynamics simulations that refine experimental NMR structures.
  • 2.1 AMPS-NMR: The portal also supports molecular dynamics simulations of biomacromolecular systems generally.
  • 2.1 AMPS-NMR: AMPS-NMR refinement consistently improves rotamer distributions, backbone normality, and steric-clash occurrence without affecting agreement with experimental data.
  • 2.1 AMPS-NMR: The refinement effect is especially relevant in protein regions with scarce experimental information.
  • 2.1 AMPS-NMR: A predefined multi-step rMD protocol removes the need for users to tune numerous parameters in a complex molecular-dynamics calculation.
  • 2.1 AMPS-NMR: AMPS-NMR automatically handles the most commonly used formats for experimental data.

2.2 DISVIS · 2.3 FANTEN · 2.4 HADDOCK

DISVIS visualizes distance-restrained interaction spaces from experimental interaction data, FANTEN determines paramagnetic NMR anisotropy tensors and generates protein–protein models, and HADDOCK performs experimentally guided integrative docking with tiered parameter access.

  • 2.2 DISVIS: DISVIS uses residue-participation and distance information from XL-MS and FRET to visualize the three-dimensional distance-restrained possible interaction space.These experiments can identify participating residues and provide distances or upper distance limits between reacting groups.
  • 2.2 DISVIS: DISVIS therefore presents experimentally constrained molecular interaction possibilities in a three-dimensional visualization.The constraints derive from interaction measurements such as cross-linking mass spectrometry and FRET.
  • 2.3 FANTEN: FANTEN determines anisotropy tensors associated with NMR paramagnetic pseudocontact shifts and residual dipolar couplings.The server supports analyses based on pcs and/or rdc data.
  • 2.3 FANTEN: FANTEN can determine protein–protein adduct structures from pseudocontact shifts using a rigid-body approach.The resulting models can be submitted directly to HADDOCK for flexible refinement.
  • 2.4 HADDOCK: HADDOCK is an integrative docking platform that uses experimental and theoretical information to dock up to 20 macromolecules.It is described as a pioneer of integrative docking and is named High Ambiguity Driven biomolecular DOCKing.
  • 2.4 HADDOCK: HADDOCK separates users into access tiers because exposing 400+ parameters could hinder usability.The Easy tier exposes and permits modification of a limited set of parameters, while the Guru tier provides full control upon specific request.
  • 2.4 HADDOCK: The Guru tier grants full access and control of the HADDOCK docking protocol after a specific user request.Users request this tier through the registration page; the HADDOCK webserver is available online.

2.5 MetalPDB · 2.6 PDB-tools web

MetalPDB organizes metal-binding sites in biological macromolecule structures as fold-independent 3D templates and provides statistical analyses of metal usage. PDB-tools web addresses the continued use of the strict, flat-text PDB format despite wwPDB’s shift to the mmCIF dictionary.

  • 2.5 MetalPDB: MetalPDB is a database of metal-binding sites found in 3D structures of biological macromolecules.It collects information on sites present in experimentally determined macromolecular structures.
  • 2.5 MetalPDB: Metal-binding sites are represented as 3D templates of local environments around metal ion(s).The templates describe local metal-site environments independently of the overall macromolecular fold.
  • 2.5 MetalPDB: The site templates are independent of the overall macromolecular fold.This representation focuses on the local environment surrounding the metal ion(s).
  • 2.5 MetalPDB: MetalPDB provides detailed statistical analyses of metal usage in proteins with known 3D structures.The passage identifies statistical analysis of metal usage as one of MetalPDB’s features.
  • 2.6 PDB-tools web: The classic PDB format remains a flat text representation used by many structural biology software packages.It represents the spatial coordinates of macromolecular structures.
  • 2.6 PDB-tools web: The wwPDB has shifted its official format to the Crystallographic Information Framework mmCIF dictionary.The classic PDB format nevertheless continues to be used by many structural biology software tools.
  • 2.6 PDB-tools web: Despite its apparent simplicity, the flat-text PDB format follows strict formatting rules.The passage notes that these rules can be easily overlooked.

2.7 Powerfit · 2.8 proABC-2

Powerfit is presented as a Python package for fast, sensitive rigid-body fitting of molecular structures into cryo-EM density maps when resolution prevents direct model building. proABC-2 predicts antibody residues involved in antigen contacts and characterizes their hydrophobic or hydrophilic interactions.

  • 2.7 Powerfit: Powerfit addresses cryo-EM cases where map resolution is insufficient to build molecular models directly.Existing three-dimensional structures can instead be fitted into the resulting density maps.
  • 2.7 Powerfit: Powerfit is implemented as a Python package for fitting molecular structures into EM density maps.The package was developed for rigid-body fitting.
  • 2.7 Powerfit: Powerfit is designed for fast and sensitive rigid-body fitting of molecular structures.Its stated application is fitting existing three-dimensional structures into cryo-EM maps.
  • 2.8 proABC-2: Monoclonal antibodies are established therapeutic tools because of their high affinity and specificity toward targets.This context motivates tools for analyzing antibody–antigen interactions.
  • 2.8 proABC-2: proABC-2 predicts which antibody residues can form intermolecular contacts with an antigen.It is the updated version of proABC.
  • 2.8 proABC-2: proABC-2 provides insight into the chemical components of antibody–antigen interactions.The passage describes this as an additional function beyond contact-residue prediction.
  • 2.8 proABC-2: proABC-2 differentiates hydrophobic from hydrophilic interactions at antibody–antigen interfaces.This interaction classification is part of the tool’s reported characterization of intermolecular contacts.

2.9 Prodigy · 2.10 SpotON

Prodigy predicts protein–protein binding affinity from three-dimensional complex structures, while SpotON identifies interface hot spots that contribute substantially to binding. Together, the tools support computational characterization of interaction strength and residue-level interface contributions.

  • 2.9 Prodigy: Prodigy predicts the binding affinity of protein–protein complexes from their three-dimensional structures.It addresses the need to estimate interaction strength after structural characterization.
  • 2.9 Prodigy: Prodigy is designed to characterize the strength of interactions in biomolecular complexes.Binding affinity is described as another key property alongside structural characterization.
  • 2.9 Prodigy: Prodigy operates as the PROtein binDIng enerGY Prediction tool.The name identifies its focus on protein-complex binding-energy prediction.
  • 2.10 SpotON: SpotON focuses on determining which interface regions contribute most to binding.These regions are described as interaction “hot spots.”
  • 2.10 SpotON: Interaction hot spots are residues whose alanine mutation changes complex binding free energy by more than 2.0 kcal/mol.The passage defines hot spots using the mutation-induced binding free-energy difference.
  • 2.10 SpotON: SpotON addresses the limitations of experimental hot-spot identification, which is typically low throughput, laborious, and expensive.Experimental identification requires repeated point mutations and binding-energy evaluations.

3 INFRASTRUCTURE

WeNMR portals combine local, distributed HTC, and GPU resources, with EOSC/EGI access governed by registration and certificate policies. Their infrastructure is supported by EGI services, formal resource agreements, federated SSO, and a workload manager for distributed computing.

  • Execution infrastructure: WeNMR portals run on local resources or harvest distributed EOSC/EGI HTC and GPU resources, with some portals requiring registration.GPU resources can accelerate computations.
  • Registration and certificates: EOSC/EGI HTC portals require X509 certificates, while robot certificates can remove the need for users’ personal certificates.The infrastructure distinguishes local, HTC, and GPU execution modes.
  • EGI Federation: >1 millions of CPU-cores and >900 PB of disk and tape storage support EGI’s worldwide, multidisciplinary research infrastructure.EGI federates hundreds of resource centers and serves more than 70,000 researchers.
  • Authentication: All WeNMR portals integrate EGI Check-in as an SSO mechanism, allowing users to register with shared credentials for WeNMR services.The portals are compliant with GDPR and provide clear conditions of use.
  • Workload management: The EGI Workload Manager provides a single user-friendly interface for accessing distributed computing resources of various types through the EOSC Compute Platform.It is built on software from the DIRAC Interware project.

4 COMMUNITY AND USAGE

WeNMR services serve over 150,000 users through tutorials, support mechanisms, and feedback-driven improvements. Usage of EGI HTC resources has grown substantially, reaching ~5,300 CPU-years in 2020, while COVID-related HADDOCK submissions comprised about one-third of submissions since April 2021.

  • Community and usage: Over 150,000 users access the combined WeNMR services, supported by tutorials, request channels, and question-based assistance.User feedback and reported issues drive service improvements and new feature implementation.
  • Community and usage: ~5,300 CPU-years of computing were used in 2020, continuing an increase in EGI HTC resource use since 2009.The increase is shown by normalized monthly CPU usage over time.

5 Conclusions

Over more than a decade, WeNMR has developed advanced computational services into Thematic Services in the European Open Science Cloud using EGI’s high-throughput infrastructure. Its worldwide community has submitted more than 12 million jobs, accounting for approximately 4,000 CPU-years.

  • 5 Conclusions: >12 million jobs submitted in 2020 accounted for ~4,000 CPU-years.These figures quantify the computational activity supported by WeNMR through EGI’s high-throughput computing infrastructure.
  • 5 Conclusions: WeNMR has facilitated advanced computational tools for more than a decade and developed them as Thematic Services in the European Open Science Cloud.The project’s services take advantage of EGI’s high-throughput compute infrastructure and have been continuously improved in collaboration with EGI.

6 Conflict of Interest · 8 Funding

The work was co-funded by three Horizon 2020 projects and supported by an NWO-ENW computing grant. No conflict-of-interest information is provided in the supplied passage.

  • 6 Conflict of Interest: The supplied passage reports no conflict-of-interest statement.
  • 8 Funding: The work was co-funded by EOSC-hub, EGI-ACE, and BioExcel under Horizon 2020.The cited grant numbers are 777536, 101017567, 823830, and 675728, respectively.
  • 8 Funding: A computing grant from NWO-ENW also supported the work.The project number was 2019.053.
Loading 2107.01056v1…