Source-linked AI summary

Deep learning for drug repurposing: methods, databases, and applications

Xiaoqin Pan, Xuan Lin, Dongsheng Cao, Xiangxiang Zeng, Philip S. Yu, Lifang He, Ruth Nussinov, Feixiong Cheng

arXiv:2202.05145v1q-bio.BMcs.LG

TL;DR

Drug development is costly and slow, while integrating large biomedical datasets for repurposing remains challenging. This review surveys databases, representations, deep learning models, applications, and future challenges, highlighting the field's potential and current validation gap.

  • Problem

    Drug development is time-consuming and expensive, and comprehensively obtaining and integrating biomedical data remains challenging for drug repurposing.

  • Method

    The review synthesizes drug-repurposing databases, sequence- and graph-based representations, deep learning models, and applications.

  • Results

    Deep learning methods have shown great potential for drug repurposing, including applications aimed at COVID-19.

  • Takeaways & Limitations

    The review provides guidelines for using deep learning methodologies and tools in drug repurposing and identifies future challenges.

  • Takeaways & Limitations

    Predictions in the reviewed COVID-19 applications were not validated.

Abstract

from arXiv · show

Drug development is time-consuming and expensive. Repurposing existing drugs for new therapies is an attractive solution that accelerates drug development at reduced experimental costs, specifically for Coronavirus Disease 2019 (COVID-19), an infectious disease caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). However, comprehensively obtaining and productively integrating available knowledge and big biomedical data to effectively advance deep learning models is still challenging for drug repurposing in other complex diseases. In this review, we introduce guidelines on how to utilize deep learning methodologies and tools for drug repurposing. We first summarized the commonly used bioinformatics and pharmacogenomics databases for drug repurposing. Next, we discuss recently developed sequence-based and graph-based representation approaches as well as state-of-the-art deep learning-based methods. Finally, we present applications of drug repurposing to fight the COVID-19 pandemic, and outline its future challenges.

Graphical/Visual Abstract and Caption

The deep learning drug-repurposing pipeline moves from data creation and feature representation through model evaluation to repurposing predictions.

  • High-quality data sources are created for compounds, proteins, and diseases.
  • Informative feature vectors are generated using graph-, sequence-, and text-based representations.
  • Deep learning models are built and evaluated before repurposing tasks are conducted.
  • The pipeline supports predictions of drug-target binding affinity, drug-target interaction, compound-protein interaction, and drug-disease associations.

1. INTRODUCTION

Drug repurposing can reduce the time, cost, and risk of developing therapies, while deep learning offers tools for integrating expanding biomedical data. This review surveys databases, representations, models, applications, and future challenges.

  • Developing a candidate drug typically takes 10-15 years and 0.8-1.5 billion dollars.
  • Drug repurposing uses existing drugs to seek new indications and can bypass many preapproval tests required for newly developed compounds.
  • Drug repurposing offers lower failure risk, less investment, and a shorter development timeframe.
  • Large-scale chemical, genomic, phenotypic, and omics data create a critical need for effective domain-data exploration in repurposing.
  • Deep learning automatically learns multiple representation levels from input data and has shown potential for drug repurposing, although applications remain at an infancy stage.
  • The review integrates drug-repurposing databases, sequence- and graph-based representations, target- and disease-based models, applications, and future challenges.

2. DATABASE

Drug-repurposing databases span chemical, biomolecular, drug-target interaction, and disease information. The review emphasizes accessible resources and illustrates how heterogeneous data can support prediction.

  • PubChem contains 110 million compounds, 271 million substances, and 297 million bioactivities in three dynamically growing primary databases.
  • The reviewed databases fall into four categories: chemical, biomolecular, drug-target interaction, and disease databases.
  • Researchers should prioritize publicly available datasets that are easy to download or access through an API for deep learning integration.
  • Database selection can involve choosing inputs from multiple sources and conducting comparative analysis across databases.
  • CDRscan integrates CCLE genomic data, GDSC drug-response assays, structural fingerprints, and DrugBank QSAR information to predict anticancer activity.
  • CDRscan identified 14 oncology and 23 non-oncology drugs with new potential cancer indications from 1,487 approved drugs.

3. REPRESENTATION LEARNING

Representation learning enables deep learning models to extract useful features from raw drug, protein, and biological-network data. The review organizes these representations into sequence-based and graph-based approaches, describing their encodings, applications, and limitations.

  • Overview: Representation learning automatically discovers representations for feature extraction or classification from raw data, supporting end-to-end deep learning.Deep learning integrates representation learning into feature design so useful information can be extracted from input data.
  • Sequence-based representation: Sequence-based methods encode drugs as SMILES or molecular fingerprints and proteins as amino-acid sequences, distance maps, or learned sequence embeddings.RNNs, CNNs, and language-model-inspired methods learn latent or distributed features from molecular and protein sequences.
  • Limitations and future directions: 1D and 2D compound representations lose bond-length and 3D-conformation information that may matter for drug–target binding details.The review identifies 3D representations as a future direction, although they require protein or target structural data and costly molecular-docking simulations.
  • Limitations and future directions: Protein sequence representations can ignore physical, chemical, biological, and three-dimensional structural information despite capturing sequence order.Alternative featurization uses protein physical, chemical, and biological properties, while graph approaches incorporate structural information.
  • Graph-based representation: Graph-based representation approaches have emerged as a solution for improving drug-repurposing performance because compounds, proteins, and their associations naturally form graphs or networks.Large-scale heterogeneous biological networks provide opportunities for graph-based learning in drug repurposing.
  • Graph-based representation: Graph-based methods represent molecular atoms and bonds, protein atoms, or biological-network structure as nodes and edges for vector learning and downstream prediction.GNNs and network-embedding methods capture local or global structural properties for tasks including node classification, clustering, and link prediction.

4. DEEP LEARNING MODELS FOR DRUG REPURPOSING

Deep-learning drug-repurposing tools predict drug-target or drug-disease interactions through target-centered and disease-centered approaches, using sequence, graph, network, and other representations. Reported methods combine heterogeneous biomedical data with deep models and include computational predictions supported by experimental, clinical, or literature validation.

  • Drug-repurposing tools usually predict unknown drug-target or drug-disease interactions through target-centered or disease-centered methods.
  • Target-centered models: Target fishing encodes drug chemical structures to screen targeted proteins, but a single predicted target cannot fully describe disease characteristics.
  • Target-centered models: Sequence-based interaction models combine neural architectures such as GNNs, CNNs, LSTMs, GCNs, and transformers to represent compounds and proteins.
  • Target-centered models: TransformerCPI achieved the best performance under more rigorous label-inversion evaluation, addressing limitations associated with sequence splitting and hidden ligand bias.
  • Target-centered models: Network and knowledge-graph methods learn drug and target representations from heterogeneous biomedical networks to predict novel drug-target interactions.
  • Target-centered models: deepDTnet achieved an AUC-ROC of 0.963 for identifying novel molecular targets, exceeding random forest (0.911), SVM (0.869), k-nearest neighbors (0.839), and Naive Bayes (0.783).
  • Validation and applications: Predicted candidates from several methods were supported by clinical trials, published studies, experimental assays, or associations with approved treatments.
  • Disease-centered models: Disease-centered methods integrate drug-disease associations and disease features, with deepDR achieving AUROC 0.908 and a 4.6% absolute gain over DTINet.

5. APPLICATIONS OF DRUG REPURPOSING

Deep learning-based repurposing approaches have prioritized host-targeting and antiviral candidates for COVID-19, with several predictions supported by computational, registry, or clinical evidence. These applications span knowledge graphs, multimodal models, and neural networks, but predicted drugs still require experimental and randomized-trial validation.

  • Host-targeting therapies: Autoencoder integration of transcriptomic, proteomic, and structural data prioritized doxapram, dasatinib, and ribavirin for older individuals with COVID-19.Serine/threonine and tyrosine kinases were highlighted as potential targets intersecting SARS-CoV-2 and aging pathways.
  • Antiviral therapies: MathDL ranked binding affinities across 137 SARS-CoV-2 Mpro crystal structures and identified 71 candidate covalent-bonding inhibitors.The method combines algebraic topology with deep learning for antiviral drug discovery.
  • Validation requirements: Predicted COVID-19 candidates had not been validated by preclinical models or randomized controlled clinical trials and therefore require experimental assays and trials before patient use.This validation boundary applies to recommendations for use in patients with COVID-19.
  • Clinical and real-world evidence: Baricitinib was identified using BenevolentAI’s knowledge graph and was later associated with reduced mortality in hospitalized adults in a phase 3 randomized placebo-controlled trial.The passage describes this as the first successful example of deep learning approaches for COVID-19 drug-repurposing development.

CONCLUDING REMARKS AND FUTURE CHALLENGES

The review identifies data quality, interpretability, clinical validation, and uncertainty about deep learning’s superiority as continuing challenges. It recommends improving data sharing and quality while matching methods to the specific repurposing problem and preserving comparative evaluation.

  • Data quality and availability: Deep learning requires large-scale datasets, but biomedical data are often noisy, incomplete, inaccurate, and expensive to annotate manually.The review presents community efforts and open-data platforms as potential ways to increase data reuse and improve compound bioactivity resources.
  • Methodological directions: Transfer learning, active learning, and semi-supervised few-shot learning are discussed for heterogeneous, scarce, or weakly labeled biomedical data.Active learning iteratively queries important unlabeled samples and labels them for subsequent training.
  • Precision medicine: Large-scale omics data can support patient stratification and identification of subtype-specific repurposable drugs for precision medicine.The data sources include genetics, genomics, transcriptomics, proteomics, and metabolomics from patient samples or disease models.
  • Data quality and availability: DREAM communities showed that data usage can sometimes contribute more to prediction accuracy than model choice.The review links progress to open data sharing and collaboration among chemistry, pharmacology, data science, computer science, and drug-discovery experts.
  • Interpretability: Interpretability remains a major challenge because complex neural networks often do not provide biological explanations for drug-discovery communities.Attention mechanisms and visualization of networks or internal mechanisms are proposed directions.
  • Comparative performance: Deep learning’s superiority over other machine-learning methods remains unresolved; for structured input descriptors, it appears at least on par with other methods.The review advises investigating the advantages and limitations of specialized architectures rather than relying exclusively on deep learning.
  • Clinical translation: Clinical-trial use of deep learning in drug repurposing has yet to be demonstrated, and method selection may depend on the problem and modeler familiarity.The review states that firm conclusions about superiority are premature.
Loading 2202.05145v1…