Source-linked AI summary

Network Medicine Framework for Identifying Drug Repurposing Opportunities for COVID-19

Deisy Morselli Gysi, Ítalo Do Valle, Marinka Zitnik, Asher Ameli, Xiao Gan, Onur Varol, Susan Dina Ghiassian, JJ Patten, Robert Davey, Joseph Loscalzo, Albert-László Barabási

arXiv:2004.07229v2q-bio.MNcs.LGq-bio.QMstat.ML

TL;DR

The paper addresses the need for rapid, reliable prioritization of clinically approved drugs for SARS-CoV-2 infection. It compares artificial-intelligence, network-diffusion, and network-proximity rankings of 6,340 drugs against experimental and clinical-trial evidence, then develops a consensus approach. The consensus ranking consistently exceeds individual pipelines, while most effective screened drugs act through host-network perturbations rather than direct viral-target binding.

  • Problem

    Reliable methods for rapidly prioritizing clinically approved drugs for COVID-19 were needed because de novo drug development takes a decade or longer.

  • Method

    The study applies three network-medicine algorithm families to rank 6,340 drugs and fuses their predictions through a multimodal consensus approach.

  • Results

    The consensus ranking consistently exceeds individual pipelines, with a 9% hit rate among the top 100 drugs and 58 of 77 strong- or weak-effect drugs among the top 800.

  • Takeaways & Limitations

    Network-based strategies can prioritize effective COVID-19 drugs that direct binding-based methods would miss, supporting repurposing for future pathogens and neglected diseases.

  • Takeaways & Limitations

    Only 918 of the 6,340 drugs prioritized by CRank were experimentally screened, so evaluation covered a selected subset of the ranked library.

Abstract

from arXiv · show

The current pandemic has highlighted the need for methodologies that can quickly and reliably prioritize clinically approved compounds for their potential effectiveness for SARS-CoV-2 infections. In the past decade, network medicine has developed and validated multiple predictive algorithms for drug repurposing, exploiting the sub-cellular network-based relationship between a drug's targets and disease genes. Here, we deployed algorithms relying on artificial intelligence, network diffusion, and network proximity, tasking each of them to rank 6,340 drugs for their expected efficacy against SARS-CoV-2. To test the predictions, we used as ground truth 918 drugs that had been experimentally screened in VeroE6 cells, and the list of drugs under clinical trial, that capture the medical community's assessment of drugs with potential COVID-19 efficacy. We find that while most algorithms offer predictive power for these ground truth data, no single method offers consistently reliable outcomes across all datasets and metrics. This prompted us to develop a multimodal approach that fuses the predictions of all algorithms, showing that a consensus among the different predictive methods consistently exceeds the performance of the best individual pipelines. We find that 76 of the 77 drugs that successfully reduced viral infection do not bind the proteins targeted by SARS-CoV-2, indicating that these drugs rely on network-based actions that cannot be identified using docking-based strategies. These advances offer a methodological pathway to identify repurposable drugs for future pathogens and neglected diseases underserved by the costs and extended timeline of de novo drug development.

Introduction

The paper evaluates network-medicine algorithms for rapidly prioritizing repurposable COVID-19 treatments and finds that no individual method performs consistently across ground-truth datasets. A multimodal consensus approach improves prediction reliability while identifying effective drugs through network-based actions rather than direct viral-target binding.

  • Motivation: The pandemic motivates drug repurposing because de novo drug development lasts a decade or longer, making rapid prioritization of approved compounds necessary.The authors note that unreliable methodologies concentrated clinical-trial resources on hydroxychloroquine or chloroquine rather than a wider range of candidates.
  • Algorithm evaluation: Most algorithms showed predictive power, but performance varied across datasets and metrics, preventing identification of one consistently trustworthy algorithm.Proximity performed better for experimental outcomes, whereas AI pipelines performed strongly for clinical-trial drugs.
  • Consensus prediction: A multimodal ensemble seeks consensus across predictive methods and improves the accuracy and reliability of drug-ranking predictions.The approach extracts joint predictive power from all pipelines because no individual pipeline performs consistently best across outcomes.
  • Network-based mechanisms: 76 of 77 experimentally effective drugs do not directly bind SARS-CoV-2-targeted proteins, instead acting indirectly by perturbing the host subcellular network.These network drugs cannot be identified by traditional binding-based methods but are successfully prioritized by network-based methods.
  • Network proximity: Effective strong- and weak-effect drug targets cluster near the COVID-19 disease module, while ineffective drug targets are farther away than expected by chance.Strong-effect targets are closer than weak-effect targets, making network proximity a positive predictor of efficacy.

Discussion

The discussion presents CRank as a resource-efficient strategy for prioritizing experimentally testable COVID-19 drug candidates, while emphasizing that screening and cell-model evidence remain incomplete. Combining algorithmic consensus with expert knowledge improved enrichment and highlighted approved drugs for possible clinical consideration.

  • Resource prioritization: 0.8% was the hit rate for brute-force screening, compared with 9% among the top 100 drugs prioritized by CRank.The top 800 CRank-ranked drugs contained 58 of the 77 S&W drugs.
  • Candidate drugs: Azelastine and digoxin were highly ranked drugs with strong outcomes but were not yet in clinical trials, supporting consideration of both for clinical trials.Azelastine ranked CRank #10 and digoxin #33; folic acid, methotrexate, omeprazole, fluvastatin, ivermectin, and sildenafil were additional highlighted candidates.
  • Candidate mechanisms: Methotrexate’s anti-inflammatory mechanism was presented as potentially relevant to profound infection-associated hyperimmune responses.Omeprazole was also described as altering lysosome acidification and, with other benzimidazoles, binding nsp3, which interferes in viral formation.
  • Broader implications: The methodological advances provide an algorithmic toolset for identifying potential treatments for future diseases underserved by conventional de novo drug-discovery costs and timelines.The discussion frames the approach as relevant beyond COVID-19 drug candidates.
  • Scope boundaries: Only 918 of the 6,340 CRank-prioritized drugs were screened, and selection driven by compound availability left many potentially efficacious FDA-approved drugs untested.The discussion therefore treats the experimental dataset as an incomplete sample of the prioritized drug space.
  • Model and screening boundaries: Some drugs inactive in VeroE6 cells may show efficacy in human cells, as illustrated by loratadine’s activity in Caco-2 cells.Ritonavir was top-ranked but inactive in the screen despite being explored in more than 42 clinical trials.

Authors Contribution

The authors’ contributions covered study design, drug prediction, experimentation, data analysis, candidate curation, and manuscript preparation. Individual contributors were assigned specialized analytical, experimental, and coordination roles.

  • Experiments and comorbidities: I.D.V analyzed disease comorbidities, and R.D. and J.J.P. performed the drug experiments and experimental screens.These contributions covered comorbidity analysis and experimental validation.
  • Data analysis and curation: A.A., D.M.G., I.D.V., M.Z., O.V., X.G., and J.L. analyzed the data, while J.L. manually curated drug candidates.O.V. also curated the list of drugs in COVID-19 clinical trials.
  • Writing and oversight: A.L.B., D.M.G., I.D.V., and X.G. wrote the paper with input from all authors, and all authors approved the manuscript.S.D.G. guided A.A. in designing diffusion-based similarity implementations.
  • Study design and prediction: A.L.B designed the study, while A.A., D.M.G., M.Z., and X.G. performed drug predictions.The study design and prediction activities were assigned to distinct contributors.

Declaration of interests

The paper discloses author affiliations, consulting relationships, and a scientific-founder relationship involving Scipher Medicine. It also presents the study’s drug-repurposing framework, experimental validation, and predictive pipelines.

  • J.L. and A.L.B. are co-scientific founders of Scipher Medicine, Inc.
  • A.L.B. is the founder of Nomix Inc. and Foodome, Inc., which apply data science to health.
  • O.V. and D.M.G. are scientific consultants for Nomix Inc., while I.D.V. is a scientific consultant for Foodome Inc.
  • The study combines artificial intelligence, network diffusion, and network proximity into 12 predictive drug-ranking pipelines.
  • The framework validates predictions using 918 experimentally screened drugs and clinical-trial drug lists.
  • Figure 2 reports a 208-protein connected component among SARS-CoV-2 targets, while Figure 4 evaluates pipeline and rank-aggregation performance.

Supplementary Information: Network Medicine Framework for

The supplementary information organizes the paper’s network, disease, algorithmic, experimental, and validation procedures. It also includes a section on prediction-algorithm complementarity.

  • The supplementary information covers the human interactome, SARS-CoV-2 and drug targets, lung gene expression, and disease comorbidities.
  • It describes drug-repurposing prediction algorithms based on artificial intelligence, diffusion, and proximity.
  • It examines network properties of prediction algorithms, including explanatory subgraphs and complementarity.
  • The experimental procedures include cell cultivation, virus infection inhibition assays, drug-response classification, and biological interpretation of effective drugs.

3 Statistical Validation ............................................................................................................................ 23

The paper’s statistical validation section evaluates drug-repurposing predictions using ROC curves, precision, and recall. These metrics provide the stated framework for performance evaluation.

  • Statistical validation uses ROC curves to evaluate predictive performance.
  • Precision is included among the performance metrics for evaluating drug-repurposing predictions.
  • Recall is included alongside ROC curves and precision in the evaluation framework.

4 Rank Aggregation Algorithms (RAAs) .................................................................................................. 27

The paper evaluates several rank-aggregation algorithms for combining drug-repurposing rankings. The section covers average rank, Borda, Dowdall, CRank, and comparisons among these methods.

  • The average rank method is presented as a rank-aggregation algorithm.
  • The Borda method is presented as a rank-aggregation algorithm.
  • The Dowdall method is presented as a rank-aggregation algorithm.
  • CRank is presented as a rank-aggregation algorithm.
  • The paper includes a comparison of rank-aggregation algorithms.

1 Network-Based Drug Repurposing For COVID-19

The study builds network-based representations of SARS-CoV-2, human disease modules, drugs, and their targets to identify repurposing opportunities. Network analyses found limited direct overlap with major disease modules and complementary predictive subgraphs across artificial-intelligence, diffusion, and proximity methods.

  • Disease comorbidity: 110 of 332 SARS-CoV-2-targeted proteins were implicated in other diseases, but their gene overlap was not statistically significant.The reported Fisher’s exact test used FDR-BH padj > 0.05.
  • Disease comorbidity: The closest disease modules included cardiovascular diseases and cancers, while network evidence also linked COVID-19 targets to neurological diseases.The neurological connection was consistent with host-target expression in the brain and reported neurological manifestations.
  • Implications for repurposing: The host targets did not overlap with proteins associated with any major disease, motivating drug-target mapping without restricting targets to a particular disease module.The authors therefore emphasize network-based relationships beyond direct disease-module overlap.
  • Complementary predictive methods: AI, proximity, and diffusion methods explored complementary regions of the protein-interaction network and selected different gene neighborhoods.AI methods overlapped drug targets but were separated from the COVID-19 module, whereas proximity-based methods showed the opposite pattern; diffusion methods avoided both.

2 Experimental Validation

Experimental screening in VeroE6 cells identified drugs that reduced SARS-CoV-2 infection, followed by target and pathway analyses of the effective compounds. These analyses found no shared drug category or universal target-binding pattern explaining efficacy.

  • Screening outcome: 77 drugs showed strong or weak effects in the high-throughput SARS-CoV-2 screening.Strong and weak effects were subsequently analyzed for shared target and pathway profiles.
  • Target and category analysis: No ATC drug category was enriched among strong-, weak-, or strong-and-weak-effect drugs.The enrichment test reported FDR-BH padj > 0.05.
  • Target-profile clustering: Hierarchical clustering failed to identify binding patterns shared by all effective drugs and revealed only four small groups with shared targets.Three groups contained drugs from multiple categories, while one comprised seven nervous-system-related drugs.
  • Pathway enrichment: 42 of the 77 strong- or weak-effect drugs belonged to three groups associated with common pathways, including 20 drugs with diverse indications.The pathway analysis therefore identified partial rather than universal grouping among effective drugs.
  • Interpretation: Neither pathway nor target analysis revealed patterns that could explain the efficacy of the 77 strong- or weak-effect drugs.Some enriched targets involved membrane receptors such as ADRA1A, HTR2A, and HRH1, many associated with nervous-system disorders.

3 Statistical Validation

The study evaluated drug-ranking pipelines against experimental screening and clinical-trial ground truths using whole-list and top-K metrics. Most methods showed predictive power, while rank aggregation produced stronger performance than individual pipelines in the reported comparisons.

  • Evaluation metrics: AUC values range from 0 to 1, with 1 indicating perfect performance and 0.5 indicating a random classifier.The study used ROC curves and AUC scores to assess predictive power.
  • Evaluation metrics: AUC measured whole-ranked-list discrimination, while precision@K and recall@K evaluated positive drugs among the top-K predictions.These metrics address both overall ranking quality and prioritization near the top of each list.
  • Ground truths: The CT415 ground truth comprised 67 drugs in ongoing COVID-19 clinical trials as of April 2020.Trial drug names were matched to DrugBank records.
  • Evaluation caveats: Reported precision and recall were conservative estimates because rankings covered all 6,340 drugs rather than only experimentally screened compounds.The authors state that restricting analysis to screened drugs could yield higher values.
  • Pipeline comparison: Rank aggregation methods showed strong predictive performance, with CRank generally outperforming individual pipelines in top precision and recall comparisons.The reported comparison covered E74 and CT615 ground truths, and CT06 usually had higher hit rates, precision, and recall than E74.

4 Rank Aggregation Algorithms (RAAs)

The section introduces rank aggregation methods for combining pipeline rankings into consensus drug lists, then evaluates their agreement with experimental and clinical ground truths. CRank accommodates uncertainty and variable pipeline importance, achieving the strongest agreement for top-ranked drugs in the reported comparisons.

  • Consensus ranking: Kemeny consensus maximizes pairwise agreements between a final ranking and its input rankings but is NP-hard to compute.This motivates heuristic and approximate rank aggregation methods.
  • Average Rank: Average Rank computes each drug’s mean position across the 12 pipeline rankings, but the authors report that average ranks can be a poor aggregation approach.The method is described as a straightforward, ad-hoc strategy.
  • Borda: Borda assigns each drug points according to how many drugs it outranks in each ranking, sums those points, and sorts drugs by the resulting count.The method is theoretically a 5-approximation algorithm for the Kemeny-optimal ranking.
  • Dowdall: Dowdall scores a drug by the reciprocal of its rank in each pipeline and sorts candidates by their total score.A first-place rank receives 1, second place 1/2, and third place 1/3.
  • Evaluation: For K<250, CRank had the best agreement with experimental outcomes, including KS = 0.2679 versus KS = 0.4545 for Dowdall, while Borda reached KS = 0.7131 at K = 100 in E918.The optimal Kemeny consensus remains unknown because it is NP-hard, so comparisons use experimental and clinical ground-truth rankings.

5 Supplementary Tables

The supplementary tables document the human and SARS-CoV-2 interactomes, drug-target and gene-expression resources, learned embeddings, evaluation datasets, and pipeline rankings used in the study.

  • Interactome data: Table S3 contains 332,749 human protein-protein interactions among 18,508 human proteins, while Table S4 lists interactions between 29 SARS-CoV-2 proteins and human proteins.These tables provide the human and viral-human interactome resources.
  • Drug and gene data: Table S5 lists drugs and their respective DrugBank targets, and Table S6 lists 17,222 differentially expressed genes from exposure of 793 drugs across cell types.The supplementary resources support drug-target and drug-induced gene-expression analyses.
  • Network overlap: Table S7 reports network overlap between 299 diseases and SARS-CoV-2 targets using the Svb measure.The table describes disease–viral-target network overlap.
  • Embedding vectors: Tables S8 and S9 provide disease and drug embedding vectors learned by the GNN model, with each row representing one disease or drug embedding.These tables document the learned representations used by the artificial-intelligence pipeline.
  • Evaluation datasets: Tables S10–S13 contain the E918 and E74 experimental datasets and the C415 and C615 clinical-trial drug lists, including experimental outcomes in VeroE6 cells.The clinical-trial lists are dated April and June 2020, respectively.
  • Drug rankings: Table S14 provides rankings from the 12 pipelines and their aggregation, while the top 10% of purchasable predictions were available for purchase.The table records the individual and aggregated drug rankings.
Loading 2004.07229v2…