Source-linked AI summary

Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans

Michael Roberts, Derek Driggs, Matthew Thorpe, Julian Gilbey, Michael Yeung, Stephan Ursprung, Angelica I. Aviles-Rivero, Christian Etmann, Cathal McCague, Lucian Beer, Jonathan R. Weir-McCall, Zhongzhao Teng, Effrossyni Gkrania-Klotsas, James H. F. Rudd, Evis Sala, Carola-Bibiane Schönlieb

arXiv:2008.06388v4cs.LGcs.CVeess.IVstat.ML

TL;DR

It was unclear which machine-learning models using chest radiographs or CT scans could support COVID-19 diagnosis or prognosis. This systematic review assessed their methodological quality and found that none were currently suitable for clinical use.

  • Problem

    The clinical utility of machine-learning models for COVID-19 diagnosis and prognosis from chest radiographs or CT scans remained unclear despite their promise.

  • Method

    The authors systematically reviewed 61 studies, assessed bias and methodological quality, and developed recommendations for data, evaluation, reproducibility, reporting, and peer review.

  • Results

    None of the reviewed models met the robustness, reproducibility, and external-validation standards needed to support clinical practice or wider clinical translation.

  • Takeaways & Limitations

    Higher-quality datasets, reproducible documentation, and external validation are needed before these models can advance toward clinical trials and implementation.

  • Takeaways & Limitations

    For 38 of 61 studies using deep-learning-derived features, predictor biases could not be judged because the imaging features were unknown and abstract.

Abstract

from arXiv · show

Machine learning methods offer great promise for fast and accurate detection and prognostication of COVID-19 from standard-of-care chest radiographs (CXR) and computed tomography (CT) images. Many articles have been published in 2020 describing new machine learning-based models for both of these tasks, but it is unclear which are of potential clinical utility. In this systematic review, we search EMBASE via OVID, MEDLINE via PubMed, bioRxiv, medRxiv and arXiv for published papers and preprints uploaded from January 1, 2020 to October 3, 2020 which describe new machine learning models for the diagnosis or prognosis of COVID-19 from CXR or CT images. Our search identified 2,212 studies, of which 415 were included after initial screening and, after quality screening, 61 studies were included in this systematic review. Our review finds that none of the models identified are of potential clinical use due to methodological flaws and/or underlying biases. This is a major weakness, given the urgency with which validated COVID-19 models are needed. To address this, we give many recommendations which, if followed, will solve these issues and lead to higher quality model development and well documented manuscripts.

Introduction

COVID-19 imaging can complement RT-PCR, while machine learning offers potential for diagnosis and prognostication from chest radiographs and CT scans. This review examines methodological challenges and proposes recommendations for dataset creation, algorithm development, reproducibility, reporting, and peer review.

  • Motivation: Imaging can complement RT-PCR for diagnostic certainty or serve as a surrogate where RT-PCR is unavailable.Chest radiograph abnormalities may appear after an initially negative RT-PCR, and chest CT has shown higher sensitivity in several studies.
  • Motivation: Machine learning could improve COVID-19 diagnosis against RT-PCR while providing insight into patient prognostication from multimodal data.The models are presented as potentially exploiting the large amount of multimodal patient data collected in clinical care.
  • Review contribution: The review focuses on challenges specific to classical and deep learning models developed from COVID-19 imaging data.It extends earlier broad reviews by emphasizing the unique issues researchers face when using imaging data.
  • Review contribution: The recommendations address public imaging datasets, algorithm methodology, reproducibility, manuscript documentation, and peer review.These five domains are intended to improve model development and the documentation of published studies.

PPV: 1·00

The section reports high diagnostic performance for several CT-based deep-learning models, including perfect or near-perfect internal metrics. External validation and prognosis results were more variable, with lower discrimination and specificity in some evaluations.

  • Diagnostic CT models: AUC was 0·99, with Sensitivity 1·00 and Specificity 0·99, in another diagnostic CT model.
  • Diagnostic CT models: Another external evaluation reported AUC 0·81 [0·71,0·84], below its internal validation performance.
  • Diagnostic CT models: AUC reached 1·00 with Accuracy, Sensitivity, and Specificity all 1·00 in one diagnostic CT model.
  • Prognosis and severity: For prognosis and severity assessment, reported results included Correlation 0.98 and AUC 0.91, Accuracy 0.90, Sensitivity 0.83, and Specificity 0.90.
  • Diagnostic CT models: A further diagnostic model reported Accuracy 0·76, Sensitivity 0·74, and Specificity 0·79.

Wang et al. 54

The section describes machine-learning approaches using radiomic features and clinical factors, with internal and external validation reported for COVID-19 imaging studies. The studies include prognosis tasks involving severity, lung opacity, and lung involvement.

  • Models incorporated radiomic features and clinical factors.
  • 709 images, including 560 COVID-19, were used in one validation set.
  • Prognosis: Prognosis studies evaluated severity using chest radiographs and deep learning on images of differing severities.
  • External validation included 111 images and 113 images in separate studies.
  • Prognosis: Other methods addressed lung opacity and the extent of lung involvement, using CNN-derived features extracted at various layers.

Yue et al. 83

The section concerns prognosis for COVID-19 patients, including whether they convert to a severe stage and their hospital stay.

  • Hospital stay is also considered for COVID-19 patients.
  • The prognosis concerns whether COVID-19 patients will convert to a severe stage.

Zhu et al. 77

The section summarizes prognostic models using chest imaging-derived radiomic features and clinical data to predict severe COVID-19 outcomes. Reported evaluations included internal and external validation, with one model achieving an AUC of 0·86±0·02.

  • The prognostic model predicted risk of death, need for ventilation, or requirement for overventilation.
  • 47 patients of varying severity were included in the reported cohort.
  • AUC: 0·86±0·02, Accuracy: 0·86±0·02, Sensitivity: 0·77±0·03, Specificity: 0·88±0·015.These results were reported for internal validation involving 6 patients, all long-term.
  • Short-term prognosis outcomes included intubation and death, using radiomic features and clinical data.

Wu et al. 80

The section concerns prognosticating whether patients require admission to an intensive care unit (ICU).

  • The model addresses prognosis for ICU admission.
  • The prognostic target is admission to an ICU.
  • ICU admission is the clinical outcome considered in this section.

Zheng et al. 81

Zheng et al. developed a COVID-19 prognostic model using radiomic features and clinical data to predict severe short-term outcomes, including mechanical ventilation or death. Reported performance varied across analyses, with external validation yielding a C-index of 0·89.

  • The model combined radiomic features with clinical data for COVID-19 prognosis.
  • Severe short-term outcomes included use of mechanical ventilation or death.
  • Precision (w): 0·94, Sensitivity (w): 0·94, Specificity (w): 0·81.
  • Precision (w): 0·77, Sensitivity (w): 0·94, Specificity (w): 0·82.
  • AUC: 0·86, Sensitivity: 0·80, Specificity: 0·86.
  • External validation reported a C-index of 0·89, although the paper's description was unclear.

Chen et al. 82

Chen et al. 82 used radiomic features and clinical data with internal holdout validation. The extracted results report diagnostic and prognostic performance across several image and patient cohorts.

  • The models combined radiomic features with clinical data.
  • Accuracy was 0·88, Sensitivity was 0·55, and Specificity was 0·95.
  • The extracted cohorts included 247 images with 36 severe cases, 105 images with 15 severe cases, 24 images with unclear severe-case counts, and 135 patients with unclear non-survivor counts.
  • AUC was 0·93, Accuracy was 0·91, Sensitivity was 0·81, and Specificity was 0·95.
  • C-index ranged from [0.92, 0.95], while Accuracy ranged from [0.85, 0.87], Sensitivity from [0.71,0.76], and Specificity from [0.91,0.92].

Appendix · 1. Search strategy

The search began with date-filtered arXiv papers containing COVID-19-related terms and all papers from the Living Evidence on COVID-19 resource. Identified papers were then refined using machine-learning and chest-imaging terms, with the search widened slightly from the original PROSPERO description.

  • 1. Search strategy: ArXiv papers were initially extracted when their titles or abstracts contained “ncov”, “coronavirus”, “covid”, “sars-cov-2” or “sars-cov2”.The arXiv extraction covered the relevant date ranges.
  • 1. Search strategy: All papers in the appropriate date range were downloaded from the “Living Evidence on COVID-19” resource.This was part of the initial extraction stage.
  • 1. Search strategy: The refined search required titles or abstracts to contain machine-learning terms such as “ai”, “deep”, “learning”, “machine” or “neural”.Additional terms included “intelligence”, “prognos”, “diagnos”, “classification” and “segmentation”.
  • 1. Search strategy: The refined search also required chest-imaging terms including “ct”, “cxr”, “x-ray”, “xray”, “imaging”, “image” or “radiograph”.Only “ai”, “ct” and “cxr” had to appear as complete words.
  • 1. Search strategy: The search differed slightly from the original PROSPERO description because it was widened to identify additional papers.The paper provides the full searching and filtering history through its source code.
  • 1. Search strategy: The full history of searching and filtering is available in the cited source code repository.The repository documents the search and filtering process.

2. Supplementary author list Name Affiliation

The supplementary author list spans hospitals, universities, research centers, and companies across the UK, Europe, and China. Cambridge-based institutions and organizations are particularly prominent among the listed affiliations.

  • Cambridge affiliations include Royal Papworth Hospital, Qureight Ltd, Addenbrooke’s Hospital, and the University of Cambridge.
  • The list includes clinical and academic affiliations in Vienna, Helsinki, London, Macau, Munich, and Wuhan.
  • Industry affiliations include AINOSTICS Ltd, SparkBeyond UK Ltd, contextflow GmbH, AstraZeneca, and Aladdin Healthcare Technologies Ltd.
  • Additional affiliations represent biomedical imaging, radiology, medical computing, chemistry, astronomy, and systems oncology research.
Loading 2008.06388v4…