Source-linked AI summary

Artificial intelligence in digital pathology: a systematic review and meta-analysis of diagnostic test accuracy

Clare McGenity, Emily L Clarke, Charlotte Jennings, Gillian Matthews, Caroline Cartlidge, Henschel Freduah-Agyemang, Deborah D Stocken, Darren Treanor

arXiv:2306.07999v3physics.med-phcs.AIcs.CVeess.IVq-bio.QM

TL;DR

Reliable evidence is needed before AI diagnostic tools applied to digital pathology images enter clinical use. This systematic review and meta-analysis synthesized diagnostic accuracy across whole slide images and disease types, finding high average performance but substantial variability and frequent risk of bias. The findings support continued evaluation using more rigorous and consistently reported study designs.

  • Problem

    Routine clinical use of pathology AI remains rare, while evidence quality, risk of bias, and diagnostic performance remain concerns.

  • Method

    The review synthesized diagnostic accuracy studies of AI applied to whole slide images across disease types, using histopathology and/or immunohistochemistry as reference standards and a bivariate random effects model.

  • Results

    AI showed mean sensitivity of 96.3% and mean specificity of 93.3% across 48 meta-analysed studies representing 50 assessments.

  • Takeaways & Limitations

    AI offers high diagnostic accuracy across whole-slide-image applications and disease types, but its performance requires more rigorous evaluation before clinical adoption.

  • Takeaways & Limitations

    Substantial heterogeneity and deficient reporting limited interpretation, and 99% of included studies had at least one area at high or unclear risk of bias.

Abstract

from arXiv · show

Ensuring diagnostic performance of AI models before clinical use is key to the safe and successful adoption of these technologies. Studies reporting AI applied to digital pathology images for diagnostic purposes have rapidly increased in number in recent years. The aim of this work is to provide an overview of the diagnostic accuracy of AI in digital pathology images from all areas of pathology. This systematic review and meta-analysis included diagnostic accuracy studies using any type of artificial intelligence applied to whole slide images (WSIs) in any disease type. The reference standard was diagnosis through histopathological assessment and / or immunohistochemistry. Searches were conducted in PubMed, EMBASE and CENTRAL in June 2022. We identified 2976 studies, of which 100 were included in the review and 48 in the full meta-analysis. Risk of bias and concerns of applicability were assessed using the QUADAS-2 tool. Data extraction was conducted by two investigators and meta-analysis was performed using a bivariate random effects model. 100 studies were identified for inclusion, equating to over 152,000 whole slide images (WSIs) and representing many disease types. Of these, 48 studies were included in the meta-analysis. These studies reported a mean sensitivity of 96.3% (CI 94.1-97.7) and mean specificity of 93.3% (CI 90.5-95.4) for AI. There was substantial heterogeneity in study design and all 100 studies identified for inclusion had at least one area at high or unclear risk of bias. This review provides a broad overview of AI performance across applications in whole slide imaging. However, there is huge variability in study design and available performance data, with details around the conduct of the study and make up of the datasets frequently missing. Overall, AI offers good accuracy when applied to WSIs but requires more rigorous evaluation of its performance.

Background

Pathology AI has expanded rapidly across whole-slide-image applications, but routine clinical use remains rare because evidence quality and reporting are inconsistent. This review evaluates diagnostic accuracy across disease areas while examining study design and bias.

  • Background: AI applications in whole-slide imaging have expanded across pathology, especially cancer, but routine clinical use remains rare.The evidence base is accompanied by concerns about performance, evidence quality, and risk of bias.
  • Objective and scope: This review assesses the diagnostic accuracy of AI applied to whole slide images across disease types using histopathological assessment and/or immunohistochemistry as reference standards.Eligible studies evaluated disease detection or subtype classification from WSIs.
  • Evidence base: 2976 abstracts yielded 100 included studies, with 48 contributing to the full meta-analysis.Study selection followed screening of 1666 records after duplicate removal and review of 296 full texts.
  • Evidence base: Over 152,000 whole slide images were represented across the 100 studies, although dataset sizes varied from 4 to nearly 30,000 WSIs.Test-set sizes were frequently unavailable; where reported, they ranged from 10 to nearly 14,000 WSIs.
  • Evidence quality: Study designs and reporting varied substantially, while 99% of included studies had at least one area at high or unclear risk of bias.Common concerns included non-random or unclear case selection, absent external validation, and unclear separation of training and testing data.
  • Diagnostic performance: The meta-analysis found mean sensitivity of 96.3% and mean specificity of 93.3% across studies and disease types.The corresponding confidence intervals were 94.1-97.7 for sensitivity and 90.5-95.4 for specificity.

DISCUSSION

AI models generally showed high diagnostic accuracy across whole slide imaging and many disease types, but substantial heterogeneity, incomplete reporting, and pervasive risk of bias limit interpretation. More rigorous, transparent evaluation with diverse datasets and external validation is needed before broader clinical use.

  • AI showed high sensitivity and specificity across diagnostic tasks and disease types in whole slide images.
  • Urological studies had the most favorable subgroup results, while breast studies showed lower diagnostic accuracy than some other specialties.
  • Liver, lymphoma, melanoma, pancreatic, brain, lung, and rhabdomyosarcoma models also demonstrated high sensitivity and specificity.
  • Sensitivity and specificity were higher with more data sources and external validation, although few studies used more than two sources.Diverse datasets are important for models intended to generalize across institutions and populations.
  • 99% of the 100 included papers had at least one area at high or uncertain risk of bias.Study designs, datasets, analysis units, metrics, and reporting detail varied substantially, complicating interpretation of diagnostic accuracy.
  • Only 48 of 100 studies entered the meta-analysis because deficient reporting frequently omitted true- and false-classification data.Future reporting guidelines should improve transparency and comparability of pathology AI diagnostic studies.

COMPETING INTERESTS & ACKNOWLEDGEMENTS

The authors report no competing interests and identify institutional and grant support for several authors.

  • The authors declare that they have no competing interests.
  • Several authors are funded by the National Pathology Imaging Co-operative.
  • The National Pathology Imaging Co-operative is supported by a £50m UKRI-managed investment.

Reference standard Data sources Training set details

The included studies used varied reference standards, data sources, image units, and training designs across pathology applications. Datasets ranged from small patch- or slide-based collections to larger multicentre and external-validation datasets.

  • Training set details: Reported dataset sizes varied from tens of images or slides to thousands of whole slide images and hundreds of thousands of patches.
  • Reference standard: Pathologist diagnoses, annotations, immunohistochemistry, and medical records served as reference-standard elements across the reported studies.
  • Training set details: Some studies used cross-validation or external validation, while others reported no external validation or unclear dataset separation.
  • Training set details: Training and test data were reported as whole slides, patches or tiles, regions of interest, and patient-level cases.
  • Data sources: Data sources included institutional hospitals, public datasets such as TCGA and Camelyon, cancer registries, and multicentre collections.

S1 – Search strategy of three databases (PubMed, EMBASE & CENTRAL)

The search strategy combined digital pathology, whole slide imaging, histopathology, and artificial-intelligence terms across PubMed, EMBASE, and CENTRAL. The PubMed strategy used database limits and Boolean combinations of pathology and AI concepts.

  • PubMed searches were limited to human and English-language records and combined digital pathology, whole slide imaging, histopathology, and AI-related terms.
  • The AI component included artificial intelligence, deep learning, machine learning, neural networks, computer vision, and support vector machines.
  • The strategy combined pathology/imaging terms with AI terms to identify potentially relevant records.

Appendix – Adapted QUADAS2 tool

The adapted QUADAS-2 tool assessed risk of bias and applicability across patient selection, index testing, reference standards, and flow and timing. Its signaling questions focused on representativeness, independence, interpretation, and case flow.

  • Patient Selection: Patient-selection bias was assessed using consecutive or random sampling and avoidance of inappropriate exclusions.
  • Concerns of Applicability: Applicability assessment examined whether included patients, index tests, and reference standards matched the review question and diagnostic criteria.
  • Index Test(s): Index-test bias was assessed through independent test data, external testing, consistent image analysis, and inclusion of all test cases.
  • Reference Standard: Reference-standard bias was assessed by whether diagnosis correctly classified the target condition and was interpreted without index-test knowledge.
  • Flow and Timing: Flow-and-timing bias considered whether the interval between reference diagnosis and slide scanning was less than 10 years and whether case flow could introduce bias.

S12 – Further details of study characteristics for al included studies

The included studies covered many diseases, pathology subspecialties, funding settings, dataset sizes, and intended uses. Study examples illustrate substantial variation in data scale, validation, openness, and diagnostic targets.

  • Funding: Funding sources included governmental, academic, hospital, charitable, and industry-linked organisations, while some studies declared no funder.
  • Study characteristics: Dataset sizes ranged from tens of whole slide images or cases to thousands of slides and hundreds of thousands of patches.
  • Study characteristics: Studies addressed detection and subtype classification across breast, lung, gastric, colorectal, renal, prostate, melanoma, and other diseases.
  • Intended use: Reported intended uses included disease detection, subtype classification, and grading or severity assessment.
  • Data sources: Several studies used multiple institutions or independent datasets, including TCGA, New York, and external hospital collections.
Loading 2306.07999v3…