Source-linked AI summary

Automated Gleason Grading of Prostate Biopsies using Deep Learning

Wouter Bulten, Hans Pinckaers, Hester van Boven, Robert Vink, Thomas de Bel, Bram van Ginneken, Jeroen van der Laak, Christina Hulsbergen-van de Kaa, Geert Litjens

arXiv:1907.07980v1eess.IVcs.CV

TL;DR

Inter-observer variability limits the reliability of Gleason grading, despite its importance for prostate cancer prognosis. The paper develops a fully automated deep learning system using semi-automatically labeled biopsy data and evaluates it against expert and external references. The system achieved high internal agreement, outperformed 10 of 15 pathologists, and showed external agreement within inter-observer variability.

  • Problem

    Gleason grading is the most important prognostic marker for prostate cancer patients but suffers from substantial inter- and intra-observer variability.

  • Method

    A fully automated deep learning system grades entire prostate biopsies using semi-automatic training labels, gland-level pattern segmentation, and interpretable grade-group determination.

  • Results

    0.918 quadratic kappa was achieved against expert consensus, while the system outperformed 10 of 15 pathologists in the observer experiment.

  • Takeaways & Limitations

    The system has potential as a first or second reader for prostate biopsy grading and can provide gland-level growth-pattern overlays for pathologists.

  • Takeaways & Limitations

    The development data came from a single center, and performance was lower on the external test set; multi-center data could improve robustness.

Abstract

from arXiv · show

The Gleason score is the most important prognostic marker for prostate cancer patients but suffers from significant inter-observer variability. We developed a fully automated deep learning system to grade prostate biopsies. The system was developed using 5834 biopsies from 1243 patients. A semi-automatic labeling technique was used to circumvent the need for full manual annotation by pathologists. The developed system achieved a high agreement with the reference standard. In a separate observer experiment, the deep learning system outperformed 10 out of 15 pathologists. The system has the potential to improve prostate cancer prognostics by acting as a first or second reader.

Introduction

Gleason grading is central to prostate cancer prognosis and treatment planning but has substantial observer variability. This study develops a fully automated deep learning system for grading the full range of Gleason grades in whole prostate biopsies without manual pixel-level annotations.

  • Clinical motivation: The Gleason score is the strongest prognostic marker for prostate cancer but varies substantially between and within pathologists.Specialized uropathologists achieve higher concordance, but their expertise is not widely available.
  • Study objective: The study investigates computational pathology and deep learning for automated Gleason grading of prostate biopsies.Deep learning is presented as potentially reproducible and capable of expert-level pathological diagnosis in other tasks.
  • Clinical motivation: Treatment planning largely relies on the biopsy Gleason score, which sums the most common and highest secondary growth patterns.The grading system classifies architectural patterns from 1 to 5; patterns 1 and 2 are now rarely reported in biopsies.
  • Clinical motivation: The introduction of five prognostically distinct grade groups did not reduce inter- and intra-observer variability.Groups range from score 3+3 and lower in group 1 to higher scores in group 5.
  • Study contribution: The system targets entire prostate biopsies, covers the full Gleason-grade range, and uses semi-automated training labels instead of manual pixel-level annotations.The study also evaluates the system against an expert consensus reference standard and on an external tissue microarray test set.

Results

The system was trained and evaluated on biopsy data using semi-automatic labeling, gland-level segmentation, and grade-group prediction. It showed high agreement with expert consensus, strong clinical-category discrimination, and pathologist-level performance, while external performance was within inter-observer variability.

  • Data and evaluation: 5834 biopsies from 1243 patients were divided into training, tuning, and independent test sets.The split comprised 4712 training biopsies, 497 tuning biopsies, and 550 test biopsies.
  • Data and evaluation: Semi-automatic labeling combined segmentation algorithms with slide-level Gleason grades because exhaustive gland- and cell-level annotation was impractical.The training and tuning sets contained 5209 biopsies.
  • System: The system predicted biopsy grade groups through gland-level Gleason-pattern segmentation followed by calculation of normalized epithelial tissue percentages.The reported proportions included benign tissue and patterns G3, G4, and G5.
  • Internal test performance: 0.918 quadratic kappa was achieved on 535 test biopsies against the expert consensus grade group.Most discrepancies occurred between grade groups 2 and 3, and between grade groups 4 and 5.
  • Internal test performance: 0.990 AUC was achieved for benign-versus-malignant classification, with 99% tumor sensitivity at 82% specificity.Using grade group 2 as the cutoff, the test-set AUC was 0.978.
  • Observer comparison: 0.854 kappa on the observer set exceeded the panel median, and the system outperformed 10 of 15 panel members.Its performance was better than pathologists with less than 15 years of experience and not significantly different from those with more than 15 years.
  • Observer comparison: The system’s observer-set performance ranked in the top three when compared with all panel members independently of the consensus reference.The panel’s median inter-rater agreement with the consensus was 0.819 quadratic kappa.
  • External evaluation: On the external tissue-microarray test set, Gleason-score quadratic kappa was 0.711 for one pathologist and 0.639 for the other.These values were lower than the comparator algorithm but within inter-observer variability.

Discussion

The fully automated system achieved pathologist-level Gleason grading, with strong agreement against expert consensus and performance comparable to pathologists across clinically relevant classifications. Its interpretable gland-level outputs may support prescreening and second reading, although broader validation and tumor-type coverage remain necessary.

  • Performance: Quadratic kappa was 0.918 against the consensus reference standard, while external-test kappa values of 0.711 and 0.639 were within expected inter-observer variability.The external results were comparable to the reference standard’s inter-observer agreement of 0.71.
  • Clinical utility: Gland-level growth-pattern overlays and volume-based grading make the system interpretable and potentially useful for prescreening or second reading.The approach mirrors pathologists’ use of growth-pattern volumes rather than adding a separate learned grading model.
  • Limitations: The study focused solely on acinar adenocarcinoma and did not explicitly detect intraductal carcinoma or other prognostically relevant information.Other tumor types and foreign tissue may also occur in prostate biopsies.

Methods

The study assembled a large biopsy dataset, created consensus and observer reference standards, and trained a U-Net system using semi-automatically generated labels. Independent evaluation included expert consensus, an international observer panel, and external testing.

  • Reference standard: An independent test set contained 550 biopsies from 210 patients, with three expert uropathologists establishing the reference standard.Cases without initial consensus underwent regrading and, when necessary, a consensus meeting.
  • Observer experiment: The observer experiment used 100 biopsies graded independently by 13 pathologists and two pathologists in training from 14 laboratories and 10 countries.Observers followed the ISUP 2014 guidelines and had varying experience levels.
  • Label quality: Automated training labels had quadratic Cohen’s kappa 0.853 against the reference on grade group, indicating label noise in the training data.The retrieved labels had grade-group accuracy 0.720 and Gleason-score accuracy 0.675.
  • Labeling and training: Training labels were generated semi-automatically from pathology reports and refined using an initial model trained on pure Gleason-score biopsies.Hard-negative patches were added to improve benign, inflammatory, and tumorous tissue discrimination.

Additional information

The paper reports competing interests and funding relationships involving several authors and external organizations.

  • Competing interests: Most listed authors declared no conflict of interest.The statement identifies W.B., H.P., T.d.B., C.H.-v.d.K., R.V., H.v.B., and B.v.G.
  • Competing interests: J.v.d.L. reported advisory-board memberships and research funding involving Philips, ContextVision, and Sectra.
  • Competing interests: G.L. reported research funding from Philips Digital Pathology Solutions and consultancy and research relationships involving Novartis.
Loading 1907.07980v1…