Source-linked AI summary

Development and Validation of a Deep Learning Algorithm for Improving Gleason Scoring of Prostate Cancer

Kunal Nagpal, Davis Foote, Yun Liu, Po-Hsuan, Chen, Ellery Wulczyn, Fraser Tan, Niels Olson, Jenny L. Smith, Arash Mohtashamian, James H. Wren, Greg S. Corrado, Robert MacDonald, Lily H. Peng, Mahul B. Amin, Andrew J. Evans, Ankur R. Sangoi, Craig H. Mermel, Jason D. Hipp, Martin C. Stumpe

arXiv:1811.06497v1cs.CVcs.LG

TL;DR

Gleason scoring is subjective and poorly reproducible despite its importance for prostate cancer prognosis and management. This study develops a deep learning system for whole-slide Gleason scoring and reports higher accuracy than generalist pathologists.

  • Problem

    Gleason scoring is subjective and variable despite its important role in prostate cancer prognostication and patient management.

  • Method

    The study develops a deep learning system for region-level Gleason classification on whole-slide prostatectomy images.

  • Results

    The DLS was more accurate than a cohort of 29 generalist pathologists in Gleason scoring whole-slide images from prostatectomy patients.

  • Takeaways & Limitations

    The reported findings support further consideration of DLS-based whole-slide Gleason scoring.

  • Takeaways & Limitations

    The region-level system does not account for context beyond its input image.

Abstract

from arXiv · show

For prostate cancer patients, the Gleason score is one of the most important prognostic factors, potentially determining treatment independent of the stage. However, Gleason scoring is based on subjective microscopic examination of tumor morphology and suffers from poor reproducibility. Here we present a deep learning system (DLS) for Gleason scoring whole-slide images of prostatectomies. Our system was developed using 112 million pathologist-annotated image patches from 1,226 slides, and evaluated on an independent validation dataset of 331 slides, where the reference standard was established by genitourinary specialist pathologists. On the validation dataset, the mean accuracy among 29 general pathologists was 0.61. The DLS achieved a significantly higher diagnostic accuracy of 0.70 (p=0.002) and trended towards better patient risk stratification in correlations to clinical follow-up data. Our approach could improve the accuracy of Gleason scoring and subsequent therapy decisions, particularly where specialist expertise is unavailable. The DLS also goes beyond the current Gleason system to more finely characterize and quantitate tumor morphology, providing opportunities for refinement of the Gleason system itself.

Introduction

Gleason scoring is central to prostate cancer prognosis and management but remains subjective, with substantial interobserver and intraobserver variability. This study developed a deep learning system (DLS) for Gleason scoring and quantitation of prostatectomy whole-slide sections, comparing it with pathologists and specialist-defined references while exploring finer-grained grading.

  • Clinical motivation: Gleason scoring guides prognosis and standardized patient management but remains subjective and has 30-53% reported discordance.The Gleason score and tumor stage are powerful prognostic predictors, while Gleason scoring suffers from interobserver and intraobserver variability.
  • Prior work and limitation: Prior artificial intelligence studies addressed prostate cancer detection in needle biopsies and Gleason grading in research tissue microarrays, not clinically diagnosed specimens.Tissue microarrays comprise selected tumor sub-regions outside routine clinical workflow.
  • Study rationale: A whole-slide Gleason scoring tool could reduce grading variability, improve prognostication, and optimize patient management.The rationale was that accurate scoring on whole-slide sections used in clinical workflows could address variability.
  • Study contribution: The study developed a DLS to perform Gleason scoring and quantitation on prostatectomy specimens.DLS accuracy was compared with pathologists using a reference standard defined by genitourinary specialist pathologists.
  • Study objectives: The study also examined DLS, pathologist, and specialist-reference risk stratification for predicting disease progression and explored finer-grained tumor grading.The finer-grained measures were intended to support more precise prognostication.

Results

On an independent 331-slide validation set, the deep learning system (DLS) outperformed general pathologists in whole-slide Gleason scoring and quantitated Gleason patterns more accurately. Its region-level predictions aligned strongly with pathologist consensus, while finer-grained pattern representations showed improved prognostic discrimination.

  • Overview of the Deep Learning System (DLS) and Data Acquisition: 112 million image patches from 912 slides formed the DLS training dataset, described as roughly 4× larger in annotated tissue area than Camelyon16.The dataset covered approximately 115,000 mm^2 of tissue and required approximately 900 pathologist hours to annotate.
  • Comparison of DLS to pathologists on Whole-Slide Gleason Scoring: 0.70 was the DLS accuracy versus 0.61 for 29 pathologists on whole-slide grade-group classification (p=0.002).The DLS accuracy had a 95% CI of 0.65-0.75, compared with 0.56-0.66 for pathologists.
  • Comparison of DLS to pathologists on Gleason Pattern Quantitation: 4-6% lower mean absolute error was achieved by the DLS than the average pathologist when quantitating Gleason patterns 3 and 4.For grade groups 2 and 3, the DLS achieved 8% lower mean absolute error.

Discussion

The DLS was more accurate than 29 pathologists for Gleason scoring and may reduce variability through finer-grained pattern classification and direct quantitation. Clinical implementation remains contingent on further validation across settings, specimen types, and larger clinically annotated datasets.

  • Clinical performance: 66% Gleason score concordance and 61% Gleason grade group concordance were achieved by pathologists against genitourinary specialist pathologists, while the DLS was more accurate than the cohort.These concordances were at the high end of reported inter-pathologist Gleason score concordances of 47%-70%.
  • Region-level pattern classification: The DLS reflected ambiguity in discordant regions through its prediction scores and demonstrated potential for finer-grained Gleason pattern assignment.Finer-grained categorization could mitigate variability from applying coarse categories to a continuous histologic spectrum and support more precise risk stratification.
  • Pattern quantitation: The DLS directly quantified Gleason patterns from region categorizations, bypassing variability from visual quantitation and agreeing more accurately with the specialist-adjudicated reference standard than pathologists.This suggests an opportunity for more precise prognostication.
  • Model development: >30,000 DLS stage-1 inferences per second enabled the quasi-online hard-negative mining approach used in this study.The authors anticipate that continuous hard-negative mining may benefit development of other histopathology deep learning algorithms.
  • Limitations and future work: Clinical implementation requires addressing the study’s digital-review setting and validating the DLS for biopsies, other histologic variants, prognostic categorizations, and larger clinically annotated datasets.The study’s sensitivity analysis excluding cases for which pathologists preferred additional resources or consults showed qualitatively similar results.

Conclusions

The DLS outperformed 29 generalist pathologists in Gleason scoring of prostatectomy whole-slide images. It also enabled more accurate Gleason-pattern quantitation, finer differentiation, and potentially improved risk stratification and treatment decisions.

  • The DLS outperformed a cohort of 29 generalist pathologists in Gleason scoring of prostatectomy whole-slide images.
  • The system provided more accurate quantitation of Gleason patterns and finer-grained discretization across the well-to-poor differentiation spectrum.
  • The DLS created opportunities for better risk stratification and could enhance the clinical utility of the Gleason system.
  • These capabilities could support better treatment decisions for patients with prostatic adenocarcinoma.

Competing interests

Several authors are Google employees who own Alphabet stock. The article also disclaims official government endorsement and identifies military-service-member authors whose work was prepared as part of official duties.

  • Several named authors are employees of Google Inc. and own Alphabet stock.
  • The authors’ views do not necessarily reflect the official policy or position of the Department of the Navy, Department of Defense, or U.S. Government.
  • A.M., N.O., J.L.S., and J.H.W. are military Service members, and this work was prepared as part of their official duties.
  • U.S. Government work prepared by military Service members as part of official duties is not eligible for copyright protection under Title 17, U.S.C.

Supplementary Information

Supplementary analyses describe dataset construction, variant handling, validation exclusions, comparative accuracy metrics, Gleason-pattern quantitation, and clinical-event modeling. These analyses further characterize the DLS relative to pathologists and the reference standard.

  • Datasets: Training and tuning used 1–7 slides per patient, each reviewed by 3–5 pathologists, with ungradable slides excluded.Slide-level Gleason scores and region-level Gleason-pattern annotations were collected for overlapping subsets.
  • Variant handling: Slides with non-Gleason-gradable cancer or histological variants were excluded, except intraductal carcinoma; other variants followed ISUP 2014 recommendations.
  • Validation exclusions: 20 slides were excluded because the specialist adjudicator lacked confidence; among 12 cases reaching consultant consensus, the DLS was concordant on 9.The supplementary analysis compared DLS classifications with the adjudicated consultant consensus.
  • Accuracy metrics: Supplementary tables compare DLS and pathologists on unadjusted accuracy, adjusted grade-group accuracy, Cohen’s kappa, and Gleason-score accuracy.Comparisons include the cohort of 29 pathologists and 10 individual pathologists; adjusted accuracy uses a population-level grade-group distribution of 7397:8353:3106:1968.
  • Supplementary analyses: The DLS was more accurate than 14 of 19 pathologists on overlapping validation subsets, while additional analyses examined Gleason-pattern quantitation and Cox models for progression or biochemical recurrence.Cox models used quantified or fine-grained Gleason patterns as input features on the validation set of 331 slides.

Supplementary Fig. 1: Confusion Matrices for the DLS and two pathologist subgroups

Supplementary Fig. 1 compares DLS confusion matrices with two pathologist subcohorts on the validation set. The DLS was more accurate for GG1, GG2, and GG4-5, but less accurate for GG3.

  • Confusion-matrix comparisons: The confusion matrices compare the DLS with 10 pathologists who individually annotated every validation slide and 19 pathologists who collectively provided three reviews per slide.Both comparisons use the validation set.
  • Confusion-matrix findings: The DLS shows greater accuracy in classifying slides as GG1 and GG2 than the two pathologist subcohorts.
  • Confusion-matrix findings: The DLS shows greater accuracy for GG4-5 but lower accuracy for GG3 on the validation set.

Grading

Gleason grading was derived from predominant and next-most-common glandular patterns, while annotations and labels were standardized through pathologist review, majority voting, and specialist adjudication. The DLS then converted patch-level classifications into slide-level grade-group predictions using calibrated features and k-nearest-neighbor models.

  • Consensus grading: Training-slide labels used the most common annotation from 3–7 pathologists, with ties resolved toward the more severe grade.Tuning slides were initially reviewed by 3–5 pathologists and adjudicated by one of three genitourinary specialists.
  • Slide-level grading: Slide-level Gleason scores, such as 3+4, were derived from the predominant and next-most-common glandular patterns because directly provided scores showed inconsistent tertiary replacement.Grade groups were then determined directly from the derived Gleason score using published definitions.
  • Region annotation: Regions with ambiguous or difficult boundaries received mixed-grade labels, while artifacts were labeled separately and uncertain regions received a “consult” label.Perineural and lymphovascular invasive tumor and intraductal carcinoma were labeled non-Gleason-gradable.
  • DLS grading model: The second DLS stage summarized patches as %Tumor, %GP3, %GP4, and %GP5, then trained separate k-nearest-neighbor models for grade-group tasks.The final selected hyperparameters were k=24 with uniform neighbor weighting.

Fine-grained Gleason Pattern (GP)

The DLS generates a quantitative Gleason pattern that smoothly interpolates among existing patterns 3, 4, and 5. It derives this value from calibrated likelihoods and visualizes or retrieves regions according to the resulting quantitative pattern.

  • Fine-grained Gleason Pattern (GP): The quantitative Gleason pattern smoothly interpolates between existing Gleason patterns 3, 4, and 5 using calibrated DLS predictions.The method processes calibrated DLS-predicted likelihoods for each pattern.
  • Fine-grained Gleason Pattern (GP): 0.78 is the interpolated weight for Gleason pattern 3 when predictions are [0.7, 0.2, 0.1], yielding a quantitative Gleason pattern of 3.78.The weight is calculated as 0.7 / (0.7 + 0.2) = 0.78, followed by 3 + 0.78 = 3.78.
  • Fine-grained Gleason Pattern (GP): CIELAB color space was used to visualize quantitative Gleason patterns because it is designed to be perceptually uniform for numerical values.The visualization examples include Fig. 4a.
  • Fine-grained Gleason Pattern (GP): Validation-dataset image patches were selected when their computed quantitative Gleason patterns most closely matched a desired value, such as 3.5.This procedure was used for Fig. 4c and Supplementary Fig. 3.

Statistical Analysis

The study used a modified permutation test to compare DLS accuracy with the cohort-of-29 pathologists and balanced pathologist representation in risk-stratification analyses. Confidence intervals were estimated by bootstrapping slides and annotators with replacement.

  • Statistical Analysis: The cohort-of-29 comprised 10 pathologists annotating all 331 slides each and 19 pathologists collectively annotating all slides 3 times.The 19 pathologists contributed about 50±10 annotated slides each, while the 10 full-slide annotators were selected based on reviewing speed and availability.
  • Statistical Analysis: Risk-stratification analyses sampled annotations to approximate equal representation of each pathologist, using subgroup-specific probabilities for the 10- and 19-pathologist groups.Each subgroup-of-10 annotation had 1/29 probability, while each subgroup-of-19 annotation had (19/29)*(1/3) probability.

Supplementary Results · Supplementary References

The DLS’s first-stage errors arose from imprecise spatial localization, ambiguous histology, and true prediction mistakes. Its second stage was fairly robust to these errors by summarizing slide-wide regional predictions into a small feature set.

  • DLS Region-level Errors: The first stage sometimes localized predicted Gleason pattern regions imprecisely.Errors included inaccurate spatial extent and misclassification of intervening non-tumor tissue between nearby tumor regions.
  • DLS Region-level Errors: The DLS had difficulty delineating the precise stroma–tumor interface, particularly for GP5 and stroma.GP5 may appear as individual tumor cells within connective tissue, making cell-level outlining impractical.
  • DLS Region-level Errors: Imperfect underlying region-level annotations limited the DLS’s precision.The passage attributes this limitation to annotation impurity.
  • DLS Region-level Errors: Ambiguous histology also caused errors, such as tangentially cut GP3 resembling the fused-gland pattern defining GP4.The DLS interpreted the image patch surrounding the region rather than context beyond its input image.
  • DLS Region-level Errors: True prediction mistakes were expected to improve naturally with more data.These errors were distinct from those caused by spatial localization or ambiguous histology.
  • DLS Region-level Errors: The second stage was fairly robust to first-stage errors by summarizing predictions from all slide regions as a small number of features.This slide-level aggregation reduced sensitivity to individual region-level errors.
Loading 1811.06497v1…