Source-linked AI summary
Deep Learning-based Computational Pathology Predicts Origins for Cancers of Unknown Primary
Ming Y. Lu, Melissa Zhao, Maha Shady, Jana Lipkova, Tiffany Y. Chen, Drew F. K. Williamson, Faisal Mahmood
TL;DR
Determining the primary site of metastatic tumors is difficult, particularly for cancers of unknown primary. The paper presents TOAD, a whole-slide histopathology deep-learning algorithm that predicts tumor origin and metastatic status, with predictions concordant with assigned primary differentials.
Problem
Assigning primary origins to cancers of unknown primary remains a difficult tumor-assessment problem.
Method
TOAD uses whole-slide histopathology to simultaneously predict metastatic status and assign a differential tumor-origin diagnosis.
Results
Model predictions from whole-slide images were concordant with the primary differentials assigned to cases.
Takeaways & Limitations
TOAD offers an automated approach that can be tested on digitized histology images without manual slide or region-of-interest annotation.
Takeaways & Limitations
Evaluation is constrained by limited and weak ground-truth labels for challenging metastatic cases.
Abstract
from arXiv · showhide
Cancer of unknown primary (CUP) is an enigmatic group of diagnoses where the primary anatomical site of tumor origin cannot be determined. This poses a significant challenge since modern therapeutics such as chemotherapy regimen and immune checkpoint inhibitors are specific to the primary tumor. Recent work has focused on using genomics and transcriptomics for identification of tumor origins. However, genomic testing is not conducted for every patient and lacks clinical penetration in low resource settings. Herein, to overcome these challenges, we present a deep learning-based computational pathology algorithm-TOAD-that can provide a differential diagnosis for CUP using routinely acquired histology slides. We used 17,486 gigapixel whole slide images with known primaries spread over 18 common origins to train a multi-task deep model to simultaneously identify the tumor as primary or metastatic and predict its site of origin. We tested our model on an internal test set of 4,932 cases with known primaries and achieved a top-1 accuracy of 0.84, a top-3 accuracy of 0.94 while on our external test set of 662 cases from 202 different hospitals, it achieved a top-1 and top-3 accuracy of 0.79 and 0.93 respectively. We further curated a dataset of 717 CUP cases from 151 different medical centers and identified a subset of 290 cases for which a differential diagnosis was assigned. Our model predictions resulted in concordance for 50% of cases (\k{appa}=0.4 when adjusted for agreement by chance) and a top-3 agreement of 75%. Our proposed method can be used as an assistive tool to assign differential diagnosis to complicated metastatic and CUP cases and could be used in conjunction with or in lieu of immunohistochemical analysis and extensive diagnostic work-ups to reduce the occurrence of CUP.
Tumor origin assessment via deep learning
TOAD uses routinely acquired H&E whole-slide images in an interpretable multi-task framework to predict metastatic status and assign differential diagnoses for tumor origin. It achieved strong performance on known-primary and metastatic cases and showed concordance with expert differentials in CUP cases.
- Method: TOAD analyzes scanned H&E whole-slide images to simultaneously predict whether a specimen is metastatic and assign differential diagnoses for primary origin.The framework is designed to assist pathologists without immunohistochemistry, genomic testing, or extensive diagnostic screening.
- Known-primary test performance: 83.6% overall accuracy and 0.988 micro-averaged AUC ROC were achieved on previously seen test cases, with 94.4% top-3 and 97.8% top-5 accuracy.High-confidence predictions were especially accurate: 2,440/4,932 had confidence ≥0.95 and reached 98.5% accuracy.
- External validation: 78.5% accuracy, 92.6% top-3 accuracy, 95.9% top-5 accuracy, and 0.981 AUC ROC were obtained on an independent external test set.Metastatic-versus-primary classification achieved 87.3% accuracy and 0.922 AUC, indicating generalization across diverse unseen data sources.
- Metastatic and CUP assessment: 62.6% overall accuracy, 84.9% top-3 accuracy, and 92.3% top-5 accuracy were achieved across 882 metastatic cases, with 70.9% sensitivity for metastatic identification.For difficult metastatic cases, accuracy remained 56.0% with 79.0% top-3 accuracy and 91.5% top-5 accuracy.
- CUP concordance: 145/290 CUP cases achieved direct concordance with the primary differential (50.0%; κ=0.397), rising to 74.5% with top-3 predictions and 90.0% with top-5 predictions.Among 53 predictions with confidence ≥0.9, agreement was substantial (κ = 0.705).
Discussion
TOAD is a histology-based deep learning model for predicting tumor origins in cancers of unknown primary using routinely acquired whole-slide images and limited patient metadata. Its predictions showed meaningful concordance with post-workup differentials and may support diagnostic decision-making, particularly where clinical expertise, immunohistochemistry, or molecular testing is limited.
- Contribution: TOAD predicts tumor origins from whole-slide histopathology for cancers of unknown primary, a task usually requiring extensive clinical work-ups.The model uses histology and can be tested on digitized images without fine-grained manual slide or ROI annotation.
- Diagnostic performance: Using only histology and patient gender, TOAD made fairly accurate top-3 or top-5 primary differential predictions, including for challenging metastatic cases.These cases often required extensive IHC testing and clinical or radiologic correlation for diagnosis.
- CUP evaluation: After IHC work-up, model predictions were concordant to a meaningful degree with primary differentials in extremely challenging CUP cases where H&E slides were insufficient for human experts.The model used only patient gender and morphological information from whole-slide images.
- Multitask classification: 95.0% accuracy (n = 446) was achieved for distinguishing gliomas from other tumors metastasized to the brain, with similarly high accuracy at GI metastatic sites.The reported GI results were colorectal: 92.9% accuracy for n = 406 and esophagogastric: 94.1% accuracy for n = 273.
- Clinical applicability: High-confidence and top-k predictions can narrow possible origins and support TOAD as an assistive second reader, especially in low-resource settings.Attention heatmaps and high-attention patches may be used with probability scores for human interpretability and validation.
Online Methods · Dataset Description
The study assembled a large, anonymized histopathology dataset spanning 18 cancer origins, with patient-level stratified partitions and geographically diverse external and CUP cohorts. CUP cases were clinically reviewed to define differential-diagnosis subsets for model validation, while acknowledging that definitive ground truth was unavailable.
- Dataset Description: 14,518 WSIs came from consented internal Brigham and Women’s Hospital cases collected between 2010 and 2019.Slides were scanned at 20× using an Aperio scanner, and each WSI generally represented a unique patient.
- Dataset Description: 24,885 FFPE H&E digitized diagnostic slides comprised 20,413 primary and 4,472 metastatic WSIs from 23,297 patient cases.The dataset included 54.7% female and 45.3% male patients and represented approximately 21.8 terabytes of raw data.
- Dataset Description: 70% of cases formed training, 10% validation, and 20% test sets through random, class-stratified patient-level partitioning.All slides from the same patient remained in the same partition.
- Dataset Description: 662 external consult cases came from 202 medical centers across 34 US states and 19 international centers in 8 other countries.Their slides were prepared at the originating institutions using varied tissue preparation, processing, and staining protocols.
- Dataset Description: 717 consented CUP cases came from 146 US medical centers across 22 states and 5 international centers in 2 other countries.Records incorporated pathology, laboratory, patient-history, oncology, radiology, endoscopy, and autopsy reports where available.
- Dataset Description: CUP differentials served as an assessment reference because ground truth could not be obtained for these cases.The analysis evaluated the model’s ability to assign appropriate differentials using histology alone.
Multi-task Weakly-Supervised Computational Pathology
TOAD uses weakly supervised multi-task learning to predict tumor origin and primary-versus-metastatic status directly from whole-slide images. Attention pooling identifies informative regions automatically, enabling slide-level prediction without manual tumor-region annotations.
- Multi-task prediction: The model simultaneously predicts tumor origin and whether each tumor is primary or metastatic from whole-slide images.Tumor origin is an 18-class task, whereas primary-versus-metastatic status is binary.
- Weak supervision: Multiple instance learning treats each whole-slide image as a bag of image regions and trains directly with slide-level labels without manually extracted regions of interest.This approach incorporates information from the entire slide while avoiding manual region selection.
- Feature extraction: Each 256 × 256 RGB patch is encoded into a 1024-dimensional feature vector using a pretrained CNN before slide-level aggregation.Dimensionality reduction improves computational efficiency while transfer learning supplies the initial patch representations.
- Attention pooling: Task-specific attention scores weight patch representations, allowing the network to identify informative regions and form slide-level histology features.High attention scores indicate regions informative for the classification task, while low scores indicate limited diagnostic value.
- Attention pooling: The attention-based aggregation predicts the primary without requiring detailed annotations outlining precise tumor regions.The model automatically learns which subset of slide regions supports the slide-level prediction.
- Late-stage fusion: Patient sex is fused with the slide-derived deep features as an additional binary covariate before final task-specific classification.Concatenation produces a 513-dimensional vector for the final classification layer.
Additional Experiments
Additional experiments trained histology-based classifiers for adenocarcinoma and squamous cell carcinoma across selected tumor origins, as well as site-specific models for metastatic tumors in lymph nodes and liver. These models used 70/10/20 train/validation/test splits and largely matched the main network’s architecture and training setup.
- Classification of Adenocarcinoma and Squamous Cell Carcinoma: The adenocarcinoma model used 8292 WSIs spanning five tumor origins, while the SCC model used 1707 WSIs spanning four origins.The adenocarcinoma origins were Lung, Colorectal, Esophagogastric, Prostate, and Pancreatic; SCC origins were Lung, Head Neck, Cervix, and Esophagogastric.
- Experimental design: All additional experiments used 70/10/20 training, validation, and testing splits.The split was specified for both carcinoma classification and metastatic-site experiments.
- Experimental design: The additional models retained the main network’s architecture, learning schedule, and hyperparameters, but metastatic-site models disabled the primary-versus-metastatic attention branch.The attention branch was disabled because all liver and lymph cases were metastatic.
Computational Hardware and Software
Whole-slide images were processed on Intel Xeon CPUs and 16 NVIDIA P100 and 2080 Ti GPUs using a custom publicly available Python pipeline. Models used PyTorch, while analyses and visualizations used Python, R, and specified scientific packages.
- Hardware: Whole-slide images were processed on Intel Xeon multi-core CPUs and 16 NVIDIA P100 and 2080 Ti GPUs.
- Deep learning pipeline: The custom, publicly available CLAM12 whole-slide processing pipeline was implemented in Python, and models were trained on multiple GPUs with PyTorch 1.5.
- Analysis and visualization: Plots were generated in Python 3.7.5 with matplotlib 3.1.1, while NumPy 1.18.1 supported vectorized numerical computation.
- Specialized visualizations: Geographic diversity maps used pyshp 2.1.0, basemap 1.1.0, and geopy 1.22.0, whereas confusion matrices were created in R 3.6.3 with ComplexHeatmap 2.5.3.
- Statistical analysis: AUC ROC was estimated with scikit-learn 0.22.1 using the Mann-Whitney U-statistic, and 95% confidence intervals used DeLong’s method via pROC 1.16.2 in R.
WSI Processing
Whole-slide images were segmented to isolate tissue, exhaustively tiled into nonoverlapping patches, and encoded into compact 1024-dimensional feature vectors for downstream analysis.
- Segmentation: Tissue regions were automatically segmented with CLAM using thresholding, contour smoothing, artifact suppression, and area filtering.Thresholding was applied to the HSV saturation channel, followed by median blurring and morphological closing to produce the final mask.
- Patching: 280 million 256 × 256 patches were cropped without overlap at 20× magnification, with a median of 10320 patches per slide.When 20× images were unavailable, 512 × 512 patches were cropped from 40× images and downscaled to 256 × 256.
- Feature Extraction: Feature extraction used up to 16 GPUs in parallel with a batch-size of 128 per GPU.This parallelization supported efficient processing of the large patch bags generated from each whole-slide image.
Interpreting Model Prediction via Attention Heatmap
The study interprets model predictions by computing patch-level attention scores for primary-origin prediction and converting overlapping-crop scores into normalized heatmaps over the original whole-slide image. The resulting heatmaps are displayed as semitransparent overlays and made available in figures and an interactive demo.
- Interpreting Model Prediction via Attention Heatmap: Attention scores were first computed for nonoverlapping 256 × 256 WSI patches for primary-origin prediction.This established the reference attention-score distribution.
- Interpreting Model Prediction via Attention Heatmap: Up to 90% overlap was then used to generate more fine-grained heatmaps.Scores from overlapping crops were converted into normalized percentile scores from 0.0 for low attention to 1.0 for high attention using the reference distribution.
- Interpreting Model Prediction via Attention Heatmap: Normalized scores were registered to each patch’s spatial location, while overlapping-region scores were accumulated and averaged.This produced a spatially aligned attention representation on the original WSI.
- Interpreting Model Prediction via Attention Heatmap: A colormap was applied, and the attention heatmap was displayed as an overlay layer with transparency value 0.5.The maps were shown in Figure 3 and Extended Data Figures 6, 8, and 9, and could also be visualized in an interactive demo.
Data Availability
TCGA digitized whole slide images and corresponding diagnoses are publicly accessible through the NIH genomic data commons. Requests for additional data and materials undergo review and may be restricted by intellectual property, confidentiality, or institutional requirements.
- TCGA digitized, high-resolution diagnostic whole slide images and corresponding diagnoses are publicly accessible through the NIH genomic data commons.
- Requests for in-house raw and analyzed data and materials are reviewed by the authors for intellectual property or confidentiality obligations.
- Shareable data requests are processed through formal channels under institutional and departmental guidelines, with patient-related data potentially subject to confidentiality.
Code Availability
The TOAD experiments were implemented in Python with PyTorch, and the code and reproduction scripts are publicly available under the GNU GPLv3 license.
- The code was implemented in Python using PyTorch as the primary deep learning package.
- Code and scripts to reproduce the paper’s experiments are available at the TOAD GitHub repository.Repository: https://github.com/mahmoodlab/TOAD
- All source code is provided under the GNU GPLv3 free software license.