Source-linked AI summary
Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation
Yumi Lee, Harim Oh, Hyoryung Kim, Minji Kim, Eunsu Kim, Hyeseong Lee, Junya Fukuoka, Andrey Bychkov, Jijgee Munkhdelger, Rajiv Kumar Kaushal, Ayushi Sahay, Rajni Yadav, Bharathi Prabakaran, Sulen Sarioglu, Serdar Balcı, Ilknur Turkmen, Yuri Tolkach, Christian Harder, Julian Westerdorf, Reinhard Buettner, Audun Ljone Henriksen, Sepp De Raedt, Byung Hyun Lee, Sungjin Lim, Joohoon Lee, Gwanghyun Kim, Se Young Chun, Suryakant Singh, Saarthak Kapse, Prateek Prasanna, Kyung A Kim, Yousun Kang, Sehwan Yoo, Sungman Hong, Shubham Innani, Michael Feldman, Spyridon Bakas, Ujjwal Baid, Prasad Dutande, Suhas Gajare, Bhakti Baheti, Serkan Sökmen, Ece Tuğba Cebeci, Ahmet Halıcı, Musa Balcı, Kardelen Peçenek, Srividhya Sainath, Kyongseok Jang, Messi H. J. Lee, Noorul Wahab, Bodong Du, Jiaming Zhang, Qixiang Zhang, Jang-Hwan Choi, Sangjeong Ahn
TL;DR
WSI pathology report generation lacks large paired datasets and must bridge spatially distributed visual patterns with structured clinical text. The study introduces a clinically curated Pan-Asia dataset and REG 2025 benchmark, analyzing diverse multimodal models. Results favor structured report representations and multimodal grounding, while revealing quantitative instability and diagnostic overspecification.
Problem
WSI report generation remains limited by scarce large-scale paired datasets and the difficulty of mapping spatially distributed slide patterns to structured clinical reports.
Method
The study constructs a clinically curated Pan-Asia WSI–report dataset and analyzes diverse submitted models through the REG 2025 MICCAI challenge benchmark.
Results
Top-performing approaches benefited from structured pathology-report representations and multimodal grounding, while VLM-based frameworks generalized across organs and an independent European cohort.
Takeaways & Limitations
REG 2025 provides a benchmark for systematic evaluation of pathology report generation and multimodal diagnostic reasoning.
Takeaways & Limitations
Models showed instability in quantitative attribute estimation and frequently overspecified diagnoses instead of preserving meaningful uncertainty.
Abstract
from arXiv · showhide
The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the scarcity of large-scale WSI--report datasets and the complexity of mapping spatially distributed visual patterns to structured clinical text. To address this, we introduce a clinically curated Pan-Asia WSI--report dataset of approximately 10,500 pairs from five institutions and establish the REG 2025 benchmark through a MICCAI challenge for systematic evaluation of multimodal models. We analyze submitted methods spanning pretrained VLMs, multiple-instance learning frameworks, hierarchical expert models, retrieval-augmented generation, and cross-modal Transformers. Rather than indicating that VLM use alone was sufficient for superior performance, the results suggest that top-performing methods benefited from structured report representations, hierarchical diagnostic decomposition, and effective multimodal grounding. We identify key limitations, including instability in quantitative attribute estimation (e.g., numeric hallucination) and a tendency toward diagnostic overspecification, with some errors resembling known diagnostic pitfalls in routine pathology. These findings establish REG 2025 as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology models.
1. Introduction
WSI-based pathology report generation is constrained by gigapixel images, spatially distributed patterns, weak image–text alignment, and limited paired datasets. The paper introduces a clinically curated Pan-Asia dataset and REG 2025 benchmark, then analyzes model characteristics and limitations.
- Motivation: WSI report generation must translate gigapixel images with weakly localized patterns into abstract clinical reasoning and diagnostic conclusions.This creates a substantial spatial–semantic gap between whole-slide images and reports.
- Motivation: Existing WSI image–text benchmarks remain extremely limited compared with benchmarks for classification, segmentation, and survival prediction.Available examples include TCGA, HISTAI, and HANCOCK, with TCGA reports often requiring extensive preprocessing.
- Contributions: The study introduces a clinically curated Pan-Asia WSI–report dataset and establishes REG 2025 for systematic multimodal report-generation evaluation.The dataset follows CAP-guided reporting principles, while the benchmark was conducted through a MICCAI challenge.
- Contributions: Submitted approaches include pretrained VLMs, pathology-specific MIL frameworks, and cross-modal Transformer architectures.The challenge attracted 389 participants from 40 countries, with 24 teams submitting final models.
- Findings: The analysis links stronger performance with multimodal fusion, structured report learning, clinical precision, linguistic coherence, and complete pathological reasoning.Pan-Asia-trained models also showed strong generalization to independent European cohorts.
- Findings: REG 2025 jointly evaluates pathology report generation and vision–language understanding while exposing structured reasoning requirements and quantitative instability.The benchmark provides practical guidance for developing future multimodal pathology models.
2. Materials and Methods
The study constructs and evaluates a standardized multi-institutional WSI–report dataset through a two-phase challenge. Its clinically grounded evaluation combines lexical, keyword, and semantic measures to assess report quality and diagnostic adequacy.
- Dataset construction: Reports were standardized with a unified template, reviewed by pathology trainees and board-certified pathologists, and aligned with WHO and CAP terminology.The curated reports include organ, procedure, histologic type, and applicable grade or organ-specific cancer attributes.
- Dataset construction: Two-stage screening removed technically inadequate slides and diagnostically equivocal cases to reduce label noise from borderline lesions.Exclusion criteria included insufficient tissue, poor focus or staining, cautery artifact, and uncertain diagnoses after expert review.
- Dataset construction: The dataset combines WSIs and reports from five institutions using three scanner platforms, covering seven organs and diverse diagnostic categories.The final dataset contains 10,494 WSI–report pairs spanning malignant, premalignant, benign, and non-neoplastic conditions.
- Challenge design: The challenge provided approximately 8,500 training pairs and two test phases, each requiring reports for 1,000 WSI-only samples under patient-level partitioning.Phase 2 combined 500 Pan-Asia and 500 European samples to assess cross-domain robustness.
- Evaluation metrics: Evaluation combined ROUGE-L, BLEU-4, clinically relevant keyword similarity, and biomedical embedding similarity into a composite ranking score.Keyword and semantic components address clinically important equivalence and negation that lexical overlap can miss.
3. Results
The REG 2025 evaluation compared participating teams across Pan-Asia and German cohorts, finding strong overall performance and evidence of cross-regional generalization. Qualitative analysis showed high diagnostic concordance but persistent instability in quantitative estimation and diagnostic specificity.
- 389 participants from 40 countries joined the challenge, with 20 teams submitting in Phase 1 and 22 in Phase 2.
- Phase 1 was led by ICGI at 0.8098, while Phase 2 was led by IMAGINE Lab at 0.8494; the top three final teams all exceeded 0.8000.
- German-subset scores had higher means and lower variance than Pan-Asia scores, while relative team rankings were largely preserved across subsets.The authors attribute this contrast to dataset composition and variability rather than inherent population-level difficulty.
- Diagnostic concordance: Models showed high diagnostic concordance and preserved clinically relevant attributes in representative cases, including uncommon diagnoses such as malignant melanoma.
- Diagnostic discordance: Quantitative attribute estimation was unstable: in a prostate case, models correctly identified diagnosis and grading but reported tumor volume as 80% instead of the 5% ground truth.
- Diagnostic discordance: Models frequently over-specified uncertain diagnoses, including fibroepithelial lesions as fibroadenoma or phyllodes tumor and ADH as DCIS.
- Comparative analysis: Top-performing methods used structured clinical representation, multimodal grounding, or intermediate diagnostic reasoning to preserve hierarchical and clinically grounded report information.
4. Overview of Submitted Algorithms
Submitted systems combined standardized WSI preprocessing with diverse report-processing and model-architecture choices. The strongest performance was associated with category-aware report structuring, while hierarchical decomposition improved organization but introduced routing-related failure modes.
- Preprocessing strategies: Most teams used standardized MIL-based WSI feature-extraction pipelines, with no significant performance differences attributable to preprocessing alone.
- Report processing: Top-performing teams partitioned pathology reports into distinct categories rather than treating them as plain text sequences.ICGI aligned tagged report sections with tile features, ICL_PathReport used a hierarchical annotation tree, and IMAGINE Lab generated category-specific concept prompts.
- Report processing: The three leading category-structured teams reached average Phase 2 scores of ROUGE 0.8572 and BLEU 0.6853.
- Report processing: The structured approaches collectively achieved a markedly superior average KEY score of 0.7990 compared with other participants.
- Report processing: Category-aware methods improved by +0.09 BLEU and +0.13 KEY from Phase 1 to Phase 2, compared with approximately +0.01 and +0.02 for tokenization-based methods.
- Model architectures: ICGI used an n-gram-constrained Transformer decoder to suppress hallucinations and demonstrated correct predictions in several rare cases.
- Model architectures: ICL_PathReport’s Tree-of-Experts modeled organ-to-diagnosis-to-subtype hierarchy, but an incorrect organ-level routing decision could constrain subsequent diagnosis selection.
- Model architectures: Hierarchical and specialized approaches achieved higher accuracy on rare cases but showed reduced robustness on majority cases.
5. Discussion
The challenge shows that pathology report generation requires clinically grounded, structured multimodal reasoning rather than superficial image-to-text translation. It also exposes weaknesses in current evaluation metrics and VLM frameworks, while pointing to reasoning-aware and knowledge-integrated designs.
- Evaluation limitations: KEY and EMB aligned more consistently with final rankings than ROUGE and BLEU, which sometimes diverged from clinical correctness.Some models scored highly on ROUGE while missing critical diagnostic terms, whereas others preserved clinically important attributes despite lower BLEU.
- Evaluation limitations: Lexical-overlap metrics inadequately capture pathology-report accuracy because reports require structured diagnostic reasoning and precise concept representation.Important attributes include Gleason score, tumor proportion, and histologic subtype.
- Model limitations: Current VLM report generators show recurring numeric hallucination and diagnostic misclassification, especially for tumor proportion, grading score, and subtype.Models may produce semantically plausible text without sufficient grounding in visual evidence.
- Reasoning-aware generation: VQA-based structured generation reduced hallucination and improved slot-level concept grounding relative to conventional end-to-end captioning approaches.The discussion identifies slot-based prediction, structured generation, and diagnostic decomposition as reasoning-augmented directions.
- Reasoning-aware generation: Retrieval-augmented generation and hierarchical expert architectures contributed to lower hallucination and stronger semantic grounding.These approaches leverage external evidence or structured expert knowledge instead of relying only on internal model representations.
- Overall implications: The challenge frames report generation as a multimodal reasoning problem involving structured diagnosis, quantitative estimation, and uncertainty-aware decisions.The discussion recommends clinically grounded metrics, reasoning-integrated generation, and structured pathology knowledge.
6. Conclusion
The REG 2025 Challenge constructs a large Pan-Asia WSI–report dataset and benchmarks diverse pathology report-generation models. Its analyses show clinically coherent reporting, European-cohort generalization, and advantages for structured representations with multimodal grounding.
- Dataset and benchmark: Approximately 10,500 WSI–report pairs form the first large-scale Pan-Asia multi-institutional pathology dataset released through REG 2025.The dataset was constructed and publicly released through the challenge.
- Dataset and benchmark: REG 2025 systematically compares submitted pathology report-generation models using the constructed dataset.The challenge supports comparative analysis across participating teams.
- Benchmark findings: VLM-based frameworks generated clinically coherent pathology reports across diverse organs and pathological categories.The conclusion reports this pattern across the participant approaches analyzed.
- Benchmark findings: Models maintained stable performance on an independent European cohort, indicating cross-domain generalization capability.The cohort was independent of the Pan-Asia training data.
- Benchmark findings: Structured pathology-report representations and multimodal grounding achieved superior performance, underscoring the importance of structured multimodal modeling.The conclusion identifies these design choices among the approaches associated with stronger performance.
- Benchmark significance: REG 2025 provides a benchmark for systematic evaluation of multimodal diagnostic understanding in pathology report generation.The challenge explicitly frames the task as more than image captioning.
CRediT authorship contribution statement
The contribution statement lists author-specific roles spanning data curation, resources, analysis, investigation, project administration, and writing.
- Author contributions: Yumi Lee contributed data curation, formal analysis, investigation, project administration, and manuscript drafting and revision.
- Author contributions: Harim Oh contributed resources, data curation, investigation, and manuscript drafting.
- Author contributions: Hyoryung Kim and Minji Kim contributed data curation and manuscript review.
- Author contributions: Eunsu Kim and Hyeseong Lee contributed data curation.
- Author contributions: Junya Fukuoka, Andrey Bychkov, Jijgee Munkhdelger, Rajiv Kumar Kaushal, and Ayushi Sahay contributed resources.
A.1. ICGI (Fig. S2-(a))
The ICGI pipeline extracts tile-level foundation-model features, aggregates them into slide-level representations, and uses task-specific heads for classification and text generation.
- Feature extraction: The ICGI team uses H-optimus-1 to extract tile-level embeddings after color correction and tissue detection.
- Prediction heads: The slide embedding feeds two linear classifiers for anatomical site and procedure prediction.
- Prediction heads: A Transformer decoder generates slide descriptions from the aggregated embedding.
- Slide representation: An eight-head ABMIL backbone concatenates weighted tile sums into a slide-level representation.
- Inference: An inference-time n-gram mask filters token sequences not observed in the training corpus.
A.2. ICL_PathReport (Fig. S2-(b))
ICL_PathReport proposes a Tree-of-Experts framework that hierarchically predicts structured diagnostic report components. Its training uses feature augmentation to address limited data for rare or complex diagnoses.
- Contributions: The Tree-of-Experts framework hierarchically predicts organs, procedures, histological types, and other detailed diagnostic attributes through specialized MLP experts.The system aggregates patch features into slide-level embeddings before applying the expert hierarchy.
- Method: Noisy Feature Mixup combines noise perturbation and weighted feature mixing to improve generalization and calibration for rare or complex diagnoses.The augmentation is applied to slide features during training.
A.11. REG_Path (Fig. S4-(k))
REG_Path uses adaptive preprocessing and a two-stage training strategy to generate clinically grounded pathology reports. The pipeline first establishes visual-textual alignment, then performs instruction tuning for full report generation.
- Method: REG_Path dynamically adjusts segmentation thresholds according to each slide’s statistical distribution for more accurate feature extraction.The adaptive preprocessing is built into the SlideChat-based pipeline.
- Method: The two-stage framework performs domain alignment followed by instruction tuning for end-to-end report generation.Domain alignment uses 4.2K WSI-VQA pairs, while instruction tuning uses 176K WSI-caption pairs.
B.3. Pre-submission Evaluation
The challenge evaluation was designed to limit submission opportunities while supporting transparent local assessment and reproducibility checks. Teams could evaluate methods using publicly released code before submitting results.
- Evaluation policy: Each team could submit at most two submissions per phase, while the official evaluation code enabled local method evaluation before submission.Leaderboard access was restricted, but the evaluation code was publicly released.
- Reproducibility: Top-five teams had to provide source code for independent reproducibility verification before receiving awards.Submitted code was used internally and was not publicly released.
B.5. Data Usage and License Policy
REG2025 data use is restricted to attributed, non-commercial research and related activities under CC BY-NC-SA 4.0. The surrounding materials also document expert annotation and representative pathology concepts and cases used to examine model behavior.
- Data Usage and License Policy: The anonymized WSI–report dataset supports non-commercial research, education, development, benchmarking, and publication with attribution under CC BY-NC-SA 4.0.Commercial use, re-identification attempts, and clinical use are prohibited, subject to Grand Challenge terms and ethics approvals.
- Qualitative Analysis: Representative analyses cover organ-specific histopathological concepts across breast, lung, and prostate tissues.These concepts were curated to interpret model predictions qualitatively.
- Representative Cases: Supplementary cases examine attribute fidelity, diagnostic concordance, rare-entity recognition, site mismatches, and fine-grained histologic grading.Examples include breast, prostate, colon, stomach, lung, and colorectal-type morphology cases.