Source-linked AI summary

Evaluating the Effects of Inter-Observer and Model Variability on Radiological Peritoneal Cancer Index Assessment

Savvas Saragiotis, Pieter C. Gort, Lotte J. S. Fleurkens-Ewals, Anna F. van Herwijnen, Marion Tops-Welten, L. D. Kampmeijer, Joost Nederend, Fons van der Sommen

arXiv:2608.28716v1eess.IVcs.CV

TL;DR

The paper asks whether geometric segmentation metrics capture clinically meaningful effects on rPCI-based decisions. It compares human and nnU-Net variability on multi-observer CT annotations and propagates region differences through probabilistic metastasis simulations. Geometric deviations were larger for the model in several regions, but simulated PCI differences were generally small and decision changes concentrated near PCI 20.

  • Problem

    It remains unclear whether improvements or differences in Dice, HD95, and ASD translate into clinically meaningful changes in downstream rPCI-based decisions.

  • Method

    The study compares four observers with a published nnU-Net model on ten CT scans and simulates metastases on majority-vote rPCI maps to evaluate PCI variability at the PCI 20 threshold.

  • Results

    Model–observer agreement was lower than inter-observer agreement in several regions, while simulated PCI-score deviations were generally less than one point and misclassifications clustered near PCI 20.

  • Takeaways & Limitations

    rPCI-derived scoring is generally robust to typical region-segmentation variability, but borderline cases near clinical thresholds require careful expert review.

  • Takeaways & Limitations

    The simulation isolates region-assignment variability and does not account for missed lesions or acquisition-related factors, so the reported deviations are a lower bound on total rPCI uncertainty.

Abstract

from arXiv · show

Deep learning segmentation models are often evaluated using geometric metrics such as Dice, HD95, and ASD, yet it remains unclear to what extent improvements in these metrics translate into clinically meaningful changes in downstream decision-making. The metric-to-decision gap is examined using radiological Peritoneal Cancer Index (rPCI) region segmentation on contrast-enhanced CT, where a consensus definition provides anatomically grounded 3D regions and the clinically used PCI 20 threshold enables decision-level evaluation. Inter-observer variability is quantified across four experts on ten abdominal CT scans, and a published nnU-Net based rPCI segmentation model is benchmarked against this human reference using Dice, HD95, and ASD across all 13 regions. To relate geometric differences to clinical impact, a probabilistic peritoneal metastasis simulation is implemented on majority-vote rPCI maps, propagating region-boundary variability into variability of derived (r)PCI scores and classification at the PCI 20 cutoff. Observers showed high agreement (mean Dice $0.87$), while the model matched human performance in most regions but deviated more in regions 4, 8, and the small-bowel regions (9-12). Across simulations, score differences were typically small (mean $Δ$rPCI $\approx 0.3$-$0.6$) for both observers and the model, and decision flips occurred predominantly when the reference score was near 20. These results suggest that rPCI-derived scoring is generally robust to typical segmentation variability, while highlighting borderline cases as the main setting where expert review remains essential.

1 Introduction

The study addresses whether geometric segmentation metrics reflect clinically meaningful changes in rPCI-based treatment decisions. It compares human and model variability and propagates segmentation differences into PCI scores at the clinical PCI 20 threshold.

  • Motivation: Geometric metrics quantify mask agreement but do not directly show whether segmentation improvements change downstream clinical decisions.The paper frames this as a metric-to-decision gap in medical image segmentation.
  • Clinical setting: PCI divides the abdomen into 13 regions scored from 0–3, producing totals from 0–39; PCI 20 is commonly used for colorectal treatment-eligibility decisions.This threshold enables decision-level analysis of segmentation variability.
  • Clinical setting: rPCI regions provide consensus-based, non-invasive estimates on contrast-enhanced CT, but delineation is particularly challenging in small-bowel regions.Inter-observer variability is needed to benchmark automated segmentation beyond geometric metrics.
  • Research question: The key question is how model-induced rPCI variability compares with human inter-observer variability and affects decisions at the PCI 20 cutoff.The comparison uses Dice, HD95, and ASD alongside decision-level outcomes.
  • Approach: The study quantifies variability across four radiologists and evaluates how probabilistic peritoneal-metastasis simulations propagate segmentation differences into PCI scores and threshold classifications.The analysis uses ten abdominal CT scans and examines classification around PCI 20.

2 Methods

The study uses ten contrast-enhanced CT scans with multi-observer rPCI annotations to compare human variability with a published segmentation model. A probabilistic metastasis simulation tests how region-assignment differences affect derived PCI scores.

  • Data and annotations: Ten contrast-enhanced abdominal CT scans were divided into cohorts A and B, with three independent annotations per cohort across four clinical researchers.Observers segmented all 13 rPCI regions using consensus definitions.
  • Clinical variability simulation: Simulated CT examples display nodules using colored overlays for PCI lesion-size classes LS1–LS3.The figure illustrates the lesion configurations used in the simulation.
  • Reference construction: Majority-vote label maps assigned each voxel the rPCI region selected by at least two of three annotators.These maps served as the reference for the clinical variability simulation.
  • Clinical variability simulation: The simulation sampled lesion voxels within rPCI regions and queried observer or model region assignments to quantify disagreement.This propagates segmentation variability into simulated clinical scoring.
  • Qualitative assessment: Qualitative agreement heatmaps and contour overlays provide a visual assessment of inter-observer segmentation differences.These visualizations complement the quantitative variability analysis.

3 Results

Observers showed high overall rPCI agreement, while model–observer agreement was lower and especially weak in small-bowel and selected abdominal regions. The qualitative overlays visualize these regional discrepancies.

  • Inter-observer variability: 0.87 ± 0.04 mean inter-observer Dice indicated high agreement across the 13 rPCI regions.Overall inter-observer HD95 was 8.8 ± 3.0 mm and ASD was 2.4 ± 1.1 mm.
  • Inter-observer variability: 0.86 ± 0.01 Dice in small-bowel regions 9–12 was lower than 0.88 ± 0.05 in regions 0–8.Small-bowel regions also had higher HD95 and ASD, with disagreement spatially localized around the bowel.
  • Human–model variability: 0.82 ± 0.06 mean model–observer Dice was lower than inter-observer agreement across all metrics.Model–observer HD95 was 20.8 ± 17.0 mm and ASD was 4.4 ± 2.4 mm.
  • Human–model variability: 0.77 ± 0.04 model–observer Dice in regions 9–12 was lower than 0.84 ± 0.06 in regions 0–8.The largest deviations occurred in regions 4, 8, 11, and 12, with increased HD95 and ASD.
  • Qualitative comparison: Axial, coronal, and sagittal overlays compare majority-vote reference segmentations, model predictions, and segmentation errors.The figure provides qualitative context for the reported model–observer discrepancies.

4 Discussion

The study frames segmentation evaluation around decision-level consequences rather than geometry alone, finding that regional model errors can exceed human variability while simulated PCI classifications remain generally robust. Borderline cases near PCI 20 remain the main setting requiring expert attention.

  • Decision-level framing: The framework propagates segmentation variability into PCI variability and tests its impact at a clinically used treatment threshold.The study is designed to support deployment decisions and identify when expert review is most critical.
  • Regional variability: Table 1 contrasts human inter-observer and model–observer agreement using Dice, HD95, ASD, and regional differences.The comparison highlights whether model variability stays within the human reference range.
  • Decision-level outcomes: Table 2 summarizes region-level accuracy, rPCI differences, and classification performance at PCI 20.These measures connect geometric and region-assignment variability to clinical threshold outcomes.
  • Decision-level outcomes: Predicted-versus-reference PCI plots categorize cases at the cutoff into true negatives, true positives, false positives, and false negatives.The figure is intended to show how threshold classifications correspond to reference scores across cohorts A and B.
  • Regional variability: Model errors exceeded human variability in some regions, especially regions 4 and 8 and the small-bowel segments, despite strong agreement in most regions.The model’s regional pattern only partly matches known definitional ambiguity, suggesting additional sensitivity to local anatomical factors.
  • Decision-level outcomes: Misclassifications cluster near PCI 20, where small rPCI shifts can change treatment eligibility.The model’s higher specificity and lower sensitivity suggest a conservative bias around the threshold.

5 Conclusion

The study quantifies human and model-induced rPCI variability and propagates it into PCI scores and threshold classifications. Geometric differences generally produced small score deviations, but borderline cases near PCI 20 remained most vulnerable to misclassification.

  • The study introduces a framework combining inter-observer variability measurement with simulation of its impact on PCI-based decisions.The approach evaluates segmentation variability against a clinical decision threshold rather than relying only on geometric agreement.
  • Model-induced variability was generally within the human inter-observer range in well-defined regions but exceeded it in several regions.
  • Simulated PCI scores deviated by less than one point on average despite geometric differences.
  • Most predictions fell on the correct side of the PCI 20 threshold, with misclassifications concentrated near borderline cases.The simulation isolates region-assignment variability and therefore represents a lower bound on total rPCI uncertainty.
  • Cases near clinically relevant thresholds still require careful expert review, while future validation should use larger multi-center cohorts and surgical PCI.
Loading 2608.28716v1…