Source-linked AI summary

Model Effect or Label Effect? Refined Annotations and a Human-Referenced Benchmark for Pulmonary Embolism Segmentation

Qihang Sun, Zhongxiao Liu, Bailiang Jian, Shenman Qiu, Jingyuan Wang, Lei Zhang, Lixiang Xie, Jiazhen Pan, Christian Wachinger

arXiv:2608.24486v1eess.IVcs.CV

TL;DR

The study asks whether annotation quality or model training choices drive measured PE segmentation performance. It compares fixed predictions and training configurations, finding annotation changes at least as influential as training changes and establishing a human-referenced framework.

  • Problem

    Public PE segmentation datasets contain reported label errors, making differences in measured model performance difficult to distinguish from differences in annotation fidelity.

  • Method

    The study refines annotations across three public datasets, compares pretrained models against original and refined labels, evaluates model effects with fixed annotations, and benchmarks nnPE against human raters.

  • Results

    Annotation replacement changed DSC by 0.143–0.188, versus 0.028 from training-dataset composition, while refined annotations showed lower attenuation variability and nnPE remained below human raters.

  • Takeaways & Limitations

    FairPE, nnPE, and the evaluation framework provide a shared reference for attributing PE segmentation performance differences to models rather than annotations.

  • Takeaways & Limitations

    The multi-rater analysis covered only 15 cases, refined annotations came from one primary rater, and validation lacked histopathological and prospective multicenter confirmation.

Abstract

from arXiv · show

Purpose: To quantify how evaluation annotations influence measured pulmonary embolism (PE) segmentation performance relative to model training changes, and to establish a human-referenced framework. Materials and Methods: This retrospective study screened 166 voxel-annotated CT pulmonary angiography cases from CADPE (n=91), FUMPE (n=35), and READ (n=40); 149 were included. A primary rater annotated PE by protocol, and a senior thoracic radiologist reviewed and revised all segmentations. Three additional raters at three centers annotated a 15-case subset. The label effect was measured by evaluating two pretrained nnU-Net models (nnU-Net-A, nnU-Net-B) against original and refined annotations. The model effect was measured by comparing the same architecture trained on different dataset combinations with annotations fixed. The benchmark model (nnPE) was trained with leave-one-dataset-out and pooled five-fold cross-validation. Four metric categories were analyzed with case-paired Wilcoxon signed-rank tests, Benjamini-Hochberg correction, and bootstrap 95% CIs. Results: Changing only the annotation increased mean DSC by 0.143 (0.122-0.166) for nnU-Net-A and 0.188 (0.163-0.213) for nnU-Net-B (both P < .001), whereas changing training-dataset composition changed DSC by 0.028. The label effect exceeded the model effect on CADPE and FUMPE and was 0.045 on READ. Within-mask attenuation SD fell in all three datasets after re-annotation (all P < .001). nnPE reached DSC 0.72 +/- 0.22 on pooled cross-validation but scored below all four annotators across 52 paired comparisons (all corrected P < .05). Conclusion: Evaluation annotations affected measured PE segmentation performance at least as much as model training choices. A human-referenced evaluation framework is publicly available for future study.

1. Introduction

This study addresses whether differences in PE segmentation benchmarks reflect model capability or annotation quality. It introduces refined annotations and a human-referenced framework to separate these effects on shared cases.

  • Motivation: Accurate PE segmentation supports reliable lesion-volume quantification and downstream risk assessment from CTPA.CTPA is the standard imaging modality for PE diagnosis, and thrombus burden correlates with right ventricular dysfunction and risk stratification.
  • Motivation: Reported benchmark differences can arise from variable evaluation protocols, data leakage, and restrictive private-dataset selection.These factors include evaluating only positive slices and splitting anatomically adjacent slices from the same patient across partitions.
  • Research gap: Public PE datasets contain label errors, making annotation fidelity difficult to distinguish from model capability in reported comparisons.Reported errors include omitted emboli and partial-volume artifacts labeled as distal PE.
  • Contributions: The study quantifies label and model effects on the same 149 cases using an identical metric set and case-level pairing.It compares fixed predictions scored against original versus refined annotations and compares training-data compositions with annotations held fixed.
  • Contributions: The authors evaluate original masks using physical, human, and model-based criteria independent of the refined reference.These criteria include within-mask CT attenuation, agreement among four annotation sets on 15 cases, and externally trained models.
  • Contributions: FairPE and nnPE provide refined annotations, a human-referenced evaluation framework, and a baseline for future studies.The framework includes annotation protocols, multi-rater uncertainty ranges, multidimensional metrics, and a baseline trained on refined annotations.

2. Methods

The study re-annotated public CTPA datasets under a predefined protocol, evaluated annotation and model effects with paired metrics, and established cross-validation procedures for benchmarking.

  • Data and annotation: 149 cases from CADPE, FUMPE, and READ were included after screening 166 publicly available voxel-annotated CTPA cases.The datasets contributed 91, 35, and 40 cases, respectively; all were de-identified and CC BY 4.0 licensed.
  • Data and annotation: Annotations were drawn directly from CTPA images under a predefined protocol, with original masks used only as navigation aids.Delineations were performed in 3D Slicer 5.6.2.
  • Data and annotation: A primary rater created annotations for all cases, and a senior thoracic radiologist reviewed and revised them; three additional raters annotated 15 cases.The subset was stratified by the most proximal clot level, and raters were blinded to one another’s delineations.
  • Data and annotation: Original annotations were assessed for over-segmentation, under-segmentation, and anatomically implausible errors.Error categories included pulmonary artery wall, bronchi, lung parenchyma, pulmonary veins, and under-segmented PE; categories were not mutually exclusive.
  • Effect measurement: The label effect was measured by holding model predictions fixed while evaluating them against original and refined annotations.Two externally trained nnU-Net models were applied without modification to all three public datasets, and the per-case metric difference defined the label effect.
  • Evaluation: Performance was evaluated using overlap, boundary, volumetric, and embolus-level detection metrics.Metrics included DSC, ASSD, NSD at τ = 1 mm, absolute volume error, and lesion precision, recall, and F1 at three overlap thresholds.
  • Statistical analysis: Paired effects and confidence intervals were analyzed with case-level Wilcoxon tests and bootstrap resampling, while inter-rater agreement used ICC and pairwise spatial metrics.Bootstrap confidence intervals used 5000 resamples, and STAPLE generated a consensus reference for model and rater comparisons.
  • Study design: No formal a priori sample-size calculation was performed because the design included all publicly available pixel-annotated PE CTPA cases known at analysis time.The 15-case subset balanced rater workload across centers while preserving embolus-volume stratification.

3. Results

Across 149 included cases, refined annotations changed embolus measurements, inter-annotation agreement, and evaluated model performance. Annotation changes produced larger DSC shifts than training-set composition changes, while multi-rater comparisons established a human reference.

  • 3.1. Public dataset characteristics: 149 cases were included: CADPE 76, FUMPE 33, and READ 40.Central pulmonary-artery attenuation did not differ significantly across datasets (P = .477).
  • 3.2. Annotation characteristics: −5.53 mL mean bias showed that original annotations systematically overestimated embolus volume versus refined annotations.The 95% CI was −7.03 to −4.02 mL, with P < .001.
  • 3.2. Annotation characteristics: Mask HU SD decreased significantly after re-annotation in CADPE, FUMPE, and READ.The respective changes were 131.38 → 65.57, 121.46 → 61.03, and 101.77 → 76.76; all P < .001.
  • 3.3. Multi-rater agreement: ICC(2,1) = 0.961 indicated excellent case-level embolus-volume agreement among four annotations in the 15-case subset.Agreement between the original annotation and each refined annotation was below the pairwise range among refined annotations across reported metrics.
  • 3.4. Performance of pretrained weights against the two annotations: Lesion F1 improved after refinement for both pretrained models at the 10% overlap threshold.nnU-Net-A changed from 0.53 ± 0.28 to 0.67 ± 0.28, and nnU-Net-B from 0.43 ± 0.32 to 0.57 ± 0.31; both P < .001.
  • 3.4. Performance of pretrained weights against the two annotations: 0.143 and 0.188 pooled DSC increases resulted from changing only annotations for nnU-Net-A and nnU-Net-B, respectively.Both changes were significant at P < .001; nnU-Net-A improved from DSC 0.48 ± 0.25 to 0.63 ± 0.25.
  • 3.5. Label effect versus model effect: 0.028 was the DSC shift from varying training-set composition, smaller than the annotation effect; the label effect exceeded the Control effect in CADPE and FUMPE.In READ, the label effect was 0.045 (0.034–0.059), while the Control effect was not significant.

4. Discussion

Annotation quality influenced measured PE segmentation performance as strongly as model configuration, while the refined reference was supported by physical and human agreement analyses. The study also identifies measurement limits and scope boundaries for interpreting performance.

  • Evidence for refined annotations: Refined annotations were supported by lower within-mask attenuation variability and stronger agreement with independent refined annotation sets than with the original annotations.The multi-rater analysis used 15 cases, and attenuation variability decreased across all three datasets.
  • Annotation versus model effects: Annotation changes altered DSC by 0.143–0.188, several times more than the 0.028 change from training-dataset composition.With scans and predictions held constant, replacing original annotations with refined annotations changed DSC; varying training composition produced the smaller difference.
  • Human-referenced interpretation: Human reproducibility was mostly sub-millimetric, so a 1-mm NSD tolerance separates clinically meaningful boundary error from disagreement already present between experts.The 2 mL size threshold was selected near the center of the same subset’s volume distribution without a clinical consensus boundary.
  • Human-referenced interpretation: DSC combines a human disagreement floor on small lesions with model-specific deficits that remain visible in NSD and lesion-level F1.On smaller-volume cases, DSC decreased and dispersed for both raters and model, whereas NSD and lesion-level F1 remained tight among raters but fell for the model.
  • Limitations: Multi-rater estimates were limited by annotation of 15 cases, and refined labels remained subject to partial-volume uncertainty without histopathological confirmation.The refined annotations were produced by one primary rater and reviewed by a senior thoracic radiologist.
  • Limitations: The 149-case refined dataset may not represent the full range of scanner vendors, reconstruction kernels, and contrast protocols in routine practice, and prospective multicenter validation is absent.
  • Implications: FairPE, nnPE, and the evaluation framework provide a shared reference standard for attributing performance differences to models rather than annotations.

Funding

The study reports no specific grant funding from public, commercial, or not-for-profit agencies.

  • No specific grant supported this research from public, commercial, or not-for-profit funding agencies.

Supplementary Material

The supplementary annotation protocol specifies the imaging conditions, PE inclusion rules, label definitions, boundary handling, difficult-case exclusions, and software output format.

  • Imaging and reading conditions: CTPA pulmonary arterial-phase images were read primarily in a pulmonary artery window of W = 700 and L = 100, adjustable case by case.
  • PE identification criteria: PE was identified as an intraluminal pulmonary-artery filling defect, including complete, partial, and mural hypodense lesions.
  • Annotation definition: Foreground labels included occlusive, partially occlusive, mural, and saddle emboli, while remaining voxels were background.
  • Boundary rules: Lesion boundaries followed the visually separable thrombus–contrast interface rather than a global HU threshold, with distal lesions requiring identification on at least two consecutive slices.
  • Difficult cases: Flow and mixing artifacts were excluded, as were poorly opacified cases below 200 HU and distal emboli obscured by respiratory motion.
  • Tools and output: Annotations were created in 3D Slicer 5.6.2 and saved in NRRD format.

S2. Original annotation errors

Original annotation errors included over-segmentation, annotation noise, and missed regions. Evaluation used complementary voxel-, boundary-, volume-, and embolus-level measures, including symmetric lesion-level agreement.

  • S2. Original annotation errors: Representative errors included over-segmentation into pulmonary arteries or veins, internal voids, and missed regions.These examples are shown in Figure S2.
  • S2. Original annotation errors: Segmentation performance was assessed using voxel-level overlap, boundary accuracy, volumetric agreement, and embolus-level detection.The metric set included DSC, ASSD, NSD, absolute volume error, and lesion precision, recall, and F1.
  • S2. Original annotation errors: ASSD measures mean bidirectional surface distance in millimeters, while NSD measures bidirectional surface agreement within a 1-mm tolerance.Both metrics quantify boundary accuracy rather than voxel overlap.
  • S2. Original annotation errors: Embolus-level detection classifies predicted and reference emboli using overlap thresholds of 1 pixel, 10%, or 20%.Lesion precision and recall are derived from the two-stage volume-aware framework, and lesion F1 combines the two directional recalls.
  • S2. Original annotation errors: The symmetric lesion-level formulation also supports inter-rater agreement analyses between raters of equivalent status.The two directional recalls are used to recover the conventional F1 score.

S4. Basic characteristics of the public datasets

Table S4 summarizes patient demographics and CT acquisition characteristics for CADPE, FUMPE, and READ. Demographic values are reported from available source data when reporting was incomplete.

  • S4. Basic characteristics of the public datasets: Table S4 presents patient demographics and CT acquisition characteristics for the three public CTPA datasets.The datasets are CADPE, FUMPE, and READ.
  • S4. Basic characteristics of the public datasets: Demographic statistics were calculated only from available reported data when source publications provided incomplete reporting.Continuous variables are summarized as mean ± SD and categorical variables as n (%), unless otherwise indicated.
  • S4. Basic characteristics of the public datasets: NR denotes values not reported in the source publication.

S5. Multi-metric inter-rater agreement heatmaps

Figure S5 visualizes pairwise agreement among four annotation sets and the original annotation across six metrics. Agreement is shown across all jointly annotated cases and separately by embolus volume.

  • S5. Multi-metric inter-rater agreement heatmaps: Each heatmap panel is a 5 × 5 matrix of pairwise agreement among Annotation 1–4 and the Original annotation.
  • S5. Multi-metric inter-rater agreement heatmaps: The heatmaps cover DSC, ASSD, NSD at 1 mm, and lesion F1 at 1-pixel, 10%, and 20% tolerances.
  • S5. Multi-metric inter-rater agreement heatmaps: Panels a–f summarize lower-triangular pairwise means across all 15 jointly annotated cases.
  • S5. Multi-metric inter-rater agreement heatmaps: Panels g–l stratify the same metrics by embolus volume after splitting cases into smaller- and larger-volume halves.

S6. Lesion-level detection performance of the pretrained weights on the two annotations

The supplementary materials report lesion-level detection performance for nnU-Net weights evaluated on Original versus Refined annotations and across training-set combinations. Results use defined overlap hit criteria and dataset-specific training labels.

  • S6. Lesion-level detection performance of the pretrained weights on the two annotations: Table S6 reports lesion-level detection performance for different nnU-Net weights on Original versus Refined annotations.
  • S6. Lesion-level detection performance of the pretrained weights on the two annotations: Within-dataset Original-versus-Refined comparisons use the Wilcoxon signed-rank test, with significance indicated at P < .05, P < .01, and P < .001.Results are reported as mean ± SD.
  • S6. Lesion-level detection performance of the pretrained weights on the two annotations: The 1px and 20% criteria denote at least 1-voxel overlap and overlap exceeding 20% of reference lesion volume, respectively.
  • S6. Lesion-level detection performance of the pretrained weights on the two annotations: Table S7 reports lesion-level detection across datasets under different training-set combinations at the ≥1-pixel and ≥20% overlap thresholds.
  • S6. Lesion-level detection performance of the pretrained weights on the two annotations: Dataset labels A, B, and C correspond to CADPE, FUMPE, and READ, while ABC denotes pooled internal five-fold cross-validation.

S8. Paired comparison of nnPE against the four human raters

On a 15-case test set, nnPE was compared with four human raters across lesion-level metrics using paired, multiple-comparison-corrected tests. The model performed worse than the raters across the reported comparisons.

  • Comparison framework: nnPE scored below all four human raters across 52 paired comparisons, with globally Benjamini–Hochberg-corrected P values indicating significant differences.The comparison used one-sided Wilcoxon signed-rank tests with the alternative hypothesis that the model performed worse than the rater.
  • Lesion-level metrics: Lesion Recall (1px) was 0.76 ± 0.23 for nnPE versus 0.94–0.99 for the annotations, with corrected P < 0.01 across all four comparisons.The matched-pairs rank-biserial correlations were −0.96 to −1.00.
  • Lesion-level metrics: Lesion Precision (1px) was 0.79 ± 0.24 for nnPE versus 0.88–0.97 for the annotations, with corrected P values from 0.02 to 0.04.All reported matched-pairs rank-biserial correlations were negative, ranging from −0.56 to −0.69.
  • Threshold sensitivity: The same pattern held at the 10% and 20% lesion thresholds, where nnPE had lower F1 and recall than the human annotations in the reported comparisons.For example, Lesion F1 (10%) was 0.65 ± 0.25 for nnPE versus 0.88–0.95 for annotations, while Lesion Recall (10%) was 0.69 ± 0.30 versus 0.94–0.98.

S9. Qualitative results and representative failure modes

Figure S9 presents qualitative model outputs and representative failure modes for a model trained jointly on all three datasets. Each case juxtaposes the CTPA image, reference annotation, prediction overlay, and a magnified highlighted region.

  • Qualitative results: Each row shows one representative case with the input CTPA image, GroundTruth annotation, Prediction overlay, and Zoomed highlighted region.The figure is organized to support qualitative comparison between reference embolus annotations and model predictions.
  • Visual encoding: Green denotes the reference embolus annotation, whereas red denotes the model prediction.
  • Representative failure modes: Red dashed boxes and a red arrow identify highlighted regions used for magnified inspection.
Loading 2608.24486v1…