Source-linked AI summary

ImageCAS-X: a dataset and benchmark for coronary artery segmentation and centerline extraction in coronary CT angiography

Kit M. Bransby, Esther Øksnebjerg, Kristoffer Kjær, Jacob Kirkeby, Yasmin El Youssef, Aïda Jiménez, Philip R. Pedersson, Martina C. de Knegt, Klaus F. Kofoed, Rasmus R. Paulsen

arXiv:2608.30404v1cs.CVcs.AI

TL;DR

Coronary lumen segmentation in CCTA is difficult to validate because manual annotation is labour-intensive and large, high-quality public datasets are scarce. ImageCAS-X supplies voxel-wise lumen and segment labels, centerlines, and meshes for 800 scans, then benchmarks automated methods against inter-observer variability across clinical and anatomical factors. The benchmark shows automated methods remain below human-level agreement, while performance varies systematically with vessel characteristics and supports clinically contextualised evaluation and downstream modelling.

  • Problem

    Large, high-quality publicly available datasets are lacking for rigorous validation of coronary lumen segmentation and centerline methods in CCTA.

  • Method

    The authors created voxel-wise lumen and segment annotations, centerlines, and mesh surfaces for 800 ImageCAS scans and benchmarked automated methods across clinical and anatomical strata.

  • Results

    Automated methods did not yet match human-level agreement, while segmentation performance was higher with greater lumen attenuation and diameter and lower at more distal positions.

  • Takeaways & Limitations

    The dataset supports independent benchmarking and development of lumen, plaque, perivascular, and haemodynamic analysis methods using clinically meaningful labels.

Abstract

from arXiv · show

Accurate segmentation of the coronary vessel lumen is a prerequisite for quantitative assessment of atherosclerotic plaque and perivascular adipose tissue in coronary computed tomography angiography (CCTA). Cardiologists rely on semi-automated methods for this task because manual vessel tracing and segmentation are labour-intensive. Although many automated methods have been proposed, their validation remains limited by the lack of large, high-quality publicly available datasets. We provide a new dataset of voxel-wise annotations of the vessel lumen and coronary segments, alongside centerlines, and mesh surfaces for 800 scans from the publicly available ImageCAS dataset. Using this dataset, we benchmark established lumen segmentation methods against inter-observer variability, stratifying performance by disease, image quality, coronary dominance, coronary segment, vessel diameter, and lumen attenuation. These labels allow segmentation accuracy to be described in anatomical and clinical context rather than reported as a single aggregate score. The dataset supports the development and validation of methods for lumen segmentation, plaque and perivascular quantification, and haemodynamic modelling.

Background & Summary

ImageCAS-X addresses limited public resources for validating coronary lumen segmentation by providing high-quality labels for anatomical and clinically stratified evaluation. The dataset supports benchmarking against human agreement and downstream plaque, perivascular, and haemodynamic analyses.

  • CCTA lumen delineation is technically difficult because of artefacts, limited resolution, and the need to preserve continuous branching topology.Expert annotation is labour-intensive and variable, while guidelines still recommend semi-automated workflows.
  • Earlier public coronary benchmarks were small or unavailable, limiting large-scale validation of segmentation and tracing methods.CAT08 and the 2012 challenge contained 32 and 48 cases, respectively.
  • ImageCAS-X provides voxel-wise coronary lumen segmentations with mesh surfaces and centerlines for the publicly available ImageCAS dataset.Additional labels support assessment by anatomical segment, disease, and coronary dominance.
  • Public, balanced annotations enable independent benchmarking against inter-observer variability and support plaque, perivascular tissue, and haemodynamic modelling applications.Mesh surfaces additionally support CT-FFR, haemodynamic simulation, and stent-planning algorithms.

Methods

The authors retrospectively selected and quality-screened ImageCAS CCTA scans, then created corrected, segment-labelled lumen masks, centerlines, and mesh surfaces through a semi-automated annotation workflow. The resulting 800 cases were partitioned for model development and blinded inter-observer assessment.

  • The source cohort comprised 1,000 ImageCAS CCTA scans collected at Guangdong Provincial People’s Hospital between 2012 and 2018.The cohort included 586 males and 414 females with documented vascular disease histories.
  • 200 scans were excluded for non-diagnostic quality, leaving 800 cases graded from poor to excellent.The main exclusions were motion artefacts (114) and step artefacts (76).
  • Scan Characteristics: The retained scans included 729 right-dominant, 41 left-dominant, and 30 co-dominant coronary trees, with 388 diseased and 412 non-diseased cases.Disease labels were based on identifying calcified, noncalcified, or mixed atherosclerotic lesions.
  • Vessel Segmentation: Analysts generated and corrected centerlines, named them under the 18-segment model, and used corrected centerlines to guide lumen annotation.A 3D U-Net trained on 100 manually annotated cMPR vessel volumes produced initial lumen surface predictions for all scans.
  • Postprocessing: Corrected masks were skeletonised and smoothed to generate centerlines, with segment assignments reviewed manually and mesh surfaces produced using marching cubes and Taubin smoothing.The workflow also identified centerline start points near the segmented aorta.
  • Dataset Split: The 800 cases were split into 560 training, 80 validation, and 160 test cases, with blinded re-annotation of every test case by another analyst.The additional labels enabled inter-observer variability assessment.

Data Records

The ImageCAS-X repository is publicly accessible and organised around segmentation masks, centerlines, mesh surfaces, file lists, and scan-level descriptors. It preserves patient identity and subset membership while providing structured anatomical and clinical metadata.

  • The repository provides segmentation masks, centerlines, mesh surfaces, and file lists through publicly accessible Zenodo records.Masks use compressed NIfTI format, while centerlines and mesh surfaces use VTK format.
  • Centerlines encode start, bifurcation, and end vertices plus categorical coronary-segment assignments.Segment labels are stored as centerline attributes for each vertex.
  • Each patient has a unique ImageCAS-matching ID, and train, validation, test, and exclude membership is supplied in text files.Descriptors.xlsx contains labels for coronary dominance, disease, and image quality.

Technical Validation

The benchmark evaluates automated coronary lumen segmentation and centerline extraction against analysts across clinical, anatomical, and image-quality factors. CAS-Net performed best among automated methods, but automated predictions remained below human-level agreement and contained topological errors despite high overlap scores.

  • 91.9 versus 93.6 DSC, with plaque-associated scans performing significantly worse than scans without plaque (p < 0.001).Plaque can resemble contrast material and produce more complex lumen morphology.
  • Main coronary branches achieved DSC 84.8–95.3, whereas side branches achieved 70.9–83.6, with significant differences across segments for all metrics (p < 0.001).Smaller diameters, lower attenuation, and ambiguity about distal tracing make side branches more difficult to segment.
  • Local DSC increased with lumen attenuation (ρ = + 0.90) and diameter (ρ = + 0.89), but decreased with distal position (ρ = - 0.36); all trends had p < 0.001.Distal vessels tend to have lower contrast concentration and smaller diameters, complicating tissue differentiation.
  • CAS-Net was the best-performing automated method, achieving a DSC of 91.2, HD95 of 2.99 mm, and βerr of 1.9 for lumen segmentation.
  • Topological errors such as vessel breaks occurred in all model predictions across worst, median, and best performance examples despite high DSC and clDice.
  • All models predicted lumen segmentations in < 2 minutes, compared with analysts’ average of 35 minutes per scan; CAS-Net was also the most computationally efficient.

Data Availability

The dataset is available through Zenodo under CC BY 4.0, while the original ImageCAS volumes are distributed separately through Kaggle under Apache 2.0.

  • The generated dataset is available for download via Zenodo under a CC BY 4.0 license.
  • Original ImageCAS volumes are not redistributed and are provided by the original authors through Kaggle under an Apache 2.0 license.

Funding

The research was funded by Novo Nordisk A/S, which had no role in the study’s design, data collection, analysis, or manuscript preparation.

  • Novo Nordisk A/S provided funding for the research.
  • The funder had no role in study design, data collection, analysis, or manuscript preparation.

A Segmentation accuracy in the ImageCAS dataset

The study used analyst-refined coronary trees, operational labelling rules, and a 3D U-Net-assisted workflow to address the time and ambiguity involved in lumen annotation.

  • A Segmentation accuracy in the ImageCAS dataset: Three supplementary visual examples explain why the ImageCAS (2023) dataset is unsuitable for training and validating CCTA lumen segmentation methods.
  • A Segmentation accuracy in the ImageCAS dataset: ImageCAS segmentation issues include plaque inclusion, false-positive pulmonary vessels, and false-positive coronary veins.
  • A Segmentation accuracy in the ImageCAS dataset: Side branches were omitted when their connections or distal extents were unclear, when they could not be delineated, or when their distal diameter was below 1 mm.
  • A Segmentation accuracy in the ImageCAS dataset: At main-branch bifurcations, continuation versus side-branch labels were assigned according to arterial direction and course rather than size.
  • A Segmentation accuracy in the ImageCAS dataset: Analysts resolved ambiguous cases using a guide for anatomical variation in the left circumflex artery.
  • A Segmentation accuracy in the ImageCAS dataset: A 3D U-Net generated initial lumen contours from manually annotated cMPR volumes, reducing the labour required for segmentation from scratch.

E Shared training and inference framework

All methods share fixed preprocessing, augmentation, optimisation schedules, training budgets, post-processing, test-time augmentation, and metrics, while retaining method-specific design choices for fair comparison.

  • E Shared training and inference framework: The framework fixes shared components for fair comparison while retaining method-specific architectures, losses, input sampling, batch sizes, and learning rates.
  • E Shared training and inference framework: Every scan is resampled to isotropic 0.5 mm spacing, clipped to [−200, 1000] HU, and mapped linearly to [0, 1].
  • E Shared training and inference framework: Training augmentation includes noise, blur, brightness, contrast, simulated low resolution, gamma transformations, and coronal and sagittal axis flips.
  • E Shared training and inference framework: Each epoch comprises 250 training and 50 validation batches sampled with replacement, with common training duration and optimisation schedules across methods.
  • E Shared training and inference framework: Inference uses single-pass processing for whole-volume methods and 50%-overlap sliding-window tiling for patch- and crop-based methods.

F Method-specific implementation notes

Baseline implementations used standardized hyperparameter reporting and adapted training procedures to accommodate computational constraints and dataset-specific performance. Ensemble construction was also modified based on validation findings.

  • Baseline hyperparameters are summarized in Supplementary Table 5, with implementation details provided for each method.
  • The benchmark reports CE, DS, HD, and MSE as abbreviations for cross-entropy, deep supervision, Hausdorff distance, and mean squared error.
  • The authors omitted an axial-slice identification step because foreground over-sampling in the patch data loader already addressed the same preprocessing need.
  • Pure dice loss replaced a proposed multi-loss recipe after producing suboptimal results on the authors’ data.
  • Key-value average pooling reduced the key set by 64×, while dice and cross-entropy replaced unspecified weighted cross-entropy for connectivity prediction.
  • Probability ensembling outperformed binary-mask majority voting, and excluding the stage 1 coarse segmentation model further improved performance.
Loading 2608.30404v1…