Source-linked AI summary
CHAOS Challenge -- Combined (CT-MR) Healthy Abdominal Organ Segmentation
A. Emre Kavur, N. Sinem Gezer, Mustafa Barış, Sinem Aslan, Pierre-Henri Conze, Vladimir Groza, Duc Duy Pham, Soumick Chatterjee, Philipp Ernst, Savaş Özkan, Bora Baydar, Dmitry Lachinov, Shuo Han, Josef Pauli, Fabian Isensee, Matthias Perkonigg, Rachana Sathish, Ronnie Rajan, Debdoot Sheet, Gurbandurdy Dovletov, Oliver Speck, Andreas Nürnberger, Klaus H. Maier-Hein, Gözde Bozdağı Akar, Gözde Ünal, Oğuz Dicle, M. Alper Selver
TL;DR
Abdominal-organ segmentation lacks comprehensive benchmarks spanning CT, MRI, cross-modality learning, and multi-organ tasks. CHAOS addresses this gap with five tasks on unpaired healthy-subject CT and MR data, consensus expert annotations, and hidden-test evaluation. Single-modality models achieve reliable volumetric performance but remain limited on distance measures, while cross-modality performance declines and repeated submissions expose challenge-design concerns.
Problem
Existing abdominal-segmentation evidence is limited across CT, MRI, cross-modality, and multi-organ settings, motivating broader benchmark evaluation.
Method
CHAOS evaluates five complementary segmentation tasks using unpaired CT and MR scans, consensus-based expert annotations, hidden manual test delineations, and multiple metrics.
Results
Cross-modality performance decreases substantially, while single-modality models provide reliable volumetric analysis but remain limited on distance-based metrics.
Takeaways & Limitations
Cross-modality segmentation and clinically integrated post-processing remain important directions for applying abdominal-organ segmentation in real-world workflows.
Takeaways & Limitations
Iterative submissions can enable peeking at test data and tuning algorithms without access to ground truth, making peeking a challenge shortcoming.
Abstract
from arXiv · showhide
Segmentation of abdominal organs has been a comprehensive, yet unresolved, research field for many years. In the last decade, intensive developments in deep learning (DL) have introduced new state-of-the-art segmentation systems. In order to expand the knowledge on these topics, the CHAOS - Combined (CT-MR) Healthy Abdominal Organ Segmentation challenge has been organized in conjunction with IEEE International Symposium on Biomedical Imaging (ISBI), 2019, in Venice, Italy. CHAOS provides both abdominal CT and MR data from healthy subjects for single and multiple abdominal organ segmentation. Five different but complementary tasks have been designed to analyze the capabilities of current approaches from multiple perspectives. The results are investigated thoroughly, compared with manual annotations and interactive methods. The analysis shows that the performance of DL models for single modality (CT / MR) can show reliable volumetric analysis performance (DICE: 0.98 $\pm$ 0.00 / 0.95 $\pm$ 0.01) but the best MSSD performance remain limited (21.89 $\pm$ 13.94 / 20.85 $\pm$ 10.63 mm). The performances of participating models decrease significantly for cross-modality tasks for the liver (DICE: 0.88 $\pm$ 0.15 MSSD: 36.33 $\pm$ 21.97 mm) and all organs (DICE: 0.85 $\pm$ 0.21 MSSD: 33.17 $\pm$ 38.93 mm). Despite contrary examples on different applications, multi-tasking DL models designed to segment all organs seem to perform worse compared to organ-specific ones (performance drop around 5\%). Besides, such directions of further research for cross-modality segmentation would significantly support real-world clinical applications. Moreover, having more than 1500 participants, another important contribution of the paper is the analysis on shortcomings of challenge organizations such as the effects of multiple submissions and peeking phenomena.
1. Introduction
Biomedical imaging benchmarks provide common datasets and structured comparisons for learning-based systems, but challenge design can limit how well results reflect true capability. CHAOS addresses gaps in abdominal-organ segmentation benchmarking by evaluating CT and MR data across complementary tasks.
- Benchmarks enable standardized training, testing, and comparison of medical image-analysis approaches on clinically important tasks.
- Dataset construction, observer variation in ground truth, and evaluation criteria can prevent challenges from establishing their true potential.
- Existing abdominal-organ challenges are dominated by CT scans and tumor or lesion classification, with few benchmarks containing abdominal MRI series.
- MRI advances in resolution, dynamic range, and speed enable joint analyses of CT and MR modalities.
- CHAOS organized an ISBI 2019 challenge with unpaired CT and MR scans, consensus-based expert annotations, hidden test delineations, and multiple evaluation metrics.
2. Related Work
Prior abdominal-imaging challenges largely emphasize CT and disease-focused tasks, while newer benchmarks broaden evaluation toward multiple organs, generalizability, and cross-modality learning. CHAOS builds on these directions with unpaired CT–MR segmentation tasks.
- The abdominal-imaging challenge landscape includes 12 challenges, with pioneering benchmarks such as SLIVER07 focused on liver segmentation under deliberately varied CT conditions.
- Table 1 summarizes upper-abdomen challenges and their tasks, while excluding other structures from the overview.
- Medical Segmentation Decathlon evaluated structures across diverse tasks to study the generalizability, translatability, and transferability of deep-learning systems.
- Recent work shifts from organ- or disease-specific segmentation toward multi-organ tasks that represent complex and flexible abdominal anatomy.
- CHAOS targets cross-modality learning and multi-modal segmentation of multiple organs from unpaired CT and MR datasets, including two MR pulse sequences.
3. CHAOS Challenge
CHAOS provides healthy abdominal CT and MR datasets with expert annotations for five related segmentation tasks spanning single-modality, cross-modality, and multi-modal settings. The challenge uses unpaired data, varied organ outputs, and hidden testing to examine robustness in clinically relevant workflows.
- Data Information and Details: The CHAOS dataset contains 80 healthy subjects: 40 underwent CT and 40 underwent MR with two pulse-sequence families.
- CT Data Specifications: CT data were acquired in the contrast-enhanced portal venous phase and annotated only for liver segmentation.
- MRI Data Specifications: MRI data comprise 120 DICOM datasets from registered T1-DUAL in-phase and opposed-phase images plus unregistered T2-SPIR images.
- Aims and Tasks: The five tasks cover liver segmentation in CT, MRI, and mixed CT–MRI data, plus multi-organ segmentation in MRI and mixed modalities.
- Aims and Tasks: Task 5 extends MRI liver segmentation to four abdominal organs, whereas Task 4 combines CT–MRI input while CT provides liver labels and MRI provides four-organ labels.
- Ground Truth Generation: Three radiology experts labeled all 2D slices, with majority voting determining final reference shapes and joint decisions resolving exceptional cases.
- Dataset Organization: Training and testing each used 20 CT and 20 MRI sets, with stratification intended to balance resolution, slice thickness, and patient-age characteristics.
- Dataset Organization: The challenge distributed anonymized DICOM data and image-series ground truth under a CC-BY-SA 4.0 license for long-term academic study.
4. Evaluation
CHAOS evaluates segmentation with multiple complementary metrics, combines normalized metric scores into an overall ranking, and examines variability from repeated and independent annotations.
- Evaluation metrics: Four metrics assess overlapping, volumetric, and spatial differences between segmentation results and ground truth.The metrics are DICE, RAVD, ASSD, and MSSD.
- Evaluation metrics: DICE measures overlap, with larger values indicating better segmentation.It is computed from the intersection and cardinalities of the segmentation and ground-truth voxel sets.
- Evaluation metrics: RAVD measures relative absolute volume difference as a percentage, with smaller values indicating better segmentation.It compares the absolute volume difference against the ground-truth volume.
- Evaluation metrics: ASSD and MSSD measure symmetric surface distances in millimeters, with smaller values indicating better segmentation.ASSD averages border-voxel Hausdorff distances, whereas MSSD uses the maximum distance.
- Scoring system: Metric outputs are transformed to [0, 100], thresholded, averaged across four metrics and test cases, and used to calculate each team’s overall task score.Values outside threshold ranges receive zero points, as do missing test cases.
- Annotation variability: Annotation variability depends on modality and metric: CT shows narrower changes than MRI, while DICE varies less than the other metrics.The analysis compares repeated labels by the same annotator with labels from different annotators.
5. Participating Methods
Participating methods were predominantly U-Net variants, with additional adversarial, attention, residual, multi-task, multi-modal, and ensemble strategies represented across the challenge submissions.
- Method landscape: Most participating methods used U-Net variations, while two studies relied on ensembles.The reported methods include conference participants and selected post-conference online submissions.
- Architectural strategies: PKDIA combined conditional generative adversarial networks with cascaded partially pretrained encoder-decoder networks extending standard U-Net.Its encoder was replaced with a deeper VGG-19-based network.
- Ensembles: MedianCHAOS used an averaged ensemble of five networks, including DualTail-Net and four U-Net architecture variants.The component models included TernausNet, LinkNet34, ResNet-50, and SE-ResNet50-based networks.
- Multi-modal strategies: CIR MPerkonigg trained one framework jointly across modalities using IVD-Net, augmentation, and modality dropout.Modality dropout was used when training with multiple modalities to reduce overfitting on particular modalities.
- Multi-modal strategies: nnU-Net submitted ensembles of three-dimensional full-resolution U-Nets for Tasks 3 and 5 without external data.The models originated from cross-validation, and Task 3 predictions were generated by isolating the liver label from Task 5.
6. Results
CHAOS results show strong single-modality segmentation, but performance and robustness decline for MR, cross-modality learning, and some multi-organ settings. The challenge also exposed variability across cases, architectural limitations, and differences between automatic and interactive approaches.
- Participation: 550 submitted results came from more than 1500 registered participants, with multiple submissions contributing to this disparity.The challenge imposed no registration restrictions, allowing passive and highly active participants.
- CT liver segmentation: 0.98±0.00 DICE was achieved by both Task 2 winners, while interactive approaches were outperformed by a large margin.Deep learning methods approached inter-expert quality for volumetric analysis and average surface differences, although maximum-error metrics remained weaker.
- MR liver segmentation: 70.71±6.40 was the on-site winning score for MR liver segmentation, while MR standardization, artifacts, and image-quality variation complicated performance.The online winner scored 75.10±7.61; DICE and ASSD were strong, but RAVD and MSSD were lower than CT results.
- Cross-modality segmentation: Cross-modality models performed worse than models trained on individual modalities for both CT and MR data.The comparison indicates that a single solution for multiple modalities still requires improvement.
- Cross-modality segmentation: 0.88±0.15 DICE accompanied the 55.78±19.20 winning score for liver cross-modality segmentation, with other measures lowering the overall grade.The winning score was below the method’s CT score of 61.13±19.72 but above its MR score of 41.15±21.61, while PKDIA showed a dramatic performance drop.
7. Discussions and Conclusion
CHAOS evaluated abdominal organ segmentation across five modality and organ settings, finding strong single-modality volumetric performance but substantially greater difficulty for cross-modality tasks. The discussion also identifies clinical integration, model interpretation, dataset limitations, and challenge-design issues as important boundaries.
- Five tasks evaluated single-modality, cross-modality, and multi-modal segmentation using four metrics.
- Task-Based Conclusions: Deep learning methods achieved strong CT liver segmentation, reaching inter-expert variability for DICE and volumetry, while distance-based metrics still required improvement.Top methods were qualitatively judged usable in real-life solutions with little post-processing.
- Task-Based Conclusions: MR liver models performed almost as well as interactive methods for DICE but remained weaker on distance-based measures.Clinical use would require integration into an accessible workstation or DICOM viewer, with minimal post-processing interaction potentially needed.
- Task-Based Conclusions: MRI average MSSD was 61.01 mm for liver versus 44.31 mm for right kidney, 46.57 mm for left kidney, and 44.22 mm for spleen.The authors note that multi-organ performance gains may be slight and may not justify the effort required.
- Task-Based Conclusions: Cross-modality performances for Task 1 and Task 4 were clearly lower than the single-modality results.Multi-organ cross-modality segmentation remained the most challenging setting, with further progress needed for real-world clinical application.
- Participating Models: High score variance across predominantly U-Net-based systems reflects the influence of architecture, implementation, optimization, tuning, and evaluation choices.These interacting factors make it difficult to explain why particular models perform well across heterogeneous teams and environments.
- Multiple Submissions, Peeking and Ensembles: Online challenge submissions can produce misleading performance through peeking, and available precautions do not completely resolve the problem.The authors used source code, method documents, or prior-use references to support verification of online results.