Source-linked AI summary

AbdomenCT-1K: Is Abdominal Organ Segmentation A Solved Problem?

Jun Ma, Yao Zhang, Song Gu, Cheng Zhu, Cheng Ge, Yichi Zhang, Xingle An, Congcong Wang, Qiyuan Wang, Xin Liu, Shucheng Cao, Qi Zhang, Shangqing Liu, Yunpeng Wang, Yuhui Li, Jian He, Xiaoping Yang

arXiv:2010.14808v2cs.CV

TL;DR

Existing abdominal organ segmentation results may not generalize beyond narrow benchmark distributions. The paper introduces AbdomenCT-1K, evaluates SOTA methods across diverse conditions, and establishes four benchmarks to advance practical segmentation research.

  • Problem

    Existing datasets often lack diverse centers, phases, vendors, diseases, and comprehensive evaluation of generalization and annotation-efficient segmentation tasks.

  • Method

    The paper constructs AbdomenCT-1K, evaluates nnU-Net with DSC and NSD, and develops fully supervised, semi-supervised, weakly supervised, and continual learning benchmarks.

  • Results

    SOTA performance does not generalize well to unseen datasets containing new CT phases, medical centers, and challenging unseen diseases.

  • Takeaways & Limitations

    AbdomenCT-1K and its benchmarks provide a platform for studying more challenging and practical abdominal organ segmentation problems.

  • Takeaways & Limitations

    Existing SOTA methods perform strongly when test data resemble training data and DSC is used, but remain limited on diverse challenging clinical cases.

Abstract

from arXiv · show

With the unprecedented developments in deep learning, automatic segmentation of main abdominal organs seems to be a solved problem as state-of-the-art (SOTA) methods have achieved comparable results with inter-rater variability on many benchmark datasets. However, most of the existing abdominal datasets only contain single-center, single-phase, single-vendor, or single-disease cases, and it is unclear whether the excellent performance can generalize on diverse datasets. This paper presents a large and diverse abdominal CT organ segmentation dataset, termed AbdomenCT-1K, with more than 1000 (1K) CT scans from 12 medical centers, including multi-phase, multi-vendor, and multi-disease cases. Furthermore, we conduct a large-scale study for liver, kidney, spleen, and pancreas segmentation and reveal the unsolved segmentation problems of the SOTA methods, such as the limited generalization ability on distinct medical centers, phases, and unseen diseases. To advance the unsolved problems, we further build four organ segmentation benchmarks for fully supervised, semi-supervised, weakly supervised, and continual learning, which are currently challenging and active research topics. Accordingly, we develop a simple and effective method for each benchmark, which can be used as out-of-the-box methods and strong baselines. We believe the AbdomenCT-1K dataset will promote future in-depth research towards clinical applicable abdominal organ segmentation methods. The datasets, codes, and trained models are publicly available at https://github.com/JunMa11/AbdomenCT-1K.

1 INTRODUCTION

Abdominal organ segmentation has advanced substantially, but its clinical robustness remains uncertain because existing evaluations often lack diversity and comprehensive testing. AbdomenCT-1K addresses these gaps with a diverse dataset, a large-scale SOTA study, and four challenging benchmarks.

  • Abdominal CT organ segmentation is difficult because soft-tissue contrast is low, organs have complex structures and lesions, and scanners or phases change organ appearance.
  • The paper introduces benchmarks for fully supervised, semi-supervised, weakly supervised, and continual learning to target challenging practical problems.
  • Existing methods and benchmarks lack large-scale diverse data, comprehensive cross-center evaluation, annotation-efficient tasks, and boundary-based metrics.
  • 1112 CT scans from 12 medical centers form AbdomenCT-1K, covering multi-center, multi-phase, multi-vendor, and multi-disease cases with annotations for four organs.
  • The study evaluates nnU-Net for single- and multi-organ segmentation using DSC and NSD, adding boundary accuracy to the assessment.
  • SOTA segmentation succeeds in some ideal or easy settings but remains unsolved for new medical centers and unseen abdominal cancer cases.

2 RELATED WORK

Related work spans model-based and learning-based segmentation methods, annotation-efficient learning, and increasingly diverse abdominal CT benchmarks. The review highlights persistent weaknesses in weak-boundary segmentation, catastrophic forgetting, annotation burden, and dataset scale or diversity.

  • Abdominal organ segmentation methods comprise classical model-based approaches and modern learning-based approaches.
  • Model-based methods use energy minimization, shape templates, or atlases but often fail on weak boundaries and low contrasts while requiring high computational cost for 3D CT.
  • Deep CNN methods learn discriminative features from annotated CT scans and have achieved SOTA performance without handcrafted features or anatomical correspondences.
  • Supervised work includes single-organ and multi-organ segmentation, with pancreas segmentation treated as more challenging and often addressed using cascaded approaches.
  • Semi-supervised, weakly supervised, and continual learning reduce reliance on full annotations, while continual learning must address catastrophic forgetting or interference.
  • Existing benchmarks vary in scale, organs, modalities, centers, and annotation scope, including BTCV, NIH Pancreas, VISCERAL Anatomy, and CT-ORG.

3 ABDOMENCT-1K DATASET

AbdomenCT-1K expands abdominal CT segmentation data to 1112 scans with multi-center, multi-phase, multi-vendor, and multi-disease diversity, while adding four-organ annotations and quality-controlled labeling. The study evaluates segmentation with complementary region- and boundary-based metrics.

  • Dataset composition: The dataset includes multi-center, multi-phase, multi-vendor, and multi-disease cases, addressing diversity limitations in existing abdominal segmentation datasets.It was collected from 12 medical centers and annotates liver, kidney, spleen, and pancreas for all cases.
  • Dataset composition: AbdomenCT-1K contains 1112 3D CT scans from five existing datasets and a new Nanjing University dataset, covering diverse abdominal cases.The Nanjing dataset includes patients with pancreas, colon, and liver cancer.
  • Annotations: Unlike the original single-organ datasets, AbdomenCT-1K provides annotations for four organs across every case.These multi-organ versions are termed plus datasets, such as LiTS Plus.
  • Annotations: Annotations combine existing labels with inferred masks, manual refinement by 15 junior annotators, and verification by a senior radiologist.Annotation was performed on axial images under radiologist supervision.
  • Annotation quality control: The annotation workflow reduces variability through shared protocols, senior-radiologist review, and double-checking cases with low five-fold cross-validation DSC or NSD scores.Obvious label errors are corrected, and low-scoring cases are rechecked.
  • Evaluation metrics: DSC measures region overlap, whereas NSD measures surface proximity at a specified tolerance; this study uses a tolerance of 1mm.NSD can expose boundary errors that DSC may not reflect and ignores small boundary deviations.

4 A LARGE-SCALE STUDY ON FULLY SUPERVISED

The fully supervised study tests whether strong single-organ methods generalize beyond same-center, same-distribution evaluation. Performance remains strong for liver but degrades substantially for kidney, spleen, and pancreas across new datasets and conditions.

  • Study design: The evaluation compares same-source testing with three testing datasets from new medical centers to assess cross-dataset generalization.The benchmark context is dominated by single-center datasets and high same-distribution SOTA performance.
  • Single-organ results: 94.9% to 96.5% DSC is achieved for liver segmentation on three new testing datasets, but KiTS performance is 2.5% lower than on LiTS.The reported drop is attributed mainly to arterial-phase scans in KiTS versus predominantly portal-phase scans in LiTS.
  • Single-organ results: Up to 15% DSC and 19% NSD declines occur for kidney segmentation on datasets outside KiTS.The largest declines occur for LiTS and Pancreas datasets, where CT phases differ from KiTS.
  • Single-organ results: 10.6% DSC and 17.9% NSD drops are observed for spleen segmentation on KiTS testing data, indicating poor generalization across CT phases.
  • Single-organ results: Pancreas segmentation shows improved NSD by 9.3% and 11% on LiTS and Spleen datasets, where most cases have healthy pancreases.The study reports better generalization on healthy pancreas cases than on pathology cases.
  • Overall finding: High same-distribution performance degrades when single-organ models are tested on data from new medical centers.

4.2 Multi-organ segmentation

Cross-center multi-organ evaluation reveals stable region overlap for liver and spleen but unstable boundary accuracy, large kidney variation, and consistently difficult pancreas segmentation. Challenging cases are associated with lesions, noise, and degraded image quality.

  • Experimental setup: The multi-organ experiments train nnU-Net on one dataset and test it on three datasets from different medical centers.Four experiment groups evaluate cross-dataset generalization with four organ annotations.
  • Liver and spleen: Liver and spleen DSC scores remain above 90%, while their NSD scores fluctuate from 77.4% to 92.1% and 86.0% to 97.0%, respectively.The contrast between stable DSC and variable NSD highlights differing region and boundary behavior.
  • Kidney: Kidney performance varies by more than 10% across testing sets, including 96.0% versus 85.6% DSC and 92.4% versus 78.9% NSD.These values compare Pancreas Plus and Spleen Plus results in the first group of experiments.
  • Pancreas: Pancreas segmentation remains lower than the other organs across all multi-organ experiments, marking it as a continuing challenge.
  • Challenging cases: Well-segmented cases have clear boundaries, good contrast, and no severe artifacts or lesions, whereas challenging cases commonly contain heterogeneous lesions and image-quality degradation.Examples include liver and pancreas lesions, noise, and other degradations described for Figure 7.

4.3 Is abdominal organ segmentation a solved problem?

Abdominal organ segmentation appears solved only under region-based evaluation, matched train-test distributions, and trivial cases for several organs. The paper argues it remains unsolved for boundary accuracy, cross-center or cross-phase generalization, and unseen or severe diseases, motivating broader benchmarks.

  • Conditions for apparent solution: Liver, kidney, and spleen segmentation appears solved under DSC evaluation when testing data match training distributions and cases are not severely diseased or low quality.
  • Unsolved settings: Segmentation remains unsolved when NSD evaluates organ boundaries or when testing data come from new medical centers with different distributions.
  • Benchmark scope: The new benchmarks target region and boundary errors, cross-center and cross-phase generalization, and cases with unseen or severe diseases.Boundary errors are emphasized because they matter in clinical applications such as surgical planning for transplantation.
  • Benchmark scope: The paper introduces abdominal organ benchmarks for semi-supervised, weakly supervised, and continual learning in addition to fully supervised segmentation.These settings address active research topics and can reduce dependence on annotations.
  • Benchmark design: Each new benchmark uses 50 challenging cases and 50 random cases for testing, limiting inference cost and reducing performance bias.

5.1 Fully supervised abdominal organ segmentation benchmark

The fully supervised benchmark tests multi-organ segmentation under challenging cases and contrast-phase shifts. Its 3D nnU-Net baseline performs better when training includes shared phases, but difficult lesions and boundaries remain unresolved.

  • The benchmark targets liver, kidney, spleen, and pancreas segmentation while addressing unsolved problems identified in the large-scale study.
  • The base training set uses MSD Pan. Plus (281), with additional datasets forming two subtasks that test generalization across cases and phases.
  • The final testing set contains 100 non-overlapping cases, including 50 challenging cases selected by low average DSC and NSD and 50 randomly selected cases.
  • Shared contrast phases improve performance: all organs score lower in subtask 1 than in subtask 2 using the 3D nnU-Net baseline.
  • Liver DSC exceeds 95% in both subtasks, while NSD is 83% and 85.8%, indicating remaining boundary errors despite strong region overlap.
  • Challenging fatty-liver, kidney-tumor, and spleen-tumor cases produce incomplete, under-segmented, or incorrect organ predictions that current benchmarks do not highlight.

5.2 Semi-supervised organ segmentation benchmark

The semi-supervised benchmark evaluates whether unlabelled CT data can reduce annotation demand for abdominal organ segmentation. Its teacher-student baseline improves as unlabelled data is added and can reduce pathology-related errors.

  • The benchmark addresses the absence of a medical-image segmentation benchmark for using unlabelled data to reduce manual annotation demand.
  • Training uses 41 labelled Spleen Plus cases alongside progressively incorporated unlabelled cases, with lower- and upper-bound fully supervised comparisons.
  • The teacher-student method trains a teacher on labelled data, generates pseudo labels, trains and fine-tunes a student, then iterates with the student as teacher.
  • Adding unlabelled data progressively increases average DSC and NSD for multi-organ segmentation.
  • Unlabelled data reduces misclassification in challenging kidney-tumor, cholangiectasis, and liver-spleen appearance cases, with errors gradually corrected as more data is used.

5.3 Weakly supervised abdominal organ segmentation benchmark

The weakly supervised benchmark studies organ segmentation from sparse slice-level annotations at 5%, 15%, and 30% rates. Performance improves with more annotations, but gains become less linear and CRF refinement provides little improvement.

  • The benchmark uses weak annotations to generate full segmentation results, providing sparse labels on only part of the training slices.
  • Spleen Plus (41) is selected as training data to reflect settings with limited well-annotated cases in many medical centers.
  • Three subtasks use roughly uniform slice annotations at 5%, 15%, and 30% rates, with 100-case testing sets selected from inferred remaining cases.
  • The baseline combines 2D nnUNet with fully connected CRF refinement using probability maps and CT-attenuation-based Gaussian pairwise potentials.
  • More annotations improve performance; with 15% annotations, liver average DSC exceeds 90%, while gains from 15% to 30% are smaller than gains from 5% to 15%.
  • CRF refinement does not produce remarkable performance improvements, leaving energy-based refinement as an open challenge for inaccurate CNN segmentations.

5.4 Continual learning benchmark for abdominal organ segmentation

The continual-learning benchmark builds a multi-organ model from sequential single-organ datasets without revisiting previous task data. Its simple pseudo-labeling baseline underperforms full-annotation training and still forgets earlier tasks.

  • The benchmark addresses the lack of a public continual-learning benchmark for medical image segmentation and supplies a baseline solution.
  • Training uses sequential single-organ datasets, requires multi-organ prediction, and prevents access to previous task datasets when switching tasks.
  • The baseline progressively expands nnU-Net from liver to kidney, spleen, and pancreas by generating pseudo labels for previously learned organs.
  • Performance with single-organ datasets is lower than full-annotation training, indicating that the model still forgets part of previous tasks when learning new ones.

5.5 Evaluation and comparison on the common testing set

The four benchmarks use different testing sets, so a shared 50-case NJU dataset enables direct comparison. Fully supervised learning performs best for three organs, while semi-supervised learning approaches it overall with far fewer labeled cases and weak supervision minimizes annotation burden.

  • The four benchmarks use different testing sets, so the 50-case NJU dataset provides an apple-to-apple comparison.
  • Fully supervised learning achieves the best average DSC and NSD scores for kidney, spleen, and pancreas because it uses many labelled cases.
  • With 41 labelled and 800 unlabelled cases, semi-supervised learning achieves the best liver performance and nearly matches fully supervised overall performance using 361 labelled cases.
  • Weakly supervised methods achieve the lowest performance but require the least annotation burden.

6 CONCLUSION

The study concludes that strong benchmark performance does not guarantee generalization to diverse clinical data, motivating benchmarks that test scanner, center, disease, and boundary-related challenges. The dataset is broad but remains limited to four large organs, while the authors provide baseline methods to support further progress.

  • 6 CONCLUSION: SOTA methods perform well when testing data resemble training data and evaluation uses DSC, but fail to generalize reliably to challenging unseen datasets.
  • 6 CONCLUSION: The four new benchmarks include fully supervised, semi-supervised, weakly supervised, and continual learning settings.
  • 6 CONCLUSION: Testing across distinct scanners and medical centers, including unseen or rare diseases such as huge tumors, targets clinically challenging generalization.
  • 6 CONCLUSION: The benchmarks emphasize NSD alongside DSC because boundary errors matter in preoperative planning for tumor resections and organ transplantation.
  • 6 CONCLUSION: The dataset primarily covers four large abdominal organs, although 50 cases include eight extra organs and 663 cases include pseudo tumor labels.
  • 6 CONCLUSION: The dataset and out-of-the-box baseline methods are intended to help move abdominal organ segmentation toward real clinical practice.
Loading 2010.14808v2…