Source-linked AI summary
Annotation-efficient deep learning for automatic medical image segmentation
Shanshan Wang, Cheng Li, Rongpin Wang, Zaiyi Liu, Meiyun Wang, Hongna Tan, Yaping Wu, Xinfeng Liu, Hui Sun, Rui Yang, Xin Liu, Jie Chen, Huihui Zhou, Ismail Ben Ayed, Hairong Zheng
TL;DR
Medical image segmentation often depends on large, high-quality annotations that are difficult to obtain. The paper introduces AIDE, an annotation-efficient framework for imperfect datasets, and reports comparable segmentation using only 10% of training annotations, including in breast tumor data.
Problem
The lack of large and high-quality labeled datasets limits supervised deep learning for medical imaging.
Method
AIDE uses cross-model co-optimization and self-correction to learn from imperfect datasets across segmentation settings.
Results
Using 10% of training annotations, AIDE achieves segmentation results comparable to fully supervised models and independent radiologists.
Takeaways & Limitations
AIDE can save almost 90% of manual annotation effort compared with conventional deep learning training.
Takeaways & Limitations
The study focuses on medical image segmentation, and AIDE’s distance metrics were not largely improved in one domain-adaptation setting.
Abstract
from arXiv · showhide
Automatic medical image segmentation plays a critical role in scientific research and medical care. Existing high-performance deep learning methods typically rely on large training datasets with high-quality manual annotations, which are difficult to obtain in many clinical applications. Here, we introduce Annotation-effIcient Deep lEarning (AIDE), an open-source framework to handle imperfect training datasets. Methodological analyses and empirical evaluations are conducted, and we demonstrate that AIDE surpasses conventional fully-supervised models by presenting better performance on open datasets possessing scarce or noisy annotations. We further test AIDE in a real-life case study for breast tumor segmentation. Three datasets containing 11,852 breast images from three medical centers are employed, and AIDE, utilizing 10% training annotations, consistently produces segmentation maps comparable to those generated by fully-supervised counterparts or provided by independent radiologists. The 10-fold enhanced efficiency in utilizing expert labels has the potential to promote a wide range of biomedical applications.
Introduction
Medical image segmentation supports research and clinical care, but supervised deep learning is constrained by scarce, costly, time-intensive, and noisy annotations. AIDE is introduced to address limited, absent target-domain, and noisy labels through annotation-efficient learning.
- Medical image annotation can require minutes to hours per image and is time-consuming, labor-intensive, and expensive.
- Real-world annotations inevitably contain noise from annotator errors and inter-annotator variation, while biases can transfer to learned models.
- Large, high-quality labeled datasets are the primary limitation of supervised deep learning for medical imaging.
- SSL, UDA, and NLL represent common clinical challenges involving limited annotations, missing target-domain labels, or noisy annotations.
- AIDE targets all three challenges by transforming SSL and UDA into noisy-label learning and applying cross-model self-correction.
- Experiments on public datasets and breast tumor data evaluate AIDE under imperfect annotations, including 11,852 images from three medical centers.
Results
AIDE handles imperfect datasets by using a unified framework for SSL, UDA, and NLL. Its training combines cross-model optimization, local filtering, global label correction, and consistency-based learning.
- AIDE addresses SSL, UDA, and NLL within one framework by optimizing models with problematic labels.
- For SSL and UDA, low-quality labels are generated for unlabeled data using models trained on limited labeled or source-domain data.
- Two networks exchange information during parallel training to support cross-model co-optimization.
- Suspected noisy samples are filtered locally, augmented, and assigned pseudo-labels distilled from augmented-input predictions.
- After each epoch, labels with low similarity to network predictions can be updated through a global correction criterion.
- Segmentation performance is characterized using DSC, RAVD, ASSD, and MSSD, with higher DSC and lower distance or difference values indicating greater accuracy.
Enhancement of AIDE compared to conventional fully-supervised learning
AIDE is designed to exploit imperfect labels rather than discard them, using consistency constraints and self-correction to improve learning. Across scarce and noisy-label settings, it achieves strong segmentation performance with substantially fewer annotations.
- AIDE uses transformation-consistent predictions and progressive label correction to reduce the effects of low-quality annotations.
- AIDE exploits suspected noisy samples through correction instead of dropping them, preserving useful information from imperfect inputs.
- Cross-model co-optimization reduces the risk of error propagation and overfitting to a network’s own pseudo-labels.
- Limited annotations: Increasing the number of unlabeled training samples continued to improve AIDE performance.
- Noisy label learning: 97% noise with 30 labeled and 954 unlabeled samples produced results comparable to training with 331 high-quality annotated samples.
- Limited annotations: 83.1% average DSC was achieved on CHAOS test data using 1 labeled case and 9 unlabeled cases.
Unsupervised domain adaptation with large domain discrepancies
AIDE improves segmentation across domain adaptation and multi-annotator settings without target-domain or multiple high-quality annotations. Benefits are strongest for difficult small-object tasks, while some distance metrics and highly diverse domains remain challenging.
- Domain adaptation: AIDE improves DSC and decreases RAVD when adapting segmentation models across domains without target-domain high-quality annotations.
- Domain adaptation: For transfer from Domain 1 to Domain 2, AIDE increased DSC from 45.8% to 80.0%.
- Domain adaptation: Distance metrics were not largely improved, and performance on the combined Domain 3 dataset was worse than on Domains 1 and 2.
- Domain adaptation: AIDE achieved better performance than pseudo-label and co-teaching methods in the evaluated UDA experiments.
- Multiple annotators: AIDE’s improvements were minor for relatively large regions but substantial for challenging small-object segmentation tasks.
- Multiple annotators: Using one annotator’s labels, AIDE achieved average DSCs of 93.2%, 86.6%, 93.7%, and 89.8% across four tasks.
- Multiple annotators: AIDE may relax the demand for multiple annotators by self-adjusting labels and correcting possible observer biases.
Real-life case study for breast tumor segmentation
AIDE was evaluated for breast tumor segmentation on clinical DCE-MR data from three medical centers, using substantially fewer training annotations than fully supervised models. With limited labels, it achieved performance comparable to fully supervised models and independent radiologists, while improving over respective baselines.
- Study design: Three datasets from different medical centers were used to evaluate AIDE on breast tumor segmentation in clinical DCE-MR images.The datasets included GGH, GPPH, and HPPH cases.
- Comparison with baselines: 7.7% absolute DSC improvement was observed for GGH, while GPPH showed an 18.2% absolute increase over its respective baseline.These gains were reported under the same experimental settings.
- Annotation efficiency: AIDE used 10% of the annotations used by fully supervised methods for GGH and GPPH, and 9.2% for HPPH.For HPPH, this corresponded to 25 training cases.
- Comparison with fully supervised models: AIDE achieved average DSC values of 0.690 ± 0.251, 0.654 ± 0.221, and 0.731 ± 0.196 for GGH, GPPH, and HPPH, respectively.The corresponding fully supervised values were 0.722 ± 0.208, 0.678 ± 0.260, and 0.738 ± 0.227, with no significant differences reported for these comparisons.
- Comparison with radiologists: AIDE produced comparable or better segmentation than independent radiologists across the three datasets.For GGH, AIDE exceeded manual annotations; for GPPH and HPPH, no significant differences were observed.
- Clinical scope and limitations: The breast tumor task remains difficult because tumors occupy small regions and background signals are confounded by organs and dense glandular tissue.The reported DSC values were lower than those for some other segmentation tasks, and AIDE is intended to assist rather than replace radiologists.
Discussion
AIDE is designed to address imperfect training datasets across semi-supervised learning, unsupervised domain adaptation, and noisy-label learning. Experiments report better performance than conventional fully-supervised models under matched conditions, comparable results with perfect datasets, and satisfactory breast-tumor segmentation using 10% of training annotations.
- Open-dataset evaluation: AIDE performs substantially better than conventional fully-supervised models under the same training conditions and remains comparable to models trained on corresponding perfect datasets.These findings were reported from experiments with open datasets containing imperfect annotations.
- Clinical evaluation: AIDE generated satisfactory segmentation results on three independent clinical breast datasets using only 10% of the training annotations.The clinical experiments used data from three medical centers and were presented as evidence of robustness and generalization to real clinical data.
- Noisy-label learning: AIDE combines selective, orderly label updating with filtered-label segmentation and consistency loss for suspected noisy data.The consistency-loss contribution increases during training, while local filtering and global correction progressively emphasize pseudo-labels.
- Scope and outlook: The current work addresses medical image segmentation, while applicability to other medical image-analysis tasks remains for future evaluation.The authors specifically mention image classification as an example of a possible extension.
- Framework scope: AIDE targets three imperfect-data settings: semi-supervised learning, unsupervised domain adaptation, and noisy-label learning.The framework is intended to work with imperfect training datasets rather than build a more sophisticated fully-supervised model.
- Annotation efficiency: In the breast-tumor case study, AIDE saved almost 90% of the manual effort required to annotate training data compared with conventional DNN training.The paper describes this as a 10-fold improvement in efficiency in utilizing expert labels.
Methods
The experiments evaluate AIDE across semi-supervised, domain-adaptation, and noisy-label settings using public medical-imaging datasets. Dataset construction varies by task, with limited labels, multiple acquisition domains, unlabeled target data, or multiple expert annotations used to represent imperfect supervision.
- Semi-supervised learning: The CHAOS abdomen-MR setting uses one labeled case with 30 labeled image samples, pseudo-labels for 29 remaining cases, and 10 labeled cases for testing.The pseudo-labels are generated by a network trained on the labeled samples.
- Unsupervised domain adaptation: Unsupervised domain adaptation combines labeled source-domain data with unlabeled target-domain data and evaluates performance on target-domain testing data.Pseudo-labels are generated for target-domain training data, which is combined with labeled source data and noisy target labels.
- Noisy-label learning: The QUBIQ experiments use four datasets and four tasks, treating each expert annotation set as a noisy label set for model training.The datasets cover prostate, brain growth, brain tumor, and kidney segmentation, with multiple annotations per case.
- Evaluation: QUBIQ model performance is measured by thresholding continuous predictions and averaged-expert ground-truth labels at probability levels from 0.1 to 0.9, then averaging DSCs.The final metrics are obtained by averaging the DSC values across thresholds.
Clinical datasets
The study assembles multi-center breast MRI datasets and develops AIDE to learn from limited or noisy annotations through filtering, correction, and self-correction.
- Clinical datasets: Experienced radiologists independently delineated tumors, with a third radiologist reviewing and finalizing the annotations.Three radiologists with more than 10 years of experience provided the image annotations.
- Clinical datasets: 11,852 image samples from three medical centers comprise the breast MRI experiments.The datasets include 300 GGH images, 4,902 GPPH images, and 6,650 HPPH images.
- AIDE framework: AIDE standardizes semi-supervised learning and unsupervised domain adaptation as noisy-label learning.Limited annotations support pseudo-label generation for unlabeled data before noisy-label learning is applied.
- AIDE framework: Local label filtering combines segmentation and consistency losses for samples suspected of having low-quality labels.Pseudo-labels are formed from augmented inputs after temperature sharpening, and consistency is measured with MSE.
- AIDE framework: Global label correction updates labels for 25% of training samples with the smallest DSCs when the update criteria are met.Updates occur before the warm-up threshold and every 10 epochs thereafter.
- AIDE framework: AIDE's filtering and correction schedule is motivated by the observed pattern that networks learn low-noise samples before memorizing noisy labels.The study reports a similar memorization pattern in medical image segmentation.
- Scope: All experiments use 2D processing; extending the approach to 3D requires considerably more effort to construct suitable implementation datasets.The 3D extension is outside the study's scope.
Evaluation metrics
Segmentation performance is evaluated with overlap, area or volume, and surface-distance metrics, with statistical comparisons assessed using paired tests.
- Metrics: Dice score, RAVD, ASSD, and MSSD quantify segmentation overlap, area or volume difference, and surface distances.DSC is the Dice similarity coefficient; ASSD and MSSD measure average and maximum symmetric surface distance.
- Metrics: ASSD uses boundary points from predicted and reference segmentations to quantify symmetric surface distance.The equation compares nearest boundary-point distances in both directions.
- Metrics: DSC is computed from true-positive, false-positive, and false-negative predictions.These quantities represent correct positive, incorrect positive, and missed positive predictions, respectively.
- Statistics: Differences between experiments and model results versus human annotations are tested with two-sided paired t-tests using P≤0.05.The statistical comparison is paired across independent cases.
Statistics and reproducibility
The study supports reproducibility by releasing code and models and repeating experiments across datasets, tasks, and hospital cohorts.
- Reproducibility: The training code and models are publicly available for reproducibility.The implementation uses PyTorch and publicly available packages.
- Reproducibility: Experiments were repeated three times for CHAOS and the breast datasets, six times for prostate domain adaptation, and multiple times across QUBIQ subtasks.The four QUBIQ subtasks were repeated 6, 7, 3, and 3 times, respectively.
Data Availability
Open datasets are available through challenge websites, while clinical breast data are de-identified but restricted by patient-privacy considerations.
- Data access: Raw data from CHAOS, PROMISE12, and QUBIQ are accessible through their respective official challenge websites.Access follows standard procedures.
- Data access: Clinical breast data are de-identified and not publicly available because of patient-privacy considerations.Academic-use requests can be submitted to the corresponding authors for review.
- Figure legends: Figure mappings distinguish training from testing domains and compare conventional optimization with AIDE-related settings.Vertical axes indicate training datasets and horizontal axes indicate testing datasets.
- Figure legends: Figure examples compare high-quality, low-quality, and self-corrected labels, along with segmentation results across annotation settings.Displayed DSC values and box plots summarize the comparisons, with significance marked by asterisks.
- Figure legends: LQA and HQA denote low- and high-quality annotations, while labels such as LQA200 specify the high-quality and low-quality data composition.LQA200 means 20 high-quality labeled and 180 low-quality labeled data are used.
- Figure legends: Breast-dataset visualizations compare LQA, LQA_Ours, and independent radiologist contours against high-quality annotations.The figure uses red, magenta, green, and yellow contours for the respective annotation sources or results.
Tables
The tables organize segmentation evaluations across SSL settings, domain shifts, and four QUBIQ tasks, with AIDE compared against conventional fully supervised approaches. Results indicate stronger gains on more challenging tasks and conditions.
- Table 1 reports segmentation results under different SSL settings.
- Table 2 compares networks trained and tested on prostate datasets from different domains.
- Tables 3 and 4 report DSC (%) for prostate segmentation and brain growth segmentation, respectively.
- Table 5 reports DSC (%) for brain tumor and kidney segmentation tasks.