Source-linked AI summary

The Liver Tumor Segmentation Benchmark (LiTS)

Patrick Bilic, Patrick Christ, Hongwei Bran Li, Eugene Vorontsov, Avi Ben-Cohen, Georgios Kaissis, Adi Szeskin, Colin Jacobs, Gabriel Efrain Humpire Mamani, Gabriel Chartrand, Fabian Lohöfer, Julian Walter Holch, Wieland Sommer, Felix Hofmann, Alexandre Hostettler, Naama Lev-Cohain, Michal Drozdzal, Michal Marianne Amitai, Refael Vivantik, Jacob Sosna, Ivan Ezhov, Anjany Sekuboyina, Fernando Navarro, Florian Kofler, Johannes C. Paetzold, Suprosanna Shit, Xiaobin Hu, Jana Lipková, Markus Rempfler, Marie Piraud, Jan Kirschke, Benedikt Wiestler, Zhiheng Zhang, Christian Hülsemeyer, Marcel Beetz, Florian Ettlinger, Michela Antonelli, Woong Bae, Míriam Bellver, Lei Bi, Hao Chen, Grzegorz Chlebus, Erik B. Dam, Qi Dou, Chi-Wing Fu, Bogdan Georgescu, Xavier Giró-i-Nieto, Felix Gruen, Xu Han, Pheng-Ann Heng, Jürgen Hesser, Jan Hendrik Moltz, Christian Igel, Fabian Isensee, Paul Jäger, Fucang Jia, Krishna Chaitanya Kaluva, Mahendra Khened, Ildoo Kim, Jae-Hun Kim, Sungwoong Kim, Simon Kohl, Tomasz Konopczynski, Avinash Kori, Ganapathy Krishnamurthi, Fan Li, Hongchao Li, Junbo Li, Xiaomeng Li, John Lowengrub, Jun Ma, Klaus Maier-Hein, Kevis-Kokitsi Maninis, Hans Meine, Dorit Merhof, Akshay Pai, Mathias Perslev, Jens Petersen, Jordi Pont-Tuset, Jin Qi, Xiaojuan Qi, Oliver Rippel, Karsten Roth, Ignacio Sarasua, Andrea Schenk, Zengming Shen, Jordi Torres, Christian Wachinger, Chunliang Wang, Leon Weninger, Jianrong Wu, Daguang Xu, Xiaoping Yang, Simon Chun-Ho Yu, Yading Yuan, Miao Yu, Liping Zhang, Jorge Cardoso, Spyridon Bakas, Rickmer Braren, Volker Heinemann, Christopher Pal, An Tang, Samuel Kadoury, Luc Soler, Bram van Ginneken, Hayit Greenspan, Leo Joskowicz, Bjoern Menze

arXiv:1901.04056v2cs.CV

TL;DR

Manual delineation of liver lesions in 3D CT scans remains challenging, motivating evaluation of automated liver and tumor segmentation methods. LiTS presents a benchmark setup and reports tumor-segmentation results across events, while noting annotation-related limitations.

  • Problem

    Manual delineation of target lesions in 3D CT scans and fully automated segmentation of the liver and its lesions remain challenging.

  • Method

    LiTS evaluates state-of-the-art automated liver and liver-tumor segmentation methods using reference segmentations, with annotations verified by additional blinded readers.

  • Results

    0.739 Dice was achieved by the top-performing teams in MICCAI 2018, compared with 0.702 in MICCAI 2017 and 0.674 in ISBI 2017.

  • Takeaways & Limitations

    The benchmark evaluates the state of automated liver and liver-tumor segmentation methods across multiple events.

  • Takeaways & Limitations

    Annotations from only one rater at each medical center may introduce label bias, especially for small-lesion segmentation.

Abstract

from arXiv · show

In this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedical Imaging (ISBI) 2017 and the International Conferences on Medical Image Computing and Computer-Assisted Intervention (MICCAI) 2017 and 2018. The image dataset is diverse and contains primary and secondary tumors with varied sizes and appearances with various lesion-to-background levels (hyper-/hypo-dense), created in collaboration with seven hospitals and research institutions. Seventy-five submitted liver and liver tumor segmentation algorithms were trained on a set of 131 computed tomography (CT) volumes and were tested on 70 unseen test images acquired from different patients. We found that not a single algorithm performed best for both liver and liver tumors in the three events. The best liver segmentation algorithm achieved a Dice score of 0.963, whereas, for tumor segmentation, the best algorithms achieved Dices scores of 0.674 (ISBI 2017), 0.702 (MICCAI 2017), and 0.739 (MICCAI 2018). Retrospectively, we performed additional analysis on liver tumor detection and revealed that not all top-performing segmentation algorithms worked well for tumor detection. The best liver tumor detection method achieved a lesion-wise recall of 0.458 (ISBI 2017), 0.515 (MICCAI 2017), and 0.554 (MICCAI 2018), indicating the need for further research. LiTS remains an active benchmark and resource for research, e.g., contributing the liver-related segmentation tasks in \url{http://medicaldecathlon.com/}. In addition, both data and online evaluation are accessible via \url{www.lits-challenge.com}.

1. Introduction

LiTS addresses the need for reproducible automated liver and tumor segmentation by creating a public multicenter dataset and benchmarking state-of-the-art methods across three challenge events.

  • Accurate liver-lesion segmentation supports cancer diagnosis, treatment planning, response monitoring, and lesion localization for several treatment options.
  • Manual delineation of target lesions in 3D CT scans is time-consuming, poorly reproducible, and operator-dependent.
  • Technical challenges: Automated segmentation must handle variable lesion contrast, lesion types and appearances, chronic liver disease, and disease-specific changes in lesion and liver morphology.
  • Contributions: The LiTS challenge evaluated automated liver and liver-tumor segmentation in three events associated with ISBI 2017, MICCAI 2017, and MICCAI 2018.
  • Contributions: The work introduced a public multicenter dataset of 201 abdominal CT volumes with reference segmentations for liver and liver tumors.
  • Contributions: The paper describes the benchmark setup, evaluates and ranks submitted algorithms, analyzes tumor detection, and discusses technical trends, challenges, limitations, and future work.

2. Prior Work: Datasets & Approaches

Prior liver datasets were limited in size, annotation, or lesion representation, motivating broader benchmarking. Existing approaches spanned shape models, atlas and spatial methods, traditional learning, and deep learning, with deep learning increasingly adopted after 2016.

  • Datasets: Available liver datasets often had few images, lacked lesion-focused cohorts, or provided no reference segmentations.
  • Datasets: Earlier challenges included SLIVER07, LTSC’08, ImageCLEF 2015, VISCERAL, and CHAOS, covering liver, tumor, reporting, or healthy-organ tasks.
  • Liver segmentation: Liver segmentation methods were organized around shape and geometric priors, intensity distribution and spatial context, or deep learning.
  • Liver segmentation: 0.49 salience is reported for the finding that statistical shape models achieved the best results in SLIVER07.
  • Liver segmentation: Deep learning methods use data-driven end-to-end optimization, and top-performing methods commonly combine 3D CNN segmentation with post-processing.
  • Liver tumor segmentation: Liver tumors are more challenging because they vary broadly in shape, size, contrast, location, and boundary clarity, with contrast-agent uptake adding variability.
  • Liver tumor segmentation: Tumor segmentation approaches included thresholding and spatial regularization, local features and learning algorithms, and deep learning.
  • Liver tumor segmentation: Spatially constrained and learning-based methods used morphology, shape or surface information, clustering, classifiers, texture features, and supervoxels.

3. Methods

LiTS established a recurring online benchmark using multi-institution abdominal CT data, standardized access and evaluation, and segmentation metrics complemented by lesion-detection metrics. Its dataset spans diverse tumor types and imaging conditions, while the train–test split does not test generalization to unseen centers.

  • Benchmark organization: LiTS was organized across ISBI 2017, MICCAI 2017, and MICCAI 2018, with later editions adding liver segmentation and joining the Medical Segmentation Decathlon.
  • Benchmark organization: The online CodaLab platform distributed annotated training and unannotated test data, automatically computed scores, and hosted the benchmark evaluation.
  • Dataset: 201 CT images were collected from seven clinical sites, with 194 scans containing lesions.
  • Dataset: The cohort included primary and secondary tumors, varied lesion-to-background ratios, pre- and post-therapy scans, different scanners, protocols, artifacts, and image resolutions.
  • Dataset: Tumor counts ranged from 0 to 12, tumor sizes from 38 mm3 to 1231 mm3, and axial slices from 42 to 1026.
  • Dataset limitations: 183 salience is reported for the limitation that the training and test sets had similar center distributions, so generalizability to unseen centers was not tested.
  • Evaluation: The benchmark ranked submissions primarily by per-case averaged Dice for volumetric overlap, while also reporting surface distance and volume similarity.
  • Evaluation: LiTS introduced lesion-detection metrics alongside segmentation evaluation because of the clinical relevance of detecting lesions.

4. Results

Submitted LiTS methods predominantly used automated, multi-stage U-Net-based pipelines with preprocessing, augmentation, varied losses, and post-processing. Across events, architectures increasingly incorporated ensembles, residual connections, hybrid 2D/3D designs, and eventually 3D models.

  • Algorithms and architectures: 73 submissions were fully automated, while one used an interactive approach.
  • Algorithms and architectures: U-Net-derived architectures were overwhelmingly used, commonly in coarse-to-fine cascades for liver and tumor segmentation.
  • Critical components: Common preprocessing included HU-value clipping, normalization, standardization, and geometric data augmentation.
  • Critical components: Training used ADAM or stochastic gradient descent with momentum, alongside cross-entropy, Dice, Jaccard, Tversky, L2, and ensemble losses.
  • Post-processing: Post-processing commonly formed connected tumor components and overlaid liver masks to discard tumors outside the liver.
  • Technique evolution: In MICCAI 2018, 3D deep learning models generally outperformed 2.5D or 2D models without more sophisticated preprocessing.

4.2. Results of inidividual challenges

Liver segmentation achieved high performance across the challenges, whereas tumor segmentation and lesion detection remained more difficult. Tumor metrics and detection recall improved across events, but ranking depended on the task and metric.

  • Liver segmentation: 0.963 was the best liver segmentation Dice score, with almost all methods exceeding 0.920 per case.
  • ISBI 2017: 0.674 was the winning liver tumor segmentation Dice in ISBI 2017, followed by 0.652 and 0.645.
  • Tumor detection: Small-lesion detection remained difficult, with top-performing teams achieving only around 0.10 F1 score.
  • MICCAI 2017: 0.702 was MICCAI 2017’s highest average tumor Dice per case, compared with 0.674 on the same ISBI test set.
  • MICCAI 2018: 0.739 and 0.721 were the top MICCAI 2018 tumor Dice scores, compared with 0.702 previously.
  • Tumor detection: 0.554 was the best MICCAI 2018 lesion recall, improving over 0.479 in MICCAI 2017 and 0.458 in ISBI 2017.

4.3. Meta Analysis

The meta-analysis examined performance trends, annotation variability, and relationships between segmentation and detection. It found improving benchmark performance, substantial annotation differences, and incomplete alignment between segmentation quality and tumor detection.

  • Segmentation and detection: Top-performing methods did not consistently achieve good tumor-detection scores across the three LiTS challenges.
  • Inter-rater agreement: 70.2% was the median Dice agreement between new and existing consensus annotations, versus 95.2% for a board-certified radiologist and existing annotations.
  • Inter-rater agreement: The models were optimized on R1, while the best leaderboard model achieved 82.5%, leaving room for improvement.
  • Performance over time: MSD’18 showed significant improvement over ISBI’17 on both evaluated metrics.
  • Performance over time: 2022 submissions achieved significantly better scores than 2021 submissions, indicating that LiTS remained active and contributed to methodology development.

4.4. Technique trend and recent advances

LiTS supported continuing method development, with a major shift toward 3D deep learning and evidence of improving results over time. Performance remained sensitive to lesion size and tumor–liver contrast, especially for small or low-contrast tumors.

  • Recent advances: The released LiTS dataset contributed to novel methodology development in medical image segmentation.
  • Lesion size: State-of-the-art methods performed well on large tumors but struggled with smaller tumors, especially single tumors below 10 mm^3.
  • Image contrast: Methods performed best when tumor–liver contrast was high, particularly when focal lesions were 40–60 HU above background liver.
  • Image contrast: The worst results occurred when tumor–liver contrast was below 20 HU, including tumors less dense than the liver.

5. Discussion

The discussion identifies annotation, metric, imaging, and benchmark-design limitations that affect LiTS evaluation and interpretation. It also highlights tumor-size difficulty and the dataset’s broader research value.

  • Limitations: Single-rater annotations may introduce label bias, especially for small-lesion segmentation; consensus annotations could reduce label noise.
  • Evaluation: Dice alone may not distinguish top-performing teams because large tissue dominates; combining multiple metrics provides better discrimination.
  • Data limitations: Scanner and demographic information were unavailable across centers, although the authors state these data are essential for in-depth analysis.
  • Challenge design: Public test-data release enables rapid evaluation but cannot prevent overfitting or cheating through iterative submissions and manual correction.
  • Tumor segmentation: Low tumor-to-liver HU differences challenge segmentation, while small, ambiguous lesions remain particularly difficult to delineate.
  • Research value: LiTS supports benchmark methodology, domain adaptation, federated-learning studies, shape modeling, and other research beyond segmentation evaluation.

xxxii

This section lists prior work and related resources spanning medical-image segmentation, tumor analysis, deep learning, and evaluation criteria.

  • Clinical assessment: Related work also covers tumor burden, response evaluation, volumetry, and automated liver-tumor analysis from CT images.
  • Segmentation methods: Prior studies address liver and tumor segmentation using shape models, active contours, registration, graph cuts, and learning-based methods.
  • Deep learning: Deep-learning references include convolutional networks, dense U-net architectures, self-supervised learning, and multi-organ segmentation.

CRediT Author Statement

The CRediT statement assigns contributors responsibilities across conceptualization, writing, data curation, visualization, software, investigation, validation, supervision, and correspondence.

  • Contributions: Contributors are credited for conceptualization, writing and revision, data curation, visualization, software, validation, investigation, supervision, and correspondence.
  • Contributions: The author statement distributes these roles across a large group of named contributors.
  • Contributions: Several contributors are specifically credited with writing, review and editing, data curation, visualization, and investigation.

Appendix A. Segmentation performance w.r.t tumor size and number of tumors.

Segmentation performance varies with tumor burden: methods perform best on volumes containing larger tumors and worst when tumors are small and isolated.

  • Large tumors are associated with better participating-method performance, whereas small tumors produce worse results.
  • Volumes containing a single small tumor below 15mm3 yield the worst results.
  • The best results occur when volumes contain fewer than 6 tumors with overall tumor volume above 40mm3.

Appendix B. Segmentation performance w.r.t. HU value differences.

Segmentation performance varies with the HU contrast between tumors and surrounding liver tissue, with higher contrast associated with better results and low contrast producing the worst outcomes.

  • Higher tumor–liver HU contrast is associated with better performance across participating methods.The analysis clusters test volumes by HU differences between tumor and non-tumor liver tissue.
  • Below 20 HU contrast produces the worst results, including tumors with lower HU values than the liver.

Appendix C. Correspondence algorithm

The correspondence algorithm maps connected components between reference and prediction masks despite split and merge errors, then assigns metrics while preserving the reference components’ identity.

  • Error types: Split errors occur when one reference component is predicted as multiple components, while merge errors occur when multiple reference components share one prediction.
  • Reference mapping: The algorithm first converts many-to-many overlap into a many-to-one mapping by merging reference components linked to the same predicted component.
  • Unmatched components: Unmatched reference components are labeled false negatives, whereas unmatched predicted components are labeled false positives.
  • Prediction mapping: Remaining predicted components are associated with the merged reference region having the largest total intersected area, and components sharing that region are merged.
  • Metric attribution: Metrics computed on merged reference components are attributed to each constituent component to avoid exaggerating errors from separate evaluation.For example, a combined Dice score of 0.7 is assigned to each constituent reference component.

Appendix D. Automated tumor burden analysis in MICCAI-LiTS 2017

Automated liver and tumor segmentation enables volumetric tumor-burden analysis relevant to disease progression, treatment assessment, and surgical planning. In MICCAI-LiTS 2017, tumor burden was generally predicted accurately, with low RMSE among the best methods.

  • Clinical relevance: Fully volumetric liver and tumor segmentation supports tumor-burden computation and can simplify surgical liver resection planning.
  • Definition and motivation: Tumor burden is defined as the fraction of the liver afflicted by cancer, or the liver/tumor ratio.
  • Results: The best-performing method achieved an RMSE of 0.015 and a maximum error of 0.033 for tumor-burden estimates.
  • Results: Methods with high Dice per-case scores generally also obtained lower RMSE values, although RMSE and maximum-error rankings showed only slight correlation.
Loading 1901.04056v2…