Source-linked AI summary

The KiTS21 Challenge: Automatic segmentation of kidneys, renal tumors, and renal cysts in corticomedullary-phase CT

Nicholas Heller, Fabian Isensee, Dasha Trofimova, Resha Tejpaul, Zhongchen Zhao, Huai Chen, Lisheng Wang, Alex Golts, Daniel Khapun, Daniel Shats, Yoel Shoshan, Flora Gilboa-Solomon, Yasmeen George, Xi Yang, Jianpeng Zhang, Jing Zhang, Yong Xia, Mengran Wu, Zhiyang Liu, Ed Walczak, Sean McSweeney, Ranveer Vasdev, Chris Hornung, Rafat Solaiman, Jamee Schoephoerster, Bailey Abernathy, David Wu, Safa Abdulkadir, Ben Byun, Justice Spriggs, Griffin Struyk, Alexandra Austin, Ben Simpson, Michael Hagstrom, Sierra Virnig, John French, Nitin Venkatesh, Sarah Chan, Keenan Moore, Anna Jacobsen, Susan Austin, Mark Austin, Subodh Regmi, Nikolaos Papanikolopoulos, Christopher Weight

arXiv:2307.01984v1cs.CVcs.AIcs.LG

TL;DR

Kidney-tumor segmentation needs accurate, reproducible annotations and evaluation that generalizes beyond a single institution. KiTS21 addresses this through transparent multiple-annotation procedures, an external test set, and challenge-wide analysis; its top team surpassed KiTS19 performance despite geographic and institutional shift. The challenge also exposed important limitations in annotation consistency and representation.

  • Problem

    Kidney-tumor segmentation requires accurate spatial delineation, but progress is constrained by limited high-quality annotations and concerns about single-institution evaluation and subgroup performance.

  • Method

    KiTS21 used multiple volumetric annotations produced through a transparent web-based process, an outside-institution test set, revised evaluation metrics, and meta-analysis of submitted methods.

  • Results

    The top-performing team surpassed the prior KiTS19 state-of-the-art performance despite evaluation on a test set from a different institution and geographic area.

  • Takeaways & Limitations

    KiTS21 advanced kidney-tumor segmentation while providing challenge-level evidence about methods, case characteristics, and performance under institutional and geographic shift.

  • Takeaways & Limitations

    Annotation quality remains affected by both genuine uncertainty and fatigue-related mistakes during precise delineation of many structures across many axial frames.

Abstract

from arXiv · show

This paper presents the challenge report for the 2021 Kidney and Kidney Tumor Segmentation Challenge (KiTS21) held in conjunction with the 2021 international conference on Medical Image Computing and Computer Assisted Interventions (MICCAI). KiTS21 is a sequel to its first edition in 2019, and it features a variety of innovations in how the challenge was designed, in addition to a larger dataset. A novel annotation method was used to collect three separate annotations for each region of interest, and these annotations were performed in a fully transparent setting using a web-based annotation tool. Further, the KiTS21 test set was collected from an outside institution, challenging participants to develop methods that generalize well to new populations. Nonetheless, the top-performing teams achieved a significant improvement over the state of the art set in 2019, and this performance is shown to inch ever closer to human-level performance. An in-depth meta-analysis is presented describing which methods were used and how they faired on the leaderboard, as well as the characteristics of which cases generally saw good performance, and which did not. Overall KiTS21 facilitated a significant advancement in the state of the art in kidney tumor segmentation, and provides useful insights that are applicable to the field of semantic segmentation as a whole.

1 Introduction

KiTS21 addresses the need for accurate, reproducible kidney-tumor segmentation by introducing a larger, more transparent challenge with multiple annotations, independent testing, and expanded clinical targets. It builds on prior challenge-based progress while supporting broader analysis of segmentation methods and performance.

  • 1.2 Kidney Tumor Radiomics: Accurate spatial delineation is essential for radiomics, but manual segmentation is time-consuming and subject to interobserver variability.These limitations motivate highly accurate automatic semantic segmentation methods for kidney tumors.
  • 1.2 Kidney Tumor Radiomics: Deep learning segmentation remains limited by the need for large, high-quality annotated datasets and many unresolved design choices.These constraints are especially relevant to kidney cancer, where little annotated data is publicly available.
  • 1.3 The KiTS21 Challenge: Challenge reports enable head-to-head comparisons and population-level meta-analyses of design decisions across participating methods.KiTS21 used reviewed method manuscripts to support analysis beyond identifying a single winning method.
  • 1.3 The KiTS21 Challenge: KiTS21 extends KiTS19 with public-view annotation, multiple annotations per region of interest, and renal cysts as an independent class.The challenge also required participating teams to submit method papers for review before participation.
  • 1.3 The KiTS21 Challenge: The test set came from a separate institution in a different geographic area, providing an external setting for evaluating generalization.This design addresses concerns that single-institution validation can inflate apparent performance.

2.1 The KiTS21 Dataset

The KiTS21 dataset combines a 300-case training cohort with a separate 100-case test cohort, using corticomedullary-phase CT and unified segmentation labels. Although treatment was centered at one academic institution, scans came from diverse sources, while the cohorts retained a substantial male predominance.

  • 2.1 The KiTS21 Dataset: KiTS21 comprised distinct training and test cohorts collected at separate time points for different purposes but annotated through one unified labeling effort.The cohort construction separated patient identification and image collection processes for the two sets.
  • 2.1 The KiTS21 Dataset: 300 cases were selected for the training set after reviewing 544 eligible nephrectomy cases for complete corticomedullary-phase kidney and tumor CT scans.The initial institutional query identified 962 patients, of whom 799 underwent detailed review for suitable imaging.
  • 2.1 The KiTS21 Dataset: The test set contained 100 Cleveland Clinic cases selected for available preoperative corticomedullary-phase CT after partial or radical nephrectomy for suspected renal malignancy.This cohort was collected separately from the training cohort.
  • 2.1 The KiTS21 Dataset: Training patients were treated at one academic center, but their preoperative scans came from varied community hospitals and clinics, creating scanner and protocol heterogeneity.The geographic distribution of scanning institutions is shown in Figure 2.

2.2 Data Annotation Process

KiTS21 introduced a transparent, purpose-built annotation process to reduce delineation mistakes and quantify residual variability. Trainees localized and guided each region of interest, after which laypeople produced three independent delineations under shared annotation intent.

  • Annotation transparency: The process was designed to improve clarity about how medical-image annotations were defined, produced, and instructed.A public website exposed training-case annotation status, annotation views, and the instructions used by annotators.
  • Sources of annotation error: Annotation uncertainty arises from indistinct tumor borders, partial-volume artifacts, and mistakes during extensive slice-by-slice delineation.The paper distinguishes genuine boundary uncertainty from fatigue-related deviations between annotators’ intentions and completed delineations.
  • Three-phase workflow: The workflow comprised localization, guidance, and delineation phases.A trainee placed a 3D bounding box, added T-shaped pins along intended boundaries, and a layperson used those cues to produce the delineation.
  • Multiple annotations: Three independent layperson delineations were collected for each region after trainee localization and guidance were reviewed and refined by an expert when needed.Using a shared annotation intent enabled quantifying and controlling for delineation mistakes.
  • Final segmentation classes: The dataset prioritized kidney, tumor, and cyst labels because producing sufficiently high-quality labels for ureter, artery, and vein was not feasible before release.Complex structures required substantial correction and refinement time, so those additional regions were left for future work.

2.3 Challenge Design Decisions

KiTS21 redesigned evaluation and challenge operations to test generalization, account for hierarchical segmentation errors, and improve reporting quality. It used a separate-institution test cohort, expanded metrics and ranking, peer-reviewed submission papers, and ultimately revised private-evaluation plans because of resource constraints.

  • Separate-institution test set: The test set used patients treated at Cleveland Clinic, providing a properly separate institutional cohort for external evaluation.This design addressed concerns that single-institution validation can overestimate performance relative to deployment populations.
  • Submission reporting: KiTS21 required peer review of challenge-submission papers to improve clarity and completeness of method reporting.A template guided expected content; 27 of 28 submitted papers were approved, most after revisions and repeat review.
  • Metrics: Sørensen-Dice ranking was supplemented with Surface Dice because Sørensen-Dice alone can mishandle object detection and relative weighting across multiple objects.The concern included cases where smaller objects receive disproportionate influence despite clinical importance of detecting a smaller contralateral tumor.
  • Metrics: Hierarchical Evaluation Classes grouped all regions, tumor-plus-cyst masses, and tumor alone to reflect nested target relationships and avoid double penalization.The hierarchy targets difficult distinctions between tumors and cysts and between masses and healthy kidney.
  • Ranking procedure: Final rankings averaged Sørensen-Dice and Surface Dice across randomly sampled composite segmentations for each hierarchical class, with tumor Sørensen-Dice breaking ties.Three annotations per region yielded 3^N possible composite masks, where N was the number of regions of interest.
  • Submission procedure: Private Docker-based evaluation was abandoned after more than half of submitted containers exceeded cloud-system time or memory limits.The planned procedure was therefore impractical with the available resources.

3.1 Performance and Ranking

KiTS21 evaluated 25 teams on 100 external test cases, combining tumor-segmentation performance with statistical analysis of leaderboard stability. The top teams approached inter-rater agreement, while case-level boundary variation remained important for clinical applications.

  • 25 teams were included in the final leaderboard after exclusions and withdrawals from the 29 teams that submitted papers.
  • 0.86 tumor Dice was achieved by the top team, compared with 0.88 inter-rater agreement; the top five teams had no complete tumor misses across 100 cases.The remaining top-five values were 0.83, 0.83, 0.82, and 0.81.
  • Tumor predictions generally targeted the correct kidney region, but teams varied substantially in delineating the tumor–kidney boundary.The boundary was described as more difficult than tumor detection itself.
  • Boundary errors matter for surgical planning because removing too much healthy parenchyma can impair renal function, while leaving tumor behind can increase positive margins and recurrence risk.
  • The first-place team was not statistically superior to any of the top five at α = 0.05, but it was superior to nearly every team thereafter except the ninth-place team.Pairwise analyses used multiple-testing correction, underscoring uncertainty in a static ranking based on a finite test set.
  • Challenge rankings are limited for identifying one definitively best method, but they support population-level analysis of which methods perform better and which design choices are common among submissions.

3.2 Methods Used

The top of the KiTS21 leaderboard was dominated by nnU-Net and, to a lesser extent, coarse-to-fine frameworks and transfer learning. The winning system combined nnU-Net with cross-entropy and surface loss.

  • nnU-Net dominated the top half of the leaderboard, while coarse-to-fine frameworks and transfer learning were also overrepresented among high-performing teams.
  • The most popular loss was the sum of cross entropy and Dice loss, but the winning team used nnU-Net with a sum of cross entropy and surface loss.
  • Teams that did not use nnU-Net often adopted attention-based architectures, including two visual transformer networks, but these generally underperformed nnU-Net.

3.3 Hidden Strata Analysis

Hidden-strata analyses found worse performance for non-white patients among the top five teams, while women performed better than men. Transformer-based teams showed a different pattern, with tumor size more influential and race and sex nonsignificant.

  • 3.3.1 Hypothesis-Driven Analysis: Training-set underrepresentation motivated testing whether segmentation performance varied by race and sex.
  • 3.3.1 Hypothesis-Driven Analysis: Non-white patients had significantly worse average tumor Dice than white patients among the top five teams, whereas women had significantly better performance than men.The analysis used multivariate regression with race, gender, and two additional covariates.
  • 3.3.1 Hypothesis-Driven Analysis: For the two transformer teams, race and gender were nonsignificant, while tumor size appeared to play a larger role.The authors state that this could indicate greater robustness to hidden strata, while noting that the top-five teams still performed better on the weakest transformer subpopulations.
  • 3.3.2 Unsupervised Analysis: Test-case clustering identified groups where nearly every team performed poorly, nearly every team performed well, or only selected teams performed well.The clusters were formed using each team’s performance metrics as a feature vector.

3.4 Methods Used by Top 3 Teams

The three highest-performing KiTS21 teams used cascaded or ensemble 3D U-Net strategies, with coarse-to-fine localization, multi-resolution processing, and task-specific postprocessing. The top three achieved mean volumetric Dice scores of 0.908, 0.896, and 0.8944, respectively.

  • First place: The first-place method used four nnU-Net networks in a coarse-to-fine pipeline that segmented kidneys, then tumor and mass regions, before aggregating predictions.The coarse kidney mask generated crops for finer kidney, tumor, and mass segmentation; masses represented tumor and cyst unions.
  • First place: A surface loss was added after Dice and cross-entropy objectives plateaued to align optimization with the surface Dice leaderboard metric.Surface Dice was used alongside volumetric Dice for final ranking.
  • First place: 0.908 average volumetric Dice and 0.826 average surface Dice earned the first-place method the KiTS21 leaderboard lead.Its kidney-region volumetric Dice was 0.86, close to the reported KiTS19 interobserver agreement of 0.88.
  • Second place: The second-place submission ensembled three single-stage 3D U-Nets with one two-stage cascade using low- and high-resolution processing.Its postprocessing removed implausible tumor and cyst findings outside kidneys and small kidney fragments surrounded by another class.
  • Second place: 0.896 mean Dice and 0.816 mean Surface Dice placed the ensemble second on the KiTS21 leaderboard.The ensemble included a regularized-loss model, a differently seeded model, and models initialized from LiTS training.
  • Third place: The third-place method used a two-stage coarse-to-fine 3D U-Net, first delineating kidneys at downsampled resolution and then segmenting all three classes at full resolution.The second stage was guided by the first-stage segmentation maps.
  • Third place: 0.8944 mean sampled average Dice and 0.8140 mean sampled average surface Dice produced third place on a 100-scan test set.The cascade achieved average Dice values of 0.9748 for kidney, 0.8813 for tumor, and 0.8710 for cyst.

4 Conclusions

KiTS21 advanced kidney tumor segmentation through expanded, more transparent annotations and a test set from a different institution and geographic area. Its analysis found nnU-Net remained dominant, while hidden-strata results showed that leaderboard rank did not guarantee uniform performance across subpopulations.

  • Challenge design: KiTS21 introduced multiple volumetric annotations per region of interest and a fully transparent web-based annotation process.The challenge attracted 25 full submissions from teams worldwide.
  • Challenge results: The top-performing team surpassed KiTS19 state-of-the-art performance despite evaluation on a test set from an entirely different institution and geographic area.
  • Method analysis: The meta-analysis found continued dominance of nnU-Net, alongside substantial interest in transformer and other attention-based methods.
  • Method analysis: Hidden-strata analysis showed that the highest-ranked teams were not necessarily the most uniform performers across test-set subpopulations.
  • Future directions: Future KiTS editions aim to improve dataset quality while adding heterogeneity and complexity to move evaluation toward more realistic real-world settings.KiTS23 adds the venous contrast phase to the corticomedullary phase used in KiTS19 and KiTS21.
Loading 2307.01984v1…