Source-linked AI summary

Why rankings of biomedical image analysis competitions should be interpreted with care

Lena Maier-Hein, Matthias Eisenmann, Annika Reinke, Sinan Onogur, Marko Stankovic, Patrick Scholz, Tal Arbel, Hrvoje Bogunovic, Andrew P. Bradley, Aaron Carass, Carolin Feldmann, Alejandro F. Frangi, Peter M. Full, Bram van Ginneken, Allan Hanbury, Katrin Honauer, Michal Kozubek, Bennett A. Landman, Keno März, Oskar Maier, Klaus Maier-Hein, Bjoern H. Menze, Henning Müller, Peter F. Neher, Wiro Niessen, Nasir Rajpoot, Gregory C. Sharp, Korsuk Sirinukunwattana, Stefanie Speidel, Christian Stock, Danail Stoyanov, Abdel Aziz Taha, Fons van der Sommen, Ching-Wei Wang, Marc-André Weber, Guoyan Zheng, Pierre Jannin, Annette Kopp-Schneider

arXiv:1806.02051v2cs.CV

TL;DR

Biomedical image-analysis challenges are important validation tools, but the field lacks consistent quality control and reporting. The paper evaluates 150 challenges and finds that rankings and interpretations are sensitive to design choices, motivating guidelines and further research.

  • Problem

    Biomedical image-analysis challenges have high scientific importance, but commonly respected quality-control processes for their design and organization do not exist.

  • Method

    The paper comprehensively evaluates biomedical image-analysis challenges and examines their reporting, design, and ranking practices.

  • Results

    Challenge reporting is often insufficient for reproducibility and interpretation, while rankings are sensitive to metrics, aggregation schemes, test data, and annotating observers.

  • Takeaways & Limitations

    The authors recommend international guidelines, rigorous challenge assessment, better data feedback, incentives for high-quality challenges, and research on unresolved design issues.

  • Takeaways & Limitations

    Statistical significance may not reflect clinical or biological relevance, and clinically relevant performance differences may be nonsignificant with small samples.

Abstract

from arXiv · show

International challenges have become the standard for validation of biomedical image analysis methods. Given their scientific impact, it is surprising that a critical analysis of common practices related to the organization of challenges has not yet been performed. In this paper, we present a comprehensive analysis of biomedical image analysis challenges conducted up to now. We demonstrate the importance of challenges and show that the lack of quality control has critical consequences. First, reproducibility and interpretation of the results is often hampered as only a fraction of relevant information is typically provided. Second, the rank of an algorithm is generally not robust to a number of variables such as the test data used for validation, the ranking scheme applied and the observers that make the reference annotations. To overcome these problems, we recommend best practice guidelines and define open research questions to be addressed in the future.

1. Introduction

Challenges have become important validation tools in biomedical image analysis, but inconsistent reporting and design make results difficult to reproduce, interpret, and compare.

  • Challenge performance influences paper acceptance, community impact, scientific careers, and potential clinical translation, yet no commonly respected quality-control processes exist.
  • The paper presents a comprehensive evaluation of 150 biomedical image analysis challenges conducted through the end of 2016.
  • Reproduction, interpretation, and cross-comparison are often impossible because challenges report only a fraction of relevant information and use heterogeneous designs.
  • Algorithm rankings are sensitive to test data, annotating observers, performance metrics, and value-aggregation methods.

2. Results

Analysis of 150 biomedical image analysis challenges found incomplete reporting, heterogeneous design, and rankings sensitive to metrics, aggregation, test data, and annotators. Questionnaire respondents broadly supported improved design, quality control, and comprehensive reporting.

  • Challenge landscape: 150 challenges comprised 549 tasks, mostly segmentation (70%) and classification (10%), with half organized in the MICCAI context.56% published results in journals or conference proceedings.
  • Reporting: A task reported a median of 64% of 53 relevant challenge parameters, leaving substantial information unavailable for interpretation and reproduction.Only 6% of parameters were reported for all tasks.
  • Challenge design: 97 performance metrics were used, with metric choice typically unjustified (77%) and 51% of metrics applied to only one task.Metric heterogeneity was also substantial across comparable challenges.
  • Ranking robustness: Changing HD to HD95 ranked the HD-based tenth-place algorithm first in one 2015 segmentation challenge.DSC was used in 92% of segmentation tasks, while HD was used in 47%.
  • Ranking robustness: Ranking outcomes changed with aggregation choices, observers, and test data: winners remained stable in only 21%, 11%, and 9% of tasks for DSC, HD, and HD95, respectively.Different observers produced different rankings in 15%, 46%, and 62% of pairwise comparisons for DSC, HD, and HD95; mean and metric-based aggregation were more robust.
  • Missing data: 82% of tasks gave no information about missing-data handling, and 25% of non-winning algorithms could have ranked first by submitting plausible results systematically.In 9% of 56 tasks, every participating team could have ranked first under the evaluated missing-data behavior.
  • Community feedback: Among 295 questionnaire participants from 23 countries, 92% supported improving challenge design, 87% wanted best-practice guidelines, and 71% supported more quality control.Concerns covered data, annotation, evaluation, and documentation.
  • Recommendations: The authors recommend comprehensive reporting, robust ranking procedures, clinically relevant metrics, and multiple metric results with visualizations.Structured MICCAI 2018 submissions instantiated a median of 100% of essential parameters, versus 53% in previous tasks.

3. Discussion

Challenges are essential benchmarking tools in biomedical image analysis, but poor reporting and heterogeneous design undermine reproducibility, interpretation, and ranking stability. The paper recommends standardized parameters, improved incentives, and further research to address these systemic hurdles.

  • Challenges cover a broad range of biomedical image analysis problems, algorithm classes, and imaging modalities.
  • Poor reporting prevents adequate interpretation and reproducibility of challenge results.
  • Heterogeneous challenge design and ranking choices make algorithm rankings sensitive to metrics, aggregation schemes, test data, and observers.
  • A list of 53 challenge parameters is recommended before execution to support fairness, transparency, interpretability, and reproducibility.
  • The authors call for international guidelines, rigorous assessment, open feedback on challenge data, incentives for high-quality challenges, and research on unresolved design issues.
  • Challenges are essential to biomedical image analysis, but systemic hurdles must be overcome to realize their potential as valid benchmarking tools.

Definitions

The paper defines challenges, tasks, cases, metrics, and aggregation schemes as related components of biomedical image analysis competitions. These terms specify how problems are organized, evaluated, and converted into algorithm rankings.

  • A challenge is an open competition organized around a dedicated biomedical image analysis problem, potentially containing multiple separately assessed tasks.
  • A task is a challenge subproblem with its own ranking or leaderboard and potentially distinct assessment metrics.
  • A case is a dataset containing one or more biomedical images and typically a gold-standard annotation for testing.
  • A metric measures algorithm performance for a case, often normalized from 0 for worst performance to 1 for best performance.
  • Metric-based aggregation first combines metric values across test cases, whereas case-based aggregation first ranks algorithms within each test case.

Inclusion criteria

The study aimed to capture biomedical image analysis challenges conducted through 2016 and analyzed qualifying challenges and tasks using defined inclusion criteria. The resulting corpus included 150 challenges and 549 tasks.

  • The inclusion effort targeted biomedical image analysis challenges conducted up to 2016, excluding 2017 challenges because publication of scientific reports could be delayed.
  • The study separately specified inclusion criteria for challenge-level and task-level analyses.

Challenge parameter list

The paper developed a structured parameter list to describe challenge design and results comprehensively, supporting interpretation and reproducibility. The list was iteratively refined through analysis, tooling, expert input, and ontological modeling.

  • The parameter-list effort aimed to formalize challenge design and results comprehensively to facilitate interpretability and reproducibility.
  • An initial parameter set was based on parameters for reference-based validation studies.
  • Challenge-specific parameters were added during analysis of challenge websites and scientific papers.
  • A formalization tool prompted further refinement of the parameter list during use.
  • An international questionnaire finalized the list by soliciting comments on parameter names, descriptions, importance, instantiations, and additions.
  • The final list was represented as an ontology and used for structured submission of MICCAI 2018 biomedical challenges.

Statistical methods

The analysis quantified ranking agreement and variability using Kendall’s tau, bootstrapping, and leave-one-out analyses rather than statistical testing.

  • Statistical methods: Kendall’s tau quantified agreement between rankings produced by different aggregation methods or metric variants.Tau ranges from 1 for identical rankings to -1 for reverse rankings.
  • Statistical methods: Bootstrapping measured ranking variability by repeatedly applying a ranking scheme to 1000 samples drawn from a task’s datasets.The original ranking and winner were compared with rankings from each bootstrap sample.
  • Statistical methods: Leave-one-out analysis assessed ranking stability after removing one dataset and reapplying the ranking scheme.The same summary measures used for bootstrapping were determined for the reduced dataset subsets.
  • Statistical methods: The study avoided statistical tests because dataset counts and algorithm numbers vary across tasks, making results incomparable by design.The authors also found cases where statistical significance and bootstrap stability gave conflicting conclusions.

Experiment: Comprehensive reporting

The comprehensive challenge analysis examined the field’s role, design practices, and whether reporting supports reproducibility and interpretation, using structured observer formalization.

  • Experiment: Comprehensive reporting: The analysis asked what role biomedical image analysis challenges play across fields, algorithm categories, and imaging modalities.This question included how many challenges had been conducted to date.
  • Experiment: Comprehensive reporting: It investigated common challenge-design practices, including metrics, ranking methods, training and test images, and annotation procedures.The analysis also assessed whether common standards exist.
  • Experiment: Comprehensive reporting: It evaluated whether challenge reporting enables reproducibility and adequate interpretation of results.This was treated as a distinct research question from challenge role and design practice.
  • Experiment: Comprehensive reporting: Five engineers and one medical student formalized qualifying challenges using a parameter-list instantiation tool.Each challenge was independently formalized by two observers, with a third consulted when disagreements remained.

Experiment: Sensitivity of challenge ranking

The ranking-sensitivity experiments tested how challenge rankings change with metrics, aggregation choices, observers, test data, and missing-value handling.

  • Experiment: Sensitivity of challenge ranking: The experiments addressed whether rankings depend on specific test cases, metric variants, aggregation methods, and reference observers.They also examined whether rankings vary across commonly applied metrics and ranking schemes.
  • Experiment: Sensitivity of challenge ranking: Organizers of 2015 segmentation challenges were asked to provide per-dataset DSC, HD, and HD95 results for 124 tasks.Segmentation was selected because it represented 70% of biomedical image analysis challenges.
  • Experiment: Sensitivity of challenge ranking: For 56 eligible segmentation tasks, rankings were compared across metric variants, aggregation operators, aggregation categories, and observers.The default single-metric rankings used DSC and HD, with Kendall’s tau evaluating changes.
  • Experiment: Sensitivity of challenge ranking: Ranking robustness was quantified for DSC, HD, and HD95 using bootstrapping and leave-one-out analysis.The study also compared metric-based with case-based aggregation and mean with median aggregation.
  • Experiment: Sensitivity of challenge ranking: The study investigated whether ignoring missing values, a practice reported for 82% of tasks, could be exploited to manipulate rankings.This analysis used per-dataset assessment data from qualifying 2015 segmentation challenges.

International Survey

The international survey gathered challenge-organization and design information through a questionnaire distributed across research communities and challenge networks.

  • International Survey: The questionnaire covered challenge identity, venue, timetable, ethics approval, data usage, interaction level, organizer participation, training-data usage, pre-evaluation, and submission procedures.These parameters were illustrated with representative challenge instantiations and response options.
  • International Survey: The survey represented one-time, open-call, and repeated challenge formats across examples including brain tumor segmentation, instrument tracking, and literature image classification.Examples included challenges associated with DREAM, ImageCLEF, and online competitions.
  • International Survey: Challenge policies varied regarding data redistribution, algorithm interaction, organizer participation, additional training data, pre-evaluation, and result submission.Examples ranged from restricted redistribution and automatic-only participation to reusable data and semi-automatic algorithms.
  • International Survey: Reference-estimation error sources included acquisition errors, user errors, preprocessing errors, and inter- or intra-observer variability.For literature image classification, ambiguity and rapid annotation were identified as important error sources, with inter-observer variability measured separately.
  • International Survey: Only complete questionnaires and questionnaires with more than 50% of answers were included in the analysis.The retained sample included 117 complete questionnaires and 12 questionnaires with more than half of the answers.

Data

Challenges face substantial data-related concerns, especially limited representativeness and small numbers of participating institutes. Acquisition barriers, heterogeneous data, infrastructure problems, and restricted accessibility further constrain their usefulness.

  • Data: 33% of concerns involved data representativeness, including selection bias, imbalance, unrealistic artefacts, and typically few centers, vendors, or devices.The median number of institutes involved was 1 (IQR: (1, 1)).
  • Data: Small challenge data sets make it impossible to distinguish an algorithm’s effect from the effect of participants’ supplementary training data.
  • Data: 17% of concerns involved data acquisition, with legal barriers and high costs limiting organizers; 22% encountered acquisition problems.Access to validation data motivated 30% of challenge participants.
  • Data: Heterogeneity, infrastructure issues, post-challenge accessibility, documentation gaps, data quality, and overfitting or cheating were additional reported problems.Heterogeneity, infrastructure issues, accessibility, and documentation were each reported at 8%; data quality at 6%; overfitting/cheating/tuning at 4%.

Annotation

Reference annotations are constrained by subjectivity, limited quality control, and inconsistent generation practices. Participants also reported transparency, resource, and standardization problems surrounding annotation.

  • Annotation: 33% of concerns involved reference-data quality, including subjective or biased annotations and insufficient quality control.Single observers and automatic annotation or initialization tools were cited as contributing factors.
  • Annotation: 16% of concerns involved reference-generation methods, with expert annotations varying significantly and multiple-expert merging reported for only 73% of tasks.Additionally, 27% of challenge organizers encountered problems generating reference data.
  • Annotation: 15% of concerns involved annotation transparency, including requests for raw annotations, observer variability, and documentation of reference generation.
  • Annotation: 14% of concerns involved resources because high-quality annotation was challenging and logistically difficult with improper tools.
  • Annotation: 10% of concerns involved missing annotation and data-format standards, including guidelines for merging annotations.

Evaluation

Challenge evaluation is affected by disputed metrics, inconsistent standards, and insufficient transparency. Ranking methods, quality control, uncertainty handling, and excessive focus on ranks were additional concerns.

  • Evaluation: 20% of concerns involved metric choice, including weak clinical linkage, rare runtime consideration, difficult aggregation with missing data, and 23% of organizers struggling with metric selection.
  • Evaluation: 19% of concerns involved missing standards for metrics and evaluation frameworks, with identical metrics sometimes named or applied differently.
  • Evaluation: 12% of concerns involved missing evaluation documentation, including nontransparent criteria that could allow organizers to influence final rankings.
  • Evaluation: Additional difficulties included inadequate evaluation infrastructure, ranking methods for missing values, metric quality control, excessive focus on sensitive ranks, and uncertainty handling.These concerns were reported at 7%, 7%, 7%, 5%, and 4%, respectively.

Documentation

Challenge reporting is often incomplete and insufficiently transparent across design, results, and methods. Publication delays, changing websites, missing code, and inconsistent reporting standards further impair documentation.

  • Documentation: 47% of concerns involved completeness and transparency, with comprehensive reporting of challenge design, results, and methods currently uncommon.Method performance may depend crucially on parameters, while only 4% of participants stated that all parameters were provided to participants.
  • Documentation: 13% of concerns involved publication delays and discrepancies between challenge websites and corresponding publications.
  • Documentation: 10% of concerns involved challenge lifetime and dynamics, because changing website content makes proper referencing difficult and optimal challenge lifetime remains an open question.Overfitting was cited as one consideration in determining challenge lifetime.
  • Documentation: Further documentation problems included missing open-source code, structured-reporting standards, post-challenge information accessibility, and acknowledgement of contributors.These concerns were reported at 9%, 7%, 5%, and 3%, respectively.
Loading 1806.02051v2…