Source-linked AI summary

Understanding metric-related pitfalls in image analysis validation

Annika Reinke, Minu D. Tizabi, Michael Baumgartner, Matthias Eisenmann, Doreen Heckmann-Nötzel, A. Emre Kavur, Tim Rädsch, Carole H. Sudre, Laura Acion, Michela Antonelli, Tal Arbel, Spyridon Bakas, Arriel Benis, Matthew Blaschko, Florian Buettner, M. Jorge Cardoso, Veronika Cheplygina, Jianxu Chen, Evangelia Christodoulou, Beth A. Cimini, Gary S. Collins, Keyvan Farahani, Luciana Ferrer, Adrian Galdran, Bram van Ginneken, Ben Glocker, Patrick Godau, Robert Haase, Daniel A. Hashimoto, Michael M. Hoffman, Merel Huisman, Fabian Isensee, Pierre Jannin, Charles E. Kahn, Dagmar Kainmueller, Bernhard Kainz, Alexandros Karargyris, Alan Karthikesalingam, Hannes Kenngott, Jens Kleesiek, Florian Kofler, Thijs Kooi, Annette Kopp-Schneider, Michal Kozubek, Anna Kreshuk, Tahsin Kurc, Bennett A. Landman, Geert Litjens, Amin Madani, Klaus Maier-Hein, Anne L. Martel, Peter Mattson, Erik Meijering, Bjoern Menze, Karel G. M. Moons, Henning Müller, Brennan Nichyporuk, Felix Nickel, Jens Petersen, Susanne M. Rafelski, Nasir Rajpoot, Mauricio Reyes, Michael A. Riegler, Nicola Rieke, Julio Saez-Rodriguez, Clara I. Sánchez, Shravya Shetty, Maarten van Smeden, Ronald M. Summers, Abdel A. Taha, Aleksei Tiulpin, Sotirios A. Tsaftaris, Ben Van Calster, Gaël Varoquaux, Manuel Wiesenfarth, Ziv R. Yaniv, Paul F. Jäger, Lena Maier-Hein

arXiv:2302.01790v4cs.CV

TL;DR

Image-analysis validation metrics are often poorly matched to research questions, while reliable guidance on their pitfalls remains difficult to access. This paper used expert consensus and community feedback to create a comprehensive, domain-agnostic resource, showing that common metric choices were frequently unjustified in 24% of surveyed competitions.

  • Problem

    Validation metrics in image analysis depend on the underlying task, but reliable information about their strengths, weaknesses, and pitfalls remains difficult to access.

  • Method

    A multidisciplinary consortium used a multi-stage Delphi process and community feedback to compile and taxonomize validation-metric pitfalls across four classification-related image-analysis tasks.

  • Results

    The resource provides a comprehensive, domain-agnostic collection of validation-metric pitfalls, while 24% of surveyed competitions justified metric choices by community common practice.

  • Takeaways & Limitations

    The work offers researchers a single point of access to structured information for understanding validation-metric pitfalls in image analysis.

  • Takeaways & Limitations

    The work examines reference-based metric pitfalls only for image-level classification, semantic segmentation, instance segmentation, and object detection.

Abstract

from arXiv · show

Validation metrics are key for the reliable tracking of scientific progress and for bridging the current chasm between artificial intelligence (AI) research and its translation into practice. However, increasing evidence shows that particularly in image analysis, metrics are often chosen inadequately in relation to the underlying research problem. This could be attributed to a lack of accessibility of metric-related knowledge: While taking into account the individual strengths, weaknesses, and limitations of validation metrics is a critical prerequisite to making educated choices, the relevant knowledge is currently scattered and poorly accessible to individual researchers. Based on a multi-stage Delphi process conducted by a multidisciplinary expert consortium as well as extensive community feedback, the present work provides the first reliable and comprehensive common point of access to information on pitfalls related to validation metrics in image analysis. Focusing on biomedical image analysis but with the potential of transfer to other fields, the addressed pitfalls generalize across application domains and are categorized according to a newly created, domain-agnostic taxonomy. To facilitate comprehension, illustrations and specific examples accompany each pitfall. As a structured body of information accessible to researchers of all levels of expertise, this work enhances global comprehension of a key topic in image analysis validation.

MAIN … A common taxonomy enables domain-agnostic categorization of pitfalls

The paper identifies widespread metric-selection and application pitfalls in image analysis, driven by complex research objectives, inaccessible knowledge, and unjustified common practices. A multidisciplinary Delphi process and community feedback yield a structured, domain-agnostic taxonomy intended as a central resource for understanding these pitfalls.

  • MAIN: Image-analysis metrics must match the research question, because tasks may require object localization, classification, or exact structure boundaries.A metric suitable for drawing a bounding box may be unsuitable for tracing boundaries for fluorescent signal quantification.
  • Information on metric pitfalls is largely inaccessible: Researchers frequently choose validation metrics inadequately because metric properties and limitations are difficult to access in a central resource.The paper describes this accessibility gap as a major bottleneck for image-analysis validation.
  • Historically grown practices are not always justified: 24% of surveyed 2022 competitions based metric choices on community common practice, although such practices were often unjustified and could propagate poor methods.The paper also highlights incorrect naming and inconsistent mathematical formulation of a cell-instance-segmentation metric as a representative example.
  • A multidisciplinary Delphi process reveals numerous pitfalls in biomedical image analysis validation: A consortium of 62 multidisciplinary experts used a multi-stage Delphi process to build a reliable collection of biomedical image-analysis metric definitions, limitations, and pitfalls.The process was designed for consensus building and future information access.
  • A common taxonomy enables domain-agnostic categorization of pitfalls: The taxonomy groups pitfalls into inadequate problem-category choice, poor metric selection, and poor metric application.Poor metric selection is subdivided according to domain interest, target structures, data-set properties, and algorithm-output properties.
  • A common taxonomy enables domain-agnostic categorization of pitfalls: The taxonomy is domain-agnostic, reflecting that metric-related pitfalls generalize across imaging domains and modalities.It links pitfall categories with individual metrics and supports rapid identification of which metrics are affected.
  • A common taxonomy enables domain-agnostic categorization of pitfalls: Metric suitability can be undermined by target-structure, data-set, and algorithm-output properties, including structure size, class imbalance, small samples, annotations, and overlapping predictions.For example, Balanced Accuracy may score highly despite many False Positive samples in an imbalanced setting.
  • A common taxonomy enables domain-agnostic categorization of pitfalls: Poor metric application includes nonstandardized implementation, inappropriate aggregation, noncomplementary rankings, poorly informative reporting, and misinterpretation of scores.Different implementations can produce substantially different scores, while aggregation can hide information about images, classes, hierarchy, missing values, and bias.

The first illustrated common access point to metric definitions and pitfalls

The work provides a structured, illustrated common access point to metric definitions and pitfalls, addressing the limited coverage of existing literature and online resources. Dedicated illustrations and metric profiles are designed to make key information understandable and accessible across expertise levels.

  • Coverage of existing resources: 68% of identified metric-related pitfalls were located in existing research literature, while 11% appeared in online resources and 8% in both.The search covered Google Scholar, Google, and online resources such as blog posts.
  • Structured access point: The resource presents metric pitfalls in a highly structured and easily understandable form.This structure is intended to provide a common access point to information on metric pitfalls.
  • Illustrated pitfalls: A dedicated illustration accompanies each discussed pitfall to facilitate comprehension and accessibility regardless of expertise level.The illustrations are contained in SUPPL. NOTE 2.

DISCUSSION · EXTENDED DATA

The work identifies widespread difficulties and pitfalls in selecting and interpreting validation metrics for biomedical image analysis, and organizes them into an accessible, domain-agnostic resource. Extended Data illustrates and tabulates these pitfalls across classification, segmentation, object detection, and instance segmentation tasks.

  • DISCUSSION: Poor validation-metric choices impede clinical translation and undermine assessment of scientific progress in biomedical image analysis.These flaws frequently arise when researchers disregard the specific properties and limitations of individual metrics.
  • DISCUSSION: Researchers face an overwhelming, fragmented literature with no common entry point to reliable information about metric pitfalls and limitations.Finding relevant information often requires knowing the exact pitfall and related keywords in advance.
  • DISCUSSION: A multi-stage Delphi process combined distributed expertise and community feedback to produce a comprehensive, practically relevant illustrated collection of metric pitfalls.Sharing early results as a dynamic preprint maintained proximity to issues arising in practical applications.
  • DISCUSSION: The pitfalls generalize across imaging modalities and application domains and are organized by underlying sources into a domain-agnostic taxonomy.The taxonomy provides structure for the large number of identified pitfalls while avoiding domain-specific categorization.
  • DISCUSSION: The complementary Metrics Reloaded framework uses a domain-independent problem fingerprint to guide metric selection for specific tasks, informed by the pitfalls described here.The fingerprint abstracts from specific domain knowledge while capturing properties relevant to metric selection.
  • DISCUSSION: The scope covers image-level classification, semantic segmentation, instance segmentation, and object detection, while regression and registration were excluded as beyond scope.The included tasks share similarities because they can be viewed as classification at image, object, or pixel levels.
  • DISCUSSION: The resulting resource is intended to improve validation quality, accelerate translation into practice, and raise awareness of flawed AI validation beyond biomedical imaging.The expert consortium primarily focused on biomedical applications, although the pitfalls’ generalization suggests broader applicability.
  • EXTENDED DATA: Extended Data presents metric-specific pitfall sources and illustrations across classification, semantic and instance segmentation, object detection, and localization criteria.Examples include small structures substantially affecting Dice Similarity Coefficient and Intersection over Union, and overlapping predictions motivating instance rather than semantic segmentation.

CODE AVAILABILITY STATEMENT

Reference implementations for all Metrics Reloaded metrics are provided in the MONAI open-source framework and are publicly accessible on GitHub.

  • Reference implementations for all Metrics Reloaded metrics are provided within the MONAI open-source framework.They are accessible at https: //github.com/Project-MONAI/MetricsReloaded.

COMPETING INTERESTS

The authors disclose extensive competing interests, including employment, shareholdings, advisory roles, grants, external funding, consulting fees, and patent royalties. These disclosures involve multiple healthcare, pharmaceutical, technology, and medical-imaging companies.

  • Competing interests: Authors report employment or ownership ties to Siemens AG, Thirona, HeartFlow Inc, Kheiron Medical Technologies Ltd, Lunit, Aiosyn BV, Histofy, and Nvidia GmbH.Additional disclosures include an advisory-board role at Canon Healthcare IT and an Nvidia GPU Grant.
  • Competing interests: One author reports funding from GSK, Pfizer, and Sanofi, plus fees from Travere Therapeutics, Stadapharm, Astex Therapeutics, Pfizer, and Grunenthal.Another author receives patent royalties from iCAD, ScanMed, and Philips.

SUPPLEMENTARY METHODS · Literature search · Delphi process

The supplementary methods combined structured literature searches with a multi-stage Delphi process and community feedback to identify, verify, illustrate, and organize metric-related pitfalls. Searches targeted metric-specific limitations in scholarly and online sources, while expert surveys established consensus and a taxonomy.

  • Literature search: Literature searches used Google Scholar with patents included, citations excluded, and otherwise default settings unchanged.
  • Literature search: Each metric search combined quoted metric names, synonyms, and acronyms with “metric” and quoted terms for pitfalls, limitations, and flaws.
  • Literature search: The DSC search string included DSC synonyms such as Dice Similarity Coefficient, Sørensen–Dice coefficient, F1 score, and DCE.
  • Literature search: A second search using Google Scholar and Google assessed whether proposed pitfalls appeared in research literature or online resources and whether they were visually presented.
  • Delphi process: More than 60 international biomedical image analysis experts contributed to a multi-stage Delphi process supplemented by general scientific community feedback.
  • Delphi process: The Delphi stages compiled pitfall sources, collected concrete pitfalls, incorporated social media-based feedback, and obtained final agreement on inclusion.
  • Delphi process: The final collection was illustrated, metric values were verified by two independent observers, and the pitfalls were organized into an expert-approved taxonomy.

Expert consortium

The expert consortium comprised 70 researchers from 65 institutions, spanning 19 countries and five continents. Most experts were professors or postdoctoral researchers, with a median h-index of 31.5 and median academic age of 18 years.

  • Consortium composition: 70 researchers from 65 institutions formed the expert consortium, representing 19 countries across five continents.The consortium included 70% male and 30% female researchers.
  • Professional backgrounds: 50% of experts were professors and 39% were postdoctoral researchers.These were the two largest reported professional groups in the consortium.
  • Research experience: The consortium had a median h-index of 31.5 and median academic age of 18 years.The reported h-index ranged from 6 to 113, while academic age ranged from 3 to 42 years.

SUPPLEMENTARY NOTES · SUPPL. NOTE 1 · METRIC FUNDAMENTALS

Metric validation in biomedical image analysis is grounded mainly in confusion-matrix cardinalities, but metric terminology and computation vary across task types. The notes distinguish classification, segmentation, detection, and instance-segmentation settings, including thresholding, localization, assignment, and calibration considerations.

  • METRIC FUNDAMENTALS: Biomedical image-analysis metrics commonly derive from TP, FN, FP, and TN cardinalities in confusion matrices.These cardinalities underlie problems interpreted as image-, object-, or pixel-level classification tasks.
  • METRIC FUNDAMENTALS: Counting metrics operate directly on confusion-matrix cardinalities, while multi-class and per-class variants differ in class handling.In segmentation, counting metrics are typically called overlap-based metrics.
  • METRIC FUNDAMENTALS: Sensitivity, True Positive Rate (TPR), and Recall are equivalent terms, as are DSC and the F1 Score.Metric terminology varies with context and community, including image-level classification and semantic segmentation.
  • 1.1 Image-level Classification: Image-level classification thresholds predicted class scores to assign positive or negative outcomes, after which counting and multi-threshold metrics can be computed.Calibration metrics additionally assess whether predicted confidence scores reflect true outcome probabilities.
  • 1.2 Semantic Segmentation: Semantic segmentation assigns labels to pixels and commonly uses overlap-based metrics, complemented by boundary-focused distance-based metrics.Biomedical applications may also require problem-specific volume metrics such as Absolute or Relative Volume Error.
  • 1.3 Object Detection: Object-detection validation requires object representation, a localization criterion, assignment resolution, and metric computation.Localization determines TP, FP, and FN outcomes, while TNs are not defined for object detection tasks.
  • 1.3 Object Detection: Object-detection metrics include fixed-threshold counting metrics and range-based multi-threshold metrics such as Average Precision (AP) and Free-Response Receiver Operating Characteristic (FROC) Score.The assignment step resolves ambiguities that can arise when multiple predictions or references are matched.
  • 1.4 Instance Segmentation: Instance segmentation distinguishes instances of the same class, combining object-level detection with pixel-level correspondence for validation.It therefore shares characteristics with both object detection and semantic segmentation.

Localization criteria:

Localization criteria include Boundary IoU, Mask IoU, and IoR, while PQ can jointly assess detection and segmentation. Instance segmentation is often formulated as semantic segmentation followed by post-processing such as connected component analysis.

  • Localization criteria:: Boundary IoU, Mask IoU, and Intersection over Reference (IoR) are listed as localization metrics.The cited figures are Fig. SN 3.73 for Boundary IoU, Fig. SN 3.76 for Mask IoU, and Fig SN 3.75 for IoR.
  • Localization criteria:: PQ can assess detection and segmentation performance simultaneously in a single score.PQ is presented as an additional counting metric for joint assessment.
  • Localization criteria:: Instance segmentation problems are often phrased as semantic segmentation problems with an additional post-processing step.Connected component analysis is given as an example of this post-processing.

SUPPL. NOTE 2 · METRIC PITFALLS

The note organizes image-analysis metric pitfalls around inadequate problem-category choice, poor metric selection, and poor metric application. Its examples show that metric validity depends on the property of interest, data and algorithm outputs, implementation choices, aggregation, ranking, reporting, and interpretation.

  • 2.1 Pitfalls related to an inadequate choice of the problem category: Metrics applied at the image level can mislead object-detection validation because aggregation discards localization and object-count information.Image-level AUROC-based validation does not measure localization or distinguish errors in object matching.
  • 2.1 Pitfalls related to an inadequate choice of the problem category: Overlap-based segmentation metrics can misrepresent domain interests such as volumetric ratios when the validation goal is not structure overlap.Using DSC for volume-ratio accuracy may be misleading when no common metric directly captures the property of interest.
  • 2.2 Pitfalls related to poor metric selection: Metric selection must match the domain interest because volume-, boundary-, overlap-, center-, and calibration-based metrics capture different properties and can overlook relevant errors.Volume-based metrics miss mislocalization, boundary-based metrics can miss holes, overlap-based metrics can miss center alignment, and common calibration metrics can imply perfect calibration incorrectly.
  • 2.2.3 Pitfalls related to disregard of the properties of the data set and algorithm output: Data and algorithm-output properties can invalidate metric interpretation, including small-sample AUROC instability, inter-rater variability, empty images, overlapping predictions, and missing class scores.The note reports six samples per data set for unstable AUROC comparisons, NSD with τ=1 capturing annotation variability, division by zero for empty images, perfect semantic-segmentation scores hiding instance errors, and less interpretable multi-threshold scores without class scores.
  • 2.3.1 Pitfalls related to inadequate metric implementation: Metric implementations are not standardized, so FPPI ranges, calibration bins, IoU thresholds, and global class thresholds can substantially change reported scores and comparability.FROC scores differ when FPPI ranges are [0, 1], [0, 2], or [0, 4], while ECE and MCE scores are substantially affected by the number of bins.
  • 2.3.3–2.3.4 Pitfalls related to inadequate ranking scheme and metric reporting: Ranking and reporting can overstate differences or hide uncertainty because related metrics add redundant information, similar performance can yield identical ranks, and fixed seeds do not guarantee reproducibility.DSC and IoU typically produce the same ranking, while HD may differ; ranking tables conceal uncertainty, and identical training conditions can produce different results even with fixed seeds.
  • 2.3.5 Pitfalls related to inadequate interpretation of metric values: Metric values require context because grid resolution, unattainable theoretical bounds, and numerically different but irrelevant scores can distort interpretation and rankings.The MCC theoretical minimum of -1 may be unattainable in a multi-class setting, and extremely small aggregate differences can assign different ranks despite lacking biomedical relevance.

SUPPL. NOTE 3 · METRIC PROFILES

The metric profiles organize metrics relevant to image-analysis validation by describing their formulas, value ranges, interpretation, characteristics, and relevant pitfalls. They cover counting, multi-threshold, distance-based, calibration, and localization metrics, with examples spanning common biomedical image-analysis tasks.

  • METRIC PROFILES: Each profile provides a metric’s description, formula, value range, direction of better performance, confusion-matrix cardinalities, prevalence dependency, and relevant pitfalls.The profiles are intended as a common, structured reference for metric-related information.
  • METRIC PROFILES: The profiles use confusion-matrix concepts for binary and C-class settings and may include a positive cost matrix with weights w_ij > 0.The binary example illustrates cardinalities for predictions of triangles and circles.
  • 3.1.1 Counting metrics.: Counting metrics include Accuracy, Balanced Accuracy, clDice, DSC, Expected Cost, Fβ Score, FPPI, IoU, MCC, Net Benefit, NPV, PQ, LR+, PPV, Sensitivity, Specificity, and WCK.The profiles indicate whether higher or lower values are preferable; Expected Cost is lower-is-better, while the other listed profiles use upward value-range arrows.
  • 3.1.2 Multi-threshold metrics.: Multi-threshold evaluation is represented by AUROC, Average Precision, and FROC, with FROC linked to False Positives per Image.These profiles use upward arrows to indicate that higher values are better than lower values.
  • 3.1.3 Distance-based metrics.: Distance-based profiles include ASSD, Boundary IoU, HD, MASD, NSD, and the Xth Percentile of HD, combining upward- and downward-oriented metrics.ASSD, HD, MASD, and the Xth Percentile of HD are lower-is-better, whereas Boundary IoU and NSD are higher-is-better.
  • 3.2 Calibration metrics: Calibration metrics comprise Brier Score, CWCE, ECE, ECEKDE, KCE, NLL, and RBS, generally evaluating lower values as better.The KCE profile does not state a value-range direction in the supplied passage.
  • 3.3 Localization criteria: Localization criteria include Boundary IoU, Center Distance, IoR, Mask/Box/Approx Intersection over Union, Mask IoU > 0, and Point inside Mask/Box/Approximation.Boundary IoU, IoR, and the intersection-over-union criteria are higher-is-better when stated, while Center Distance is lower-is-better.

ACRONYMS

This section defines abbreviations used throughout the paper, including artificial intelligence, biomedical image analysis, validation metrics, and segmentation terminology.

  • Acronym definitions: The acronym list includes AI for artificial intelligence, BIAS for Biomedical Image Analysis Challenges, and metrics such as AP, ASSD, AUC, AUROC, BA, DSC, and clDice.It also defines calibration and agreement measures including BS, BSS, CK, CWCE, and WCK.
  • Acronym definitions: The section defines ROC as Receiver Operating Characteristic, SemS as Semantic Segmentation, and abbreviations for true negatives and positives, including TN, TNR, TP, and TPR.WCK is defined as Weighted Cohen’s Kappa.
Loading 2302.01790v4…