Source-linked AI summary
Metrics reloaded: Recommendations for image analysis validation
Lena Maier-Hein, Annika Reinke, Patrick Godau, Minu D. Tizabi, Florian Buettner, Evangelia Christodoulou, Ben Glocker, Fabian Isensee, Jens Kleesiek, Michal Kozubek, Mauricio Reyes, Michael A. Riegler, Manuel Wiesenfarth, A. Emre Kavur, Carole H. Sudre, Michael Baumgartner, Matthias Eisenmann, Doreen Heckmann-Nötzel, Tim Rädsch, Laura Acion, Michela Antonelli, Tal Arbel, Spyridon Bakas, Arriel Benis, Matthew Blaschko, M. Jorge Cardoso, Veronika Cheplygina, Beth A. Cimini, Gary S. Collins, Keyvan Farahani, Luciana Ferrer, Adrian Galdran, Bram van Ginneken, Robert Haase, Daniel A. Hashimoto, Michael M. Hoffman, Merel Huisman, Pierre Jannin, Charles E. Kahn, Dagmar Kainmueller, Bernhard Kainz, Alexandros Karargyris, Alan Karthikesalingam, Hannes Kenngott, Florian Kofler, Annette Kopp-Schneider, Anna Kreshuk, Tahsin Kurc, Bennett A. Landman, Geert Litjens, Amin Madani, Klaus Maier-Hein, Anne L. Martel, Peter Mattson, Erik Meijering, Bjoern Menze, Karel G. M. Moons, Henning Müller, Brennan Nichyporuk, Felix Nickel, Jens Petersen, Nasir Rajpoot, Nicola Rieke, Julio Saez-Rodriguez, Clara I. Sánchez, Shravya Shetty, Maarten van Smeden, Ronald M. Summers, Abdel A. Taha, Aleksei Tiulpin, Sotirios A. Tsaftaris, Ben Van Calster, Gaël Varoquaux, Paul F. Jäger
TL;DR
Biomedical image-analysis validation often uses metrics that do not reflect domain interests. Metrics Reloaded provides problem-aware metric-selection guidance grounded in expert consensus, with 93% median agreement across its subprocesses. The framework offers systematic guidance across imaging tasks while noting that calibration metrics may require re-validation for new cohorts.
Problem
Biomedical image-analysis algorithms lack reliable, objective performance assessment because commonly used validation metrics may not reflect domain interests.
Method
Metrics Reloaded uses a multi-stage Delphi process and problem fingerprinting to guide metric selection across biomedical image-analysis tasks.
Results
93% median agreement was achieved across the framework’s Delphi subprocesses.
Takeaways & Limitations
The framework provides systematic, problem-aware guidance for choosing validation metrics across different biomedical imaging tasks.
Takeaways & Limitations
Calibration metrics are prevalence-dependent, so calibration quality may require re-validation for each new study cohort.
Abstract
from arXiv · showhide
Increasing evidence shows that flaws in machine learning (ML) algorithm validation are an underestimated global problem. Particularly in automatic biomedical image analysis, chosen performance metrics often do not reflect the domain interest, thus failing to adequately measure scientific progress and hindering translation of ML techniques into practice. To overcome this, our large international expert consortium created Metrics Reloaded, a comprehensive framework guiding researchers in the problem-aware selection of metrics. Following the convergence of ML methodology across application domains, Metrics Reloaded fosters the convergence of validation methodology. The framework was developed in a multi-stage Delphi process and is based on the novel concept of a problem fingerprint - a structured representation of the given problem that captures all aspects that are relevant for metric selection, from the domain interest to the properties of the target structure(s), data set and algorithm output. Based on the problem fingerprint, users are guided through the process of choosing and applying appropriate validation metrics while being made aware of potential pitfalls. Metrics Reloaded targets image analysis problems that can be interpreted as a classification task at image, object or pixel level, namely image-level classification, object detection, semantic segmentation, and instance segmentation tasks. To improve the user experience, we implemented the framework in the Metrics Reloaded online tool, which also provides a point of access to explore weaknesses, strengths and specific recommendations for the most common validation metrics. The broad applicability of our framework across domains is demonstrated by an instantiation for various biological and medical image analysis use cases.
MAIN
Metrics Reloaded is a consensus-based framework for problem-aware selection of validation metrics in biomedical image analysis, addressing unreliable assessment and supporting scientific progress and translation into practice. It combines problem fingerprinting, recommendations across four image-analysis task categories, biomedical use cases, and an open online tool.
- MAIN: The framework uses problem fingerprinting to represent domain, target-structure, dataset, and algorithm-output properties relevant to metric selection.It was developed through a multi-stage Delphi process for consensus building and educated metric-choice decisions.
- MAIN: Metrics Reloaded includes Metric Cheat Sheets and an open-source MONAI implementation, alongside an online tool and applications to common biomedical use cases.These components were introduced to improve user experience and provide metric-specific recommendations and implementations.
- MAIN: The common framework treats classification, detection, and segmentation as related classification tasks at different scales while warning that category similarities can cause incorrect task selection.It addresses image-level classification, object detection, semantic segmentation, and instance segmentation in one framework.
- MAIN: The final Delphi recommendations achieved the required >75% consensus threshold across all ten core components, with disagreement ranging from 0% to 7%.The recommendations are represented in Fig. 2 and Subprocesses S1-S9, and were designed to address metric pitfalls identified in related work.
- MAIN: The framework’s use cases show that shared problem properties can yield nearly identical metric recommendations across domains, including modality-independent recommendations for semantic segmentation.For segmentation, target-object size relative to the image grid matters more than the specific image modality.
- MAIN: Metrics Reloaded provides guidelines and tools for choosing performance metrics in a problem-aware manner, developed by an international consortium of more than 70 experts.The consortium spanned biomedical image analysis, machine learning, statistics, epidemiology, biology, and medicine.
ACRONYMS
The framework combines expert Delphi consensus with problem fingerprinting to guide metric selection across image-analysis tasks. It also highlights important metric-specific limitations, especially for calibration, dense instance segmentation, and imbalanced classification.
- Delphi process: The Metrics Reloaded recommendations were developed by an international expert consortium through a multi-stage Delphi process, with strong final consensus on the framework’s core components.The final recommendation received strong support after two rounds of revisions, while the other nine components achieved consensus in the first round.
- 1.3 Generation of the problem fingerprint: Problem fingerprinting structures the properties relevant to metric selection into user-instantiated binary or categorical items.Fingerprint items are referenced using the notation FPX.Y.
- 2.2 Recommendations for Image-level Classification: For image-level classification, calibration metrics may be selected alongside discrimination metrics when calibration of predicted class scores is part of the evaluation goal.The recommendation framework distinguishes calibration assessment from discrimination capabilities.
- 2.4 Recommendations for Object detection: Object detection evaluates localization and category-specific instance matching by comparing reference objects with predicted objects.True positives are matched predictions, whereas false positives are predictions without an associated reference object.
- 2.5 Recommendations for Instance segmentation: Standard reference-based metrics can fail for instance segmentation images with extremely dense, complex-shaped structures because overlap may not establish unique correspondences.Specialized metrics that do not rely on one-to-one correspondences may be needed in such cases.
- 2.6 Recommendations for Calibration of Predicted Class Scores: Calibration metrics generally depend on prevalence, so calibration quality may need re-validation on each new study cohort when cohort prevalence differs from the target population.The framework also distinguishes calibration of score behavior from overall performance measurement and warns that predicted scores should not automatically be interpreted as true posterior probabilities.
- DG2.1: Weighted Cohen’s Kappa (WCK) versus Expected Cost (EC): Weighted Cohen’s Kappa and Expected Cost, including normalized Expected Cost, measure disagreement differently despite their conceptual similarity.Normalized Expected Cost is the comparable random-performance-normalized counterpart to Weighted Cohen’s Kappa.
- DG2.3: Balanced Accuracy (BA) versus Matthews Correlation Coefficient (MCC) versus normalized EC (ECN): BA, MCC, and ECN can rate the same classifier as near-perfect, fairly good, and random/naive, respectively, illustrating metric-dependent assessments.The example reports BA = 0.99, MCC = 0.7, and ECN = 1; BA can also remain near-perfect despite a PPV of 0.09 because it does not consider predictive values.
Prediction
The prediction guidance shows that metric choice must reflect dataset size, operating-range definitions, calibration goals, and whether interest centers on top-label decisions or all predicted scores. It contrasts AP and FROC for detection evaluation and compares calibration metrics’ interpretability, bias, and configuration requirements.
- Prediction metrics: FROC accounts for the number of images through False Positives per Image, whereas AP gives identical scores for data sets D1 and D2.For D2, the lower FPPI yields a higher FROC score.
- Prediction metrics: FROC scores change when different FPPI ranges define the curve’s x-axis, so the selected boundaries affect evaluation.
- Calibration: For comparing calibration across classifiers, KCE is unbiased but difficult to interpret and configure, whereas ECEKDE is interpretable and straightforward to configure but inherently biased.Both estimate canonical calibration error using different distance functions; KCE uses maximum mean discrepancy, while ECEKDE uses the ℓp norm.
- Calibration: For comparing re-calibration methods, Brier Score is attractive when calibration preserves accuracy, while KCE enables relative comparison but requires nontrivial kernel configuration.If re-calibration changes discrimination, Brier Score conflates calibration with discrimination and must be applied with care.
- Calibration scope: Top-label calibration is appropriate when the research question focuses on classifier decisions, whereas all-score calibration may better reflect the canonical calibration condition and clinical interest in other probabilities.
2.7.7 Decision guide S8.
Decision guide S8 compares Mask IoU, Boundary IoU, and IoR for instance segmentation, emphasizing their differing treatment of overlap, boundaries, small structures, and touching reference objects. A custom localization criterion based on the target segmentation metric may also be appropriate, but some metrics make cutoff selection challenging.
- Boundary and overlap criteria: Mask IoU can over-penalize small structures and predictions when reference objects frequently touch, while Boundary IoU is preferable for boundary-focused assessment.The small-structure limitation is especially relevant when structure sizes vary substantially; touching objects can create non-split errors that IoU penalizes heavily.
- Custom localization criterion: Localization in instance segmentation may use a criterion aligned with the target segmentation metric, although metrics without fixed upper bounds such as HD make adequate cutoffs difficult.For example, an NSD target metric could define the localization criterion accordingly.
- Boundary and overlap criteria: Mask IoU measures general structural overlap, whereas Boundary IoU focuses on boundary correctness but can produce a perfect value of 1.0 for imperfect predictions.Boundary-focused evaluation therefore introduces a specific failure mode in which imperfect boundaries are not penalized.
- IoR: IoR reduces penalties for non-split errors by measuring reference-object area covered by a prediction and allowing multiple true-positive matches to the same prediction.IoR is mainly used in dense cell-segmentation images, but it can be deceived by large predictions and shares Mask IoU’s behavior for boundaries and small structures.
Large structure · Small structure
Boundary IoU improves on Mask IoU by penalizing boundary errors and remaining more invariant to structure size across large and small structures.
- Large structure: Boundary IoU specifically penalizes boundary errors compared with Mask IoU.
- Small structure: Boundary IoU is more invariant to structure sizes than Mask IoU.
- Small structure: The comparison evaluates Boundary IoU using two different thresholds.
- Large structure: The size-invariance comparison includes large structures.
- Small structure: The size-invariance comparison includes small structures.
- Large structure: Boundary IoU is presented in the third and fourth columns of the figure.
Overlapping boundary pixels
Overlapping boundary pixels can make Boundary IoU appear perfect despite an imperfect prediction, while localization criteria and thresholds introduce additional ambiguity and task-dependent trade-offs.
- Overlapping boundary pixels: Boundary IoU can equal 1.00 for a prediction with a central hole when the distance-to-border region contains all mask pixels, whereas Mask IoU detects the defect.The example uses distance = 2.
- Localization criteria: Using IoU > 0 as a detection criterion can accept very large predictions, making the predicted localization ambiguous.This loose criterion is recommended only when reference annotations provide exact outlines.
- Localization thresholds: Localization thresholds should reflect the task: lower thresholds suit existence-focused or small, variable, three-dimensional, or uncertain-reference structures, while higher thresholds suit precise localization.Metrics are commonly averaged over multiple cutoff values, with IoU defaults from 0.5 to 0.9 in steps of 0.05, but problem properties can limit cutoff relevance.
3.1 Metrics Cheat Sheets
This section presents Metrics Reloaded cheat sheets that describe relevant metrics, including their formulas, value ranges, characteristics, and recommendations. Many metrics are based on the confusion matrix, illustrated for binary and multiclass settings.
- Cheat sheet framework: The cheat sheets summarize each metric’s formula, value range, applicable problem categories, prevalence dependence, and recommended use.They also indicate whether higher or lower values are preferable.
- Counting metrics: Balanced Accuracy is recommended as a multi-class counting metric in Subprocess S2.
- Overlap metrics: Centerline Dice is recommended as an overlap-based metric in Subprocess S6.
- Overlap metrics: Intersection over Union is recommended as an overlap-based metric in Subprocess S6.
- Counting metrics: Positive Likelihood Ratio is recommended as a per-class counting metric in Subprocess S3.
4.1 Image-level classification
The framework was instantiated for biomedical image-level classification, covering seven concrete use cases from sperm motility and disease classification to cardiac disease classification. Resulting recommendations are presented in Fig. SN 4.1, with detailed metric-selection guidance in Figs. SN 4.5–SN 4.7.
- 4.1 Image-level classification: Fig. SN 4.1 presents the resulting metric recommendations for the instantiated image-level classification problems.The framework’s recommendations are also detailed for metric-selection Subprocesses S2–S5 in Figs. SN 4.5–SN 4.7.
- 4.1 Image-level classification: The instantiations demonstrate the framework’s application across diverse biomedical image-level classification problems.The examples include video, dermoscopic, cellular, ultrasound, MRI, and mammography images.
- 4.1 Image-level classification: Seven biomedical image-level classification use cases span microscopy, dermoscopy, cell-state, ultrasound, MRI, multiple sclerosis, mammography, and cardiac disease applications.These include sperm motility, dermoscopic disease [33], autophagy-stage, ultrasound plane [11], multiple-sclerosis lesion [79], breast-cancer, and cardiac-disease classification.
4.2 Semantic segmentation
The framework was instantiated for five biomedical semantic segmentation use cases, with metric recommendations provided in Fig. SN 4.2 and detailed recommendations for subprocesses S6 and S7 in Figs. SN 4.10–SN 4.11.
- 4.2 Semantic segmentation: Metric recommendations for these problems are presented in Fig. SN 4.2, with detailed recommendations for metric-selection subprocesses S6 and S7 in Figs. SN 4.10–SN 4.11.
- 4.2 Semantic segmentation: Five semantic segmentation use cases span embryo microscopy, liver CT, breast WSI lesion labeling, cortical 3D MRI structures, and aneurysm TOF-MRA segmentation.These instantiations are identified as SemS-1 through SemS-5.
4.3 Object detection
The framework was instantiated for biomedical object-detection problems, producing use-case-specific metric recommendations. These recommendations cover diverse applications, including cell, lesion, polyp, mitosis, and lung-nodule detection.
- Object detection: The framework provides metric recommendations for concrete biomedical object-detection use cases, with an overview in Fig. SN 4.3 and detailed guidance in Figs. SN 4.6–SN 4.9.The detailed figures address metric-selection Subprocesses S3–S4 and S8–S9.
- Object detection: The instantiated cases span cell detection and tracking in time-lapse microscopy, MS lesion detection in multimodal brain MRI, and polyp detection with predefined sensitivity of 0.95.The use cases are identified as ObD-1, ObD-2, and ObD-3; the associated studies are cited as, [79], and.
- Object detection: Additional applications include mitosis detection in histopathology images and lung-nodule detection in CT images.The cited studies are [8] for mitosis detection and for lung-nodule detection.
4.4 Instance segmentation
The framework was instantiated for four biomedical instance segmentation problems spanning microscopy, colonoscopy, cell tracking, and brain MRI. Resulting metric recommendations are presented in Extended Data Fig. SN 4.4, with detailed subprocess guidance in additional figures.
- Metric selection: Detailed recommendations for the use cases are provided for metric-selection Subprocesses S3–S4 and S6–S9 in Extended Data Figs. SN 4.6–SN 4.11.
- Use cases: Four instance segmentation use cases cover fruit-fly neuron imaging, colonoscopy instrument segmentation [98], cell-nuclei tracking, and multiple-sclerosis lesion segmentation [79].
- Framework instantiation: Extended Data Fig. SN 4.4 instantiates the framework with metric recommendations for these concrete biomedical instance segmentation problems.
5.1 Symbol References
This section provides an overview of the symbols used in the process diagrams, whose notation originates from Business Process Model and Notation (BPMN).
- 5.1 Symbol References: The process diagrams use symbols based on Business Process Model and Notation (BPMN).Extended Data Fig. SN 5.1 gives an overview of the symbols used in these diagrams.
5.2 Expected formats of reference and algorithm output
The metric mapping defines expected reference and algorithm-output formats for image-level classification, semantic segmentation, object detection, and instance segmentation. Inputs generally pair class labels with scores, locations, or pixel maps, while requiring compatible coordinate systems and boundary representations where applicable.
- Image-level Classification: Image-level classification uses per-image class labels or multi-label indicators, with algorithm outputs matching this format or providing predicted class scores.For C classes, references are either y_I ∈ {1, ..., C} or y_I ∈ {0, 1}^C; score-based outputs are expected when available.
- Semantic Segmentation: Semantic segmentation requires reference and predicted labels for each pixel in the same coordinate system with identical spacing.Each pixel may have a single class assignment or per-class multi-label indicators, with corresponding prediction formats.
- Object Detection: Object detection represents each object as a class-and-location tuple, with predictions optionally including a class score in [0, 1].Location information may be a box, center point, radius, or another supported representation.
- Instance Segmentation: Instance segmentation represents each object with a class and image-sized binary pixel map, optionally adding a predicted class score.Reference and prediction boundaries should be supplied as separate boundary-pixel lists for each instance.
- Instance Segmentation: Semantic annotations can be converted to instance format by connected-component analysis when instances are non-touching and connected.When references deviate from the expected format, matching may also be achieved by measures such as aggregating pixel-level references to image-level references.
5.3 Acronyms
This section defines the abbreviations used throughout the paper, spanning artificial intelligence, imaging, calibration, segmentation, detection, and validation metrics.
- Acronyms: The glossary defines core technical abbreviations, including AI, ML, CT, MRI, DSC, AUC, AUROC, AP, ASSD, and calibration metrics such as ECE and Brier Score.It also includes task- and performance-related terms such as semantic and instance-analysis metrics, expected cost, Cohen’s Kappa, and centerline Dice Similarity Coefficient.
- Acronyms: Additional abbreviations cover machine-learning tools, clinical concepts, detection and segmentation metrics, decision analysis, and statistical measures, including MONAI, MS, ObD, PQ, NB, NLL, PPV, ROC, and RI.The list also defines PHI, PR, PSR, RBS, NSD, NPV, and the observed-to-expected ratio.
5.4 Glossary
The glossary defines core image-analysis validation terms, including bounding boxes, calibration plots, metrics, and training/test cases. It also clarifies specialized metric notation and related structural terminology.
- Definitions: A bounding box is typically the smallest rectangle completely surrounding an object to be detected, while a calibration plot visualizes deviations from perfect calibration.Calibration plots diagnose whether model outputs are generally overconfident or underconfident.
- Definitions: Metrics quantify and validate algorithm performance according to domain-specific validation goals and properties of interest, with distinctions such as reference-based and non-reference-based metrics.Metrics can also be grouped into families according to their mathematical properties.
- Definitions: Training and test cases are datasets used for algorithm development and validation, where a case provides the data needed to produce one result.Training cases include reference annotations and are used for algorithm training; test cases support validation.