Source-linked AI summary
Metrics reloaded: Recommendations for image analysis validation
Lena Maier-Hein, Annika Reinke, Patrick Godau, Minu D. Tizabi, Florian Buettner, Evangelia Christodoulou, Ben Glocker, Fabian Isensee, Jens Kleesiek, Michal Kozubek, Mauricio Reyes, Michael A. Riegler, Manuel Wiesenfarth, A. Emre Kavur, Carole H. Sudre, Michael Baumgartner, Matthias Eisenmann, Doreen Heckmann-Nötzel, Tim Rädsch, Laura Acion, Michela Antonelli, Tal Arbel, Spyridon Bakas, Arriel Benis, Matthew Blaschko, M. Jorge Cardoso, Veronika Cheplygina, Beth A. Cimini, Gary S. Collins, Keyvan Farahani, Luciana Ferrer, Adrian Galdran, Bram van Ginneken, Robert Haase, Daniel A. Hashimoto, Michael M. Hoffman, Merel Huisman, Pierre Jannin, Charles E. Kahn, Dagmar Kainmueller, Bernhard Kainz, Alexandros Karargyris, Alan Karthikesalingam, Hannes Kenngott, Florian Kofler, Annette Kopp-Schneider, Anna Kreshuk, Tahsin Kurc, Bennett A. Landman, Geert Litjens, Amin Madani, Klaus Maier-Hein, Anne L. Martel, Peter Mattson, Erik Meijering, Bjoern Menze, Karel G. M. Moons, Henning Müller, Brennan Nichyporuk, Felix Nickel, Jens Petersen, Nasir Rajpoot, Nicola Rieke, Julio Saez-Rodriguez, Clara I. Sánchez, Shravya Shetty, Maarten van Smeden, Ronald M. Summers, Abdel A. Taha, Aleksei Tiulpin, Sotirios A. Tsaftaris, Ben Van Calster, Gaël Varoquaux, Paul F. Jäger
TL;DR
Biomedical image-analysis validation에서는 domain interest를 반영하지 못하는 metric이 흔히 사용된다. Metrics Reloaded는 전문가 합의에 기반해 problem-aware metric-selection guidance를 제공하며, subprocess 전반에서 median agreement 93%를 보였다. 이 framework는 imaging task 전반에 체계적인 guidance를 제공하지만, calibration metric은 새로운 cohort에서 재검증이 필요할 수 있음을 지적한다.
문제
Biomedical image-analysis algorithm은 흔히 사용되는 validation metric이 domain interest를 반영하지 못할 수 있어 신뢰할 수 있고 객관적인 performance assessment가 어렵다.
방법
Metrics Reloaded는 다단계 Delphi process와 problem fingerprinting을 사용해 biomedical image-analysis task 전반에서 metric selection을 안내한다.
결과
Framework의 Delphi subprocess 전반에서 median agreement 93%를 달성했다.
시사점 및 한계
이 framework는 서로 다른 biomedical imaging task에서 validation metric을 선택할 수 있도록 체계적이고 problem-aware한 guidance를 제공한다.
시사점 및 한계
Calibration metric은 prevalence에 의존하므로, 새로운 study cohort마다 calibration quality를 재검증해야 할 수 있다.
Abstract
from arXiv · showhide
Increasing evidence shows that flaws in machine learning (ML) algorithm validation are an underestimated global problem. Particularly in automatic biomedical image analysis, chosen performance metrics often do not reflect the domain interest, thus failing to adequately measure scientific progress and hindering translation of ML techniques into practice. To overcome this, our large international expert consortium created Metrics Reloaded, a comprehensive framework guiding researchers in the problem-aware selection of metrics. Following the convergence of ML methodology across application domains, Metrics Reloaded fosters the convergence of validation methodology. The framework was developed in a multi-stage Delphi process and is based on the novel concept of a problem fingerprint - a structured representation of the given problem that captures all aspects that are relevant for metric selection, from the domain interest to the properties of the target structure(s), data set and algorithm output. Based on the problem fingerprint, users are guided through the process of choosing and applying appropriate validation metrics while being made aware of potential pitfalls. Metrics Reloaded targets image analysis problems that can be interpreted as a classification task at image, object or pixel level, namely image-level classification, object detection, semantic segmentation, and instance segmentation tasks. To improve the user experience, we implemented the framework in the Metrics Reloaded online tool, which also provides a point of access to explore weaknesses, strengths and specific recommendations for the most common validation metrics. The broad applicability of our framework across domains is demonstrated by an instantiation for various biological and medical image analysis use cases.
주요 내용
Metrics Reloaded는 biomedical image analysis에서 문제를 고려한 validation metric 선택을 위한 합의 기반 framework로, 신뢰하기 어려운 평가에 대응하고 과학적 진보와 실제 적용을 지원한다. 이는 problem fingerprinting, 네 가지 image-analysis task 범주 전반의 권고안, biomedical use case, 공개 online tool을 결합한다.
- 주요 내용: 이 framework는 metric 선택과 관련된 domain, target structure, dataset, algorithm output의 특성을 나타내기 위해 problem fingerprinting을 사용한다.합의 형성과 근거에 기반한 metric 선택을 위해 다단계 Delphi process를 통해 개발됐다.
- 주요 내용: Metrics Reloaded에는 Metric Cheat Sheets와 open-source MONAI implementation이 포함되며, online tool 및 일반적인 biomedical use case 적용도 제공한다.이 구성 요소들은 사용자 경험을 개선하고 metric별 권고안과 implementation을 제공하기 위해 도입됐다.
- 주요 내용: 이 공통 framework는 classification, detection, segmentation을 서로 다른 규모에서 수행되는 연관된 classification task로 다루지만, 범주 간 유사성이 잘못된 task 선택을 초래할 수 있음을 경고한다.하나의 framework에서 image-level classification, object detection, semantic segmentation, instance segmentation을 다룬다.
- 주요 내용: 최종 Delphi 권고안은 열 개의 핵심 구성 요소 모두에서 필요한 >75% consensus threshold를 달성했으며, 불일치율은 0%에서 7%까지였다.권고안은 Fig. 2와 Subprocesses S1-S9에 제시됐으며, 관련 연구에서 확인된 metric pitfalls를 다루도록 설계됐다.
- 주요 내용: 이 framework의 use case는 공통된 문제 특성이 domain 전반에서 거의 동일한 metric 권고안을 도출할 수 있음을 보여주며, semantic segmentation에 대해서는 modality-independent 권고안도 제시한다.segmentation에서는 특정 image modality보다 image grid에 대한 target object 크기의 상대적 비율이 더 중요하다.
- 주요 내용: Metrics Reloaded는 70명 이상의 전문가로 구성된 international consortium이 개발한 것으로, 문제를 고려한 방식으로 performance metric을 선택하기 위한 guideline과 tool을 제공한다.이 consortium에는 biomedical image analysis, machine learning, statistics, epidemiology, biology, medicine 분야의 전문가가 참여했다.
약어
이 프레임워크는 expert Delphi 합의와 problem fingerprinting을 결합해 image-analysis task 전반에서 metric 선택을 안내한다. 또한 calibration, dense instance segmentation, imbalanced classification에서 특히 중요한 metric별 한계를 강조한다.
- Delphi process: Metrics Reloaded 권고안은 국제 expert consortium이 multi-stage Delphi process를 통해 개발했으며, 프레임워크의 핵심 구성요소에 대해 최종 합의가 강하게 이루어졌다.최종 권고안은 두 차례의 수정 후 강한 지지를 받았고, 나머지 9개 구성요소는 첫 번째 라운드에서 합의에 도달했다.
- 1.3 문제 fingerprint 생성: Problem fingerprinting은 metric 선택과 관련된 속성을 사용자가 인스턴스화하는 binary 또는 categorical item으로 구조화한다.Fingerprint item은 FPX.Y 표기법을 사용해 참조한다.
- 2.2 image-level classification 권고: image-level classification에서는 예측 class score의 calibration이 평가 목표에 포함될 때 discrimination metric과 함께 calibration metric을 선택할 수 있다.권고 프레임워크는 calibration assessment와 discrimination capability를 구분한다.
- 2.4 object detection 권고: Object detection은 reference object와 predicted object를 비교해 localization과 category-specific instance matching을 평가한다.True positive는 매칭된 prediction이고, false positive는 연관된 reference object가 없는 prediction이다.
- 2.5 instance segmentation 권고: 표준 reference-based metric은 구조가 extremely dense하고 complex-shaped인 instance segmentation image에서 실패할 수 있다. overlap만으로는 고유한 대응관계를 확립할 수 없기 때문이다.이러한 경우 one-to-one correspondence에 의존하지 않는 specialized metric이 필요할 수 있다.
- 2.6 predicted class score calibration 권고: Calibration metric은 일반적으로 prevalence에 의존하므로, cohort prevalence가 target population과 다를 때는 새로운 study cohort마다 calibration quality를 재검증해야 할 수 있다.이 프레임워크는 score behavior의 calibration과 overall performance measurement도 구분하며, predicted score를 true posterior probability로 자동 해석해서는 안 된다고 경고한다.
- DG2.1: Weighted Cohen’s Kappa (WCK) versus Expected Cost (EC): Weighted Cohen’s Kappa와 normalized Expected Cost를 포함한 Expected Cost는 개념적으로 유사하지만 disagreement를 서로 다르게 측정한다.Normalized Expected Cost는 Weighted Cohen’s Kappa에 대응하는, random-performance-normalized counterpart다.
- DG2.3: Balanced Accuracy (BA) versus Matthews Correlation Coefficient (MCC) versus normalized EC (ECN): BA, MCC, ECN은 동일한 classifier를 각각 near-perfect, fairly good, random/naive로 평가할 수 있어 metric-dependent assessment를 보여준다.이 예에서는 BA = 0.99, MCC = 0.7, ECN = 1로 보고되었다. BA는 predictive value를 고려하지 않으므로 PPV가 0.09여도 near-perfect로 유지될 수 있다.
예측
예측 지침은 metric 선택이 데이터셋 크기, operating range 정의, calibration 목표, 그리고 관심이 최고 label 결정에 있는지 또는 모든 예측 score에 있는지를 반영해야 함을 보여준다. Detection 평가에서 AP와 FROC를 대조하고, calibration metric의 해석 가능성, bias, configuration 요구사항을 비교한다.
- 예측 metric: FROC는 Image당 False Positives를 통해 이미지 수를 반영하는 반면, AP는 데이터셋 D1과 D2에 동일한 score를 부여한다.D2에서는 더 낮은 FPPI가 더 높은 FROC score를 산출한다.
- 예측 metric: 곡선의 x축을 정의하는 FPPI 범위가 달라지면 FROC score도 변하므로, 선택한 경계가 평가에 영향을 준다.
- Calibration: Classifier 간 calibration을 비교할 때 KCE는 unbiased하지만 해석과 configuration이 어렵고, ECEKDE는 해석 가능하고 configuration이 간단하지만 본질적으로 biased하다.둘 다 서로 다른 distance function을 사용해 canonical calibration error를 추정한다. KCE는 maximum mean discrepancy를 사용하고, ECEKDE는 ℓp norm을 사용한다.
- Calibration: Calibration이 accuracy를 보존할 때에는 Brier Score가 재-calibration 방법 비교에 매력적이며, KCE는 상대적 비교를 가능하게 하지만 상당한 kernel configuration이 필요하다.재-calibration이 discrimination을 변화시키면 Brier Score는 calibration과 discrimination을 혼동하므로 주의해서 적용해야 한다.
- Calibration 범위: 연구 질문이 classifier 결정에 초점을 둘 때는 top-label calibration이 적절한 반면, all-score calibration은 canonical calibration condition과 다른 probability에 대한 임상적 관심을 더 잘 반영할 수 있다.
2.7.7 Decision guide S8.
Decision guide S8은 instance segmentation에서 Mask IoU, Boundary IoU, IoR을 비교하며, overlap, boundaries, small structures, touching reference objects를 서로 다르게 처리하는 방식을 강조한다. Target segmentation metric에 기반한 custom localization criterion도 적절할 수 있지만, 일부 metric은 cutoff 선택을 어렵게 한다.
- Boundary 및 overlap criteria: Mask IoU는 reference objects가 자주 touching할 때 small structures와 predictions를 지나치게 과도하게 벌점화할 수 있는 반면, Boundary IoU는 boundary-focused assessment에 더 적합하다.이 small-structure limitation은 structure sizes가 크게 다를 때 특히 중요하며, touching objects는 IoU가 크게 벌점화하는 non-split errors를 유발할 수 있다.
- Custom localization criterion: Instance segmentation의 localization에는 target segmentation metric에 맞춘 criterion을 사용할 수 있지만, HD와 같이 고정된 upper bound가 없는 metric은 적절한 cutoff를 정하기 어렵게 한다.예를 들어, NSD target metric이 그에 맞는 localization criterion을 정의할 수 있다.
- Boundary 및 overlap criteria: Mask IoU는 전반적인 structural overlap을 측정하는 반면, Boundary IoU는 boundary correctness에 초점을 맞추지만 불완전한 predictions에도 perfect value인 1.0을 산출할 수 있다.따라서 boundary-focused evaluation은 불완전한 boundaries에 벌점을 부과하지 않는 특정 failure mode를 유발한다.
- IoR: IoR은 prediction이 포함하는 reference-object area를 측정하고 동일한 prediction에 여러 true-positive matches를 허용하여 non-split errors에 대한 penalties를 줄인다.IoR은 주로 dense cell-segmentation images에 사용되지만, large predictions에 의해 왜곡될 수 있으며 boundaries와 small structures에 대해서는 Mask IoU와 동일한 동작을 보인다.
큰 구조 · 작은 구조
Boundary IoU는 경계 오류에 페널티를 부여하고 큰 구조와 작은 구조에서 구조 크기에 더 불변이므로 Mask IoU보다 개선된 성능을 보인다.
- 큰 구조: Boundary IoU는 Mask IoU와 비교해 경계 오류에 특별히 페널티를 부여한다.
- 작은 구조: Boundary IoU는 Mask IoU보다 구조 크기에 더 불변이다.
- 작은 구조: 이 비교에서는 두 가지 서로 다른 threshold를 사용해 Boundary IoU를 평가한다.
- 큰 구조: 크기 불변성 비교에는 큰 구조가 포함된다.
- 작은 구조: 크기 불변성 비교에는 작은 구조가 포함된다.
- 큰 구조: Boundary IoU는 그림의 세 번째 열과 네 번째 열에 제시된다.
경계 픽셀 중첩
경계 픽셀이 중첩되면 예측이 완벽하지 않아도 Boundary IoU가 완벽해 보일 수 있으며, localization 기준과 threshold가 추가적인 모호성과 task-dependent trade-off를 유발한다.
- 경계 픽셀 중첩: distance-to-border 영역에 모든 mask 픽셀이 포함되면 중앙에 구멍이 있는 예측에서도 Boundary IoU가 1.00이 될 수 있지만, Mask IoU는 이 결함을 검출한다.이 예시는 distance = 2를 사용한다.
- Localization 기준: IoU > 0을 detection 기준으로 사용하면 매우 큰 예측도 수용할 수 있어 predicted localization이 모호해진다.이 느슨한 기준은 reference annotation이 정확한 outline을 제공할 때만 권장된다.
- Localization threshold: Localization threshold는 task를 반영해야 한다. 존재 여부에 초점을 둔 구조나 작고, 변이가 크며, 3차원이거나 reference가 불확실한 구조에는 lower threshold가 적합하고, 정밀한 localization에는 higher threshold가 적합하다.Metric은 일반적으로 여러 cutoff 값에 대해 평균을 내며, IoU의 기본값은 0.5에서 0.9까지 0.05 간격으로 설정되지만, 문제의 특성에 따라 cutoff의 관련성이 제한될 수 있다.
3.1 Metrics Cheat Sheets
이 절에서는 관련 metric의 공식, 값의 범위, 특성 및 권고사항을 설명하는 Metrics Reloaded cheat sheet를 제시한다. 많은 metric은 confusion matrix를 기반으로 하며, binary 및 multiclass 설정에서 이를 예시로 보여준다.
- Cheat sheet framework: 각 metric의 공식, 값의 범위, 적용 가능한 문제 범주, prevalence 의존성 및 권고 사용법을 요약한다.또한 높은 값과 낮은 값 중 어느 쪽이 바람직한지도 나타낸다.
- Counting metrics: Balanced Accuracy는 Subprocess S2에서 multi-class counting metric으로 권고된다.
- Overlap metrics: Centerline Dice는 Subprocess S6에서 overlap-based metric으로 권고된다.
- Overlap metrics: Intersection over Union은 Subprocess S6에서 overlap-based metric으로 권고된다.
- Counting metrics: Positive Likelihood Ratio는 Subprocess S3에서 per-class counting metric으로 권고된다.
4.1 영상 수준 분류
이 framework는 정자 운동성과 질병 분류부터 심장 질환 분류까지 7개의 구체적 use case를 포함하는 생물의학 영상 수준 분류에 적용되었다. 그 결과로 도출된 권고사항은 Fig. SN 4.1에 제시되며, 세부적인 metric 선택 지침은 Figs. SN 4.5–SN 4.7에 제시된다.
- 4.1 영상 수준 분류: Fig. SN 4.1은 적용된 영상 수준 분류 문제에 대한 결과 metric 권고사항을 제시한다.framework의 권고사항은 Figs. SN 4.5–SN 4.7에서 metric 선택 Subprocesses S2–S5에 대해서도 상세히 제시된다.
- 4.1 영상 수준 분류: 이 적용 사례들은 다양한 생물의학 영상 수준 분류 문제에 대한 framework의 적용을 보여준다.예시에는 video, dermoscopic, cellular, ultrasound, MRI, mammography 영상이 포함된다.
- 4.1 영상 수준 분류: 7개의 생물의학 영상 수준 분류 use case는 microscopy, dermoscopy, cell-state, ultrasound, MRI, multiple sclerosis, mammography 및 심장 질환 응용을 포괄한다.여기에는 정자 운동성, dermoscopic 질환 [33], autophagy-stage, ultrasound plane [11], multiple-sclerosis lesion [79], breast-cancer 및 cardiac-disease 분류가 포함된다.
4.2 세맨틱 segmentation
이 framework는 5개의 biomedical 세맨틱 segmentation use case에 적용되었으며, metric recommendation은 Fig. SN 4.2에, subprocesses S6 및 S7에 대한 상세 recommendation은 Figs. SN 4.10–SN 4.11에 제시된다.
- 4.2 세맨틱 segmentation: 이 문제들에 대한 metric recommendation은 Fig. SN 4.2에 제시되며, metric-selection subprocesses S6 및 S7에 대한 상세 recommendation은 Figs. SN 4.10–SN 4.11에 제시된다.
- 4.2 세맨틱 segmentation: 5개의 세맨틱 segmentation use case는 embryo microscopy, liver CT, breast WSI lesion labeling, cortical 3D MRI structures, aneurysm TOF-MRA segmentation을 포괄한다.이러한 instantiation은 SemS-1부터 SemS-5로 식별된다.
4.3 객체 검출
이 framework는 biomedical object-detection 문제에 적용되어 use case별 metric recommendation을 제시한다. 이러한 recommendation은 cell, lesion, polyp, mitosis, lung-nodule detection을 포함한 다양한 응용을 다룬다.
- Object detection: 이 framework는 구체적인 biomedical object-detection use case에 대한 metric recommendations을 제공하며, 개요는 Fig. SN 4.3에, 상세 지침은 Figs. SN 4.6–SN 4.9에 제시된다.상세 그림은 metric-selection Subprocesses S3–S4 및 S8–S9를 다룬다.
- Object detection: 적용된 사례는 time-lapse microscopy에서의 cell detection 및 tracking, multimodal brain MRI에서의 MS lesion detection, 그리고 predefined sensitivity of 0.95를 적용한 polyp detection을 포괄한다.이 use case들은 ObD-1, ObD-2, ObD-3으로 식별되며, 관련 연구는 각각, [79], 로 인용된다.
- Object detection: 추가 응용으로는 histopathology image에서의 mitosis detection과 CT image에서의 lung-nodule detection이 있다.인용된 연구는 mitosis detection에 [8], lung-nodule detection에 다.
4.4 인스턴스 분할
이 프레임워크는 현미경검사, 대장내시경, 세포 추적, 뇌 MRI를 아우르는 네 가지 생의학 인스턴스 분할 문제에 적용되었다. 도출된 metric 권고안은 Extended Data Fig. SN 4.4에 제시했으며, 세부 subprocess 지침은 추가 그림에 제시했다.
- Metric 선택: Extended Data Figs. SN 4.6–SN 4.11에는 metric 선택 Subprocesses S3–S4 및 S6–S9에 대한 사용 사례별 상세 권고안이 제시되어 있다.
- 사용 사례: 네 가지 인스턴스 분할 사용 사례는 초파리 뉴런 영상화, 대장내시경 기구 분할 [98], 세포핵 추적, 다발성 경화증 병변 분할 [79]을 포괄한다.
- 프레임워크 적용: Extended Data Fig. SN 4.4는 이러한 구체적인 생의학 인스턴스 분할 문제에 대한 metric 권고안과 함께 프레임워크를 적용한 결과를 제시한다.
5.1 기호 참조
이 절에서는 Business Process Model and Notation (BPMN)에서 유래한 표기법을 사용하는 프로세스 다이어그램의 기호를 개괄한다.
- 5.1 기호 참조: 프로세스 다이어그램은 Business Process Model and Notation (BPMN)에 기반한 기호를 사용한다.Extended Data Fig. SN 5.1에 이 다이어그램에서 사용되는 기호가 개괄되어 있다.
5.2 참조 및 알고리즘 출력의 예상 형식
Metric mapping은 image-level classification, semantic segmentation, object detection, instance segmentation에서 예상되는 참조 및 알고리즘 출력 형식을 정의한다. 입력은 일반적으로 class label과 score, 위치 또는 pixel map을 짝지으며, 해당되는 경우 호환 가능한 좌표계와 경계 표현을 요구한다.
- Image-level Classification: Image-level classification은 image별 class label 또는 multi-label indicator를 사용하며, 알고리즘 출력은 이 형식과 일치하거나 predicted class score를 제공한다.C개 class에서 참조값은 y_I ∈ {1, ..., C} 또는 y_I ∈ {0, 1}^C이며, 가능한 경우 score 기반 출력을 기대한다.
- Semantic Segmentation: Semantic segmentation은 동일한 좌표계와 동일한 spacing에서 각 pixel에 대한 참조 label과 predicted label을 요구한다.각 pixel은 하나의 class assignment 또는 class별 multi-label indicator를 가질 수 있으며, 이에 대응하는 prediction 형식을 사용한다.
- Object Detection: Object detection은 각 object를 class-and-location tuple로 나타내며, prediction에는 [0, 1] 범위의 class score를 선택적으로 포함할 수 있다.위치 정보는 box, center point, radius 또는 지원되는 다른 표현일 수 있다.
- Instance Segmentation: Instance segmentation은 각 object를 class와 image-sized binary pixel map으로 나타내며, predicted class score를 선택적으로 추가한다.참조 및 prediction 경계는 각 instance에 대해 별도의 boundary-pixel list로 제공해야 한다.
- Instance Segmentation: Instance가 서로 접하지 않으면서 연결되어 있을 때는 connected-component analysis를 통해 semantic annotation을 instance 형식으로 변환할 수 있다.참조값이 예상 형식에서 벗어나는 경우에도 pixel-level reference를 image-level reference로 집계하는 등의 방법으로 matching을 수행할 수 있다.
5.3 약어
이 절에서는 artificial intelligence, imaging, calibration, segmentation, detection, validation metrics 전반에 걸쳐 논문 전체에서 사용되는 약어를 정의한다.
- 약어: 용어집은 AI, ML, CT, MRI, DSC, AUC, AUROC, AP, ASSD와 ECE 및 Brier Score 같은 calibration metrics를 포함한 핵심 technical abbreviations를 정의한다.또한 semantic 및 instance-analysis metrics, expected cost, Cohen’s Kappa, centerline Dice Similarity Coefficient와 같은 task- 및 performance-related terms도 포함한다.
- 약어: 추가 약어에는 machine-learning tools, clinical concepts, detection 및 segmentation metrics, decision analysis, statistical measures와 관련된 MONAI, MS, ObD, PQ, NB, NLL, PPV, ROC, RI가 포함된다.목록은 PHI, PR, PSR, RBS, NSD, NPV와 observed-to-expected ratio도 정의한다.
5.4 용어집
용어집은 bounding box, calibration plot, metric, training/test case를 비롯한 핵심 영상 분석 검증 용어를 정의한다. 또한 특수한 metric 표기법과 관련 구조 용어를 명확히 설명한다.
- 정의: bounding box는 일반적으로 검출 대상 object를 완전히 둘러싸는 가장 작은 직사각형이며, calibration plot은 완벽한 calibration에서 벗어난 정도를 시각화한다.Calibration plot은 model output이 전반적으로 과신 또는 과소신인지 진단한다.
- 정의: Metric은 domain-specific validation goal과 관심 특성에 따라 algorithm performance를 정량화하고 검증하며, reference-based metric과 non-reference-based metric 같은 구분이 있다.Metric은 수학적 특성에 따라 family로 묶을 수도 있다.
- 정의: Training 및 test case는 algorithm 개발과 검증에 사용되는 dataset이며, case 하나는 하나의 결과를 생성하는 데 필요한 data를 제공한다.Training case에는 reference annotation이 포함되며 algorithm training에 사용되고, test case는 validation을 지원한다.