Source-linked AI summary

The METRIC-framework for assessing data quality for trustworthy AI in medicine: a systematic review

Daniel Schwabe, Katinka Becker, Martin Seyferth, Andreas Klaß, Tobias Schäffter

arXiv:2402.13635v1cs.LGcs.AI

TL;DR

Medical AI needs trustworthy datasets because training-data quality shapes system behaviour and can contribute to bias and unfairness. The paper synthesises data-quality research for medical ML and proposes the METRIC-framework, whose 15 awareness dimensions guide assessment of medical training data. The framework is intended to support dataset understanding and systematic assessment relevant to future regulation, while its scope is limited to ML-focused, application-specific dataset content.

  • Problem

    Medical AI requires trustworthy data because training-data quality shapes system behaviour and biased data can contribute to unfair learned patterns.

  • Method

    The paper transfers existing data-quality research to medical ML and synthesises it into a specialised framework for medical training data.

  • Results

    The METRIC-framework comprises 15 awareness dimensions for investigating medical training datasets with respect to a specific task.

  • Takeaways & Limitations

    Systematic dataset-quality assessment can help developers understand AI behaviour and may support future medical-AI regulation and approval.

  • Takeaways & Limitations

    The framework assesses fixed dataset content for a specific application and limits its results to ML approaches, particularly DL.

Abstract

from arXiv · show

The adoption of machine learning (ML) and, more specifically, deep learning (DL) applications into all major areas of our lives is underway. The development of trustworthy AI is especially important in medicine due to the large implications for patients' lives. While trustworthiness concerns various aspects including ethical, technical and privacy requirements, we focus on the importance of data quality (training/test) in DL. Since data quality dictates the behaviour of ML products, evaluating data quality will play a key part in the regulatory approval of medical AI products. We perform a systematic review following PRISMA guidelines using the databases PubMed and ACM Digital Library. We identify 2362 studies, out of which 62 records fulfil our eligibility criteria. From this literature, we synthesise the existing knowledge on data quality frameworks and combine it with the perspective of ML applications in medicine. As a result, we propose the METRIC-framework, a specialised data quality framework for medical training data comprising 15 awareness dimensions, along which developers of medical ML applications should investigate a dataset. This knowledge helps to reduce biases as a major source of unfairness, increase robustness, facilitate interpretability and thus lays the foundation for trustworthy AI in medicine. Incorporating such systematic assessment of medical datasets into regulatory approval processes has the potential to accelerate the approval of ML products and builds the basis for new standards.

Introduction

Trustworthy AI in medicine depends critically on data quality because training data shapes system behaviour and can transmit bias. This review transfers data-quality research to medical ML and proposes the METRIC-framework for evaluating medical training data.

  • Motivation: Training-data quality fundamentally influences AI behaviour, while biased data can produce unfair learned patterns that neural networks may amplify at test time.These concerns make dataset assessment relevant to both AI developers and regulators.
  • Scope: Trustworthy AI encompasses ethical, societal, security, privacy, robustness, interpretability, explainability, transparency, and accountability concerns.The paper narrows its results to ML approaches, particularly DL, while using the broader AI vocabulary common in healthcare discussions.
  • Research question: The review asks which characteristics should be evaluated when datasets are used for trustworthy AI in medicine.It focuses on dataset content rather than surrounding data-governance or infrastructure processes.
  • Contribution: The METRIC-framework transfers existing data-quality knowledge to medical ML and provides 15 awareness dimensions for investigating medical training datasets.The dimensions are intended as a comprehensive framework for medical training data with respect to a specific task.
  • Scope: The framework is intended for a fixed dataset assessed against a specific application, not for data quality in isolation or for training strategies that compensate for poor data.The review excludes data governance, data management, and infrastructure-focused quality considerations.

Results

The review synthesizes 62 studies on data quality for trustworthy AI in medicine, spanning general, big-data, and ML-focused frameworks and applications. It finds that ML research usually examines isolated data-quality dimensions and robustness, while broader frameworks and effects on fairness or explainability remain underrepresented.

  • Results: 62 records from 2362 identified studies formed the systematic-review literature corpus on data quality for trustworthy AI in medicine.The review followed PRISMA guidelines and searched PubMed and the ACM Digital Library.
  • Results: The corpus comprises 28 general-data, 7 big-data, and 27 ML-data studies, reflecting the field’s historical development.The categories represent changing perceptions of data quality across publication periods.
  • General data quality: Medical data-quality frameworks evolved from EHR-focused work toward specialised frameworks for immunisation, public-health, and other data types, amid inconsistent terminology.Medical frameworks addressed characteristics including accuracy, completeness, timeliness, and concordance between differing data sources.
  • ML data quality: Deep-learning studies mainly evaluate isolated dimensions such as data amount, missingness, feature contamination, and label noise rather than broad theoretical frameworks.The literature spans tabular data, time series, images, natural language, and molecular data, with data amount showing performance benefits that saturate.
  • ML data quality: Most ML studies assess robustness to erroneous or limited inputs, whereas generalisability and predictive uncertainty receive comparatively little attention.The corpus includes only a few studies on generalisability and one notable study that additionally examines predictive uncertainty.
  • ML data quality: The review identifies limited attention to theoretical frameworks and underrepresentation of fairness and explainability, a potential shortcoming for safety-critical medical diagnosis applications.The literature’s emphasis is skewed toward robust predictive performance after data manipulation.

METRIC-framework for medical training data

The METRIC-framework is a specialised framework for evaluating medical training-data quality through awareness dimensions rather than a one-model-fits-all scheme. It organises data-quality characteristics into clusters, dimensions, and subdimensions spanning acquisition, timeliness, representativeness, informativeness, and consistency.

  • METRIC-framework: The METRIC-framework is specialised for medical training data because data quality depends on the application and strongly influences ML behaviour.It presents awareness dimensions for evaluation rather than guidelines for measuring data quality.
  • METRIC-framework: The framework uses three detail levels: clusters group related dimensions, dimensions describe individual quality characteristics, and subdimensions provide finer-grained attributes.Frequently mentioned dataset properties are separated into a data-management cluster rather than included in the framework itself.
  • Measurement process: Measurement process covers device and human acquisition errors, completeness, and source credibility, with noise evaluated against expected deployment conditions rather than assumed to be universally harmful.When medical ground truth is unavailable, training-data noise should be compared with deployment noise; adding noise can sometimes improve performance.
  • Timeliness: Timeliness evaluates whether dataset creation and updating remain appropriate for use as medical knowledge, diagnoses, and coding systems change.The relevant comparison is between when data was created or updated and when it is used for the task.
  • Representativeness: Representativeness examines population coverage, demographic variety, data depth, and target-class balance for the intended medical application.Rare classes may be deliberately overrepresented to help models learn their patterns, even when that differs from real-world prevalence.
  • Informativeness and consistency: Informativeness assesses whether data conveys useful information clearly and compactly, including understandability, redundancy, informative missingness, and feature importance.Consistency separately addresses rule-based presentation, logical contradictions, and statistical distribution differences across subsets.

Discussion

The METRIC-framework organizes 15 awareness dimensions for task-specific medical training-data assessment, while highlighting measurement type, use-case dependence, and practical priorities. The framework is a starting point for systematic assessment, but complete evaluation still requires task-appropriate measures, prioritization, and aggregation.

  • Discussion: The METRIC-framework provides 15 awareness dimensions for investigating medical training data with respect to a specific task.These dimensions help developers understand data characteristics relevant to AI-system behaviour and improvement.
  • Discussion: Overall data-quality assessment requires selecting measures, obtaining dimension-specific results, evaluating task appropriateness, and eventually combining outcomes.The framework does not rank dimensions, and a complete assessment process must account for importance, measurability, use-case dependence, and required expertise.
  • Discussion: About half of the dimensions are mostly quantitatively measurable, whereas roughly one fifth require predominantly qualitative inspection.Quantitative measures can improve objectivity and automation, while qualitative or mixed assessment may require costly medical-domain expertise.
  • Discussion: Representativeness, timeliness, device error, and feature importance are use-case-dependent dimensions requiring additional knowledge, work, and time during assessment and improvement.Appropriate coding standards, data currency, feature relevance, and noise levels depend on the intended application rather than being universally optimized.
  • Discussion: Representativeness, feature importance, distribution consistency, and human-induced error are estimated as crucial overall-quality factors, with six dimensions recommended for prioritization.Except for feature importance, the crucial dimensions are mostly quantitative and therefore potential targets for dataset-quality software tools; their effects remain to be quantified across ML problems.
  • Discussion: Systematic dataset-quality assessment could support regulatory evaluation and potentially accelerate approval of medical ML products.The authors present METRIC and related considerations as a starting point for incorporating data quality into regulation and certification.

Methods

The study used a PRISMA-guided systematic review to identify literature on data quality frameworks and their effects in medical ML. The authors then synthesized terminology and concepts from eligible records into the METRIC-framework.

  • Eligibility criteria: Eligible studies either proposed broad or medical data-quality frameworks or examined how training-data quality dimensions affect deep-learning behaviour.Frameworks specific to non-medical fields and studies focused only on isolated dimensions without ML relevance were excluded.
  • Screening: Two authors independently screened titles and abstracts, resolving disagreements through discussion or consultation with a third author when necessary.This workflow reduced the record set to 104 before snowballing and subsequent assessment.
  • Literature review process: 62 records passed all screening steps from the systematic-review corpus.The initial search, screening, snowballing, and full-text assessment produced 62 eligible entries.
  • Framework construction: The authors extracted data-quality terms from all selected records, grouped detailed terms into subdimensions, and hierarchically organized them into METRIC dimensions and clusters.Terms were filtered for transferability and relevance, while authors reached consensus on framework definitions; the subdimension variety of data sources was retained despite limited systematic quantification.

Competing interests

The authors report no competing interests and acknowledge funding from the EU TEF-Health project.

  • The authors declare no competing interests.
  • The study received funding from the EU TEF-Health project within the Digital Europe Programme.
  • Extracted literature data and terms underlying the METRIC-framework are provided in a supplementary Excel file.
Loading 2402.13635v1…