Source-linked AI summary
Explainable Medical Imaging AI Needs Human-Centered Design: Guidelines and Evidence from a Systematic Review
Haomin Chen, Catalina Gomez, Chien-Ming Huang, Mathias Unberath
TL;DR
Transparent ML in medical image analysis is often developed without sufficient attention to the users and contexts that determine whether explanations are intelligible. This systematic review identifies design and validation gaps and proposes INTRPRT, which centers formative research, iterative evidence, and user testing. The authors conclude that transparency efforts should extend beyond computational advances to establish whether systems afford transparency for intended users.
Problem
Medical-imaging ML research often treats transparency as a model property while lacking formative user research, empirical user studies, and explicit attention to users and contexts.
Method
The paper conducts a systematic review of transparent ML in medical image analysis and distills the findings into the INTRPRT guideline.
Results
The review finds that contemporary studies disproportionately prioritize algorithmic development, with no reported formative user research or empirical testing to inform and validate design choices.
Takeaways & Limitations
Designers should ground transparency choices in end-user needs and domain requirements, then empirically assess whether systems achieve human-factors goals.
Takeaways & Limitations
The review notes that human-centered design in healthcare is constrained by barriers to end-user involvement, and its guidelines may require refinement as methods mature.
Abstract
from arXiv · showhide
Transparency in Machine Learning (ML), attempts to reveal the working mechanisms of complex models. Transparent ML promises to advance human factors engineering goals of human-centered AI in the target users. From a human-centered design perspective, transparency is not a property of the ML model but an affordance, i.e. a relationship between algorithm and user; as a result, iterative prototyping and evaluation with users is critical to attaining adequate solutions that afford transparency. However, following human-centered design principles in healthcare and medical image analysis is challenging due to the limited availability of and access to end users. To investigate the state of transparent ML in medical image analysis, we conducted a systematic review of the literature. Our review reveals multiple severe shortcomings in the design and validation of transparent ML for medical image analysis applications. We find that most studies to date approach transparency as a property of the model itself, similar to task performance, without considering end users during neither development nor evaluation. Additionally, the lack of user research, and the sporadic validation of transparency claims put contemporary research on transparent ML for medical image analysis at risk of being incomprehensible to users, and thus, clinically irrelevant. To alleviate these shortcomings in forthcoming research while acknowledging the challenges of human-centered design in healthcare, we introduce the INTRPRT guideline, a systematic design directive for transparent ML systems in medical image analysis. The INTRPRT guideline suggests formative user research as the first step of transparent model design to understand user needs and domain requirements. Following this process produces evidence to support design choices, and ultimately, increases the likelihood that the algorithms afford transparency.
Introduction
Transparent ML for medical imaging must be designed around users and clinical context, not treated as a purely computational property. The review identifies limited user involvement and validation as major shortcomings that threaten intelligibility and relevance.
- Designing transparent ML requires human-factors work because explanation mechanisms may suit one user group but not another.
- Transparency is an affordance: it depends on the relationship between an ML algorithm and the user processing its information.
- Iterative empirical user studies and rapid prototyping help ground design choices in users’ needs, knowledge, and context before commitment to one approach.
- Healthcare complicates human-centered design through knowledge mismatch, restricted user access, complex clinical data and tasks, and limited designer training in human factors.
- The systematic review finds absent formative research, scarce empirical user studies, and insufficient attention to users and contexts, risking unintelligible and irrelevant systems.
- The paper proposes INTRPRT to help designers involve end users and justify transparency-related choices during design, construction, and validation.
An overview of current trends in transparent machine learning development
The review organizes transparent medical-imaging ML around six interacting themes and finds substantial gaps in multidisciplinary participation, user specification, and user-centered justification.
- The review structures 68 included studies around six themes: incorporation, interpretability, target, reporting, prior, and task.
- Only 33 studies involved multidisciplinary clinician-engineering teams, and no study reported formative user research before model construction.
- Only 30 studies specified end users, all targeting clinical care providers despite the variety of healthcare stakeholders.
- Clinical priors or guidelines inspired about half of the selected articles, with n=28 reporting this basis for transparent systems.
- Prediction tasks dominated the reviewed applications, appearing in 57/68 studies.
INTRPRT guideline
INTRPRT translates human-centered design principles into guidance for specifying clinical contexts, justifying transparency, communicating with users, and validating both task performance and transparency.
- INTRPRT distills guidance from interactions among six themes to support transparent ML design and validation across healthcare applications.
- Guideline 1: Guideline 1 requires specifying the clinical scenario, constraints, requirements, and end users before designing the healthcare ML algorithm.
- Guideline 2: Guideline 2 requires justifying the transparency choice and identifying its evidence level, especially where designers and users differ in expertise and context.
- Evidence levels: Evidence levels range from Level 0 with no user investigation to Level 3 with iterative refinement through communication and feedback between designers and end users.
- The framework encourages explicit awareness of evidence strength so development effort can be balanced between ML methods and richer support for user-related assumptions.
- Guideline 3: Guideline 3 requires consistency between the transparency technique and the assumptions used to justify it.
- Guideline 4: Guideline 4 emphasizes that information format, communication channel, and interactivity can affect users’ experience and performance.
- Guidelines 5–6: Guidelines 5 and 6 require quantitative task-performance evaluation plus validation of technical correctness and human factors such as trust, reliance, satisfaction, mental models, and acceptance.
Discussion
The discussion argues that transparent medical-imaging ML must be designed and validated with end users because transparency is context-dependent. The review motivates the INTRPRT guideline and emphasizes formative research, empirical testing, and broader stakeholder consideration.
- INTRPRT guideline: The INTRPRT guideline organizes considerations into six themes and provides a step-by-step framework for designing transparent ML systems with clinical end users.The themes are incorporation, interpretability, target, reporting, prior, and task.
- Human-centered design: Formative user research and empirical testing are critical for understanding user needs, iterating design choices, and validating whether systems afford transparency.The authors report that contemporary studies generally lacked both formative research and empirical testing.
- Human-centered design: Most reviewed studies prioritized technological changes to complex ML systems while assuming those changes would achieve transparency and human factors goals.The authors caution that such assumptions are unlikely to be reliable without formative research or empirical tests.
- Diverse stakeholders: Transparent ML research focused heavily on clinicians, although nurses, technicians, administrators, insurers, and patients may have distinct needs, knowledge, and expectations.Different stakeholder groups may therefore require different technological and human-factors approaches.
- Task context: Tasks with established human workflows and guidelines provide Level 2 evidence for transparency, whereas tasks without human baselines require empirical validation of proposed transparency mechanisms.Human-defined baselines also facilitate data collection and annotation by specifying intermediate outputs in advance.
- Implications: The review identifies the lack of explicit formative research as the largest barrier to capitalizing on transparent ML’s benefits in medical image analysis.The authors note that the guideline may require refinement as human-centered development matures.
Methods
The authors systematically reviewed transparent ML methods for medical image analysis published from January 2012 through July 2021. They searched three databases, screened records using PRISMA procedures, and independently extracted data from 68 included articles.
- Scope: The review covered transparent ML methods for medical image analysis published after January 2012, when interest in learning-based image processing accelerated.The stated aim was to survey the field’s current state.
- Search strategy: The authors searched PubMed, EMBASE, and Compendex for relevant articles from January 2012 through July 2021 using PRISMA procedures.Titles, abstracts, and keywords were screened during the search.
- Study selection: After duplicate removal, title-and-abstract screening, and full-text eligibility review, 68 articles met the inclusion criteria.The passage reports 1,731 records after duplicate removal and 217 after initial prescreening.
- Data extraction: Two authors independently coded all 68 included articles using a shared extraction template, after which one author merged their reports into a consensus document.The template summarized information across the six INTRPRT themes.
- Limitations: The review may have missed relevant studies because transparency terminology was often absent from article metadata and sometimes from the article text.Its restriction to published manuscripts, long articles, and novel approaches also introduces potential publication bias.
Detailed analysis of findings during systematic review
The 68-study review found diverse technical approaches to transparency, with attention mechanisms most common and clinical expertise influencing model incorporation of prior knowledge. However, formative user research was not explicitly reported.
- IN: Interpretability: Transparency methods included human-understandable features, combined deep and traditional ML, visualization, clustering, uncertainty estimation, feature relations, and custom techniques.The review counted 11 human-understandable-feature studies, 7 combined-model studies, 5 visualization studies, 4 clustering studies, and 20 custom-technique studies.
- IN: Interpretability: Attention mechanisms were the most common transparency technique, appearing in 15 studies and supporting pixel-attribution visualizations for class-specific importance.In segmentation, multi-resolution features were aggregated to improve outcomes in applications such as fetal MRI.
- IN: Interpretability: Human-understandable features were created through hand-crafted morphological or radiomic features and clinical variables, often analyzed with a separate classifier.Other studies analyzed deep encoded features using clinically informed decision trees, rules, or regression methods.
- IN: Interpretability: Feature-based methods clustered deep features or estimated feature importance, while region-importance methods used image occlusion to identify informative sub-regions.Occluded regions could be blank or healthy-looking in classification and detection tasks.
- IN: Interpretability: Architecture-modification methods embedded clinical knowledge by pruning networks, mirroring clinical workflows, or aggregating multiple imaging views.Examples included scale-invariant pruning, ten ultrasound-image branches, and three-view mammography processing.
T: Targets
Among 68 reviewed studies, none targeted users beyond care providers, and fewer than half explicitly identified clinicians as intended end users.
- T: Targets: None of the selected articles aimed to build transparent systems for users other than care providers.
- T: Targets: Less than half of the articles explicitly specified clinicians as intended end users (n=30).Among the remaining 38 articles, 17 implied clinicians, while 21 did not specify target users.
- T: Targets: 47% of articles specifying or implying clinicians implemented clinical prior knowledge, compared with 18% of articles without end-user information.
R: Reporting
The review found that transparent medical-imaging systems were usually evaluated through task performance or qualitative visualization, with limited direct assessment of human-factors outcomes and incomplete justification of design choices.
- R: Reporting: Functional evaluation used auxiliary tasks such as detection or segmentation to quantify explanation quality, but required additional manual ground-truth annotations.
- R: Reporting: Qualitative validation was the most common transparency-evaluation approach, appearing in 40 studies through attribution overlays, feature rankings, and narrative observations.Such methods did not inherently account for human factors and have been criticized for limited fidelity and specificity.
- R: Reporting: Direct user studies evaluated how target users interacted with transparent systems, including radiologists’ understanding of example-based and feature-based explanations.One cited study involved 8 radiologists and assessed understanding by measuring prediction of the AI diagnosis for a target image.
- R: Reporting: 91% of articles evaluated task performance (n=62), while 49 reported no metrics beyond the main task and 9 did not discuss transparency.Among studies comparing transparent systems with non-transparent baselines, 36 reported improved performance and 5 comparable results.
- R: Reporting: 93% of articles incorporating clinical-knowledge priors directly implemented them in model structure or inference, compared with 68% using computer-vision priors.
- R: Reporting: Clinical knowledge was incorporated through human-understandable features, biomarkers, clinical guidelines, and workflows, whereas computer-vision priors were generally less application-specific.
- R: Reporting: None of the included articles formally described how priors were formulated to achieve transparency.The authors argue that unmatched priors and insufficiently justified design choices risk producing inaccurate, unintelligible, or irrelevant insights for end users.
T: Task
The reviewed literature concentrated on transparent ML for classification, detection, and selected segmentation tasks, especially in 3D radiology and pathology, usually for tasks already performed by human experts.
- T: Task: 57 articles proposed transparent ML algorithms for classification and detection problems.The most common modalities were 3D radiology images (n=24) and pathological images (n=15).
- T: Task: Segmentation was a major application field represented by 9 articles, mainly involving brain and cardiac MRI.
- T: Task: 60 articles addressed tasks routinely performed by human experts in current clinical practice.Only 4 articles targeted more difficult tasks without a human baseline.
Data Availability
Figure 2 contains images from the ORIGA and BraTS2020 public datasets.
- Data Availability: Figure 2 contains images from the ORIGA and BraTS2020 datasets.