Source-linked AI summary
What Clinicians Want: Contextualizing Explainable Machine Learning for Clinical End Use
Sana Tonekaboni, Shalmali Joshi, Melissa D McCradden, Anna Goldenberg
TL;DR
Clinical ML systems may remain unused even when accurate because clinicians need trustworthy, clinically usable explanations. The authors surveyed ICU and ED clinicians and connected their needs to explainability research, finding that trust is tied to clinical alignment, context, and awareness of model limitations.
Problem
Clinical ML lacks objective definitions of explanation quality and usability, while accurate systems may still fail to gain clinicians’ endorsement and sustained use.
Method
The authors conducted an exploratory pilot with 10 clinical stakeholders and synthesized their feedback with explainability literature to identify explanation classes and evaluation metrics.
Results
Clinicians valued explanations that justify decisions, identify relevant features, provide operating context, reveal model shortcomings, and align predictions with clinically significant changes.
Takeaways & Limitations
Clinical explainability should be evaluated through domain-appropriate representations and consistent explanations that support reliable actionability and clinician trust.
Takeaways & Limitations
The findings come from ICU and ED specialists familiar with clinical ML and may have limited applicability to other model classes and outpatient settings.
Abstract
from arXiv · showhide
Translating machine learning (ML) models effectively to clinical practice requires establishing clinicians' trust. Explainability, or the ability of an ML model to justify its outcomes and assist clinicians in rationalizing the model prediction, has been generally understood to be critical to establishing trust. However, the field suffers from the lack of concrete definitions for usable explanations in different settings. To identify specific aspects of explainability that may catalyze building trust in ML models, we surveyed clinicians from two distinct acute care specialties (Intenstive Care Unit and Emergency Department). We use their feedback to characterize when explainability helps to improve clinicians' trust in ML models. We further identify the classes of explanations that clinicians identified as most relevant and crucial for effective translation to clinical practice. Finally, we discern concrete metrics for rigorous evaluation of clinical explainability methods. By integrating perceptions of explainability between clinicians and ML researchers we hope to facilitate the endorsement and broader adoption and sustained use of ML systems in healthcare.
1. Introduction
Clinical ML adoption is constrained not only by technical barriers and high stakes, but also by clinicians’ need to trust and endorse model outputs. This study frames explainability as a practical, measurable way to understand clinician needs and support sustained clinical use.
- Technical barriers to clinical ML adoption include limited robustness, complex modeling tasks, and high-stakes decisions.
- Accurate ML systems may still fail to gain routine clinical endorsement or sustained use, motivating research on how they can earn clinicians’ trust.
- Existing explainability research lacks objective definitions for validating explanation quality and usability in clinical practice.
- The paper defines healthcare ML explainability as measurable, quantifiable, and transferable attributes that help clinicians calibrate trust in an ML system.
- An exploratory pilot with 10 clinical stakeholders examined explainability’s role, useful explanation classes, ML implications, and evaluation metrics.
- The work connects clinicians’ explanation needs with technical properties and existing explainability research to identify gaps relevant to clinical translation.
2. Study Design
The study used upstream qualitative stakeholder engagement to explore explainability before model implementation. Ten ICU and ED clinicians discussed hypothetical prediction tools and how they would assess alerts, trust, and clinical actionability.
- The researchers conducted interviews before model implementation because qualitative methods could capture conceptual complexity and clarify clinicians’ reasoning.
- Ten clinicians from ICU and ED settings were interviewed to identify explainability needs for reliable ML-supported clinical practice.
- The convenience sample included end-user stakeholders familiar with ML-based clinical tools and ongoing developments in the field.
- The sample comprised 6 ICU clinicians and 4 ED clinicians, evenly split between senior and junior participants and between men and women.
- Interviews began by eliciting clinicians’ meanings and expectations of explainability, then introduced specialty-specific hypothetical ML scenarios.
- Follow-up questions examined alert validity, clinical actionability, comparison with non-ML tools, and clinicians’ existing practices.
3. Results
Clinicians described explainability as supporting understanding, rationalization, trust calibration, and validation of ML predictions in acute care. They prioritized clinically relevant features, patient trajectories, transparent designs, context-aware representations, and rigorous evaluation across uncertainty and model-design conditions.
- Clinician needs: Clinicians use explanations to understand and rationalize predictions, requiring clinically relevant features, operating context, and awareness of situations where models may fall short.These needs support clinical decision-making and help users assess whether predictions align with evidence-based practice.
- Explanation classes: Feature importance helps clinicians compare model decisions with clinical judgment and focus attention on patient characteristics, especially in time-constrained emergency settings.Clinicians requested both patient-specific and population-level variable importance, although patient-level importance remains comparatively underexplored.
- Explanation classes: Instance-level explanations based on similar patients may inform diagnostic reasoning or intervention choices, but their usefulness depends on the clinical task and similarity definition.Clinicians noted that patients with similar outcomes may have different trajectories, while similar cases may still be useful for examining actions and intervention outcomes.
- Explanation classes: Patient trajectories and influential state changes were identified as important temporal explanations, but higher-order temporal dependencies and inconsistent attention-based explanations remain challenges.ICU clinicians particularly wanted to see how changes in an individual patient’s state contributed to a prediction.
- Explanation classes: Transparent designs that mirror evidence-based medical reasoning can facilitate rationalization, particularly when clinicians need to inspect prediction-driving features or discrepancies with clinical protocols.Decision-tree-like designs may help, but excessive detail can reduce usefulness and unexplained protocol discrepancies can undermine confidence.
- Evaluation metrics: Clinical explanations should be evaluated for domain-appropriate representation and consistency, while also accounting for data uncertainty, model misspecification, robustness, and induced biases.Some explanation methods may work well for particular data types, such as images, but require substantial adaptation for other data and complex clinical settings.
4. Related Work
Prior work spans interpretable models, feature-level explanations, clinical visualization, attention mechanisms, and formal evaluation of explainability. However, the literature includes limited stakeholder-focused clinical studies and methods may require clinical adaptation.
- The authors state that no prior work had specifically used target-stakeholder studies to identify explainability challenges in clinical ML.
- General explainability research includes simpler model classes, LIME, and Anchors for deriving feature-level explanations.
- Table 1 summarizes explainable ML methods contextualized for clinical applicability.
- Prior clinical ML studies examined visualization systems, neural attention, tree-based regularization, and rule-based or sparse linear models.
- Interpretability evaluation has been approached through application-specific frameworks, user studies, and statistical robustness methods.
5. Discussion
The discussion frames clinician-centered explainability as a way to identify technical requirements for trust and clinical translation. The study also acknowledges limited generalizability beyond its surveyed specialties and prediction-focused setting.
- The survey involved diverse ICU and ED stakeholders to identify when explainability can enhance trust in ML models.
- Clinician needs were mapped to specific technical challenges and evaluated against explainable ML literature.
- The authors propose rigorous evaluation of clinical explainability methods using metrics described in the paper.
- The survey was restricted to ICU and ED specialists, and the identified challenges may have limited applicability to other model classes and outpatient settings.
- The authors aim to develop a conceptual framework for consistent and high-quality explainability requirements across hospital ML tools.
Additional Questions
The interview questions elicited clinicians’ views on information that supports confidence, estimates patient outcomes, conveys prediction confidence, highlights influential measurements, and presents similar cases.
- Clinicians were asked what information could increase confidence in managing a patient.
- Clinicians were asked what information they currently use to estimate how a patient will do.
- The interviews assessed whether knowing confidence for an individual model prediction would affect clinicians’ reactions.
- The survey examined whether displaying measurements and history most influential to a model decision could help clinicians and save decision-making time.
- Clinicians were asked whether similar patients and their trajectories or outcomes would be helpful, and how they would use that information.
6. Are the set of top measurements that are driving the prediction useful?
The question asks whether identifying the measurements driving a prediction would provide useful and actionable information to clinicians in ICU or ED scenarios.
- Clinicians were asked which added information would be most useful and actionable after considering all proposed options.
Appendix B. Qualitative Synopsis and Representative Quotes from Stakeholders
Stakeholders viewed clinical ML as valuable for surveillance and risk-based attention allocation, but explainability was tied to understanding and justifying model predictions in practice.
- Clinical ML was valued as an adjunct for improving patient outcomes in both ICU and ED settings.Clinicians described uses including continuous physiologic surveillance and support for care decisions.
- Clinicians valued ML systems that could mimic the clinical acuity of experienced specialists.
- In the ED, clinicians wanted ML to allocate attention by directing them to the right patient at the right time.They associated this need with uncertainty about individual patient trajectories and reactive intervention.
- Clinicians viewed explainability primarily as a way to justify clinical decision-making about model predictions to patients and colleagues.
- Understanding the variables and data driving a prediction was considered essential, especially when the model differed from clinical protocol.Such discrepancies made clinicians nervous when the reason for the model's behavior was unknown.