Source-linked AI summary
CheXplain: Enabling Physicians to Explore and UnderstandData-Driven, AI-Enabled Medical Imaging Analysis
Yao Xie, Melody Chen, David Kao, Ge Gao, Xiang 'Anthony' Chen
TL;DR
Data-driven medical AI can perform diagnostic imaging analysis but remains difficult for physicians to understand because it operates as a black box. The paper develops CheXplain through a paired survey, physician co-design, and prototype evaluation, finding that interaction supported more detailed and accurate explanations of AI. It concludes with recommendations organized around motivation, constraint, explanation, and justification.
Problem
Data-driven medical AI promises diagnostic-imaging performance but remains difficult for physicians to understand, while prior explainability work rarely addresses physicians’ domain-specific needs and practices.
Method
The authors used a paired physician–radiologist survey, co-designed a low-fidelity prototype with three physicians, and evaluated a high-fidelity CheXplain prototype with six medical professionals.
Results
Physicians provided more detailed and accurate explanations of the underlying AI after interacting with CheXplain.
Takeaways & Limitations
The paper proposes physician-centered design recommendations for explainable medical AI organized around motivation, constraint, explanation, and justification.
Takeaways & Limitations
The available chest X-ray dataset lacked patients’ medical history and additional laboratory results, limiting further exploration with CheXplain.
Abstract
from arXiv · showhide
The recent development of data-driven AI promises to automate medical diagnosis; however, most AI functions as 'black boxes' to physicians with limited computational knowledge. Using medical imaging as a point of departure, we conducted three iterations of design activities to formulate CheXplain---a system that enables physicians to explore and understand AI-enabled chest X-ray analysis: (1) a paired survey between referring physicians and radiologists reveals whether, when, and what kinds of explanations are needed; (2) a low-fidelity prototype co-designed with three physicians formulates eight key features; and (3) a high-fidelity prototype evaluated by another six physicians provides detailed summative insights on how each feature enables the exploration and understanding of AI. We summarize by discussing recommendations for future work to design and implement explainable medical AI systems that encompass four recurring themes: motivation, constraint, explanation, and justification.
INTRODUCTION
CheXplain addresses the difficulty physicians face understanding black-box, data-driven medical AI by using iterative, physician-centered design for interactive chest X-ray analysis. Three iterations—survey, co-design, and prototype evaluation—produced eight features and showed that interaction helped physicians provide more detailed and accurate explanations of AI.
- INTRODUCTION: AI’s low-level neural-network constructs make its conclusions difficult for medical professionals to understand, limiting clinical adoption despite promising performance.
- INTRODUCTION: CheXplain was developed through a survey, physician co-design, and high-fidelity evaluation to help physicians explore and understand AI-enabled chest X-ray analysis.The survey included 39 referring physicians and 38 radiologists; the high-fidelity prototype was evaluated with six medical professionals.
- INTRODUCTION: The low-fidelity co-design produced eight features covering targeted inquiries, urgency-based complexity, hierarchical explanations, contrastive examples, probabilities, contextual information, and image comparisons.
- INTRODUCTION: After interacting with CheXplain, physicians provided more detailed and accurate explanations of the underlying AI than before interaction.
- INTRODUCTION: The paper’s recommendations organize explainable medical AI around motivation, constraint, explanation, and justification.
BACKGROUND: AI AND RADIOLOGY
Medical AI has progressed from rule-based expert systems toward deep-learning analysis of chest X-rays, while explainability research has pursued rule extraction, attribution, and intrinsically interpretable models. These approaches target model transparency but do not by themselves establish how physicians understand AI outputs.
- BACKGROUND: AI AND RADIOLOGY: CheXpert uses a deep neural network to analyze individual chest X-rays and output labels for radiographic observations.
- BACKGROUND: AI AND RADIOLOGY: Explainable-AI research commonly uses rule-extraction, attribution, or intrinsic methods to make model behavior more interpretable.
Explainable Artificial Intelligence (XAI) Research in HCI
HCI research frames explainable AI around users, context, cognition, learnability, and visualization, while clinical decision-support research seeks understandable systems for medical professionals. However, AI and HCI research remain insufficiently connected in medical applications.
- Explainable Artificial Intelligence (XAI) Research in HCI: HCI research treats explainability as a user-centered problem involving context awareness, cognitive explanations, software learnability, and visualization.
- Explainable Artificial Intelligence (XAI) Research in HCI: The HCI and AI communities’ research often remains disconnected, leaving limited interdisciplinary work on explainable AI.
- Explainable Artificial Intelligence (XAI) Research in HCI: Clinical decision-support systems are intended to enhance physician decision-making across screening, treatment planning, medication decisions, and diagnosis.
- Explainable Artificial Intelligence (XAI) Research in HCI: AI-based clinical decision-support use is limited when systems fail to provide relevant information or capture physicians’ mental models.
Communication between Radiologists and Physicians
Communication between radiologists and referring physicians provides the clinical context for designing explanations in AI-enabled radiology. Because prior work offered little evidence about how such explanations are used, the paper begins with a paired survey of both groups.
- Communication between Radiologists and Physicians: Prior studies suggest that precise medical histories, informative requests, direct communication, and patient questioning can improve radiologist–physician communication.
- Communication between Radiologists and Physicians: The authors found little prior research on how explanations are used between referring physicians and radiologists, especially in AI-enabled scenarios.
- Communication between Radiologists and Physicians: A paired survey asked referring physicians and radiologists parallel questions to compare both parties’ perspectives on explanations.
- Communication between Radiologists and Physicians: The survey included 39 referring physicians and 38 radiologists and also asked referring physicians about hypothetical AI-generated results.
Whether Explanation Is Needed
Explanation needs are driven by information gaps, low trust in AI, and disagreement or high-stakes findings, while time cost constrains seeking it.
- Time cost constrains explanation seeking: waiting ranges from 15 minutes to one day, and only 56% of participants were satisfied with the efficiency.
- Only 1/3 of referring physicians trusted an AI radiologist, compared with 38/39 likely or very likely to trust a human radiologist.Participants expected AI to explain itself with probabilities and professional knowledge.
- Referring physicians seek explanations when questions remain unanswered or findings are atypical or critical.
- Physicians are more likely to seek explanations when their own hypotheses conflict with AI results or they are unsure which is correct.
- Prior images, other modalities, differential diagnoses, annotations, and regional comparisons are commonly expected explanatory resources.
- Justifications use external information such as prior images or other modalities, whereas explanations describe the intrinsic process generating results.
KEY SYSTEM DESIGNS
A physician-centered, low-fidelity co-design process translated survey requirements into CheXplain’s system designs, spanning targeted, hierarchical, comparative, probabilistic, contextual, and urgency-aware interaction.
- A co-design with three physicians developed CheXplain’s system requirements into eight specific features for exploring and understanding AI radiologists.The participants represented novice, competent, and expert radiological knowledge levels.
- Specific inquiries: Specific inquiries let physicians provide case-relevant questions so AI can focus examination on relevant CXR regions.
- Hierarchical explanations: Hierarchical explanations connect low-level examinations, mid-level observations, and high-level impressions, supporting bottom-up and top-down reasoning.More knowledgeable physicians tended to reason bottom-up, whereas less knowledgeable physicians often reasoned top-down.
- Contextual and comparative designs: Contrastive examples, probability calibration, prevalence and traces, temporal comparisons, and cross-patient comparisons provide contextual ways to inspect AI results.
#3 Contextualizing observations with contrastive examples
CheXplain contextualizes observations through contrastive examples and calibrated probabilities, helping physicians inspect findings against normal, abnormal, and differently scored cases.
- Contrastive examples let physicians compare a selected CXR region with normal and abnormal cases to explore what-if questions and shared patterns.The design uses images of the same observation so physicians can assess whether the current case resembles a common pattern or an outlier.
- Probability values are useful only when mapped to medical workflows, because physicians reported confusion about interpreting numeric confidence directly.
- The prototype calibrates observation probabilities by showing same-observation CXRs sorted across unlikely, likely, and very likely levels.An 80% edema example is presented as similar to a near-100% case.
- Physicians also proposed contextualizing impressions with epidemiologic baselines and factors such as age, risk factors, and smoking history.
- The system adds prevalence and traces back to observations that raise or lower an impression’s probability.
#6 Comparison across time
CheXplain uses prior-image and cross-patient comparisons to contextualize findings, while an urgency mode reduces interaction demands when physicians have limited time.
- Physicians valued comparing current images with the same patient’s prior images to identify persistence, change, improvement, worsening, and new observations.
- The temporal comparison design provides side-by-side current and prior cases with annotations and filtering for differences.
- Physicians also suggested showing other patients with similar inquiries, implemented as cross-patient cases responding to the same inquiry.
- Because time cost constrains explanation seeking, participants proposed urgent and non-urgent modes, with hurried users focusing first on important abnormalities.
- The urgent mode shows annotations for observations that are both high-confidence and associated with critical impressions.
- The eight features were integrated into a high-fidelity prototype evaluated with six medical professionals of varied radiological knowledge.
Apparatus
The high-fidelity CheXplain prototype was evaluated through paired interpretation tasks comparing physicians’ understanding before and after interaction with the system. The study used real AI results with manually generated explanatory information and assessed feature helpfulness.
- System apparatus: CheXplain integrated eight explanatory features into a web-based front end using CheXpert cases, AI-generated observations, and manually generated supporting information.Input images, regional annotations, prior images, and across-patient images came from CheXpert; remaining information was based on guidelines and verified by a radiologist.
- Evaluation procedure: Physicians interacted remotely through browsers and Zoom while researchers observed, recorded, transcribed, and coded their sessions.A second experimenter reviewed the coding and resolved disagreements through discussion.
- Evaluation procedure: The study compared each physician’s interpretation after viewing CheXpert results alone with their interpretation after exploring the same case in CheXplain.Physicians first interpreted the AI output, then interacted with CheXplain and described their understanding again.
- Evaluation procedure: After interacting with CheXplain, all physicians described AI as comparing cases with normal, abnormal, or prior images, whereas nearly all could not interpret AI from CheXpert results alone.Three physicians also identified where AI was looking in the image.
- Evaluation procedure: Participants rated feature helpfulness from 1 to 5, and Figure 5 grouped the features for comparison.The evaluation also collected think-aloud responses and in situ feedback about how each feature supported understanding.
Overall feedback: how CheXplain enables understanding AI
Physicians found image-based comparisons most helpful for understanding AI, while interconnected information supported navigation across detail levels and filters focused attention. The evaluation also exposed confusion about which images AI actually used and differing preferences for information breadth.
- Image-based comparisons: Image-based comparisons were rated the most helpful feature category because examples provided evidence for interpreting AI results that otherwise appeared in isolation.Comparisons across time revealed changes from prior images, while cross-patient comparisons surfaced common features for the same observation.
- Image-based comparisons: Comparing regions, prior images, and other patients helped physicians assess AI observations and identify differences or recurring features across cases.Contrastive regions helped physicians judge normality and similarity, while temporal comparisons highlighted changes that were difficult to notice in isolation.
- Image-based comparisons: Image sources could confuse physicians when they were unsure whether prior images were inputs to AI, and cross-patient features did not always generalize across patients.The authors recommend clearly separating AI inputs from contextual comparisons and grouping cross-patient cases by variance.
- Interconnection among information: Interconnected information gave physicians a roadmap for navigating CXR results at different levels of detail, with low-level annotations considered especially helpful.Impressions were more controversial because tracing them to lower-level observations did not always match physicians’ expectations of explanation.
- Information filters: Information filters helped physicians focus on relevant observations and reduce interaction time, but physicians differed in how much information they wanted to see.The authors suggest making urgency-based filtering customizable to personal preferences.
DESIGN RECOMMENDATIONS FOR MEDICAL AI SYSTEM
The authors organize medical AI design recommendations around motivation versus constraint and explanation versus justification. They recommend clinically targeted, customizable access to heterogeneous evidence at multiple abstraction levels while recognizing unresolved imaging-AI implementation challenges.
- Motivation vs. Constraint: Physicians should be able to target AI understanding to a specific clinical problem rather than pursue general, open-ended understanding.Question input can narrow the scope of analysis, while overly deep technical explanations may constrain attention to clinically relevant aspects.
- Motivation vs. Constraint: Medical AI should expose diverse data sources because differential diagnosis seldom relies on a single modality.The authors contrast this with separate XAI systems for multiple input modalities and recommend a mixed-modality approach.
- Motivation vs. Constraint: Physicians should control how much information they process, from critical results to hierarchical explanations and regional, prior, or cross-patient examples.This recommendation extends beyond a simple urgency toggle to accommodate different time and information preferences.
- Explanation vs. Justification: Explanations describe intrinsic processes producing outputs, whereas justifications draw on extrinsic information such as prevalence, contrastive images, and patient history.The authors treat both categories as methods for helping physicians understand AI.
- Explanation vs. Justification: Medical data analysts may need more explanation, while medical data consumers may need more justification for AI results.Radiologists tended to value low-level annotations, whereas referring physicians cared more about extrinsic validity.
- Explanation vs. Justification: Explanation and justification should be available at multiple abstraction levels, but medical imaging AI still must explain how it interprets regions and why images are clinically similar or different.The paper identifies natural-language accounts of regional interpretation and clinically meaningful image comparison as implementation challenges.
LIMITATION
The study is constrained by speculative survey and evaluation settings, limited participant availability, and incomplete CXR data. These boundaries may produce slight differences from actual diagnostic workflows.
- AI-radiologist survey responses were inevitably speculative because AI adoption in current medical practice remains limited.
- The participant count was limited by the scarcity of medical professionals’ availability.
- The CXR dataset lacked patients’ medical history and additional laboratory results, restricting further exploration with CheXplain.
- The final system evaluation was speculative rather than conducted in a real workplace because of strict regulations.
- The results may differ slightly from the actual diagnostic workflow.