Source-linked AI summary
Metrics for Explainable AI: Challenges and Prospects
Robert R. Hoffman, Shane T. Mueller, Gary Klein, Jordan Litman
TL;DR
The paper asks how to determine whether XAI explanations work and whether users achieve a pragmatic understanding of AI. It integrates research literatures with psychometric evaluations to propose measurement approaches, concluding that evaluation should address explanation quality, user understanding, curiosity, trust, reliance, and human-XAI performance. The paper also emphasizes that measures and their decision thresholds remain subject to development and policy.
Problem
The paper addresses how to measure whether XAI explanations work well and give users a pragmatic understanding of AI systems.
Method
The authors integrate research literatures with conceptual measurement guidance and psychometric evaluations of explanation assessment instruments.
Results
The paper presents a measurement framework spanning explanation goodness, satisfaction, mental models, curiosity, appropriate trust and reliance, and human-XAI performance.
Takeaways & Limitations
XAI evaluation should examine whether explanations support appropriate user understanding, trust, reliance, and performance rather than treating explanation quality as the only outcome.
Takeaways & Limitations
The authors state that their ideas and methods are not final, and that XAI still needs policy-derived metrics for interpreting measurements and making decisions.
Abstract
from arXiv · showhide
The question addressed in this paper is: If we present to a user an AI system that explains how it works, how do we know whether the explanation works and the user has achieved a pragmatic understanding of the AI? In other words, how do we know that an explanainable AI system (XAI) is any good? Our focus is on the key concepts of measurement. We discuss specific methods for evaluating: (1) the goodness of explanations, (2) whether users are satisfied by explanations, (3) how well users understand the AI systems, (4) how curiosity motivates the search for explanations, (5) whether the user's trust and reliance on the AI are appropriate, and finally, (6) how the human-XAI work system performs. The recommendations we present derive from our integration of extensive research literatures and our own psychometric evaluations.
1 Introduction
XAI addresses the difficulty of justifying increasingly complex and autonomous AI decisions by making explanation measurable. The paper connects explanation, mental models, performance, trust, and reliance within a conceptual evaluation process.
- 1 Introduction: Complex machine-learning and deep-network systems make it difficult for decision makers to understand and justify AI-generated results.This creates a need for explanations that support reasonable judgments about the full process.
- 1 Introduction: XAI builds on longstanding work on explanation across expert systems, intelligent tutoring, philosophy, psychology, education, and human factors.The paper positions this prior literature as a source of ideas for current XAI research.
- 1 Introduction: The paper asks how to measure whether explanations work well and whether users achieve a pragmatic understanding of AI systems.It focuses on measurement concepts for evaluating XAI systems and human-machine performance.
- 1.1 Key Measurement Concepts: Initial instruction forms a user’s mental model, while later experience and explanations refine it toward better performance and appropriate trust and reliance.Figure 1 presents this as a conceptual process linking instruction, experience, mental models, performance, trust, and reliance.
2 Explanation Goodness and Satisfaction
The paper distinguishes explanation goodness, judged independently from explanation content, from explanation satisfaction, judged by users in context. It proposes checklists and psychometrically evaluated scales to assess these dimensions.
- 2 Explanation Goodness and Satisfaction: Explanation is interactional: what counts as an explanation depends on the user’s knowledge, needs, goals, and context.Users’ questions can function as triggers expressing the kind of explanation they seek.
- 2.1 Explanation Goodness: The Explanation Goodness Checklist supports a priori evaluation of clarity and precision by independent judges who did not create the XAI system.It can guide explanation design or evaluate explanations generated by an XAI system.
- 2.2 Explanation Satisfaction: Explanation Satisfaction is users’ contextualized, a posteriori judgment of how well they feel they understand the explained AI system or process.Satisfaction can differ from an explanation’s independently judged goodness.
- 2.2 Explanation Satisfaction: The Explanation Satisfaction scale targets understandability, satisfaction, detail, completeness, usefulness, accuracy, and trustworthiness.These attributes were incorporated into an initial pool of Likert-scale items and evaluated for reliability and validity.
- 2.3 Scale Validation: Content Validity: Cronbach’s alpha was 0.86, with item-total correlations from 0.41 to 0.76 and an average inter-item correlation of 0.71, indicating strong internal consistency.The scale was brief and easy to administer and score.
- 2.4 Scale Validation: Discriminant Validity: Content and discriminant validity analyses led the authors to conclude that the Explanation Satisfaction scale is valid, although one confusing reverse-scored item was reconsidered.Eight volunteers’ ratings were generally higher for good explanations (M = 42.75) than bad explanations (M = 30.00), with a large effect size (Cohen’s d = 1.5).
3. Measuring Mental Models
In XAI, mental models are users’ understandings of AI systems, and evaluating them requires eliciting, representing, and analyzing those understandings. The paper recommends combining multiple elicitation methods because performance on one task may not align with performance on another.
- 3.1 Users' Models of the Computer versus the Computer's Models of Users: Mental models in XAI represent users’ understanding of an AI system, distinct from computer-generated user models that adapt system operations or interactions.The paper focuses on eliciting information about users’ mental models rather than automatically generating user models.
- 3.2 Empirical Assertions: Empirical evidence, guided reflection, and convergent elicitation methods can reveal people’s understanding despite limits in their ability to report it completely or clearly.Diagrams can also support understanding of dynamic and complex systems.
- 3.3 Overview of Mental Model Elicitation Methods: Mental-model elicitation methods include self-explanation, glitch detection, prediction, ShadowBox comparison, and diagramming.These tasks respectively probe users’ understanding, errors in explanations, predicted outcomes, differences from expert understanding, and concepts or relations.
- 3.3 Overview of Mental Model Elicitation Methods: Elicitation methods trade off speed, analytical effort, and completeness: some provide a quick window into mental models, while free responses require content analysis and may be incomplete.Methods can also take time to create and require clear rationale for selecting prediction cases.
- 3.4 Application to the XAI Context: For XAI, the preferred method elicits mental models quickly, produces analyzable data, and helps users recognize both sound and limited understanding.The paper frames this recognition as insight into knowledge shields that preserve reductive understandings.
- 3.4 Application to the XAI Context: Because different task performances may diverge, XAI evaluations should use more than one mental-model elicitation method.The paper cautions that adequacy of a user-generated diagram did not necessarily match better prediction-task performance.
4. Measuring Curiosity
Curiosity motivates explanation seeking when users recognize gaps in their knowledge, but explanations can either support insight or suppress curiosity. The paper therefore recommends measuring situation-specific curiosity during XAI use.
- 4.2 The Nature of Curiosity: Explanations may promote curiosity and support insight and better mental models, but they can also suppress curiosity and reinforce flawed mental models.Suppression may occur when explanations overwhelm users, restrict questioning, induce reticence, or increase confusion and complexity.
- 4.2 The Nature of Curiosity: Assessing users’ feelings of curiosity can inform evaluation of XAI systems.The paper treats curiosity as relevant because explanation seeking is driven by curiosity.
- 4.1 Introduction: Curiosity in XAI arises when learners recognize gaps, violated expectations, or incomplete understanding that explanations might resolve.Closing the gap and achieving insight can make successful explanation or self-explanation seem feasible.
- 4.3 Measuring Curiosity in the XAI Context: General curiosity scales are poorly suited to XAI because XAI curiosity is situation- or task-specific and concerns computational-device workings.The paper notes that apparently relevant instruments may measure numerical fluency, pervasive curiosity style, or everyday information seeking instead.
- 4.3 Measuring Curiosity in the XAI Context: A brief questionnaire administered when users request explanations can identify the triggers motivating their curiosity.The paper presents this as a “quick window” into curiosity during XAI use.
- 4.3 Measuring Curiosity in the XAI Context: Curiosity responses can constrain explanation generation, identify what needs explaining, reveal curiosity suppression, and quantify curiosity depth by the number of triggers selected.These uses connect curiosity measurement to both system adaptation and evaluation.
5 Measuring Trust in the XAI Context
Trust in XAI should be measured as an evolving, context-dependent relationship that distinguishes justified trust, mistrust, and reliance. The paper recommends adapting existing scales while repeatedly assessing how users’ trust changes across trials and situations.
- Trust and reliance: XAI users should know when and why to trust, mistrust, or cautiously rely on system recommendations.Trust and mistrust may both be justified, depending on tasks, goals, contexts, and problem situations.
- Trust as exploration: Trusting XAI is exploratory rather than a single stable state or decontextualized metric.Users may move among justified trust, unjustified mistrust, and other states as they encounter explanations, errors, and changing conditions.
- Trust measurement scales: Existing trust scales can be reliable, but many are designed for interpersonal or highly specific application contexts.Some scales concern robots, simulations, medical providers, or terrain-image detectors and therefore require adaptation for XAI.
- Trust and reliance: At minimum, trust measurement can separately ask whether users trust machine outputs and whether they would follow machine advice.These correspond to trust and reliance as distinct aspects of the relationship.
- Measurement recommendations: The paper distills overlapping items from existing scales into a recommended set for XAI research.Most items come from the Cahour-Fourzy scale, with additional items from other scales.
- Measurement recommendations: Trust measurement should be repeated after individual trials, explanations, intermediate trial series, and final trials.Episodic measures can track how users maintain trust and move toward appropriate trust over time.
6 Measuring Performance
Performance evaluation should assess the human-machine work system, not the XAI alone. Recommended measures cover task success, user understanding and prediction of AI behavior, trust-linked reliance, controllability, correctability, productivity, learning, and adoption.
- Performance hypotheses: User performance is hypothesized to improve with satisfying explanations and depend on mental-model quality and epistemic trust.Appropriate reliance is expected when users explore the AI system’s competence envelope.
- Performance scope: XAI performance cannot be evaluated separately from user performance or the human-machine system as a whole.The goal is to determine how successfully the combined system conducts its designed tasks.
- Primary task performance: Primary task performance can be measured by successful trials within a prespecified time period.Examples include image categorization, action identification, and emergency rescue operations.
- User performance: User-level evaluation should measure response speed, prediction correctness, hits, errors, misses, false alarms, and explanations of unusual or anomalous AI outputs.Users’ predictions of machine outputs should be examined for both typical and atypical cases.
- Trust and reliance: Appropriate trust and reliance emerge as users encounter difficult cases near the AI’s competence boundary.Appropriate reliance requires knowing when and when not to follow system outputs.
- Work-system performance: Work-system analysis can measure controllability and correctability: whether users can produce intended outcomes and improve machine outputs.These measures address alignment with objective states or users’ judgments of what the machine should determine.
- Work-system performance: Work-system evaluation can compare productivity with current practice, analyze learning curves, and assess how readily stakeholders adopt the XAI.Adoption is described as a particularly direct indicator of practical performance.
7. Prospects
The paper distinguishes explanation goodness from learner satisfaction and advocates a multi-method measurement approach for the complex human-XAI interaction. It also identifies unresolved challenges around interpreting measurements as application-specific metrics.
- The paper treats its proposed measurement ideas and methods as provisional and open to refinement and extension.
- The authors distinguish explanation goodness as an a priori researcher evaluation from explanation satisfaction as an a posteriori learner evaluation.
- A multi-method approach is advocated because human-XAI interaction is a complex cognitive system.
- The paper does not address metrics, distinguishing measurements from application-context thresholds used to interpret them.A measurement describes how to measure; a metric supports decisions such as whether performance is superior, acceptable, or poor.
- Metrics do not emerge directly or easily from theoretical concepts or operationalized measures; they come from policy.
- Resolving the metrics challenge may depend on more XAI projects reaching full performance evaluation.
Appendix A Explanation Goodness Checklist
Appendix A presents a Goodness Checklist for judging explanation quality independently and in advance, based on features identified across explanation research literatures.
- The Goodness Checklist catalogs literature-based features that make explanations good as statements.
- Researchers or domain experts can use the checklist to evaluate explanations generated by other researchers or XAI systems.
Appendix B
Appendix B contains materials for evaluating the Explanation Satisfaction Scale, including questions about how users understand and assess explanations of everyday systems.
- Appendix B includes materials used to evaluate the discriminant validity of the Explanation Satisfaction Scale.
- The materials ask how cell phones provide directions and include both detailed and simplified explanations of that process.
- The materials ask how automobile cruise control works and provide a detailed explanation involving speed sensing, control signals, the throttle, and engine operation.
- The cruise-control explanation states that the system maintains engine speed until disengaged and adjusts for the car gear.
- The materials ask how computers predict hurricanes and provide explanations based on atmospheric modeling and historical hurricane paths.
Appendix C
Appendix C presents the Explanation Satisfaction Scale, a validated user-focused instrument covering understanding, satisfaction, detail, completeness, usability, accuracy, and trust judgments.
- The Explanation Satisfaction Scale asks users whether they understand how the software, algorithm, or tool works.
- Users rate whether the explanation is satisfying, sufficiently detailed, and complete.
- Users rate whether the explanation tells them how to use the system and supports their goals.
- Users rate whether the explanation shows system accuracy and helps them judge when to trust or not trust it.
APPENDIX D
Appendix D presents representative trust-scale material for XAI, including definitions, factors, sample items, and an experience-based administration condition.
- Trust and distrust are defined as sentiments shaped by knowledge, beliefs, emotions, and experience that generate positive or negative expectations about system interactions.
- The scale analyzes trust into reliability, predictability, and efficiency.
- Representative items ask whether users feel confident in the XAI system and judge it predictable, reliable, safe, and efficient.
- The scale assumes considerable prior experience with the XAI system, making it more appropriate after use than immediately after an explanation.
3. Is the [tool] reliable? Do you think it is safe?
This section reviews trust scales for reliability and safety, comparing their factors, item content, overlap, and suitability for evaluating XAI systems.
- Several scales contain overlapping or redundant items, including repeated questions about confidence, reliability, and trust.
- The authors recommend incorporating Jian et al.’s item 4 into the XAI version of the Cahour-Fourzy Scale because it lacks a counterpart there.
- The reviewed scales measure trust through factors including reliability, technical competence, understandability, confidence, reliance, safety, and familiarity.
- The Madsen-Gregor Scale’s technical competence and understandability items may reference users’ mental models rather than trust alone.
- The human-robot collaboration scale is problematic for XAI because it uses anthropomorphic behaviors and overlaps with explanation-satisfaction evaluation.
- A scale merging trust and reliance presupposes prior experience and is therefore considered inappropriate when users are first learning an XAI system.
Appendix E
Appendix E presents a recommended XAI trust scale covering confidence, predictability, reliability, safety, efficiency, wariness, performance, and liking, while noting its experience requirement and evidential basis.
- The recommended scale directly assesses whether users find the XAI confident, predictable, reliable, efficient, and believable.
- The scale is intended for administration after considerable XAI use rather than immediately after an explanation and before use experience.
- Most recommended items are adapted from the Cahour-Fourzy Scale, with additional items drawn from the Jian, Schaefer, and Madsen-Gregor scales.
- The authors infer reliability and content validity from overlap and semantic similarity with existing scales reported as reliable.
- Its items additionally assess safety, wariness, performance relative to a novice human user, and liking the system for decision making.