Source-linked AI summary
Should We Trust (X)AI? Design Dimensions for Structured Experimental Evaluations
Fabian Sperrle, Mennatallah El-Assady, Grace Guo, Duen Horng Chau, Alex Endert, Daniel Keim
TL;DR
XAI evaluation lacks a structured way to compare how users, explanations, models, tasks, environments, and study designs shape trust and bias. The paper reviews comparative and application studies to derive design dimensions and a dependency model, finding that explanation trust, fidelity, and realistic high-impact settings remain under-evaluated. It concludes that these dimensions and the dependency model can guide more holistic XAI evaluations.
Problem
XAI evaluation must determine whether explanations are valid, calibrate trust appropriately, introduce bias, and apply across users, data types, tasks, and environments.
Method
The paper reviews XAI evaluation research and application papers, derives five design-dimension groups, and develops a dependency model of bias propagation and trust building.
Results
The review finds that studies often evaluate model trust rather than explanation trust, rarely manipulate explainer fidelity, and commonly use low-impact settings rather than realistic high-impact environments.
Takeaways & Limitations
The design dimensions and dependency model support more holistic evaluation of XAI systems by exposing under-covered sources of bias and trust propagation.
Takeaways & Limitations
The review relies on keyword searches of titles and abstracts, does not disambiguate terminology, and does not perform a complex meta-evaluation of reported effects.
Abstract
from arXiv · showhide
This paper systematically derives design dimensions for the structured evaluation of explainable artificial intelligence (XAI) approaches. These dimensions enable a descriptive characterization, facilitating comparisons between different study designs. They further structure the design space of XAI, converging towards a precise terminology required for a rigorous study of XAI. Our literature review differentiates between comparative studies and application papers, revealing methodological differences between the fields of machine learning, human-computer interaction, and visual analytics. Generally, each of these disciplines targets specific parts of the XAI process. Bridging the resulting gaps enables a holistic evaluation of XAI in real-world scenarios, as proposed by our conceptual model characterizing bias sources and trust-building. Furthermore, we identify and discuss the potential for future work based on observed research gaps that should lead to better coverage of the proposed model.
1 Introduction
The paper argues that XAI evaluation needs structured comparison across users, explanations, models, tasks, environments, and study designs because trust and validity depend on more than explanation generation alone.
- XAI spans the machine-learning process from raw data through model behavior and outputs presented to users, motivating evaluation across the complete pipeline.Explanations can support trust and help discover, improve, control, or justify machine-learning models.
- Evaluation must examine whether explanations are valid, calibrate trust appropriately, introduce bias, and apply across data types, tasks, users, and environments.Answering these questions requires comparing how explanations are generated, designed, and presented to different user types.
- The authors propose five design-dimension groups—personal characteristics, explanation, model, task and environment, and study setup—for structured XAI evaluation.These groups are derived from a literature review and are intended to characterize evaluation designs and identify research gaps.
- The review distinguishes comparative studies from application papers to clarify how different research communities evaluate XAI methods.The paper reviews evaluation research in human-computer interaction and visualization alongside application papers presenting user evaluations of XAI effectiveness.
2 Background and Related Work
Prior work treats trust as multidimensional and offers evaluation taxonomies and design guidance, but these contributions leave open how trust-related factors should be evaluated across XAI studies.
- Trust is described as a multidimensional construct requiring calibration and recalibration over time.Robbins instead frames trust through actor beliefs, another actor’s trustworthiness, the matter at hand, and unknown outcomes.
- Meta-analyses identify factors influencing trust in robots and automation, but do not emphasize how those factors can be evaluated for human trust in AI.The paper suggests these dimensions may nevertheless be relevant to AI trust and can be grouped with the proposed XAI design dimensions.
- Existing XAI work includes cognitive-bias-aware design guidance and a taxonomy of application-grounded, human-grounded, and functionally-grounded evaluation studies [12].These works discuss how evaluation types may be appropriate, while the present paper focuses on structuring experimental design dimensions.
3 Literature Review
The literature review identifies design dimensions for structured XAI evaluation and exposes gaps between intended properties, measured outcomes, and realistic usage contexts.
- Methodology: The review covers comparative studies and application papers, coding dimensions across XAI design, evaluation, users, tasks, and environments.It draws on selected HCI, visualization, machine learning, and AI publications, using iterative coding refined through coder agreement.
- Limitations: The review’s definitions form a union of author-reported concepts because conflicts, ambiguities, and overlaps were not resolved.Application papers also often state goals such as trustworthiness without evaluating whether systems satisfy them.
- Comparative studies: Expertise was measured in 8 studies, with 7/8 published in 2019, while information content was the most common explanation dimension at 9/17.The review therefore identifies opportunities to vary explanation detail and personalization for different users.
- Comparative studies: Only two studies manipulate explainer fidelity, and Eiband et al. find little difference in trust between real and placebic explanations.The review also notes a gap between reported understanding and understanding demonstrated through tasks or quizzes.
- Discussion of findings: Comparative experiments mainly examine low-impact recommendations or social feeds, leaving higher-impact, higher-criticality settings and realistic usage conditions underrepresented.The authors expect better simulation of actual conditions to affect results, especially when impact and criticality are high.
- Discussion of findings: Explanation fidelity should be matched to audience and task: high fidelity is essential for expert systems, while example-based explanations may suit casual recommendation users.The authors caution that wrong explanations can mislead users, and incomplete demographic reporting limits generalization across cultures.
- Application papers: Application papers usually present dataset case studies or proof-of-concept demonstrations, with little direct testing of explanations and only one paper discussing explanation fidelity.They nevertheless support more complex tasks, and 18/35 papers target users with high machine-learning expertise.
4 Design Dimensions for Experimental Evaluations
The paper synthesizes design dimensions for structured XAI evaluation across users, explanations, models, and task environments. These dimensions support comparable study designs by specifying who uses the system, what is explained, which models are involved, and how use is situated.
- 4.1 User Attributes: User attributes distinguish immutable personal characteristics from experience, including expertise that should be assessed through years of practice and knowledge tests rather than self-rating alone.Participants can be classified as novice, intermediate, proficient, or expert, while domain, technical, and machine-learning expertise capture different aspects of experience.
- 4.2 Explanations: Explanation dimensions cover availability, information content, trustworthiness, effectiveness, fidelity, reasoning strategy, and transparency.These dimensions support varying whether explanations are present, what information they contain, how trustworthy or convincing they are, whether they match model decisions, and how their reasoning is presented.
- 4.3 Models: Model dimensions ask which AI models are used and include simulatability and interpretability, with interpretability defined as understanding why a system behaves as it does under given circumstances.Interpretability involves forming and checking a mental model of system behavior rather than treating it as a single monolithic property.
- 4.4 Tasks and Environment: Task and environment dimensions characterize how and where models and explanations are used, including user tasks and the costs of misuse.The framework represents task requirements alongside impact and criticality, while treating effort as an additional cost dimension.
- 4.4 Tasks and Environment: The task taxonomy orders seven general tasks by increasing difficulty: application (A), understanding (U), simulation (S), diagnosis (D), refinement (R), justification (J), and comparison (C).The taxonomy is intended to simplify comparisons while distinguishing explanation requirements across different kinds of work.
5 Structuring the Design Space
The paper structures XAI evaluation through a dependency model linking stakeholders, data, models, explainers, users, bias, and trust. Bias can propagate through dependencies, whereas trust develops through users’ interactions with outputs and moves backward through the process.
- 5 Structuring the Design Space: The model supports holistic evaluation by making bias sources, dependencies, and trust-building relationships explicit without requiring excessive modeling complexity or loss of detail.Its central design aim is broad coverage of the XAI process while retaining a manageable level of abstraction.
- 5.1 Dependency Model: The dependency model covers stakeholders and building blocks in XAI systems, but deliberately omits interactions and iterative feedback loops.Model and explainer providers, instances, training data, outputs, and users occupy connected stages whose relationships define the evaluation context.
- 5.2 Bias Propagation: Biases from data, model, explainer, or user-specific experience can enter at different blocks, accumulate, and propagate along the dependency arrows.A biased training dataset prevents a fair, unbiased, high-quality model; explanations may reveal bias without reducing it and can miscalibrate trust even when they faithfully represent an insufficient model.
- 5.3 Trust Building: Trust develops from concrete interactions with data, model outputs, and explanations before extending to the model and its providers.Users first assess whether outputs work and explanations make sense, then build trust in the model and eventually in the model instance provider, with personal biases affecting these judgments.
- 5.3 Trust Building: The review reports a disconnect between communities, with HCI emphasizing trust-building in presentation and AI/ML emphasizing model correctness.The authors argue that evaluations should broadly cover the XAI process and avoid introducing information about one dependency component without accounting for related biases and cross-effects.
6 Discussion and Implications
The discussion proposes broader, better-defined, and less biased XAI evaluations that cover the full process and distinguish trust in explanations from trust in models. It also identifies review limitations, including keyword-based coverage, unresolved terminology, and a preliminary dependency framework.
- Better Coverage of the XAI Process: End-to-end XAI studies should span multiple process stages because machine learning, HCI, and application-focused visual analytics currently cover different parts of the process.Application studies provide rich interactions but are problem-specific and often collect only qualitative feedback.
- Clearly Defined Terminology: Researchers should adopt clearly defined terminology and measure whether claimed properties such as interpretability or accountability are actually achieved.Shared terminology would help align field-specific goals with measurable dimensions, approximations, and proxies.
- Trust and Bias: XAI evaluations should account for bias in trust-building and distinguish trust in explanations from trust in the underlying models.Unverifiable explanations may merely shift uncertainty from the model to another black-box explanation stage.
- Fidelity and Motivation: Evaluation design should consider explainer fidelity and users’ motivation because model complexity, placebic explanations, task criticality, and effort can affect responses.The discussion also recommends designing study constellations that reflect why users interact with XAI and how much effort they are willing to exert.
- Limitations: The review is limited by keyword-based venue coverage, unresolved terminology, absence of meta-evaluation, and a preliminary dependency framework that omits model implementations and interaction affordances.The authors propose broader reference-based reviews and framework extensions as future work.
- Evaluation Measures: Most comparative studies assess model trustworthiness with Likert questionnaires, whereas simulatability and weight of advice provide implicit alternatives that reduce bias from users’ preconceptions.Implicit measures avoid adding a post-usage questionnaire step between system use and trust assessment.
7 Conclusion
The paper derives design dimensions from a literature review to structure user-centered XAI evaluation and identifies gaps in how abstract goals are assessed. Its dependency model maps stages where bias and trust can propagate, guiding future evaluations.
- Conclusion: The review distills design dimensions for comparing user-centered evaluations and identifies research gaps, especially the limited assessment of abstract goals such as interpretability and accountability.The dimensions and dependency model together guide future evaluations of XAI systems.