Source-linked AI summary
The Who in XAI: How AI Background Shapes Perceptions of AI Explanations
Upol Ehsan, Samir Passi, Q. Vera Liao, Larry Chan, I-Hsiang Lee, Michael Muller, Mark O. Riedl
TL;DR
The paper addresses how users’ AI backgrounds shape their perceptions and interpretations of AI explanations. Using quantitative and qualitative analyses, it finds that both groups developed unwarranted faith in numbers for different reasons and valued different explanations beyond their intended design.
Problem
Understanding who interacts with an AI system remains underexplored, despite the importance of explainability for informed decisions.
Method
The paper combines quantitative analyses of explanation perceptions with qualitative analysis of how AI background influences interpretation.
Results
Both groups showed unwarranted faith in numbers for different reasons, while AI background produced significant differences in how explanations were perceived and interpreted.
Takeaways & Limitations
The findings motivate XAI design interventions to mitigate over-reliance on numbers and reimagine explanations for different forms of appropriation.
Takeaways & Limitations
The study’s insights should be scoped because its focus on AI students limits the population represented.
Abstract
from arXiv · showhide
Explainability of AI systems is critical for users to take informed actions. Understanding "who" opens the black-box of AI is just as important as opening it. We conduct a mixed-methods study of how two different groups--people with and without AI background--perceive different types of AI explanations. Quantitatively, we share user perceptions along five dimensions. Qualitatively, we describe how AI background can influence interpretations, elucidating the differences through lenses of appropriation and cognitive heuristics. We find that (1) both groups showed unwarranted faith in numbers for different reasons and (2) each group found value in different explanations beyond their intended design. Carrying critical implications for the field of XAI, our findings showcase how AI generated explanations can have negative consequences despite best intentions and how that could lead to harmful manipulation of trust. We propose design interventions to mitigate them.
1 INTRODUCTION
Explainable AI must account not only for how systems are opened, but also for who interprets their explanations. This study examines how AI background shapes perceptions of explanations and identifies risks and design implications.
- AI background is an important user characteristic because differences between creators and end-users can produce usability failures, irresponsible design, and inequities.
- Current XAI deployments often serve AI engineers, creating a consumer-creator gap between intended explanations and users’ actual perceptions.
- The study compares people with and without AI backgrounds to explain how and why their perceptions of AI explanations differ.
- Both groups showed unwarranted faith in numbers for different reasons, while each found value in explanations beyond their intended designs.
- The findings motivate design implications addressing overreliance on numbers, over-trust, AI education, and potentially harmful manipulation of user perceptions.
- The mixed-methods study evaluates natural-language justifications, action descriptions, and numerical explanations across confidence, intelligence, understandability, second chance, and friendliness.
2 BACKGROUND
Background work motivates a human-centered, pluralistic approach to XAI that attends to recipients’ characteristics and cognitive processes. The paper addresses a gap in empirical understanding of how AI background shapes interpretation and use.
- XAI explanation-generation methods make opaque models accessible by producing approximations of behavior, often privileging understandable explanations over direct inspection of mechanisms.
- Rationale generation produces natural-language accounts of agent behavior and is intended to support functional understanding, especially for non-AI experts.
- Users have divergent explanation needs, and explanations can impose cognitive burden, create false security, and engender over-trust, including when justificatory content is absent.
- Heuristics can produce cognitive biases when applied inappropriately, including associations between explanations and AI competence that contribute to over-trust.
- Human-centered XAI calls for pluralistic explanation design that considers different consumers and their needs rather than treating users as homogeneous.
- Prior work leaves limited empirical understanding of differences among XAI users and actionable design guidance, particularly for people with or without AI backgrounds.
- The paper extends user-characteristics research by examining how AI background influences what users perceive as an explanation and how they interpret and use it.
3 STUDY DESIGN AND METHODS
The study uses a 2x3 within-subjects experiment comparing participants with and without AI backgrounds across three explanation types, followed by quantitative and qualitative analysis. Identical robot behavior isolates differences in explanation format and interpretation.
- The experiment compares two participant groups across three explanation types using a 2x3 factorial design and measures perceptions quantitatively.
- Participants rank explanation preferences and provide open-ended justifications that are qualitatively analyzed for group differences.
- The navigation task uses a sequential environment and tabular Q-learning, addressing an under-explored setting for explainability research.
- All three robots follow the same reinforcement-learning action trace and differ only in how they explain their actions.
- The Rationale-Generating robot gives natural-language justifications, the Action-Declaring robot states actions without justification, and the Numerical-Reasoning robot outputs Q-values.
- Q-values expose relative action utility but do not explain why one action has higher utility than another, and the study leaves them unlabeled.
- The study establishes high-contrast AI and non-AI groups using differences in AI knowledge and self-reported programming knowledge: p< 2.2 × 10−16 for both tests.
4 QUANTITATIVE RESULTS
Quantitative analyses found a consistent overall ranking of RG above AD above NR, while AI background significantly shaped preferences for AD and NR across explanation dimensions.
- 4.1 Within-group Comparisons: The AI-background group unambiguously preferred RG over the other robots across dimensions, particularly over AD.RG was the only robot consistently preferred by the AI group in within-group comparisons.
- 4.1 Within-group Comparisons: The non-AI group showed no preference between RG and AD in four of five dimensions; RG won only on Friendliness.The non-AI group preferred AD over NR in every dimension except Intelligence.
- 4.1 Within-group Comparisons: NR exceeded AD on Intelligence only for the AI group, whereas the non-AI group showed no preference between them on that dimension.This result motivated the qualitative analysis of AI participants’ preference for numerical representations.
- 4.2 Between-group Comparisons: Between groups, RG showed no significant preference difference across dimensions, while AD consistently favored the non-AI group.All RG p-values exceeded 0.05, whereas AD p-values were below 0.05 and its odds ratios exceeded 1.0.
- 4.2 Between-group Comparisons: For NR, AI participants ranked it higher than non-AI participants on Confidence, Second Chance, and Understandability.The corresponding odds ratios were below 1.0; no significant group difference appeared for Friendliness or Intelligence in this analysis.
- 4.3 Conclusions of Quantitative Analyses: Rank_RG > Rank_AD > Rank_NR overall, but AI background changed perceptions of explanations across the identified dimensions.The groups were indistinguishable on Friendliness, while differing especially in preferences for AD and NR.
5 QUALITATIVE ANALYSIS & FINDINGS
Qualitative findings show that both groups placed unwarranted faith in numerical explanations, but for different reasons and with different intended uses. Non-AI participants valued AD’s affirmatory statements, whereas AI participants attributed diagnostic value to NR’s numbers.
- 5 QUALITATIVE ANALYSIS & FINDINGS: Both groups exhibited unwarranted faith in numbers, while also finding value in explanations beyond their intended design.The qualitative analysis links these differences to distinct explanatory intents associated with AI background.
- 5.1 Faith in Numbers: The AI group preferred NR over AD for Intelligence, while preferring AD for Understandability, treating numerical representation as evidence of intelligence despite its difficulty.This contrasts with the non-AI group’s greater preference for AD on Confidence and Understandability.
- 5.1 Faith in Numbers: AI participants associated numerical representations with algorithms, logic, intelligence, and trustworthiness even without fully understanding them.Numbers could make NR seem smarter, more real, or theoretically more likely to succeed.
- 5.1 Faith in Numbers: AI participants saw NR’s numbers as potentially actionable for debugging or predicting behavior, even when the numbers’ meanings were unclear.Some believed they could infer patterns or eventually use the numbers despite not understanding them immediately.
- 5.1 Faith in Numbers: Non-AI participants also treated numbers as signals of intelligence, often interpreting their opacity as evidence of precision or higher-order thinking.This helps explain their lack of preference between AD and NR for Intelligence despite preferring AD on other dimensions.
- 5.2 Unanticipated Explanatory Value: Thus, non-AI participants sought confirmation from AD, whereas AI participants over-ascribed diagnostic value to NR’s numbers.The two groups assigned different uses to explanations that were not designed primarily for those purposes.
- 5.2 Unanticipated Explanatory Value: Non-AI participants valued AD’s declarative statements as affirmatory information because its actions and statements appeared aligned and consistent.AD’s concise, direct language also increased perceived understandability and confidence.
6 DISCUSSION & IMPLICATIONS
AI background shaped how participants appropriated explanations: both groups placed faith in numbers, but through different heuristics and explanatory intents. These interpretations produced unanticipated uses and motivate appropriation-aware, user-centered XAI design.
- Heuristics and misplaced faith: AI participants associated numbers with logical intelligence and imagined using them to manipulate, diagnose, or reverse engineer the robot.The paper warns that this heuristic is risky because the numbers were Q-values with limited actionability beyond assessing available actions.
- Heuristics and misplaced faith: Non-AI participants associated their inability to understand complex numbers with higher-order intelligence and lacked the background for deliberative reasoning about them.Thus, different AI-related heuristics led both groups to the same outcome: faith in numbers.
- Appropriation and explanatory intent: Both groups appropriated explanations beyond their intended design, even during passive, one-way interactions where participants were not asked to act on them.AI participants envisioned troubleshooting scenarios, while non-AI participants treated inaccessible numbers as in-actionable and valued declarative confirmation.
- Appropriation and explanatory intent: AI participants often assigned diagnostic value to unclear numbers, whereas non-AI participants used declarative statements as signals of stable performance.These differing explanatory intents were connected to participants’ AI backgrounds and their own sense of how explanations could be used.
- Contributions to XAI: The paper extends XAI research by connecting heuristics and appropriation to AI background and by treating explanations as both products and interpretive processes.It examines explanation products from three robots alongside how AI background influences their interpretation.
- Explanation design implications: Users preferred RG’s language-rich explanations for their reasoning depth, variety, relatability, personality, and perceived humanlikeness, but such explanations may not support mechanistic understanding.RG-style explanations also require human explanation data, which can be difficult to collect in complex environments.
- Explanation design implications: Because users appropriate explanations regardless of careful design, XAI should support rather than control end-users and provide resources that mitigate harmful appropriations.Suggested interventions include awareness of group differences, exposure that calibrates trust in numbers, and appropriation-aware design that learns from user appropriation.
LIMITATIONS & FUTURE WORK
The study offers a formative account of how AI background shapes interpretations of AI explanations, while limiting its claims to specific participants, tasks, explanation types, and a quasi-experimental design.
- The AI-background group came from a largely representative introductory AI course, but differences in curricula may affect participants’ backgrounds and explanation perceptions.
- The study focuses on an RL-based agent performing sequential decision-making in a controlled environment, so findings may not extend to other agents, tasks, or sociotechnical contexts.
- The quasi-experimental setup lacked random assignment, so future randomized controlled trials are needed to establish causal relationships.
- Although the groups were screened to differ in AI background while controlling for age and education, technological familiarity and AI use may remain confounds.
- The study examined three explanation types, which are informative but not exhaustive; future work could compare additional user characteristics and more differentiated AI backgrounds.
8 CONCLUSIONS
The paper examines how people with and without AI backgrounds perceive AI explanations through a mixed-methods study. It finds that preferences overlap, while interpretations, appropriation, and trust in numbers differ across groups.
- The study investigates how people with and without AI backgrounds perceive different AI explanations using a mixed-methods user study.
- Both groups preferred natural-language justificatory rationales, but the reasons behind that preference were nuanced.
- Different explanatory perceptions sometimes reflected different appropriations, such as diagnostic versus affirmatory intent.
- Both groups showed unwarranted faith in numbers, but numerical and incomprehensible reasoning were associated with intelligence through different heuristics.
- The paper proposes design implications for reducing over-reliance on numbers and argues that focusing on users can support more pluralistic, human-centered XAI.
A.1 Best practices for data integrity and participant engagement
The study used multiple participant-engagement and data-integrity practices, including fair compensation, multimodal orientation, attention checks, manual review, and scheduling across time zones.
- Payment: Participants received compensation calibrated to task duration and effort, with the study aiming to meet or exceed the local minimum wage.
- Most participants completed the task in about 30 minutes, while the authors report receiving high-quality data.
- Task environment and setup: Multimodal orientation, attention checks, acknowledgment requirements, timers, and response review were used to support engagement and data quality.
- Task environment and setup: The researchers standardized robot appearances while preserving distinctions to reduce preferential treatment and excessive recall demands.
- Deployment and Review: Manual response review and staggered task releases across a 24-hour cycle supported global participation and screening for low-quality responses.
A.2 Participant Screening
Participants were screened and grouped to create measurably different AI-background populations using knowledge tests, self-reported knowledge, and AI-class history.
- Group assignment combined a five-question knowledge test, self-reported programming and AI knowledge, and confirmation of prior AI classes.
- The non-AI group required no reported programming or AI knowledge and no prior AI classes, whereas AI-group thresholds required higher knowledge and self-reported experience.
- The AI-background group included 96 students taking an AI class, while the non-AI group included 83 adults recruited through Amazon Mechanical Turk.
- AI-group participants scored an average of 4.73 out of 5 on the knowledge test, compared with 0.91 for the non-AI group.
- Statistical tests found the groups significantly different on the knowledge test, with p < 2.2 × 10^-16 after Bonferroni correction.
A.2.1 Screening questionnaire.
The screening questionnaire tested foundational knowledge across programming, machine learning, reinforcement learning, and Markov decision processes. It included output-prediction, task-classification, objective, and Markov-assumption questions.
- The questionnaire screened programming knowledge by asking participants to predict the output of Python programs.
- It tested machine-learning knowledge by asking which task is unsupervised learning.The listed options included classification, clustering without labels, and regression.
- It assessed reinforcement-learning knowledge by asking about its general goal.The response options contrasted maximizing expected reward or punishment with reaching goals or avoiding obstacles.
- It assessed understanding of Markov decision processes by asking what the Markov assumption means.The options concerned how the current state depends on previous states and actions.
Computer Programming Background Knowledge.
The survey measured computer-programming background with a five-level self-report scale ranging from no knowledge to extensive application or cutting-edge software creation.
- Programming background was measured on a five-level scale from no knowledge to a lot of knowledge.
- The scale distinguished awareness without coding, basic concepts without application, and prior coding experience.
- Its highest categories captured frequent application of programming concepts and creation of cutting-edge software.
AI Background Knowledge.
The survey measured AI background using a five-level self-report scale and included a question about AI coursework. The appendix also reports OLR summaries across robot types, groups, and perception dimensions.
- AI background was measured on a five-level scale from no knowledge to frequent application or cutting-edge software creation.
- The scale distinguished awareness without AI knowledge, basic concepts without application, prior coding experience, and more frequent AI use.
- Participants were also asked whether they had taken or were currently taking AI classes.
- OLR Summary Tables: The appendix organizes OLR summaries by robot type and AI-group status across confidence, friendliness, intelligence, potential, understandability, and ranking.