Source-linked AI summary

Towards Human-centered Explainable AI: A Survey of User Studies for Model Explanations

Yao Rong, Tobias Leemann, Thai-trang Nguyen, Lisa Fiedler, Peizhu Qian, Vaibhav Unhelkar, Tina Seidel, Gjergji Kasneci, Enkelejda Kasneci

arXiv:2210.11584v5cs.AIcs.HC

TL;DR

XAI user evaluations need better evidence about users’ needs and the human effects of explanations, because automated measures are difficult to compare and may not reflect human preferences. The paper systematically reviews 97 recent core papers, categorizes their evaluations, and finds inconsistent effectiveness across applications while identifying sparse integration of cognitive and social sciences. It proposes guidelines for human-centered user studies and future research directions linking psychological science with XAI.

  • Problem

    Automated XAI evaluation measures are difficult to compare and may not reflect human preferences, while human-subject evaluations remain limited.

  • Method

    The paper conducts a systematic literature review of 97 recent core papers with human-based XAI evaluations and categorizes them by trust, understanding, usability, and human-AI collaboration performance.

  • Results

    The review finds that XAI effectiveness in users’ interactions with ML models is inconsistent across applications and that relevant cognitive and social-science perspectives are underrepresented.

  • Takeaways & Limitations

    The paper provides general guidelines for more transparent, comparable, and human-centered XAI user studies and highlights psychological science as a future research direction.

  • Takeaways & Limitations

    The link between proxy-task evaluations and real-world applications remains insufficiently explicit.

Abstract

from arXiv · show

Explainable AI (XAI) is widely viewed as a sine qua non for ever-expanding AI research. A better understanding of the needs of XAI users, as well as human-centered evaluations of explainable models are both a necessity and a challenge. In this paper, we explore how HCI and AI researchers conduct user studies in XAI applications based on a systematic literature review. After identifying and thoroughly analyzing 97core papers with human-based XAI evaluations over the past five years, we categorize them along the measured characteristics of explanatory methods, namely trust, understanding, usability, and human-AI collaboration performance. Our research shows that XAI is spreading more rapidly in certain application domains, such as recommender systems than in others, but that user evaluations are still rather sparse and incorporate hardly any insights from cognitive or social sciences. Based on a comprehensive discussion of best practices, i.e., common models, design choices, and measures in user studies, we propose practical guidelines on designing and conducting user studies for XAI researchers and practitioners. Lastly, this survey also highlights several open research directions, particularly linking psychological science and human-centered XAI.

1 INTRODUCTION

This paper addresses the limited systematic understanding of human-based XAI evaluations and develops practical guidance for conducting user studies. It reviews recent work across evaluation methods, study designs, findings, and future directions for human-centered XAI.

  • Black-box AI systems create an evaluation dilemma in high-stakes domains because their decision-making processes are often not understandable.
  • Automated explanation measures are difficult to compare and may not reflect human preferences, while only about 20 % of XAI evaluation projects consider human subjects.
  • The review examines recent user studies from top-tier HCI and XAI venues using searches that combine explainable-AI and user-study keywords.
  • The paper analyzes 97 core papers, their references, and follow-up citations to identify foundational works, application domains, study designs, measures, and research findings.
  • The review organizes human-based evaluations around trust, understanding, usability, and human-AI collaboration performance.
  • It synthesizes best practices into user-study guidelines and identifies cognitive and psychological sciences as under-investigated areas relevant to human-centered XAI.

3 METHODOLOGY

The methodology categorizes XAI user studies by their objectives and analyzes their foundations, application domains, measures, and findings. This framework highlights both the technical focus of the literature and the limited representation of cognitive and social-science perspectives.

  • The review categorizes collected XAI user studies into groups based on their objectives, distills research questions, summarizes measurement methods, and examines foundational and follow-up papers.
  • Categorization of User-Study Objectives: The four measurement categories are trust, understanding, usability, and human-AI collaboration performance.
  • Categorization of User-Study Objectives: Understanding concerns users’ mental models of how an ML model operates, while usability concerns successful, efficient, and satisfactory task completion.
  • Foundational Domain: The bibliometric analysis identifies model explanations, interpretability, LIME, SHAP, attribution methods, GradCAM, and saliency maps as prominent foundational topics.
  • Foundational Domain: Cognition-related references are comparatively scarce, indicating that research domains focused on human understanding remain underrepresented.
  • Application Domains: Follow-up studies span many applications, with recommendation systems prominent, trust studied in medical diagnosis and transportation, and usability influential in visualization, software development, and education.

4 COMPREHENSIVE USER STUDY ANALYSIS

The survey analyzes XAI user studies across models, explanation techniques, evaluation dimensions, and experimental designs. It distinguishes how studies assess trust, understanding, usability, and model-behavior detection while documenting participant and design choices.

  • Scope and organization: The survey organizes covered studies by AI models, explanation techniques, application domains, evaluation measures, experimental designs, and analysis tools.Table 3 categorizes model types by explanation types, while the analysis covers four measured quantities and experimental designs.
  • Models and explanations: Feature-based explanations, including SHAP and LIME, are the most frequently used explanation type, with local and global scopes distinguished.The survey also identifies intrinsically interpretable white-box models, black-box models, and other explanation types such as rules and game strategies.
  • Evaluation dimensions: User evaluations are categorized by trust, understanding, usability, and human-AI collaboration performance.Trust is measured through self-reported questionnaires or observed agreement with model decisions.
  • Understanding: Objective understanding is assessed with proxy tasks such as forward simulation, in which users predict the model’s output for an input.Relative simulation and manipulation or counterfactual simulation are additional proxies for understanding model behavior.
  • Usability: Usability measures include helpfulness, workload, satisfaction, ease of use, and detection of undesired system behavior.Studies examine whether explanations help users identify unfairness, bias, and incorrect model decisions using objective performance measures or fairness perceptions.
  • Experimental designs and participants: Slightly above 55 % of surveyed user studies use between-subjects designs, while within-subjects studies commonly recruit fewer participants and often target restrictive expert populations.Between-subjects studies range from around 30 participants to 1070 across three conditions or 1250 across five conditions; expert studies may include fourteen medical professionals or five radiologists.

5 FINDINGS OF USER STUDIES

Across 97 core papers, explanations show mixed effects across trust, understanding, usability, and human-AI collaboration. Benefits depend on explanation type, evaluation measure, user expertise, and task conditions.

  • Cross-cutting findings: Explanations improve subjective understanding, but their effects on trust and usability remain inconsistent.The survey identifies a positive trend for perceived understanding, while trust and usability findings include positive, null, negative, and mixed effects.
  • Understanding: Objective understanding findings vary by explanation technique, data modality, and proxy task.Saliency maps, counterfactuals, feature importance, and some LIME evaluations improve objective understanding, whereas other techniques do not outperform baselines.
  • Understanding: Users may overestimate their understanding because applying explanations in practice can reduce perceived understanding.The survey links this pattern to the illusion of explanatory depth and distinguishes subjective from objective understanding.
  • Usability: Explanation effects on usability are mixed, with satisfaction improving in some studies but showing no significant change or declining in others.Effects also vary by setting: explanations improved trust and satisfaction in a simulated driving environment but not with real-world data.
  • Undesired behavior detection: Explanations can help users detect unfairness, bias, and model failures, but detection is not guaranteed and assessments of feature relevance vary.Perceived fairness increases when decisions are explained with feature scores or a combination of scores and highlighted features; several studies report no significant explanation effect.
  • Human-AI Collaboration Performance: Viewing explanations can improve human decision accuracy, especially with feature-based explanations for text inputs, but example-based explanations may not help text classification.Performance gains differ between novices and experts, and explanations can improve novices’ performance while decreasing experts’ performance.

6 A GUIDELINE FOR XAI USER STUDY DESIGN

The guideline organizes XAI user-study practice into decisions made before, during, and after data collection. It emphasizes defining appropriate measures, using rigorous comparisons and proxy tasks, maintaining data quality, and reporting study details transparently.

  • Before the User Study: Design studies around clearly defined quantities, using established constructs such as trust or application-specific outcomes such as human-AI collaboration performance.Definitions should reflect both social-science and technical conceptualizations where relevant.
  • Before the User Study: Compare explanation conditions with a baseline without explanations to establish the overall strength of XAI rather than only selecting a winning technique.Random explanations can serve as an additional comparative baseline when appropriate.
  • Before the User Study: Choose proxy tasks that are manageable yet preserve important characteristics of the target application, while monitoring their difficulty.Forward simulation has been criticized as unrealistically complex in some computer-vision settings; feature-importance queries and manipulatability checks are alternatives.
  • Before the User Study: Use self-reported and observed measures in parallel because subjective ratings may not reflect behavior and different understanding measures can diverge.Users may report trusting a model while not following its suggestions, and objective and subjective understanding can also be weakly correlated.
  • Before the User Study: Plan studies with preregistered variables, hypotheses, exclusion criteria, and sample sizes, and use expert interviews or think-aloud pre-studies to refine designs.Preregistration helps provide evidence against selective reporting or p-hacking, while preliminary qualitative work can inform explanation-system and study design.
  • During and After the User Study: Execute and analyze studies systematically by planning contingencies, matching participant numbers and backgrounds to the design, checking attention, randomizing within-subject conditions, and using appropriate statistical and reliability tests.The guideline covers logistics before participation, quality checks during collection, statistical analysis after collection, and reliability assessment for aggregated measures.
  • After the User Study: Report participant allocation, recruitment, consent, incentives, treatment conditions, and descriptive data so readers can assess the study’s explanatory power.Transparent reporting should include total participants and group assignments as well as characteristics of the collected data.

7 FUTURE RESEARCH DIRECTIONS

The paper proposes increasingly human-centered XAI research through user-centered design, psychological theories, careful evaluation, and emerging interactive explanation methods. It identifies confounders, personalization needs, proxy-task limitations, and the current inability of simulated evaluations to replace human studies.

  • Towards Increasingly User-Centered XAI: User-centered methods should shape XAI design as well as evaluate finished systems, so solutions better respond to user needs.
  • Psychology and pedagogy: Psychological and pedagogical frameworks can inform explanations by modeling how users learn, form mental models, and differ in motivation and reasoning.The paper discusses expectancy-value motivation theory, theory of mind, and hybrid teaching.
  • Emerging explanation methods: LLMs open research directions for interactive natural-language explainers and for using textual explanations as subsequent inputs.
  • Evaluation design: Trust evaluations require controlling confounders because model accuracy and displayed accuracy can influence trust beyond explanation faithfulness.
  • Evaluation design: Explanations should calibrate trust, but revealing model weaknesses can produce negative explanation ratings and complicate fairness evaluation.
  • Personalization: One-size-fits-all explanations overlook personal bias, while representative samples and user models can support more tailored XAI evaluation and design.
  • Sequential interaction: Sequential explanations may shape perception and understanding differently from one-time exposure, especially in recommendation systems.The distinction between single-use and sequential settings remains insufficiently investigated.
  • Evaluation alternatives: Proxy-task outcomes may diverge from real-world decision-making, and SimEvals can help select explanations but cannot yet replace human evaluation.Simulated evaluations do not adequately capture factors such as cognitive biases.

8 CONCLUSION

The survey finds that XAI effectiveness in interaction with machine-learning models is inconsistent across applications. It synthesizes prior designs and findings into guidelines and future directions for more transparent, human-centered research.

  • XAI effectiveness in users’ interaction with machine-learning models is not consistent across applications.
  • The paper analyzes prior design patterns and findings to propose general guidelines and future research directions for human-centered XAI user studies.It presents this work as a starting point for more transparent and human-centered XAI research.

APPENDIX A DATA-DRIVEN BIBLIOMETRIC ANALYSIS

The bibliometric analysis maps the foundations and influence of XAI user studies by examining references and follow-up citations. It identifies relevant research domains and nascent areas for future work, including cognition-driven analysis tools.

  • The analysis groups papers by topic using references extracted from the studied papers and keywords assigned to each paper.
  • Figure 4 maps foundational research domains and the research domains influenced by human-centered XAI core papers.
  • Examining foundations and impact reveals pertinent future areas such as cognition-driven analysis tools in XAI.
  • The survey analyzes over 3000 references and focuses on sources cited by at least ten core papers, approximately 50 papers.
  • The paper distinguishes references as sources contained in core-paper bibliographies and citations as follow-up works referencing core papers.
  • The analyzed literature includes works motivating XAI through mental-model support, prior user-study templates, and broader research on user trust.

APPENDIX B MODELS AND EXPLANATIONS IN XAI USER STUD-

Black-box models dominate the surveyed human-AI interaction research, with local feature explanations such as LIME and SHAP used frequently. Explanation types also vary by application domain.

  • Black-box models are studied more frequently, and local feature explanations including LIME and SHAP are popular in XAI user studies.
  • Recommendation systems use application-specific explanation types, including content-based and hybrid explanations.

C.1 Trust

Trust is measured primarily through self-reported questionnaires, with 7-point and 5-point Likert scales commonly used. Objective trust is often assessed through human agreement rates.

  • Self-reported trust is most often measured with questionnaires using 7-point or 5-point Likert scales.
  • Many studies create their own questionnaires to measure user trust.
  • Objective trust is frequently measured using the agreement rate of human participants.

C.2 Usability

Usability evaluations cover workload, helpfulness, satisfaction, undesired-behavior detection, and ease of use. Subjective perceptions are commonly measured with questionnaires, while debugging tasks can use objective accuracy and completion-time measures.

  • C.2 Usability: Usability is divided into workload, helpfulness, satisfaction, undesired behavior detection, and ease of use.
  • C.2 Usability: Questionnaires commonly measure subjective perceptions of workload, helpfulness, satisfaction, and ease of use.
  • C.2 Usability: Debugging usability can be measured objectively through answer-confirmation accuracy and task-solving time.

C.3 Understanding of Explanations

Understanding evaluations test whether users can process novel or cognitively challenging explanations as intended. Common approaches include comprehension questions, assignment tasks, and natural-language description of discovered concepts.

  • C.3 Understanding of Explanations: Novel or cognitively challenging explanations are evaluated by testing whether users can use the information provided.
  • C.3 Understanding of Explanations: Understanding tests are often combined with other measures to assess whether explanations are correctly processed by users.
  • C.3 Understanding of Explanations: Conceptual explanations are evaluated using questions about semantic coherence, assignment tasks, and natural-language describability.
  • C.3 Understanding of Explanations: Contrastive-learning feature vectors can produce clusters reported as almost as interpretable as human labels.

APPENDIX D FINDINGS

The reviewed studies often compare explanation types without an explanation-free baseline. The paper connects transparent AI evaluation to pedagogical and educational-psychology perspectives on how users learn model behavior.

  • APPENDIX D FINDINGS: Many studies compare explanation types without including a control group that receives no explanation.
  • APPENDIX D FINDINGS: Some researchers justify omitting explanation-free controls by arguing that explanations have already demonstrated usefulness.
  • APPENDIX D FINDINGS: The paper reviews pedagogical frameworks to inform future transparent AI systems and human-centered evaluations.
  • APPENDIX D FINDINGS: Human interaction with XAI interfaces involves learning about model inner workings through explanations and achieving model understanding.

E.2 Theory of Mind

Theory of Mind in XAI concerns how users form mental models of machine-learning systems from observed explanations and examples. Related work uses explanation sequencing and teaching strategies to shape or support these user models, while the survey identifies unresolved questions about cognitive realism.

  • E.2 Theory of Mind: Users form mental models of machine-learning algorithms from explanations or examples, generalizing observations from a few cases to the broader system.
  • E.2 Theory of Mind: Theory of Mind describes the human ability to infer, rationalize, and summarize the decisions of other intelligent agents.
  • E.2 Theory of Mind: Monte Carlo tree search can identify an informative explanation sequence under the assumption that some explanations are more effective initially.
  • E.2 Theory of Mind: The survey reports that XAI still lacks evidence on whether BToM is robust or realistic relative to human cognitive processes.
  • E.2 Theory of Mind: The authors advocate probabilistic and computational cognitive models, interdisciplinary expertise, and grounded user-centered solutions for systems such as robots and autonomous vehicles.
  • E.2 Theory of Mind: XAI teaching methods can either present representative decision examples directly or provide tools for users to explore an AI system independently.
Loading 2210.11584v5…