Source-linked AI summary
Explainable AI: Beware of Inmates Running the Asylum Or: How I Learnt to Stop Worrying and Love the Social and Behavioural Sciences
Tim Miller, Piers Howe, Liz Sonenberg
TL;DR
The paper argues that explainable AI risks being designed for researchers rather than intended users. It surveys related work and social-science foundations, finding limited grounding in explanation science and rare behavioural evaluation, while acknowledging the survey’s limited scope.
Problem
Explainable AI may prioritize researchers’ understanding of models over intended users’ needs for explanations.
Method
The paper conducts a lightweight literature scan and presents relevant explanation research from philosophy, psychology, cognitive science, and human factors.
Results
The scan found limited use of social-science research in explainable AI and rare human behavioural experiments.
Takeaways & Limitations
Explainable-AI researchers and practitioners should collaborate with social and behavioural scientists to inform model design and behavioural evaluation.
Takeaways & Limitations
The literature survey was illustrative rather than comprehensive, covering a small, workshop-derived set of papers.
Abstract
from arXiv · showhide
In his seminal book `The Inmates are Running the Asylum: Why High-Tech Products Drive Us Crazy And How To Restore The Sanity' [2004, Sams Indianapolis, IN, USA], Alan Cooper argues that a major reason why software is often poorly designed (from a user perspective) is that programmers are in charge of design decisions, rather than interaction designers. As a result, programmers design software for themselves, rather than for their target audience, a phenomenon he refers to as the `inmates running the asylum'. This paper argues that explainable AI risks a similar fate. While the re-emergence of explainable AI is positive, this paper argues most of us as AI researchers are building explanatory agents for ourselves, rather than for the intended users. But explainable AI is more likely to succeed if researchers and practitioners understand, adopt, implement, and improve models from the vast and valuable bodies of research in philosophy, psychology, and cognitive science, and if evaluation of these models is focused more on people than on technology. From a light scan of literature, we demonstrate that there is considerable scope to infuse more results from the social and behavioural sciences into explainable AI, and present some key results from these fields that are relevant to explainable AI.
1 Introduction
The paper argues that explainable AI risks designing explanations for expert researchers rather than intended users. It proposes grounding explanation models and evaluations in social and behavioural science.
- Causal explanation is fundamentally social interaction: someone explains something to someone else through conversation.This view makes explanation sensitive to conversational rules rather than treating it solely as a model output.
- Explainable AI risks repeating software-design failures when experts decide what constitutes a good explanation for complex models.The paper compares this risk with programmers designing software for themselves rather than target users.
- The paper scans explainable-AI literature for influence from philosophy, psychology, cognitive science, and human factors, and for human behavioural evaluations.The scan examines 23 articles listed as workshop-related work.
- The paper presents relevant bodies of social and behavioural research on explanation and discusses their potential impact on explainable AI.These areas include research on how people generate, select, present, and evaluate explanations.
2 Explainable AI Survey
A lightweight survey of workshop-related papers found limited use of social-science research and rare human behavioural evaluation in explainable AI. The authors therefore argue for stronger integration of these sciences while acknowledging the survey’s limited scope.
- 2.1 Selected Papers: The survey examined 23 workshop-related articles, excluding one cognitive-science survey because it was not an explainable-AI paper.The list was community-compiled and objective from the authors’ perspective, but far from comprehensive.
- 2.2 Categorisation: The lightweight survey scored papers for topical relevance, social-science grounding, and validation using human behavioural studies.Data-driven scores distinguished references to social-science explanation research from algorithms explicitly derived from it.
- 2.3 Results: Only four on-topic papers referenced relevant social-science research, and only one built a model on it; serious human behavioural experiments were rare.Off-topic papers showed similarly limited social-science input and behavioural experimentation.
- 2.4 Discussion: The results provide evidence that many explainable-AI models do not build on current scientific understanding of explanation, while human behavioural experiments remain rare.The authors connect this gap to the need for more useful explanatory agents.
- 2.4 Discussion: The authors do not equate limited social-science grounding with poor explainable-AI research, noting that some work developed its own understanding through behavioural experiments.They present social-science research as a sound starting point when developing such an understanding is not feasible.
3 Where to? A Brief Pointer to Relevant Work
The paper argues that explainable AI should draw more heavily on social and behavioural science, especially research on how people generate, select, present, and evaluate explanations. It highlights contrastive questions, attribution, causal simulation, explanation selection, conversational norms, and empirically grounded evaluation as relevant foundations.
- Contrastive explanation: Why-questions are contrastive: people seek explanations for why P occurred rather than an expected alternative Q.The contrast case frames relevant answers, although eliciting it from an observer may be difficult.
- Contrastive explanation: Answering a contrastive question can require understanding only the difference between two cases, rather than identifying every cause in the full causal chain.This can make contrastive explanations easier to provide than complete causal attributions.
- Attribution theory: Attribution theory models how people explain behaviour through beliefs, desires, intentions, traits, and distinctions between failed and successful actions.Malle’s framework is linked to deliberative reasoning, belief-desire-intention models, and AI planning.
- Causal connection: People connect causes through mental simulation, using heuristics such as proximity, abnormality, and controllability when full causal chains are infeasible.These heuristics can help explainable AI models skip or discount events while remaining consistent with explainee expectations.
- Explanation selection: People select a small number of causes, preferring proximal events but sometimes tracing toward human actions or abnormal events.For causal chains with more than a handful of causes, explanation selection can simplify or prioritise explanations.
- Explanation evaluation: People judge explanations using pragmatic criteria including simplicity, generality, coherence, and usefulness, sometimes preferring simpler explanations over more likely ones.These criteria may be incorporated as objective criteria for explainable AI models when understanding and acceptance are important goals.
- Explanation as interaction: Social-science research treats explanations as interactive, conversational acts shaped by the explainer, the explainee, and conversational norms.Visual explanations should have similar properties to verbal explanations, with quality, quantity, relation, and manner offering objective criteria.
- Relevant theoretical foundations: The authors recommend building explainable AI models on newer, widely accepted cognitive and social-science models rather than outdated deductive or co-variation theories.The paper describes the logically deductive model and co-variation model as no longer valid models of human explanation in these fields.
4 Conclusions
The paper concludes that human models of explanation are highly relevant to explainable AI but insufficiently represented in the field. It recommends collaboration with social and behavioural scientists and greater emphasis on human evaluation, while allowing proxy studies and computational work where appropriate.
- Conclusions: Existing models of how people generate, select, present, and evaluate explanations are highly relevant to explainable AI.The authors’ brief survey provides evidence that little explainable-AI research draws on such models.
- Conclusions: The authors encourage collaboration with social and behavioural scientists to inform explainable-AI model design and human behavioural experiments.They support human-in-the-loop techniques and respect for the time and effort required for intensive evaluations.
- Conclusions: Proxy studies remain valid for early explanation models, and computational problems remain legitimate subjects of explainable-AI research.The authors do not argue that every explainable-AI paper must include human behavioural experiments.
- Conclusions: The paper hopes researchers adopt existing models and methods to reduce the risk that explainable AI is designed primarily by experts for themselves.
A List of Papers Surveyed
This appendix lists the papers included in the explainable-AI workshop’s related-work survey. The entries span explanation, interpretability, planning, robotics, classifiers, cognitive science, and related AI topics.
- Surveyed papers: The list contains papers on explainable agency, ontology-stream diagnosis, interpretable classifiers, explanation research, document classification, and autonomous-system narration.
- Surveyed papers: Other entries address explainable robots, object recognition, deep neural networks, plan explanations, case-based reasoning, and knowledge systems.
- Surveyed papers: The list also includes work on tactical behaviour, DQNs, plan explicability and predictability, multimedia event detection, interpretability science, visual explanations, explanatory capabilities, interactive machine learning, and concept learning.
B Detailed Results
The detailed-results section presents columns for paper, topical relevance, data-driven grounding, validation, and comments.
- Detailed results: The results table is organized by paper, on-topic status, data-driven status, validation, and comments.
- Detailed results: The table includes an on-topic field for assessing whether each paper concerns explainable AI.
- Detailed results: The table includes data-driven and validation fields alongside comments.