Source-linked AI summary

Explaining Explanations in AI

Brent Mittelstadt, Chris Russell, Sandra Wachter

arXiv:1811.01439v1cs.AI

TL;DR

The paper examines a gap between xAI methods and explanation research, arguing that simplified models approximate complex decision systems rather than straightforwardly provide explanations. It reviews xAI methods alongside work in philosophy, cognitive science, and social science, concluding that interactive, contrastive, selective, and socially responsive approaches could better support accountability and contestation.

  • Problem

    The paper addresses the gap between xAI methods and explanation research, particularly whether approximate models should be treated as explanations.

  • Method

    The paper reviews xAI explanation methods and compares them with accounts of explanation from philosophy, cognitive science, and the social sciences.

  • Results

    The review finds that xAI methods are generally more akin to scientific modelling than explanation giving, while human explanations tend to be contrastive, selective, and socially interactive.

  • Takeaways & Limitations

    The paper argues that xAI should develop interactive post-hoc interpretability methods that facilitate contestation and informed dialogue among users, developers, systems, and other stakeholders.

  • Takeaways & Limitations

    The relevance of contrastive outputs depends on whether the returned data point corresponds to a fact of interest to the user, and model usefulness depends on the fitted domain capturing relevant examples.

Abstract

from arXiv · show

Recent work on interpretability in machine learning and AI has focused on the building of simplified models that approximate the true criteria used to make decisions. These models are a useful pedagogical device for teaching trained professionals how to predict what decisions will be made by the complex system, and most importantly how the system might break. However, when considering any such model it's important to remember Box's maxim that "All models are wrong but some are useful." We focus on the distinction between these models and explanations in philosophy and sociology. These models can be understood as a "do it yourself kit" for explanations, allowing a practitioner to directly answer "what if questions" or generate contrastive explanations without external assistance. Although a valuable ability, giving these models as explanations appears more difficult than necessary, and other forms of explanation may not have the same trade-offs. We contrast the different schools of thought on what makes an explanation, and suggest that machine learning might benefit from viewing the problem more broadly.

1 INTRODUCTION

The paper frames explainable AI as an accountability response to opaque, complex machine-learning decisions, but argues that simplified approximations are often models rather than explanations. It examines this gap and calls for interactive interpretability methods that support contestation and dialogue.

  • Automated decision-making raises accountability questions for builders about system performance and compliance, and for affected people about fairness and how to obtain better outcomes.
  • Machine-learning systems rely on highly complex black-box functions whose interdependent internal values can make their decisions difficult to explain.
  • xAI and the explanation sciences use “explanation” differently, creating a gap between simplified model approximations and established accounts of explanation.
  • The paper reviews xAI explanation methods and compares them with explanation research in philosophy, cognitive science, and the social sciences.
  • The paper argues that interactive post-hoc interpretability should help stakeholders understand, contest, and discuss algorithmic decisions.

2 A BRIEF PRIMER ON EXPLANATIONS IN PHILOSOPHY

The primer distinguishes interpretability from explanation and situates explanation as information exchanged about model functionality or decision rationale. It also emphasizes that explanations serve different stakeholders and purposes, while the paper focuses on explaining particular models and decisions rather than complete scientific accounts.

  • Interpretability concerns the degree to which a black-box model or decision is comprehensible to humans.
  • Explanation concerns exchanging information about a model’s functionality or a decision’s rationale and criteria with different stakeholders.
  • Explanations can support legal compliance, system debugging, learning, and trust, and can be directed to developers, professionals, or people affected by system outputs.
  • Philosophical accounts distinguish explanations by completeness, including whether they explain an event’s full causal chain and necessity.
  • The paper focuses on explanations for particular decisions, events, models, or applications rather than full scientific explanations of general relationships.

3 EXPLAINABLE AI

xAI commonly uses interpretable approximations and post-hoc methods to describe black-box model behaviour, but these approximations resemble scientific models more than complete explanations. Their usefulness depends on domain, accuracy, and recipient expertise, while local methods can mislead when their limitations are unknown.

  • 3 EXPLAINABLE AI: Interpretability concerns human comprehensibility, whereas explanations exchange information about model functionality or decision rationale with different stakeholders.The paper distinguishes transparency, which concerns internal functioning, from post-hoc interpretation, which concerns observed behaviour.
  • 3 EXPLAINABLE AI: xAI often retrofits simplified linear, gradient-based, or decision-tree models onto complex algorithms as global or local approximations.Global approximations cover all possible datapoints, whereas local approximations describe a restricted slice or a few datapoints.
  • 3.1 Scientific Modelling and Explainable AI: These approximations are useful for expert prediction and pedagogy over restricted domains but can mislead lay recipients when presented as explanations of complete model functioning.The paper compares them to scientific models: coarse representations that can be useful without capturing full system behaviour.
  • 3.1 Scientific Modelling and Explainable AI: Local approximations can answer what-if and contrastive questions within a reliable domain, but their value depends on documenting and understanding where they break down.They can function as an explanation kit for expert exploration, prototyping, or debugging.
  • 3.2.1 Linear Models in Continuous Spaces: Linear approximations miss curvature and variable interdependencies, so their accuracy depends on the selected domain and variable values.A feature’s sensitivity may change with the size of the change, while dependencies determine how features jointly affect outcomes.
  • 3.2.2 Gradient Sensitivity verses Binarization: Binarization and image baselines impose domain choices that can substantially change feature importance and the resulting interpretation.Contrasting images with grey or blurred versions can respectively suggest that colour or texture is irrelevant.
  • 3.3 Exploring alternatives to scientific modelling: Because local approximations face generalizability and arbitrariness problems, their reliability for non-experts and affected individuals is highly questionable.The paper argues that other explanation methods may be preferable from the user’s perspective.

4 CONTRASTIVE EXPLANATIONS

Contrastive explanations answer why one outcome occurred rather than another, matching empirical accounts of how people seek and evaluate explanations. In xAI, directly computed alternatives can avoid some approximation problems, but their usefulness depends on selecting relevant contrasts.

  • Human explanations are contrastive: People commonly seek explanations for why event P occurred instead of event Q, reflecting a psychological preference for contrastive explanations.Perceived abnormality also influences which events prompt explanation requests.
  • Human explanations are selective: Everyday explanations are selective: explainers emphasize a manageable subset of causes according to the recipient’s context and purpose.In xAI, this selection often appears as emphasized features or evidence weighted by influence on a prediction.
  • Human explanations are social: Explanations are socially interactive, requiring information to be tailored through dialogue or other forms of exchange between explainers and explainees.This interaction can include disagreement about which causes are relevant.
  • Contrastive explanations in xAI: Contrastive xAI methods provide an alternative data point directly instead of asking users to interpret a model approximating functional values over a restrictive domain.These alternatives can be computed exactly, avoiding approximation-quality and domain-limit challenges to a comparable degree.
  • Contrastive explanations in xAI: The main limitation is relevance: a single contrastive data point may not address the fact or justification the user needs.The same issue affects fitted models when their chosen domain omits examples relevant to the intended audience.

5 TOWARDS COMMUNICATIVE, CONTRASTIVE EXPLANATIONS

The paper argues that useful explanations should support not only information exchange but also critical discussion of whether algorithmic decisions are justified. This requires communicative methods that are contrastive, selective, social, and contestable.

  • Communicative risks: Explanation choices can shape recipients’ beliefs, and conflicts may arise when explainers seek trust while explainees seek non-intuitive causes or grounds for evaluation.The paper warns that malicious explainers could use this asymmetry to discourage critical questioning.
  • Explanation as argumentation: Conversational explanations function as argumentation by offering causes while supporting claims about their truth or relevance.This creates a link between explanation and justification as forms of discourse.
  • Justification and contestability: Because justification is tied to comprehending and contesting decisions, algorithmic explanations should make the justifiability of specific decisions debatable.The paper connects this requirement to ethical acceptability and democratic resolution of legitimate disagreements.
  • Justification and contestability: xAI rarely addresses the link between interpretability and contestability, leaving a gap in explanations intended for affected parties and other non-insiders.The paper identifies contesting and post-hoc auditing abnormal events as requirements for closing this gap.
  • Requirements for xAI: Approximation-based explanations should disclose their limitations, addressed domain, and rationale for choosing that domain.The paper also calls for retained records that support contesting errors or inaccurate input data.
Loading 1811.01439v1…