Source-linked AI summary

Explainable AI is Dead, Long Live Explainable AI! Hypothesis-driven decision support

Tim Miller

arXiv:2302.12389v3cs.AIcs.HC

TL;DR

The paper argues that recommendation-driven XAI can fail because people either distrust or blindly follow machine advice, while explanations may not engage them or preserve their control. It proposes Evaluative AI, a machine-in-the-loop framework that supplies evidence for and against human-selected options. The paper concludes that this approach better supports weighing alternatives and human expertise, while acknowledging unresolved engagement and cognitive-load challenges.

  • Problem

    Recommendation-driven XAI can lead to under-reliance or over-reliance, and contrastive explanations mainly compare alternatives with the machine’s recommendation rather than supporting broader trade-offs.

  • Method

    The paper proposes Evaluative AI, in which decision support tools help people filter options, generate hypotheses, and request evidence for and against their own judgements.

  • Results

    The paper concludes that Evaluative AI gives decision makers control to explore the strengths and weaknesses of options while retaining explainable AI as a broader paradigm.

  • Takeaways & Limitations

    AI-assisted decision-support tools should follow Prudence by supporting human critique and weighing alternatives rather than only justifying machine recommendations.

  • Takeaways & Limitations

    The framework still assumes that people will attend to machine-provided evidence, and it may require more engagement and cognitive effort than recommendation-driven approaches.

Abstract

from arXiv · show

In this paper, we argue for a paradigm shift from the current model of explainable artificial intelligence (XAI), which may be counter-productive to better human decision making. In early decision support systems, we assumed that we could give people recommendations and that they would consider them, and then follow them when required. However, research found that people often ignore recommendations because they do not trust them; or perhaps even worse, people follow them blindly, even when the recommendations are wrong. Explainable artificial intelligence mitigates this by helping people to understand how and why models give certain recommendations. However, recent research shows that people do not always engage with explainability tools enough to help improve decision making. The assumption that people will engage with recommendations and explanations has proven to be unfounded. We argue this is because we have failed to account for two things. First, recommendations (and their explanations) take control from human decision makers, limiting their agency. Second, giving recommendations and explanations does not align with the cognitive processes employed by people making decisions. This position paper proposes a new conceptual framework called Evaluative AI for explainable decision support. This is a machine-in-the-loop paradigm in which decision support tools provide evidence for and against decisions made by people, rather than provide recommendations to accept or reject. We argue that this mitigates issues of over- and under-reliance on decision support tools, and better leverages human expertise in decision making.

1 Introduction

The paper contrasts recommendation-driven XAI, which persuades people to accept machine answers, with Evaluative AI, which supports human-controlled exploration through evidence for and against options. It argues that this approach better aligns decision support with human decision making, while remaining most suitable for decisions where people have accountability and time to explore.

  • Motivation: The paper uses Bluster and Prudence to contrast recommendation-giving with feedback that helps decision makers assess strengths and weaknesses while retaining control.A survey during a talk found that just three of over 100 people preferred Bluster, while the remaining audience preferred Prudence.
  • Problem: Recommendation-driven AI can produce under-reliance or over-reliance because people struggle to calibrate trust in imperfect decision aids.Under-reliance leaves tools ineffective, whereas over-reliance leads people to follow recommendations even when they are wrong.
  • Proposed framework: Evaluative AI avoids machine recommendations and instead presents evidence for or against human-selected options, allowing decision makers to determine which hypotheses to explore.The framework is intended to support filtering unlikely options, generating hypotheses, and examining trade-offs across options.
  • Proposed framework: By giving decision makers control over which options receive feedback, Evaluative AI is proposed as a better fit for cognitive decision processes than persuasive explanations of machine recommendations.The paper presents this as a paradigm shift from persuading users to accept an answer toward helping them critique their own ideas.
  • Scope: Evaluative AI is positioned for medium- and high-stakes decisions with human accountability, especially when decision makers have time to explore options.The paper does not intend this framework for every scenario; Bluster-like recommendations may remain useful for low-stakes or time-limited decisions.

2 Background and related work

The paper frames good decision support around helping people evaluate options and understand tool behavior, then grounds explainable AI in abductive reasoning and cognitive decision research. It reviews evidence that recommendation-driven XAI often fails to improve decisions, while arguing that decision support should leverage both intuition and structured thinking.

  • Decision making and decision support: Good decision support should help people decide across multiple cardinal issues, while remaining understandable about how and why it works and where it fails.Understandability is presented as important for calibrating trust and finding mistakes.
  • Decision making and decision support: A decision support tool supports the process of deciding but need not provide answers, extending beyond the recommendations and judgements typically emphasized in XAI.The criteria are intended to support decision makers rather than take over the decision.
  • Cognitive processes for decision making: Abductive reasoning involves forming hypotheses and judging their likelihood to explain observations, with the process potentially iterating as new evidence appears.The paper treats this process as a foundation for conceptualizing explainable AI.
  • Cognitive processes for decision making: Expert decision makers use intuition to narrow options, then search for evidence supporting and refuting each hypothesis, a process the framework builds around the Data/Frame model.The paper connects this sensemaking account to abductive reasoning.
  • Explainable/interpretable AI and decision making: Studies report little or inconsistent decision-making impact from recommendation-driven XAI, alongside over-reliance, under-reliance, and limited engagement with explanations.Some studies found explanations increased over-reliance, while observational work found no statistically significant improvement in decision making.
  • Explainable/interpretable AI and decision making: The paper cautions against prioritizing system 2 thinking over intuition because naturalistic decision-making research describes intuition as a powerful source of fast and sometimes better decisions.It instead argues that decision support should exploit intuition, expertise, and structured thinking.

3 (Explainable) AI as decision support

The paper evaluates four recommendation-centered decision-support paradigms and finds that each provides limited support beyond the machine’s preferred option. Explainability and cognitive forcing add understanding or engagement, but remain recommendation-driven and leave important decision-making needs unsupported.

  • Approaches evaluated: The paper compares four paradigms: recommendations without explanations, recommendations with explanations, intrinsically interpretable models, and cognitive forcing.These approaches are evaluated against criteria for good decision support.
  • Giving recommendations with no explanatory information: Recommendation-only decision support can produce unwarranted distrust or trust, causing people to ignore useful advice or accept wrong decisions.The approach also provides no support for scrutinising why the decision was made.
  • Giving recommendations with explanatory information: Explainability adds understanding of the machine decision, but contrastive explanations mainly compare the recommendation with one foil rather than support broader trade-offs.The paper describes this as persuasive and limited to comparing each option with the recommended one.
  • Intrinsically interpretable models: Interpretable models provide understanding of the machine decision, but do not themselves calculate options, possibilities, stakeholder values, or trade-offs for the decision maker.The paper treats this as overlapping with explainability rather than sufficient decision support on its own.
  • Cognitive forcing: Cognitive forcing withholds recommendations to prompt engagement, partially supports option filtering and trade-offs, and provides understanding of the machine decision.Its weaknesses are that it remains recommendation-driven and offers little decision-making information for alternatives to the recommendation.
  • Transition to evaluative AI: The paper therefore proposes building on cognitive forcing while grounding decision support in the cognitive science of decision making.This proposal leads into the evaluative AI framework developed in the next section.

4 Evaluative AI: A conceptual framework of explainable decision support

The paper introduces evaluative AI as a conceptual framework for hypothesis-driven explainable decision support. It is designed to support human cognitive decision making while giving decision makers greater control over which options and information to explore.

  • Framework: Evaluative AI is a conceptual framework for hypothesis-driven explainable decision support.The framework is presented as the paper’s proposed model.
  • Design criteria: Its first design criterion is to support the properties of good decision making by supporting the decision maker’s cognitive process.The framework is explicitly grounded in the cognitive decision-making process.
  • Design criteria: Its second design criterion is to provide the decision maker with better internal locus of control over which options to explore and when.Control concerns both the choice of options and the timing of exploration.
  • Framework vision: The framework’s vision is to help decision makers access the information they want and need to evaluate a hypothesis when they want it.Information access is organized around the decision maker’s hypothesis and timing.

4.1 Conceptual framework

Evaluative AI supports decision makers by filtering options, generating hypotheses, and providing evidence for and against hypotheses while preserving human control over exploration.

  • Evaluative AI framework: Evaluative AI filters out unlikely options or generates new hypotheses, then provides evidence for and against the hypotheses selected by the decision maker.Evidence may also compare one hypothesis with another.
  • Evaluative AI framework: The framework keeps the decision maker in control of which hypotheses to explore rather than necessarily providing recommendations.This is intended to align decision support with human decision making.
  • Option support: Presenting multiple context-specific options can reduce fixation by avoiding a single recommendation.A probabilistic classifier could highlight options within a selected probability range or use uncertainty estimates.
  • Trade-off support: Unlike persuasive contrastive explanations, evaluative AI explains trade-offs between any two option sets and provides evidence for and against each option.The evidence is not restricted to the option judged most likely.
  • Trade-off support: Good decision makers test initial conclusions by seeking negative evidence, whereas recommendation-driven approaches typically emphasize the recommendation and omit evidence against it.Research with anaesthesiology residents found that testing an initial conclusion for negative evidence was associated with the best decisions.

4.2 Example: Diagnosis

The diagnosis example illustrates how evaluative AI narrows seven possible skin-disease diagnoses to likely hypotheses and presents supporting and opposing evidence for the decision maker’s final judgment.

  • Example: Diagnosis: For skin-lesion diagnosis, the system begins with seven disease categories and filters them to three likely hypotheses: melanoma, basal cell carcinoma, or actinic keratosis.The example uses a dermoscopic image and metadata such as lesion location.
  • Example: Diagnosis: Showing several likely diagnoses instead of one is intended to mitigate fixation on a single option.The prototype highlights the retained hypotheses with bold text.
  • Example: Diagnosis: For basal cell carcinoma, lesion location, colour, scarring, and occasional bleeding support the diagnosis, while asymmetry, itchiness, and recent colour change count against it.The decision maker integrates these forms of evidence to make the final decision.

4.3 Summary

The paper characterizes evaluative AI as human-centred decision support that helps explore options, support judgment, understand machine decisions, and perform trade-offs without determining stakeholder values.

  • 4.3 Summary: Evaluative AI helps provide new options or filter unlikely ones, identify possibilities, support judgment, enable trade-offs, and explain machine decisions.It does not help determine stakeholder values.
  • 4.3 Summary: Evaluative AI explicitly supports options, judgment, understanding, and trade-offs, distinguishing it from other explainable decision-support approaches.Table 4 compares five approaches to explainable decision support.
  • 4.3 Summary: The framework follows the Data/Frame model by allowing decision makers to explore hypotheses rather than receive only information justifying a machine recommendation.It also hands control over which hypotheses are investigated and prioritized to the decision maker.
  • 4.3 Summary: Evaluative AI aims to retain the benefits of cognitive forcing while letting decision makers examine strengths and weaknesses of any option.This contrasts with examining only recommendation strengths and alternative weaknesses.

4.4 Long live explainable AI!

The paper proposes shifting from recommendation-driven to hypothesis-driven decision support without rejecting explainable AI, because evaluative AI remains an XAI paradigm and existing techniques remain useful.

  • 4.4 Long live explainable AI!: The paper proposes a pivot from recommendation-driven decision support to hypothesis-driven decision support, while retaining explainable AI.The authors describe recommendation-driven XAI as unsuitable for some decision-making situations, not universally.
  • 4.4 Long live explainable AI!: Recommendation-driven approaches may remain appropriate for some applications, including decisions made at scale.The paper also notes that XAI has applications beyond decision support, such as verification, regulation, and scientific discovery.
  • 4.4 Long live explainable AI!: Evaluative AI requires an underlying decision-making model, alongside techniques such as machine learning, planning, optimization, and their explainability methods.This requirement is presented as an assumption for machines to judge decisions.
  • 4.4 Long live explainable AI!: Human-to-human explanation properties—contrastive, selected, interactive, and causal—also apply to providing evidence and evaluating trade-offs in evaluative AI.The framework therefore remains connected to established cognitive and social aspects of explainability.
  • 4.4 Long live explainable AI!: Existing XAI tools such as contrastive explanation, Weights of Evidence, feature importance, and case-based reasoning can contribute evidence to evaluative AI.The paper presents existing work as a foundation for developing new XAI models.

4.5 Challenges and Limitations

Evaluative AI assumes people will attend to machine-provided evidence, but it may require more engagement and cognitive effort than recommendation-driven support. Its design challenge is balancing human control with manageable information demands.

  • Evaluative AI still assumes that decision makers will pay attention to what a machine says.
  • Evaluative AI may not reduce cognitive load because it forces more engagement with decisions.This could make it less preferred by decision makers, despite supporting cognitive reasoning.
  • Abductive reasoning can reduce information demands by presenting likely hypotheses and prioritising important information.The authors identify balancing reduced information with increased engagement as an unresolved challenge.

4.6 An incomplete research agenda

The proposed research agenda develops evaluative AI around abductive reasoning, from observing events and generating options to judging plausibility and recording decisions. It identifies interface, uncertainty, evidence, and documentation mechanisms for supporting these stages.

  • The agenda frames evaluative AI as an incomplete research programme based on the framework.
  • Observing events: Interfaces can clarify events and data while highlighting anomalous behaviour that may matter for decisions.
  • Generating options: Machine-generated options can reveal overlooked possibilities and accelerate time-sensitive decisions by filtering or systematising assessment.
  • Generating options: Decision support can provide probabilities, likely option sets, and uncertainty measures without necessarily issuing a recommendation.Likely option sets may narrow hypotheses while reducing fixation compared with a single recommendation.
  • Judging plausibility: Plausibility assessment can combine explanations, evidence weights, epistemic and aleatoric uncertainty, and evidence selection.
  • Recording decisions: Decision records should capture outcomes, supporting and opposing evidence, rejected alternatives, and human-used evidence unavailable to the tool.Further research is needed on supporting decision makers during reevaluation.

5 Conclusion

The conclusion presents evaluative AI as a machine-in-the-loop alternative to recommendation-driven decision support, while acknowledging that it increases cognitive work. The authors argue that this trade-off may be necessary for improving human engagement with AI-assisted decisions.

  • Evaluative AI puts decision makers in control and helps them weigh different options better than recommendation-driven support.The conclusion still allows that recommendation-driven support can be useful for low-stakes or time-limited decisions.
  • Evaluative AI is not a panacea because its demand for more work may lead people to prefer easier decisions with worse results.The authors describe this as a possible side effect of cognitive engagement that may need to be accepted.
  • The paper concludes that justified recommendations represent a dead paradigm for some situations, while hypothesis-driven explainability keeps explainable AI relevant.
Loading 2302.12389v3…