Source-linked AI summary
Understanding the Role of Human Intuition on Reliance in Human-AI Decision-Making with Explanations
Valerie Chen, Q. Vera Liao, Jennifer Wortman Vaughan, Gagan Bansal
TL;DR
The study examines how decision-makers reconcile intuition with AI predictions and explanations when deciding whether to rely on or override AI. Using a think-aloud, mixed-methods study, it identifies three intuition-driven pathways and finds that example-based explanations outperformed feature-based explanations, achieving complementary human-AI performance.
Problem
Empirical evidence on whether AI explanations improve appropriate reliance is mixed, including indications that explanations can increase reliance on incorrect predictions.
Method
A think-aloud, mixed-methods study examined feature- and example-based explanations across two prediction tasks, analyzing how participants used intuition to assess AI predictions.
Results
Example-based explanations reduced overreliance and helped achieve human-AI complementary performance, whereas feature-based explanations did not improve outcomes and increased overreliance on incorrect predictions.
Takeaways & Limitations
Three intuition-driven pathways—about outcomes, features, and AI limitations—describe how decision-makers override AI predictions and inform explanation-method design.
Takeaways & Limitations
The study used two relatively low-stakes, nonrepresentative tasks, a relatively small sample, and a feature-based explanation method that was not completely faithful to the model.
Abstract
from arXiv · showhide
AI explanations are often mentioned as a way to improve human-AI decision-making, but empirical studies have not found consistent evidence of explanations' effectiveness and, on the contrary, suggest that they can increase overreliance when the AI system is wrong. While many factors may affect reliance on AI support, one important factor is how decision-makers reconcile their own intuition -- beliefs or heuristics, based on prior knowledge, experience, or pattern recognition, used to make judgments -- with the information provided by the AI system to determine when to override AI predictions. We conduct a think-aloud, mixed-methods study with two explanation types (feature- and example-based) for two prediction tasks to explore how decision-makers' intuition affects their use of AI predictions and explanations, and ultimately their choice of when to rely on AI. Our results identify three types of intuition involved in reasoning about AI predictions and explanations: intuition about the task outcome, features, and AI limitations. Building on these, we summarize three observed pathways for decision-makers to apply their own intuition and override AI predictions. We use these pathways to explain why (1) the feature-based explanations we used did not improve participants' decision outcomes and increased their overreliance on AI, and (2) the example-based explanations we used improved decision-makers' performance over feature-based explanations and helped achieve complementary human-AI performance. Overall, our work identifies directions for further development of AI decision-support systems and explanation methods that help decision-makers effectively apply their intuition to achieve appropriate reliance on AI.
1 INTRODUCTION
The paper examines how decision-makers’ intuition interacts with AI predictions and explanations when deciding whether to rely on or override AI. It identifies three intuition-driven pathways and contrasts feature-based and example-based explanations.
- Prior explanation studies report mixed effectiveness and sometimes increased overreliance on incorrect AI predictions.
- The study addresses a literature gap by examining decision-making processes and the role of intuition rather than measuring aggregate outcomes alone.
- Here, intuition means beliefs or heuristics based on domain knowledge, experience, instinct, or pattern recognition.
- Feature-based explanations did not improve outcomes and increased overreliance, whereas example-based explanations reduced overreliance and supported complementary human-AI performance.
- The analysis identifies outcome intuition, feature intuition, and intuition about AI limitations as bases for overriding AI predictions.
2 RELATED WORK AND RESEARCH QUESTIONS
Related work frames explanations as tools for appropriate reliance but finds unresolved effectiveness and overreliance concerns. The paper responds by studying intuition and real-time decision processes through a think-aloud mixed-methods approach.
- 2.1 Overview of XAI: XAI methods include interpretable models and post-hoc explanations, which may be global or local depending on whether they summarize overall or instance-specific behavior.
- 2.1 Overview of XAI: Feature-based explanations assign contribution values to features, while example-based explanations show representative or similar training samples and may include predictions and labels.
- 2.2 Related Work: Empirical work has generally not confirmed that explanations improve appropriate reliance, with several studies finding greater overreliance on wrong AI predictions.
- 2.2.1 Can XAI methods improve appropriate reliance?: A central unresolved question is how decision-makers decide to rely on or override AI, even when they engage with explanations.
- 2.2.2 What is the role of human intuition on reliance?: The paper studies intuition as cognitively stored knowledge and beliefs, including domain expertise, and examines how it helps people reason about explanations and model errors.
- 2.3 Research Questions: Using think-aloud data during interaction, the study investigates why feature-based and example-based explanations improve or inhibit appropriate reliance.
3 METHODS
The study uses an online, within-subjects think-aloud design with two prediction tasks and two explanation types. Participants first decide without AI, then evaluate the same instances with randomized AI-supported conditions.
- Study design: The study uses a think-aloud, mixed-methods protocol to examine participants’ decision processes with AI predictions and explanations.
- Participants: Participants were recruited in the US through convenience sampling and diversified across education, job role, and self-reported ML/XAI backgrounds.
- Participants: The sample included 26 participants and was skewed toward highly educated and ML-experienced people.
- Prediction tasks: The income task asks whether a profile indicates income above $50,000, using a random forest trained on 2018 US Census data with 80.8% held-out accuracy.
- Prediction tasks: The biography task asks participants to infer one of five professions from an online biography.
- Explanations and design: Participants received feature-based and example-based explanations in a within-subjects design, with explanation order randomized.
- Procedure: Participants first completed all instances without AI and then saw the same instances with AI support in random order.
- Procedure: Each participant judged 16 instances, including 10 correct and 6 incorrect AI predictions, producing an experienced AI accuracy of 62.5%.
4 RESULTS
Participants used intuition about outcomes, features, and AI limitations to decide when to override AI predictions. Across two tasks, example-based explanations supported complementary performance, whereas feature-based explanations did not.
- Intuition about the outcome: Participants formed outcome intuitions from prior experience, pattern recognition, prominent features, prototypes, and unusual feature values.Strong outcome intuition led participants to discount AI predictions; weak intuition led them to inspect explanations more closely and sometimes defer to AI.
- Intuition about features: Feature-based explanations prompted participants to evaluate feature relevance and weights, using disagreements with model reasoning as evidence against predictions.With example-based explanations, feature intuition also guided judgments about whether examples were genuinely similar and sometimes updated participants’ own feature beliefs.
- Intuition about AI limitations: Participants recognized AI limitations through explanation signals indicating prediction unreliability, including incorrect predictions on similar examples and potentially biased feature patterns.These signals differed by explanation type, with example-based explanations often revealing unreliability when predictions were wrong on similar examples.
- Pathways to non-reliance: The study summarizes three intuition-driven pathways to non-reliance: disagreeing outcome intuition, feature-based critique of explanations, and recognition of AI limitations.These pathways were frequently mentioned in think-aloud data, so their individual effects were not isolated or quantified.
- Exploratory quantitative analysis: 61.1% ± 2.9%, 60.6% ± 4.2%, and 71.1% ± 2.9% were the income-task accuracies for no AI, feature-based explanations, and example-based explanations; biography-task values were 60.0% ± 4.4%, 64.4% ± 4.3%, and 71.2% ± 4.6%.The trained AI models had 62.5% accuracy on the presented instances, making example-based performance higher than human or AI alone on both tasks.
5 DISCUSSION
The discussion interprets three intuition-driven pathways for overriding AI predictions and uses them to explain explanation effects, individual differences, and design recommendations. It also qualifies the findings because think-aloud observation, repeated instances, task and explanation choices, and sample characteristics constrain interpretation and generalizability.
- Intuition-driven pathways: Three pathways support appropriate non-reliance: strong outcome intuition, feature-based reasoning that discredits AI, and recognition of AI limitations through unreliability signals.These pathways provide a framework for understanding when and what explanations may help decision-making.
- Individual differences: For participants with low domain knowledge, example-based explanations improved average phase 3 accuracy from 43.8% phase 2 accuracy to 70.0%, whereas feature-based explanations reached 52.5%.The low-domain-knowledge group comprised five participants, with N=16 used for the overall phase 2 proxy measure.
- Design recommendations: The authors recommend adapting AI support to varied intuition strengths, including personalized information that reflects decision-makers’ confidence.People may have strong intuition for some instances but not others, creating different opportunities for AI support.
- Design recommendations: Explanations should match natural decision rationales because participants sometimes mistakenly overrode AI when feature explanations highlighted trivial disagreements.The paper notes that natural human explanations are predominantly qualitative.
- Design recommendations: Explanation design should help users identify evidence that discredits predictions and form intuition about AI limitations, including through model failure cases or global explanations.Participants naturally sought multiple similar examples with different or incorrect predictions, while uncertainty information may also support limitation awareness.
- Limitations: Think-aloud observation may reduce realism and cannot isolate intuition effects or establish precise causal relations between intuition and decisions.The authors therefore emphasize quantitative trends and themes from observational data.
- Limitations: Repeating the same instances across study phases may have strengthened prior intuition more than a realistic human-AI setting.The authors do not expect the identified intuition types themselves to depend on this design.
- Limitations: The low-stakes tasks, nonrepresentative feature types, unfaithful post-hoc explanations, small sample, and participant profile limit the findings’ completeness and generalizability.Participants were not task experts, and the sample was more highly educated and ML-experienced than the general population.
6 CONCLUSION
The conclusion presents a mixed-methods study of how decision-makers reconcile intuition with AI predictions and explanations. It identifies three override pathways and uses them to explain why the example-based explanations studied better supported appropriate reliance than the feature-based explanations.
- Study focus: The mixed-methods study examined how human decision-makers reconcile their intuition with AI predictions and explanations.The analysis focused on participants’ think-aloud data.
- Findings: Three pathways reduced overreliance: strong outcome intuition, feature-based reasoning that discredits AI, and recognition of prediction unreliability.These pathways describe how participants applied intuition to override AI predictions.
- Findings: Example-based explanations better supported appropriate reliance and complementary human-AI performance than the feature-based explanations used.The paper attributes this to less disruption of outcome intuition, stronger inductive reasoning support, and accurate unreliability signals.
A PARTICIPANT INFO
The participant-information materials comprise demographic information and a rating scale for participants’ self-assessed ML and XAI knowledge.
- Participant information: Table 2 contains demographic information about the study participants.
- Knowledge ratings: Participants rated ML and XAI knowledge from 0, meaning none, to 3, meaning expert.The intermediate ratings represented limited experience at 1 and frequent or day-to-day experience at 2.
B POST-STUDY INTERVIEW QUESTIONS
The post-study interviews asked participants about decision strategies, the effect of AI, reasoning with predictions and explanations, explanation preferences, and comparative views of human and AI performance.
- Interview topics: Participants were asked to describe their decision strategies without AI and the difference introduced by AI in phase 3.
- Interview topics: Interview questions examined how participants reasoned with AI predictions and explanations and compared feature contributions with similar examples.
- Interview topics: Participants were also asked whether they considered themselves better than the AI and would still want AI assistance.
C USER STUDY INSTRUCTIONS
The study instructions present interfaces for feature-based and example-based explanations across the income prediction and biography classification tasks. The variants include cases where participants saw example-based explanations before feature-based explanations.
- Some instruction variants presented example-based explanations before feature-based explanations.
- Figures 4 and 5 show the income prediction interfaces for feature-based and example-based explanations, respectively.
- Figures 6 and 7 show the biography classification interfaces for feature-based and example-based explanations, respectively.