Source-linked AI summary
Explanations Can Reduce Overreliance on AI Systems During Decision-Making
Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael Bernstein, Ranjay Krishna
TL;DR
Overreliance on incorrect AI predictions threatens human-AI decision-making, and prior work has not reliably shown that explanations reduce it. This paper proposes a cost-benefit framework and tests it across five maze-task studies, finding that explanations reduce overreliance under conditions that make verification more worthwhile.
Problem
Human-AI teams often fail to achieve complementarity because people accept incorrect AI predictions, while prior studies have not found that explanations reduce overreliance.
Method
The paper formalizes reliance as a cost-benefit framework and tests it across five maze-solving studies with a simulated AI.
Results
Explanations reduce overreliance when task difficulty, explanation understandability, or monetary benefits make engaging with the task more attractive.
Takeaways & Limitations
The findings support viewing human-AI collaboration as a strategic allocation of effort responsive to the relative costs and benefits of verification.
Takeaways & Limitations
The crowd-sourced studies lack long-term user engagement.
Abstract
from arXiv · showhide
Prior work has identified a resilient phenomenon that threatens the performance of human-AI decision-making teams: overreliance, when people agree with an AI, even when it is incorrect. Surprisingly, overreliance does not reduce when the AI produces explanations for its predictions, compared to only providing predictions. Some have argued that overreliance results from cognitive biases or uncalibrated trust, attributing overreliance to an inevitability of human cognition. By contrast, our paper argues that people strategically choose whether or not to engage with an AI explanation, demonstrating empirically that there are scenarios where AI explanations reduce overreliance. To achieve this, we formalize this strategic choice in a cost-benefit framework, where the costs and benefits of engaging with the task are weighed against the costs and benefits of relying on the AI. We manipulate the costs and benefits in a maze task, where participants collaborate with a simulated AI to find the exit of a maze. Through 5 studies (N = 731), we find that costs such as task difficulty (Study 1), explanation difficulty (Study 2, 3), and benefits such as monetary compensation (Study 4) affect overreliance. Finally, Study 5 adapts the Cognitive Effort Discounting paradigm to quantify the utility of different explanations, providing further support for our framework. Our results suggest that some of the null effects found in literature could be due in part to the explanation not sufficiently reducing the costs of verifying the AI's prediction.
1 INTRODUCTION
Human-AI teams often fall short of complementarity because people accept incorrect AI decisions without verification. This paper argues that explanations can reduce overreliance when their cognitive costs and benefits make verification worthwhile.
- Human-AI teams should outperform either humans or AIs alone, but empirical work has found limited evidence of such complementarity.
- Overreliance occurs when people accept incorrect AI decisions without verifying them and is especially concerning in high-stakes domains.
- Prior experiments found no reduction in overreliance when explanations accompanied predictions, motivating alternative explanations beyond inevitable cognitive bias.
- The proposed cost-benefit framework predicts that explanations reduce overreliance when tasks are more effortful or explanations make verification easier.
- Five experiments with a simulated AI in maze-solving tasks test how task difficulty and explanation engagement affect overreliance.
- The results indicate that overreliance is a strategic decision responsive to the costs and benefits of engaging with the task.
2 BACKGROUND
Background research links overreliance to weak human-AI complementarity and possible cognitive biases, while explanations have not reliably improved performance. The paper therefore asks when people engage effortfully with explanations to verify AI predictions.
- Studies often find that human-AI teams do not achieve complementary performance, even when explanations are provided.
- When AI accuracy exceeds human accuracy, improved team performance may reflect blind trust rather than better understanding of model capabilities.
- In domains with roughly equal human and AI performance, explanations have not improved performance over prediction-only baselines and can increase reliance on incorrect predictions.
- The paper investigates when people engage in effortful thinking to verify AI predictions and reduce overreliance.
- Dual-process theory distinguishes fast intuitive reasoning from deliberate analytical reasoning, with the latter able to override default biases and heuristics.
- Cognitive forcing functions can reduce overreliance on incorrect predictions but also reduce reliance on correct predictions.
3 FRAMEWORK
The framework treats reliance on AI as a strategic choice shaped by the relative costs and benefits of engaging with a task, verifying an explanation, or relying on the AI. It predicts that harder tasks, easier explanations, and higher rewards can reduce overreliance.
- 3 FRAMEWORK: The framework rejects inevitable overreliance and models it as a strategic choice responsive to the benefits of correctness and the costs of extracting the answer.
- 3 FRAMEWORK: People compare cognitive effort and time against benefits such as task performance, professional accomplishment, monetary rewards, or stakes when choosing a decision strategy.
- 3 FRAMEWORK: A strategy becomes more attractive when an element of the task reduces its costs or increases its benefits relative to alternatives.
- 3 FRAMEWORK: Explanations may fail to reduce overreliance when they leave the cost of engaging with the task roughly equal to prediction-only or human-only strategies.
- 3 FRAMEWORK: When tasks are more difficult and explanations reduce engagement costs, verifying with the AI becomes relatively more attractive than either relying or solving the task alone.
- 3.1 Predictions of our cost-benefit framework: The framework predicts greater reductions in overreliance with explanations as task difficulty increases.
- 3.1 Predictions of our cost-benefit framework: It also predicts lower overreliance when explanations are easier to understand and when monetary benefits for correct completion increase.
- 3.1 Predictions of our cost-benefit framework: Subjective utility reflects how a strategy balances benefits against costs, including the cognitive effort reduction supplied by an explanation.
4 METHODS
The study uses a maze-based human-AI collaboration task to manipulate task difficulty, explanation modality, and cognitive effort while measuring overreliance. A simulated AI, controlled for comparable human accuracy, provides predictions and explanations whose formats and errors are experimentally varied.
- Human-AI collaboration task: Participants choose the correct exit from four maze options, making random-choice accuracy 25% and random-guessing overreliance 25% when the AI is incorrect.The task is multi-class classification and is designed to remain difficult enough that AI assistance could help.
- Explanation conditions: The task supports highlight explanations that are easy to check and written explanations that are more difficult to check.Highlight explanations visually indicate the path from the start point to the AI-predicted exit.
- Task difficulty: Maze dimensions of 10 × 10, 25 × 25, and 50 × 50 represent easy, medium, and hard task conditions.Questions were pre-tested on crowdworkers to calibrate difficulty.
- Simulated AI: The experiments use a simulated AI with accuracy controlled at exactly 80%, while pilot human accuracy was 83.5% across difficulty conditions.Incorrect AI predictions are distributed across maze path lengths so participants do not notice a bias in the errors.
- General study design: The study reduces trust-related influences by avoiding anthropomorphism, calling outputs “suggestions,” withholding immediate feedback, and emphasizing that suggestions can be rejected.These design choices aim to attribute differences in overreliance to cost-benefit manipulations rather than individual differences in trust.
5 STUDY 1 – MANIPULATING COSTS VIA TASK DIFFICULTY
Study 1 tested whether task difficulty changes the effect of explanations on overreliance in a maze-solving collaboration with a simulated AI. Explanations reduced overreliance and increased exploratory accuracy only for hard tasks, while easy tasks showed no difference and medium tasks did not support the predicted reduction.
- Design: Study 1 manipulated maze difficulty while varying whether participants received only an AI prediction or an accompanying highlighted path.Easy, medium, and hard mazes were 10 × 10, 25 × 25, and 50 × 50, respectively.
- Results: Overreliance increased as tasks became more difficult in the prediction-only condition.The authors interpret this pattern through the higher cognitive effort required to verify predictions or solve difficult tasks independently.
- Results: In easy tasks, explanations produced similar overreliance levels to predictions alone.The authors argue that explanations did not meaningfully reduce the cost of engaging with an easy task.
- Results: Medium-difficulty tasks did not show the predicted lower overreliance with explanations.The authors suggest the task may not have been difficult enough for explanations to produce the expected cost reduction.
- Results: In hard tasks, explanations reduced overreliance compared with predictions alone.The highlighted assistance substantially reduced the relative effort of engaging with the task.
- Results: Exploratory analysis found that explanations increased decision-making accuracy in the hard task condition.The figure-level result accompanies the reduction in overreliance for hard tasks.
- Summary: Overall, explanations reduced overreliance only when tasks were substantially complex and effortful.The study found no observable reduction in easy or medium tasks but did find one in hard tasks.
6 STUDY 2 – MANIPULATING COSTS VIA INCREASING EXPLANATION DIFFICULTY
Study 2 tested whether the cognitive difficulty of understanding an explanation affects overreliance. Easier-to-parse highlighted explanations reduced overreliance relative to written explanations, while hard-to-understand explanations did not differ from prediction-only assistance across task difficulties.
- Design: The study hypothesized that easier-to-understand explanations would produce lower overreliance than harder-to-understand explanations.The manipulation targeted the cognitive cost of engaging with the task through the explanation.
- Design: Study 2 manipulated explanation difficulty using highlighted paths that were easier to parse and written path descriptions that required more effort.Written explanations listed directional steps alongside the maze, whereas highlights overlaid the proposed path directly on it.
- Results: Overreliance was lower with highlight explanations than with written explanations in medium-difficulty tasks.This supported Hypothesis 2a.
- Results: Overreliance was lower with highlight explanations than with written explanations in hard tasks.This supported Hypothesis 2b.
- Results: Across easy, medium, and hard tasks, written explanations did not differ from prediction-only assistance in overreliance.The authors state that written explanations therefore did not function as an additional trust signal in these exploratory analyses.
- Summary: Study 2 shows that overreliance responds both to the effort of completing a task and to the effort of understanding an explanation.The result extends the cost-based account beyond task difficulty alone.
7 STUDY 3 – MANIPULATING COSTS VIA DECREASING EXPLANATION DIFFICULTY
Study 3 tested whether making explanation mistakes easier to notice reduces overreliance. In an exploratory hard-task experiment, salient explanations produced the lowest overreliance, reaching an average rate of 0%.
- Study design: Study 3 placed participants in a hard maze task with prediction-only, highlight, written, incomplete, or salient explanations.Salient explanations highlighted the incorrect continuation after the AI crossed a wall; incomplete explanations stopped at the wall.
- Study design: The exploratory study asked how much overreliance could be reduced by explanations that make AI mistakes easy to spot.The study was not preregistered because it explored a possible floor effect and followed prior confirmation of explanation-difficulty hypotheses.
- Results: The figure comparison shows lower overreliance for more salient explanations and no observed floor effect across conditions.Figure 11 reports that participants did not overrely when they knew the AI was wrong.
- Results: Salient explanations produced less overreliance than prediction, highlight, written, and incomplete explanations.Incomplete explanations also produced less overreliance than prediction-only and written explanations, but more than salient explanations.
- Results: 0% was the average overreliance rate in the salient explanation condition, with no observed floor effect preventing further reduction.The authors qualify this initial result by noting that more participants are needed to assess whether it generalizes.
8 STUDY 4 – MANIPULATING BENEFIT VIA MONETARY BONUS
Study 4 tested the benefit component of the cost-benefit framework by varying monetary rewards for correct maze completion. Higher bonuses reduced overreliance, and highlight explanations again reduced overreliance in hard tasks.
- Study design: Study 4 manipulated correct-completion bonuses and AI information in a hard maze task.Participants received either prediction-only or highlight explanations and experienced both bonus levels in a blocked within-subjects design.
- Study design: The low bonus paid $0.01 per correctly completed maze, while the high bonus paid $0.50.The bonus was awarded for correct answers, making the benefit of completing the task properly explicit.
- Hypotheses: The study predicted lower overreliance with high bonuses and with highlight explanations than with prediction-only information.The explanation comparison retested the earlier hard-task finding.
- Results: Higher monetary benefits reduced overreliance, supporting the benefit component of the cost-benefit framework.The authors explain this as increasing the utility of engaging with the task relative to relying on an AI likely to produce mistakes.
- Results: In hard tasks, highlight explanations reduced overreliance compared with receiving only the AI’s prediction.This replicated the Study 1 finding despite the bonus manipulation.
- Summary: Overreliance responded to both the costs of strategies and the benefits received for expending those costs.The authors note that bonus structures may affect findings in crowd-sourced studies.
9 STUDY 5 – MEASURING THE UTILITY OF XAI METHODS
Study 5 adapted Cognitive Effort Discounting to measure the subjective utility of AI explanations across task and explanation difficulties. Participants assigned higher utility to explanations in harder tasks and when explanations were easier to understand.
- Aim: Study 5 measured how task difficulty and explanation difficulty affect the subjective utility participants assign to AI assistance.It tested whether explanations have greater utility in harder tasks and when they are easier to understand.
- Method: COG-ED inferred utility from the reward difference at which participants were approximately indifferent between the high- and low-effort options.The final converged reward difference quantified how much reward participants would forgo for the easier AI-assisted task.
- Method: The adapted COG-ED task asked participants to choose between completing a task alone for a fixed reward or with AI assistance for a shifting reward.The high-effort prediction-only option paid a fixed 100 credits, while the low-effort explanation option initially paid 50 credits.
- Method: Participants compared highlight explanations across easy and medium tasks and compared highlight with written explanations in the medium task.These within-subject comparisons operationalized the task- and explanation-difficulty predictions.
- Results: AI explanations had higher utility in harder tasks because they increased correctness while reducing the time and cognitive effort needed to complete the task properly.The authors contrast this with easier tasks, where using the AI added little help relative to completing the task alone.
- Results: AI explanations had higher utility when they were easier to understand, consistent with lower cognitive-effort and time costs.The study therefore supports the framework’s prediction that people strategically choose whether to engage with AI assistance.
10 DISCUSSION
The paper argues that overreliance is strategically shaped by task costs, explanation costs, and benefits, identifying conditions where explanations reduce it. The discussion also bounds these findings to controlled maze tasks, selected explanation formats, and short-term crowd-sourced users.
- Implications: Explanations reduce overreliance when they lower the cognitive effort required to verify AI predictions, especially in difficult tasks.The framework links explanation use to the relative cost of verification versus relying on the AI.
- Implications: Task difficulty, explanation difficulty, and monetary reward act as knobs that modulate overreliance.These factors proxy broader costs and benefits in human-AI interaction.
- Implications: Explanations that are difficult to understand provide no assistance over prediction-only baselines.The authors recommend explanations that are easy to map to predictions, while noting that human-interpretable explanations remain difficult to generate.
- Limitations and Future Work: The findings may not generalize beyond the controlled, low-stakes maze-solving environment.The authors specifically identify high-stakes, generative, non-game-playing, uncertain, and incompletely explained prediction tasks as future settings.
- Limitations and Future Work: The study used primarily perfect procedural explanations, whereas practical systems may provide imperfect or format-limited explanations that do not establish prediction veracity.The authors hypothesize effects of lower explanation quality but do not empirically validate them.
- Limitations and Future Work: Crowd-sourced users and limited long-term engagement may not reflect practitioners’ needs, values, or changing reliance in high-stakes contexts.The paper calls for co-design with relevant stakeholders and evaluation in real contexts.
- Limitations and Future Work: The paper does not address how explanations affect trust or whether explanation-induced trust reflects understanding versus blind reliance.Prior findings on explanations and trust are conflicting.
11 CONCLUSION
The conclusion presents a cost-benefit framework explaining when AI explanations reduce overreliance. Across five studies, changing task costs, explanation understandability, and rewards altered reliance on AI predictions.
- 11 CONCLUSION: Prior work found that explanations did not empirically reduce overreliance, motivating a cost-benefit framework for identifying when they do.The paper validates this framework experimentally.
- 11 CONCLUSION: Increasing the cost of doing the task alone or using predictions only led people to use explanations and reduced overreliance versus prediction only.The cost was manipulated through task difficulty.
- 11 CONCLUSION: Decreasing explanation understandability increased the cost of engaging with the task and increased overreliance.This supports the framework’s prediction that people use explanations less when verification is more effortful.
- 11 CONCLUSION: Increasing the benefit of completing the task correctly through a monetary bonus decreased overreliance.The result validates the benefit component of the framework.
- 11 CONCLUSION: The Cognitive Effort Discounting studies found that participants assigned higher utility to harder-task and more understandable-explanation conditions associated with larger overreliance reductions.Participants sometimes forewent monetary rewards for those AI conditions.
- 11 CONCLUSION: The experiments suggest that overreliance is at least partly a strategic choice rather than an immutable cognitive inevitability.This conclusion follows from participants’ responsiveness to costs and benefits.
12 APPENDIX
The appendix documents survey measures, analysis procedures, task materials, and supplementary Wikipedia question-answering studies. It also reports that those preregistered studies ended early after showing small effects and likely failed to create sufficient task difficulty.
- Appendix Measures: Participants completed measures of AI competence, trust, confidence, dependability, reliance, task experience, AI use, perceived accuracy, and task preference.The appendix also included a Need for Cognition survey.
- Appendix Analysis: The appendix used Bayesian models with random intercepts for participants and mazes, followed by post-hoc estimated marginal means tests.Fixed effects included AI condition, task difficulty, and explanation condition.
- Supplementary Studies: The supplementary studies replaced the maze task with Wikipedia question answering, using paragraph length as the task-difficulty manipulation and inline highlights as explanations.The easy condition used one paragraph and the hard condition used five paragraphs.
- Supplementary Studies: The preregistered Wikipedia studies ended early because their hypotheses did not show a large effect, so the authors do not report those values.The final samples were substantially smaller than preregistered.
- Supplementary Studies: At least 7/17 and 17/57 participants in the two hard prediction conditions reported skimming, scanning, or searching the longer passages.The authors interpret this as evidence that the hard condition was too easy to do.
- Supplementary Studies: The authors revised their conceptual framing from explanations increasing easy-task cost to explanations not decreasing easy-task cost.This revision motivated switching from frequentist to Bayesian statistics to represent no-difference hypotheses.