Source-linked AI summary
Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations
Max Schemmer, Niklas Kühl, Carina Benz, Andrea Bartos, Gerhard Satzger
TL;DR
The paper addresses ambiguity in defining, measuring, and explaining reliance on AI advice by introducing Appropriateness of Reliance (AoR) as a two-dimensional measurement concept. It analyzes explanations for AI advice and reports effects on reliance-related outcomes, while noting that generalizability is limited by the explanations studied.
Problem
Current research on reliance on AI advice remains ambiguous regarding its definition, measurement, and impact factors.
Method
The paper introduces the two-dimensional Appropriateness of Reliance (AoR) metric and analyzes explanations for AI advice, including their mediators.
Results
Explanations affect R_AIR, while trust has no significant effect on R_AIR or R_SR in the study.
Takeaways & Limitations
AoR provides a two-dimensional concept for describing and measuring reliance on AI advice.
Takeaways & Limitations
The generalizability of the experimental findings is limited by the choice of explanations.
Abstract
from arXiv · showhide
AI advice is becoming increasingly popular, e.g., in investment and medical treatment decisions. As this advice is typically imperfect, decision-makers have to exert discretion as to whether actually follow that advice: they have to "appropriately" rely on correct and turn down incorrect advice. However, current research on appropriate reliance still lacks a common definition as well as an operational measurement concept. Additionally, no in-depth behavioral experiments have been conducted that help understand the factors influencing this behavior. In this paper, we propose Appropriateness of Reliance (AoR) as an underlying, quantifiable two-dimensional measurement concept. We develop a research model that analyzes the effect of providing explanations for AI advice. In an experiment with 200 participants, we demonstrate how these explanations influence the AoR, and, thus, the effectiveness of AI advice. Our work contributes fundamental concepts for the analysis of reliance behavior and the purposeful design of AI advisors.
1 INTRODUCTION
AI advice is increasingly used in complex decisions where errors make unquestioning acceptance risky. The paper addresses ambiguity in appropriate reliance by introducing AoR and experimentally examining how explanations affect reliance behavior.
- Motivation: Complex AI applications increase both the number and severity of potential errors, making unconditional acceptance inappropriate.The introduction illustrates this risk with medical diagnosis, where incorrect AI advice could cause a physician to miss cancer.
- Motivation: Human-AI complementarity requires decision-makers to distinguish when to follow AI advice from when to rely on their own judgment.Simply accepting even superior AI advice can fail to exploit complementary human capabilities and achieve complementary team performance.
- Research gap: Current research on appropriate reliance remains ambiguous in its definition, measurement, and impact factors.The literature uses appropriate reliance both as a binary target state and as a degree of appropriateness, without a unified measurement concept.
- Study design: The study uses a behavioral experiment with 200 participants to examine how explanations influence AoR and related reliance processes.The analysis considers reliance and confidence as mediators and uses AoR to examine factors associated with changes in overall performance.
- Contribution: The paper provides a theoretical foundation, measurement concept, and design guidance for analyzing and developing AI advisors.The authors position the experimental findings as a starting point for further evaluations of factors affecting AoR.
2 RELATED WORK
Prior work treats appropriate reliance inconsistently, links it to team performance or case-by-case behavior, and reports mixed effects of explanations. The paper identifies the need for a precise definition and unified measurement concept.
- Appropriate reliance in automation: Automation and robotics research commonly defines appropriate reliance through its contribution to human-automation team performance.Some accounts define it as the reliance pattern most likely to produce the best team performance.
- Appropriate reliance in human-AI decision-making: Human-AI studies examine prediction learning, acceptance of correct and incorrect advice, and detecting when models make mistakes.These approaches connect improved prediction of AI behavior with more appropriate reliance but do not consistently formalize the concept.
- Related concepts: The literature also uses related ideas such as appropriate trust and often evaluates reliance on a case-by-case basis.Some measures assess following correct recommendations while rejecting incorrect ones, whereas other studies examine over- and self-reliance qualitatively.
- Research gap: Overall, prior research lacks a precise definition, unified measurement concept, and clear account of when and why explanations influence appropriate reliance.The related-work review explicitly identifies all three gaps.
- Explainable AI and appropriate reliance: Explainable AI is studied as a means to improve reliance, but findings differ across explanation types and decision contexts.Explanations may support understanding and correct reliance, yet some can increase reliance on wrong AI recommendations.
3 CONCEPTUALIZATION OF APPROPRIATE RELIANCE
The paper defines reliance as observable behavior and introduces Appropriateness of Reliance (AoR) as a two-dimensional measure distinguishing appropriate AI and self-reliance across sequential decisions.
- Measurement gap: Existing appropriate-reliance measures do not fully capture the initial human decision and therefore lose information about human discrimination ability.The sequential paradigm records an initial decision, AI advice, and a potentially revised decision.
- Reliance and Appropriateness: Reliance is defined as the decision-maker’s observable behavior following an advisor’s advice, rather than a feeling or attitude.
- Error setting: The framework focuses on systematic-error settings, where humans may evaluate AI advice case by case, rather than relying uniformly on average AI performance.Random errors cannot be distinguished, whereas systematic errors may be identifiable and support complementary human-AI performance.
- Reliance outcomes: The framework distinguishes four outcomes: correct AI reliance, incorrect self-reliance, correct self-reliance, and incorrect AI reliance.These outcomes separate under-reliance from over-reliance while excluding confirmation cases where the AI and human initially agree.
4 THEORY DEVELOPMENT AND HYPOTHESES
The paper develops hypotheses about how explanations affect AoR through self-reliance, AI reliance, confidence change, and trust. The direction of the explanation effect on self-reliance is left unspecified because explanations may either reveal errors or signal competence.
- Explanations and self-reliance: Explanations may help people detect incorrect advice, but they may also signal competence and increase over-reliance; therefore, the effect on RSR is hypothesized without a specified direction.
- Explanations and AI reliance: Explanations are hypothesized to increase relative AI reliance (RAIR) by helping initially incorrect decision-makers learn and validate patterns supporting correct advice.The paper argues that explanations can inspire knowledge extensions and help users assess whether those extensions make sense.
- Confidence: Explanations are hypothesized to increase the change in human self-confidence after receiving AI advice.The proposed mechanism concerns post-advice confidence rather than initial confidence.
- Confidence pathways: An increased change in self-confidence is hypothesized to increase both relative self-reliance (RSR) and relative AI reliance (RAIR).
- Trust: Explanations are hypothesized to increase trust, while trust is hypothesized to decrease RSR and increase RAIR.The model distinguishes trust’s association with greater AI reliance from the need for trust to remain calibrated to AI capability.
5 EXPERIMENTAL DESIGN
The study tests explanation effects in a between-subject online experiment using deceptive hotel-review classification, an 86%-accuracy SVM advisor, and LIME feature-importance explanations.
- Task and data: The experiment uses deceptive-versus-genuine hotel-review classification with reviews drawn from a dataset containing 400 deceptive and 400 genuine examples.
- AI advisor and explanations: The AI advisor is a Support Vector Machine with 86% accuracy, and the explanation condition uses LIME feature-importance explanations displayed by highlighting influential words.
- Conditions: Participants are randomly assigned between a control condition with AI advice alone and an explanation condition with AI advice plus feature importance.
- Procedure: Each participant completes two training reviews followed by 16 main reviews in a sequential process: initial classification, AI advice, and possible decision revision.The main tasks provide no performance feedback, while training tasks include feedback.
- Scope: The authors caution that recruiting crowd workers may limit generalizability and suggest studying professional deception-detection screening services.
- Participants and measures: The study includes 200 participants, recruited through Prolific, and measures AoR from revised decisions alongside confidence change and trust ratings.One participant was excluded from the feature-importance condition after failing the manipulation check.
6 RESULTS
The behavioral experiment finds that explanations increase appropriate reliance on AI advice without significantly changing selective reliance. Structural equation modeling further links explanations to confidence and appropriate reliance, with confidence partially mediating that effect.
- AoR and AR: Explanations significantly increased participants’ R_AIR compared with the control condition (t = −1.95, p = 0.05).
- AoR and AR: The change in confidence was significant, with participants feeling more confident after receiving AI advice with explanations (t = −2.33, p = 0.02).
- AoR and AR: Explanations did not significantly change R_SR, which remained 71.87% (±3 pp) in control and 69.45% (±3 pp) with feature-importance explanations.
- AoR and AR: The human-AI team did not significantly outperform human accuracy, so the criterion for appropriate reliance was not reached.
- Structural equation modeling: SEM confirmed significant effects of explanations on R_AIR and confidence, partial mediation through confidence, and trust effects on both R_AIR and R_SR.
7 DISCUSSION
The discussion presents AoR as a theoretical and measurement foundation for appropriate reliance, while reporting how explanations affect reliance and confidence. It also identifies important boundaries involving task type, sequential measurement, explanation design, and experimental scope.
- Theoretical foundation of appropriate reliance: The paper defines appropriate reliance and develops Appropriateness of Reliance (AoR) as a two-dimensional measurement concept.AoR accounts for the initial human decision, enabling distinctions between advice effects and confirmation.
- Implications for appropriate reliance and explainable AI: The research model examines how explanations influence AoR and tests the model in a deception-detection task.The study analyzes explanations as a design feature of AI advice and considers reliance-related mechanisms.
- Implications for appropriate reliance and explainable AI: Explanations significantly affect R_AIR, while they do not significantly affect R_SR in this study.The findings therefore do not support a universal claim that explanations reduce over-reliance across all task types.
- Implications for appropriate reliance and explainable AI: Trust significantly affects both R_AIR and R_SR, and the effect of explanations on R_AIR partially depends on changes in confidence.The study also reports that explanations may increase task knowledge, which is offered as a possible reason for their effects.
- Limitations: The study is limited by its deception-detection setting, selected explanation form, classification-task focus, and sequential task setup.The sequential setup may mentally prepare participants, induce anchoring, reduce over-reliance, and contribute to the experiment’s overall low R_AIR; the findings are therefore an approximation of real human behavior.
- Future Work: Future work should compare alternative reliance measures and investigate additional explanation designs, including counterfactual and global explanations.The paper also proposes modeling newly learned knowledge as a mediator and developing techniques to distinguish incorrect AI advice.
8 CONCLUSION
The paper argues that AI use should shift from adoption and acceptance toward ensuring appropriate reliance during use. It provides the AoR definition and measurement concept, plus initial insights into how explanations influence reliance.
- Appropriate reliance is presented as the next milestone after research focused on AI adoption and acceptance.
- The paper argues that AI use requires ways to ensure appropriate reliance and effective use after deployment.
- The authors provide a definition and measurement concept—Appropriateness of Reliance (AoR)—to guide future research.
- The paper generates initial insights into how explanations influence Appropriateness of Reliance.
- The authors hope this research inspires future work and practice that support effective AI use.