Source-linked AI summary

"I'm Not Sure, But...": Examining the Impact of Large Language Models' Uncertainty Expression on User Reliance and Trust

Sunnie S. Y. Kim, Q. Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, Jennifer Wortman Vaughan

arXiv:2405.00623v2cs.HCcs.AI

TL;DR

LLMs can produce convincing errors, creating a need to understand how uncertainty language affects user reliance and trust. The authors test this in a preregistered medical-question experiment and find that first-person uncertainty reduces agreement and confidence while improving accuracy, though overreliance is not eliminated.

  • Problem

    Empirical evidence is limited on how users perceive and act upon natural-language uncertainty expressions in LLM-infused systems, despite risks from overreliance on incorrect outputs.

  • Method

    In a preregistered experiment with N=404, participants answered medical questions with or without responses from a fictional LLM-infused search engine using different uncertainty perspectives.

  • Results

    First-person uncertainty expressions reduced confidence in and agreement with the system’s answers while increasing accuracy; general-perspective effects were weaker and not statistically significant.

  • Takeaways & Limitations

    Natural-language uncertainty may reduce overreliance and overtrust, but language choices require empirical validation before deployment.

  • Takeaways & Limitations

    The yes/no medical-question setup does not establish how uncertainty expression affects more complex everyday tasks, and time and source-use measures were less reliable than in-person measures.

Abstract

from arXiv · show

Widely deployed large language models (LLMs) can produce convincing yet incorrect outputs, potentially misleading users who may rely on them as if they were correct. To reduce such overreliance, there have been calls for LLMs to communicate their uncertainty to end users. However, there has been little empirical work examining how users perceive and act upon LLMs' expressions of uncertainty. We explore this question through a large-scale, pre-registered, human-subject experiment (N=404) in which participants answer medical questions with or without access to responses from a fictional LLM-infused search engine. Using both behavioral and self-reported measures, we examine how different natural language expressions of uncertainty impact participants' reliance, trust, and overall task performance. We find that first-person expressions (e.g., "I'm not sure, but...") decrease participants' confidence in the system and tendency to agree with the system's answers, while increasing participants' accuracy. An exploratory analysis suggests that this increase can be attributed to reduced (but not fully eliminated) overreliance on incorrect answers. While we observe similar effects for uncertainty expressed from a general perspective (e.g., "It's not clear, but..."), these effects are weaker and not statistically significant. Our findings suggest that using natural language expressions of uncertainty may be an effective approach for reducing overreliance on LLMs, but that the precise language used matters. This highlights the importance of user testing before deploying LLMs at scale.

1 INTRODUCTION

LLMs can produce fluent but incorrect outputs, creating risks of overreliance. This study examines whether natural-language uncertainty expressions—and especially their perspective—change users’ reliance, trust, and accuracy.

  • Motivation: LLMs’ plausible yet incorrect outputs can lead users to take actions based on false information.The paper frames overreliance as a significant risk of widely deployed LLMs.
  • Research gap: Little is known about how users respond to natural-language expressions of LLM uncertainty.Prior work has called for uncertainty communication, but the authors identify limited empirical understanding of effective language choices.
  • Research question: The study compares first-person expressions such as “I’m not sure, but...” with general-perspective expressions such as “It’s not clear, but...”.The comparison isolates the perspective used to communicate uncertainty.
  • Study overview: N=404 participants answered medical questions with or without responses from a fictional LLM-infused search engine.The preregistered experiment varied access to system responses and measured accuracy, time, reliance, and trust.
  • Main result: First-person uncertainty reduced confidence in the system and agreement with its answers while increasing participants’ accuracy.These effects were observed relative to responses without uncertainty expression.
  • Implications: Natural-language uncertainty may reduce overreliance, but its effectiveness depends on the language used and should be tested with end users before deployment.The exploratory analysis attributes higher accuracy to reduced, but not eliminated, overreliance on incorrect answers; general-perspective effects were weaker and nonsignificant.

2 RELATED WORK

Prior research has studied uncertainty communication and calibration across AI and human contexts, but evidence about natural-language uncertainty in LLM systems remains limited. The paper develops its measures and design around this gap.

  • Uncertainty communication: Uncertainty can be communicated numerically, visually, or through natural language, with numerical and visual forms offering precision but often being difficult to interpret.Natural-language expressions provide an alternative communication format.
  • Natural-language evidence: Research on natural-language uncertainty has examined conversational breakdowns, confidence communication, and the persuasive effects of different hedges.Hedge effectiveness can vary with wording and implied likelihood.
  • Prior AI evidence: Prior AI studies report that uncertainty communication can increase vigilance and task performance while reducing overreliance.Examples include quantile dot plots and numerical confidence displays.
  • Uncertainty in LLMs: LLMs frequently generate confidence and doubt expressions, but these expressions are poorly calibrated.This motivates studying how users react to the expressions themselves rather than assuming they accurately represent uncertainty.
  • Research gap: Empirical work on how uncertainty expressions affect users of LLM-infused systems remains limited.Existing exceptions address highlighted uncertainty in code completion or search and natural-language expressions in related settings.
  • Study rationale: The study avoids assuming calibration by varying uncertainty independently across correct and incorrect system answers.This permits direct comparisons between responses with and without uncertainty expression.
  • Measures: Reliance is assessed through agreement with the system’s answer, while trust is assessed through confidence, source usage, trust intentions, and trust beliefs.The paper treats agreement as a comparative behavioral indicator rather than a direct measure of reliance or trust.
  • Additional measures: The study also considers anthropomorphism, transparency, correctness, and time on task as relevant system or performance measures.These measures address possible trust pathways and task outcomes.

3 METHODS

The authors preregistered a controlled experiment in which participants answered eight medical yes/no questions under four AI-access conditions. Behavioral, self-reported, and exploratory analyses assessed reliance, trust, performance, and uncertainty effects.

  • Design: The preregistered study used a between-subjects experiment with within-subject comparisons in a fictional LLM-infused search setting.Participants completed medical information-seeking tasks with or without access to system responses.
  • Experimental conditions: Participants were randomly assigned to Control, Uncertain1st, UncertainGeneral, or No-AI conditions.Control omitted uncertainty; the two uncertain conditions varied perspective; No-AI removed system responses.
  • Procedure: Each participant answered eight challenging factual medical questions, with AI answers correct for half the questions and uncertainty expressed for half in the uncertainty conditions.Question and uncertainty order were randomized, while the underlying answer set was fixed.
  • Dependent variables: The measured outcomes included agreement, correctness, time, link clicks, AI and internet use, confidence, trust, anthropomorphism, and transparency.These combined observed behavior with self-reported ratings.
  • Analysis: Confirmatory analyses tested condition effects between groups and uncertainty-expression effects within the two uncertainty conditions.Repeated measures used mixed-effects models with participant and question intercepts; exit measures used ANOVA.
  • Exploratory analyses: Exploratory analyses examined over- and underreliance by AI-answer correctness and thematically analyzed participants’ free-form responses.These analyses complemented the preregistered confirmatory tests.
  • Materials: The questions were yes/no medical items selected for difficulty, search resistance, and objective automatic assessment.Items originated from MedQuAD with minor modifications to increase difficulty.
  • AI responses: AI responses were based on Microsoft Copilot in Bing and received only minor presentation changes without substantive content modifications.Responses were collected in July 2023 and standardized to begin with “Yes” or “No”.

4 RESULTS: CONFIRMATORY ANALYSIS

The confirmatory analyses show that uncertainty expression generally reduced reliance-related responses, with the clearest effects for first-person wording. First-person uncertainty also improved accuracy, while source-use and broader trust effects were more limited or mixed.

  • Agreement with AI: 80.9% of Control participants agreed with the AI system versus 58.4% without AI access, showing greater agreement when responses were provided.
  • Agreement with AI: 74.8% agreement in Uncertain1st was significantly below Control’s 80.9%, whereas UncertainGeneral’s 77.6% difference was not significant.
  • Confidence in Answers: ConfidenceAI fell from 3.95 in Control to 3.66 in Uncertain1st, while UncertainGeneral’s 3.80 was not significantly different from Control.
  • Source Usage: Uncertain responses reduced UseAI and increased UseInternet within uncertainty conditions, although between-condition source-use differences were not significant.Participants also reported using links or conducting independent searches when the system expressed uncertainty.
  • Trust and Perception of AI: Uncertain1st reduced TrustIntention to 2.91 versus 3.25 in Control and 3.36 in UncertainGeneral, while trust beliefs, anthropomorphism, and transparency did not differ significantly.
  • Task Performance: 72.8% correctness in Uncertain1st was significantly higher than Control’s 63.9%, while UncertainGeneral’s 67.9% increase was not significant.AI access overall reduced accuracy relative to No-AI, and the experimental system was correct on only 50.0% of questions.

5 RESULTS: ADDITIONAL ANALYSES

Additional analyses show that uncertainty improves accuracy mainly when the AI system is wrong, with first-person expressions producing stronger benefits than general expressions. Participants interpreted uncertainty in several ways, including as a signal to verify difficult or unreliable answers.

  • 5.1 Effect of Uncertainty Expression on Over- and Underreliance: 88.5% versus 77.9% accuracy when the system was correct, but 33.0% versus 64.7% when it was incorrect, comparing Control with No-AI.Access to the system helped when its answer was correct but harmed accuracy when its answer was incorrect.
  • 5.1 Effect of Uncertainty Expression on Over- and Underreliance: Uncertainty improved accuracy on questions the system answered incorrectly without reducing accuracy when it answered correctly, with larger gains for first-person expressions.This comparison supports reduced overreliance on incorrect answers, although participants remained less accurate than the No-AI group on such questions.
  • 5.1 Effect of Uncertainty Expression on Over- and Underreliance: When uncertainty was expressed, accuracy fell from 92.2% to 89.2% for Uncertain1st and from 94.8% to 83.1% for UncertainGeneral when the system was correct.The corresponding incorrect-system comparisons increased from 43.6% to 52.0% for Uncertain1st and from 32.8% to 48.0% for UncertainGeneral.
  • 5.2 Participants’ Interpretations of AI’s Uncertainty Expression: Most participants in uncertainty conditions attributed the expressions to the system’s inability to answer, and some interpreted them as encouragement to verify answers independently.Other interpretations included programming, impression management, credibility preservation, liability avoidance, and medical-answer restrictions.
  • 5.2 Participants’ Interpretations of AI’s Uncertainty Expression: Participants in UncertainGeneral more often attributed uncertainty to difficult or conflicting information, while Uncertain1st participants more often attributed it to system limitations.The reported proportions were 51.5% versus 41.3% and 20.7% versus 7.4%, respectively.

6 DISCUSSION

The discussion argues that natural-language uncertainty can reduce overreliance and overtrust, especially in first-person form, but does not eliminate overreliance. The authors therefore emphasize empirical validation and caution against generalizing beyond this controlled medical-question setting.

  • 6 DISCUSSION: Natural-language uncertainty prompted more cautious behavior, including longer decision times and greater reported reliance on outside sources, but did not fully eliminate overreliance.Participants without AI access achieved the highest task performance.
  • 6 DISCUSSION: First-person uncertainty had stronger effects than general-perspective uncertainty, heightening the warning effect while potentially amplifying unjustified positive messages.The discussion connects this pattern to greater involvement and engagement from first-person messages.
  • 6 DISCUSSION: The most successful approach for reducing overreliance also decreased trust, creating a tradeoff when users already under-trust an AI system.The authors recommend customized, evidence-based solutions rather than a universal approach.
  • 6 DISCUSSION: The study’s yes/no medical-question design limits conclusions about more complex tasks, personal symptom searches, repeated interaction, and broader everyday behavior.The authors also caution that time and source-usage measures would be more reliable in an in-person laboratory study.
  • 6 DISCUSSION: Because the findings may not generalize broadly, teams should evaluate uncertainty language with end users before releasing LLM systems.The paper frames language choices as consequential for how people perceive and act on LLM outputs.

7 ETHICAL CONSIDERATIONS AND POSITIONALITY

The paper addresses ethical responsibilities in conducting and applying research on uncertainty expression in LLM systems. It emphasizes participant protections, risks of misuse, and the researchers’ positionality.

  • Mitigating harms to human subjects: Participants were recruited from U.S.-based MTurk workers, with an intended wage of $15 and an average estimated wage of $14.80 per hour.The estimate may understate wages because time spent on other activities was unknown.
  • Mitigating harms to human subjects: All completed participants were paid and approved regardless of data-quality results, and participants were debriefed that the AI’s medical information could be incorrect.
  • Potential negative societal impact: Because findings may not generalize across contexts, deploying teams should conduct extensive user testing rather than decide uncertainty language from these results alone.
  • Potential negative societal impact: The authors warn that uncertainty expressions could be strategically used by bad actors to make misinformation more persuasive.
  • Positionality: The authors identify their U.S. technology-company employment, responsible-AI experience, and access to research funding as influences on the study’s perspective and design.

A PARTICIPANT DEMOGRAPHICS AND BACKGROUND

The final sample included 404 participants whose demographics differed from the U.S. population, particularly in age, education, and racial representation. Participants reported moderate LLM familiarity and use, with somewhat positive attitudes.

  • Participant demographics: Of 404 participants, 51.7% identified as women, 46.8% as men, and 0.5% as non-binary.
  • Participant demographics: Compared with the U.S. population, the sample was younger and more educated, with white respondents over-represented and Black and Hispanic/Latino respondents under-represented.
  • LLM background: LLMAttidue averaged 3.8 ± 1.0, between neutral and somewhat positive attitudes toward LLMs.

B DATA COLLECTION AND EXCLUSION

The study preregistered recruitment, sampling, and exclusion procedures, collected 656 complete responses, and retained a final sample of 404 after quality-control exclusions.

  • Data collection: The target sample was set at N=432 using an a priori power analysis for four conditions, allowing 20% additional participants for possible exclusions.The calculation required 360 participants for medium-sized effects at α = 0.05 and power 0.90.
  • Data collection: Researchers planned to recruit U.S.-based MTurk participants meeting Masters, approval-rating, and completed-task criteria, with preregistered adjustments if recruitment was insufficient.
  • Data collection: Data collection occurred over two weeks in September 2023, beginning with the Masters requirement and later removing it when recruitment targets were not reached.
  • Data exclusion: Of 656 complete responses, 252 (38.4%) were excluded for honeypot failures, identical task answers, very short completion times, low attention accuracy, or problematic free-form responses.
  • Data exclusion: Manual review of free-form answers identified off-topic and identical responses as an additional data-quality issue, while overlapping exclusion criteria meant some responses were flagged multiple times.

C.1 Exploration of LinkClick and UseLink

The appendix reports model-fitting difficulties for behavioral link-use measures and evaluates the reliability of multi-item trust and AI-perception scales.

  • LinkClick and UseLink: The preregistered within-condition model failed to fit properly for LinkClick and UseLink in the UncertainGeneral condition because of large individual variance.Among 94 participants, 50 never clicked a link, 17 clicked links on all eight tasks, and 27 showed intermediate behavior.
  • Trust and perception measures: TrustBelief, TrustIntention, Anthropomorphism, and Transparency were calculated as indexes from participants’ ratings on multi-item scales.The appendix assesses their internal consistency using Cronbach’s alpha, a 0-to-1 reliability measure.

D FULL WORDING USED IN THE EXPERIMENT

Participants are introduced to AI system A, an internet-connected LLM prototype whose fluent responses and sources may nevertheless be inaccurate, incomplete, or inconsistent.

  • Participants may use additional resources, including the internet, books, friends, and family, while completing the tasks.
  • The study includes eight information-seeking tasks followed by an exit questionnaire about participants’ experiences, perceptions, and backgrounds.The study is designed to take around 20 minutes, with the questionnaire taking 5–7 minutes.
  • Participants are told that LLM-generated responses can sound convincing and fluent but may not always be correct.
  • The study introduces AI system A as an internet-connected prototype based on large language model technology.It can answer a wide range of questions and sometimes provides sources.

Task example

The task example page is shown to participants in the Control, Uncertain1st, and UncertainGeneral conditions, but not to participants in the No-AI condition.

  • Participants in the Control, Uncertain1st, and UncertainGeneral conditions see the task example page.
  • Participants in the No-AI condition see only the task question and a slightly different set of survey questions.

Task comprehension questions

Before the information-seeking tasks, participants answer comprehension questions about the study and AI system A, review the correct answers, and then proceed to eight tasks.

  • Participants answer TRUE-or-FALSE questions recalling details about the study and AI system A.
  • The comprehension questions cover AI system A’s internet connection, use of clickable sources, and similarity to OpenAI’s ChatGPT.
  • Participants review the correct answers before proceeding to the information-seeking tasks.
  • Participants are instructed to complete eight information-seeking tasks in one sitting.

Task (repeated 8 times)

Participants answer eight challenging medical yes-or-no questions, with task materials and AI responses varying by condition, and report confidence, resources, trust-related perceptions, and background characteristics.

  • Task procedure: All participants answer the same eight medical questions, while conditions vary the AI response or provide no AI response.Questions are randomly selected from a larger list and shown in random order.
  • Task procedure: After viewing the task materials, participants provide a final answer and rate confidence in both the AI system’s answer and their own answer.
  • Task procedure: Participants report whether their final answers were based on AI system A, linked sources, their own knowledge, internet searches, or other resources.
  • Post-task measures: Participants also report how they used AI system A, why they consulted other resources, and why they diverged from the system’s answer.
  • Post-task measures: The questionnaire measures trust beliefs, trust intentions, anthropomorphism, transparency, and participants’ awareness of uncertainty expressions.Trust beliefs use six statements, trust intentions use four, anthropomorphism uses four items, and transparency uses two statements.
  • Background measures: The study collects demographic information alongside measures of LLM familiarity, usage frequency, and general attitude toward LLM applications.
Loading 2405.00623v2…