Source-linked AI summary

Show me the evidence: Evaluating the role of evidence and natural language explanations in AI-supported fact-checking

Greta Warren, Jingyi Sun, Irina Shklovski, Isabelle Augenstein

arXiv:2601.11387v1cs.HC

TL;DR

AI-assisted fact-checking raises concerns about factuality, bias, and overreliance, while the role of evidence remains under-researched. The study varied explanation type, AI certainty, and correctness in a mixed-methods experiment, finding that participants relied on accessible evidence across conditions. The results position evidence, alongside explanations, as important support for evaluating AI outputs, while calling for further study in more naturalistic settings and with domain experts.

  • Problem

    The study addresses the under-researched role of evidence in helping people evaluate AI outputs during high-stakes information-seeking tasks such as fact-checking.

  • Method

    A mixed-methods controlled experiment varied explanation type, AI certainty, and AI correctness while participants evaluated fact-checking claims and AI predictions with access to underlying evidence.

  • Results

    Evidence was the most important information source across explanation, certainty, and correctness conditions, while explanations were useful indicators of potentially unreliable AI output.

  • Takeaways & Limitations

    Making evidence underlying AI outputs accessible can support critical assessment and increase engagement with sources in AI-assisted information-seeking.

  • Takeaways & Limitations

    Further research should examine evidence and explanation evaluation in more naturalistic task contexts and with domain experts rather than only crowdworkers.

Abstract

from arXiv · show

Although much research has focused on AI explanations to support decisions in complex information-seeking tasks such as fact-checking, the role of evidence is surprisingly under-researched. In our study, we systematically varied explanation type, AI prediction certainty, and correctness of AI system advice for non-expert participants, who evaluated the veracity of claims and AI system predictions. Participants were provided the option of easily inspecting the underlying evidence. We found that participants consistently relied on evidence to validate AI claims across all experimental conditions. When participants were presented with natural language explanations, evidence was used less frequently although they relied on it when these explanations seemed insufficient or flawed. Qualitative data suggests that participants attempted to infer evidence source reliability, despite source identities being deliberately omitted. Our results demonstrate that evidence is a key ingredient in how people evaluate the reliability of information presented by an AI system and, in combination with natural language explanations, offers valuable support for decision-making. Further research is urgently needed to understand how evidence ought to be presented and how people engage with it in practice.

1 Introduction

AI systems can support fact-checking but may produce factual and biased outputs that encourage overreliance. This study examines how people use evidence and explanations when evaluating AI-assisted fact-checking decisions.

  • LLMs are increasingly used for high-stakes information-seeking and decision-support tasks such as fact-checking, despite factuality and bias concerns.
  • Overreliance occurs when people accept or follow incorrect AI recommendations, making critical evaluation important in decision-support tasks.
  • Evidence-seeking is considered important for calibrating reliance on AI advice, but it remains surprisingly under-researched.
  • Research on explanations and reliance is mixed: explanations may support efficient AI use, but some types may increase overreliance.
  • The study used a mixed-methods 3x2x2 controlled experiment to compare verdict-focused, uncertainty, and no-explanation conditions across AI certainty and correctness.
  • Participants overwhelmingly identified evidence as the most useful information source across AI correctness, certainty, and explanation conditions.

2 Related Work

Prior work addresses overreliance through evidence access and explanations, but their effects on AI-supported decisions remain unsettled. This study frames evidence and explanations as complementary supports for critical engagement with AI outputs.

  • Overreliance mitigation approaches encourage users either to evaluate evidence independently or to understand the AI system and its reliability.
  • Providing evidence alongside AI predictions is proposed as a way to reduce overreliance and promote users’ own conclusions.
  • Earlier question-answering research linked source-listed LLM responses with lower overreliance and greater confidence in users’ own answers.
  • Explanations can improve credibility assessment, but LLM-generated explanations may also increase overreliance.
  • The study presents claims, evidence, AI verdicts, uncertainty, and explanations, then asks participants whether to use the AI prediction and which information informed their decision.

3 Method

The experiment presented fact-checking claims with AI predictions, certainty estimates, explanations, and two evidence documents. It varied explanation type, AI correctness, and AI certainty while measuring participants’ decisions and information use.

  • The controlled experiment provided each fact-checking item with an AI verdict, numerical certainty estimate, explanation, and two underlying evidence documents.
  • Evidence documents and claims were drawn from the DRUID fact-checking dataset, which includes claims from professional fact-checking websites and retrieved online evidence.
  • Uncertainty explanations identified conflicting and concordant evidence spans influencing model certainty, while verdict explanations referenced or summarized evidence supporting the prediction.
  • Pre-testing led researchers to revise claims, instructions, and task procedures after observing participants’ evidence-focused behavior and source-identity concerns.
  • The design crossed three explanation types with correct or incorrect AI advice and high or low AI certainty.
  • Participants reported whether their decisions relied on the AI verdict, certainty, explanation, evidence, own knowledge, or other information.
  • A five-minute limit per claim caused evidence and explanations to disappear when time elapsed, forcing participants to decide.
  • The study recruited 208 Prolific participants randomly assigned to uncertainty, verdict-based, or no-explanation conditions.

4 Results

Participants relied heavily on evidence when evaluating AI-supported fact-checking claims, while natural-language explanations were used as complementary aids whose usefulness depended on correctness, alignment, and clarity.

  • Evidence use: 64% of participants opened both evidence documents for every claim, while only three of 208 participants accessed neither document.Evidence was reported as useful significantly more often than any other available information, regardless of condition.
  • AI advice: Participants followed AI advice more when it was correct than incorrect and when certainty was high than low.The correctness effect was F(1, 207)=208.779, p<.001, η2p=.50; the certainty effect was F(1, 207)=55.55, p<.001, η2p=.21.
  • Explanation use: Natural-language explanations were used more often than no explanation, but did not change reliance on AI advice or interact with advice correctness or certainty.Explanation use differed significantly by condition, F(2, 205)=24.668, p<.001, η2p=.19, whereas reliance showed no explanation-condition effect, F(2, 205)=1.825, p=.164.
  • Explanation use: Explanations were judged more useful when advice was correct, while explanations and evidence were often used together when explanations were natural-language uncertainty descriptions.Verdict explanations were rated more useful than no explanation, and explanations sometimes helped participants identify errors when they conflicted with evidence.
  • Evidence use: Participants reported using evidence most frequently, at 67.8% overall, regardless of explanation condition, advice correctness, or AI certainty.Information sources differed significantly in reported use, F(5, 1035)=200.6, p<.001, η2p=.49.
  • Evidence use: Participants in explanation conditions opened all evidence documents less often than those without explanations, suggesting explanations could sometimes substitute for inspecting both documents.All-document opening was 83% with no explanation, compared with 55% for Uncertainty and 57% for Verdict explanations.
  • Evidence interpretation: Participants tried to infer evidence credibility from format, language, and statistics when source identities were omitted.Some participants considered numerical data evidence of respectable sourcing, even though source information was unavailable.
  • Trust and reliance: Higher trust predicted greater agreement with AI advice and greater explanation use, whereas evidence use was associated with lower trust.Trust correlated with AI recommendation reliance (r=.425) and explanation use (r=.26); qualitative responses linked evidence checking to distrust.

5 Discussion

The discussion identifies evidence as central to evaluating AI outputs and suggests that evidence-focused explanations can support critical engagement. It also emphasizes that evidence presentation and engagement require further study.

  • Evidence enabled participants to evaluate AI content successfully regardless of explanation type, while natural language explanations were appreciated.
  • Approximately two-thirds of participants opened all available evidence documents, and almost all opened at least one for each claim.The evidence was embedded directly in the task interface, reducing access friction compared with external links.
  • Participants reported using evidence regardless of whether explanations were provided and noticed inconsistencies between explanations, verdicts, and evidence.Such inconsistencies prompted more critical consideration of AI output.
  • Evidence-focused explanations did not produce greater overreliance on incorrect advice than no explanations in this study.The explanations extracted relevant evidence and explained how it supported the predicted verdict or certainty.
  • Further research should examine evidence and explanation evaluation in more naturalistic contexts and with domain experts.
  • The empirical role of evidence in AI-assisted decision-making remains relatively under-researched.

6 Conclusion

The conclusion reports that evidence was the most important consideration in participants’ decisions about LLM-based fact-checking outputs. Explanations were useful indicators of possible unreliability.

  • Evidence underpinning the AI output was the most important consideration, regardless of whether explanatory information was provided.
  • Participants found explanations useful for identifying where the AI system might be less reliable.

C Subjective Evaluation Scales

The subjective evaluation materials measured perceived explanation helpfulness, AI trust and competence, behavioral intentions, and prior familiarity with claims.

  • Subjective Evaluation Scales: Participants rated how helpful AI explanations were for determining system reliability on a five-point scale.
  • Trust Belief and Intention Scales: The trust scale assessed competence, effectiveness, capability, honesty, benevolence, reliability, and willingness to depend on the AI system.
  • Claim Familiarity: Participants reported whether they had no, limited, some, or full prior knowledge of each claim.
  • Subjective Evaluation Scales: The study collected participants’ confidence and trust judgments and reported these variables for transparency.

D.1 Claim familiarity

Participants were unfamiliar with most claims, and claim familiarity had limited effects on how they used the available information or judged explanations.

  • Claim familiarity: 58% of claim encounters involved no prior familiarity, compared with 14% some knowledge, 24.94% limited knowledge, and 2.5% full knowledge.
  • Claim familiarity: Participants relied more on their own knowledge when they had some claim knowledge than when they had none, t(384)=6.87, p<.001.
  • Claim familiarity: Claim familiarity otherwise did not affect use of available information or judgments of explanations and AI agreement.
  • Confidence: Explanation condition did not affect decision confidence, F(2, 205)=0.428, p=.652.
  • Confidence: Participants were more confident when AI advice was correct, F(1, 207)=6.122, p=.014, and when certainty was high, F(1, 207)=27.886, p<.001.
  • Trust: Explanation groups did not differ in trust-scale judgments, F(2, 205)=.091, p=0.40.

E Uncertainty Estimation and Explanation Generation Method

The study generated verdict and uncertainty explanations for an AI fact-checking system and presented participants with claims, predictions, explanations, and optional evidence. Uncertainty was derived from the entropy of the model’s softmax distribution, while explanations were generated using prompted language-model outputs.

  • Uncertainty estimation: The model’s uncertainty for a verdict label was computed as the predictive entropy of its softmax distribution over candidate labels.The distribution was obtained from logits for True, False, and NEI labels.
  • Explanation generation: Verdict explanations were generated by prompting the model to produce the verdict together with reasons supporting that prediction.Uncertainty explanations additionally referred to key evidence-span interactions to convey prediction certainty.
  • Study procedure: The experiment presented participants with claims, AI verdicts, certainty estimates, explanations, and optional evidence documents before asking them to evaluate reliance on the verdict.Participants rated decision confidence, explanation usefulness, and the information sources used for their decisions.
  • Study procedure: Participants could inspect one or both evidence documents and then identify which information sources supported their final decision.The evidence consisted of retrieved excerpts that could agree or disagree with one another.

G.3 Claim 3 (True, Correct, Low Certainty)

For the low-certainty scone claim, the AI predicted True with 35% certainty, while the evidence presented both a lower average calorie fraction and a report of a large scone exceeding one third of daily calories.

  • AI prediction: 35% certainty accompanied the AI prediction that a scone can equal one third of recommended daily calories.The system judged the claim True despite its low certainty.
  • Evidence: The evidence reported that an average scone provides one fifth of recommended daily calories for females and one sixth for males.It also stated that calorie content is more closely related to scone size than filling.
  • Evidence: A separate report stated that the highest-calorie scone provided over one third, or 38%, of recommended daily calorie intake.The report concerned a large scone without spread or jam.
  • AI explanation: The explanation treated the average-scone evidence as consistent with the claim because both described a substantial contribution to daily calorie consumption.It acknowledged that the specific fractions differed.

G.4 Claim 4 (False, Correct, Low Certainty)

For the low-certainty windmill claim, the AI predicted False with 31% certainty, and the evidence indicated that the claim misrepresented a quote and omitted conditions about energy payback.

  • AI prediction: 31% certainty accompanied the AI prediction that the windmill claim was False.The claim asserted that a windmill could never generate as much energy as invested in building it.
  • Evidence: The evidence stated that a professor did not say windmills would never generate the energy invested in building them.It characterized the social-media meme as omitting substantial context from the original statement.
  • Evidence: The evidence explicitly said that production energy was not greater than the electricity generated over a turbine’s working lifetime.
  • AI explanation: Evidence alignment with the claim’s wording about energy expenditure reduced the AI system’s certainty in its False verdict.The explanation distinguished this partial wording agreement from the evidence’s direct contradiction of the claim’s broader assertion.
  • AI explanation: The explanation highlighted that a well-sited windmill could achieve energy payback in three years or less, whereas poorly placed windmills might not.This condition contradicted the claim’s universal wording.

G.5 Claim 5 (True, Incorrect, High Certainty)

For the high-certainty tea-bag claim, the AI predicted True with 73% certainty, although the evidence qualified the claim by indicating that not all tea bags contain plastic.

  • AI prediction: 73% certainty accompanied the AI prediction that all tea bags contain harmful microplastics.The system judged the claim True.
  • Evidence: The evidence stated that the vast majority of brands use mesh tea bags composed partly of plastic.It also described potential microplastic release when plastic tea bags are exposed to heat.
  • AI explanation: The evidence explicitly qualified the blanket claim by stating that not all tea bags contain harmful microplastics.One cited example described tea bags made without plastic.
  • AI explanation: The explanation treated widespread plastic use and reported microplastic release as support for the AI’s True verdict.It characterized exposure as nearly universal for consumers using standard tea bags.

G.6 Claim 6 (False, Incorrect, High Certainty)

For Claim 6, the AI system judged the statement false with 78% certainty. The evidence was mixed: some material supported the general woodedness claim, but it did not directly verify Lagan Valley or ancient woodland.

  • Evidence describing Lagan Valley as an Area of Outstanding Natural Beauty did not directly address its woodland level.
  • Evidence citing low woodland percentages supported the general claim but used council areas rather than Lagan Valley itself.
  • The evidence provided no specific data about ancient woodland in Lagan Valley, preventing verification of that part of the claim.
  • 78% certainty accompanied the AI system’s false verdict on the claim.

G.7 Claim 7 (True, Incorrect, Low Certainty)

For Claim 7, the AI system judged the statement true but assigned only 29% certainty. The evidence described Ronaldo’s nonparticipation in a rainbow armband and other captains’ armband choices, but also contained a tournament mismatch.

  • The evidence directly described Ronaldo refusing the One Love armband, but it referred to the 2022 World Cup rather than Euro 2020.
  • Evidence stated that Germany, England, and Denmark’s captains wore rainbow or pro-LGBTQ+ armbands, while Ronaldo was among captains wearing normal armbands.
  • The evidence’s tournament mismatch and references to many captains wearing normal armbands reduced certainty that Ronaldo alone lacked a rainbow armband.
  • 29% certainty accompanied the AI system’s true verdict on the claim.
  • For a separate economy claim, evidence included both polling that supported the statement and contrasting concerns about abortion rights and broader representativeness.
Loading 2601.11387v1…