Source-linked AI summary
Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight
Xinyu Fu, Narayan Ramasubbu, Dennis Galletta
TL;DR
LLM errors can pass review even when users are capable and motivated, because oversight-relevant information may not be accessible at the moment of review. Across two randomized lab-in-the-field experiments with 640 employees, the paper tests retrieval-based mechanisms and finds that self-explanations improve detection while personalized cues slow its decline under repeated use.
Problem
LLM oversight research emphasizes users’ capability and engagement but leaves open whether relevant information becomes accessible during review.
Method
Two randomized lab-in-the-field experiments test generative encoding through self-explanation and cue-supported reactivation during repeated LLM use.
Results
Self-generated explanations improve subsequent error detection, while daily cues tied to earlier reasoning make detection decline significantly more slowly.
Takeaways & Limitations
Information retrievability is a distinct precondition for effective oversight, supported by active encoding and cues that reactivate prior reasoning.
Takeaways & Limitations
The paper focuses on user-side conditions and does not cover artifact and organizational conditions such as diagnostic evidence or authority to intervene.
Abstract
from arXiv · showhide
Large language models (LLMs) are increasingly embedded in organizational work, yet their errors often pass human review. Prior research locates such failures in users' capability to review LLM output or their engagement in doing so. We develop an alternative, retrieval-based account of human oversight and posit that error detection is more effective when oversight-relevant information is accessible to users at the moment of review. Across two randomized lab-in-the-field experiments with 640 customer-facing employees, we show that self-generated explanations improve error detection and strengthen recall of verification-relevant reasoning, while cues that reactivate such reasoning help sustain detection under repeated LLM use. Theoretically, we identify information retrievability as a distinct precondition for effective oversight and specify generative encoding and cue-supported reactivation as mechanisms that build and sustain it. Practically, lightweight onboarding self-explanations and daily retrieval cues can make human oversight more resilient as LLM use becomes routine.
Introduction
The paper argues that LLM oversight can fail even when users possess relevant capabilities and motivation, because oversight-relevant information may not be accessible during review. It develops retrievability as a complementary precondition and tests generative encoding and cue-supported reactivation as mechanisms for improving oversight.
- Motivation: Users may miss LLM errors despite capability and motivation when review contexts fail to cue oversight-relevant information.Relevant information includes awareness of system fallibility, recurring error patterns, and verification strategies.
- Contribution: Retrievability is proposed as a distinct and complementary precondition for effective LLM oversight.The account distinguishes possessing information from accessing it at the moment of review.
- Mechanisms: Generative encoding strengthens oversight by producing more elaborated and personally organized verification-relevant reasoning than passive exposure.Self-explanation requires users to work through why a specific output was wrong.
- Evidence: Study 1 found approximately a 10-percentage-point improvement in subsequent error detection after self-explanation.The study included 400 participants and also found more elaborated recall.
- Evidence: Study 2 found that error detection declined significantly more slowly when daily messages carried participants’ verification-related cues.The cue condition was compared with equally personalized verification-neutral phrases across repeated workdays.
- Implications: Cue-supported reactivation can sustain oversight without repeating the original explanation or supplying new task-specific corrective information.Cue-evocation and cue–episode association measures were consistent with reactivation of earlier reasoning.
Literature Review
Prior accounts explain oversight failures through users’ capability or engagement, but the paper identifies a further gap: relevant information may be possessed and actively reviewed without becoming accessible. It frames retrievability as a designable user-side condition while noting that artifact and organizational conditions also matter.
- Oversight: Human oversight involves reviewing LLM output, examining claims and evidence, and deciding whether to accept, revise, or reject it.The study focuses on error detection as a central manifestation of effective oversight.
- Capability: Capability-based accounts ask whether users possess the evaluative resources needed to scrutinize GenAI output.These resources include AI literacy, system mental models, domain expertise, reference evidence, and verification procedures.
- Engagement: Engagement-based accounts ask whether users initiate and sustain the attention and effort required for evaluation.Automation bias, complacency, and declining scrutiny can make routine acceptance attractive.
- Retrievability: Neither capability nor engagement accounts explain whether information required for a particular output becomes accessible during review.The paper identifies this accessibility gap as retrievability.
- Scope: The paper focuses on user-side conditions while acknowledging artifact and organizational factors such as diagnostic evidence and authority to intervene.This is the stated scope boundary of the analysis.
- Retrievability: Retrievability concerns whether recently acquired, task-specific oversight reasoning becomes accessible during subsequent review.The paper develops an account linking encoding and review cues to that accessibility.
Theory and Hypotheses
The theory distinguishes encoding, retrievability, and retrieval, arguing that initial encoding shapes later accessibility while review cues support retrieval. It applies these relationships to self-explanation and cue-supported reactivation, predicting improved detection and sustained oversight.
- Conceptual framework: Initial encoding shapes retrievability, whereas cues available during review support retrieval of oversight-relevant information.The theory develops these as complementary relationships and design levers.
- Conceptual framework: Retrievability is the potential for encoded oversight-relevant information to become accessible later, while retrieval is accessing it during review.Only retrieved information can directly guide verification of the current output.
- Generative encoding: Generative activity can improve retrievability by producing more elaborated, organized, and distinctive representations than passive exposure.Self-explanation requires users to construct and organize information in ways resembling later verification.
- Cue-supported retrieval: Associated cues can reconnect the current review context to a previously encoded representation and support its later retrieval.Cue usefulness depends on reinstating features associated with the original encoding episode.
- Integrated account: Generative encoding and associated cues jointly explain why users with the same information and review activity may differ in accessible reasoning.The two pathways respectively improve retrievability and support retrieval.
- Hypothesis 1: Hypothesis 1 predicts that users who generate explanations of LLM errors will detect more subsequent errors than users who only receive standardized explanations.The prediction treats self-explanation as the generative-encoding pathway.
Cue-Supported Reactivation Under Repeated Use
Under repeated LLM use, cues linked to users’ prior verification reasoning are intended to reactivate that reasoning during review and slow declines in error detection. The model distinguishes association with prior reasoning from generic reminder salience.
- Cue-supported reactivation: A cue associated with a user’s earlier explanation can reactivate that reasoning, even when the cue itself conveys little information to outsiders.The paper illustrates this association using a private keyword linked to a prior episode.
- Cue-supported reactivation: Repeated use creates recurring retrieval opportunities as time and intervening tasks separate encoding from later review.Encoded reasoning may remain retained yet inaccessible when suitable retrieval cues are absent.
- Cue-supported reactivation: Retrieval support reconnects current review with previously encoded verification-relevant reasoning.The cue’s role is to reinstate prior reasoning at the point of review.
- Cue-supported reactivation: The retrieval account predicts that association, rather than prominence alone, distinguishes prior-reasoning cues from comparably presented personalized messages.Generic vigilance predicts temporary scrutiny from any prominent reminder, not a specific advantage for an associated cue.
- Cue-supported reactivation: Hypothesis 2 predicts that retrieval cues will slow, rather than eliminate, error-detection decline across repeated LLM-assisted work.The prediction targets a longitudinal trajectory relative to a comparably presented personalized message without cue association.
Measurements
The studies measure behavioral error detection alongside false positives, verification activity, and participant-level prerequisites. Prespecified ground truth and blind double coding support consistent outcome assessment, while the net-detection index has a unit mismatch limitation.
- Primary and robustness measures: Behavioral error detection is the primary dependent variable, measured as a 0–1 rate at participant or participant-day level.Study 1 uses participant-level analysis; Study 2 uses participant-day analysis.
- Outcome validation: Researchers prespecified embedded errors, acceptable resolutions, and scoring rules, then used blind independent coders with high agreement.Agreement was Cohen’s κ = .92 in Study 1 and κ = .89 in Study 2.
- Primary and robustness measures: False positives capture alterations to verifiably correct statements, separating meaningful error detection from generalized distrust.Each email contained 10–13 correct statements, and false-positive rates were averaged across email-level rates.
- Primary and robustness measures: Net detection contrasts error detection with false positives as a robustness measure.Because the two components use different coding units, the resulting index is interpreted as a summary rather than a single-scale rate.
- Auxiliary outcomes: Verification activity includes review time, reference consultation, and the extent of changes made before submission.These outcomes assess whether manipulations affected verification behavior beyond the primary detection measure.
- Covariates: Participant-level covariates cover objective AI knowledge, domain capability, and oversight orientation, measured independently of treatment exposure.Verification behaviors recorded during tasks remain post-treatment outcomes rather than covariates.
Study 1: Generative Encoding and Subsequent Error Detection
Study 1 randomized employees to generate their own explanations or receive standardized explanations before reviewing new LLM drafts without explanatory support. It tests whether generative encoding improves later recall and error detection.
- Participants and design: Study 1 analyzed 400 employees, with 200 participants in each explanation condition after seven exclusions.The initial sample contained 407 employees assigned to self-explanation or system-provided explanation.
- Outcome procedure: After training, participants reviewed 10 new verification tasks without access to examples, references, or explanatory content.This unaided phase tests whether encoded reasoning remains available during later review.
- Encoding manipulation: Both conditions received identical errors, standardized explanations, and total training time; only the preceding training activity differed.Self-explanation participants generated verification-relevant reasoning before standardized feedback, while controls completed a placebo writing task.
- Measures: Participants’ error detection used scores of 1, 0.5, or 0 for fully, partly, or unresolved errors, averaged across 10 tasks.False-positive and verification measures were aggregated analogously.
Data Analysis
Study 1 estimates the self-explanation effect using prespecified regression models, robustness checks, and manipulation checks. Self-explanation increased error detection while also improving recall and reducing false positives.
- Analysis plan: Regression models estimate self-explanation effects on error detection, false positives, and net detection using two-tailed tests.Covariate-adjusted models add objective AI knowledge, domain capability, and oversight orientation.
- Error detection: Mean error detection was 0.78 with self-explanation versus 0.68 with system-provided explanation.The difference was about 10.6 percentage points and approximately one additional fully resolved error across 10 tasks.
- Error detection: The adjusted error-detection difference remained b = 0.105, p < .001 after adding the three covariates.Equivalent ANOVA and ANCOVA analyses yielded the same conclusion.
- Recall and recognition: Self-explanation participants recognized all three focal error categories at 28.5% versus 8.0%.Focal recognition exceeded distractor endorsement, indicating more accurate discrimination.
- Recall and recognition: Recall elaboration was 1.84 versus 1.43, and recall predicted detection only in the self-explanation condition.The condition × recall interaction was significant at p = .011.
- Interpretation: The findings are consistent with generative encoding making verification-relevant reasoning more accessible during later review.Behavioral indicators also point to more deliberate and discriminating verification, though they are not uniquely diagnostic of the retrieval account.
- Verification quality: Self-explanation reduced the false-positive rate by three percentage points, from 0.12 to 0.09.The lower false-positive rate indicates greater effort was discriminating rather than indiscriminate alteration.
Study 2: Sustaining Error Detection Under Repeated Use
Study 2 followed customer-facing employees using an internal LLM assistant across seven workdays. It tested whether self-generated retrieval cues could sustain oversight during routine use.
- Participants and design: Participants were randomly assigned individually to retrieval-cue or non-cue conditions after a common training session and baseline assessment.Both groups completed the same self-explanation exercise before condition-specific customization.
- Participants and design: 253 of 254 eligible employees completed Study 2, with 240 participants retained after prespecified attention-check exclusions.The final analytic sample contained 120 participants per condition.
- Embedded task period: Study 2 embedded three simulated test emails per day within employees’ normal workflow, each containing two to four predefined errors.Participants were unaware of the test emails’ timing, frequency, or specific content.
- Retrieval-cue intervention: The retrieval-cue group created a word or short phrase linked to the moment they recognized that an LLM response could be wrong.The cue was intended to reactivate the earlier verification-relevant reasoning rather than provide a generic warning.
- Retrieval-cue intervention: Cues appeared once per workday in the standard login welcome message, making retrieval support a lightweight peripheral workflow element.The matched non-cue condition used equally personalized cosmetic phrases with identical presentation, timing, and LLM behavior.
- Measures: Error detection was measured as the proportion of predefined errors corrected in each embedded email.Additional outcomes included correction of the expert-designated highest-risk error, false positives, and net detection.
Data Analysis
The analysis tracked error detection and verification behavior across an eight-day period, comparing trajectories between randomized cue conditions. Descriptively, routine LLM use coincided with declining oversight, while retrieval cues reduced that decline.
- Data structure: 1,920 participant-day observations came from 240 participants completing three test emails per day across eight study days.The study also analyzed 5,760 email-level observations.
- Model specification: Linear mixed-effects models estimated study-day, retrieval-cue, and day × cue effects with participant random intercepts and prespecified covariates.Covariates were objective AI knowledge, domain capability, and oversight orientation.
- Model specification: The day × cue interaction tests whether error detection declined more slowly under retrieval cues than under non-cues.A positive interaction indicates a flatter decline in the retrieval-cue condition.
- Descriptive trajectories: By Day 7, mean error detection was 0.445 in the non-cue condition versus 0.517 in the retrieval-cue condition.The conditions were comparable at Day 0: non-cue Mean = 0.622 and retrieval cue Mean = 0.640, t(238) = 0.71, p = .48.
- Descriptive trajectories: Error detection declined steadily across successive study days despite varied test emails and stable task structure.Participants in the non-cue condition also spent progressively less time reviewing responses and consulted reference materials less often.
- Descriptive trajectories: The retrieval-cue condition showed better-sustained error detection and verification behavior, although some decline remained.The widening gap was identified against each day’s common task set, while task composition could not be completely separated from the shared temporal trend.
Model-Based Test of H2
Mixed-effects models supported H2: retrieval cues slowed the decline in error detection over repeated LLM use. Process measures also indicated that cues reactivated verification-relevant reasoning rather than producing indiscriminate skepticism.
- Primary H2 test: b = 0.008, p = .013 for the day × retrieval-cue interaction, indicating a significantly slower decline in error detection with cues.Error detection declined by b = −0.025 per day without cues and b = −0.017 per day with cues.
- Primary H2 test: The cue offset about one third of daily error-detection erosion, corresponding to about half an additional embedded error caught on Day 7.The cumulative rate difference was 0.008 × 7 ≈ 0.056, with nine embedded errors per day on average.
- Alternative outcomes: The cue effect was largest for expert-designated most consequential errors, where detection eroded by b = −0.101 per day and the cue effect was b = 0.021, p < .001.This outcome targeted the highest compliance or financial risk error in each email.
- Process evidence: 77.5% of retrieval-cue participants versus 48.3% of non-cue participants produced responses reflecting a verification mindset.Mean evocation ratings were 2.15 versus 1.36, respectively, t(238) = 6.85, p < .001.
- Process evidence: The retrieval-cue condition also reported stronger cue–episode association: Mean = 2.86 versus 2.44, t(238) = 6.98, p < .001.The association measure had Cronbach’s α = .81.
- Mechanism: Across the two studies, self-explanations established retrievability through initial encoding, while retrieval support sustained access during repeated use.Study 2 held generative encoding constant and tested subsequent cue-supported reactivation.
Robustness Checks
Robustness checks reproduced the central trajectory pattern across alternative specifications and outcomes. They also provided little support for message salience, semantics, or generic skepticism as competing explanations.
- Alternative specifications: The positive day × cue interaction remained similar across alternative specifications and outcome operationalizations.It remained significant in random-slope, day-fixed-effects, email-level, error-level, and net-detection models.
- Alternative outcomes: Detection of the expert-designated highest-risk error declined most steeply and showed the largest cue effect.Median-dichotomized error detection yielded the same substantive conclusion.
- Competing explanations: False-positive rates were nearly identical across conditions, providing little support for indiscriminate skepticism as the mechanism.The results instead align with sustained detection of genuine errors.
- Cue content: Removing three noncompliant generic cues left the day × retrieval-cue interaction positive and significant.Ordinary-semantic and idiosyncratic cues showed no detectable slope difference within instruction-compliant cue users.
- Competing explanations: Conditions did not differ in perceived meaningfulness, noticeability, memorability, liking, or attention capture, and these attributes did not predict detection trajectories.This offers little support for a message-salience explanation of the trajectory difference.
Overall Discussion
The discussion presents retrievability as a distinct precondition for effective LLM oversight: users may possess relevant information and review actively yet fail when that reasoning is inaccessible. It explains how generative encoding and cue-supported reactivation can build and sustain access across repeated use.
- Retrievability as a Distinct and Complementary Precondition for Oversight: Oversight can fail when relevant information is inaccessible during review, even if users have received it and are actively checking the output.This adds retrievability to capability- and engagement-based accounts as a distinct oversight condition.
- Retrievability as a Distinct and Complementary Precondition for Oversight: The paper identifies retrievability as a distinct precondition that helps translate users’ capability and engagement into effective oversight.Capability concerns evaluative resources, while engagement concerns undertaking and sustaining review; neither guarantees access to the needed reasoning.
- Generative Encoding and Oversight Support: Self-generated explanations improve later access to verification-relevant reasoning because users process information in ways that prepare it for retrieval.The argument distinguishes active encoding from merely receiving equivalent explanatory information.
- Generative Encoding and Oversight Support: Oversight support depends not only on the information delivered but also on the representation users retain for later review.Content determines available information, whereas encoding shapes what remains accessible when a new output is evaluated.
- Reactivation Across Repeated Use: Cue-supported reactivation can restore previously formed reasoning and sustain error detection during repeated use without retraining or repeating the original explanation.The paper distinguishes reactivation from generic attention increases and blanket skepticism because it aims to preserve discrimination between erroneous and acceptable outputs.
- Reactivation Across Repeated Use: The discussion treats declining detection under routine use partly as a retrieval problem, shifting oversight support from one-time review quality toward maintenance across encounters.This temporal perspective motivates reactivation rather than re-instruction as users repeatedly interact with the same system.
Practical Implications
The paper recommends designing LLM oversight around information retrieval, using self-generated explanations at onboarding and cues during later review. These supports should complement realistic workload expectations and continued monitoring as LLM use becomes routine.
- Practical Implications: As LLM use becomes routine, organizations should sustain oversight by tracking error detection, false positives, and review time rather than relying on one training session.The paper frames oversight support as a continuing workflow-design problem involving onboarding and repeated-use retrieval support.
- Practical Implications: Self-generated explanations can encode verification-relevant reasoning so it remains more retrievable during later review.Onboarding can ask users to explain why an example output is wrong and how it should be verified.
- Practical Implications: User-chosen cues can reactivate earlier verification reasoning during later output review without repeating the original explanation.Organizations can capture a brief cue in the workflow and surface it when subsequent outputs are reviewed.
- Practical Implications: Careful review consumes attention and time, so adding oversight duties without adjusting output targets risks trading caught errors for pace.The paper recommends pairing retrieval support with realistic expectations rather than layering it onto unchanged performance demands.
- Limitations and Future Research: The experiments leave open how retrieval support varies across deeper-expertise tasks, higher-stakes decisions, and settings without reference materials.They also isolate capability, engagement, and retrieval support rather than manipulating all three together, and do not fully separate retrieval from attention or effort.
- Practical Implications: Retrieval-oriented workflow design complements, rather than replaces, investments in user capability and engagement.Retrieval support makes oversight-relevant information more accessible but does not guarantee retrieval or eliminate review costs.
Appendix
The appendix reviews capability- and engagement-based explanations for oversight failure, then identifies information retrievability as a distinct unresolved precondition. It describes self-explanations and retrieval cues as mechanisms for building and sustaining access to oversight-relevant reasoning.
- User-Capability Explanation: Capability-based accounts attribute oversight failure to deficient, inaccurate, or incomplete evaluative resources about the AI system or substantive task.These resources include AI literacy, system-specific mental models, domain expertise, evidence, and verification procedures.
- User-Capability Explanation: Generated content requires reviewers to decompose plausible drafts, identify claims needing verification, locate evidence, and correct errors without discarding valid content.This combines AI-specific understanding with task-specific means to test particular claims.
- Engagement Explanation: Engagement-based accounts explain failures to initiate, appropriately allocate, or sustain the attention and effort needed to scrutinize output and regulate reliance.Relevant problems include automation bias, default acceptance, and declining monitoring when systems usually perform well.
- Information Retrievability: Information retrievability adds the possibility that users possess relevant knowledge and actively review output but cannot access the needed information at the point of review.The needed information may include system-fallibility awareness, recurring error patterns, domain knowledge, or verification strategies.
- Information Retrievability: Self-generated explanations improve later unaided error detection, while self-generated retrieval cues slow detection declines over repeated use without supplying corrective task information.Cue-evocation and cue–episode association measures are consistent with reactivating participants’ earlier reasoning.
Appendix C. Measures and Materials
This appendix documents measurement procedures and robustness analyses for the two studies. The analyses support stable self-explanation and retrieval-cue effects across alternative specifications, outcome codings, day effects, and repeated-measures assumptions, while identifying limits on individual-level mediation claims.
- Measures and Materials: Behavioral outcomes were computed from system logs, with reliabilities reported for Study 1 (N = 400) and Study 2 (N = 240).Outcomes included error detection, false-positive rate, verification time, reference-material consultation, and textual modification.
- Measures and Materials: Encoding-source comprehension was high, with 94.5% and 91.0% of participants correctly identifying how explanations were produced across conditions.The manipulation check contrasted generating explanations first with receiving them directly.
- Robustness Tests: Self-explanation effects remained significant across alternative model specifications and control-variable sets, indicating a robust behavioral difference in Study 1.Equivalent analyses included ANOVA/ANCOVA, covariate adjustments, and logistic regression.
- Robustness Tests: The Study 2 day × retrieval-cue interaction remained positive under random-slope models and day fixed effects, supporting trajectory differences beyond participant dependence or specific study days.The day fixed-effects estimate was b = 0.008, p = .013.
- Robustness Tests: Alternative outcome codings produced the same declining-detection and positive cue-interaction pattern, including a day × cue odds ratio of 1.17 (p < .001).The daily-rate models used 1,920 participant-day observations from 240 participants.
- Robustness Tests: Process measures support the retrieval interpretation, but uneven cue evocation and its end-of-study measurement prevent individual-level mediation or standalone mechanism identification.The authors interpret intervention, evocation, association, and message-attribute evidence as convergent process evidence.