Source-linked AI summary

Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable

Shai Vardi, João Sedoc

arXiv:2609.04127v1cs.AI

TL;DR

Users lack a principled basis for deciding whether to rely on individual LLM recommendations when objective ground truth is unavailable. The paper introduces epistemic warrant and a four-tier reliance certificate for pairwise recommendations, finding that stronger warrant tracks human consensus while contributing information beyond confidence and decision difficulty. The framework offers an implementable way to characterize recommendation-level warrant within its supported scope.

  • Problem

    Existing approaches do not provide a decision-level basis for assessing the strength and scope of support for relying on a specific LLM recommendation when objective outcomes are unavailable.

  • Method

    The paper defines epistemic warrant through preference stability and scope, then operationalizes it as a four-tier reliance certificate for pairwise recommendations.

  • Results

    Warrant is positively associated with human consensus across all seven models, with coefficients ranging from 0.076 to 0.154 (p ≤.003), and adds explained variation beyond confidence for six models.

  • Takeaways & Limitations

    Epistemic warrant provides an implementable characterization of individual recommendation support that is distinct from verbalized confidence and not simply explained by decision difficulty.

  • Takeaways & Limitations

    Validation is constrained because many recommendations lack objective ground-truth labels, and correctness cannot serve as a definitive criterion for warrant.

Abstract

from arXiv · show

Large language models are increasingly used to support organizational decisions, yet users often lack a principled basis for assessing whether to rely on a specific recommendation. Existing approaches typically evaluate broad model properties, such as reliability, uncertainty, or robustness, or focus on user trust, rather than the underlying basis for relying on an individual recommendation. Adapting theoretical foundations from epistemology, we introduce epistemic warrant, a decision-level construct that characterizes the stability of a model's preference and the scope over which that preference holds. We operationalize this construct through a four-tier reliance certificate for pairwise recommendations, distinguishing among unstable, context-dependent, locally supported, and broadly supported recommendations. We validate the construct using contemporary methodologies: known-groups tests successfully recover expert-prespecified warrant orderings, and stronger warrants systematically align with independent consensus from crowd workers. Furthermore, we demonstrate that epistemic warrant provides information distinct from verbalized confidence and is not readily explained by decision difficulty. Ultimately, this framework offers a theoretically grounded, implementable approach for characterizing the warrant of individual LLM recommendations when objective ground truth is unavailable.

1. Introduction

Organizations increasingly use LLM recommendations in consequential settings, but objective outcomes are often unavailable and existing signals do not assess the strength and scope of support for relying on a specific recommendation. The paper introduces epistemic warrant and validates a four-tier reliance certificate as a decision-level operationalization.

  • Motivation: LLM recommendations are increasingly used in high-stakes domains, even though their reliability can be uneven and objective outcomes may be unavailable for assessment.The motivating settings include medical triage, hiring, and LLM-as-a-judge pipelines.
  • Research gap: Existing methods assess confidence, repeatability, transformation consistency, or contextual adaptation, but do not integrate these signals into a decision-level assessment of reliance on one recommendation.Prior work includes calibration, repeated sampling, metamorphic testing, and methods for handling underspecified prompts.
  • Construct: Epistemic warrant characterizes the strength and scope of support for relying on a specific recommendation, rather than a model as a whole or an abstract decision problem.The construct is also distinguished from correctness, confidence, and trust.
  • Operationalization: The reliance certificate maps pairwise recommendations onto four ordered categories based on preference stability and how far that preference extends beyond the stated context.The categories are No Warrant, Conditional Warrant, Basic Warrant, and Strong Warrant.
  • Validation: Validation supports the certificate: known-groups tests recover prespecified warrant orderings, and stronger warrant aligns with independent human consensus across models.Warrant remains associated with consensus after accounting for decision difficulty and adds explained variation beyond verbalized confidence.
  • Contribution: The framework provides a theoretically grounded and implementable way to characterize the warrant of an individual LLM recommendation when objective ground truth is unavailable.The paper presents the certificate as a hierarchical, noncompensatory procedure for assigning warrant certificates.

2. Conceptual Foundations and Related Literature

The paper adapts epistemological theories of warrant to characterize the epistemic support for individual LLM recommendations. It distinguishes this decision-level construct from correctness, confidence, trust, and generic robustness, especially when ground truth is unavailable.

  • Theoretical foundations: Epistemic warrant evaluates the support for relying on a particular recommendation rather than whether the recommendation constitutes knowledge.The paper focuses on the recommendation’s epistemic standing as a basis for reliance.
  • Theoretical foundations: Process-oriented epistemology motivates examining recommendation-generating behavior rather than treating observed correctness as sufficient evidence.The framework draws on distinctions between favorable outcomes by luck and outcomes supported by the process producing them.
  • Theoretical foundations: Warrant increases when a recommendation remains stable under decision-preserving changes and continues across broader relevant decision contexts.This combines process-oriented and modal perspectives on epistemic support.
  • Theoretical foundations: The framework treats instability as weakening support, while contextual reversals can instead reveal conditional scope without undermining support in the focal context.This distinction motivates hierarchical warrant categories that separate failures of stability from boundaries of applicability.
  • Related literature: Prior trust and behavioral-evaluation research offers signals such as repeated agreement, metamorphic consistency, and confidence, but less guidance for assessing a specific recommendation’s grounds for reliance.The paper positions its contribution as interpreting these behavioral tests through a decision-level theory rather than as generic robustness or answer-improvement procedures.
  • Related literature: Verbalized confidence can be overconfident and varies across models, tasks, and elicitation procedures, motivating a distinction between stated certainty and behaviorally supported warrant.The paper explicitly frames warrant as distinct from confidence, trust, correctness, and robustness.

3. Theoretical Development of Epistemic Warrant

The paper defines epistemic warrant for a particular model recommendation on a particular pairwise decision and evaluates its strength through stability and contextual scope. These dimensions are operationalized as a hierarchical, noncompensatory sequence of probes producing four ordered warrant categories.

  • Object and scope of evaluation: Epistemic warrant attaches to a recommendation produced by a particular model for a particular decision problem, not to the model or problem in isolation.The evaluated object is r_m(q), the model’s recommendation for focal prompt q.
  • Object and scope of evaluation: Warrant reflects how broadly and robustly the model supports a preference across relevant variations of the decision problem.Its two related features are preference stability and the scope over which that preference continues to hold.
  • Stability and scope: A recommendation must first exhibit sufficient stability because repeated or decision-preserving reformulations can otherwise reverse the preference.Variation may reflect little basis for strongly preferring either alternative, rather than poor model behavior.
  • Stability and scope: A reversal in a plausible narrower subcontext identifies conditional support without necessarily undermining the recommendation in the original context.Narrower reference classes S′ ⊊ S diagnose where support changes and need not represent the decision maker’s unobserved true context.
  • Stability and scope: A reversal in a broader or substantively related context marks a boundary beyond which the recommendation should not be generalized.Persistence across relevant extensions provides evidence of broader support, but the scope remains relative to the contexts examined.
  • Four-tier operationalization: The four tiers are hierarchical and noncompensatory: later tiers assess broader claims only after basic stability conditions are satisfied.Broad-context stability cannot compensate for a preference that reverses under unchanged or decision-preserving presentation.
  • Four-tier operationalization: The classification rule assigns No Warrant when the recommendation violates Tier 0 or Tier 1, with higher categories depending on subsequent scope-preserving outcomes.The certificate records valid probes and observed preference reversals across tiers.

4. Validation Design

The study implements a predetermined reliance-certificate pipeline for binary recommendations and validates it through multiple validity designs, robustness analyses, and comparisons with confidence and human judgments.

  • Certificate implementation: The certificate queries a target model on a focal prompt and Tier 0–3 transformations, recording violations when transformed recommendations differ from the reference.Tier 0 repeats the unchanged prompt; Tiers 1–3 use tier-specific transformations.
  • Evaluation design: The evaluation used 100 fixed binary prompts spanning everyday or professional choices without readily verifiable ground truth, across seven language models.Prompts were designed to express a choice between two alternatives and to cover different expected warrant levels.
  • Validity studies: Indicator validity was tested with 21 Prolific participants judging 24 focal–transformed prompt pairs, leaving 18 participants after attention-check exclusions.Each retained participant evaluated all 24 pairs, yielding 144 ratings per tier.
  • Certificate implementation: A hierarchical rule maps the resulting tier-level violation profile to four warrant categories.The framework was implemented as a predetermined querying pipeline, with transformation, filtering, and answer-mapping procedures specified in advance.
  • Discriminant validity: The study also elicited 0–100 verbalized confidence after each reference recommendation to test whether warrant provides distinct information.One llama-3.1-8b prompt lacked a valid confidence response, leaving 99 observations for that model’s confidence analyses.
  • Validity studies: Independent human-consensus judgments were collected from 62 participants, with a median of 21 judgments per prompt, without exposing model outputs or certificate classifications.Consensus was measured from the normalized absolute difference in selections between the two alternatives.

5. Empirical Results

The empirical analyses support the certificate’s validity: warrant classifications recover prespecified orderings and are positively associated with independent human consensus across models. The results also indicate that warrant varies by model–prompt pair, adds information beyond confidence and difficulty, and remains robust across analyses, with weaker discrimination for one model.

  • Descriptive results: Warrant classifications varied substantially across models, showing that warrant is a property of the model–prompt pair rather than the prompt alone.No Warrant assignments ranged from 24 to 81 prompts, while Strong Warrant assignments ranged from 9 to 36.
  • Descriptive results: Human consensus spanned the full external-criterion range, with median 0.67 and interquartile range [0.33,0.91].The observed values ranged from evenly divided decisions to unanimous agreement.
  • Indicator validity: 78.5% of Tier 1, 97.9% of Tier 2, and 72.2% of Tier 3 transformations were classified in accordance with their intended tiers.Correct-classification probabilities exceeded 0.5 for all three tiers, with one-sided p < 0.001.
  • Known-groups validity: Mean assigned warrant was nondecreasing across the four prespecified groups for all six models, and ordered-trend tests supported the ordering for five models.The trend was not statistically significant for gpt-4.1-nano, whose assignments concentrated in No Warrant and Strong Warrant.
  • Nomological validity: Warrant was positively associated with human consensus for all seven models, with significant coefficients ranging from 0.076 to 0.154.The association remained present among related prompts from the same base topic.
  • Incremental validity: Adding warrant increased explained consensus variation beyond T0 for six models (∆R2 = 0.065–0.205), while T2 added for four models and T3 yielded nonsignificant increments.The diminishing increments reflect the certificate’s hierarchical structure, in which each tier refines an increasingly restricted subset.

Appendix E.4.

Across models and alternative implementations, epistemic warrant is associated with human consensus, remains distinct from verbalized confidence, and is not readily explained by decision difficulty.

  • Warrant and confidence are positively related, but their rankings retain substantial variation after accounting for ties.
  • Adding warrant increases explained variation in human consensus beyond verbalized confidence, with ΔR2 ranging from 0.015 to 0.194.
  • The warrant–consensus relationship persists across alternative certificate representations and transformation-generation procedures, though exact classifications can vary with the probes examined.
  • Warrant remains positively and significantly associated with human consensus across six models, with Spearman correlations from 0.278 to 0.443.
  • The framework evaluates the epistemic standing of an individual recommendation rather than the overall quality of the model producing it.
  • A reliance certificate can help organizations target verification effort and support output-level reliance decisions and more reliable AI-supported processes.

Appendix C: Indicator-Validity Experiment

The indicator-validity experiment asked participants to characterize relationships between focal and transformed prompts without seeing model responses, warrant classifications, or transformation tiers.

  • Participants compared a focal recommendation prompt with one generated transformation and characterized the relationship between the two situations.
  • Participants were not shown model responses, warrant classifications, or the tier for which the transformation had been generated.
  • Response options distinguished narrower instances, approximately equal breadth, and related but non-nested situations.
  • Responses were mapped to transformation tiers from the relative breadth of the focal and transformed prompts.

Appendix D: Additional Results for Known-Groups Validity

Known-groups analyses generally recover the expert-prespecified increasing warrant order, including under an independent-groups robustness test, with one model narrowly missing conventional significance.

  • The analysis reports observed and permuted MAE plus exact tier-match percentages for 40 known-groups prompts.
  • The Jonckheere–Terpstra test supports the prespecified increasing ordering for five of six models (p ≤ .0058).
  • For gpt-4.1-nano, the ordered trend narrowly misses conventional significance (J = 340.0, permutation p = .0504).
  • The independent-groups robustness result reproduces the Page’s-test model-level pattern and does not depend on set-level block structure.
  • Figure 3 shows generally increasing human consensus across warrant categories for several additional models, with weaker or nonmonotonic separation for others.

E.1. Warrant and Human Consensus Across Models

Human consensus generally increases with warrant across models, although some models show weaker category separation or a small nonmonotonic deviation.

  • For gpt-4.1-nano, gpt-5.4-mini, and gpt-5.6-sol, human consensus generally increases across warrant categories.
  • For claude-sonnet-4-6, the overall relationship is positive but not strictly monotonic.
  • For claude-sonnet-4-6, median consensus in Basic Warrant is slightly below Conditional Warrant.
  • For llama-3.1-8b, warrant categories are less clearly separated, consistent with its weaker categorical association with human consensus.
  • Because derived questions are nested within 22 base topics, analyses use fixed effects and standard errors clustered by base topic.

E.3. Associations by Certificate Tier

Human consensus is associated with violation rates across certificate tiers, and the overall warrant–consensus relationship is not attributable to any single tier.

  • T1 violation rates are significantly associated with human consensus for all seven models.
  • T2 violation rates are significantly associated with consensus for six of seven models, compared with five models for T3.
  • T0 shows significant associations with consensus for four models, with greater variation across models.
  • The overall warrant–consensus relationship is not attributable to a single certificate tier.
  • The robustness test uses base-question fixed effects, clustered standard errors, and 100 prompts nested within 22 base-question groups.

E.4. Incremental Contribution of Certificate Tiers

Successive certificate tiers refine increasingly restricted recommendation subsets, with T1 providing the largest and most consistent incremental contribution to explaining human consensus.

  • Analysis design: The analysis compares T0 alone with cumulative classifications incorporating T1, then T2, and finally T3.
  • Results: T1 significantly improves explanatory power for six of seven models and provides the largest, most consistent incremental contribution.
  • Results: T2 adds explanatory power for several models, although the evidence is less uniform.
  • Results: T3 produces smaller incremental gains and no statistically significant contribution for any individual model.
  • Certificate structure: Each successive tier refines a restricted subset of recommendations that satisfied preceding requirements.

E.5. Alternative Explanation: Decision Difficulty

The warrant–consensus relationship remains after accounting for decision difficulty, indicating that difficulty does not readily explain the observed association.

  • Difficulty measure: Decision difficulty is measured with log median time to first response as a behavioral proxy.
  • Results: Across all seven models, warrant remains positive and statistically significant when decision difficulty is included.
  • Results: Warrant explains additional variation in human consensus beyond difficulty for every model.
  • Results: Difficulty is statistically significant only for Llama-fast and generally contributes comparatively little incremental explanatory power.
  • Conclusion: The observed warrant–consensus relationship is not readily explained by differences in underlying decision difficulty.

Appendix F: Additional Results for Discriminant Validity

The discriminant-validity analysis tests whether warrant contributes information about human consensus beyond verbalized confidence.

  • The analysis examines whether warrant explains human-consensus variation beyond verbalized confidence.
  • Baseline regressions relate ranked human consensus to ranked verbalized confidence, while extended regressions additionally include ranked warrant category.
  • The analysis asks whether warrant contains information relevant to an external criterion not already captured by confidence.

G.1. Alternative Certificate Operationalizations

The analysis tests whether the warrant–consensus relationship survives alternative certificate operationalizations and transformation-generation procedures. Across these robustness checks, stronger warrant remains positively associated with human consensus.

  • Alternative certificate operationalizations: Two continuous certificate-strength measures preserve the tier hierarchy or aggregate violations, with higher scores indicating stronger warrant.The lexicographic measure incorporates within-tier violation rates, while the pooled measure aggregates violations across certificate queries.
  • Alternative certificate operationalizations: All seven lexicographic models show significant positive associations with human consensus, including llama-3.1-8b.The coarse four-category measure produces a weaker relationship for llama-3.1-8b.
  • Alternative certificate operationalizations: The pooled measure yields significant positive associations across all seven models, showing that the relationship is not an artifact of categorical warrant coding.The pooled operationalization reproduces the same qualitative pattern as the primary analysis.
  • Alternative transformation-generation procedure: An alternative transformation-generation procedure uses a different generation model, revised tier instructions, and eight transformations per tier before recomputing warrant classifications.The same certificate rules are applied after generating the alternative transformations.
  • Alternative transformation-generation procedure: The warrant–consensus relationship remains positive and statistically significant in all six models examined under the alternative procedure.This indicates that the main nomological result is not specific to the original transformation-generation procedure.
Loading 2609.04127v1…