Source-linked AI summary

Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes

Meng Li, Michael Vrazitulis, David Schlangen

arXiv:2506.01512v1cs.CLcs.AI

TL;DR

Current LLMs must express evidence-sensitive distinctions between fact, imagination, and confidence, but their underlying linguistic knowledge of uncertainty is not well established. The paper evaluates epistemic modality with typologically informed controlled stories and finds limited, unreliable performance, motivating richer semantic representations of epistemic modality.

  • Problem

    The paper addresses the lack of careful evidence about whether LLMs possess the linguistic knowledge needed to generate epistemic uncertainty expressions.

  • Method

    The study evaluates open-weight LLMs with typologically informed controlled stories that manipulate evidence, modal semantics, and commitment conditions.

  • Results

    LLMs show limited performance in generating appropriate epistemic expressions, with higher accuracy for necessity than possibility and for facts than beliefs.

  • Takeaways & Limitations

    Uncertainty-aware LLMs require improved semantic representations of epistemic modality in addition to confidence estimation and calibration.

  • Takeaways & Limitations

    The study focuses on English and does not directly test human participants or examine modality’s morphological encoding in low-resource languages.

Abstract

from arXiv · show

Rational speakers are supposed to know what they know and what they do not know, and to generate expressions matching the strength of evidence. In contrast, it is still a challenge for current large language models to generate corresponding utterances based on the assessment of facts and confidence in an uncertain real-world environment. While it has recently become popular to estimate and calibrate confidence of LLMs with verbalized uncertainty, what is lacking is a careful examination of the linguistic knowledge of uncertainty encoded in the latent space of LLMs. In this paper, we draw on typological frameworks of epistemic expressions to evaluate LLMs' knowledge of epistemic modality, using controlled stories. Our experiments show that the performance of LLMs in generating epistemic expressions is limited and not robust, and hence the expressions of uncertainty generated by LLMs are not always reliable. To build uncertainty-aware LLMs, it is necessary to enrich semantic knowledge of epistemic modality in LLMs.

1 Introduction

The paper argues that LLMs must distinguish facts, imagination, and evidence strength, but their linguistic knowledge of epistemic modality has not been carefully examined. It evaluates this knowledge with typologically informed controlled stories and finds limited performance in selecting appropriate epistemic expressions.

  • 1 Introduction: LLMs should match epistemic expressions to evidence and confidence when distinguishing fact from imagination.Epistemic expressions encode necessity and possibility according to available evidence.
  • 1 Introduction: Existing confidence-estimation work assumes fluent LLMs already know how uncertainty expressions form and mean, leaving this assumption untested.Prior approaches focus on eliciting and calibrating verbalized uncertainty rather than examining the underlying linguistic knowledge.
  • 1 Introduction: Epistemic meanings comprise evidence about information sources and commitment representing a speaker’s degree of certainty.Evidence divides into direct and indirect evidence, while commitment spans full, partial, and neutral support.
  • 1 Introduction: English expresses uncertainty through modal adjectives and adverbs, attitude verbs, and modal auxiliaries or semi-auxiliaries.These forms include probability expressions, mental-state descriptions, and necessity or possibility modals.
  • 1 Introduction: Controlled stories address limitations of knowledge-intensive datasets by controlling evidence types, commitment degrees, prompt formats, and modal semantics.They simplify reasoning while testing how these factors affect epistemic-expression generation.
  • 1 Introduction: The study reports limited epistemic-modality performance and finds effects of model size, prompt format, and modal semantics, including better necessity than possibility performance.For attitude verbs, models perform better on facts than beliefs under different certainty levels.

2 Related Work

Related work addresses confidence estimation, uncertainty expressions, theory of mind, and modal reasoning in LLMs. The paper positions its contribution as a behavioral evaluation of epistemic linguistic knowledge using controlled stories rather than a test of general logical reasoning.

  • 2 Related Work: Confidence-estimation research includes logit-based, internal-state, verbalized, consistency-based, and surrogate methods for eliciting or measuring uncertainty.Verbalized approaches prompt models to output uncertainty in words or numbers.
  • 2 Related Work: Studies also show that uncertainty expressions affect LLM performance and that models may avoid expressing uncertainty when answers are wrong.Other work defines metrics for faithful response uncertainty.
  • 2 Related Work: Probability adjectives and adverbs have received particular attention, while neural language models struggle with probability words.Their popularity partly reflects the ease of mapping expressions to explicit numerical responses.
  • 2 Related Work: Theory of mind research links understanding epistemic modality to representing mental states and uses templated information-asymmetry stories to evaluate language models.This connects children’s semantic development with model question-answering evaluations.
  • 2 Related Work: The paper studies linguistic knowledge in simple controlled stories without assuming a specific logical or probabilistic reasoning theory.Logical reasoning complexity is intentionally limited rather than treated as the evaluation focus.

3 Method

The method infers semantic knowledge from observable model behavior on experimentally controlled, contextualized stories. Stimuli combine manual writing with template-based generation while controlling prompt format, modal semantics, and story type across eight instruction-tuned open-weight models.

  • 3 Method: The study infers semantic knowledge from responses to specific stimuli in experimentally controlled settings.Meaning is treated as indirectly observable through behavior rather than directly measurable.
  • 3 Method: Contextualized stories are used because modal-word meanings are abstract and highly context-dependent.The design follows paradigms from hidden-object tasks and linguistic fieldwork storyboards.
  • 3 Method: Prompt format, modal semantics, and story types are controlled to create targeted contexts for testing specific linguistic forms.This enables hypotheses about epistemic expressions to be evaluated behaviorally.
  • 3 Method: Stimuli combine manual writing with template-based generation.This places the study between manually authored and procedurally generated datasets.
  • 3 Method: The evaluation covers eight open-weight instruction-tuned models from the Llama3, Llama3.1, Qwen2, and Qwen2.5 families.Models range from 7–8B to 70–72B parameters and use greedy decoding.

4 Experiment 1: Modal Auxiliaries

Experiment 1 tests whether LLMs distinguish possibility from necessity in controlled modal scenarios. It examines how model size, story type, and question-answer format relate to response accuracy.

  • The experiment evaluated eight instruction-tuned models on stories requiring selection between may/might and must/have to.Models were grouped into small (7–8B) and medium (70–72B) parameter categories.
  • Experimental Design: Across five story templates and 150 stories, necessity conditions ruled out alternatives, whereas possibility conditions left multiple outcomes plausible.The design included base, 1-shot, and counterfactual stories, with 50 stories of each type.
  • Experimental Design: Responses were collected through direct-slot, indirect-slot, and indirect-sentence formats, and performance was assessed using accuracy and paired accuracy.Paired accuracy counted a response as correct only when both a necessity and a possibility question in a pair were answered correctly.
  • Large models showed substantially higher accuracy than small models, with descriptive accuracies of 79.6%–95.3% versus 55.1%–72.3%.
  • Statistical Analysis and Results: Logistic regression showed significant effects of parameter count and modal condition across all model families, with higher accuracy for necessity than possibility trials.The parameter-by-condition interaction was significant across models but differed in sign and magnitude, while story type and QA format had small, inconsistent effects.

5 Experiment 2: Attitude Verbs

Experiment 2 tests whether LLMs can select attitude verbs that match facts or beliefs and differing certainty levels in controlled Theory-of-Mind stories. Medium-sized models generally perform better, but certainty effects vary by model family and statement type.

  • Experimental Design: The experiment asks LLMs to choose know, believe, or doubt for facts and beliefs with different certainty levels in Theory-of-Mind stories.Thirty stories yield contrastive statements covering previous and current facts and agents’ beliefs.
  • Results: Medium-sized models achieve about 20 percentage points higher accuracy than small models, at 60.8–72.8% versus 36.4–56.9%.Table 2 reports mean and paired accuracies for Experiment 2.
  • Statistical Analysis: The analysis evaluates parameter size, epistemic certainty, statement type, and question-answer format using logistic regression.It also tests interactions between parameter size and the remaining factors.
  • Results: For all assessed LLM classes except Qwen2, increasing parameters significantly improves response accuracy.Qwen2 shows no parameter effect, whereas Qwen2.5, Llama3, and Llama3.1 show significant increases.
  • Results: Higher-certainty verbs usually yield higher accuracy, but this pattern reverses for medium-sized Llama3 and Llama3.1 models.For small models and medium-sized Qwen models, believe or know outperform doubt; medium-sized Llama models show the opposite pattern.
  • Results: Fact-based statements receive substantially higher accuracies than belief-based statements across all LLM classes.The effect is large and highly significant for Qwen2, Qwen2.5, Llama3, and Llama3.1.

6 Discussion

The discussion situates the findings within research on modal and attitude-verb learning and identifies boundaries for extending epistemic-modality benchmarks. Future work should address low-resource-language morphology, multimodal evidence, and conflicting evidence.

  • Epistemic Modal Verbs: Children’s and LLMs’ performance on epistemic expressions can depend on syntactic-semantic context and task format.Ozturk and Papafragou (2015) found children better at evaluating statements than answering questions, while the paper examines related format effects in LLMs.
  • Attitude Verbs: The study raises whether current training pipelines can learn attitude-verb meanings and whether language models can learn them without consciousness.These questions connect computational-linguistic evaluation with theories of language acquisition.
  • Forms of Modality in Low-Resource Languages: Modal morphology in low-resource languages, including modal affixes and modal case, remains untested for truthful multilingual LLMs.These cross-linguistic forms pose challenges for multilingual evaluation and modeling.
  • Enriching Benchmarks with Multimodal Evidence and Complex Reasoning: Both experiments use text-based stories, leaving epistemic reasoning in multimodal environments for future evaluation.The discussion connects this gap to the growing importance of embodied intelligence.
  • Enriching Benchmarks with Multimodal Evidence and Complex Reasoning: The controlled tasks limit unnecessary world knowledge and reasoning complexity, but LLMs still need to improve handling of conflicting evidence.The paper cites Kazemi et al. (2023) and Wan et al. (2024) in this context.

7 Conclusion

The paper evaluates epistemic-modality knowledge in open-weights LLMs with controlled stories and finds limited performance in generating appropriate epistemic expressions. It concludes that reliable uncertainty-aware systems require richer semantic representations of epistemic modality.

  • Conclusion: Controlled-story experiments show limited performance by open-weights LLMs in generating appropriate epistemic expressions.The evaluation targets semantic knowledge of epistemic modality.
  • Conclusion: The findings imply that LLM responses containing epistemic uncertainty may be unreliable.The paper identifies insufficient semantic knowledge of epistemic modality as a potential reason for this unreliability.
  • Conclusion: Building rational LLMs requires enriching semantic representations of epistemic modality alongside improving uncertainty estimation and calibration.The conclusion presents both directions as necessary for uncertainty-aware systems.

Limitations

The paper evaluates epistemic modality behaviorally but identifies limits in its evidence base and scope. It also frames the experiments through a semantic map relating evidence and commitment.

  • The study did not test human participants directly, relying on indirect evidence from child acquisition profiles and adult-group literature.
  • The analysis did not use logit probabilities to create intrinsic-uncertainty metrics or map model logits to self-reported human responses.
  • The study focuses on English and does not examine morphological modality encoding in low-resource languages or subword-tokenization effects.
  • The paper uses a semantic map to represent cross-linguistic semantic regularities and relate the evidence and commitment categories to its two experiments.

C.1 Experiment 1

The experiments model accuracy with logistic regressions covering parameter count and task factors, while testing whether interactions with scale alter response patterns. AIC-based selection determines whether random effects and some interactions are retained.

  • The analyses assess main effects of parameter count, story type, modal condition, and QA format, plus parameter-count interactions with the other factors.
  • The models test parameter scaling because larger models might display qualitative response differences that modify other factor effects.
  • AIC-based model selection found that assessed random effects did not substantially improve fit, so simple logistic regressions without random effects were fitted.
  • Some fixed-effect interactions were dropped when their fit improvement did not offset the AIC penalty for additional parameters.
  • Factors use sum-to-zero effect coding, with Experiment 1 results summarized in Tables 7–10 and ROC curves shown in Figure 10.
  • Experiment 2 models assess parameter count, epistemic certainty, statement type, QA format, and their interactions with parameter count.

D.1 Experiment 1

Figures 6–9 report response accuracy with 95% confidence intervals across model classes, parameter counts, and experimental task conditions. The figures separate story type and QA format in Experiment 1 from statement type and modal condition in Experiment 2.

  • Figures 6 and 7 show Experiment 1 response accuracy by LLM and parameter count, split respectively by story type and QA format.
  • All four figures display error bars representing 95% confidence intervals.
  • Figures 8 and 9 show Experiment 2 response accuracy by LLM and parameter count, split respectively by statement type and modal condition.

E.1 Experiment 1

Experiment 1 uses controlled stories and contrastive questions to test whether models select epistemically appropriate possibility or necessity expressions. The materials vary story templates, certainty conditions, and response formats.

  • Experiment 1 generates stories from five templates: hidden object, Whodunnit, city travel, grocery-store promotion, and fictional characters.
  • Some templates were inspired by Wu et al. (2024) and Jin et al. (2024), while the others were designed by the authors.
  • The hidden-object materials contrast cases where evidence leaves alternatives open with cases where only one location remains possible.
  • The stimulus set includes possibility and necessity judgments about whether a peanut may be in a box or has to be there.
  • Counterfactual stories ask whether a peanut would have been in another box if a prior move had not occurred, requiring a choice between may and has to.
  • Responses use direct words, numbered word choices, or numbered sentence choices across direct-slot, indirect-slot, and indirect-sentence formats.

F.1 Experiment 1

This section presents logistic-regression analyses of LLM response data across two experiments and multiple Qwen and Llama model variants. The analyses include AIC-selected optimal models, with ROC curves summarizing predictive performance.

  • F.1 Experiment 1: Tables 7–10 report AIC-selected logistic-regression results for Qwen2, Qwen2.5, Llama3, and Llama3.1 data from Experiment 1.
  • Experiment 2: Tables 11–14 report AIC-selected logistic-regression results for Qwen2, Qwen2.5, Llama3, and Llama3.1 data from Experiment 2.
  • Predictive analyses: Figures 10 and 11 show ROC curves for logistic-regression models predicting LLM response data in Experiments 1 and 2, respectively.
Loading 2506.01512v1…