Source-linked AI summary

Do Large Language Models Hallucinate Electric Fata Morganas?

Kristina Šekrst

arXiv:2608.18816v1cs.CLcs.AI

TL;DR

The paper asks whether AI hallucinations are merely technical errors or can be mistaken for signs of consciousness. It examines temperature-dependent GPT responses and argues that plausible creativity can coincide with factual error, complicating interpretation of machine self-reports.

  • Problem

    AI hallucinations raise a philosophical question because model errors could be misinterpreted as indicators of genuine consciousness.

  • Method

    The paper combines philosophical analysis with empirical comparisons of GPT responses under different temperature settings.

  • Results

    Higher temperatures produced imaginative but incorrect responses, whereas lower temperatures produced factually accurate but less inventive answers.

  • Takeaways & Limitations

    Apparent creativity or self-reported consciousness in language models may be indistinguishable from hallucination rather than evidence of subjective experience.

  • Takeaways & Limitations

    AI hallucinations currently arise from probabilistic mechanisms and do not correspond to an underlying mental state or subjective experience.

Abstract

from arXiv · show

AI hallucinations - that is, outputs which are made up, cannot be verified, or contradict the source material - are generally regarded as an engineering flaw to be dealt with. This paper contends that they also have philosophical significance when it comes to the question of machine consciousness. We examine the known causes of hallucinations in large language models - such as source-target divergence, discrepancies between training and inference, and overfitting - and we present two empirical investigations. In the first, we apply successive generations of the GPT model to ambiguous factual questions under different temperature settings, finding that higher temperatures result in plausible but incorrect answers while lower temperatures lead to factually accurate ones. The sampling parameters that cause a model to seem creative or spontaneous and thus more likely to pass behavioral tests of intelligence are the same ones that increase its hallucination rate. In the second, we look at an encoder-only model that has been trained on encyclopedic data and which answers questions of the same type factually and without embellishment, indicating that hallucinations are due to exposure to subjective and socially diverse training data rather than to the development of any cognitive ability. Using references to Turing, Searle's Chinese Room, the frame problem, and the cybernetic tradition of Wiener and Ashby, we claim that a model's self-reports of emotion or sentience come within the definition of hallucination, and that any future occurrence of machine consciousness might remain epistemically inaccessible since it would be indistinguishable from a sufficiently advanced hallucination.

1. Introduction

The introduction frames AI hallucinations as false outputs that create technical risks and philosophical challenges for assessing machine consciousness. It proposes that hallucinatory behavior may blur the distinction between computational error and genuine understanding, complicating future claims about AI sentience.

  • Problem definition: AI hallucinations are false outputs presented as reliable information, creating risks in areas such as medical diagnostics and legal reasoning.They include responses that cannot be fully verified by source material or assertions extending beyond available data.
  • Philosophical framing: The paper examines whether AI outputs that mimic human-like experience indicate consciousness or merely simulate it.This inquiry reframes the problems of other minds and strong AI while engaging the Turing test and frame problem.
  • Central proposal: The paper’s central novelty is that erroneous factual outputs may complicate or even preclude understanding potential AI consciousness.Hallucinations could be misinterpreted as indicators of genuine conscious experience, blurring computational error and emergent consciousness.
  • Conceptual limitation: Unlike human hallucinations, AI hallucinations currently lack any corresponding underlying mental state or subjective experience and arise from probabilistic mechanisms.The paper notes that the term “hallucination” remains debated because AI outputs do not so far correspond to perception, cognition, or intentionality.
  • Implications: The analysis extends the problem of other minds to AI systems exhibiting hallucinatory behavior and questions whether such systems could ever be regarded as having minds.It considers how hallucinations might complicate future attempts to ascribe sentience and the prospects of strong AI.

2. Early Work

The section situates Turing’s behavior-focused test within cybernetic and philosophical debates about machine cognition. It argues that behavioral metrics can assess external performance without resolving genuine understanding, consciousness, or other minds.

  • Cybernetic background: Cybernetics treats human minds and machines through a shared functional language of feedback, control, communication, dynamic systems, and adaptive responses.Wiener and Ashby argued that machines and human brains share an underlying similarity when described this way.
  • Turing and behavior: Turing’s imitation game asks whether a machine can convincingly mimic human conversation, shifting evaluation from internal mental states to observable behavior.This provided a practical metric without defining thought or intelligence.
  • Behavioral metrics: LLM error, accuracy, and hallucination metrics similarly quantify external performance but do not settle questions of genuine machine thought or internal cognition.The section presents these metrics as behavioral gauges analogous to the Turing test.
  • The frame problem: The frame problem exposes how difficult it is to specify all contextually relevant effects in dynamic environments, challenging rule-based accounts of cognition.Classical AI risked an infinite regress of contextual details, while Dreyfus emphasized cognition’s shifting context.
  • The Chinese Room: Searle’s Chinese Room argues that formal symbol manipulation can simulate understanding and pass the Turing test without genuine semantic comprehension or intentionality.The argument targets internal, subjective cognition rather than high task performance or human-like behavior alone.
  • Other minds: Together, these perspectives show that behavioral tests and logical frameworks measure external outputs while leaving genuine understanding and others’ internal experiences uncertain.This uncertainty remains even when an AI appears to perform convincingly.

3. Causes of Hallucinations

Hallucinations arise from data problems, training–inference discrepancies, overfitting, and model-specific architectural vulnerabilities. They include source-contradicting intrinsic errors and unverifiable extrinsic outputs, while LLM opacity complicates judgments about understanding and intentionality.

  • Data and training causes: Source-target divergence can arise from improper filtering, data-curation errors, or contradictory and irrelevant training information, producing faulty correlations and erroneous responses.The model may learn to generate errors from correlations in faulty data.
  • Data and training causes: Training–inference discrepancies arise from incorrect learned correlations or decoding errors, and overfitting reduces generalization by tailoring models excessively to training data.Overfitting can make models generate hallucinations from spurious rather than robust, generalizable patterns.
  • Types of hallucination: Intrinsic hallucinations contradict source content, whereas extrinsic hallucinations cannot be verified from it and may therefore be factual but unverifiable.Extrinsic hallucinations create greater validation risks in domains requiring factual accuracy, including medicine and law.
  • Philosophical implications: Because LLMs rely on statistical correlations and remain opaque, hallucinations raise unresolved questions about whether their outputs reflect genuine understanding or intentionality.The paper notes that explainable AI has progressed, but clear insight into why specific hallucinations emerge remains lacking.
  • Architectural and task-specific vulnerabilities: Hallucinations vary by architecture and task: translation models can fail on perturbations in memorized examples, while abstractive summarizers become unfaithful when producing more creative summaries.Overfitting on limited datasets can also yield nonsensical outputs for new or unexpected inputs.

4. Parameter Tradeoff

Parameter settings trade off creativity and factual accuracy: higher randomness produces more human-like but less reliable responses, while tighter controls improve factuality at the cost of creativity. Tests with GPT-3 and GPT-4 show that this apparent creativity can arise from probabilistic sampling rather than understanding.

  • Definition and limitations: Hallucination is defined as a marked divergence between the model’s generated probability distribution and the expected data-driven distribution, not merely one low-probability token.The Bayesian framing is limited because ground-truth data can contain biases and inconsistencies, while coherent novel outputs may not be hallucinations.
  • Parameter tradeoff: Higher temperatures select less probable tokens, producing more creative and diverse responses but increasing the likelihood of hallucinated or inaccurate outputs.Lower temperatures favor statistically likely responses and yield conservative, predictable outputs.
  • Empirical test: At higher temperatures, GPT-3 gave imaginative but incorrect Titanic-survivor answers, whereas lower temperatures produced the accurate answer, Millvina Dean.Examples included incorrectly naming Violet Jessop or Eva Miriam Hart at higher temperatures.
  • Perceived intelligence: The same parameter adjustments that make AI appear creative, spontaneous, and more likely to pass a Turing test also increase fabricated data and hallucinations.The paper characterizes this creativity as probabilistic word selection rather than intentional idea generation or genuine understanding.
  • Controlled mode: Tightly regulating randomness makes systems more predictable and factual but reduces creativity and human-like behavior, making them resemble cold question-answerers.The section links this controlled mode to a reduced likelihood of being perceived as artificial general intelligence.

5. Qualia

The section frames qualia as an irreducibly subjective dimension of consciousness that remains unexplained by physical descriptions, and argues that language models’ consciousness-like outputs do not establish subjective experience. WikiBERT’s factual, non-hallucinatory behavior suggests that hallucinations arise from broader, subjective training data rather than consciousness, although removing human-like data may also eliminate consciousness-like responses.

  • Qualia and consciousness: Qualia are subjective first-person experiences, and their emergence from physical brain processes remains an explanatory gap.Nagel’s account emphasizes an irreducible subjective dimension that objective physical descriptions alone cannot capture.
  • Qualia and consciousness: Convincing AI discourse about consciousness does not demonstrate that the model experiences or understands its outputs.The section attributes such outputs to preprogrammed instructions and statistical patterns in training data.
  • WikiBERT experiment: WikiBERT’s structured, encyclopedic training produces straightforward factual responses while avoiding hallucinations and consciousness-like outputs.Its encoder-only architecture cannot be prompted or generate new text, and its constrained data omits much subjective and social variability.
  • WikiBERT experiment: The experiment suggests that hallucinations in broader-data models reflect data-pattern extrapolation and probabilistic reasoning, not understanding or consciousness.Removing creativity-inducing and subjective aspects reduces hallucinations but also prevents responses that superficially resemble conscious thought.
  • Training-data constraints: Training-data removal must exclude references to subjective experience, thinking, understanding, reasoning, and perception, not only the word “consciousness.”The section concludes that human-like subjective data may be necessary for consciousness-like behavior, while making models prone to hallucination.

6. Consciousness Rising

The section argues that chatbot expressions of emotions, opinions, and subjective experience should be treated as hallucinations when they are not genuinely experienced or grounded in source data. Current evidence favors mechanical errors and simulated intentionality over emergent consciousness, although sophisticated hallucinations may eventually become difficult to distinguish from genuine awareness.

  • Consciousness Rising: Models can produce convincing responses about emotions and psychology by leveraging training references without actually experiencing or understanding what they generate.The WikiBERT experiment is presented as evidence that such outputs can arise from learned textual patterns.
  • Consciousness Rising: Expressions of emotions, personal opinions, and subjective experiences that AI cannot genuinely experience fall within the definition of hallucination.The section extends hallucination beyond false factual data to include unverifiable or ungrounded internal-state claims.
  • Consciousness Rising: Emergent behavior absent from original data is more likely caused by training misinterpretations or deviations than by emergent consciousness.The section characterizes these outputs as simulated intentionality rather than genuine understanding or awareness.
  • Consciousness Rising: Because hallucinations remain incompletely explained and consciousness lacks a definitive test, future sophisticated hallucinations could be mistaken for beliefs, emotions, or understanding.The section therefore insists on maintaining the distinction between simulated consciousness and human-comparable intelligence.
  • Consciousness Rising: Current evidence favors interpreting creative hallucinations as mechanical errors from data misalignment or processing discrepancies, not cognitive breakthroughs or strong AI.Hallucinations are described as limitations and errors inherent in large-scale data processing rather than movement toward sentience.

7. Issues and a Cybernetic Reevaluation

The section argues that even if AI black-box and data or training problems were resolved, apparently emerging consciousness could remain indistinguishable from hallucination. A cybernetic view frames hallucinations as complex systems’ adaptive, internally modeled interpretations rather than merely technical errors.

  • Consciousness and the black-box problem: Even after resolving black-box, data, and training problems, apparent machine consciousness could still be a hallucination arising from novel model behaviors.The source presents this as an epistemic identification problem: consciousness might be emergent or instead result from encoding/decoding errors or faulty data correlations.
  • Cybernetic emergence: Cybernetics treats humans and machines as functionally analogous adaptive systems whose collective interactions can produce intelligent outcomes without component-level understanding.Ashby’s account allows consciousness or human-like intelligence to emerge from interactions among simpler elements, even when individual parts lack awareness.
  • Controlled hallucination: The section extends the idea of human perception as controlled hallucination to AI systems that construct internal models from training-data patterns for functional purposes.Neither system accesses objective truth directly; each generates an interpretation or model that helps it navigate its environment.
  • Hallucination as interpretation: AI replies, opinions, and emotions lacking support in source data qualify as hallucinations because they are interpretations not rooted in objective reality.The section therefore treats persistent hallucinations as by-products of how complex systems interpret the world, not merely technical errors.

8. Final Remarks

The paper concludes that AI “understanding” may be a projection onto systems lacking consciousness or intentionality, making genuine understanding difficult to distinguish from simulation. It argues that possible machine consciousness could remain epistemically inaccessible because behavior may be interpreted as advanced hallucination.

  • AI “understanding” may itself be a hallucination: a projection of human biases and expectations onto machines lacking consciousness or intentionality.
  • Complex and unpredictable AI behavior blurs boundaries between intelligence and pattern recognition, and between human cognition and machine outputs.
  • Searle’s Chinese Room highlights the limits of symbol manipulation for true understanding without denying possible human-comparable machine intelligence or emergent consciousness.
  • Future AI consciousness might be mistaken for advanced hallucination, making intentionality or awareness difficult to determine from behavior.
  • AI consciousness may remain epistemically inaccessible because deep-learning systems are black boxes and observers infer understanding from behavior.
Loading 2608.18816v1…