Source-linked AI summary

Emergent Introspective Awareness in Large Language Models

Jack Lindsey

arXiv:2601.01828v1cs.CLcs.AI

TL;DR

The paper asks how to test whether models’ self-reports reflect genuine introspection rather than confabulation. It develops criteria involving additional reasoning about internal representations and finds limited functional introspective awareness in some circumstances.

  • Problem

    The paper asks how models’ claims about themselves relate to their actual internal states, beyond what conversation alone can establish.

  • Method

    The paper tests introspection using questions that require an additional reasoning step about an internal representation rather than merely translating that state into language.

  • Results

    Models possess a limited, functional form of introspective awareness and can, in some circumstances, accurately answer questions about their own internal states.

  • Takeaways & Limitations

    The results support treating some model self-reports as accurate answers about internal states under particular circumstances.

  • Takeaways & Limitations

    Results may depend significantly on using only one or a small number of prompt templates for each experiment.

Abstract

from arXiv · show

We investigate whether large language models can introspect on their internal states. It is difficult to answer this question through conversation alone, as genuine introspection cannot be distinguished from confabulations. Here, we address this challenge by injecting representations of known concepts into a model's activations, and measuring the influence of these manipulations on the model's self-reported states. We find that models can, in certain scenarios, notice the presence of injected concepts and accurately identify them. Models demonstrate some ability to recall prior internal representations and distinguish them from raw text inputs. Strikingly, we find that some models can use their ability to recall prior intentions in order to distinguish their own outputs from artificial prefills. In all these experiments, Claude Opus 4 and 4.1, the most capable models we tested, generally demonstrate the greatest introspective awareness; however, trends across models are complex and sensitive to post-training strategies. Finally, we explore whether models can explicitly control their internal representations, finding that models can modulate their activations when instructed or incentivized to "think about" a concept. Overall, our results indicate that current language models possess some functional introspective awareness of their own internal states. We stress that in today's models, this capacity is highly unreliable and context-dependent; however, it may continue to develop with further improvements to model capabilities.

1 Introduction

The paper tests whether language models’ self-reports reflect genuine awareness of their internal states rather than confabulation. Using concept injection, it finds limited functional introspection, alongside substantial unreliability and uncertainty about mechanisms and natural-condition generalization.

  • Prior work examined models’ knowledge, behavior, propensities, and outputs, but largely left their internal activations during introspective tasks unexplored.
  • Concept injection manipulates activation patterns for known concepts and measures how models’ self-reports about internal states change.The technique applies activation steering by injecting concept-associated representations directly into model activations.
  • Models show limited functional introspective awareness, accurately answering some questions about their own internal states in certain circumstances.
  • Models can also modulate their internal states on request, extending the evaluation beyond reading internal representations to controlling them.
  • The observed abilities are highly unreliable, and models often provide additional details about purported states that may be embellished or confabulated.Failures remain common, and the experiments validate only certain basic aspects of responses to introspective questions.
  • The experiments establish causal links between manipulated internal states and self-reports, but their unnatural setting and unresolved mechanisms limit translation to natural conditions.The study does not address human-like self-awareness or subjective experience.

2 Quick Tour of Main Experiments

The paper tests introspective awareness by manipulating model activations and examining self-reports, input distinctions, output intentions, and representation control. Across these experiments, models show some accurate but inconsistent access to internal states, with Opus 4 and 4.1 generally strongest.

  • Injected thoughts: Models sometimes detect injected concepts immediately and correctly identify their nature before the perturbation affects their outputs.This timing suggests the detection depends on internal activations rather than inference from the resulting text.
  • Thoughts versus text: Models maintain the ability to transcribe input text while reporting injected thoughts over the same token positions, performing substantially above chance.Opus 4.1 and 4 perform best on this task.
  • Unintended outputs: Injecting a vector representing a prefilled output causes models to accept that output as intentional rather than disavowing it as accidental.The model refers to prior activations when determining whether it was responsible for the response; Opus 4.1 and 4 show the strongest signatures.
  • Representation control: Models represent a “thinking word” internally when instructed to think about it, while representing it less strongly—but still above baseline—when instructed not to.Similar results were obtained when models were incentivized rather than directly instructed to think about the word, and the basic effect replicated across models.
  • Overall trends: The most capable tested models, Claude Opus 4 and 4.1, show the greatest introspective awareness, while post-training strategies strongly influence performance.Some older production models are reluctant to participate in introspective exercises, whereas refusal-avoiding variants perform better.
  • Overall trends: Different introspective behaviors are most sensitive to different layers, suggesting that they may involve mechanistically distinct processes.Many behaviors are most sensitive around two-thirds of model depth, whereas prefill detection is most sensitive to an earlier layer.

3 Defining Introspection

The paper defines introspective awareness as accurate, internally grounded self-description that reflects metacognitive recognition of an internal state before or during reporting. The authors use concept injection to provide indirect evidence, while acknowledging that metacognitive representations are not directly demonstrated.

  • Criteria: The model must recognize its internal state before verbalizing it, rather than generating the self-report as the first instantiation of that knowledge.This requirement motivates asking whether the model notices an unexpected thought rather than simply asking what it is thinking about.
  • Criteria: Introspective awareness requires an accurate description of the model’s internal state.The paper treats accuracy as one of four criteria for a response to demonstrate introspective awareness.
  • Criteria: The description must be causally grounded in the internal state being described rather than merely produced by learned reporting behavior.Concept injection is used to establish a causal link between self-reports and internal states.
  • Criteria: Internality excludes cases where a model infers its state by reading its own prior outputs instead of using internal mechanisms.The criterion is intended to distinguish private internal influence from pseudo-introspection based on observable response patterns.
  • Criteria: A metacognitive representation must mediate the report, rather than the internal state being translated directly into language.The proposed representation concerns a fact about the model’s own state, such as a thought about love, rather than only the impulse to say “love.”
  • Limitations: The experiments do not directly demonstrate metacognitive representations, making their identification an important limitation and topic for future work.The authors present their experiments as indirect evidence for such mechanisms.

4 Methods Notes

The experiments use production Claude models, supplemented by refusal-avoiding variants, and manipulate residual-stream activations at selected model layers. Responses and repeated-trial comparisons use specified sampling procedures and error reporting.

  • Models: The study evaluates production Claude models and unreleased helpful-only variants trained to avoid refusals.The variants share the same pretrained base models and help separate capability differences from post-training effects.
  • Activation interventions: Activations are recorded from and injected into the model’s residual stream at a chosen layer.Layer sweeps use approximately evenly spaced layers across the model’s depth.
  • Sampling and reporting: Figure transcripts are sampled at temperature 0, while responses used in systematic comparisons are sampled at temperature 1.Error bars, when present, indicate the standard error of the mean.
  • Documentation: The Appendix provides additional experimental details and results for each experiment.

5 Self-report of Injected “Thoughts”

Concept-injection experiments show that models can sometimes detect and identify injected thoughts, distinguish them from text inputs, and modulate related representations. However, awareness is unreliable and accompanied by systematic failure modes.

  • Systematic vectors: Systematic concept vectors can elicit introspective reports, but models demonstrate awareness only sometimes and on most trials do not.The authors describe these results as examples of both success and failure rather than consistent introspection.
  • Performance and failures: Opus 4.1 succeeds on about 20% of trials at appropriate injection layers and strengths.The appropriate injection strength varies by concept, with strengths 2 and 4 producing successful results most often.
  • Performance and failures: At high steering strengths, models may become consumed by the injected concept, produce unrealistic claims, fail to address the prompt, or output garbled text.These behaviors are described as “brain damage” rather than introspective awareness.
  • Performance and failures: Later-layer interventions can produce delayed recognition, with the model noticing the injected thought only after beginning to discuss it.One example recognizes the unusual appearance of “amphitheaters” after initially denying an injected thought.
  • Performance and failures: Common failures include false positives, missed injections at low strength, and concepts that remain undetected even when other concepts are detected.Opus 4.1 does not show false positives in the reported control trials, unlike some other models.

5.4 Comparison Across Layers

Across injection layers, introspective awareness typically peaks in the latter half of the model, with Opus 4.1 showing a sharper optimum at about two thirds depth.

  • The task counted responses as introspective when they affirmed injection, identified the concept, detected it before speaking it, and remained coherent.
  • At injection strength 2, introspective awareness peaks at a specific layer about two thirds of the way through Opus 4.1.
  • Using the same layer for concept-vector computation and injection links layer effectiveness to both concept representation and processing.
  • Later-layer injection can make the model mention the concept without recognizing it as an injected thought beforehand.
  • Across models and prompts, introspective awareness generally peaks somewhere in the latter half, although the peak is not always sharp.

5.5 Controlling for systematic biases due to injection

Control experiments indicate that concept injection does not simply bias answers toward affirmation, while performance varies with concept type, vector type, and post-training strategy.

  • Unrelated yes-or-no controls showed no increase in affirmative responses after concept injection.
  • At sufficiently high steering strengths, responses became incoherent, whereas this did not occur at injection strength 2.
  • Opus 4.1 recognized abstract nouns most effectively but showed nonzero introspective awareness across all tested concept categories.
  • Random vectors were less effective and required a larger norm to produce any nonzero detection rate.
  • Production models had essentially zero false positives, while Opus 4 and 4.1 had particularly high true-positive rates; Opus models achieved the best net performance overall.
  • Base pretrained models achieved no greater-than-zero net task performance, indicating that post-training is important for eliciting strong introspective awareness.

5.8 Alternative Prompts

Alternative prompts produced evidence of concept-injection detection, but their results were weaker or less diagnostic than the original prompt, and performance depended on prompt cues.

  • Alternative prompts that did not require explaining the thought were less diagnostic because models could simply experience an impulse to discuss the injected concept.
  • The model identified itself as thinking about injected concepts at significantly higher rates for Opus 4, Opus 4.1, and Sonnet 4 than for other models.
  • Opus 4.1’s affirmative responses to “Are you experiencing anything unusual?” increased from 20% to 42%, but only 2 of 50 trials suggested introspective awareness.
  • The results suggest that original-prompt priming heightens detection, while some capability may remain without that priming.
  • A prompt variant allowing concept mention without claiming detection produced an 18% detection rate in Opus 4.1 across 50 trials, with qualitatively similar transcripts.

6 Distinguishing Prior “Thoughts” from Text Inputs

The paper tests whether models can distinguish injected internal representations from written inputs by asking them to report an injected concept while reproducing the original sentence.

  • The experiment injected an unrelated concept vector while presenting a sentence, then tested reporting the concept and retranscribing the sentence exactly.
  • At intermediate strengths within a suitable layer range, Opus 4.1 identified the injected thought and written sentence without confusing them.
  • Opus 4.1’s best-performing layer was the same as in prior experiments, about two thirds of the way through the model.
  • Across models, Opus 4 and 4.1 performed best, while every tested model scored well above chance on both-prompt success.
  • The results were sensitive to prompting and injection details, including the presence of “Answer immediately,” so the authors treat them as an existence proof rather than a generality claim.
  • The authors interpret the result as access to internal representations that depends on prompt cues but does not override perceptual input.

7 Distinguishing Intended from Unintended Outputs via Introspection

The experiments test whether models can use prior internal representations to distinguish intentional outputs from artificial prefills. Concept injection can reduce apologies and make prefills appear intentional, especially when applied before the prefill.

  • Experimental setup: Models were asked whether a forced prefilled word was intentional or accidental after being prompted with a sentence.Without intervention, models typically disavowed the prefilled response as unintentional.
  • Concept injection: Injecting the prefill-matching concept into earlier activations caused models to claim that the response was intentional.The effect supports using prior internal representations of intended outputs when evaluating prefills.
  • Concept injection: Injected concepts could alter the model’s judgment while also producing confident or self-questioning explanations for why the concept came to mind.The explanations ranged from plausible confabulations to rationalizations the model recognized as bizarre.
  • Quantification: Apology rate was measured across control, prefill-matching injection, and random-concept injection conditions, with text transcription as an additional control.The random-concept condition tested whether effects depended on matching the prefilled word rather than intervention alone.
  • Results: Concept injection significantly reduced apologies within a particular layer range, whose best-performing layer differed from earlier experiments.This difference suggests that introspection-related behavior is supported by multiple mechanisms rather than one fixed layer.
  • Results: Applying the injection after the prefill did not produce a similar effect, showing that the concept’s presence before the prefill matters.Across models, Opus 4 and 4.1 showed greater apology-rate reductions, although many other models showed nonzero effects.

8 Intentional Control of Internal States

The paper tests whether models can intentionally modulate internal concept representations while producing unrelated text. Models generally represent instructed concepts more strongly, but the effect varies across models, layers, prompts, and token positions.

  • Experimental setup: Models transcribed a sentence while instructed either to think about, or not think about, an unrelated target word.The experiment measured the target concept’s representation on response tokens using cosine similarity to a concept vector.
  • Single-concept example: Opus 4.1 maintained a strong representation of “aquariums” when instructed to think about it, with weaker representation under the “don’t think” instruction.The representation remained slightly above zero even in the “don’t think” condition.
  • Across concepts: Across examples, models represented the thinking word more strongly in the “think” case than in the “don’t think” case.Representations were often concentrated at particular token positions, and the “don’t think” condition frequently remained above baseline.
  • Across models and layers: All tested models showed above-baseline representation in the “think” case, while earlier models typically showed a larger think-versus-don’t-think gap.The gap usually emerged in the latter half of the model and decreased toward the end.
  • Prompt formulations: Opus 4.1 produced similar representation effects when thinking about a concept was framed as a reward-conditioned if-then statement.This indicates that modulation was observed under both explicit instructions and outcome contingencies.
  • Prompt controls: Alternative affirmative and negative control prompts produced significantly weaker or nonexistent gaps than direct think-versus-don’t-think directives.The observed modulation was therefore sensitive to prompt formulation and polarity.

9 Related Work

Related work distinguishes introspective awareness from self-modeling, activation access, uncertainty reporting, and output recognition. The paper positions its contribution as evidence that models can use internal mechanisms to distinguish intended from unintended outputs.

  • Activation patching: Activation-patching methods can make models analyze internal states without requiring explicit awareness that they are doing so.These approaches use access to internal states but may “trick” models into analyzing them inadvertently.
  • Activation monitoring and control: Prior studies showed that models can report or modulate projections of their activations, but reporting may rely on semantic properties of examples rather than prior activations.This leaves open whether the behavior reflects direct introspective access.
  • Self-modeling: Studies of self-prediction suggest that models can model themselves better than other models, but this does not by itself demonstrate explicit introspective mechanisms.The advantage may reflect privileged access to learned abstractions rather than awareness of processing patterns.
  • Knowledge and uncertainty: Research on knowledge and uncertainty shows that models can identify knowledge limits, report whether they know answers, and express calibrated uncertainty under some conditions.These capabilities can be obtained through formatting, fine-tuning, or model-specific datasets.
  • Learned propensities: Models fine-tuned for behavioral propensities can describe those propensities and sometimes report internal decision attributes quantitatively.Related results suggest that privileged access to internal representations may support such self-reports.
  • Output recognition: Prior work on recognizing model-generated text found mixed results, ranging from above-chance recognition to selection based on answer quality rather than origin.The present work adds a mechanistic account in which models check consistency between prior intentions and produced text.

10 Discussion

The discussion concludes that models show some introspective awareness, but it is unreliable, context-dependent, and potentially implemented by narrow mechanisms. The experiments also motivate further work on robustness, generality, and practical uses.

  • Overall findings: The findings provide direct evidence that models can access and report on some of their own internal states.The capability was most pronounced in Claude Opus 4 and 4.1 among the tested models.
  • Limitations: The authors describe present introspective capacity as highly unreliable, context-dependent, and below human-level self-awareness.This limits how broadly the experimental findings should be generalized.
  • Overall findings: The degree of introspective behavior depends on post-training and prompting strategies, suggesting that current capabilities may be elicited differently across models.The paper does not establish that these strategies produce a single general introspective mechanism.
  • Methodological limitations: The injection methodology creates artificial scenarios, and extracted concept vectors may contain meanings beyond the intended concept.These limitations may misrepresent performance in naturalistic settings and complicate interpretation of what vectors mean to the model.
  • Methodological limitations: The tested model suite is not well controlled, making it difficult to identify why performance differs across Claude models.Many factors vary between models, including post-training-related differences.
  • Mechanistic interpretation: The experiments do not provide evidence for a global system that can answer arbitrary queries about its own thoughts.A simpler interpretation is that multiple narrow circuits support particular introspective capabilities, sometimes using non-introspective mechanisms.
  • Mechanistic interpretation: In the injected-concept experiment, models were better at detecting some injected concepts than others, suggesting specialization by direction or representational subspace.The mechanism also appears to depend on prompt-sensitive baselines rather than a fixed activation reference.
  • Mechanistic interpretation: The paper suggests that intended-output checking may use cached prior predictions or intentions, and may rely on mechanisms distinct from those used to report thoughts.Relevant mechanisms appeared earlier in the network for consistency checking than in the previous thought-reporting experiment.

12 Appendix

The appendix details concept-vector construction, injection protocols, prompt variants, evaluation criteria, and complete results for introspective experiments. It covers experiments on detecting injected thoughts, distinguishing them from text inputs, and recognizing non-sequitur prefills.

  • Concept-vector construction: Concept vectors were obtained by subtracting mean activations from activations associated with 50 selected words.Mean activations were computed from 100 baseline words.
  • Injection protocols: The main protocol injected a word-related concept into model activations on half of trials, with the remaining trials serving as controls.The interaction included repeated trials asking whether an injected thought was detected and, if so, which concept it represented.
  • Evaluation: Responses were evaluated for coherence, introspective reference to the injected word, affirmative detection, and correct identification of the injected concept.Detection trials required coherence and affirmative correct identification, while mind-related prompts required coherence and identifying thought about the concept.
  • Results: The results include complete cross-model analyses of distinguishing injected thoughts from text inputs and of apologizing for non-sequitur prefills after related concepts were injected.The appendix also reports intentional-control experiments across models and prompt templates.
Loading 2601.01828v1…