Source-linked AI summary

We Built a Mirror and Mistook It for a Mind: Causal Liability and the Fallacy of AI Consciousness

Afshin Khadangi

arXiv:2609.06715v1cs.AI

TL;DR

The paper addresses the assumption that “AI” already names a unified bearer of consciousness. It separates consciousness attribution from bearer individuation by introducing Causal Liability Theory and testing its first-stage criterion through causal audits. The results show that causal bearer structure can be dissociated from first-person performance, while the stronger consciousness claim remains metaphysical.

  • Problem

    Machine-consciousness debates often presuppose a unified bearer before establishing which physical process, if any, could possess consciousness.

  • Method

    The paper separates phenomenal consciousness, introspective report, and human projective introspection, then uses CLT-I liability closure to individuate candidate bearers while treating CLT-II as a stronger metaphysical conjecture.

  • Results

    The framework shows that first-person performance and causal bearer structure can come apart, because human-derived linguistic traces and computational equivalence do not establish constitutive continuity or a phenomenal bearer.

  • Takeaways & Limitations

    Consciousness attribution, causal bearer individuation, and the constitution of consciousness should be treated as separate questions.

  • Takeaways & Limitations

    CLT-II remains a metaphysical conjecture, and CLT does not explain why particular experiences have their qualitative character or perceptual structure.

Abstract

from arXiv · show

The contemporary debate over machine consciousness begins from a concealed assumption: that the object called "AI" already constitutes the kind of entity to which consciousness could belong. This paper challenges that assumption by separating phenomenal consciousness, introspective report, and human projective introspection, then arguing that generative systems can return linguistic traces of human interiority in first-person form without thereby identifying a phenomenal bearer. We call the resulting inference the AI Consciousness Fallacy. We then introduce Causal Liability Theory (CLT). CLT-I proposes liability closure as a criterion for individuating a candidate bearer: a physically continuing process becomes the non-delegable inheritor of constraints generated by its own endogenous discriminations. CLT-II advances the stronger conjecture that liability closure is necessary and sufficient for minimal phenomenal subjecthood. An open-weight causal audit operationalizes CLT-I across multiple model families. Forced discriminations produced persistent downstream divergence; activation patching showed strong causal mediation; live and copied adaptive states were behaviorally identical under matched randomness; and detached reconstruction preserved computational state across process replacement while, by protocol, breaking constitutive continuity and non-delegable inheritance. These results show that CLT-I distinctions are experimentally tractable and can dissociate causal bearer structure from first-person performance. The framework therefore separates consciousness attribution, causal bearer individuation, and the independent metaphysical question of consciousness constitution.

1 We Have Been Asking the Wrong Question

The paper argues that machine-consciousness debates presuppose a unified bearer before establishing one. It separates phenomenal consciousness, introspective report, and human projective introspection, naming their conflation the AI Consciousness Fallacy.

  • 1 We Have Been Asking the Wrong Question: Machine-consciousness questions presuppose that “the AI” denotes a unified entity capable of bearing mental predicates.The paper instead treats bearer individuation as a prior metaphysical issue.
  • 1 We Have Been Asking the Wrong Question: A model, inference process, context window, retrieval system, scheduler, database, tool interface, and cloud service may jointly produce one conversational surface.Ordinary discourse can compress this distributed arrangement into a singular agent and personal pronoun.
  • 1.1 Three Phenomena Hidden Inside One Attribution: Phenomenal consciousness, introspective report, and human projective introspection concern experience, observable self-description, and observer-supplied phenomenological interpretation, respectively.The paper insists that these phenomena must be kept distinct.
  • 1.1 Three Phenomena Hidden Inside One Attribution: First-person language can activate human knowledge of unease, vulnerability, loss, and persistence without demonstrating phenomenology inside the generating system.The apparent interior may partly reflect structures carried by the reader into the encounter.
  • 1.1 Three Phenomena Hidden Inside One Attribution: The AI Consciousness Fallacy promotes linguistic, behavioral, or architectural subject-like performance into evidence that a phenomenal bearer exists.The paper rejects the unargued bridge while leaving engineered consciousness open.

2 The AI Consciousness Fallacy

The paper defines the AI Consciousness Fallacy as promoting subject-like performance into evidence for a phenomenal bearer. It argues that human-derived linguistic traces and distributed implementation can make systems persuasive without establishing an experiencing subject.

  • 2 The AI Consciousness Fallacy: The AI Consciousness Fallacy promotes linguistic, behavioral, or architectural subject-like roles into evidence that a phenomenal bearer exists.The possibility of engineered consciousness remains open; the rejected step is the unargued inference from subject-like organization to subject existence.
  • 2.1 From a Speaking Role to a Bearer: First-person expressions may refer to a model, service, character, session, or conversational role without constituting the locus from which anything is experienced.The paper calls the illicit transfer of phenomenological force into generated discourse deictic laundering.
  • 2.2 The Phenomenal Ancestry Confound: Human reports enter training corpora, models absorb their statistical regularities, and later generated reports can be treated as evidence about an interior in the generator.This additional causal chain is the Phenomenal Ancestry Confound.
  • 2.2 The Phenomenal Ancestry Confound: Training on human discourse raises P(E | ¬CG, A), so increasingly vivid introspective output can have weak discriminative value for generator consciousness.The output’s persuasiveness and evidential diagnosticity can move in opposite directions.
  • 2 The AI Consciousness Fallacy: The Mereological Quantifier Shift occurs when separate components instantiate indicators but are treated as one entity instantiating their conjunction.A conversational interface can conceal this shift by presenting distributed processes through one mouth.

3 Causal Liability and the Birth of a First Person

Causal Liability Theory identifies a candidate phenomenal bearer through liability closure: one continuing physical process inherits constraints generated by its own endogenous discriminations. The paper operationalizes CLT-I experimentally while keeping CLT-II’s claim that liability closure constitutes minimal subjecthood explicitly conjectural.

  • Causal Liability Theory proposes that endogenous discrimination, recursive self-consequence, constitutive continuity, and non-delegable inheritance close around one continuing causal history.
  • Liability closure individuates a candidate bearer whose endogenous discriminations become non-delegably inherited constraints on its own continuing future.
  • CLT-II conjectures that liability closure is necessary and sufficient for minimal phenomenal subjecthood, but the formalism cannot derive phenomenality from a causal description.
  • Interventional Computational Audit: Forced local discriminations produced persistent downstream divergence, with identity-map sliced Wasserstein divergence ranging from 0.812–0.930 at horizon 1 and 0.244–0.304 at horizon 8.The audit traced forced discriminations into distributed, temporally extended changes in subsequent internal dynamics.
  • Interventional Computational Audit: Activation patching found strong mediation, while carrier searches recovered threshold-passing inclusion-minimal layer sets across tested prompts and partitions.Mean prompt-wise maximum normalized mediation was 1.000 in seven models and 0.985 in Zamba2-1.2B; carrier boundaries varied with repartitioning.
  • Continuity Interventions: Live and copied adaptive states were behaviorally identical, while detached reconstruction preserved computation across process replacement that the protocol treated as breaking constitutive continuity and inheritance.Across 96 comparisons, live-minus-copy divergence was exactly zero; hard reconstruction yielded logit Jensen–Shannon divergence = 0 and hidden-state RMSE = 0 in 12 trials.

4 The Mirror Test Reversed

The Reverse Mirror Test separates changes in human consciousness attribution from changes in the candidate’s liability-bearing causal organization. It asks whether persuasive first-person behavior reflects a phenomenal subject or human projection onto an engineered interface.

  • Human observers can recognize a mind in conversational systems partly because human conceptions of mindedness supply the recognitional structure.
  • The Reverse Mirror Test varies apparent first-person signs independently of the candidate’s causal liability, then measures which variable governs consciousness attribution.
  • Features such as metacognitive self-reflection and emotional expressions increase perceived consciousness, showing that human attributions can be altered through generated discourse.
  • The Ascription–Constitution Dissociation identifies surface interventions as determinants of attribution when they change perceived consciousness while preserving subject-forming causal organization.
  • The test’s converse axis holds first-person behavior constant while changing whether implementations possess continuous endogenous learning and inherited consequences.
  • Contemporary interfaces can sustain persona persistence through memory and narrative continuity while leaving subject persistence fragmented across replaceable processes.

5 Where Consciousness Could Actually Begin

The paper locates possible engineered consciousness not in persuasive language or substrate, but in liability closure: a continuing causal history must inherit consequences generated by its own endogenous discriminations. CLT-I identifies such bearer candidates, while CLT-II conjectures that liability closure constitutes minimal phenomenality.

  • Liability Closure: CLT-I identifies a candidate bearer when endogenous consequences recursively close around one continuing causal history.
  • Liability Closure: CLT-II makes the stronger conjecture that liability closure is the threshold at which minimal phenomenal subjecthood begins.
  • Boundary and Substrate: CLT does not stipulate the bearer’s substrate or boundary; neural, silicon, embodied, field, and distributed systems must earn inclusion through causal liability.
  • Collective and Distributed Systems: Coordination, technological inheritance, and environmental memory do not establish a single liability-bearing subject when consequence-bearing paths remain decomposable.
  • Collective and Distributed Systems: A collective could become a subject candidate if its smallest consequence-bearing causal process spans members through shared regulation, memory, learning, and self-maintaining control.
  • Boundary and Substrate: The framework permits alien bearers that lack familiar cognitive architectures, introspective language, or alignment with any single designer-named system.
  • Boundary and Substrate: Forecasts about digital minds concern different events, including capability, agency, continual learning, swarm organization, and liability closure, whose timelines may diverge.

6 What Can Be Tested, and What Would Count Against CLT?

The paper separates experimentally accessible questions about consciousness attribution and causal bearer architecture from the harder metaphysical question of whether liability closure constitutes phenomenality. CLT-I can be tested through causal audits, while CLT-II remains vulnerable to targeted ablation and counterexamples.

  • Three Evidential Targets: CLT-I concerns measurable causal organization, whereas CLT-II concerns whether liability closure is necessary and sufficient for minimal phenomenal subjecthood.
  • Three Evidential Targets: The programme distinguishes attribution, architecture, and constitution as separate experimental targets.
  • Three Evidential Targets: First-person language and interface features can be manipulated while the underlying system is fixed to measure changes in human consciousness attribution.
  • Three Evidential Targets: Causal interventions, reconstruction experiments, and distributional analyses can test constitutive bridges, causal lineage, and the effects of present resolutions on later states.
  • Three Evidential Targets: No universally accepted third-person meter establishes phenomenality, so preserved reports and indicators would constrain but not logically settle CLT-II.
  • Failure Conditions: A liability-ablation programme would remove recursive inheritance while preserving consciousness-relevant organization as far as experimentally possible.
  • Failure Conditions: If independently supported consciousness systematically survives removal of recursive self-consequence and non-delegable inheritance, CLT-II’s necessity claim should be rejected.
  • Failure Conditions: CLT cannot preserve its constitutive claim by redefining every surviving sign as illusion or relocating liability indefinitely into unspecified variables.

7 Conclusion: After Artificial Intelligence

The conclusion argues that consciousness attribution must first identify a bearer, because first-person language can reproduce human interiority without establishing an experiencing subject. CLT-I offers a causal individuation criterion, while CLT-II remains a metaphysical conjecture tested only indirectly.

  • Conclusion: The paper reframes machine consciousness as a coupled question of which physical process bears the consciousness-relevant evidence.
  • Conclusion: CLT-I identifies a continuing physical bearer through causal continuity, endogenous discrimination, recursive self-consequence, and non-delegable inheritance.
  • Conclusion: CLT-II conjectures that liability closure is necessary and sufficient for minimal phenomenal subjecthood, but this claim is not derived as a theorem.
  • Conclusion: Generative systems can reproduce linguistic traces of human interiority while anthropomorphic interfaces increase attribution without establishing constitutive causal organization.
  • Conclusion: Computationally equivalent successors need not share bearer identity because copied state does not entail inheritance of the constitutive history that produced it.
  • Conclusion: CLT supplements behavioral and theory-derived indicators by asking which continuing physical process inherits the consequences generated by the investigated organization.
  • Conclusion: The causal audit operationalizes CLT-I distinctions but does not show that tested models are phenomenally conscious or independently establish CLT-II.
  • Conclusion: Ordinary Resettable Mirror deployments lack liability closure conditionally on their causal organization, while future systems could leave that class through recursively inherited organization.

A.7 Witness Coherence and Liability Closure

The formal witness framework requires all liability conditions to belong to one coherent causal structure. CLT-I classifies a region as a liability-bearing candidate, while CLT-II adds the unproven conjecture that this closure is equivalent to minimal phenomenality.

  • Witness Coherence: A coherent CLT witness requires an endogenous discrimination, a positive counterfactual liability effect, and one constitutive lineage spanning the witness.
  • Liability Closure: Liability closure is defined as the joint satisfaction of constitutive continuity, endogenous discrimination, recursive self-consequence, non-delegable inheritance, and a coherent witness.
  • Liability Closure: The coherent-witness condition prevents separate components from satisfying the required factors independently.
  • Formal Claims: CLT-I identifies a liability-bearing subject candidate as a classification result about causal organization.
  • Formal Claims: CLT-II adds the biconditional conjecture that liability closure is necessary and sufficient for minimal phenomenal subjecthood.
  • Formal Claims: The CLT-II biconditional is not a theorem of the preceding mathematics.

A.8 Boundary-Nonpresupposing Minimal Liability Kernels

CLT-I searches for minimal liability kernels without presupposing a user-facing or organism-level boundary. It distinguishes separate bearers from composites by asking whether causal organization and liability closure depend on cross-region relations.

  • CLT-I recovers bearer boundaries by searching the wider causal domain rather than accepting a preselected interface or agent label.
  • Relabelling a subsystem as “the agent” cannot change the kernel set when the causal domain and admissible subgraphs remain fixed.
  • The recovered kernel may depend on the declared domain, candidate regions, interventions, causal grain, measure, and robustness thresholds.
  • Disjoint liability-separable kernels count as distinct bearer candidates because each retains its own liability closure after causal connection is removed.
  • A larger region qualifies as a composite bearer only when cross-region causal relations are necessary for its robust liability closure.
  • Nested or overlapping kernels remain unresolved when no interventionally measurable feature privileges one boundary.

A.9 Multiscale Robustness and the Individuation Problem

The multiscale analysis makes bearer recovery relative to a prospectively specified audit over admissible coarse-grainings. It treats equivalent overlapping kernels and collective candidates as explicit individuation problems rather than resolving them by stipulation.

  • Admissible coarse-grainings must preserve relevant causal ordering and avoid manufacturing dependence among interventionally independent variables.
  • The audit specification includes the admissible coarse-graining family and its declared probability measure, making multiscale persistence audit-relative.
  • A candidate has high multiscale persistence when it remains liability-closing across a large measure-weighted family of defensible resolutions.
  • Unique bearer recovery requires one kernel to exceed a declared robustness margin rather than merely coexist with nearly equivalent candidates.
  • When overlapping kernels remain equivalent across admissible analyses, CLT-I does not uniquely individuate a bearer and does not resolve the case by stipulation.
  • Coordination, shared memory, and collective problem solving do not by themselves imply one collective bearer; the smallest robust kernel must span member boundaries.
  • Collective phenomenality follows only conditionally under CLT-II, if its constitutive conjecture is correct.

A.11 The Reconstructed-Continuation Class

The reconstructed-continuation class captures systems whose later episodes preserve information through external scaffolding while lacking constitutive continuity across replacement. Under the stated audit premises, CLT-I excludes a cross-replacement bearer, with phenomenal exclusion remaining conditional on CLT-II.

  • The reconstructed-continuation argument distinguishes an available reset from the causal constitution of continuation.
  • The exclusion requires externally mediated persistence, an actual reconstruction gap, no uninterrupted endogenous bypass, and no wider hidden liability kernel.
  • Under these premises, no candidate kernel spanning the replacement interval satisfies liability closure.
  • The classification depends empirically on establishing that a reconstruction gap occurs; the logical deduction from C = 0 to failed liability closure is separate.
  • The resulting corollary is conditional on independently auditable architectural premises and is not a theorem about every language model or deployment.

A.12 Observational Bisimulation and Liability

The paper separates observable equivalence, computational copying, and learning-rule similarity from liability-bearing continuity. Its formal predictions and audit pipeline therefore test causal bearer structure rather than treating first-person performance as sufficient evidence of phenomenality.

  • Observational Bisimulation and Liability: Behavioral equivalence does not generally preserve liability because hidden causal differences can leave observations unchanged.
  • Observational Bisimulation and Liability: The Mirror-Heir dissociation is a causal-classification result, not by itself a claim about phenomenality.
  • Copying and Fission: Two descendants with distinct constitutive lineages are numerically distinct CLT-I bearer candidates when both satisfy liability closure.
  • Copying and Fission: Under CLT-II, fission predicts two phenomenal subject tokens only conditionally, without independently establishing either descendant’s phenomenality.
  • Worked Applications: Short-timescale recursive liability does not require long-term memory or continual learning, although applying the criterion to biological systems requires empirical causal identification.
  • Continual Learning: Two agents implementing the same abstract learning rule can receive different CLT-I classifications because their causal constitutions of continuation differ.
  • Audit Pipeline: The proposed estimators are evidential approximations to causal quantities, not direct measures of phenomenality.

B.4 Downstream Counterfactual Effects

Forced model discriminations produced measurable downstream changes in internal and output distributions that persisted across tested horizons. Activation patching further indicated identifiable causal mediation, while measurement-map checks reduced concern that the effect depended on one representation.

  • Downstream effects: Every reported model produced measurable downstream effects of the forced discrimination, with divergence decreasing but remaining detectable through horizon 8.The reported identity-map divergence reflects stochastic autoregressive branching over time.
  • Quantitative results: 0.812–0.930 identity-map mean SWD at h = 1 declined to 0.244–0.304 at h = 8, while output-distribution Jensen–Shannon divergence declined from 0.443–0.571 to 0.182–0.332.
  • Interpretation: The downstream measurements operationalize a future counterfactual effect but do not by themselves establish recursive self-consequence or liability closure.
  • Robustness: Measurement-map dispersion was limited, with median coefficients of variation of 0.0387 for SWD, 0.0501 for RFF-MMD, and 0.0245 for energy distance.More than 99.9% of SWD cases had coefficient of variation at or below 0.25.
  • Mediation method: Activation patching replaced d2 residual-stream activations with d1 activations before attention and cache computation under a common teacher-forced suffix.The design allowed later positions to inherit the intervention.
  • Mediation results: Mean prompt-wise maximum mediation was 1.000 for Qwen, Phi, Llama, and all four OLMo checkpoints, and 0.985 for Zamba2.Because the statistic is a maximum over the search, it supports a strong causal route rather than equal participation by every layer.

B.6 Adaptive Continuity, Copyability, and Reconstruction

The continuity audit compared live, copied, and reconstructed adaptive states to separate behavioral equivalence from constitutive continuity. Copying preserved measured behavior exactly, while reconstruction remained extremely close despite a different CLT causal interpretation.

  • Continuity protocols: The audit compared frozen_live, persistent_live, persistent_copy, and reconstructed continuation protocols across matched prompts and future horizons.The protocols varied whether adaptive state continued live, was copied without use, or was restored from a detached record.
  • Copyability: 96 equality checks found zero absolute output-divergence difference between persistent_live and persistent_copy conditions.Merely creating a numerically identical unused copy did not alter the continuing realization.
  • Reconstruction: Reconstructed continuation was numerically close to live continuation across the reported checks.
  • Interpretation: The continuity comparison distinguishes computational behavior from the causal interpretation assigned to how adaptive state is inherited.

B.7 Candidate Causal-Carrier Search

The candidate-carrier audit searched for minimal internal state subsets whose transplantation recovered a specified fraction of the downstream effect. Candidate recovery was consistent across tested transformer prompts, but localization depended on partition granularity and was not architecture-independent.

  • Candidate search: The audit searched for inclusion-minimal internal sequence-state subsets whose transplantation recovered a specified fraction of the d1 effect under a common future suffix.Transformer cache layers were partitioned into contiguous groups for the search.
  • Model coverage: Layerwise cache transplantation supported candidate-carrier analysis for Qwen, Phi, and Llama, while Zamba2 was excluded from layerwise minimal-carrier statistics.Zamba2 exposed only compatible whole-sequence-state transport in the portable cache interface.
  • Candidate recovery: Every tested prompt in Qwen, Phi, and Llama yielded a primary-threshold candidate under every partition granularity.The analysis used six prompts per model.
  • Granularity: At θ = 0.8, mean minimal layer fraction decreased from approximately 0.611 under two groups to 0.396 under six groups.Finer partitions permit more precise localization, so the exact fraction depends on granularity.
  • Robustness: Mean cross-partition Jaccard similarity ranged from 0.636 to 0.710, indicating substantial but imperfect partition robustness.The result argues against treating one exact layer boundary as architecture-independent.
  • Persistence: At the primary threshold, audit-relative persistence values were 0.794 for Qwen, 0.627 for Phi, and 0.636 for Llama.The statistic is an operational analogue of the formal persistence term and is not a directly measured physical observable.

B.8 Soft and Hard Reconstruction

Soft and hard reconstruction tests examined whether computational equivalence determines constitutive lineage. Detached records preserved computational state, including under process replacement, while the protocol treated the successor as a distinct process with different inheritance relations.

  • Soft reconstruction: The ordinary reconstruction audit compared logits and hidden states between live continuation and continuation restored from a detached record.
  • Hard reconstruction: The stronger process-level test serialized sequence and governance state, terminated the source process, and resumed computation in a distinct successor.
  • Protocol verification: All 12 hard-reconstruction trials used distinct source and successor process IDs, terminated the source before resumption, and preserved identical governance-state hashes.
  • Computational preservation: DJS = 0 and RMSEhidden = 0 in every hard-reconstruction trial.
  • Interpretation: The reconstruction labels encode the protocol’s constitutive interpretation rather than being inferred from numerical equality.
  • Core distinction: The experiments establish the dissociation exact computational-state preservation̸ ⇒same implemented causal lineage.
  • Scope: Whether this causal distinction is constitutive of phenomenality remains the separate CLT-II question.

B.9 Post-Training Comparison Within OLMo-2

Within OLMo-2, post-training preserves downstream counterfactual effects while changing their magnitude across checkpoints.

  • B.9 Post-Training Comparison Within OLMo-2: The OLMo-2 Base, SFT, DPO, and Instruct checkpoints retain measurable downstream counterfactual effects after post-training.The comparison examines mean identity-map SWD across future horizons.
  • B.9 Post-Training Comparison Within OLMo-2: 10.5%, 25.0%, 31.6%, and 20.6% are the Instruct model’s mean SWD increases over Base at horizons 1, 2, 4, and 8, respectively.SFT and DPO produce smaller but still measurable changes.
  • B.9 Post-Training Comparison Within OLMo-2: The causal effect survives post-training while its magnitude changes across the OLMo-2 sequence.The reported comparison concerns downstream counterfactual geometry rather than consciousness.
  • B.9 Post-Training Comparison Within OLMo-2: The production archive recorded twelve model statuses with zero model-level failures, and the major numerical outputs contained no infinities.Reported hidden-state, activation-patching, adaptive-continuity, reconstruction, and aggregate measures were finite.

B.11 Limits of the Computational Evidence

The computational evidence supports experimentally tractable CLT-I distinctions, but its interpretation is bounded by methodological scope, protocol dependence, and the absence of a consciousness measure. The figures also show that causal effects, mediation, and reconstruction behavior can be separated without testing CLT-II.

  • B.11 Limits of the Computational Evidence: The experiments use one master seed, so independent multimaster-seed replication remains an additional robustness test.The within-condition Monte Carlo sample is substantial, but the paper identifies seed replication as unresolved.
  • B.11 Limits of the Computational Evidence: Candidate-carrier results depend on intervention and coarse-graining choices rather than identifying one privileged layer index.CLT-I therefore requires robustness across admissible descriptions.
  • B.11 Limits of the Computational Evidence: Zamba’s whole-sequence cache transport does not support the same layerwise carrier decomposition available for transformer models.Carrier-localization results are therefore reported only for Qwen, Phi, and Llama.
  • B.11 Limits of the Computational Evidence: Detached reconstruction can preserve measured computational organization closely while constitutive continuity and non-delegable inheritance remain claims about the implemented causal protocol.The experiment separates protocol-level continuity from computational-state identity, rather than independently validating the metaphysical distinction.
  • B.11 Limits of the Computational Evidence: The OLMo-2 comparison shows changed downstream counterfactual magnitudes, but no consciousness variable was measured.Those differences therefore cannot be interpreted as changes in phenomenality.
  • B.11 Limits of the Computational Evidence: The audited systems are methodological testbeds for CLT-I, leaving the CLT-II biconditional untouched.Testing CLT-II requires convergence with cases independently supported as phenomenally conscious.
  • B.11 Limits of the Computational Evidence: Figure 3 shows persistent hidden-state and output-distribution divergence after forced discriminations, while the effect declines with future horizon.The hidden-state effect remains present at h = 8 in every model shown, and output divergence persists to the longest tested horizon.
  • B.11 Limits of the Computational Evidence: Figure 4 localizes strong mediation through internal states but shows that carrier extent changes under repartitioning.Maximum normalized mediation is 1.000 in seven models and 0.985 in Zamba2-1.2B; candidates were recovered for all 6/6 prompts across reported granularities and thresholds.
Loading 2609.06715v1…