Source-linked AI summary

Deposon: An Auditable, Conservation-Guaranteed, Game-Theoretically Tested Scattering Layer over LLM Reasoning Paths

Qihao Yuan

arXiv:2609.09001v1cs.AIcs.LG

TL;DR

Multi-step LLM reasoning lacks a machine-recheckable record for discarded paths. The paper proposes Deposon, which binds concept-graph nodes to scattering states and audits transmission, reflection, and dissipation. Conservation is guaranteed by construction, while benchmark attribution ties the layer to a trivial rule filter and the game-theoretic equivalence claims are falsified under strengthened tests.

  • Problem

    Multi-step reasoning systems lack a recheckable account of why candidate paths were discarded, including the discarded branch’s budget and decision rationale.

  • Method

    Deposon binds LLM-generated concept-graph nodes to parameterized physical states and uses three-channel scattering with an auditable conservation identity.

  • Results

    The layer guarantees T+R+A=1 with a 2.2×10−16 audit residual, while real-benchmark performance is tied with a trivial rule filter and formal dynamical-equivalence propositions are falsified.

  • Takeaways & Limitations

    The supported differential value is machine verifiability, while the dynamical game-theoretic interpretation retains only consistency-level evidence.

  • Takeaways & Limitations

    The strengthened kill protocol falsifies P1a, P1b, and T-P1c, so the dynamics cannot be claimed to implement best response generally.

Abstract

from arXiv · show

Multi-step LLM reasoning lacks a machine-recheckable ledger: discarded reasoning paths leave no auditable record. We propose the Deposon scattering layer, which binds each node of an LLM-generated concept-decomposition graph to a two-parameter Deposon state; paths undergo three-channel scattering -- transmission, reflection, irreversible dissipation -- obeying T+R+A=1 for arbitrary parameters, with a maximum per-path energy-audit deviation of 2.2E-16 (machine epsilon). We report all three evidence tiers honestly. On synthetic trap benchmarks the path-filtering gain is closed (pre-registered): unified reaches 100% versus a decoy-capture baseline at 7%/10%. On real benchmarks the layer is indistinguishable from a trivial six-keyword rule filter (GSM8K 0.87 >= 0.85, McNemar p=0.5; StrategyQA 0.899 = 0.899); no difference is detected here, so we sharpen the claim to "the differential value lies solely in machine verifiability." Fusion yields a second negative result: convex combinations with a semantic prior never improve (physics 0.484 -> 0.452), and the apparent lambda=2 gain is an anti-field artifact; any fusion gain must be nonlinear. Modeling the reverse dynamics as a potential game on the graph, we evidence an auditable scalar's monotonicity and near-gradientness and quantify the empirical coordination ratio (ECR). The three formalized dynamical-equivalence propositions (P1a/P1b/T-P1c) are falsified under the pre-registered kill protocol, and the potential-game claim is downgraded to approximate (cyclic-graph median residual 0.669): only consistency-level evidence survives at the dynamical level. Code: github.com/zeroandcat/Deposon.

1 Introduction

The paper introduces Deposon as a physical, auditable layer for recording why reasoning paths are discarded, while separating conservation guarantees from weaker benchmark claims and reporting negative results explicitly.

  • Auditable representation: T+R+A=1 holds for arbitrary parameters, providing a conserved accounting identity for each scattering path.The construction assigns transmission, reflection, and irreversible dissipation as the three channels.
  • Auditable representation: 2.2×10−16 is the maximum per-path energy-audit residual, at double-precision machine-epsilon scale.The identity is mechanically recheckable without benchmark accuracy.
  • Auditable representation: Node-by-node elimination decisions make discarded reasoning paths attributable rather than leaving an unrecheckable pruning record.The motivation is the absence of conserved information about discarded branches, including their budget and disposal rationale.
  • Evidence boundaries: Synthetic path-filtering gains are retained as a closed, pre-registered evidence tier, while real-benchmark ties with a rule filter constrain the accuracy claim.The paper reports these evidence tiers separately rather than blending them.
  • Evidence boundaries: Convex-combination fusion never exceeds either single arm on measured λ settings, and the apparent λ=2 gain is identified as an anti-field artifact.The paper concludes that any fusion gain would require a nonlinear mechanism.
  • Game-theoretic formulation: The reverse dynamics is presented with an auditable potential and empirical coordination ratio, alongside a temperature frontier defining the audit boundary.This is offered as game-theoretic evidence rather than as a benchmark-accuracy claim.

2 Auditable Representation and Conservation Guarantees

The Deposon layer represents LLM reasoning as concept-graph paths with three-channel scattering, giving each elimination a conservation-based, machine-recheckable record. Its benchmark filtering gains are conditional, while real-task comparisons and fusion experiments relocate the contribution toward verifiability rather than accuracy.

  • Representation: Each node receives semantic coupling parameters, and BFS generates candidate paths through a directed concept graph containing operations, traps, answers, and general concepts.Trap, operation, and other nodes receive distinct parameter settings, with detuning modulating effective coupling.
  • Conservation identity: T+R+A=1 by construction, assigning every path’s energy to transmission, reflection, or irreversible dissipation.Transmitted energy recurses along paths, while reflection and dissipation accumulate as auditable quantities.
  • Per-path audit: 2.220446049250313×10−16 is the maximum observed deviation of |T + R + A −1| across variants, synthetic problems, and tested paths.The residual is at double-precision machine-epsilon scale and below the 10−6 implementation tolerance.
  • Synthetic filtering: 100% versus 7%/10% is the unified-versus-field-free greedy result on the two 100-problem synthetic benchmarks.The baseline deliberately overweights decoy edges; randomized labels reduce trap accuracy to 17.2%±6.4%, while uniform parameters yield 10%.
  • Real-benchmark boundary: 0.87 ≥ 0.85 on GSM8K and 0.899 = 0.899 on StrategyQA show that scattering is indistinguishable from the six-keyword rule filter on these real benchmarks.The paper therefore locates the differential value in machine verifiability rather than filtering performance.
  • Fusion: The semantic-prior fusion scan and the λ=2 artifact provide negative evidence against the tested convex-combination fusion hypothesis.The apparent λ=2 gain was reproduced as an anti-field artifact under the null ablation.

3 The Game-Theoretic Formulation: Potential, the Empirical Coordination Ratio (ECR), and the Audit Boundary

The section defines a graph-based potential-game formulation for reverse dynamics, then evaluates scalar monotonicity, near-gradientness, coordination, and validity boundaries under pre-registered protocols. Monotonicity and near-gradientness close under restricted protocols, while formal dynamical-equivalence claims are falsified and the potential-game interpretation is downgraded for cyclic support graphs.

  • 3.1 The model: a potential game on a finite graph: Each leave-one-out prediction edge is a player, candidate target nodes are strategies, field scores are utilities, and Φ=−E is the potential-function candidate.The reverse process ranks candidate targets after simplex natural-gradient updates on masked graph rows.
  • 3.2 Decisive evidence: monotonicity, near-gradientness, and quantification of the auditable scalar (closed (pre-registered) tier): GT-5b reports 100% potential monotonicity on 22/22 graphs, exceeding the pre-registered 80% line under the single-edge leave-out mask protocol.The narrowed claim remains separate from the failed endpoint condition and is not extrapolated to general mask structures.
  • 3.2 Decisive evidence: monotonicity, near-gradientness, and quantification of the auditable scalar (closed (pre-registered) tier): GT-6 reports a median non-potential residual of 1.594×10−29, below the 0.10 pre-registered line, while three cyclic graphs exceed the line.The residual measures how closely the per-task edge-utility vector lies in the gradient subspace; its absolute magnitude is subject to numerical-precision qualification.
  • 3.3 Formalization kill-tests: all three tiers of dynamical equivalence falsified (closed (pre-registered), systematically sampled graph families × exhaustive-state protocol): The P1a, P1b, and P1c formalized dynamical-equivalence claims are all falsified under systematically sampled small graphs and exhaustive-state tests.The better-response test records 10/338 violations, including negative potential changes beyond the peak.
  • 3.3 Formalization kill-tests: all three tiers of dynamical equivalence falsified (closed (pre-registered), systematically sampled graph families × exhaustive-state protocol): Potential completeness is downgraded to an approximate potential game: cyclic-support graphs have median residual r=0.669, whereas acyclic support graphs have r≤5.9×10−16.The downgrade is triggered because 0.924 of cyclic-support cases have r>0.30, above the pre-registered 1/3 line.
  • 3.3 Formalization kill-tests: all three tiers of dynamical equivalence falsified (closed (pre-registered), systematically sampled graph families × exhaustive-state protocol): Dissipation is not merely a payoff reparameterization: changing it moves equilibrium positions, while g_a>0 guarantees convergence when ρ(G)<1.The reported maximum fixed-point difference across g_a settings is 0.8442, and the no-dissipation cyclic case diverges when ρ(G)=1.
  • 3.4 Demarcation and boundary regularities: where the auditable advantage holds (directional evidence, not upgraded): The headline H-A1 advantage over random is killed by four reversed graphs, while H-A2 over degree survives with 19+/1−/2 ties and p=4.0×10−5.The surviving evidence contracts the validity domain to robust structural-baseline advantage plus local advantage on high-hub graphs.

4 Honest Boundaries

The paper explicitly limits its claims: formal dynamical equivalences are refuted, potential-game evidence remains consistency-level, and several evaluation settings constrain generality and interpretation.

  • Formal dynamical equivalence: All three formalized dynamical-equivalence propositions are refuted, so the game-theoretic line cannot rely on dynamics implementing best response.The kill protocol reports 0.8569 deviation for P1a, minimum cosine −1.0 for P1b, and minimum cosine −1.0 for T-P1c.
  • Consistency evidence: Positive trajectories support a potential-game reading as consistency evidence, not as a proved potential-function theorem.The paper states that this evidence should not be upgraded to a theorem.
  • Data and generator scope: Family-L graphs come from a single vendor’s LLM, leaving model-family preferences, shared-corpus contamination, and human-annotated-graph generalization unresolved.Testing non-Chinese model families and human-annotated graphs remains open.
  • Sample-size boundary: Question-bank results use n=40 per cell, making ±1 problem equivalent to ±2.5 percentage points and limiting effect-size interpretation.The paper characterizes this track as a small-sample wide interval.
  • Exploratory analyses: Exploratory demarcation regularities are weakened by feature–design circularity, low power on the hub axis, and cross-domain heterogeneity.The real_semantics axis holds in 3/4 domains, while GT-8c is mixed and programming_concepts falls below the line.
  • Dissipation boundary: Dissipation remains a motivational mechanism rather than a demonstrated task-level benefit, despite premise evidence that g_a>0 avoids divergence in 58/338 tasks.Task-level rent is negative on every measured task, with ties on synthetic benchmarks and the reported real-benchmark outcomes.

5 Conclusion

The conclusion frames the contribution as run-time, per-instance invariants supported by conservation and auditability, while sharply delimiting accuracy and dynamical claims.

  • Scope: The paper targets run-time, per-instance invariants through two equally weighted result lines: auditable conservation and game-theoretic analysis.The conclusion places these lines alongside both positive evidence and explicit negative results.
  • Auditable representation: Conservation is guaranteed by construction: T+R+A=1, the audit residual is 2.2×10−16, and elimination decisions are attributable node by node.These are presented as the auditable representation’s core properties.
  • Game-theoretic evidence: The game-theoretic evidence includes GT-5b monotonicity in 22/22 cases, GT-6 median residual 1.594×10−29, and ECR=1.333.The conclusion also states that the formal dynamics=best-response claim was killed and the potential explanation’s domain was narrowed.
  • Future work: Future work includes trainable KGE baselines, expanded exploratory regression, stronger hub-axis sampling, and pre-registered question-bank controls.The killed dynamical-equivalence claims are explicitly excluded from future-work items.

6 Data availability and evidence-strength conventions

The paper ties every experimental number and statistical verdict to frozen, mechanically evaluated artifacts, while labeling evidence strength in three tiers.

  • Traceability: Every experimental number is traceable to a field in a frozen JSON enumerated in Appendix A.Statistical verdicts are produced by pre-registered decision rules frozen before the run.

7 Artifact availability

The complete reproducibility artifact is distributed in the repository, including frozen data, verdict logic, and tests without proprietary dependencies.

  • Artifact contents: The repository contains code, frozen JSONs, verdict functions, and tests, with SHA-256 anchors and bundled scripts for mechanical reruns.The paper states that reproducing its numbers requires no proprietary dependencies.

8 AI use disclosure

The paper discloses that the corresponding author defined the Deposon framework and research direction, while KIMI-K3 assisted with core development and drafting under direct instruction.

  • KIMI-K3 assisted with the core algorithm, experimental pipeline, data-analysis scripts, and first manuscript draft.
  • The author reviewed every section, verified numerical claims against frozen repository JSONs, and accepts responsibility for the content.

A Number Traceability Table

The traceability table links reported results to frozen JSON fields and records outcomes across conservation, benchmark, fusion, statistical, and game-theoretic evaluations.

  • Audit and benchmarks: 2.220446049250313×10^-16 is the conservation-audit residual, with the audit passing under a 1e-6 tolerance.
  • Audit and benchmarks: The six-keyword rule filter scores 0.87 versus unified 0.85 on GSM8K and ties unified at 0.899 on StrategyQA.
  • Audit and benchmarks: 100% unified accuracy versus 7%/10% for the decoy-capture baseline is reported on the synthetic benchmarks.
  • Audit and benchmarks: GSM8K unified performance is 85.0% versus 97.0% for CoT, while StrategyQA reports 89.9% for unified.
  • Fusion: The four-λ scan is identical across settings, with named Hits@3=0.294 and any_lambda_pass=false.
  • Fusion: The apparent λ=2 normalized-variant score of 0.471 is contradicted by a null ablation matching the real prior at 0.1176.
  • Game-theoretic tests: GT-4 reports median ECR 1.333 across 17 finite-valued graphs, with three infinite cases excluded from the median.
  • Game-theoretic tests: GT-5b finds Φ monotone on 22/22 graphs, while GT-6 reports a median residual of 1.594e-29.
Loading 2609.09001v1…