Source-linked AI summary

Can a Dynamic Internal Field Govern a Transformer's Cognition? Certifiability, not Superiority, in Homeostatic Compute Control

Francisco M. Arrabal-Campos, Ignacio Fernandez, Francisco G. Montoya, Alfredo Alcayde

arXiv:2608.24319v1cs.AIcs.LGeess.SY

TL;DR

The paper asks whether a dynamic internal field can govern an adaptive transformer’s computation without performing cognition. It implements a PDE-governed homeostatic field, analyzes its integrator stability, and evaluates substance, structure, and certification. The results support a certifiable compute governor, but not an accuracy enhancer, while the full closed-loop certificate remains outside scope.

  • Problem

    The paper asks whether a low-dimensional, physically explicit internal field can govern how much an adaptive transformer computes, when it stops, and which modules it activates.

  • Method

    The authors evolve a PDE-governed field on the module graph alongside an adaptive-depth reasoner, and analyze its integrator with a discrete Schur–Cohn criterion.

  • Results

    The physical regime is irrelevant for accuracy, second-order structure matters only partly after equalized caps, and the field’s distinctive advantage is certifiable integrator stability rather than capability.

  • Takeaways & Limitations

    A dynamic internal field is a viable compute governor that modulates cognition without enhancing it, while matched learned recurrent governors can achieve comparable or better performance.

  • Takeaways & Limitations

    The paper does not prove closed-loop stability for the field, its host, and the interoception/modulation path considered as one system.

Abstract

from arXiv · show

An intelligent system does not merely reason: it governs its own reasoning - how much to compute, when to stop, which module to activate. Can that role be played by a dynamic internal field - a low-dimensional homeostatic state with explicit physics and certified stability - that modulates cognition without performing it? Ours is a field on the module graph governed by a family of PDEs on the graph Laplacian, advancing with an adaptive-depth reasoner. We certify the stability of the integrator of the whole family - an integrator certificate, not a closed-loop one. New, and proved here: a discrete Schur-Cohn criterion for Verlet with velocity coupling, necessary and sufficient per latent root, with no commutation hypothesis. The answer is threefold: substance no, structure only in part, certifiability yes. The type of the field's physics is irrelevant for accuracy: wave, diffusion, gated mixtures and a 2D Navier-Stokes substrate tie. A twenty-seed preregistered deconfounding campaign bounds the structural claim: at equalized caps the second-order effect is strong in one family (+0.087 [+0.042, +0.132], t=4.0) but is not detected in the other (+0.014 [-0.013, +0.040], n.s.), so part of the original contrast was capacity, not order; and a matched-interface GRU is indistinguishable in the first and nominally exceeds the field in the second (-0.035 [-0.067, -0.002]). What distinguishes the field is not capability but that its one-step operator admits an exact runtime stability check - a difference of kind, not of existence: learned recurrences carry certificates too, sufficient and conservative ones. A kill-gate with a positive control finds no evidence for the field as evidence accumulator (Delta AUC +0.0007 [-0.0065, +0.0079] vs a 0.03 threshold). A dynamic internal field is a viable, certifiable compute governor, but not an enhancer of cognition: it modulates, it does not think.

1 Introduction

The paper tests whether a physically governed internal field can regulate adaptive computation without performing reasoning. It finds no accuracy advantage from physical substance, partial evidence for second-order structure, and a stability-certification advantage.

  • Motivation: Adaptive computation is motivated by the need to match reasoning depth and halting decisions to uneven input difficulty.Existing approaches learn compute-allocation policies end-to-end, which can overfit the training distribution.
  • Approach: The Homeostatic Background Processor is a low-dimensional PDE-governed field on the module graph that modulates a recurrent reasoner through interoception and top-down signals.Its physics family includes diffusion, wave, advection, dispersion, reaction, and KdV-type dynamics.
  • Conclusion: A dynamic internal field can govern computation with certifiable integrator stability, but its distinctive contribution is not enhanced cognition.The authors position it as a metacognitive or autonomic governor rather than a reasoning module.
  • Certification: The paper introduces a necessary-and-sufficient discrete Schur–Cohn criterion for Verlet with velocity coupling, without assuming commutation, and audits exact spectral radius at trained checkpoints.The stability characterization concerns the integrator family rather than the full closed-loop system.
  • Findings: At equalized caps, the second-order effect holds in one generator family but not the other, while a matched-interface GRU achieves comparable or better performance.The introduction reports that the structural effect is confined to one of two generator families once capacity is equalized.
  • Findings: The field’s physical regime does not affect accuracy across the tested mechanisms, including wave, diffusion, non-local coupling, and a 2D Navier–Stokes substrate.The paper attributes the broad mechanism null to temporal coarse-graining of the gradient and, for the flow substrate, incompressibility.

2 Related work

The related work situates the HBP among learned halting, homeostatic modulation, expressive limitations, certified recurrent dynamics, and PDE-based graph networks. Its main distinction is a low-dimensional modulatory field with analyzable dynamics, while its closed-loop certificate remains unproved.

  • Adaptive computation: Adaptive-computation methods decouple compute depth from architectural depth, but their halting policies are generally learned end-to-end.The paper’s reasoner follows the PonderNet-style halting paradigm while anchoring control in a dynamic field.
  • Expressivity and state tracking: The S5 task is NC1-complete, motivating iterative computation and evaluation of out-of-distribution extrapolation rather than fixed-depth shortcuts.The paper also notes related expressive limitations in transformers and state-space models.
  • Memory and modulation: The gating_wm control incorporates working memory, but experiments attribute the in-distribution accuracy gain to recurrence rather than memory and not to the HBP.This separates compute governance from a memory-based explanation of the observed accuracy pattern.
  • Certified learned dynamics: The paper’s certificate is weaker than by-construction guarantees in several learned recurrent models because it uses a differentiable penalty and evaluation-time operator probing.The full closed-loop stability result for the field, host, and interoception/modulation path is explicitly not proved.
  • PDE-governed graph networks: Unlike PDE-based graph networks that transport representations through layers, the HBP transports a low-dimensional modulatory state across reasoner iterations.Its wave–diffusion mixing and odd operators require stability analysis beyond the cited GNN literature.
  • Homeostasis and neuromodulation: The HBP formulates homeostatic modulation as a physically motivated dynamical system whose effect on compute policy is experimentally isolated and whose integrator stability is analyzed.This connects neuromodulation-inspired control with explicit dynamical-system analysis.

3 A family of homeostatic fields

The HBP is a low-dimensional field on the module graph whose PDE dynamics evolve with reasoning and modulate computation through a shared interoceptive interface.

  • Architecture: The field state h spans backbone blocks, the adaptive-depth reasoner, and working memory through graph-Laplacian coupling.Each module supplies bottom-up interoception and receives top-down modulation.
  • PDE family: The family combines first- and second-order dynamics, including wave, diffusion, advection, dispersion, and saturated KdV-type nonlinearities.A convex mixture α couples the wave and diffusive branches, with optional interoception-gated physics.
  • Branch structure: The wave branch uses second-order dynamics with gyroscopic coupling, while the diffusion branch uses first-order dynamics with positional antisymmetric coupling.The paper identifies these placements as the stable choice for each branch.
  • Instantiations: The evaluated members are damped-wave, diffusive-relaxation, wave-plus-dispersion/KdV, and a gated mixture variant.For hbp_kdv, βmax=0.1 and νmax=0.3; at N=6, the system has no solitonic regime.
  • Runtime interface: At each reasoner iteration, the field modulates computation, collects interoception, advances one integration step, and biases halting and memory gates.The reasoner halts PonderNet-style after at most Nmax=24 iterations.

4 Stability analysis of the family

The paper analyzes continuous and discrete stability across the field family, proving branch-specific conditions and identifying where certificates do—and do not—extend to mixtures.

  • Continuous stability: The gyroscopic wave placement is globally asymptotically stable because antisymmetric forces do no work and damping dissipates energy.This is the Kelvin–Tait–Chetaev case for positive stiffness and damping.
  • Continuous stability: The first-order positional placement is unconditionally stable, while circulatory placement can produce flutter beyond an exact isotropic threshold.At βmax=0.1 and ρ(A3)=5.85, the circulatory threshold is violated across almost the entire admissible parameter box.
  • Discrete certificates: Verlet stability holds per latent root if and only if ˜µ^2 < g(2 − g) and q(g^2 + ˜µ^2) < 2g.The criterion is necessary and sufficient without assuming that K, C, and G commute.
  • Discrete certificates: The criterion was checked on 302 400 parameter-grid cells and 400 non-commuting operator draws, with zero discrepancies and maximum residual 6 · 10^-14.These checks validate the latent-polynomial identity against exact roots.
  • Discrete certificates: The IMEX diffusive branch is unconditionally contractive when the antisymmetric operator is included in the implicit kernel, without a commutation hypothesis.With G explicit, the unconditional claim is false.
  • Certificate scope: Stability of the convex mixture does not follow from the branchwise propositions; the gated mixture has neither an analytic nor an exact numerical certificate.A common quadratic Lyapunov search succeeds for the wave branch’s occupied region but fails for the mixture.

5 Experiments

The experiments evaluate OOD compute adaptivity under corrected dynamics, then use equalized caps and a matched-interface GRU to separate order from capacity and implementation. The deconfounded evidence supports a generator-bounded second-order effect, while recurrence—not the field—accounts for accuracy gains.

  • Experimental protocol: BF16 freezing of raw physical parameters was corrected by anchoring them to FP32 before rerunning the full protocol.The bug rounded updates below ULP/2, leaving parameters fixed at initialization.
  • In-distribution accuracy: Recurrence adds +0.071 in adjacent and +0.144 in 5-cycle, while adding working memory is nominally negative and the HBP addition is not significant at n=3.The reported recurrence and working-memory contrasts are vanilla→gating and gating→gating_wm, respectively.
  • Experimental protocol: The protocol tracks pure-OOD compute adaptivity as corr(K, E[niter]) for K ≥14, using paired seeds and preregistered evaluation controls.The task composes K generators of S5 and evaluates beyond the training range, up to K ≤24.
  • Deconfounding order: +0.087 [+0.042, +0.132] in adjacent supports a strong equalized-cap order effect, whereas +0.014 [−0.013, +0.040] in cycle_transp is not detected.The twenty-seed campaign therefore bounds the original contrast as generator-bounded rather than family-general.
  • Deconfounding order: −0.035 [−0.067, −0.002] indicates that the matched-interface GRU nominally exceeds the field in cycle_transp, while the adjacent contrast is statistically indistinguishable.The GRU uses the same interoception, forcing, and modulation heads as the physical integrator.
  • Physics variants: The KdV variant matches the pure wave, with paired contrasts null in both generator families.The reported contrasts are t=1.12, p=0.29 and t=−0.59, p=0.57.

6 A mechanism null: the field’s physics does not move accuracy

Across forced, gated, dual-regime, non-local, and flow-substrate tests, changing the field’s physical regime does not improve accuracy. The proposed explanation is that temporal coarse-graining makes transient differences invisible to the loss, while the evidence-accumulator role is separately tested by a kill-gate.

  • Forced physics: Forced wave and diffusion training produce identical accuracy across four tasks, with |∆| ≤0.012 and both regimes certified stable.The tasks are composition in S5, functional iteration, modular sum with distractors, and recall.
  • Gated physics: Gated physics converges to nearly constant physics, α →0.79 with deviation across inputs ∼10−5, and ablating D or b changes accuracy by < 0.008.A strongly input-sensitive initialization retrains to the same attractor.
  • Dual-regime tasks: Wave and diffusion remain identical within each dual-regime mode, with |∆| ≤0.005, bounding a physics-switching oracle’s possible gap on these tasks.The bound applies to the tested oscillatory-retention and monotone-settling demands and is not claimed to transfer beyond them.
  • Beyond local physics: Non-local coupling does not improve accuracy, and an incompressible flow substrate is a mixer rather than a point-delivering computer.These extensions reinforce the null through mechanisms beyond another type of local physics.
  • Interpretation: Temporal coarse-graining groups ∼5 ticks for ∼19 operations, making equilibrium set-points visible while transient differences of ∼4 ticks are not.Both branches reach forced set-points at equilibrium, whereas physics switching would require tick-aligned oscillatory demands.
  • Evidence accumulation: The kill-gate tests whether a damped second-order field acts as a leaky evidence integrator for noisy per-tick posterior signals.The probe uses a fixed pass threshold of ∆AUC ≥0.03 and frozen solvers from a distinct integration program.

7 Discussion and limitations

The HBP acts as a metacognitive compute governor rather than an accuracy enhancer: structural benefits are partial, capacity-sensitive, and matched by a learned GRU in tested settings. Its main distinctive property is certified stability, while broader architectural and task limitations remain.

  • What the homeostatic field does and does not do: The field is neutral for accuracy but can improve out-of-distribution compute allocation in only one of two generator families.The surviving structural effect concerns compute allocation rather than OOD accuracy.
  • What the homeostatic field does and does not do: The field’s physics does not determine accuracy across the tested configurations, while second-order inertia matters for compute allocation in one family.The reported null spans the instantiated physics variants, whereas the structural effect is bounded to one generator set.
  • What the homeostatic field does and does not do: A matched-interface GRU reaches the same result in one family and nominally exceeds the field where the order effect vanishes.This widens the class of viable governors beyond dynamic fields.
  • Where it fits: the field as governor, not as cortex: The HBP belongs to the metacognitive control layer: it modulates computational resources without replacing the recurrent reasoner or performing reasoning itself.A complete cognitive architecture would additionally require deliberation, memory or a world model, and drives or goals.
  • Limitations and future work: The evidence is limited to the S5 task family, small models, six-node graphs, and a structural effect reproduced by a matched-interface GRU.The deconfounding campaign equalizes coefficient caps but not controller state size, and larger-graph behavior remains open.

8 Conclusion

The paper concludes that a dynamic internal field is a viable, certifiable compute governor but does not enhance cognition. Its physics is irrelevant to tested accuracy, structural benefits are partial, and its distinctive advantage is runtime stability certification rather than capability.

  • Conclusion: A dynamic internal field is a viable and certifiable compute governor, but it modulates computation rather than thinking or improving cognition.The conclusion is explicitly bounded to the governor role within a broader intelligent architecture.
  • Conclusion: The field’s physics is irrelevant for accuracy across the tested configurations, including local, non-local, and 2D flow substrates.The reported null extends across wave, diffusion, gated mixtures, and a 2D Navier–Stokes substrate.
  • Conclusion: The HBP’s out-of-distribution compute-allocation robustness appears in one of two generator families, while a matched-interface GRU reaches the same place and nominally exceeds it in the other.The comparison bounds the structural claim and shows that the field is not uniquely capable as a governor.
  • Conclusion: The field’s distinctive contribution is an exact runtime stability check for its one-step operator, whereas learned recurrences can also have sufficient and conservative certificates.The paper distinguishes certificate type from capability rather than claiming exclusive certifiability.

A Scope and status of the stability results

The stability section separates formal theorems, sufficient operator-level certificates, and numerical verification. The sharpest claims apply only under stated hypotheses, while trained-run guarantees rely on runtime checks and are not closed-loop theorems.

  • What is proven: The Verlet criterion is necessary and sufficient per latent root, and its proof requires no commutation hypothesis.The operator-level certificate obtained through a parameter box is sufficient only and may be conservative.
  • Hypotheses and corrections: The closed-form flutter threshold is exact only when stiffness and damping are identity multiples, a regime absent from every reported implementation.Positive structural diffusion makes the published threshold sufficient but not necessary, while the headline configuration uses no antisymmetric operator.
  • What is verified numerically: Numerical audits reproduce identities and threshold locations and observe no certified-condition instability, but these checks are not theorems.They validate fixed parameter draws and long-horizon integration behavior.
  • Trained-run guarantees: For trained runs, the operative guarantee is the exact spectral-radius runtime certificate plus a training-time stability penalty.Certificates are audited on trained checkpoints rather than logged at every training step.
  • Scope of empirical verification: The evidence-accumulator kill-gate used twelve frozen solvers from a companion integration program, not the paper’s n=10 protocol.Transfer of the conclusion depends on the shared recurrent-reasoner structure.

B Proofs of the certified results

The paper proves stability results for three integrator structures and introduces a necessary-and-sufficient discrete Verlet criterion per latent root without assuming operator commutation. Operator-level certificates obtained through parameter boxes are sufficient and may be conservative.

  • Placement dichotomy: The placement dichotomy distinguishes norm-neutral first-order coupling, energy-preserving gyroscopic coupling, and potentially destabilizing circulatory coupling.Antisymmetric terms vanish from the relevant energy derivatives in the first-order and gyroscopic cases, but can pump energy under circulatory placement.
  • Flutter threshold: β ρ(A3) < 2ζω2_0 is necessary and sufficient for asymptotic stability of the isotropic unforced circulatory system.The threshold follows by reducing antisymmetric coupling to invariant eigenplanes and tracking imaginary-axis crossings.
  • IMEX kernel bound: The IMEX kernel is unconditionally contractive when the antisymmetric operator is implicit, with no timestep or commutation condition.For state-dependent forcing, the composed-map contraction additionally requires LF < ω2_0, which the paper does not assert across the operating range.
  • Discrete Verlet criterion: |z| < 1 iff ˜µ^2 < g(2 − g) and q(g^2 + ˜µ^2) < 2g for each latent Verlet root.The result applies Cohn’s criterion to the scalar quadratic induced by a latent vector and remains valid without simultaneous diagonalization.
  • Operator-level certification: The latent-root criterion is exact per root, whereas replacing the joint numerical range with a parameter box yields only a sufficient whole-operator certificate.The box check is finite but must include endpoints and interior critical points, and the paper reports marked conservatism.

C A common-Lyapunov certificate: what it buys, and where it stops

A common quadratic Lyapunov function certifies a norm bound for the wave branch and extends to mixtures and gated products within that branch. The certificate stops at loose coefficient boxes, enlarged branch mixtures, and the unclosed learned forcing loop.

  • What it buys: The common-P norm bound applies to convex mixtures and gated products within the wave branch, unlike a runtime spectral-radius probe.The stronger certificate controls products and mixtures through a common norm rather than spectral radius alone.
  • What it buys: ρ = 0.952 and κ = 2.51 define a common-P certificate for the trained wave-checkpoint envelope, whose worst vertex spectral radius is 0.732.The trained models lie well inside the certified envelope.
  • Where it stops: 40 of 64 vertices in the declared coefficient box diverge, with worst ρ = 5.24, so the nominal box is too loose for certification.The certifiable frontier collapses to c = 0 and ω0 ≤ 0.276; training, rather than the declared box, keeps the model away from divergent configurations.
  • Where it stops: The enlarged set has no common P, and under the wave branch’s P the maximum norm is 1.57 when diffusion and antisymmetric operators are included.Thus the wave-branch certificate does not extend to the full mixture.
  • Where it stops: The small-gain condition requires LF < 0.0192, but measured learned-forcing gains range from 0.367 to 0.918.Its failure does not establish closed-loop instability; it establishes only that this sufficient technique does not certify the loop.
Loading 2608.24319v1…