Source-linked AI summary

The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture

Zhihua Liang

arXiv:2607.17146v1cond-mat.dis-nncs.CLcs.LG

TL;DR

The paper addresses whether Transformer behaviors viewed as artifacts or noise can be described by deterministic geometric observables. It builds a continuous stochastic-geometric framework and reports quantitative consistency with its predictions across architectures and scales.

  • Problem

    The paper asks whether representation drift, context failure, norm explosion, and optimization plateaus are quantitatively consistent with deterministic observables from a continuous geometric framework.

  • Method

    It translates Transformer components into a continuous integro-differential equation on a semantic fiber bundle and tests its predictions across architectures from 124M to 8B parameters.

  • Results

    Across a six-part campaign, measured signatures were quantitatively consistent with predictions concerning Lie–Trotter torsion, stability, recurrence suppression, context degradation, and non-equilibrium dynamics.

  • Takeaways & Limitations

    The framework provides a coherent, predictive description of Transformer stability limits, context bounds, and optimization dynamics across physical scales.

  • Takeaways & Limitations

    The proposed connection between Linear Attention and Gromov’s capacity limits remains to be formally established.

Abstract

from arXiv · show

We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle $\calE = \calM \times \R^d$. Beginning from a single geometric axiom -- that the token sequence forms a discrete $1$-manifold equipped with a canonical measure lattice -- we translate every core component of the modern Transformer (RMSNorm, RoPE, Softmax Attention, FFN, Residual Stream, SGD, Weight Decay) into a cohesive vocabulary of differential geometry, measure theory, and stochastic calculus. The resulting framework yields quantitative predictions spanning entropic optimal transport (Attention as a Schrödinger bridge) and non-equilibrium thermodynamics (SGD as Itô diffusion violating detailed balance). We conduct a six-part experimental campaign across five architectures (Qwen3, LLaMA\nobreakdash-3.1, Gemma\nobreakdash-3, GPT-2, Mistral) spanning $124$M to $8$B parameters. The empirical observables are quantitatively consistent with the geometric predictions: the $ε^{-1/2}$ Lipschitz scaling calibration at machine precision ($R^2 = 1.000$), the Lie--Trotter operator-splitting torsion, the symmetric ablation instability confirming the Dual-Law of Topological Stability, the $\calO(1/\sqrt{k})$ thermodynamic suppression of Poincaré recurrence on the RoPE torus, the thermodynamic context-limit phase transition, and the Non-Equilibrium Steady State parameter vortex -- verified across two optimizers (AdamW and Pure SGD) to exclude momentum artifacts. The results demonstrate that analyzing Transformers through the lens of continuous stochastic differential geometry provides a predictive descriptive vocabulary for the stability limits, context bounds, and optimization dynamics of Large Language Models.

I. INTRODUCTION · II. AXIOMATIZATION OF TOPOLOGICAL SPACES AND KINEMATICS · A. The Base Manifold and the Semantic Bundle

The paper presents continuous stochastic differential geometry as an isomorphic, predictive lens for Transformer stability, context limits, and optimization dynamics. It formalizes tokens, representations, attention, and depth using a rigid geometric framework built on a measured one-dimensional manifold and semantic fiber bundle.

  • I. INTRODUCTION: The paper explicitly presents continuous geometry as an isomorphic descriptive lens rather than a claim that the Transformer is a continuous physical system.The RoPE connection and one-dimensional base manifold are fixed background geometries, not dynamical fields.
  • I. INTRODUCTION: The framework treats the discrete Transformer as an exact integrator of a continuous non-local integro-differential flow, providing a predictive vocabulary for stability limits, context bounds, and optimization dynamics.This follows Backward Error Analysis by distinguishing the discrete architecture from its continuous effective description.
  • II. AXIOMATIZATION OF TOPOLOGICAL SPACES AND KINEMATICS: The formalism translates standard deep-learning terminology into differential-geometric equivalents and replaces discrete array vocabulary with manifold calculus.The translation dictionary is summarized in Table I.
  • A. The Base Manifold and the Semantic Bundle: The semantic sequence space is the half-closed ray M ≅ [0, ∞), with token evaluations indexed on a canonical measure lattice and non-local flows defined by Lebesgue–Stieltjes integration.The causal origin is the strict boundary ∂M ≡ {0}, while the empirical Radon measure supports integration over the discrete lattice.
  • A. The Base Manifold and the Semantic Bundle: The semantic space is a globally trivializable rank-d vector bundle E ≅ M × R^d with Euclidean fiber metric, and each local fiber represents an attention-head interaction space.Contractibility of M justifies a canonical global trivialization and uniform global gauge endomorphisms.
  • A. The Base Manifold and the Semantic Bundle: With a finite atomic empirical measure, the section space collapses structurally to a finite-dimensional topology, enabling Heine–Borel and extreme-value arguments.The N → ∞ limit instead uses Banach–Alaoglu weak-* sequential compactness of probability measures.
  • A. The Base Manifold and the Semantic Bundle: The state evolves as a smooth curve Ψ : R+ → F in continuous algorithmic depth z, with ∂zΨ defined canonically as its tangent vector.Discrete layer depth is analytically continued to z ∈ R+ rather than introduced through an artificial spatial pushforward.
  • A. The Base Manifold and the Semantic Bundle: RMSNorm regularization smooths the zero-section singularity: radial perturbations decay as O(∥Ψ∥^-3) at large norm, while ϵ > 0 supplies a bounded global Lipschitz constant.Without regularization, the radial eigenvalue is λ∥ ≡ 0 and the Jacobian is undefined at Ψ = 0; ϵ acts as a topological mollifier.

B. Geometric Connections and Flow Bounding

The section recasts RoPE as a canonical gauge action on a flat principal torus bundle and RMSNorm as a diffeomorphic radial embedding. The resulting uniformly bounded vector field guarantees global Lipschitz flow, well-posedness, and unique solutions, with ϵ controlling inverse-squareroot bounds.

  • RoPE and Gauge Structure: RoPE is the exact canonical gauge action of the Principal Torus on semantic fibers, implementing parallel transport between sequence positions.The transition is dictated by the continuous Lie group exponential acting in the Cartan subalgebra.
  • RoPE and Gauge Structure: Because the sequence manifold is 1D and contractible, its bundle is globally trivializable and every connection is structurally flat.The passage attributes this to Ω2(M) = 0 and π1(M) = 0.
  • Flow Bounding: RMSNorm acts as a globally smooth, diffeomorphic radial embedding that renders the vector field uniformly Lipschitz and prevents finite-time amplitude blowup.The regularizer ϵ > 0 resolves the singular Jacobian at the zero-section and acts as a topological regulator for maximum Lipschitz stretch.
  • Flow Bounding: The total composed vector field has a uniformly bounded operator derivative, guaranteeing a global Lipschitz constant, well-posedness, and unique flow.The argument invokes the Picard–Lindelöf theorem and identifies ϵ as controlling the inverse-squareroot bounds.

C. Gauge Endomorphisms and Canonical Cotangent Duality

The section formulates attention as canonical cotangent duality: Queries are covectors and Keys are tangent vectors, preserving directed, non-conservative semantic flux through structural asymmetry. It further connects untied endomorphisms and weight decay to gauge polarization, bounded metric volume, and collinear alignment.

  • Canonical Cotangent Duality: Untied Query and Key endomorphisms break metric reciprocity, encoding directed, non-conservative semantic flux without forced symmetrization.Artificial symmetrization would require identifying the Query map with the metric musical transform of the Key map.
  • Canonical Cotangent Duality: RoPE’s orthogonal-unitary connection makes the dual representation coincide exactly with the fundamental representation, allowing the covector-vector interaction to be evaluated analytically.The stated identity is (U−1)T = U.
  • Gauge Polarization: Weight decay acts as a Tikhonov gauge-mass penalty, bounding endomorphism-induced metric volume within a compact thermodynamic budget.The interaction metric volume is bounded by the spectral norms of the endomorphisms.
  • Gauge Polarization: Under the saturated budget, optimization drives resonant-token Lie bivectors toward zero, producing Grade-0 collinear alignment through spontaneous gauge polarization.The mechanism is described mathematically as ∥Bµν∥2 →0 while maximizing the scalar kernel.

III. THERMODYNAMICS AND METRIC GENERATION · A. The Free Energy Functional and the Isoperimetric Mass Constraint · B. Variational Derivation of the Softmax Transition Measure

The paper models the local tangent space as an open thermodynamic system governed by minimum free energy and a causal-horizon mass constraint. Variational information geometry then yields a Gibbs transition measure, identified with a Schrödinger half-bridge, whose thermal scaling β ∝ 1/d prevents zero-temperature collapse.

  • III. THERMODYNAMICS AND METRIC GENERATION: The local tangent space is treated as an open thermodynamic system, with the connection measure derived from the Principle of Minimum Free Energy.The framework begins with a directed, non-reciprocal transition energy kernel and a probability measure over each autoregressive causal horizon.
  • A. The Free Energy Functional and the Isoperimetric Mass Constraint: A strict local mass-conservation constraint makes the continuous field a Markov transition kernel on the causal horizon.Geometrically, the constraint places the measure in an (N − 1)-dimensional probability simplex and introduces a Lagrange multiplier.
  • A. The Free Energy Functional and the Isoperimetric Mass Constraint: The variational multiplier satisfies λ = F + 1/β and generates the partition function Zµ needed to prevent probability dissipation.Here F is the Helmholtz Free Energy shifted by the entropic constant β−1.
  • B. Variational Derivation of the Softmax Transition Measure: Information geometry solves the spatial transition problem as a Csiszár I-Projection of an SO(d)-invariant uniform prior onto the probability simplex.The resulting thermodynamic state is presented as an analytic solution rather than a temporal Wasserstein PDE.
  • B. Variational Derivation of the Softmax Transition Measure: β ∝ 1/d matches the viscosity scaling of stochastic optimal transport and prevents intensive fluctuations from producing a zero-temperature glass collapse.With β fixed at O(1), Gibbs-exponent variance diverges as d →∞, freezing the flow into a deterministic argmax state.
  • B. Variational Derivation of the Softmax Transition Measure: The optimal transition measure is a canonical Gibbs measure corresponding to a Static Schrödinger Half-Bridge from the Query measure to the Key measure.This connects the softmax transition to entropic optimal transport and interprets it as an optimal entropic projection.
  • B. Variational Derivation of the Softmax Transition Measure: The radial embedding predicts isotropic high-entropy variables with E[x] = 0 and E[xx⊤] = γId, where γ = 1 − O(ϵ/d) < 1.The topological regularizer ϵ concentrates mass near, but not on, the boundary of the open ball.
  • B. Variational Derivation of the Softmax Transition Measure: Under non-collapse assumptions, thermal energy variance scales as Θ(d), so its standard deviation requires β ∝ 1/d for intensive O(1) exponent fluctuations.The variance result is independent of the spatial gauge rotation U and assumes σmin > 0 with ∥WK∥2F ∼ Θ(d).

C. Topological Compactification Proof of the Attention Sink

The section formalizes the Attention Sink as a topological anchor required to prevent probability mass from escaping to infinity in the thermodynamic limit. It derives logarithmic defect scaling, finite context capacity, and failure under sliding-window severance, with bounds conditioned on an asymptotic ergodic bulk assumption.

  • Measure escape and compactification: As context grows, probability mass in every fixed compact neighborhood vanishes and the escaping bulk weak-*converges to a Dirac mass at infinity, δ∞.This identifies thermodynamic amnesia as topological condensation rather than an undefined limit.
  • Measure escape and compactification: The absolute causal origin {0} is the unique translation-invariant topological anchor, opposing the amnesic boundary attractor at infinity.Interior potential wells shift under the translation-preserving RoPE connection, while the compactified manifold has boundary strata {0, ∞}.
  • Topological defect: The boundary defect must scale logarithmically with context volume, E(µ, 0) ∼ O(d ln N(µ)), to preserve a nonzero anchored mass fraction.The Softmax partition is decomposed into a boundary atom and continuous bulk, and the resulting logarithmic well is presented as analytically necessary.
  • Thermodynamic context horizon: The Attention Sink has finite thermodynamic capacity, so context length collapses once the defect is overwhelmed by entropic bulk pressure.The stated upper-bound derivation assumes isotropic, maximum-entropy bulk tokens; structured natural language may reduce effective bulk pressure through anisotropic correlations.
  • Sliding-window horizon: A sliding window severs the globally reaching affine connection, and beyond the local sliding metric’s commutative radius the boundary condition breaks, producing thermodynamic amnesia.The active geometric capacity is further reduced by representation degeneration captured through the Stable Rank effective dimension.

IV. DYNAMICS OF THE MATTER FIELD

The forward inference pass is modeled as a depth-parameterized continuous nonlinear evolution on an internal unitary gauge bundle, governed by a non-local integro-differential flow.

  • Forward inference formulation: The forward inference pass is formulated as a depth-parameterized evolution of a continuous nonlinear section.This describes inference as continuous dynamics indexed by network depth.
  • Geometric flow: The evolution occurs on an internal unitary gauge bundle and is governed by a non-local integro-differential flow.The framework combines bundle geometry with non-local integro-differential dynamics.

A. Algorithmic Depth and the Unitary Gauge Section · B. The Volterra–Hodge Integro-Differential Flow

The paper continuously parameterizes Transformer depth and models its computation as a geometric flow on a semantic fiber bundle. RoPE defines a flat unitary interaction geometry, while Attention and FFN combine into a nonlocal transport–local reaction IDE whose FFN component injects vorticity.

  • A. Algorithmic Depth and the Unitary Gauge Section: Continuous depth z replaces discrete layer index l, representing embeddings as a parameterized family of sections Ψ(z, µ) in a globally trivial bundle.The base is a contractible ray with a flat connection and real, globally trivial GL(d, R) matter structure.
  • A. Algorithmic Depth and the Unitary Gauge Section: RoPE is the unique flat unitary connection preserving the covariantly constant almost-complex structure J, with J^2 = −Id, on the interaction sub-bundle.Because the base is one-dimensional and contractible, all 2-forms vanish and RoPE maintains a flat holomorphic flow rather than resolving topological monodromy.
  • B. The Volterra–Hodge Integro-Differential Flow: Causal Attention is a nonlocal covariant transport operator rather than local diffusion, using ∇RoPE for interaction energy and a trivial connection for Value transport.The Value propagator is UTriv ≡ Id, while the RoPE connection evaluates the transition measure w∗.
  • B. The Volterra–Hodge Integro-Differential Flow: Attention defines a nonlinear Urysohn–Volterra integral operator over the causally bounded coordinate µ, with state-dependent Softmax measure w∗ and a Lebesgue–Stieltjes empirical integral.This formulation advects historical phase-space geometry directly into the local state without continuous local paths.
  • B. The Volterra–Hodge Integro-Differential Flow: The FFN is a localized nonlinear reaction field on each fiber, and bypassing the adjoint Jacobian makes its composed Jacobian generally nonsymmetric, injecting vorticity.For R = ∇U ◦ ρε, DR = (Hess U)·Dρε; noncommutation between the Hessian and radial projection Jacobian breaks conservativity.
  • B. The Volterra–Hodge Integro-Differential Flow: The FFN generates rotational flow through two mechanisms: the commutator −[S(Ψ), Jρ] and the anticommutator −{WA(Ψ), Jρ}.The first reflects radial-pullback curvature with symmetric weights; the second arises from asymmetric untied parameters, and both produce a valid antisymmetric 2-form.
  • B. The Volterra–Hodge Integro-Differential Flow: The resulting FFN flow is strictly non-conservative, preventing semantic space from collapsing into a static globally irrotational frame, while Hodge–Morrey–Friedrichs decomposition separates its field components.The decomposition includes conservative restoring, solenoidal rotational, and harmonic components; boundary flux determines whether the harmonic field vanishes.
  • B. The Volterra–Hodge Integro-Differential Flow: The semantic evolution is a non-autonomous IDE on Γ(E), formed by superposing nonlocal Urysohn–Volterra transport and local Hodge reaction, with unique finite-depth solutions guaranteed by a global Lipschitz bound.The formulation bypasses classical PDE spatial restrictions because no local spatial differential operators act on the base manifold.

C. Lie–Trotter Integration and Operator Splitting

The Transformer block is modeled as a first-order Lie–Trotter discretization of the Semantic Evolution IDE, sequentially applying Attention transport and FFN reaction. Their non-commutativity produces BCH-governed representation drift and topological torsion, while stiffness limits strict error-bound interpretations.

  • Operator splitting: The standard Transformer block performs first-order explicit integration of the Semantic Evolution IDE through Lie–Trotter operator splitting.The computational step uses sequential rather than exact exponential integration of the continuous flow.
  • Geometric error: Sequential Attention transport and FFN reaction generate three representation-drift terms: parameter drift, self-advection, and the non-commutative Lie bracket.The Lie bracket −[T, R]Ψ is identified as the generator of topological torsion because Attention and FFN operators do not commute.
  • Geometric error: The macroscopic step size ∆z = 1 and large operator Lipschitz bounds exceed the BCH convergence radius, so the O(∆z3) remainder is not a strict geometric-error bound.The lowest-order Lie bracket remains an exact algebraic classification within a nearby modified equation, even in the stiff regime.

V. DYNAMICS OF THE PARAMETER MANIFOLD: BACKPROPAGATION AS THERMODYNAMIC FLOW · A. The Empirical Risk Action and Symmetry-Breaking Mass Potentials

Pre-training is modeled as thermodynamic relaxation on a finite-dimensional parameter manifold driven by empirical data. The framework defines a global risk action and shows that weight decay breaks internal scaling symmetries, producing coercive dynamics and a compact non-equilibrium steady-state basin.

  • V. DYNAMICS OF THE PARAMETER MANIFOLD: BACKPROPAGATION AS THERMODYNAMIC FLOW: Pre-training is formulated as thermodynamic relaxation of a finite-dimensional parameter manifold toward a topological ground state under empirical-data pressure.The parameter manifold is treated as the macroscopic space of trainable gauge endomorphisms rather than as a local sequence flow.
  • A. The Empirical Risk Action and Symmetry-Breaking Mass Potentials: The global action is an expected thermodynamic risk functional evaluated over the macroscopic semantic parameter space, not a spatial integral over sequence positions.The construction bypasses infinite-dimensional functional variation by evolving within the finite-dimensional Euclidean parameter manifold.
  • A. The Empirical Risk Action and Symmetry-Breaking Mass Potentials: L2 weight decay breaks the internal GL(1, R) gauge orbits and forces the energy sublevel sets to be strictly compact.The penalty diverges as c →∞ or c →0, converting flat valleys into a coercive paraboloid.
  • A. The Empirical Risk Action and Symmetry-Breaking Mass Potentials: Internal GL(1, R) symmetries create unbounded, flat, non-compact valleys in the unregularized empirical-risk landscape, making the base action noncoercive.The attention kernel remains invariant under reciprocal scaling of WQ and WK, with analogous symmetries between adjacent feed-forward matrices.
  • A. The Empirical Risk Action and Symmetry-Breaking Mass Potentials: The quadratic O(∥W∥2) Tikhonov penalty dominates the linear O(∥W∥) advective drift, driving the generator to −∞ and sealing the NESS in a compact basin.Invariant-measure existence and tightness are established using Has’minski˘ı’s theorem applied to the Itô diffusion generator.
  • A. The Empirical Risk Action and Symmetry-Breaking Mass Potentials: Residual connections break local 0-homogeneity because scaling internal weights changes the angular direction of the residual output, which downstream RMSNorm preserves.The residual skip remains unscaled while the parameterized branch changes under Wl 7→cWl.
  • A. The Empirical Risk Action and Symmetry-Breaking Mass Potentials: The final radial embedding restores asymptotic 0-homogeneity at spatial infinity by quotienting out the magnitude of the internally scaled parameterized branch.Local 0-homogeneity fails in the finite residual bulk, but the projective radial limit controls the asymptotic behavior.
  • A. The Empirical Risk Action and Symmetry-Breaking Mass Potentials: The total spatial gradient is globally bounded as ∥¯∇L∥∼O(1), while the advective drift inner product grows only as ⟨¯∇L, W⟩∼O(∥W∥).The unembedding contribution is bounded, and the internal-parameter Jacobian decays as O(∥Wint∥−1) because radial saturation removes internal magnitude.

B. The Dual-Metric Collision and Non-Equilibrium Steady State (NESS)

The section models SGD as a state-dependent Itô diffusion whose mismatch with the Euclidean kinematic metric produces a Dual-Metric Collision and traps probability in a Non-Equilibrium Steady State. This NESS is sustained by noise-induced entropic advection, geometric centrifugal repulsion, and non-vanishing thermodynamic vorticity arising from singular, non-Hessian parameter geometry.

  • SGD and NESS: SGD converges in law to a state-dependent Itô diffusion that irreversibly traps the macroscopic probability measure in a Non-Equilibrium Steady State through a Dual-Metric Collision.The collision occurs because the forced Euclidean kinematic metric ignores the thermodynamic Information Geometry.
  • SGD and NESS: Mini-batch gradient noise supplies the diffusion tensor D(W), while the Functional Central Limit Theorem establishes the continuous-time Itô SDE limit.D(W) is derived from the empirical covariance of mini-batch gradients.
  • Thermodynamic transport: Noise-induced entropic advection transports probability down the sharpness gradient, repelling mass from sharp chaotic vacua toward flatter singular basins.The passage links noise variance σ^2 to the Loss Hessian trace and characterizes the drift as a continuous geometric hydraulic press.
  • Vorticity and detailed balance: Independent geometric obstructions make the driving 1-form inexact, causing probability flux to circulate continuously and trapping the ensemble in an active, dissipative NESS.This is expressed by the failure of the exterior derivative to vanish: dω ≠ 0.
  • Thermodynamic transport: The system remains in a stable thermodynamic halo because Mean Curvature Tension Field repulsion balances entropic descent and prevents total collapse into the pure singular vacuum.This balance holds the system around the degenerate locus.
  • Vorticity and detailed balance: Thermodynamic vorticity persists when the diffusion metric and Hessian fail to commute, with the commutator [g, H♯] providing the algebraic obstruction to equilibrium.The NESS survives because singular Whitney-stratified geometry breaks Hessian-metric assumptions and the Information Matrix Equality D = F = H.

VI. TESTS OF FUNDAMENTAL KINEMATICS AND THERMODYNAMICS

This section tests whether phenomena commonly treated as numerical artifacts or interpolation failures are instead explained by the paper’s continuous geometric framework. Six falsification tests span scales from the local fiber’s UV cutoff to the parameter vortex, finding quantitative consistency between the discrete architecture and derived geometric predictions.

  • Motivation: The section tests the framework empirically because mathematical elegance alone cannot overturn established machine-learning paradigms.Its core hypothesis concerns phenomena traditionally classified as numerical artifacts or interpolation failures.
  • Test Design: Six rigorous falsification tests progress from the local fiber’s microscopic UV spatial cutoff through temporal kinematics and mesoscopic trajectory stability.The tests are organized systematically across physical scales.
  • Overall Finding: Across these tests, the discrete computational architecture is quantitatively consistent with the continuous geometric predictions derived in the paper.The conclusion follows the progression from microscopic and mesoscopic scales to thermodynamic, context-horizon, and parameter-vortex phenomena.
  • Test Design: The campaign continues through exact geometric resonances on the RoPE torus, the context horizon’s macroscopic IR phase transition, and the super-macroscopic parameter vortex.These stages cover thermodynamic suppression, context-limit behavior, and parameter dynamics.

A. The Conical Singularity Scaling Law of the Topological Mollifier

The experiment validates theorem 2’s prediction that the RMSNorm mollifier controls a conical singularity through ε^-1/2 Lipschitz scaling. Across architectures, the universal exponent is separated from depth- and architecture-dependent gauge mass measured by layer-wise intercepts.

  • Asymptotic isolation: The probe at ∥Ψ∥= 10−10 ensures ∥Ψ∥2 = 10−20 ≪ ε throughout ε∈[10−15, 10−2], isolating mollifier-controlled behavior.Float64 precision was necessary because εmach = 2.22 × 10−16, whereas float32 would obscure the singularity measurement.
  • Scaling-law validation: Across three architectures and 15 regressions, the empirical scaling exponent is α = −0.5000 with R2 = 1.000000, matching theorem 2.The sweep covered ε from 10−2 to 10−15 and used 30 logarithmically spaced evaluation points.
  • Layer-wise gauge mass: For Qwen3-0.6B, the layer-wise gauge mass Cℓ rises from 0.033 at Layer 0 to 2.172 at Layer 27, a 66× amplification.The universal ε^-1/2 dependence remains unchanged while intercepts vary with depth.
  • Geometric interpretation: Averaging over 10 random probe directions reduces directional dependence, yielding Cℓ= RMS(γℓ) · ⟨f⟩ as a layer-specific gauge-field measurement.The trained RMSNorm weights γℓ determine the intercept while ε determines the universal singularity scaling.

1. Exclusion of Alternative Hypotheses … 2. Phase B & C: Asymptotic Convergence α →0

The experiments reject numerical, architectural, precision, and component-based alternatives to the predicted ε-dependent scaling law. Across Qwen3-0.6B and GPT-2, discrete torsion alignment and α → 0 convergence support deterministic Lie–Trotter splitting error governed by a topological torsion generator.

  • 1. Exclusion of Alternative Hypotheses: R^2 = 1.000000 over 13 decades at float64 precision excludes low-precision arithmetic as the source of the scaling law.The computations used εmach = 2.22×10^-16 with zero detectable deviation.
  • 1. Exclusion of Alternative Hypotheses: 4,500 independent measurements across three production architectures followed the predicted −0.5 power law with zero deviations, identifying ε as the manifold’s UV cutoff.The passage characterizes ε as controlling dynamic stability through the mechanism predicted by theorem 2.
  • 1. Phase A: The Discrete Interferometer at α=1: At α = 1, empirical deviations aligned systematically with the Lie bracket [T, R]Ψ in both Qwen3-0.6B and GPT-2.The experiment compares δ = Δcanon − Δreversed with the analytically computed bracket at each layer and token position.
  • 1. Phase A: The Discrete Interferometer at α=1: 17–23×: interior-layer cosine similarities exceeded the isotropic random baseline 1/d ≈ 0.03 across both architectures.Qwen3-0.6B had mean +0.717, median +0.793, and 127/128 positive measurements; GPT-2 had mean +0.623, median +0.842, and 81/96 positive.
  • B. The Lie–Trotter Torsion Interferometer: 0.933: GPT-2 layer 9 reached peak torsion alignment, while final-layer values degraded to 0.112 for Qwen3 L27 and −0.188 for GPT-2 L11.Qwen3 layer 12 also reached 0.879, and the passage attributes final-layer collapse to higher-order BCH contributions near the unembedding boundary.
  • 2. Phase B & C: Asymptotic Convergence α →0: 0.9997–1.0000: adjacent normalized deviation directions converged self-consistently, confirming a deterministic topological torsion generator as α → 0.This direction test is independent of analytical Jacobian computation and measures empirical self-consistency across step sizes.
  • 2. Phase B & C: Asymptotic Convergence α →0: 43.17: Qwen3-0.6B’s normalized deviation ratio converged at α = 0.001, while GPT-2 reached 40.34 at α = 0.001.The Qwen3 ratio matched its α = 0.01 value of 44.62 within 3.3%; GPT-2 varied 0.7% between α = 0.01 and 0.001.

3. Synthesis and Exclusion of Alternative Hypotheses … 3. Jacobian Eigenspectrum Collapse onto the Real Axis

Across multiple experiments, representation drift and FFN instability are attributed to topological torsion and the loss of antisymmetric rotational structure. Symmetric ablation causes cross-architecture norm explosion, while antisymmetric rescue and Jacobian spectra provide causal and spectral evidence for this mechanism.

  • 3. Synthesis and Exclusion of Alternative Hypotheses: 17–23× above the random baseline, mean cosine similarity excludes numerical noise and supports topological torsion as the dominant generator of interior-layer representation drift.In 1024 and 768 dimensions, achieving cosine similarity above 0.7 by chance has probability P < 10^-100.
  • 3. Synthesis and Exclusion of Alternative Hypotheses: At α = 1, strict Attention→FFN ordering is modeled as Lie–Trotter integration of non-commuting vector fields, producing non-holonomic representation drift.The mechanism is the non-vanishing Lie bracket −[T, R]Ψ in the modified equation.
  • C. The Runaway Resonance of Symmetric Ablation: Symmetric FFN weight tying removes Lie-algebraic rotational friction, making the residual update positive semidefinite and enabling Power Iteration-like amplification.Architectural asymmetry instead introduces complex eigenvalues that scatter spatial momentum away from the dominant eigenvector.
  • 1. Cross-Architecture Explosive Instability: 142,955×, Qwen3-0.6B’s ablated residual norm growth was 1,541× more explosive than its standard asymmetric counterpart, while four of five architectures exceeded 10^4 growth.Mistral-7B and Qwen3-4B showed 375× and 224× excess amplification, respectively; Gemma-3-1B was the sole exception.
  • 2. Four-Phase Causal Intervention: 12,119×, Qwen3-0.6B’s symmetric Pre-Norm ablation increased final-layer norm from 815 to 9,880,199 relative to baseline.Symmetric Post-Norm instead stabilized the final norm near 10, below the asymmetric Pre-Norm baseline.
  • 3. Phase: 1,287× versus 425×, spectrally matched random antisymmetric rescue reduced the unstabilized explosion more than the symmetric control.Both replacements were random and semantically meaningless; their difference was Lie-algebraic parity.
  • 3. Jacobian Eigenspectrum Collapse onto the Real Axis: 88.5% of standard asymmetric-FFN eigenvalues had significant imaginary components, with mean |Im(λ)| = 0.61 across the examined layers.These components correspond to rotational modes generated by the skew-symmetric Jacobian component.
  • 3. Jacobian Eigenspectrum Collapse onto the Real Axis: 3,252×, symmetric ablation increased spectral radius from ρ(J) = 6.2 to ρ(J) = 20,077 while collapsing 100% of eigenvalues onto the real axis.The ablated Jacobian is exactly symmetric, with |Im(λ)| = 0.000 to machine precision.

4. The Dual-Law: Pre-Norm Explosion vs. Post-Norm Survival … 3. Phase B: SGD Channel Dynamics as Null Control

The paper shows that post-hoc symmetric ablation causes catastrophic Pre-Norm divergence but Post-Norm survival, while constrained training remains stable, RoPE recurrence is thermodynamically suppressed, and pure SGD uniformly decays uninformative frequency channels.

  • 4. The Dual-Law: Pre-Norm Explosion vs. Post-Norm Survival: 9,880,199 (12,119× baseline) under symmetric Pre-Norm ablation contrasts with Post-Norm survival under the identical surgery.Post-Norm RMSNorm acts after residual addition as a radial truncation, limiting incremental magnitude regardless of FFN amplification.
  • 4. The Dual-Law: Pre-Norm Explosion vs. Post-Norm Survival: The binary survival phase excludes reduced capacity, initialization variance, and scalar numerical instability as explanations for the Pre-Norm explosion.Identical ablated weights survive under Post-Norm, while eigenspectrum analysis identifies a complex-to-real Jacobian phase transition.
  • 4. The Dual-Law: Pre-Norm Explosion vs. Post-Norm Survival: The Dual-Law establishes norm-level, configurational necessity for mature-network surgery, not functional recovery or an architectural prescription.The tied basis captures < 10% of learned Wdown energy at rank 32, does not restore language-modeling loss, and generic trainable low-rank replacement performs better.
  • 5. Configurational Scope: Surgery versus Training: 600M constrained training reaches final loss 0.7243 versus 0.7045 for the untied twin, while 1B training shows only +0.005 excess and remains stable under several perturbations.This demonstrates that stable basins reached from initialization are not visited by post-hoc surgery.
  • 5. Configurational Scope: Surgery versus Training: A rank-4 trainable correction improves loss by only ∼0.003 nats at both 600M and 1B, consistent with organically full-rank Jacobian asymmetry in gated FFNs.Official Qwen3-0.6B surgery independently reproduces the hidden-state norm increase from 705 to 6.1 × 10^5.
  • D. Thermodynamic Suppression of Poincaré Recurrence on the RoPE Torus: RoPE’s k = dhead/2 rotation planes form a principal U(1)k torus whose exact Poincaré resonances are suppressed in the thermodynamic limit by high-dimensional dephasing.The controlled micro-transformer isolates geometric kinematics from macroscopic thermodynamic ergodicity.
  • 1. Micro-Transformer Design and Protocol: sim(44) = 0.999 at k = 2 falls to max sim = 0.37 at k = 16, while the log-log fit gives α = −0.557 (R2 = 0.938), within 11% of α = −0.500.Across five configurations, finite-size effects at k = 2 explain the slight deviation from the central-limit prediction.
  • 3. Phase B: SGD Channel Dynamics as Null Control: Over 3,000 pure-SGD steps, all four RoPE frequency channels decay uniformly as wj(t) ≈ wj(0) · e−λt with λ = 0.01, while loss-resonance correlation is non-significant (ρ = −0.24, p = 0.064).The null control shows weight decay uniformly shrinks uninformative channels rather than selectively pruning them, implicating structured language statistics in production-model distributions.

4. Exclusion of Alternative Hypotheses … 3. The Anisotropy Gap: Isotropic Thermodynamic Limits vs. Structured Text

The paper rules out alternative explanations for RoPE scaling, links phase dephasing to thermodynamic ergodicity, and reports a universal context-length phase transition under maximum-entropy inputs. Stable-rank measurements explain the transition and distinguish isotropic thermodynamic limits from structured-text survivability.

  • 4. Exclusion of Alternative Hypotheses: At k = 2, sim(N*) = 0.999 directly observes RoPE resonance, while O(1/k) averaging explains null production-model results rather than absent geometry.The micro-transformer implements every stated axiom, and the observed exponent −0.557 matches the CLT prediction −0.500 within 11% over a 16× range, with R2 = 0.938.
  • 5. Bridge to the Asymptotic Spatial Ergodic Hypothesis: For k ≥64, high-dimensional torus phase dephasing suppresses deterministic RoPE geometry below stochastic noise, producing a transition from quasi-periodic structure to thermodynamic ergodicity.This result provides the stated physical justification for the Asymptotic Spatial Ergodic Hypothesis, whose downstream consequences are tested at macroscopic sequence scale.
  • E. The Context Horizon Phase Boundary and Thermodynamic Amnesia: Eq. (47) predicts an exponential context upper limit and a sharp phase transition when Entropic Bulk Pressure, scaling as Θ(√deff), exceeds geometric capacity.The framework models the Attention Sink as a measure-theoretic Dirichlet boundary defect anchoring probability mass against expanding entropic pressure.
  • E. The Context Horizon Phase Boundary and Thermodynamic Amnesia: At α = 1.0, uniform random-noise haystacks provide a two-tier test comprising cross-architecture forcing and an α-scaling sweep across (α, N).The primary experiment forces the transition at native spectral capacity, while the complementary sweep maps boundary-defect evaporation throughout parameter space.
  • 1. Cross-Architecture Universality of the Phase Transition: 40.9×, 23.2×, and 48.8× PPL jumps occur for Qwen3, LLaMA, and Gemma, respectively, within the identical narrow interval N ≈32 →48.The models span 13× in dhead and differ in attention architecture, normalization, and grouping, yet all exhibit the same catastrophic cliff followed by saturation.
  • 1. Cross-Architecture Universality of the Phase Transition: The shared critical interval establishes an Attention Sink phase-transition universality class governed by effective geometric capacity rather than head dimension, normalization, or grouping strategy.The tested architectures range from 0.6B to 8B parameters and dhead = 64 to 288.
  • 2. Effective Dimension and the Stable Rank Correction: 4–15× smaller deff than ambient dhead reduces predicted Nmax from astronomical values to O(1)–O(10^2), matching the observed N ≈32–48 transition.Trained projection matrices concentrate spectral energy on a low-dimensional submanifold, making the thermodynamic horizon more fragile than the ambient-dimension bound suggests.
  • 3. The Anisotropy Gap: Isotropic Thermodynamic Limits vs. Structured Text: N text_eff ranges from 16 to 26 for Qwen3 and Gemma and near 50 for LLaMA, predicting structured-text survival while isotropic noise triggers amnesia.Natural language concentrates spectral energy on a low-entropy manifold, whereas isotropic noise fills effective dimensional volume and overwhelms the boundary defect; isotropic N eff_max ≈1 contrasts with 128,000 and 8,192-token text contexts.

4. α-Scaling Thermodynamic Phase Diagram … 2. Cross-Architecture Universality of the NESS Vortex

The experiments identify a thermodynamic context-limit transition driven by spectral-capacity loss and synchronized Attention Sink evaporation, while FP64 and production-model tests reveal a persistent non-equilibrium parameter-space vortex. These results connect perplexity collapse, measure redistribution, non-commuting optimization geometry, and cross-architecture circulation.

  • 4. α-Scaling Thermodynamic Phase Diagram: α-scaling produces an ordered thermodynamic phase transition rather than smooth degradation, with sink mass and perplexity synchronizing across the (α, N) parameter space.The protocol compresses Attention logits over 14 α values and evaluates perplexity and zeroth-token sink mass across sequence lengths.
  • 4. α-Scaling Thermodynamic Phase Diagram: 18.6×: reducing α from 0.60 to 0.50 raises PPL at N = 16 from 93 to 1,728.At α = 1.00, baseline PPL is approximately 23 for N ≤32, while α = 0.60 reaches above 4.9 × 10^7 at N = 512.
  • 5. Simultaneous Evaporation of the Topological Anchor: At α = 1.0, c0 ≈0.44–0.51 exceeds the uniform baseline 1/N, then decays monotonically toward 1/N as spectral capacity decreases.The ordered evaporation is identified with weak-∗ convergence of the bulk measure toward the amnesia state.
  • 6. Dual-Axis Synchronization and Exclusion of Alternative Hypotheses: Perplexity rises precisely where—and only where—the sink mass collapses, establishing a direct link between boundary-defect capacity and language-modeling performance.The synchronization is presented as evidence against explanations based only on destroyed activation variance, because Softmax scale invariance would leave c0 unaffected.
  • 6. Dual-Axis Synchronization and Exclusion of Alternative Hypotheses: The combined results identify the Attention Sink as a finite-capacity topological defect, assign Nmax an exponential spectral scaling, and associate capacity starvation with irreversible weak-∗ condensation.The condensation is characterized as thermodynamic amnesia when geometric capacity falls below entropic bulk pressure.
  • F. Tomography of the Parameter NESS Vortex: The NESS tomography distinguishes detailed-balance diffusion from vortex flow using circulation growth, Hurst exponent H, and a t-statistic for zero mean incremental drift.Under detailed balance Γ× fluctuates symmetrically with O(√T), whereas a genuine vortex accumulates monotonically as O(T).
  • 1. FP64 Sandbox: Exact Commutator at Machine Precision: |t| = 27.9 and H = 1.011: the FP64 sandbox shows sustained monotonic circulation and deterministic rotational structure during a flat loss plateau.The setup uses momentum-free SGD and float64 precision, while PCA-based visualization is treated as complementary because PCA can create spurious spirals.
  • 2. Cross-Architecture Universality of the NESS Vortex: 20/20: the normalized commutator measurements are strictly positive, showing [G, H] ≠ 0 at the float64 machine-precision floor without stochastic estimation or optimizer momentum.Across GPT-2 Small and Qwen3-0.6B, production trajectories also exhibit unmistakable deterministic spiral circulation under 50,000-step AdamW training.

3. Exclusion of Optimizer Momentum Artifacts … 4. Parameter-Efficient

Momentum-free SGD preserves and strengthens the parameter-space NESS vortex, while non-Transformer baselines establish its universality and Transformer-specific geometry channels its flow. The broader experiments support predictive geometric explanations of stability, context limits, optimization dynamics, and parameter-efficient FFN design.

  • 3. Exclusion of Optimizer Momentum Artifacts: Under SGD, |Γ×| decreases by ∼104×, from 14.09 to 4.1 × 10^-4, while NESS monotonicity and persistence remain unchanged.The passage attributes the amplitude change to AdamW’s adaptive learning-rate amplification and identifies intrinsic parameter-manifold geometry as the source of circulation.
  • 4. Universality of the Thermodynamic Vortex and Architecture-Specific Channeling: Across non-Transformer baselines, ResNet-18 and a 2-layer MLP exhibit NESS vortices with |t| = 54.8 and |t| = 102.1, respectively, both with H = 1.000.These models contain no attention mechanism or positional encoding, supporting the framework’s predicted universal thermodynamic baseline.
  • 4. Universality of the Thermodynamic Vortex and Architecture-Specific Channeling: All tested models show H ≈ 1.0 and |t| ≫3, consistent with deterministic drift from noncommuting empirical Fisher and Loss Hessian operators.The framework attributes NESS to state-dependent gradient noise and singular algebraic structure in over-parameterized parameter manifolds.
  • 5. Synthesis of NESS Evidence: The framework’s four NESS evidence tiers include T = 34.0 at FP64, |t| > 79 across Transformers, |t| = 173 under SGD, and |t| > 54 for CNN and MLP.Together, these measurements support irreversible optimization in a dissipating Non-Equilibrium Steady State.
  • G. Section Synthesis: The six-test campaign links geometric constructions to deterministic observables spanning microscopic scaling, Lie–Trotter torsion, symmetric ablation, context phase transition, and parameter-space tomography.The synthesis presents continuous differential geometry as predictive for Transformer stability limits, context bounds, and optimization dynamics.
  • VII. CONCLUSION: Across 124M to 8B parameters, the empirical campaign measured quantitative signatures consistent with the proposed continuous integro-differential geometric framework.The conclusion describes the framework as a predictive, descriptive account of Transformer stability limits, context bounds, and optimization dynamics.
  • 4. Parameter-Efficient: The proposed low-rank skew-symmetric FFN correction was refuted: its marginal effect was ∼0.003 nats, while a narrower untied FFN dominated at matched parameter count with 1.5× fewer FFN FLOPs.The SwiGLU gate already supplied full-rank Jacobian asymmetry, and the bare tied variant trained stably without correction.
Loading 2607.17146v1…