Source-linked AI summary

A Commutator Framework for Selective Spectral Alignment in Deep Neural Networks

Kaj Nyström

arXiv:2608.22910v1stat.MLcs.LG

TL;DR

The paper addresses how spectral incompatibility is generated and transported across layers, rather than assuming training universally drives commutators to zero. It develops a finite-width commutator hierarchy with exact source decompositions and buffered localization, finding that feature-side alignment can mask internal mixing and that genuine decay requires additional geometric or damping conditions. The resulting picture is selective, mechanism-dependent alignment rather than a universal consequence of training.

  • Problem

    The paper asks how incompatibility among gates, covariance, sensitivities, AGOPs, and NFMs is generated, transported, localized, and cancelled in finite-width networks, instead of assuming universal commutator collapse.

  • Method

    The paper builds a commutator hierarchy with an exact four-source sensitivity-covariance recursion, singular-value-weighted feature transport, buffered spectral localization, and conditional stabilization and decay estimates.

  • Results

    Feature-side alignment is a spectrally filtered image of internal alignment, while analytic and numerical results exhibit source cancellation, transient growth, and selective rather than universal commutator decay.

  • Takeaways & Limitations

    Spectral alignment is a layer- and scale-dependent compatibility phenomenon governed by transport, interaction, cancellation, and possible damping.

  • Takeaways & Limitations

    The experiments are finite-time observations restricted to tested architectures, widths, datasets, and optimization horizons, and the theory requires additional hypotheses for genuine damping or decay.

Abstract

from arXiv · show

We develop a finite-width geometric framework describing how learned feature geometries are organized, transported, and selectively aligned in deep neural networks. Incompatibility among weight-generated covariance, gates, and backward sensitivities is quantified through three families of commutators: between gates and covariance, between sensitivities and covariance, and between average gradient outer products (AGOPs) and neural feature matrices (NFMs). An exact layerwise identity decomposes the sensitivity-covariance commutator into four sources: downstream transport, adjacent-layer imbalance, pointwise sensitivity fluctuations, and nonlinear gate-covariance interactions. The AGOP-NFM commutator is a singular-value-weighted transport of the internal commutator, explaining why observed feature-side alignment alone does not determine the internal geometry from which it emerges. Buffered localized energies resolve mixing between separated covariance subspaces. We establish spectral-gap, projector-evolution, and stabilization estimates, and formulate conditional Lyapunov principles that yield decay under explicit geometric error-bound or intrinsic-damping assumptions. These criteria do not follow from gradient flow alone and clarify why risk reduction need not imply commutator collapse. Analytic examples and numerical experiments exhibit factorization of spectral and activation geometry, transient growth, and cancellation among nonzero sources. In tested finite-time regimes, cancellation dominated by a negative transport-imbalance interaction persists across depths, widths, and two regression benchmarks. Spectral alignment therefore appears as a layer- and scale-dependent compatibility phenomenon governed by transport, interaction, cancellation, and possible damping, rather than a universal consequence of training.

1. Introduction

The paper frames spectral alignment as compatibility among activation, internal sensitivity, and feature-side geometries, linked by transport rather than universal training-driven collapse. It develops exact source decompositions, buffered localization, and conditional decay criteria to explain alignment, cancellation, and stabilization.

  • Motivation: The paper studies how training generates, transports, localizes, and potentially cancels incompatibility among gates, covariance, sensitivities, AGOPs, and NFMs.The framework is developed primarily for finite-width, bias-free, fully connected, scalar-output networks under population gradient flow.
  • Alignment hierarchy: The AGOP-NFM commutator is a singular-value-weighted feature-side image of the internal sensitivity-covariance commutator, so feature alignment can hide internal incompatibility.Small internal incompatibility can imply feature-side alignment, but the converse may fail through suppression in small singular directions and invisibility in covariance-null directions.
  • Localization: Buffered localized energies isolate direct coupling between separated covariance subspaces while allowing rotations or mixing inside a buffered near-degenerate cluster.Spectral-gap estimates relate localized defects to global commutator energies, and vanishing localized energy is weaker than global commutativity.
  • Dynamics: Finite variation can stabilize localized energies without making them vanish; genuine decay requires compatibility, a geometric error bound, or intrinsic transverse damping beyond gradient flow.Propagation through the hierarchy additionally requires control of imbalance, fluctuations, operator scales, and relevant spectral gaps.
  • Alignment hierarchy: The internal commutator obeys an exact backward recursion whose sources are downstream transport, adjacent-layer imbalance, sensitivity fluctuations, and gate-covariance interactions.These sources can reinforce or cancel, so the gate commutator alone does not explain the total internal incompatibility.
  • Evidence: Analytic examples and numerical diagnostics show spectral-activation factorization, transient growth, and cancellation among nonzero sources, supporting selective rather than universal commutator decay.The experiments use matched finite-time risk comparisons and identify mechanism-dependent alignment rather than risk-driven collapse.

2. Relation to existing work

The paper situates its finite-width commutator framework between AGOP/NFA feature-alignment work, deep-linear balancedness, and related dynamical and geometric approaches. It studies weaker eigenspace compatibility and decomposes its finite-width mechanisms without assuming balanced initialization or decay.

  • The framework studies eigenspace compatibility between AGOPs and NFMs rather than imposing an exact spectral functional relation between their eigenvalues.
  • Its point of departure is an exact layerwise transport identity for finite-width trajectories that decomposes internal incompatibility into four mechanisms.
  • Unlike deep-linear analyses, the framework does not assume balanced initialization or commutator decay, allowing transient growth and cancellation among nonzero sources.
  • Buffered localized energies distinguish mixing across separated covariance subspaces from mixing within nearly degenerate spectral clusters.
  • The deterministic hierarchy is complementary to asymptotic mean-field and tensor-program analyses and does not assume a universality principle for data-dependent trained-network gates.

3. Preliminaries

The preliminaries define a bias-free feedforward network, its forward covariances, backward sensitivities, AGOPs, NFMs, and three commutator diagnostics. They establish the transport and imbalance notation used to analyze alignment along population gradient flow.

  • The model is a bias-free, fully connected scalar-output network with locally Lipschitz componentwise activation and population-risk training.
  • AGOPs average outer products of output sensitivities, while pointwise and population sensitivity operators separate data-dependent backward geometry from its average.
  • Forward covariance Pℓ and NFMℓ describe outgoing and incoming weight-generated geometries and share the same nonzero eigenvalues.
  • The three commutators measure gate-covariance, sensitivity-covariance, and AGOP-NFM compatibility, with no decay assumed during training.
  • Adjacent-layer imbalance compares outgoing covariance at one layer with incoming covariance at the next, while centered sensitivity fluctuations have zero population mean.
  • Gradient-flow dissipation couples forward activation correlations with backward loss-sensitivity correlations.

4. The internal commutator hierarchy

The internal sensitivity-covariance commutator is the central, uncompressed object in an exact hierarchy. Its recursion combines downstream transport with imbalance, sensitivity-fluctuation, and nonlinear gate-covariance sources, which can cancel rather than decay automatically.

  • The exact recursion contains one transported downstream contribution and three local sources: adjacent-layer imbalance, pointwise sensitivity fluctuation, and nonlinear gate-covariance interaction.
  • The feature-side commutator is a congruence image of the internal commutator, so singular-value attenuation and covariance-null directions can conceal internal incompatibility.
  • The four source components may substantially cancel in Frobenius space, and none is asserted to decay under training.
  • A small internal commutator can reflect individually small components or cancellation among large components, whereas norm estimates control only the former mechanism.
  • Nonlinear dynamics can generate off-diagonal adjacent-layer imbalance even when neuronwise balanced initialization removes the conserved diagonal contribution.
  • Sensitivity-fluctuation smallness is an additional hypothesis or empirical observation, not a consequence of the hierarchy or gradient flow.

5. Spectral localization and covariance geometry

The localization theory resolves where covariance-spectrum mixing occurs and tracks the motion of the associated subspaces. Buffered energies, spectral gaps, projector estimates, and singular-value weights distinguish separated-block coupling from poorly resolved or feature-suppressed mixing.

  • Localization strongly detects mixing across separated spectral clusters but is weakly sensitive to mixing inside nearly degenerate clusters.
  • Buffered localized energies isolate direct coupling from a leading covariance subspace to a remote complement while allowing mixing through intermediate buffer directions.
  • Projector evolution depends inversely on the separating spectral gap and is insensitive to basis rotations within either spectral cluster.
  • A gap remains open when the accumulated spectral variation is smaller than its initial value, providing a stability condition for the selected subspaces.
  • Feature-side localization detects the same internal mixing entries but weights them by singular-value products, suppressing mixing involving small singular directions.
  • Pointwise AGOP alignment is stronger than population AGOP alignment because it additionally requires control of localized sensitivity fluctuations.
  • Higher-order localization emphasizes widely separated modes but does not produce a scale-independent monotone hierarchy.

6. Activation geometry and gate switching

Activation gates evolve through local weight and preactivation dynamics, while smooth and ReLU analyses control gate variation without supplying an intrinsic alignment mechanism. Spectral localization links gate geometry to covariance subspaces but distinguishes stabilization from vanishing misalignment.

  • Activation dynamics: The gate Dℓ is a pointwise nonlinear function whose dynamics differ from those of the averaged sensitivity operator Qℓ.Qℓ arises through backward transport and population averaging, whereas Dℓ depends pointwise on preactivation.
  • Activation dynamics: Gate motion depends on local weight velocity and activation dynamics transported from preceding layers.The forward recursion gives ∂tzℓ = Ẇℓaℓ + Wℓ∂taℓ and ∂taℓ+1 = Dℓ∂tzℓ.
  • Smooth activations: For smooth activations, bounded second derivatives control gate variation through preactivation variation but do not imply decay of the gate-covariance commutator.The estimates have no preferred sign.
  • ReLU switching: For ReLU, activation regions and their interfaces describe gate switching, while symmetric-difference identities provide a robust global formulation.The gate can be continuous in L2 without being classically differentiable.
  • Spectral localization: In the covariance eigenbasis, gate-covariance energy couples spectral separation with overlap geometry of activation regions.Smooth and ReLU estimates control variation, not its sign; convergence therefore does not force localized energies to vanish.

7. Stabilization and conditional propagation of spectral alignment

Localized commutator energies can stabilize as covariance and activation geometries vary finitely, but stabilization alone does not imply alignment. Conditional propagation requires separate control of transport, imbalance, fluctuations, spectral scales, and gaps.

  • Hierarchy: The hierarchy is not an unconditional decay chain: the gate commutator is one source of internal incompatibility, while feature-side geometry is a transported image of combined internal geometry.The relations distinguish stabilization of localized energies from decay to zero.
  • Stabilization: Localized energies converge when the relevant gates, sensitivities, and covariance projectors have controlled late-time variation, but their limiting cross blocks may remain nonzero.Vanishing occurs exactly when the limiting buffered cross block vanishes.
  • Conditional decay: A conditional dissipation criterion yields decay when a nonnegative uniformly continuous localized energy is integrable through risk dissipation.This mechanism is not a general consequence of gradient flow or bounded gate operators.
  • Global versus localized alignment: Global decay and localized decay imply one another only under covariance-scale and spectral-gap conditions.Without bounded scale and separated interfaces, the corresponding implications need not hold.
  • Propagation: Gate alignment alone does not propagate through the internal hierarchy because transported incompatibility, imbalance, and sensitivity fluctuations require independent control.The assumptions in (7.9) are explicitly independent of gate alignment.

8. Conditional coercivity and Lyapunov principles

Risk reduction controls geometric defects only under additional coercivity assumptions. The paper develops two conditional routes—risk-based geometric error bounds and intrinsic transverse damping—while emphasizing their restrictive, model-dependent scope.

  • Motivation: PL-driven exponential decay controls excess risk, not localized alignment defects.Additional coercive relations are needed to deduce selective alignment from optimization.
  • Risk-based transfer: A geometric error bound transfers risk decay to selected sensitivity or gate defects when each defect is bounded by excess risk.The resulting estimate is qi(t) ≤ KiE(t), followed by exponential decay under PL assumptions.
  • Risk-based transfer: The error-bound route is restrictive because the selected cross block must vanish on the relevant minimizing component.Interpolation and local PL conditions alone do not imply this spectral compatibility.
  • Intrinsic damping: Intrinsic damping yields exponential decay without assuming qi is controlled by excess risk, provided the sensitivity evolution contains a negative transverse term.The mechanism can remove persistent late-time defects when the model-dependent damping estimate holds.
  • Scope: The linear example verifies consistency of the Lyapunov assumptions but does not establish generic nonlinear feature alignment.Its defect decay partly results from collapse of the sensitivity operator Q(t).

9. Analytic examples

Analytic examples separate gate, internal, and feature-side alignment, showing that compatibility can arise through orthogonality, cancellation, parametrization, or transport filtering. Gradient-flow trajectories can also exhibit transient growth rather than monotone alignment.

  • Separation of alignment levels: The examples show that the three commutator levels need not align and that internal sources can cancel or grow transiently.Deep linear networks provide a reference where the nonlinear gate source vanishes identically.
  • Two-neuron geometry: Gate compatibility is not monotone in neuron correlation: orthogonality and coincident activation regions provide distinct zero mechanisms, while complementary gates maximize the example's gate energy.For fixed neuron norms, ρ = 0 and ρ = 1 yield zero energy, whereas ρ = −1 yields 2∥w1∥2∥w2∥2.
  • Transport filtering: Gate alignment can coexist with nonzero internal incompatibility, while rank collapse can hide internal mixing from the feature-side commutator.A well-conditioned weight matrix instead transmits internal incompatibility quantitatively.
  • Source cancellation: Exact sensitivity and feature-side alignment can result from cancellation among nonzero local sources rather than from small individual mechanisms.The cancellation is a concrete realization of source reinforcement and cancellation in the commutator recursion.
  • Parametrization: The same represented function can carry different commutator geometries because ReLU rescaling changes parameter geometry without changing the network function.Exact internal and feature-side alignment can be obtained through a positive-homogeneous reparametrization.
  • Transient growth: Along constrained gradient flow, gate energy may first increase and then decrease even while the trajectory converges to exact activation alignment.The initial growth reflects increasing covariance coupling; later decline reflects decreasing gate disagreement.

10. Numerical experiments

The experiments show selective, mechanism-dependent alignment rather than universal commutator decay: risk reduction leaves defects distinct, while width and benchmark changes affect geometric levels differently. Persistent transport–imbalance cancellation and parametrization sensitivity further indicate that small commutators need not reflect uniformly small source production.

  • Risk reduction and depth: At r = 10−3, the input-layer gate means are 0.4486, 0.4486, and 0.4487 for H = 2, 4, 6, while internal means are 0.0646, 0.0805, and 0.0793.Feature means are 0.0211, 0.0199, and 0.0201, and interior defects remain nonzero at order 10−2.
  • Risk reduction and depth: A hundredfold risk decrease does not produce global commutator decay: the gate defect stays essentially constant, while input internal and feature defects increase before stabilizing.At r = 10−4, the corresponding input values are 0.4487, 0.0791, and 0.0211.
  • Source cancellation: The net internal commutator is about thirty percent of the summed source magnitudes because transport and imbalance exhibit a persistent negative interaction.The interaction lies between −0.712 and −0.703 for H = 4 and between −0.716 and −0.649 for H = 6.
  • Source cancellation: Across layers, imbalance becomes dominant, transport rises to approximately 20%-25%, and fluctuation contributions decrease toward the output, maintaining small defects through reproducible cancellation.Near the input, fluctuation and nonlinear gate sources carry a substantial share of source energy, whereas transport is small.
  • Width dependence: Increasing width is associated with smaller internal and feature defects but not a smaller gate defect over the tested range.At layer 1, the internal mean decreases from 0.0635 at width 16 to 0.0185 at width 64, while the feature mean decreases from 0.0555 to 0.0170.
  • Benchmark and intervention results: The two benchmarks reproduce the separation and source-cancellation pattern, while function-preserving rescaling sharply changes source geometry without changing the initialized function or activation gates.For c = 1/2, transport and imbalance nearly cancel; for c = 2, transport carries only 0.21% of source energy.

11. Concluding remarks

The paper presents a hierarchy for analyzing selective spectral alignment and shows that alignment is governed by transport, interaction, cancellation, and possible damping rather than training risk reduction alone.

  • The hierarchy links gate-covariance, sensitivity-covariance, and AGOP-NFM commutators through structural coupling rather than logical implication.
  • The sensitivity-covariance recursion identifies downstream transport, adjacent-layer imbalance, pointwise sensitivity fluctuations, and nonlinear gate-covariance interactions as four sources.
  • Risk reduction does not impose universal monotone alignment because small net commutators can arise from cancellation among large sources or singular-value filtering.
  • Decay requires additional sufficient mechanisms, namely geometric error bounds or intrinsic transverse damping.
  • The paper identifies verifiable model-dependent localized damping conditions, stochastic-training extensions, widenetwork limits, and implicit selection among interpolation-compatible internal geometries as next problems.
  • The manuscript reports generative AI assistance for language editing, organization, formatting, consistency checking, and numerical-experiment development and debugging.
Loading 2608.22910v1…