Source-linked AI summary
Parameter-Level Attribution of Symmetry in Trained Networks Though Parameter-Wise Functional Sensitivity
Alan Muriithi, Vedanta Thapar, Torben Berndt
TL;DR
The paper asks whether a learned function’s symmetry can be realised as motion in neural-network parameter space. It develops functional-sensitivity and tangent-space projections to test this, finding that recomputed directions track symmetry orbits and reduce equivariance defects locally in classifiers and Hamiltonian networks.
Problem
It is unclear whether a function-space symmetry action can be lifted through a neural-network parametrisation and which individual parameters encode the symmetry.
Method
The paper uses functional sensitivities and projects infinitesimal symmetry generators onto the parameter tangent space, producing symmetry-directed and equivariance-directed parameter directions.
Results
Resolved directions track predicted function-space dynamics locally, while fixed directions deviate after training in both the annulus classifier and Hamiltonian-network experiments.
Takeaways & Limitations
Parameter-level differential geometry can attribute learned symmetries and distinguish motion along a symmetry orbit from motion toward equivariance.
Abstract
from arXiv · showhide
When a network has learned a function with a known symmetry, can that symmetry be moved through the parametrisation---is there a motion in parameter space realising the group action in function space? We formulate this as a lifting problem for the realisation map $Φ:θ\mapsto f_θ$, and show that a smooth parameter-space action exists only if the tangent space to the function's symmetry orbit lies within the image of $\mathrm dΦ_θ$, whose columns are the \emph{functional sensitivities} of individual parameters. This condition is also sufficient for pointwise first-order lifting. Relaxing it in least squares yields two local parameter directions: one following the symmetry orbit, one descending towards the equivariant subspace, with residuals measuring what the parametrisation cannot reach. On a rotationally invariant classifier we find these directions induce their predicted function-space motion, but only locally: recomputed directions track the orbit and reduce the equivariance defect, while directions held fixed depart from both after training. The same holds for Hamiltonian neural networks trained on a rotationally symmetric potential, even though the architecture does not explicitly enforce the symmetry.
1. Introduction
The paper introduces functional sensitivity to attribute learned-map behavior to individual parameters and studies how parameters encode symmetries. It contrasts this local differential geometry with parameter-space correlation methods.
- Functional sensitivity: Functional sensitivity measures the local response of a neural-network output map to small perturbations of an individual parameter.Aggregate sensitivity averages the magnitude over inputs using a probe distribution and a sensitivity-normalised gauge.
- Functional sensitivity: Functional sensitivity shifts analysis from comparing responses across inputs to comparing the roles of individual parameters.The object is mathematically related to the NTK, which compares pointwise sensitivities at different inputs for fixed parameters.
- Symmetry attribution: The study asks which parameters encode symmetries of learned maps in networks trained from sampled trajectory data.This motivates parameter-level attribution of symmetry generators and diagnostics for equivariance.
- Symmetry attribution: The method projects infinitesimal symmetry generators onto the parameter tangent space to obtain per-parameter symmetry attributions.The contributions also compare these quantities in Hamiltonian neural networks and identify directions along symmetry orbits and toward the equivariant subspace.
- Relation to prior work: Unlike prior parameter-space correlators, the approach analyzes local differential geometry and identifies specific parameters contributing to a given symmetry.The comparison is with Maiti et al.’s inference of symmetry manifestation through parameter-space correlators.
2. Symmetry Orbits and Infinitesimal Lifting
The paper formulates symmetry transfer as a differential lifting problem: network Jacobians determine which function-space symmetry directions parameter changes can reach. Least-squares projections separate motion along symmetry orbits from motion toward equivariance.
- Differential lifting: A global parameter-space lift may fail because transformed functions need not be realised by parameters or admit smoothly chosen representatives.The paper therefore studies the weaker infinitesimal lifting problem.
- Differential lifting: The network Jacobian determines which function-space directions are locally accessible through parameter variations.The orbit differential supplies symmetry directions, while the image of the realisation-map differential supplies reachable directions.
- Differential lifting: The tangent-space inclusion condition is necessary for lifting and sufficient for pointwise first-order lifting, but not for a smooth local or global parameter-space action.Thus pointwise infinitesimal reachability is weaker than constructing an action on parameters.
- Parameter attribution: When parameters lie in R^p, Jacobian columns are individual functional sensitivities, and coefficients of the projected direction attribute symmetry to parameters.This connects the lifting criterion directly to parameter-level attribution.
- Local alignment objectives: The method computes one locally accessible direction along a symmetry orbit and another toward the equivariant subspace.The two objectives probe distinct aspects of local geometry.
- Local alignment objectives: Residual norms quantify the parts of symmetry or equivariance-directed motions that lie outside the locally accessible tangent space.The symmetry residual measures inaccessible orbit motion, while the equivariance residual measures inaccessible motion toward the equivariant projection.
3. Experiments
Experiments test whether locally computed parameter directions produce their predicted function-space motions in classifiers and Hamiltonian neural networks. Recomputed directions work over finite trajectories, whereas fixed directions drift after training.
- Experimental design: The experiments evaluate whether local parameter-space directions induce their predicted finite changes in function space.The classifier comparison uses symmetry-directed and equivariance-directed directions with resolved and fixed trajectories.
- Annulus classifier: In the wedge-trained classifier, the equivariance-directed flow reduces the non-invariant component while the symmetry-directed flow follows the rotational orbit.The symmetry-directed flow approximately preserves distance from the equivariant subspace, and directions are recomputed at every integration step.
- Annulus classifier: Resolved classifier trajectories closely follow predicted dynamics, but fixed directions can deviate substantially after training.The discrepancy is much smaller at initialisation, showing that the construction is reliable locally but requires recomputation over finite trajectories.
- Physics-informed neural networks: The physics experiment uses a rotationally equivariant Mexican hat potential and Hamiltonian neural networks trained from trajectory data across multiple α values.The dynamics are exactly equivariant under simultaneous rotations of q and p.
- Physics-informed neural networks: For the Hamiltonian networks, walking along c⋆ shifts the potential while walking along b⋆ reduces equivariance error.With fixed directions, similar behavior appears initially but drift accumulates over time.
- Experimental design: The study frames these experiments as tests of whether symmetry actions can be lifted infinitesimally through neural-network parametrisations.This connects the empirical flows to the differential lifting problem introduced earlier.
A.1. Motivation and Setup
The paper asks whether a function-space symmetry can be lifted through a neural-network parametrisation into parameter space. Because the realised model class or smooth parameter representatives may not preserve the function-space action, it focuses on the weaker infinitesimal question.
- Function-space symmetry: The setup represents symmetry through a Lie-group action on inputs, outputs, and differentiable functions, whose fixed points are equivariant functions.The induced function-space action has equivariant functions as its fixed points.
- Neural-network parametrisation: A neural network is described by a finite-dimensional parameter manifold and a realisation map from parameters to functions.The realised model class is the image of this map.
- Global lifting: The strong lifting problem asks whether a smooth parameter-space action can make the realisation map equivariant for every group element and parameter.This would require transformed functions to be represented consistently by transformed parameters.
- Global lifting: The realised model class need not be preserved by the ambient function-space action, and even an invariant model class may lack smoothly selectable parameter representatives.Thus, a global lift need not exist.
- Infinitesimal lifting: The paper therefore studies whether each first-order symmetry variation of a realised function can be generated by a first-order parameter variation.This compares the symmetry orbit tangent directions with those accessible through the differential of the realisation map.
A.2. The Differential Lifting Criterion
The differential lifting criterion compares infinitesimal symmetry directions with function variations generated by parameter changes. Inclusion of the former in the latter is necessary for smooth lifting and sufficient for pointwise first-order lifting, but not for a smooth local or global parameter action.
- Accessible function directions: The differential of the realisation map sends parameter velocities to first-order changes in the realised function.Its image is the space of function variations accessible through the parametrisation and, at regular points, the tangent space of the realised model class.
- Infinitesimal lifting: Differentiating the lifted action yields an operator equation relating parameter velocities to infinitesimal symmetry directions in function space.The parameter velocity associated with a Lie-algebra element is mapped by the realisation-map differential to the corresponding orbit direction.
- Differential lifting criterion: Every infinitesimal function-space orbit direction must lie in the image of the realisation-map differential if a smooth parameter action exists.This is the necessary differential lifting condition.
- Differential lifting criterion: At a fixed parameter value, the same inclusion is sufficient for pointwise first-order lifting.It does not guarantee that the lifts vary smoothly, satisfy Lie-bracket relations, or meet integrability conditions.
A.3. A Least-Squares Lifting Objective
The least-squares lifting objective finds the parameter direction whose induced function variation best approximates a symmetry generator. Its residual is exactly the unreachable component outside the model tangent space, while the minimizer attributes the reachable direction to individual parameters.
- When exact lifting fails, least squares quantifies the discrepancy between a symmetry-induced function variation and the parametrisation's reachable variations.The objective is posed using Hilbert-space inner products and the orthogonal projector onto the model tangent space.
- The minimum lifting error is the symmetry-orbit component orthogonal to the model tangent space, vanishing exactly when the differential lifting criterion holds.This error measures representational reachability in function space rather than optimizer performance.
- Because the objective decouples over an orthonormal basis of the Lie algebra, each symmetry generator can be solved independently.The resulting minimum-norm solution is computed one generator at a time.
- The solution coefficients provide per-parameter attributions of each symmetry generator through the columns of the Jacobian.Those columns are the pointwise functional sensitivities of individual parameters.
A.4. The Equivariance Defect and its Attribution
The equivariance-defect construction separates motion along a symmetry orbit from motion toward the equivariant subspace. In numerical experiments, these directions are obtained from discretized functional sensitivities and a truncated, normalized Jacobian computation.
- A.4. The Equivariance Defect and its Attribution: The orbit direction preserves the equivariance defect to first order, whereas the negative defect is the steepest-descent direction toward the equivariant subspace.These objectives are complementary because the orbit direction is orthogonal to the defect direction.
- A.5. Numerical Realisation: In practice, function-space quantities are discretized on finite probe sets, turning the Jacobian, orbit direction, and defect into explicit finite-dimensional vectors.The aggregate sensitivity is computed as a root-mean-square column norm, while the full sensitivity vectors span accessible tangent directions.
- A.5. Numerical Realisation: The numerical tangent space is resolved with a truncated SVD, treating sufficiently small singular directions as noise rather than reachable variation.Without truncation, tiny nonzero singular values on finite grids can make overparameterized networks appear fully reachable.
- A.5. Numerical Realisation: Sensitivity-normalized attribution reduces dependence on parameter rescaling, while leaving the projected function direction, residual, and resolved rank unchanged.Columns below 10^-8 of the maximum norm are treated as dead and assigned zero attribution.
- A.5. Numerical Realisation: For Hamiltonian experiments, a rotationally symmetric polar probe grid and 64-point SO(2) quadrature are used to avoid grid-induced cross-talk.The grid contains 15 radii × 32 angles, or N = 480 points.
B.1. ASRNN architecture
The experiments use adaptable symplectic recurrent neural networks that represent separable Hamiltonians with separate multilayer perceptrons for kinetic and potential energies. A leapfrog integrator generates recurrent trajectory predictions trained by trajectory matching.
- ASRNNs parameterize a time-independent separable Hamiltonian as H(q, p; α) ≈ Kθ1(p) + Vθ2(q; α).Separate neural networks model kinetic and potential energies.
- The architecture introduces recurrence through a leapfrog integrator that predicts the next phase-space state from the current one.The chosen time step is denoted by Δt.
- Training minimizes a trajectory-matching loss comparing model predictions with ground-truth phase-space trajectories over observed time points.Initial conditions and physical parameters specify each trajectory.
- The potential and kinetic networks are three-hidden-layer MLPs with width 32.This follows the stated experimental hyperparameter regime.
C.1. Mexican Hat
The Mexican hat dynamics are rotationally equivariant under simultaneous rotations of position and momentum, with a radial force whose rotational Lie derivative vanishes. The ASRNN evaluation compares locally recomputed and fixed parameter directions using angular deviation profiles.
- The Mexican hat system is exactly equivariant under simultaneous rotations of q and p.The transformation is (q, p) ↦ (Rφq, Rφp).
- The radial force has vanishing rotational Lie derivative, establishing rotational symmetry as a property of the true system rather than the architecture.
- Figures 3 and 4 report angular deviation profiles for the ASRNN under Resolved and Fixed trajectories, respectively.The Resolved trajectory recomputes the local direction at every step.
C.2. Annulus
The annulus experiments compare fixed and resolved parameter-space flows for symmetry-directed and equivariance-directed directions across training conditions. Resolved directions are recomputed at every step, whereas fixed directions remain initialized once.
- Figures 5 and 6 compare fixed and resolved flows for the symmetry-directed direction c⋆ and equivariance-directed direction b⋆ across initialisation and trained models.
- After training, fixed directions can deviate substantially from their intended function-space evolution, while recomputed directions more faithfully follow the symmetry orbit or reduce the equivariance defect.
- Fixed flows compute parameter-space directions once at t = 0 and retain them throughout the trajectory.
- Resolved flows recompute the Jacobian and parameter-space directions at every step.