Source-linked AI summary

Reciprocity Separates Gradient Flow from Rotation in Conservative Physical Learning

Ruiwu Niu, Xiaowen Bi, Michaël Antonie van Wyk

arXiv:2608.30778v1cs.LGnlin.AO

TL;DR

The paper asks when conservative physical learning follows gradient descent versus a genuinely different trajectory. It analyzes matched transport responses and boundary feedback, showing that reciprocity yields preconditioned gradient flow while antisymmetric feedback enables rotation whose benefit depends on curvature and trajectory drift.

  • Problem

    The paper asks what determines whether physical learning follows gradient descent or a genuinely different learning trajectory, distinguishing conservation from reciprocity in that question.

  • Method

    The paper studies a directed layered mass-preserving transport network, derives matched-response and feedback structures, and evaluates local curvature and finite-horizon trajectory effects numerically.

  • Results

    Adjoint matching with non-negative spectral feedback produces reciprocal preconditioned gradient flow, whereas antisymmetric boundary feedback creates rotational closed-loop dynamics with instantaneous error decrease.

  • Takeaways & Limitations

    Nonreciprocal rotation should be introduced where local curvature supports it and controlled for the state drift it creates over subsequent updates.

Abstract

from arXiv · show

Physical learning lets a trainable material or network use its own physical response to carry error signals, reducing the need for a separately programmed backward computation. We ask what determines whether such a system follows conventional gradient descent or evolves along a genuinely different learning trajectory. Our canonical model is a directed layered transport network in which every node redistributes a fixed amount of flow, so learning preserves positivity and total mass. In this model, conservation constrains only the allowable learning directions. Within the matched response class studied here, adjoint matching gives the physical output response a symmetric form. Non-negative mode-wise feedback then produces a reciprocal closed-loop response and a reweighted gradient flow. Adding an antisymmetric boundary component makes the closed-loop response rotational: the learning path can turn while the error driving that update still decreases at that moment. Turning is not automatically beneficial. Its finite-step effect is set by local curvature, and its accumulated effect also depends on step selection and on the new states visited along the path. Numerical consistency checks reproduce the exact response structure, predict the sign of the local effect across new network families, and show how trajectory drift can negate a local advantage. These results separate the roles of conservation, reciprocity, and nonreciprocity in physical learning.

I. INTRODUCTION

The paper separates conservation, reciprocity, and rotational learning dynamics in a conservative physical learner. It develops matched-response theory and tests when nonreciprocal turning affects finite-step and finite-horizon learning.

  • Motivation: Conservation restricts parameter motion to positive, mass-preserving routing states but does not determine whether learning is downhill or rotational.The trainable state uses positive column-stochastic layer matrices, while score fields determine the motion within that constrained space.
  • Response geometry: Adjoint matching produces a symmetric physical response, and non-negative spectral feedback yields reciprocal, reweighted gradient flow.Within this response class, feedback changes metric and conditioning without creating independent circulation.
  • Response geometry: An antisymmetric boundary component adds a transverse closed-loop direction while preserving output mass and instantaneous selected-sample error dissipation.The resulting learning path can turn even though the error-driving component continues to descend at that moment.
  • Validation: The numerical tests verify response factorization and rotational modes, test a curvature-based sign criterion, and compare paired finite-horizon trajectories.The tests use disjoint network families and account for the final loss gap along registered reciprocal and nonreciprocal paths.
  • Implications: Nonreciprocity is therefore a design variable whose value depends on local curvature, step selection, and the later states visited by the trajectory.Turning freedom alone is not presented as a performance advantage.

III. CONSERVATION AND INTEGRABILITY

Conservation defines the feasible tangent space of simplex-valued parameters, whereas integrability determines whether the learning field derives from a scalar potential. Thus conservative motion can still contain circulation.

  • Conservation: Conservation confines each velocity to a simplex tangent space but does not itself impose a scalar potential or gradient structure.A gradient structure additionally requires an integrability condition on score differences.
  • Conservation: Every smooth conservative tangent field admits a block-replicator representation, but that representation alone does not imply loss minimization.The score is defined only up to an additive constant because the mobility operator annihilates the all-ones vector.
  • Integrability: Testing integrability requires removing one redundant simplex coordinate and checking whether the reduced score one-form is closed on a simply connected interior domain.The reduced score differences compare each coordinate with a reference coordinate.
  • Integrability: The rock–paper–scissors replicator field obeys the conservative dynamics while exhibiting periodic interior orbits, so it cannot be a gradient flow of one single-valued potential.This example demonstrates that conservation leaves circulation possible.

IV. ADJOINT-MATCHED RESPONSE AND SPECTRAL FEEDBACK

Adjoint matching gives the physical response a symmetric positive-semidefinite Gram structure at every depth. Non-negative spectral feedback preserves reciprocity and produces preconditioned gradient flow rather than independent rotational dynamics.

  • Adjoint-matched response: In the adjoint-matched model, the physical response K has a symmetric positive-semidefinite Gram form, while the boundary controller determines the closed-loop geometry.Physical-response symmetry and closed-loop reciprocity are distinct properties.
  • Adjoint-matched response: The layer-resolved response decomposes into contributions from individual layers and source columns, including cross-sample coupling.Depth changes spectrum, rank, and conditioning while preserving the Gram structure.
  • Spectral feedback: Non-negative spectral feedback makes the closed-loop response symmetric and positive semidefinite, yielding nonincreasing loss and positive-semidefinite preconditioned gradient flow.The result holds on the identifiable parameter subspace.
  • Spectral feedback: Changing the spectral function can alter descent metric and conditioning but cannot create an independent rotational mode within this feedback family.A regularized inverse is described as a metric-damped Gauss–Newton direction.

V. MINIMAL NONRECIPROCAL ESCAPE

Three outputs are the smallest single-sample setting that supports genuine rotational learning while preserving output mass. A matched physical response remains symmetric, but active boundary feedback can make the closed-loop response nonreciprocal.

  • Minimal dimension: Three outputs create the smallest single-sample zero-sum space in which genuine rotation can occur.Two outputs have only one independent mass-conserving direction and therefore remain confined to a line.
  • Controller design: P projects onto the zero-sum error plane, while C acts as a quarter-turn and α and Ω control inward relaxation and rotation.
  • Rotational response: The three-port construction preserves output mass and decreases the instantaneous selected-sample squared error.Its nonzero rotational component nevertheless gives residual eigenvalues −α ± iΩ.
  • Rotational response: The physical response K⋆ stays symmetric while the boundary controller makes the closed-loop response Rα,Ω nonreciprocal through its skew part.The symmetric component pulls inward and the skew component turns the error sideways.

VI. FINITE STEPS AND FINITE HORIZONS

Rotational learning can reduce or increase post-step loss because finite updates sample local curvature rather than only instantaneous dissipation. Over multiple steps, step policy and the changed states visited by each trajectory also affect the final loss gap.

  • Finite-step effects: A rotational displacement can land above or below the corresponding reciprocal finite step despite instantaneous loss reduction.The relevant comparison is post-step loss at the same pre-step state.
  • Finite-step effects: The second-order orientation-even effect relative to the reciprocal step is η2Ω2Q.Averaging opposite rotation orientations cancels terms odd in orientation.
  • Finite-step effects: Q can be positive or negative, with two strictly interior one-layer configurations giving Q = −0.9271875 and Q = 0.026666 . . . .Negative Q favors the mean of opposite rotational steps; positive Q favors the reciprocal update.
  • Finite horizons: A favorable local curvature sign does not guarantee a cumulative advantage because rotational updates change the states at which later updates are evaluated.
  • Finite horizons: Registered finite-horizon accounting separates orientation, step-policy, and visited-state contributions to the final loss gap.Et, Ht, St, and Dt are an exact accounting identity rather than a unique causal decomposition.

VII. NUMERICAL TESTS AND REPRODUCIBILITY

The numerical tests mirror the paper’s structural, local, and trajectory-level claims. They use new network families and paired protocols to assess response identities, curvature predictions, and finite-horizon effects.

  • Test sequence: The tests first assess structural identities and conditioning, then the local curvature sign rule, and finally registered finite-horizon accounting.
  • Network design: Depth comparisons vary the number of successive redistribution stages while keeping three outputs and hidden-layer width three.The tested depths are 2, 3, 4, and 6.
  • Reproducibility: Deterministic seeds, normalized exponential updates, derivative checks, and shared trajectory schedules support reproducible comparisons.Trials, rather than individual updates, are the independent experimental units.

A. Structural identities, conditioning, and local scaling

The structural checks reproduce the matched-response factorization at machine precision, while conditioning worsens with depth and can remove numerically usable modes. The local finite-step scaling and both signs of the curvature coefficient are also reproduced.

  • Response structure: The joint response K has six connected mass-conserving modes in the registered three-source, three-output linear family.Its numerical rank counts modes above the registered 10−10 tolerance.
  • Structural identities: The maximum factorization residual is 1.21 × 10−16 across 100 connected joint response operators.The experiment also reports maximum trajectory discrepancy 2.22 × 10−16 and maximum complex-mode residual 3.93 × 10−15.
  • Local scaling: The numerical checks reproduce instantaneous dissipation, both signs of Q, and fitted η2 finite-step slopes of 2.000 and 1.999.
  • Conditioning: All depths 2, 3, and 4 retain six modes, whereas 17 of 25 depth-6 operators retain only five.The depth-6 failures concern full numerical rank, not Gram factorization, symmetry, or positive semidefiniteness.
  • Conditioning: The depth-6 rank loss reflects weak modes approaching the finite-precision floor and may require regularization for inverse feedback.

B. Predicting the local sign in unseen network families

A selected-sample curvature ratio predicts whether rotational updates improve the orientation-averaged one-step loss across unseen network families. The diagnostic is highly accurate but remains a local predictor rather than a guarantee of cumulative advantage.

  • χs > 1 predicts Q < 0, meaning the orientation-averaged rotational step has lower one-step loss than the reciprocal step.The threshold of one marks curvature dominating displacement; Q < 0 is only a local comparison.
  • The favorable fraction rises from zero at β = 0 to 94.5% at β = 8 across four architectures.The 95% task-family interval at β = 8 is [91.0%, 97.5%].
  • The fixed predictor reaches AUC 0.998 and accuracy 97.1% on 1000 previously unused configurations.The intervals are [0.996, 0.999] for AUC and [96.2%, 98.0%] for accuracy, with 29 marked errors.
  • Analytic and finite-step signs agree in all configurations, while numerical rank can lose an identifiable response direction at depth six.Seventeen of 25 depth-6 operators lose one numerically identifiable direction, so the rank gate remains FAIL despite closed factorization.

C. Why a favorable local turn may disappear over time

Local rotational gains do not necessarily persist over a finite horizon because step allocation and the states visited later both affect the final loss. Adaptive stepping improves on step-matched nonreciprocal updates, but state drift offsets most recovered step-policy cost.

  • Adaptive stepping narrows the mean final-loss gap from 8.01×10^-3 to 4.96×10^-3 relative to the native reciprocal learner.Both gaps remain positive; the corresponding 95% trial-cluster intervals are [5.59, 10.63] × 10^-3 and [2.79, 7.41] × 10^-3.
  • State drift offsets 92.9% of the 4.35 × 10^-2 avoidable step-policy cost recovered by adaptive stepping.The decomposition separates same-state step allocation from later reciprocal progress after the policies reach different states.
  • Adaptive nonreciprocal updates improve on step-matched updates in every tested condition, without implying superiority to the native reciprocal learner.Panel (c) compares the two nonreciprocal policies directly, whereas panel (a) compares each with the native reciprocal control.
  • The framework identifies rotation as useful only when local curvature supports it and subsequent state changes remain controlled.The paper distinguishes local finite-step effects from accumulated trajectory outcomes and recommends a sign mechanism for using rotational components.
  • The formal scope is positive linear column-stochastic networks with fixed support and matched adjoint response.The effective active three-port boundary element leaves topology-local hardware realization open; nonlinear, signed, changing-support, and noisy cases remain tests of extension.

Appendix A: Proofs of the response-kernel results

The appendix derives the response-kernel structure and shows why matched reciprocal feedback remains gradient-like, whereas a skew boundary construction produces genuinely rotational modes.

  • In reduced simplex coordinates, the score differences form the covector dual to a tangent vector under the Shahshahani metric.Adding a constant to the score leaves the tangent vector unchanged, and the Poincaré lemma gives a potential when the associated one-form is closed.
  • Contracted Jacobian blocks yield K_L = J_L M_L J_L^T, so the response kernel is symmetric and positive semidefinite.The result follows by summing layer contributions weighted by local conductance.
  • The matched response factorization uses shared eigenprojectors, yielding a symmetric positive semidefinite operator H.The zero-eigenvalue contribution vanishes because the relevant projected vector lies in the kernel complement.
  • For the skew construction, P is the identity and C2 = −I on 1⊥, producing the rotational pair −α ± iΩ.A Riemannian gradient linearization has a real spectrum, so the complex pair obstructs representation as a local Riemannian gradient on the identifiable quotient.
  • Orientation averaging cancels odd terms, while rotational perturbations add Ω^2∥G∥^2 to squared first-order prediction displacement.The even second derivative contributes Ω^2H V, and Taylor expansion yields the finite-step curvature coefficient.

Appendix C: Definitions for the registered finite-horizon accounting identity

The registered accounting identity separates finite-horizon differences into orientation, step-policy, and state-drift contributions relative to a registered reference policy.

  • E_t averages the effects of the two rotation orientations, while H_t captures the realized clockwise or counterclockwise choice.These quantities are defined relative to the registered reference policy.
  • The finite-horizon comparison uses trajectory-specific notation for reciprocal and nonreciprocal policies with realized orientations and step parameters.The definitions introduce X ∈ {NR, R}, σ_t ∈ {−1, +1}, and γ_t ≥ 0.
  • The data tables and source code are contained in the accompanying reproducibility archive.A public repository identifier is slated for addition before submission.
Loading 2608.30778v1…