Source-linked AI summary

Response Renormalization for Critical Deep Equilibrium Models

Jose Luis Lima de Jesus Silva

arXiv:2608.23725v1cs.LGcs.NE

TL;DR

Near-singular residual Jacobians can amplify loss-sensitive adjoint responses and destabilize DEQ optimization. Response Renormalization selectively lifts critical denominators, and CMR and Phi-CMR preserve predictive performance while retaining uncritical response channels.

  • Problem

    Near-singular residual-Jacobian directions can make implicit DEQ gradients highly sensitive, creating a training bottleneck because conditioning can dominate parameter updates.

  • Method

    Response Renormalization modifies loss-visible near-pole responses, with CMR lifting critical denominators in a low-dimensional subspace and Phi-CMR selecting bounded positive response masses.

  • Results

    More than 98% of static and 95% of transient family-seed comparisons had CMR or Phi-CMR test errors no more than five percent higher than exact implicit differentiation.

  • Takeaways & Limitations

    Selective renormalization controls near-critical adjoint amplification without globally damping well-conditioned sensitivity or disturbing resolved modes.

  • Takeaways & Limitations

    The Phi-adaptive correction uses prescribed positive parameterizations that are not unique derivations or automatically learned corrections.

Abstract

from arXiv · show

Deep Equilibrium Models (DEQs) compute predictions from a hidden representation unchanged by the model update. Training through this equilibrium uses implicit differentiation and requires solving an adjoint system built from the residual Jacobian. If this Jacobian is nearly singular along loss-sensitive directions, small perturbations can be strongly amplified in the adjoint response, producing large, highly sensitive gradients that can make optimization unreliable. We introduce Response Renormalization, a backward-pass framework that lifts selected near-pole denominators while leaving unlifted response channels unchanged. Collective Mode Response Renormalization (CMR) applies this correction in a low-dimensional critical subspace, while Phi-adaptive CMR computes a bounded response mass from a positive susceptibility rule. We derive dense and matrix-free collective formulations, distinguish exact gradients of a modified frozen-anchor residual from backward-response surrogates, and extend the construction to Structured Implicit Layers and Vector Attractors (SILVA). Across 23 multiphysics families spanning partial differential equations, three-dimensional fields, operator maps, complex geometries, and particle systems, CMR and Phi-CMR yield test errors no more than five percent higher than those from models trained with exact implicit differentiation in more than 98% of static and 95% of transient family-seed comparisons. Solver-index experiments show convergence toward the static adjoint, while physical-time rollouts retain predictive fidelity under the evaluated conditions. These results demonstrate that selective response renormalization can control near-critical adjoint amplification without globally damping well-conditioned sensitivity. Therefore, the method can make parameter updates more reliable while preserving the useful gradient information needed for learning.

1 Introduction

DEQ training replaces backpropagation through stored solver trajectories with an adjoint sensitivity problem whose conditioning can amplify loss-relevant directions. Response Renormalization selectively lifts near-singular response channels while preserving stable and informative sensitivities.

  • DEQs compute hidden fixed points, making effective depth solver-determined rather than an explicitly stored layer stack.
  • Ill-conditioned residual-Jacobian responses can dominate parameter updates, motivating methods that control inverse response without discarding informative sensitivities.
  • Singular-value geometry captures source-to-response amplification that eigenvalues alone can miss for non-normal operators.
  • Response Renormalization identifies K^-1 as the zero-frequency response and modifies only loss-visible unstable denominators while retaining remaining response directions.
  • CMR uses a pole-selective denominator lift, while Phi-CMR and Delta-Phi provide bounded response-mass and source-conditioned variants with dense and matrix-free formulations.

2 Methods

The method treats DEQ adjoints as source-visible singular responses and selectively lifts critical denominators. It provides dense, matrix-free, collective, Phi-adaptive, and source-conditioned constructions while distinguishing exact modified-model gradients from backward surrogates.

  • The adjoint solves K^T v = g, where K = I − J is the inverse static response and singular values provide its response denominators.
  • Singular-value channels combine source projection v_i^T g, gain 1/σ_i, and output direction u_i, exposing non-normal amplification invisible to eigenvalues alone.
  • CMR marks modes with σ_i < κ as critical, lifts only loss-visible critical denominators, and leaves stable gains unchanged.
  • Dense SVDs expose the full mechanism, while matrix-free and collective formulations use low-rank singular information, Jacobian products, and GMRES.
  • Fixed-anchor equations define both an exact gradient for a modified residual and a separately evaluated surrogate obtained by applying the response correction at the original equilibrium.
  • Phi-CMR uses prescribed positive, bounded, monotone response parameterizations rather than uniquely derived or automatically learned corrections.
  • Delta-Phi adds source geometry through bounded activation, with per-mode lifting relative to Phi-CMR and collective scaling that can reduce mass toward a positive floor.

3 Results and Discussion

Controlled response tests show that CMR and Phi-CMR cap near-pole amplification while preserving stable response, and field, SILVA, and broader dataset evaluations retain predictive fidelity relative to exact implicit differentiation.

  • Experiments evaluate directional selectivity and predictive fidelity through singular-spectrum sweeps, source rotations, physical-time rollouts, and new representations.
  • CMR and Phi-CMR retain more response than Tikhonov at critical ranks one and two while maintaining higher mean directional alignment with the exact adjoint.
  • The critical denominator increases from 0.005 to 0.030 under CMR and 0.0581 under Phi-CMR, while the stable denominator remains fixed at 0.125.
  • CMR and Phi-CMR cap near-singular amplification at finite nonzero values and recover the exact response outside the critical region.
  • In the Darcy bridge, relative errors are 1.21 for Phi-CMR and Delta-Phi, 1.31 for CMR, 1.56 for Tikhonov, and 1.58 for stable-critical filtering.
  • CMR has the lowest one-step error, while Phi-CMR has lower observed error than Tikhonov in 27 of 50 paired comparisons.
  • Across 13 further datasets, the Phi-CMR error ratio is at most 1.05 in all 65 static and 57 of 60 horizon-seven comparisons, with sparse activation in every dataset.

4 Conclusion

The conclusion frames near-critical DEQ and SILVA sensitivity as a source-visible pole problem and identifies an operating regime where only loss-visible critical modes are modified. Across equations and representations, the approach preserves resolved response and predictive fidelity under the evaluated conditions.

  • CMR lifts critical denominators, Phi-CMR selects positive masses, and Delta-Phi bounds source conditioning while unlifted response channels remain unchanged.
  • Controlled source rotations and Darcy localization verify the predicted dependence on v_i^T g and distinguish backward effects from shared forward approximation error.
  • Across 23 families of PDE grids, fields, operator maps, geometries, and particle systems, the evaluations establish directional fidelity, predictive preservation, cost, and solver-index convergence.
  • The supported operating regime modifies only loss-visible poles without damping the full adjoint response or disturbing resolved modes.

Data and Software Availability

The study uses publicly available physical benchmark data and evaluates SILVA across ten physical families with shared training and evaluation protocols. Selected subsets total 120.685 GB, while dataset-specific partitions and training-only normalization define evaluation splits.

  • The physical benchmark data are publicly available, covering PDEBench, The Well, PDEArena, DynaBench, PDEGym, CFDBench, and LagrangeBench.
  • The evaluation contains ten physical families, five seeds, and 560 method evaluations across five Darcy regimes and nine additional families.
  • 120.685 GB of selected data subsets are used, excluding unused portions of the complete source collections.
  • All methods share initialization, minibatch ordering, update budgets, optimizer settings, solver settings, and tolerances, while trained parameters may differ because backward rules change updates.

A.1 All-method baselines and transient derivations

The transient analysis compares backward methods against exact implicit differentiation while separating solver-index convergence from physical-time rollout behavior. Across matched SILVA studies, Phi-CMR and CMR preserve predictive fidelity under the evaluated conditions.

  • A.1 All-method baselines and transient derivations: The transient evaluation tests causal convergence toward the static adjoint using the original J rather than a renormalized finite-frequency SILVA operator.
  • A.1 All-method baselines and transient derivations: The transient evaluation compares rollout error with exact implicit differentiation, causal-response error with the converged static adjoint, and state error with equilibrium.
  • A.1 All-method baselines and transient derivations: At horizon eight, Phi-CMR has mean error ratio 1.025, with ratios at most 1.05 in 43 of 45 family–seed pairs.
  • A.1 All-method baselines and transient derivations: The Phi-CMR mean finite-response error falls from 0.411 at N = 1 to 2.50 × 10^-4 at N = 64 as state error falls from 0.387 to 1.24 × 10^-4.
  • A.1 All-method baselines and transient derivations: The residuals lie below stated numerical tolerances, and Phi-CMR test-error ratios are at most 1.05 in 48 of 50 family–seed pairs.

B Implicit Gradients, Poles, and Baseline Filters

Implicit differentiation expresses parameter sensitivity through a shared adjoint solve, whose singular denominators can create source-visible poles. The paper compares exact, Tikhonov, and selective denominator-lifted responses, including dense and structured formulations.

  • B Implicit Gradients, Poles, and Baseline Filters: The implicit-function differential yields dL/dθa = ∂θaL + g⊤K^-1Ba, while K⊤v = g shares one adjoint solve across parameters.
  • B Implicit Gradients, Poles, and Baseline Filters: In singular coordinates, each adjoint coefficient combines the loss-source projection with the reciprocal singular value, making small denominators potential response poles.
  • B Implicit Gradients, Poles, and Baseline Filters: At an exact pole, finite adjoint solvability requires the loss source to be orthogonal to the null space; otherwise the deterministic derivative is not finite.
  • B Implicit Gradients, Poles, and Baseline Filters: CMR sets bσi = σi on stable modes and bσi = max(σi, mi) on critical modes, preserving every resolved stable channel exactly.
  • B Implicit Gradients, Poles, and Baseline Filters: A lifted adjoint can be the exact gradient of a frozen-anchor modified equilibrium or a backward-response surrogate when the original forward equilibrium is retained.
  • B Implicit Gradients, Poles, and Baseline Filters: For SILVA, Woodbury reduces the structured adjoint to local solves and a low-dimensional collective denominator, while collective singular values need not equal those of full K.

D Coordinate Metric and Threshold Stability

The response-renormalization procedures define critical modes by singular-value thresholds, select bounded target masses, and apply low-rank lifts through dense or matrix-free solves. Diagnostics track residuals, rank coverage, solver behavior, and the distinction between original and modified objectives.

  • D Coordinate Metric and Threshold Stability: A declared metric H makes the construction coordinate-consistent, while subspace selection stabilizes thresholds within near-degenerate singular clusters.
  • D Coordinate Metric and Threshold Stability: The residual Jacobian is K = I − J, and critical modes satisfy σi < κ while loss visibility requires vi⊤Ba ≠ 0.
  • D Coordinate Metric and Threshold Stability: The dense and matrix-free procedures return gated adjoints with effective denominators, residuals, solver diagnostics, and exact recovery when the active set is empty.
  • D Coordinate Metric and Threshold Stability: The modified-forward route differentiates a frozen-anchor modified equilibrium, whereas retaining the original equilibrium with a lifted adjoint is explicitly labeled a surrogate.
  • D Coordinate Metric and Threshold Stability: CMR applies ΔK = UC diag(bσi − σi)VC⊤, lifting only selected denominators while leaving unselected modes unchanged.
  • D Coordinate Metric and Threshold Stability: The structured SILVA procedure decomposes only the small collective denominator Γ and reduces the full adjoint to local solves plus an r-dimensional collective solve.

G Phi and Source-Conditioned Finite Response

The paper constructs finite critical responses by lifting selected denominators, then adapts the lift using pole proximity and loss-source visibility while preserving stable or unselected channels. The Hartree calculation motivates positive response masses, while constrained optimization and matrix-free formulations specify implementation and bounds.

  • Finite-response construction: The renormalization condition targets a positive finite static susceptibility, replacing a vanishing scalar denominator with a prescribed response mass.For scalar response G(0)=m^-1, the target is G(0)=χR>0.
  • Finite-response construction: A nonnegative lift leaves a singular channel unchanged when its correction is zero, and an empty active set recovers the exact implicit adjoint.The lifted response modifies only selected denominators rather than every response direction.
  • Phi-CMR: The target susceptibility acts as an upper bound: already-bounded channels remain unchanged, while over-responsive channels are reduced to the finite target.This selective rule controls amplification without globally damping well-conditioned response directions.
  • Phi-CMR: Phi-CMR prescribes a positive, bounded, monotone mass that acts most strongly near small denominators and is applied only when σi < κ and mΦ > σi.The threshold identifies candidate critical channels, while the positive-part rule determines whether a nonzero correction occurs.
  • Hartree motivation and scope: The Hartree closure motivates a positive response mass through self-consistent renormalization, but does not identify DEQ parameters or uniquely derive the Phi-CMR law.The paper explicitly leaves g4, D, m0, κ, and αmax undetermined by this calculation.
  • Source-conditioned extension: Source-conditioned Delta-Phi raises the target mass when source visibility is high and pole pressure is positive, but lowers it relative to Phi-CMR when visibility is low.Projection keeps the mass within the positive floor and prescribed upper cap; the adjustment vanishes at zero pole pressure or half source visibility.

H Metrics, Protocol, and Quantitative Evidence

The evaluation measures response fidelity, selective activation, computational cost, and residual consistency for candidate adjoints and collective lifts. Controlled sweeps examine rank, source alignment, and Delta-Phi gate behavior.

  • Metrics: Candidate adjoints are evaluated using amplitude, direction, projection, error, and residual metrics against the exact adjoint.Stable and critical amplitudes are computed after projection onto the corresponding subspaces.
  • Metrics: Field errors are inverse-standardized sample-relative Euclidean scores, while the 1.05 count records descriptive threshold crossings rather than noninferiority.Mixed-unit Euclidean scores are descriptive and not physical invariants.
  • Protocol: SILVA checks both the lifted collective solve residual and the residual against the original full operator, with local solves checked separately.These checks prevent local inverse inaccuracies from being hidden by aggregate residuals.
  • Controlled evidence: Controlled experiments vary collective rank, source alignment, Delta-Phi, and gate activation to assess retention, suppression, bounded mass, and strength utility.The supplied figure descriptions identify rank-two advantage, source-alignment retention, Delta-Phi sweeps, and gate ablations as the tested quantities.

J Cross-Benchmark Evaluation of SILVA Response Renormalization

The SILVA cross-benchmark protocol evaluates response renormalization across diverse physical systems and representations using matched static and transient comparisons. Phi-CMR and CMR remain near exact-implicit error levels, while lifts activate sparsely and transient responses converge toward the static adjoint.

  • Protocol: 13 additional public datasets extend SILVA evaluation across PDEs, operator maps, geometry-conditioned flow, and particle systems.The extension includes four PDEArena, six DynaBench, PDEGym, CFDBench, and LagrangeBench systems.
  • Static and transient results: Phi-CMR has mean error ratios 0.9999 statically and 1.0016 at horizon seven relative to exact implicit differentiation.Ratios at most 1.05 occur in all 65 static and 57 of 60 horizon-seven paired evaluations.
  • Static and transient results: Phi-CMR has ratios at most 1.05 in all 65 static and 57 of 60 horizon-seven paired evaluations.CMR gives mean ratios 0.9989 statically and 0.9978 at horizon seven under the same benchmark protocol.
  • Selectivity: Collective lifts occur at least once in every dataset but remain sparse over updates, supporting selective rather than globally active correction.Figure 10 separates dataset-level error, Phi-CMR–Tikhonov differences, lift frequency, and measured cost.
  • Transient evidence: At solver horizon N = 64, Phi-CMR causal-response error is 5.27 × 10^-4 relative to the static adjoint and state error is 1.49 × 10^-4 relative to equilibrium.These measurements cover 60 trained models while physical-time rollouts propagate prediction errors over seven PDE steps.

J.1 Forward-capacity extension for the highest-error families

A forward-capacity extension tests whether large errors in selected families arise from representation limits rather than response renormalization. Higher-capacity models improve prediction geometry, while a difficult Diffusion–Reaction time window remains a distinct stress test.

  • Experimental design: The extension keeps held-out trajectories, solver settings, backward methods, response parameters, and five seeds fixed while increasing forward capacity.Profiles selected on seed 123 remain fixed during the subsequent five-seed, eight-method evaluation, making the extension exploratory.
  • Capacity results: Phi-CMR improves in all 30 family–seed comparisons, with the largest reductions for Darcy and LagrangeBench.The matched compact benchmark remains the reference for comparing backward response rules.
  • Prediction geometry: Capacity changes Navier–Stokes correlation from 0.704 to 0.786 and amplitude ratio from 0.646 to 0.830 on the same held-out example.The geometry analysis separates shape recovery from amplitude calibration rather than relying on one relative-error number.
  • Prediction geometry: Corrected Maxwell Conv3D changes correlation from 0.248 to 0.666 and amplitude ratio from 0.566 to 0.906, while LagrangeBench reaches correlation 0.996.Darcy reaches correlation 0.980 but overpredicts amplitude on the first example.
  • Stress test: The Diffusion–Reaction t1 →t2 task reaches mean Phi-CMR error 0.2633, correlation 0.964, and amplitude ratio 0.972, but is excluded from the paired capacity ratio.It is a separate physical-time task from the difficult t0 →t1 stress test.

K Scope and Validity Conditions

The method’s validity depends on converged equilibria, accurate local operations, resolved critical subspaces, and informative source and parameter projections. The reported scope is limited to response fidelity and predictive behavior under matched evaluated conditions, not broad stability or physical-invariant guarantees.

  • Validity boundary: Exactness depends on whether the method modifies a frozen-anchor residual or only applies a lifted backward response at the original equilibrium.The latter is a biased surrogate, whereas no-lift cases recover the exact implicit adjoint.
  • Subspace assumptions: The dense rule requires a resolved critical singular subspace, with repeated clusters identifiable as subspaces rather than individual singular vectors.The paper recommends an unresolved interval [κ − η, κ + η] and principal-angle tracking for repeated clusters.
  • SILVA assumptions: SILVA modifies singular values of a collective matrix that need not match the full residual Jacobian, requiring accurate local solves and an informative low-rank factorization.Factorization diagnostics must account for gauge dependence of raw factor norms.
  • Scope: Neither dense nor collective renormalization repairs poor forward convergence, architecture mismatch, unresolved subspaces, or rapidly changing modes.These are explicit boundaries on what the response correction addresses.
  • Evaluation limits: The 1.05 threshold is descriptive, field errors are not conservation or PDE-residual errors, and governing-equation and invariant claims require dedicated evaluations.Quantitative conclusions therefore concern response fidelity, selective activation, convergence, and cost at the evaluated sizes.
  • Forward conditions: The forward map is contractive with ∥J∥2 ≤ 0.995 and σmin(I − J) ≥ 0.005 while still inducing near-critical response.Projection after each update enforces the spectral-norm bound.
Loading 2608.23725v1…