Source-linked AI summary
The Loss Floor of Denoising Score Matching: Fisher Geometry from Schrödinger Bridges
Avinash Raju, Kai Zhang
TL;DR
Denoising score matching reaches the correct marginal-score minimizer through a random conditional target, but the target’s variance creates an irreducible training-loss floor. The paper derives this floor from Schrödinger bridge relative entropy, identifies it exactly with integrated Fisher–Rao geometry, and connects it for diffusion corruptions to information flow, schedule weighting, loss comparability, and high-SNR behavior.
Problem
Denoising score matching uses a conditional score instead of the marginal score required for generation, leaving the geometric and quantitative meaning of the resulting excess loss to be characterized.
Method
The paper derives the ideal score objective from a Schrödinger bridge variational principle and decomposes conditional-score regression through a regular conditional endpoint family.
Results
The excess is exactly the integrated trace of the conditional endpoint family’s Fisher–Rao metric, and for corruption diffusions it factors into data-dependent information flow and schedule-dependent weighting.
Takeaways & Limitations
The floor is intrinsic to conditional denoising, so raw losses across noise ranges or weightings need not rank models consistently and should be interpreted geometrically and informationally.
Takeaways & Limitations
Closed-form evaluation as mutual-information differences assumes affine Gaussian corruption, while general corruption diffusions leave the information flow implicit and high-SNR rates may lack an information dimension.
Abstract
from arXiv · showhide
Denoising score matching trains diffusion models by regressing onto a conditional score, although generation ultimately requires the marginal score. The two objectives share the same population minimizer, but the conditional target remains random at fixed noisy state and introduces an irreducible excess in the training loss. We isolate this excess and show that, for a general corruption kernel under mild regularity assumptions, it is exactly the trace of the Fisher--Rao metric of the conditional endpoint family, integrated along the diffusion trajectory. This gives an exact conditional-variance decomposition of the denoising objective and identifies the information geometry observed in diffusion latent spaces as an intrinsic component of the training loss. We derive the result from a Schr"odinger bridge variational principle, in which the ideal objective arises as excess path-space relative entropy. For corruption diffusions, the Fisher term is proportional to the rate at which the noisy state loses mutual information about the clean data, separating the loss floor into an information flow determined by the data and a weight determined by the corruption schedule and objective. In the Gaussian case, this yields a closed form for the floor, recovers reparametrization invariance of the continuous-time objective, and relates its high-SNR divergence to the information dimension of the data. Finally, we show that raw losses obtained with different noise ranges or weightings need not rank models consistently because they contain different additive floors, and contrast the second-order geometry seen by training with the third-order conditional statistics entering numerical sampling error.
1 Introduction
The paper identifies the irreducible excess in denoising score matching as an exact Fisher–Rao geometric term and derives it through a Schrödinger bridge variational framework. For diffusion corruptions, this floor decomposes into information flow and schedule-dependent weighting, with practical consequences for loss comparison.
- Motivation: Denoising score matching uses a random conditional score whose population mean is the marginal score required for generation, so both objectives share a minimizer despite an irreducible excess.The conditional target remains random at fixed noisy state.
- Geometric characterization: The excess equals the integrated trace of the Fisher–Rao metric tensor of the conditional endpoint family.The conditional-score fluctuation is the endpoint-family score, making its second moment Fisher information.
- Geometric characterization: The Fisher geometry observed in diffusion latent spaces is an intrinsic variance term in the training loss, not an externally imposed model structure.The additive term depends only on the corruption and data, never on the model.
- Information-theoretic evaluation: For corruption diffusions, the floor separates into data-dependent information flow and a schedule-dependent weight, telescoping to mutual-information differences under matched weighting.In the Gaussian case, the corresponding identity involves differential entropy through the I–MMSE relation.
- Practical consequences: Different SNR ranges, schedules, or loss weightings can produce incomparable raw training losses because their additive floors differ, potentially reversing model rankings.Subtracting the floor repairs the exhibited ranking inversion.
2 From Path-Space Entropy to Denoising Score Matching
The paper derives diffusion score matching from Schrödinger bridge relative entropy and reduces the resulting ideal marginal-score objective to a tractable conditional regression objective. The conditional target preserves the optimizer while adding a model-independent Jensen gap.
- Bridge variational principle: The Schrödinger bridge selects an endpoint-constrained path measure by relative-entropy minimization, with optimal dynamics represented by a Doob h-transform.The reference path measure need not itself satisfy the endpoint constraints.
- From path entropy to score matching: The bridge’s path-space relative entropy decomposes into local transition divergences that become a quadratic score objective in the continuum limit.The short-time expansion preserves the diffusion coefficient and shifts the drift.
- Optimal bridge dynamics: In the score-matching sector, the optimal drift is v*(x,t) = v(x,t) + γ(t)∇ln Pt(x), so the learned score is the control field enforcing the terminal data constraint.The marginal score is the ideal regression target but is not available in closed form.
- Conditional regression: Superposition of conditional bridge fields identifies each conditional field with the corruption kernel, enabling a simulation-free objective based on conditional scores.The objective replaces the computationally prohibitive on-policy weighting with the available reference noising marginal.
- Scope of the derivation: Replacing on-policy weighting affects the projection of a restricted parametric model, even though the unrestricted functional minimizer remains unchanged.This distinction limits the weighting-invariance claim to objective values at the functional level.
- Conditional regression: The tractable denoising objective is an upper bound whose Jensen gap is an additive, model-independent term, while unbiasedness leaves its population minimizer unchanged.For affine Gaussian corruption, the conditional score is linear in the noise and recovers standard denoising score matching.
3 The Fisher Geometry of the Denoising Loss
The paper derives a regular Fisher geometry for conditional endpoint distributions from the Schrödinger bridge and identifies its integrated trace with the denoising loss floor. This geometry is the future sector of a richer latent metric and quantitatively tracks information lost under corruption.
- 3.1–3.2 Endpoint and conditional geometry: The Schrödinger bridge’s second variation defines a canonical endpoint-based quadratic form, while regularity motivates reducing it to the conditional endpoint family indexed by the noisy state.The direct conditioning of full paths is singular, so the paper uses the endpoint experiment and then conditions it on the latent state.
- 3.4 The pullback metric: In isotropic Gaussian corruption, latent spacetime has d+1 natural-parameter components, and monotone time reparametrizations move along the same SNR ray without changing its geometry.The additional coordinate corresponds to the second-moment statistic of the corruption.
- 3.3–3.4 Pullback metric: The conditional endpoint family forms a Fisher–Rao metric whose future sector is the geometry probed by denoising score matching.The full latent metric contains past and future sectors; the past sector is generally nonzero but is invisible to the clean-endpoint regression target.
- 3.4 The pullback metric: Along corruption, the radial Fisher component governs the decay of information about the data, linking the pointwise metric to the integrated loss-floor identity.The radial coordinate is the signal-to-noise ratio, and the information hierarchy is constrained by the data-processing inequality.
- 3.5 The irreducible loss floor: The conditional score is an unbiased estimator of the marginal score, so the cross term in the objective decomposition vanishes for every model.Bayes’ rule centers the conditional score under the posterior, yielding the unbiasedness identity.
- 3.5 The irreducible loss floor: The denoising objective decomposes into the ideal marginal-score error plus an irreducible model-independent floor equal to the integrated trace of the conditional Fisher information.The floor is the conditional variance of the random regression target and therefore cannot depend on model parameters.
4 The Loss Floor as Information Flow
The loss floor is locally a Fisher–Rao metric and globally an information-flow integral along corruption diffusions. Its data-dependent information flow is separated from schedule and weighting effects, yielding invariance, asymptotic, and comparison consequences.
- Local Fisher geometry: The instantaneous floor equals the Fisher information lost when endpoint-conditioned corruption kernels are mixed over the data distribution.It is the gap between Fisher information with the clean endpoint known and after marginalizing that endpoint.
- Information-flow factorization: For corruption diffusions, the accumulated floor factors into data-dependent information flow multiplied by a schedule-dependent weight.The general factorization holds for arbitrary loss weighting, while the data enters through the information-loss rate and the schedule through w/D.
- Information-flow factorization: When w = D, the floor equals I(y; x_t0) − I(y; x_t1), without requiring Gaussian, affine, or SNR assumptions.This identifies the floor with mutual information lost about the clean endpoint under corruption.
- Schedule and weighting: The information spectrum S captures data structure, with mixture features appearing near λ ∼ m^-2 and λ ∼ s^-2 and saturating at the information dimension.Equal-information schedule allocation therefore differs from uniform-in-log λ for structured data, which has bumps at structural scales.
- Gaussian corruption and invariance: Under the bridge schedule, the floor depends on the SNR path only through its endpoints, so reparametrization changes error distribution but not the total floor.The schedule is a parametrization of a fixed information-manifold curve, and the floor is an endpoint-pinned line integral.
- High-SNR asymptotics: For continuous data, the floor diverges at high SNR; its logarithmic rate is governed generally by information dimension, while Gaussian rank-k data has slope k/2.The divergence reflects unbounded mutual information from increasingly precise observation and is finite in practice only after SNR truncation.
- Thermodynamic interpretation: The entropy interpretation should not be identified with stochastic-thermodynamic entropy production except under the Gaussian-channel reading in SNR parametrization.The former is a state differential-entropy change, whereas the latter is a path-space quantity with system and environment contributions.
5 Consequences for Training and Sampling
The training objective probes second-order posterior geometry, whereas numerical sampling also depends on third-order conditional statistics. The loss floor constrains what schedule changes can accomplish and makes raw losses incomparable across configurations.
- Training versus sampling: Training uses the denoising posterior’s second cumulant, while sampler discretization error also involves its third cumulant.Training samples noise levels independently and does not discretize a trajectory; higher cumulants appear when integrating the sampler.
- Loss comparability: Raw losses can reverse model rankings across SNR ranges because different configurations integrate different additive floors.A better model on the wider range reports 1.5455 versus 1.1267 for a worse model on the narrower range; floor subtraction restores correct ordering.
- Schedule design: The total floor is fixed by the SNR-range endpoints, so schedules redistribute estimation error rather than changing the total floor.Equal-information allocation recovers the criterion used by entropic time schedulers; proportional online sampling is recorded as an open question.
- Training versus sampling: The leading first-order sampling truncation term contains the third cumulant through T[u, u] and the second-order term g du/dλ.On the analytic mixture, the third-cumulant contribution is strongly localized near mode separation and has median share 0.47 across trajectories.
6 Discussion
The paper identifies the denoising loss floor as integrated Fisher–Rao geometry and factors it into information flow and schedule-dependent weighting. It also extends the decomposition to masked diffusion while delimiting assumptions and empirical scope.
- Scope and prior work: The paper claims novelty in assembling classical ingredients into a bridge-based derivation of the latent metric, rather than introducing a new optimizer or schedule family.Its contribution is the factorization connecting the irreducible term to Fisher–Rao geometry and entropy change.
- Discrete diffusion: The same conditional-variance split applies to masked-diffusion cross-entropy, whose floor depends on endpoint masking rates and data entropy, not schedule shape.This was verified numerically for correlated tokens.
- Limitations: Closed-form floor evaluation assumes affine Gaussian corruption, while general corruption diffusions leave the information flow implicit.The high-SNR rate may lack definition for singular data distributions, and experiments are analytic or low-dimensional.
- Core conclusions: The denoising objective splits exactly into model-dependent estimation error and a model-independent integrated Fisher–Rao loss floor.The floor is the trace of the conditional endpoint family’s Fisher–Rao metric integrated along the flow.
- Core conclusions: For corruption diffusions, the floor factorizes into data-dependent information flow and schedule-dependent weight.The primary factorization is stated as Theorem 4.3.
Ethics Statement
The authors disclose using large language models for language improvement, editing, literature search, and summarization, while retaining responsibility for the manuscript.
- Use of language models: Large language models assisted with language improvement, editing, literature search, and summarization, not research ideas, derivations, or results.The authors state that every suggested change was reviewed and verified.
- Author responsibility: The authors take full responsibility for the paper’s content and any remaining mistakes.
A Variational Derivation of the Endpoint-Tilted Path Measure
The appendix derives the Schrödinger bridge by path-space relative-entropy minimization, obtains endpoint-tilted forward and backward fields, and connects the resulting optimal process to diffusion score matching.
- Bridge formulation: The bridge problem minimizes path-space relative entropy subject to prescribed endpoint marginals.The appendix explicitly avoids requiring an equivalence between marginal-propagator and path-space divergences.
- Bridge formulation: Functional variation yields an optimal path measure that factorizes into a reference path measure and endpoint functions.The endpoint multipliers have a global rescaling symmetry fixed by enforcing the endpoint marginals.
- Forward and backward fields: Forward and backward auxiliary fields obey dual recursions, with their product recovering the intermediate marginal.In the continuum limit they become dual backward Kolmogorov and forward Fokker–Planck equations.
- Optimal dynamics: The optimal process preserves the reference diffusion coefficient and adds the gradient of the log-backward field to its drift.This tilted-kernel result is the central dynamical relation of the bridge.
- Connection to score matching: In the score-matching sector, the marginal equals the conditional field average, yielding the posterior-weighted reverse kernel.The generative process averages conditional noising posteriors over the clean-data posterior given the noisy state.
- Gaussian specialization: For affine Gaussian diffusion schedules, the marginal is the forward noising marginal obtained by averaging the Gaussian reference kernel over the data distribution.
- Excess action: The per-step KL divergence between equal-diffusion processes reduces, in the refined discretization limit, to the squared drift difference underlying score matching.Higher-order drift terms do not affect the leading-order divergence.
- Endpoint sufficiency: Endpoint-measurable scores preserve the local statistical experiment relevant to Fisher information when passing from paths to endpoint pairs.The appendix establishes this through the measurability of tangent vectors and Fisher-information monotonicity.
B.3 Singularity of exact path conditioning
Exact conditioning of a continuous-path measure on an interior state produces mutually singular conditional laws, so no local Fisher expansion exists at the path level. The paper therefore extracts geometry from the conditioned endpoint experiment, where Markov structure separates past and future contributions.
- Singularity of exact path conditioning: Distinct interior-state conditionals are mutually singular because they occupy disjoint path sets, preventing any finite local Kullback–Leibler expansion.The obstruction is generic to exact point conditioning of continuous-path measures, not specific to Schrödinger bridges.
- Gaussian geometry: For anisotropic Gaussian corruption, the sufficient statistic requires a full symmetric tensor, making the family curved rather than an open exponential-family chart.Isotropy is exactly the case in which the natural-parameter dimension equals the latent-manifold dimension.
- Endpoint experiment: The endpoint conditional law forms an exponential family whose Fisher form in natural coordinates is the covariance of sufficient statistics.Composing the natural-parameter map with the latent coordinates yields the latent Fisher metric.
- Endpoint experiment: Conditioning on the intermediate state makes bridge past and future independent, so their score cross-covariance vanishes.This cancellation follows from the Markov factorization of the endpoint conditional distribution.
- Endpoint experiment: The unconditioned endpoint sectors retain a generically nonzero covariance, while denoising score matching observes only the clean-endpoint target.Thus the seed-posterior Fisher information is invisible to the denoising regression target.
B.6 Gaussian evaluation
The Gaussian conditional endpoint law admits a Hessian-form evaluation of the Fisher metric, with the posterior covariance of the clean data determining the resulting expression.
- Gaussian evaluation: The Gaussian marginal is obtained by integrating the corruption kernel q(x,t|xN) over the data distribution.Differentiating this representation yields the Tweedie relation and supports the Hessian-form calculation.
- Gaussian evaluation: The Gaussian Fisher metric equals the posterior second moment of the corruption score, which is the covariance of that score under the posterior.Using the posterior-score identity identifies this covariance with the future-sector metric.
B.7 The radial information identity
Along a fixed natural-parameter ray, the radial coordinate evolves according to the radial Fisher metric, linking monotone information loss to decreasing signal-to-noise ratio.
- The radial information identity: Along θ = r n with fixed normalized n, the radial derivative of the relative-entropy quantity is dr = r g_rr.The radial Fisher component controls the rate of change along the ray.
- The radial information identity: Because SNR_t is the radial coordinate, corruption moves along the ray while the associated quantity decreases monotonically as SNR_t falls.At zero SNR, the conditional endpoint law equals the carrier and the latent contains no information about the data.
C Proofs of the Information-Flow Identities
The information-flow identities follow by averaging score norms, applying Fokker–Planck dissipation, and converting the resulting metric trace into mutual-information or I–MMSE relations.
- Information-flow identities: Averaging conditional score norms over noisy states and data reduces the expression to the expected Fisher information of the corruption kernel.The tower property exchanges conditioning order in the averaged score norm.
- Information-flow identities: Fokker–Planck dissipation expresses entropy derivatives through Fisher information, and drift cancellation converts the difference into mutual-information flow.The cancellation uses the tower property for the conditional corruption laws.
- Gaussian specialization: In the Gaussian case, the Fisher integrand is reparameterized by SNR and the integral becomes an MMSE integral linked to mutual information by I–MMSE.The conditional entropy term is independent of the SNR parameter, so entropy differences also represent the information difference.
D Numerical and Experimental Details
The appendix reports analytic and low-dimensional numerical checks designed to verify the paper’s identities rather than benchmark performance. Across nonlinear, Gaussian-mixture, interpolant, model-ranking, trajectory, and masking examples, the reported computations closely reproduce the stated theoretical results.
- Analytic and low-dimensional experiments compute quantities exactly or by quadrature to verify identities rather than performance.The appendix explicitly frames all experiments as numerical verifications of theoretical statements.
- The nonlinear corruption test matches Eq. (4.2) with 1.7 × 10−3 median point-wise relative error and 0.4% integrated relative error.The Fokker–Planck equation was solved directly for f(x) = −ax3.
- Eq. (4.5) is numerically verified to 3 × 10−6 relative error across four distinct interpolants.The linear interpolant is evaluated analytically, while trigonometric and polynomial interpolants follow the same computation.
- For a two-mode Gaussian mixture, the accumulated floor agrees with the entropy change to 2 × 10−5 relative error.The comparison uses quadrature-computable differential entropy and accumulated MMSE.
- For rank-k Gaussian data embedded in R10, fitted floor-divergence slopes match k/2 to three decimals, confirming Eq. (4.10).
- Using different SNR ranges changes additive floors and raw losses, while floor-subtracted excesses preserve the better-versus-worse model ordering.On range A, raw losses are 1.0490 and 1.1267 with floors 1.0414; on range B, raw losses are 1.5455 and 1.8601 with floors 1.5143.