Source-linked AI summary

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz

arXiv:2608.23938v1stat.MLcs.LGmath.ST

TL;DR

The paper asks how diffusion models generalize despite finite-sample objectives whose exact minimizers memorize training data. It studies lazy denoising score matching in an inner-product-kernel RKHS under proportional high-dimensional scaling, deriving gradient-flow trajectories and their reverse-SDE behavior. The results separate interpolation from memorization and show that the lazy model’s generated law is asymptotically Gaussian, limiting non-Gaussian generation without feature learning, architectural structure, or low-dimensional data.

  • Problem

    Diffusion models can generate novel samples despite reducing distribution learning to regression tasks whose exact empirical optima memorize the training set.

  • Method

    The paper analyzes vector-valued RKHS denoising score matching with an inner-product kernel in the proportional regime n ≍d, using gradient flow and pathwise reverse-SDE analysis.

  • Results

    Overfitting does not imply memorization: localized nonlinear corrections can interpolate the denoising objective while remaining invisible to typical reverse trajectories, whose generated law is asymptotically Gaussian.

  • Takeaways & Limitations

    Lazy diffusion models separate sampling generalization, objective overfitting, and memorization, but retain only limited distributional information in proportional high dimensions.

  • Takeaways & Limitations

    In the proportional lazy regime, isotropic kernels generate at most Gaussian distributions; non-Gaussian generation requires feature learning, architectural structure, or genuinely low-dimensional data.

Abstract

from arXiv · show

Modern score-based generative models have achieved remarkable empirical success in high-dimensional tasks such as image, audio, and video synthesis. These models reduce distribution learning to a sequence of regression problems that, if solved exactly on finite data, would ultimately reproduce the training samples. Their ability to generalize must therefore arise from the implicit or explicit regularization during training. In this work, we develop a generative counterpart to the theory of benign overfitting and algorithmic regularization for overparameterized neural networks in the supervised lazy-training regime. We study denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel. In the proportional high-dimensional regime $n\asymp d$, we derive exact risk trajectories under gradient flow training. These trajectories exhibit three phases governed by qualitatively distinct estimators: a spectral estimator that generalizes, a pure-noise score with localized peaks that interpolate the training objective, and an empirical Bayes estimator that memorizes the data. We then analyze how these estimators combine along the reverse-time SDE and characterize the distribution of the resulting samples. The analysis reveals familiar mechanisms from supervised learning, including kernel linearization and self-induced regularization from the nonlinear part of the kernel, but also reveals a distinct phenomenology specific to generative modeling.

1 Introduction

Diffusion models can overfit finite-sample denoising objectives without immediately memorizing training data, but their high-dimensional lazy-training behavior is constrained by kernel-induced regularization and a Gaussian generation barrier.

  • Motivation: High-dimensional diffusion models reduce generation to regression problems, yet exact empirical optimization leads to an empirical Bayes denoiser that retrieves training samples.This creates a central tension between empirical-objective minimization and the observed generation of novel samples.
  • Limitations: In proportional high dimensions, isotropic lazy kernels have sharply limited expressive power: they learn at most polynomial structure and generate at most Gaussian approximations.Non-Gaussian generation requires feature learning, architectural structure, or genuinely low-dimensional data.
  • Distinct generative phenomenology: Denoising score matching differs from supervised regression because its empirical target is the memorizing empirical Bayes denoiser, residual noise is exponentially small, and interpolation is not benign.Generalization therefore depends on the gap between the trained estimator and the exact empirical minimizer.
  • Contributions: The paper characterizes score trajectories, training and test errors, and generated distributions through kernel linearization and reverse-SDE analysis in the proportional regime.Its framework extends kernel theory to an infinite-dimensional RKHS with operator-level linearization and non-ridge self-induced regularization.
  • Kernel mechanisms: The RKHS kernel model interpolates through localized nonlinear bumps while its linear component captures global structure and generalizes at typical test points.High-dimensional near-orthogonality makes the nonlinear correction effectively invisible away from training neighborhoods.
  • Three phenomena: The analysis separates overfitting, memorization, and population generalization: vanishing training loss need not imply memorization, but it degrades population-level generalization.The training objective can be interpolated while generated samples remain novel, because localized bumps are invisible to typical reverse trajectories.

2 Kernel denoising score matching

This section formulates diffusion denoising score matching as fixed-time regression in a vector-valued RKHS, then analyzes its empirical optimum, kernel operators, and gradient-flow dynamics. Exact empirical optimization can memorize training data, while kernel structure and training dynamics determine how fitting progresses toward population behavior.

  • 2.1 Denoising score matching: Diffusion models estimate scores through a sequence of denoising regression problems at different noise levels, whose solutions are composed by the reverse-time SDE.The analysis treats each noise level separately, although practical models often couple them through a shared network.
  • 2.1 Denoising score matching: The empirical objective replaces the unknown data distribution with the training empirical measure, producing a Gaussian-mixture noisy marginal and an unrestricted Bayes-denoising optimum.The score is learned by regressing on injected Gaussian noise, with Tweedie’s formula identifying the population minimizer as a rescaled Bayes denoiser.
  • 2.2 The empirical objective and its memorizing optimum: As the noise level vanishes, the empirical Bayes denoiser drives reverse-time dynamics toward nearest-neighbor retrieval, so exact empirical minimization memorizes rather than generalizes.This memorizing optimum exists in sufficiently rich function classes, including universal RKHSs.
  • 2.3 The RKHS model class: The model class estimates each denoiser coordinate in the product RKHS of a positive-definite inner-product kernel, which is expressive enough to represent the empirical optimum.Inner-product kernels are studied because they describe wide neural networks in the lazy-training regime.
  • 2.4 Kernel operators, gradient flow, and ridge regression: The empirical kernel operator acts on the infinite-dimensional noisy empirical space, while the train-to-population transfer operator evaluates fitted quantities on the population space.These operators provide the bridge between training-space risks and test-time risks.
  • 2.4 Kernel operators, gradient flow, and ridge regression: In high dimensions, the irreducible denoising noise is exponentially small, making the empirical DSM problem effectively noiseless compared with noisy supervised regression.The risk decomposition separates this intrinsic noise from the reducible error that optimization can remove.
  • 2.4 Kernel operators, gradient flow, and ridge regression: Gradient flow applies the spectral filter λ 7→1 −e−tλ to empirical Bayes components, fitting kernel eigencomponents on timescales set by 1/λ.The resulting risk trajectory is governed by the empirical kernel spectrum and the empirical Bayes denoiser’s alignment with its eigenvectors.

3 Learning the score along the gradient flow trajectory

Gradient flow separates score learning into generalization, overfitting, and memorization phases. Kernel linearization explains the global predictor and localized corrections, while nonlinear curvature induces sample-wise regularization and controls the transition between phases.

  • 3 Learning the score along the gradient flow trajectory: Gradient flow exhibits generalization, overfitting, and memorization phases at respectively near-polynomial, polynomial, and super-polynomial training times.The first phase behaves like a regularized linear model; the second interpolates training loss with pure-noise test behavior; the third reaches the empirical Bayes denoiser.
  • 3.1 Linearization of the kernel and localized bumps: Kernel linearization holds for delocalized noisy input pairs, while full nonlinear behavior survives on inputs sharing the same training sample.This separates a global linear component from localized nonlinear corrections around training data.
  • 3.1 Linearization of the kernel and localized bumps: The nonlinear kernel contribution adds conditional noise-averaging regularization rather than an identity ridge penalty on the empirical Gaussian space.It acts through f(x0,i, z) → E_z′[f(x0,i, z′)] and modifies the effective operator at each sample.
  • 3.2 Phase I: linearized dynamics and generalization: The generalization window ends sooner for stronger kernel curvature because localized bumps relax faster than the global covariance-learning component.Global directions learn on timescale approximately d, whereas localized corrections fit on timescale approximately δ_n^-1, producing two-stage dynamics.

4 Generation: the implicit bias of the reverse process

In the early-stopped regime, the reverse process effectively ignores nonlinear overfitting and is governed by a linearized Gaussian process. Its generated distribution is approximately Gaussian, with covariance characterized through the empirical covariance and training schedule.

  • The reverse process is analyzed directly rather than through score-error bounds, because the score error remains of constant order.The analysis focuses on early-stopped, non-memorized gradient flow and composes the learned denoisers through the reverse SDE.
  • The linearized reverse process is conditionally Gaussian because it has linear drift and Gaussian initialization.Its score has the form eS_t(x) = eΨ_t x/σ_t for a bounded negative semidefinite matrix eΨ_t.
  • The learned score and its linearization differ substantially only inside thin tubes around training points where kernel nonlinearity is active.Along the linearized trajectory, overlaps with all training directions are of order d^-1/2, establishing delocalization.
  • With high probability, the sampling trajectory remains nearly orthogonal to every training sample and never visits the memorizing bumps.Consequently, the nonlinear score and its linear surrogate are indistinguishable along the sampling path.
  • The conditional KL divergence between empirical and linearized path laws converges to zero in probability, making nonlinear overfitting invisible to the sampler.The generated distribution is effectively Gaussian, with covariance determined by the linearized score.
  • The generated samples are approximately Gaussian with limiting covariance given by a deterministic equivalent derived from the sample covariance and training schedule.The covariance flow diagonalizes in the sample-covariance eigenbasis and reduces to scalar dynamics along its eigenvalues.
  • An early-stopping schedule that is suboptimal for DSM can nevertheless yield slightly smaller KL divergence when the reverse process is stopped early.The paper identifies the relationship between score-matching performance and generation quality as a future research direction.

5 Conclusion and future directions

The conclusion separates generalization, overfitting, and memorization in lazy diffusion training, while identifying scope boundaries and directions for extending the theory.

  • Future directions: At proportional sample sizes, the learned score captures non-Gaussian structure through feature learning or architectural inductive bias, which this lazy analysis does not characterize.Coupling feature learning to reverse-process analysis is proposed as a way to determine which features the sampler encounters.
  • The training trajectory passes through a generalizing spectral estimator, a pure-noise interpolant, and an empirical Bayes estimator that memorizes the data.These phases are separated by distinct optimization timescales and statistical behaviors.
  • Overfitting does not imply memorization: the model can interpolate its training objective while still generating genuinely new samples.The separation between overfitting and memorization is exponentially large in training time.
  • Future directions: The analysis is restricted to n ≍ d, motivating extensions that could yield a hierarchy of higher-order, non-Gaussian corrections as sample scaling increases.The proposed hierarchy would progressively capture higher-order moments of π0.
  • Future directions: The assumptions exclude data supported on or concentrated near low-dimensional manifolds, where memorization and generalization interact with support geometry.Extending the analysis to structured data is identified as a natural next step.
  • Future directions: The theory trains independent RKHS models at each noise level, whereas practical diffusion models use one time-conditioned network that couples these learning problems.Understanding the effect of parameter sharing remains an important direction.

A.1 Moderate regularization

The moderate-regularization analysis replaces the nonlinear kernel estimator with an effective linearized model and expresses its train and test risks through deterministic equivalents.

  • For κ = Ω_d(d^-c), the ridge estimator is stable and the kernel linearization applies directly.The stated bounds are established under the paper’s assumptions and use the effective kernel representation.
  • The deterministic equivalents for training and test errors are explicit functions of the spectrum of Σ and can be computed from a scalar fixed-point equation.They are equivalent to the risks of the linearized effective kernel and RKHS.
  • The nonlinear kernel affects the linearized asymptotics through the self-induced shift δn in χκ.This shift enters the effective kernel operator used in the deterministic equivalents.
  • The proof linearizes the empirical risk and applies deterministic-equivalent estimates to resolvents.The resulting bounds are non-vacuous only above the scale where the empirical kernel resolvent remains sufficiently well-conditioned.
  • The appendix organizes the three regularization regimes through κ = d/t: moderate κ, polynomially vanishing κ, and the ridgeless limit κ → 0+.These correspond respectively to the generalization, interpolation, and memorization phases.

C.1 Linearization of the kernels and Bayes denoiser

The kernel and Bayes-denoiser analysis shows that, in the proportional regime, the relevant operators can be approximated by effective linearized forms, with nonlinear effects entering through a regularization shift.

  • Kernel linearization yields operator errors of order d^-3/2 for the empirical and population components under the stated assumptions.The bounds are obtained using Taylor expansions, concentration, and moment estimates.
  • The proof controls the difference between the empirical and effective kernels by decomposing it into off-diagonal and diagonal terms.Both components are shown to be small in the relevant operator or L2 norms.
  • If h(0) or h''(0) is nonzero, the linearization acquires additional constant-function contributions that can be handled separately.The paper states that these contributions do not affect the main results.
  • The nonlinear kernel contributes the shift δn = (h(ρ^2τΣ) − ρ^2τΣ)/n to the effective operator.Under the positive-definite kernel assumptions, this self-induced regularization term is non-negative.

C.2 Construction of the localized bumps

The localized-bump construction produces functions that interpolate training behavior while remaining controlled in RKHS norm, supporting the intermediate overfitting regime.

  • The bump polynomial has degree 2k0, no terms below degree k0, and is flat at both zero and one through order k0.These properties make the resulting functions localized while suppressing lower-degree components.
  • The constructed bump functions lie in the high-degree RKHS component and form an interpolation matrix whose Gaussian expectation is close to the identity.This makes the mean interpolation system invertible with a controlled inverse.
  • The localized bumps approximately interpolate the empirical Bayes denoiser while maintaining RKHS norm O_d(d).The construction is the key ingredient for analyzing polynomially vanishing regularization.
  • Off-diagonal overlaps decay rapidly because active degrees satisfy k ≥ k0 ≥ 6 and n ≍ d.Consequently, the sum of off-diagonal Gram-matrix contributions vanishes asymptotically.
  • The construction yields vanishing training error while test inputs see approximately the pure-noise score, producing the interpolation plateau.The resulting ridge estimator is analyzed through the same linearized risk and resolvent framework.

D.2 Proof of Theorem 7

Theorem 7 analyzes a regularized empirical risk estimator by comparing it with localized bump functions that interpolate the empirical Bayes denoiser. The comparison yields small training loss while controlling the estimator’s RKHS norm and isolating its linear component.

  • The proof constructs localized bump functions around each sample to interpolate the empirical Bayes denoiser with small training error and RKHS norm.
  • The regularized empirical risk adds a κ-dependent penalty to the unregularized training loss.
  • The ridge estimator’s optimality bounds its risk by the localized bump function’s risk under the condition d^(-k0+2) ≲ κ.
  • The resulting estimator has training loss bounded by κ and RKHS norm bounded by d.
  • Decomposing the estimator into linear and higher-order components allows separate control of its linear coefficients and nonlinear residual.
  • The nonlinear residual vanishes in L2(πd) after the linear coefficients are shown to approach the target vector.

E Proofs of the gradient flow results

This appendix derives the asymptotic risk equivalents for Theorems 1 and 2 while holding diffusion time fixed and following arguments analogous to those for ridge regression.

  • The proofs of the asymptotic risk equivalents for Theorems 1 and 2 assume that diffusion time t > 0 remains fixed.

E.1 Proof of Theorem 1

Theorem 1 is proved by replacing the kernel operators with linearized effective operators, deriving finite-dimensional spectral representations, and controlling the resulting approximation errors. These bounds hold uniformly over diffusion time under high-probability events.

  • The proof approximates training and test errors using linearized effective kernel operators before reducing them to finite-dimensional representations.
  • The operator semigroups are contractive because the kernel operators are positive semidefinite, enabling uniform error control over time.
  • Operator semigroup differences are bounded by t times the operator-norm difference between the original and effective kernels.
  • The spectral representation uses eigenvalues and spectral profiles to express the linearized training, test, and denoising errors.
  • The effective operators preserve linear functions and act on their coefficient vectors through finite-dimensional matrices with block spectral decompositions.
  • The final approximation statements hold with probability at least 1 − C d^(-D) for arbitrary ε and D.

F.1 Convergence of the test error under truncation

This section establishes convergence of the truncated gradient-flow estimator’s test error by proving RKHS density under the relevant Gaussian-smoothed measures and convergence to the empirical Bayes denoiser.

  • The RKHS is dense in both the noisy empirical measure’s L2 space and the population measure’s L2 space.
  • Density follows by showing that suitable weighted polynomials belong to the RKHS and that polynomials are dense under finite exponential-moment conditions.
  • As t → ∞, gradient flow converges in the noisy empirical L2 space to the empirical Bayes denoiser, while test error is asymptotically lower bounded by Bayes risk.
  • The truncated estimator’s test error converges to the empirical Bayes denoiser’s test error as t → ∞.
  • Projection onto a compact convex set preserves convergence because Euclidean projection is non-expansive.

F.2 Instability of the test error at the ridgeless limit

The section develops polynomial and kernel estimates showing that gradient flow follows finite-degree polynomial projections at suitable times, while the ridgeless-limit test risk can become unbounded.

  • Analytic control: Polynomial projections converge locally uniformly on Cd to an entire function that agrees everywhere with the RKHS target.Weierstrass convergence, L2 convergence, continuity, and the identity theorem are used in sequence.
  • Polynomial kernel structure: The truncated empirical kernel operator has range PN, the space of polynomials of degree at most N.Positive definiteness of the feature covariance operator establishes the range identity.
  • Spectral scale: The smallest positive eigenvalue λN admits bounds that decay exponentially with N, implying λN → 0 and tN → ∞.The decay follows from the upper and lower eigenvalue estimates for the truncated operator.
  • Gradient-flow approximation: At training time tN = λN^-2, gradient flow shadows the degree-N polynomial projection of its target.The approximation is formulated through the smallest positive eigenvalue λN of the truncated empirical kernel operator.
  • Main instability result: The gradient-flow risk can be unbounded for a universal inner-product kernel.This is stated as the main result of the subsection under the specialization π0 = γ1.

F.3 Test error of the empirical Bayes denoiser

The empirical Bayes denoiser is analyzed through concentration of its weights on a single training point, yielding a high-dimensional characterization of its population risk.

  • Weight concentration: Empirical Bayes weights are asymptotically concentrated on a single training point.The analysis uses a winner-take-all lemma for the competing scores associated with training samples.
  • Score separation: The two largest competing scores are attained at unique indices almost surely, and their indices fall into different balanced partitions with probability at least 1/2.This partitioning argument supports control of the gap between the largest and second-largest scores.
  • Risk decomposition: The population-risk calculation decomposes the error into two principal terms and a cross term, with the first term approaching τΣ and the middle term vanishing.The argument is conditional on the training data and uses the covariance identity for the last term.
  • Risk identification: In the fixed-dimensional setting, the infinite-training-time risk of the learned estimator equals the risk of the empirical Bayes denoiser.This identity is combined with the high-dimensional asymptotic analysis to prove the theorem’s second assertion.

G.1 Properties of the linearized reverse process

The linearized reverse process has a finite-dimensional Gaussian representation and remains bounded and delocalized under the stated assumptions.

  • Invariant-subspace structure: The sample-corrected operator preserves an (n + d)-dimensional invariant subspace, unlike the common-correction operator’s covariance-based reduction.Sample-dependent diagonal corrections prevent reduction to a spectral function of the empirical covariance.
  • Gaussian representation: Conditionally on the training data, the linearized reverse process is a centered Gaussian process.The conclusion follows from the explicit finite-dimensional representation of the linearized score dynamics.
  • Uniform boundedness: The linearized process has covariance bounded uniformly in time and dimension.This bound is stated conditionally on the input training data.
  • Delocalization: Its overlaps with training data and its Euclidean norm satisfy fixed-moment bounds, establishing delocalization.The overlap and norm controls are obtained through Gaussian-process chaining and Borell–TIS estimates.
  • Sample correction: The sample-dependent diagonal correction is essential for obtaining the d^-2 bound on the invariant subspace.Replacing it with a common correction introduces an additional term and loses this bound.

G.3 Proof of Theorem 4

The proof controls the empirical reverse SDE by comparing it with the linearized process along a delocalized path, using score-discrepancy bounds and path-space relative entropy.

  • Path-law comparison: A uniform-in-time score discrepancy along the linearized reverse process is used to control the KL divergence between the two reverse-SDE path laws.The comparison is made between the empirical reverse SDE and its linearized counterpart.
  • Operator control: Positive definiteness of the empirical and linearized operators bounds their associated propagators by τt.This contraction-type control is used in the variation-of-parameters estimates.
  • Score-error control: The score difference is bounded by decomposing its squared Euclidean norm into three terms.The proof controls these terms using the linearization estimates and operator bounds.
  • Probability statement: The uniform estimates hold with very high probability over the training sample.The proof invokes the uniformity of the preceding proposition over diffusion time and unit directions.
  • Well-posedness: The linearized reverse SDE has a unique global strong solution because its drift has uniformly bounded time-dependent linear coefficients.The bounded coefficients imply global Lipschitz continuity with linear growth.

G.4 Proof of Theorem 5

The proof establishes that the covariance of the linearized reverse process is close to that of an auxiliary spectral linearized process. It then analyzes the latter through spectral diagonalization, deterministic equivalents, and uniform contour estimates.

  • Auxiliary process: The auxiliary spectral linearized process replaces the kernel evolution eKt with the time-dependent effective empirical operator bKeff,t.This defines the spectral score estimator ˇSt and its associated covariance ˇΣt.
  • Covariance comparison: The covariance of the linearized reverse process eΣt is shown to be close to the covariance of the spectral process ˇΣt.The proof bounds the difference between the covariance ODEs and combines Lemmas 36 and 37.
  • Spectral analysis: The spectral covariance ODE is diagonalized by the eigenvectors of bΣ, enabling random-matrix analysis of the spectrum of ˇΣt.The coefficient matrices are spectral functions of bΣ, and the terminal covariance is the identity.
  • Deterministic equivalent: Lemma 37 provides a deterministic equivalent for the spectral linearized covariance under Assumptions 1, 2, and 6, with high-probability control.The statement applies to deterministic positive semidefinite test matrices A with nuclear norm at most 1 and probability at least 1 − Cd^-D.
  • Uniform estimates: Uniformity over the reverse-time schedule follows from spectral containment, contour bounds, continuity and analyticity of the relevant functions, and backward Gronwall estimates.These conditions allow the contour estimate and associated probability bound to hold uniformly over t ∈ [t0, T].
  • Conclusion: Theorem 5 follows by combining the comparison bound for the linearized and spectral processes with the deterministic-equivalent result.The two terms in the final estimate are controlled by Lemmas 36 and 37.
Loading 2608.23938v1…