Source-linked AI summary

Dynamical Regimes of Diffusion Models

Giulio Biroli, Tony Bonnaire, Valentin de Bortoli, Marc Mézard

arXiv:2402.18491v1cs.LGcond-mat.stat-mech

TL;DR

The paper addresses the theoretical behavior of diffusion models when data dimension and dataset size are large under an exact empirical score. Using statistical-physics analysis, it characterizes three backward-dynamics regimes and two transitions, with analytical and numerical evidence showing how dimensionality and dataset size govern collapse and memorization.

  • Problem

    The paper studies the open theoretical problem of understanding diffusion-model dynamics in high-dimensional spaces, where realistic data face the curse of dimensionality.

  • Method

    The authors use statistical-physics methods to analyze exact-score backward dynamics in large dimensions and large datasets, combining Gaussian-mixture solutions, general criteria, and experiments on real datasets.

  • Results

    The backward process has three regimes separated by speciation and collapse transitions, and finite-time avoidance of memorization requires a dataset exponentially large in dimension under the exact empirical score.

  • Takeaways & Limitations

    Speciation and collapse can be analyzed through data structure, while regularization can prevent collapse but may distort class proportions during speciation.

  • Takeaways & Limitations

    The speciation transition time uses a convention: the threshold Λe^-2t=1 could be replaced by another order-one value when Λ is large.

Abstract

from arXiv · show

Using statistical physics methods, we study generative diffusion models in the regime where the dimension of space and the number of data are large, and the score function has been trained optimally. Our analysis reveals three distinct dynamical regimes during the backward generative diffusion process. The generative dynamics, starting from pure noise, encounters first a 'speciation' transition where the gross structure of data is unraveled, through a mechanism similar to symmetry breaking in phase transitions. It is followed at later time by a 'collapse' transition where the trajectories of the dynamics become attracted to one of the memorized data points, through a mechanism which is similar to the condensation in a glass phase. For any dataset, the speciation time can be found from a spectral analysis of the correlation matrix, and the collapse time can be found from the estimation of an 'excess entropy' in the data. The dependence of the collapse time on the dimension and number of data provides a thorough characterization of the curse of dimensionality for diffusion models. Analytical solutions for simple models like high-dimensional Gaussian mixtures substantiate these findings and provide a theoretical framework, while extensions to more complex scenarios and numerical validations with real datasets confirm the theoretical predictions.

Introduction

The paper studies diffusion models in high-dimensional, large-dataset settings under an exact empirical score and identifies three successive regimes in backward generation. It argues that memorization and its avoidance depend critically on dimensionality, dataset size, and score regularization.

  • Introduction: The analysis targets diffusion models in the simultaneous limit of large data dimension and large dataset size, assuming the score is learned exactly from the empirical distribution.The exact-score assumption corresponds to sufficiently long training of a strongly over-parameterized network.
  • Introduction: Backward generation begins with Brownian-like motion, then specializes toward a major data class, and finally collapses onto an individual training example.The three regimes are described as pure noise, class commitment, and attraction to a memorized data point.
  • Introduction: Under the exact empirical score, avoiding memorization at finite times requires a dataset whose size is exponentially large in the dimension.Regularization and approximate score learning are presented as practical alternatives to the exact-score setting.
  • Introduction: The study identifies speciation and collapse as two crossover times separating these dynamical regimes.Speciation marks commitment to a category, while collapse marks entry into the attractor of a particular data point.
  • Introduction: The paper analyzes high-dimensional Gaussian mixtures, extends the theory to broader settings, and tests its findings numerically on CIFAR-10, ImageNet, and LSUN.The analytical and numerical studies are used to examine how dimensionality and sample count shape diffusion dynamics.

Main Results and Setting

In the large-dimension, large-dataset setting, the backward diffusion process is analyzed through three regimes: class commitment, generalization, and eventual memorization. The paper characterizes the speciation and collapse crossovers using covariance structure and excess entropy, while relating collapse to dataset scaling and statistical-physics transition phenomena.

  • Setting: The analysis considers diffusion models with data organized into distinct classes, using a principal covariance eigenvalue Λ to represent their dominant separation.The setting assumes an empirical dataset of n points in d dimensions and extends naturally beyond the simplifying two-class case.
  • Diffusion framework: Backward generation reverses a forward process whose long-time distribution is standard Gaussian, using the noisy empirical distribution and its exact empirical score.The forward process is modeled with independent Ornstein–Uhlenbeck dynamics, and the backward dynamics uses the corresponding score force and noise.
  • Dynamical regimes: The backward dynamics has three regimes: initially trajectories remain uncommitted, then specialize to a data class, and finally collapse onto an individual training point.Regimes I and II generalize when the dataset is sufficiently large, whereas regime III memorizes the training set.
  • Speciation: The speciation crossover tS marks commitment to a class and is controlled by the principal covariance eigenvalue Λ, with tS growing logarithmically with dimension when Λ diverges with d.The paper interprets this crossover as analogous to symmetry breaking and measures time from the beginning of the forward process.
  • Collapse criterion: The collapse crossover tC is located where the excess entropy density f(t)=ssep(t)−s(t) crosses zero, distinguishing separated Gaussian components from the true distribution entropy.For numerical work, the empirical estimate fe(t)=ssep(t)−se(t) approximates this criterion for t≥tC.
  • Dimensional scaling: The parameter α=(log n)/d controls collapse timing: α∼O(1) yields tC∼O(1), while avoiding memorization requires a dataset exponentially large in dimension under the exact-score setting.Increasing α shrinks regime III; practical mitigation uses a smoother approximate score together with a sufficiently large dataset.

Analysis of Gaussian mixtures

In high-dimensional Gaussian mixtures, backward diffusion undergoes speciation through dynamical symmetry breaking and later collapse into memorized data points. Analytical and numerical results connect these transitions to tS and tC, with collapse governed by dataset size relative to dimension.

  • Model: Gaussian mixtures with two clusters provide an analytically tractable model of backward diffusion in the large-dimension and large-dataset limit.The clusters have means ±m, common variance σ^2, and |m|^2=d˜µ^2.
  • Speciation: The speciation probability changes from shared-cluster outcomes to distinct-cluster outcomes around tS=(1/2)log d.The scaling variable is ˜µe^−(t−tS), and the infinite-dimensional limit produces a sharp transition at t/tS=1.
  • Speciation: After speciation, the overlap m·x(t)/d concentrates with a sign identifying a Gaussian component, so the dynamics generalizes within the selected class.The backward dynamics becomes equivalent to that for a single Gaussian centered at ±m.
  • Collapse: Collapse occurs when the single-data contribution Z1 dominates, making the score attract trajectories toward the originating training point.For t<tC, the associated random-energy model is in a glass phase corresponding to memorization.
  • Collapse: For t>tC, the collective contribution Z2...n dominates and the empirical distribution matches the population distribution at leading order in d.This non-collapsed regime is associated with the liquid, or high-temperature, phase.
  • Collapse: tC remains order one unless n is exponential in d, so avoiding collapse requires log n/d≫1.Empirical excess entropy estimates agree with the analytical collapse time even for moderately large n.

Generalization to realistic datasets

The paper extends its transition criteria beyond exactly solvable mixtures to datasets available only through samples. Speciation is inferred from covariance-spectrum structure, while collapse is estimated through entropy and volume arguments, with Gaussian-mixture predictions tested numerically.

  • General framework: The proposed criteria can be applied to realistic datasets when the population distribution is unknown or too complex for exact analysis.The approach uses criteria derived for speciation and collapse directly from available data.
  • Speciation: Speciation is identified by comparing principal-component structure Λe^−2t with noise broadening Δt.The transition criterion is Λe^−2t∼Δt, corresponding to noise blurring class structure.
  • Speciation: The speciation time is conventionally assigned where Λe^−2t reaches one, although any order-one threshold gives the same large-Λ scaling.This convention is used to state the general result for tS.
  • Collapse: Collapse is estimated by comparing the empirical distribution’s data-centered regions with the population-supported region as noise increases.A volume criterion equates the empirical and population volumes, yielding the general expression for tC.
  • Collapse: The volume argument is asymptotically valid because the relevant entropy identities have corrections that vanish as d grows.The criterion is expressed in intensive quantities and is accurate up to vanishing corrections for d→∞.
  • Statistical-physics interpretation: The framework is connected to glass physics: training-set-centered lumps correspond to glassy configurations, whereas broad population coverage corresponds to the liquid phase.The paper presents this as an analogy for the empirical distribution’s behavior before and after collapse.

Numerical experiments

Numerical experiments on realistic image datasets support the predicted speciation and collapse transitions. The measured crossover behavior agrees with the theoretical criteria for both times.

  • Experimental setup: The experiments used MNIST, CIFAR-10, downsampled ImageNet, and LSUN with two classes and varied dataset sizes and dimensions.The score was learned with a heavily over-parameterized U-Net intended to approximate the exact score.
  • Speciation time: Speciation was estimated by cloning two backward trajectories at time t and measuring whether they ended in the same class.A ResNet-18 classifier with test accuracy above 95% identified the terminal classes.
  • Speciation time: Rescaling time by the predicted tS made ϕ(t) change from 0.5 at t ≫ tS to one at t ≪ tS across markedly different image datasets.The similar rescaled curves support a common speciation phenomenon.
  • Collapse time: Collapse was studied on ImageNet16, ImageNet32, and LSUN using small training sets, including n = 200 for LSUN and n = 2000 for ImageNet.The smaller datasets were needed to represent the singular small-time behavior of the exact score.
  • Collapse time: Two collapse-time estimators agreed: cloned trajectories converged onto the same training datum, while nearest-neighbor tracking identified the last index change.The estimates crossed ϕC(t) at around 0.60 for all datasets and were compared with the excess-entropy criterion.
  • Conclusion: The experiments demonstrate speciation and collapse in realistic image datasets and validate the criteria from Eqs. (4) and (5).These criteria identify the corresponding crossover times.

Discussion

The discussion frames speciation and collapse as large-data, large-dimension transitions of the exact-score backward dynamics. It also identifies generality, practical implications, and boundaries imposed by exact-score assumptions and dimensionality.

  • Main findings: For large n and d without regularization, the backward dynamics has three regimes separated by speciation and collapse transitions.The transitions become true transitions in the large-n, large-d limit.
  • Generality: The criteria for speciation and collapse depend on Pt(x), so they apply to both deterministic flow-based and stochastic diffusion-based generative processes.Different denoising procedures share the same backward evolution of Pt(x).
  • Limitations and outlook: The analysis assumes an exact empirical score, while approximate-score and regularized settings remain important directions for further quantitative study.The paper specifically identifies the role of regularization as an open research question.
  • Practical implications: Avoiding collapse under the exact empirical score requires an exponentially large number of data in d.Regularization can circumvent this problem in practice when the dataset is sufficiently large, but it may also affect class proportions.
  • Practical implications: Accounting for the three backward-dynamics regimes could improve practical diffusion-model procedures.The paper presents this as a suggestion for improving performance and understanding implementations.

B.2 Asymptotic stochastic process in regime I and symmetry breaking

In regime I, the high-dimensional backward process is described by a reduced stochastic dynamics whose potential changes from quadratic to double-well form. This symmetry-breaking structure produces speciation near tS.

  • Reduced dynamics: During regime I, each coordinate follows a Langevin equation coupled to the other coordinates through the fluctuating rescaled overlap q(t).This is characterized as a dynamical mean-field equation.
  • Reduced dynamics: In the Gaussian-mixture analysis, the dynamics is reduced to the overlap q(t) with the direction joining the cluster centers.The resulting closed equation is a Langevin equation in an inverted potential.
  • Symmetry breaking: At times t ≫ (1/2) log d the potential is quadratic, whereas at earlier times it develops a double-well structure.This change yields symmetry breaking between the two mixture classes.
  • Speciation criterion: For large d, the clone probability ϕ(t) decreases from 1 at t = 0 to 1/2 as t approaches infinity.The Gaussian-mixture speciation estimate uses the threshold ϕ(ts) = .775.

B.4 Asymptotic stochastic process in regime II

After speciation, trajectories are committed to one Gaussian-mixture center, and the score simplifies to the backward score of that single Gaussian. The same analysis applies to either center.

  • Post-speciation scaling: After regime I, the overlap with the mixture centers grows proportionally to d, marking a change of scaling regime.Some trajectories have overlap diverging toward positive values and others toward negative values.
  • Single-center dynamics: Conditioning on trajectories committed to +m simplifies the score to that of a single Gaussian centered at +m.The analogous result holds for trajectories committed to −m.
  • Single-center dynamics: The resulting single-Gaussian backward equation ensures that trajectories committed to +m generate that Gaussian.Replacing mi by −mi gives the corresponding dynamics for the other class.

C.1 Setting and methods

The analysis treats contributions from data points as random-energy partition functions in the large-dimension, large-sample limit. Free-energy comparisons and condensation identify the transition structure of the backward diffusion.

  • Partition-function formulation: The model decomposes the data contribution into Z1, Z+, and Z−, representing one point, same-cluster points, and opposite-cluster points.The two bulk terms are analyzed as partition functions over independent energy levels.
  • Condensation analysis: Because ψ− < ψ+, the collapse transition is determined by the time when ψ+(t) equals −1/2.This condition identifies the transition in the large-dimension limit.
  • Statistical-physics interpretation: The partition functions are interpreted using a random-energy model with n−1 independent energy levels in the limit n,d → ∞ at fixed α = log n/d.This connects the diffusion calculation to the statistical-physics theory of glassy systems.
  • Partition-function formulation: The reduced energies follow a large-deviation principle, allowing the partition functions to be analyzed through rate functions and Legendre transforms.Their free energies concentrate in the large-d limit around ψ±(t).
  • Condensation analysis: The annealed free energy is exact before condensation, while afterward the partition function is dominated by the lowest-energy levels.The condensation time is characterized by equality between the minimum energy and the saddle-point energy.

C.4 Final expression

The final expression separates the backward dynamics into regimes according to which contribution dominates the probability. At the collapse time, attraction to an individual data point emerges from dominance of its term.

  • Final expression: In large dimensions, the opposite-cluster contribution Z− is irrelevant because ψ− < ψ+.The comparison therefore reduces to the individual-point term Z1 and the same-cluster term Z+.
  • Final expression: For t > tC, Z+ dominates and the dynamics is not collapsed.The probability is governed by the collective contribution of same-cluster data points.
  • Final expression: For t < tC, Z1 dominates and the backward score attracts the trajectory toward the data point a1.This is the collapsed regime.
  • Final expression: The collapse time tC equals the condensation time tcond of the random-energy model.The equality also extends beyond Gaussian mixtures when the energy distribution satisfies a large-deviation principle.

D.1 Gaussian mixtures

For two high-dimensional Gaussian clusters, the paper derives an analytic backward score and uses cloned trajectories and excess entropy to study speciation and collapse.

  • Gaussian-mixture experiments: For Gaussian clusters centered at ±m with variance σ2, the score function is analytically expressible and the backward stochastic process can be discretized.The experiment uses two clones sharing noise until a chosen time and independent noise afterward.
  • Gaussian-mixture experiments: Clone classes are recovered by projecting each final state onto m and taking sign(m · xi(0)).This provides a numerical estimate of the probability that the clones end in the same class.
  • Collapse estimation: The excess entropy density f(t) vanishes for t ≤ tC and can be derived exactly for Gaussian mixtures.It is estimated numerically by sampling noisy points associated with uniformly selected training data.
  • Collapse estimation: With n = 20 000 and n′ = 250 000, the empirical excess-entropy estimate closely fits the analytical curve across several dimensions.The procedure numerically derives tC from the dataset when the exact empirical score is used.

D.2 Realistic datasets

The realistic-dataset experiments train discrete-time DDPM denoisers on balanced two-class subsets and examine speciation and collapse across several image datasets.

  • DDPM setup: The discrete forward process uses a linear variance schedule from β1 = 10^-4 to β1000 = 2 × 10^-2.Its samples are represented from the initial state using the cumulative product of the schedule and Gaussian noise.
  • DDPM setup: The learned noise predictor generates samples by reversing the discrete process from x(T) ∼ N(0, I), with the continuous limit corresponding to the score-based diffusion.The learned noise is related to the score by −ξθ(x(t′), t′) /√1 − αt′.
  • Speciation experiment: Speciation time is estimated from the probability that cloned samples finish in the same class, using classifiers with at least 95% test accuracy.Accuracy exceeds 99% on datasets other than LSUN.
  • Collapse experiment: Collapse is tested through cloned trajectories and nearest-neighbor identities, with small training sets making attraction to training points observable.For t < tC, clones share the same nearest neighbor at very small Euclidean distance.
Loading 2402.18491v1…