Source-linked AI summary

Information Spreading in Diffusion Models from Effective Field Theory

Navonil Neogi, Nabil Iqbal

arXiv:2608.14308v1hep-thcond-mat.stat-mechcs.LG

TL;DR

Diffusion models can generate creative outputs despite training objectives that appear to favor memorization, but their dynamics remain difficult to describe simply. This paper applies effective field theory to local convolutional diffusion models and finds a diffusive regime of spatial information spreading in both toy grids and MNIST.

  • Problem

    The paper asks whether diffusion-model dynamics can be described by a universal theory whose dependence on complex training data is reduced to a few parameters.

  • Method

    The authors expand score functions using effective field theory, locality, and symmetry principles to simplify reverse-process dynamics.

  • Results

    Experiments on ±1 grids and MNIST show a diffusive regime in which spatial information spreads according to the self-similar variable x2/¯αt.

  • Takeaways & Limitations

    Effective field theory provides a quantitative description of spatial mutual-information growth in local convolutional diffusion models.

  • Takeaways & Limitations

    The theory does not analytically treat fuzziness at region boundaries, which the authors leave for future work.

Abstract

from arXiv · show

We study score-matching diffusion models with a convolutional architecture. We argue that the inductive bias of locality means that the machinery of effective field theory from physics can be usefully applied to describe the denoising dynamics. We apply this formalism first to a simple toy example which permits an analytical description, and thereafter to MNIST, and show that in both cases, the mutual information between two points grows in a manner predicted by a simple effective field theory of Brownian motion.

1 MOTIVATION

This work applies effective field theory’s scale-by-scale perspective to score-matching diffusion models, using locality and symmetry to seek simplified descriptions of reverse dynamics. It also studies spatial mutual information at the pixel level as structure develops during diffusion.

  • Motivation: Effective field theory explains long-distance behavior without requiring complete knowledge of microscopic behavior.Locality and symmetry can make short-distance physics influence long-distance dynamics through only a few relevant parameters.
  • Motivation: The paper applies this effective-theory philosophy to diffusion models that construct data structure gradually from pure noise.The authors argue that locality and symmetry can yield a simple effective field theory under certain circumstances.
  • Motivation: The study focuses on the score-matching incarnation of diffusion models and their reverse-process dynamics.Diffusion models are described as widely used generative-AI methods across images, text, and scientific applications.
  • Motivation: The work studies spatial mutual information on a pixel level over the course of diffusion.This builds on information-theoretic analyses of structure creation and mutual-information evolution in diffusion models.
  • Motivation: EFT provides the framework for modeling the reverse process at a particular length scale using constraints such as locality, symmetry, or causality.The paper presents EFT as a classic physics approach for constructing simplified models of system dynamics.

2 THEORY

Diffusion models learn a score function to reverse an Ornstein–Uhlenbeck noising process, and this paper uses effective field theory to simplify the resulting dynamics. For convolutional models, locality and translational equivariance predict diffusive growth of mutual information before a late snapping regime.

  • Diffusion-model theory: Diffusion models gradually transform data into approximately Gaussian noise through an Ornstein–Uhlenbeck process, then learn the reverse flow.The forward process runs from t = 0 at data samples to a large positive T near pure noise.
  • Diffusion-model theory: The score function π(ϕ, t) := ∇x log pt(x) is the central learned object that specifies the reverse direction from noise toward data.Neural networks learn it by score matching on a discretized forward noising process.
  • Effective field theory: Effective field theory addresses the complexity of the score function by seeking dynamics whose dependence on training data is captured by a small number of parameters.The approach treats the score as a potentially complicated functional while organizing its dynamics systematically.
  • Effective field theory: For convolutional diffusion models, spatial locality and translational equivariance suffice to apply general effective-field-theory principles to the reverse process.The paper applies this framework to the buildup of correlations in models with a convolutional backbone.
  • Diffusive regime: The EFT predicts that mutual information begins as a spatial delta function, then diffuses outward as a Gaussian whose width grows with time.The correlation length obeys ξ2 ∼ ¯αt, and plotting against |x − y|2/¯αt should collapse results at different times onto one curve.
  • Snapping regime: Near the reverse-process endpoint, individual patches decouple and follow simple exponential trajectories toward their nearest training patches, producing a memorisation or snapping regime.The EFT description also gives an explicit evolution equation for this snapping behavior.

3 EXPERIMENTS

Experiments in an analytically tractable toy model and a learned ConvNet show that reverse-process correlations follow the diffusive scaling predicted by the effective field theory. The diffusive regime ends as nonlinear snapping becomes important, with final correlation length set by locality and architecture rather than the noising schedule, while boundary effects and edge fuzziness remain limitations.

  • Analytical and learned models: Curves at different timesteps collapse under the self-similar variable, confirming diffusive scaling in the pure ELS model.The analytical example uses a constant 5 × 5 locality kernel.
  • Analytical and learned models: Curve collapse in the trained ConvNet is extremely similar to the analytical ELS model, except at very small separations where results diverge slightly.The comparison uses an equally split batch of 1 and −1 grids.
  • Transition out of diffusion: The diffusive regime ends when nonlinearities become important, but the transition continuously combines mutual-information spreading with snapping rather than occurring at a sharp point.The transition is characterized by the factor Ω, defined from the growth of the score-function gradient.
  • Transition out of diffusion: The duration of diffusion is controlled by the effective ELS kernel size ∆ and, more generally, by training-data structure and network architecture.The criterion assumes ¯αt is small enough at the end of diffusion that 1 −¯αt ≈1.
  • Final spatial correlations: Correlation length scales linearly with the ELS kernel size, while the reverse process creates spatial mutual information until a fixed level set by nonlinearity and locality-kernel size.Varying kernel size throughout reverse diffusion is expected to modify correlation length by a factor determined by that variation.
  • Mutual information: Mutual-information curves collapse at early times, but large-distance collapse is imperfect because boundary conditions interact with diffusive modes.Explicit and periodic boundary conditions produce different large-r² curves and different data-collapse quality.

4 CONCLUSION

The paper applies effective field theory to diffusion models with local convolutional architectures, simplifying reverse-process dynamics and characterizing information spreading. Experiments on ±1 grids and MNIST confirm a diffusive regime in which the noise schedule acts as diffusion time, while extensions to other architectures remain open.

  • 4 CONCLUSION: The study applies physics-inspired effective field theory to diffusion models with local convolutional architectures.The learned reverse-process score and sample evolution are expanded using locality and symmetry principles.
  • 4 CONCLUSION: EFT simplifies the ELS machine’s analytic structure near reverse-process endpoints and enables analysis of spatial mutual-information growth.The resulting equations involve only a few straightforward terms.
  • 4 CONCLUSION: Nonlinear terms are analytically difficult but unimportant near the pure-noise endpoint, and the equations are calculated explicitly for ±1-grid toy models.This conclusion follows from the structure of the ELS machine.
  • 4 CONCLUSION: Experiments on ±1 grids and MNIST show a diffusive regime where information spreads according to the self-similar variable x2/¯αt.The noise schedule plays the role of time in the diffusion equation, matching the theoretical predictions.
  • 4 CONCLUSION: Whether EFT can help analyze other architectures, including transformers, remains an open question for future research.The paper notes that locality may be inherited from data independently of the underlying architecture.

A APPENDIX · A.1 LONG-DISTANCE DESCRIPTION OF EQUIVARIANT LOCAL SCORE MACHINE

The appendix develops a derivative expansion for the equivariant local score machine, showing that its short-range structure yields an effective field theory with a weakly coupled fixed point near pure noise. The expansion explains why nonlinear terms become subleading as ᾱ_t → 0 and is validated on the ±1 grid by two matching derivations.

  • A.1 LONG-DISTANCE DESCRIPTION OF EQUIVARIANT LOCAL SCORE MACHINE: The equivariant local score machine admits a systematic derivative expansion near the pure-noise endpoint, where it has a weakly coupled fixed point.The expansion is formulated in continuum notation and organized by powers of ϕ and its derivatives.
  • A.1 LONG-DISTANCE DESCRIPTION OF EQUIVARIANT LOCAL SCORE MACHINE: Because the score depends only on local patches, its short-range kernel permits an analytic momentum-space and derivative expansion, usable when derivatives are small.Only the first few terms are worked out because the expansion rapidly becomes unwieldy.
  • A.1 LONG-DISTANCE DESCRIPTION OF EQUIVARIANT LOCAL SCORE MACHINE: The microscopic equivariant local score machine yields explicit effective-field-theory coefficients as dataset sums, while symmetries constrain which terms can appear.A Z2-symmetric dataset removes odd contributions, and rotational invariance restricts the two-tensor coefficient to a delta structure.
  • A.1 LONG-DISTANCE DESCRIPTION OF EQUIVARIANT LOCAL SCORE MACHINE: Every occurrence of ϕ(x) in the weight is multiplied by √ᾱ_t, so the weight depends on ϕ through √ᾱ_tϕ.This scaling follows after cancellation of the shared exp(−ϕ²) factor in the numerator and denominator.
  • A.1 LONG-DISTANCE DESCRIPTION OF EQUIVARIANT LOCAL SCORE MACHINE: Nonlinear EFT terms such as ϕ², ϕ³, and ϕ∇ϕ are strictly subleading to linear terms as ᾱ_t → 0, explaining their lack of visible experimental impact.Each additional power of ϕ brings at least a factor of ᾱ_t^1/2.
  • A.1 LONG-DISTANCE DESCRIPTION OF EQUIVARIANT LOCAL SCORE MACHINE: The ±1 grid example demonstrates that the EFT coefficients can be extracted meaningfully from the formalism, while distinguishing continuum-integral approximations from the exact discrete grid sum.The calculation considers φ+(x)=1 and φ−(x)=−1 and reverts to the discrete sum for the 3 × 3 neighborhood.
  • A.1 LONG-DISTANCE DESCRIPTION OF EQUIVARIANT LOCAL SCORE MACHINE: A first-principles derivation for the ±1 grid produces exactly the same EFT coefficient as the discrete calculation, providing a robust validity check.The independent derivation uses only the analytic description of the ±1 grid case.

A.2 CHECKING AGAINST EXPLICIT ELS EXPANSION IN TWO SAMPLE CASE

The explicit expansion specializes to 3×3 locality patches and, after a consistent Taylor expansion in the small-noise limit, yields a diffusion term matching the previous method. Its coefficient determines the characteristic spatial-correlation length scale.

  • Explicit ELS expansion: The derivation assumes 3×3 square locality patches around each pixel, with higher-dimensional generalizations producing higher-order derivatives.The notation later specializes the resulting operator to a 3-by-3 locality kernel.
  • Explicit ELS expansion: The reverse-process evolution is expanded using the Taylor series of tanh, with higher-order terms suppressed as O(¯α_t^3/2).This establishes consistency of the expansion in the relevant limit.
  • Explicit ELS expansion: 3¯α_t is the diffusion-term coefficient, precisely agreeing with the coefficient obtained using the previous method.The time prefactor does not affect the positional solution governing spatial correlations.
  • Explicit ELS expansion: The diffusion coefficient directly determines the characteristic length scale governing spatial correlations.The derivation reads this scale from the simplified diffusion equation in the limit ¯α_t → 0.

B DIFFUSION EQUATION EVOLUTION AND SCALING

Starting from i.i.d. standard-normal pure noise, the diffusion equation preserves zero mean while its deterministic two-point correlation evolves diffusively. Translation equivariance reduces spatial dependence to point separation, yielding self-similar behavior controlled by z2/t.

  • Scaling behavior: The resulting solutions exhibit self-similar behavior: all spatial dependence is controlled by the dimensionless variable z2/t.This scaling follows because z and t scale so that z2/t is dimensionless in the diffusion equation.
  • Initial conditions and statistics: The initial pure-noise field has independent, identically distributed standard-normal pixels, and its expectation remains ⟨u(x, t)⟩ = 0 for all time.The zero-mean property makes the two-point correlation equivalent to the connected two-point function in this setting.
  • Initial conditions and statistics: The analysis studies the deterministic two-point correlation function of the random field rather than the field itself.This correlation captures the evolving statistics of u(x, t).
  • Correlation evolution: Translation equivariance requires the correlation to depend on x and y only through their separation, represented by z.The grid and diffusion problem are translationally invariant, so the correlation is translationally equivariant.
  • Correlation evolution: The two-point correlation satisfies a diffusion equation with deterministic initial condition G(z, 0) = δ(z), inherited from the i.i.d. initial pixels.The delta-function initial correlation follows because the starting pixels are i.i.d. normally distributed.

C EXPERIMENTAL DETAILS · C.1 DIFFUSION MODEL TRAINED ON ±1 GRIDS

The experiments use convolutional DDPMs with Euler-integrated reverse ODEs, including a ±1-grid model stabilized by v-prediction. This model uses cosine noise scheduling, restricted time amortization, compact equivariant convolutions, and specified training and sampling procedures.

  • C EXPERIMENTAL DETAILS: All experiments train convolutional diffusion models on 2D image data using the standard DDPM forward process and a first-order Euler integrator for the reverse ODE.The reverse process follows the ODE formulation in (3).
  • C.1 DIFFUSION MODEL TRAINED ON ±1 GRIDS: The ±1-grid dataset contains two 64 × 64 images, one uniformly +1 and one uniformly −1, whose symmetry and zero average cause numerical instability and poor training.The model is trained on these two homogeneous samples.
  • C.1 DIFFUSION MODEL TRAINED ON ±1 GRIDS: v-prediction learns the reverse-process velocity rather than the score function π directly and is used because it is more stable for this symmetric dataset.ϵ can be obtained from v, but direct score prediction is less stable in practice.
  • C.1 DIFFUSION MODEL TRAINED ON ±1 GRIDS: The forward process ends at T = 1000 with batch size 32 and uses the standard cosine scheduling of Nichol & Dhariwal (2021).The displayed scheduling expression defines the noise schedule.
  • C.1 DIFFUSION MODEL TRAINED ON ±1 GRIDS: Training amortizes over the interval (0, 0.8T), avoiding the near-pure-noise endpoint where ϵ carries little data and is predominantly noisy.The loss is parameterized by neural-network parameters θ.
  • C.1 DIFFUSION MODEL TRAINED ON ±1 GRIDS: Training runs for 200 steps with ADAM at learning rate = 2 × 10−4, using a convolutional Z2-equivariant local network with four 2D convolutional layers and 3-sized kernels.The channel counts are (1, 2, 2, 1), and time is additively embedded after the first Conv2D layer through a two-layer SiLU MLP.
  • C.1 DIFFUSION MODEL TRAINED ON ±1 GRIDS: After learning v, the procedure recovers ϵ and ϕ0, derives the score function, and implements the Euler reverse process.The recovered quantities are used according to the usual reverse-process construction.
  • C.1 DIFFUSION MODEL TRAINED ON ±1 GRIDS: The reverse process uses 400 steps and produces an output batch size of 20 for extracting correlation statistics.These samples provide the statistics used for correlation analysis.

C.2 DIFFUSION MODEL TRAINED ON MNIST

The study trains a diffusion model on MNIST using a locality-preserving minimal UNet, standard epsilon prediction, and a linear noise schedule with a 200-step reverse process.

  • C.2 DIFFUSION MODEL TRAINED ON MNIST: The MNIST experiment uses 28×28 black-and-white digit images, a final forward-process time of T = 1000, and an output batch size of 100.The dataset contains digits 0 to 9, and the reverse process uses 200 steps.
  • C.2 DIFFUSION MODEL TRAINED ON MNIST: The model is a minimal UNet with two downsampling and two upsampling blocks, up to 128 channels, and no skip connections or ConvTranspose2D layers.ConvTranspose2D layers are omitted because they break locality and diverge from theoretical ELS machine predictions.
  • C.2 DIFFUSION MODEL TRAINED ON MNIST: Because the MNIST task is less pathological, the network learns ϵ directly with the standard loss rather than requiring v-prediction.The learned ϵ drives the reverse process through a first-order Euler integrator.
  • C.2 DIFFUSION MODEL TRAINED ON MNIST: The implementation uses a linear noise schedule with βt spanning (0.0001, 0.02), followed by the corresponding cumulative ᾱt construction.The passage states the schedule and ᾱt relation but presents the latter incompletely.

C.3 ELS IMPLEMENTATION

The ELS machine is implemented exactly as an analytic benchmark for local models, with exact reverse-process implementation for the ±1 grid and patch-based score calculation for MNIST.

  • C.3 ELS IMPLEMENTATION: Exact ELS implementations compare local-model behavior against analytically controlled results and test the theory of correlations.This provides both a behavioral check and an analytic setting with greater control than a trained model.
  • C.3 ELS IMPLEMENTATION: For the ±1 grid, analytical description (27) makes the reverse process exact and straightforward to implement.The implementation follows the analytical description directly.
  • C.3 ELS IMPLEMENTATION: For MNIST, local 5 × 5 patches are extracted with periodic boundary conditions and used in a weighted sum to calculate the score function.Patch size can be adjusted, while periodic boundaries preserve the simple theoretical ELS setting; score calculation is the most computationally intensive code component.

C.4 CALCULATING CORRELATIONS

The two-point correlation ρ(r) is computed by normalizing separation-dependent pixel correlations within each grid and averaging across the batch. Plots exclude the smallest separations and restrict r to below half the grid size because the theory and periodicity become problematic there.

  • Correlation calculation: ρ(r) is obtained by computing each grid’s pixel standard deviation, averaging two-point functions at separation (r, 0), normalizing, and then averaging across the batch.The separation-dependent mean is taken over all pixel pairs displaced by r along the x-axis.
  • Separation range: Separations near r ∼1, 2 pixels are excluded because the effective-field-theory description is not meaningful at that scale.Effective smearing across a patch contains the lengthscale under examination at these smallest separations.
  • Separation range: Only r < L/2 is retained because periodicity causes the correlation function to increase again after decaying with r beyond half the grid size.L denotes the total grid size.
Loading 2608.14308v1…