Source-linked AI summary

Theory of Speciation Transitions in Diffusion Models with General Class Structure

Beatrice Achilli, Marco Benedetti, Giulio Biroli, Marc Mézard

arXiv:2602.04404v1cs.LGcond-mat.dis-nn

TL;DR

Existing speciation analyses apply mainly when classes are spatially separated and identifiable by first moments. This paper generalizes the theory using Bayes attribution and free-entropy differences, obtaining logarithmic speciation times, hierarchical attribution, and analytical results for Ising mixtures.

  • Problem

    Previous methods apply only to target distributions whose classes are spatially well separated and identifiable by first moments.

  • Method

    The paper defines classes through Bayes classifiers and identifies speciation times by analyzing Bayes component attribution and free-entropy differences during forward diffusion.

  • Results

    Merging times scale logarithmically with sample-space dimension and are separated by finite gaps, while Ising-mixture predictions agree closely with numerical experiments.

  • Takeaways & Limitations

    The criterion extends speciation analysis to classes not identifiable by first moments and supports hierarchical attribution in broad classes of mixture models.

  • Takeaways & Limitations

    The approach assumes a proper density decomposition with concentration properties and is stated to apply to mixture models whose components have finite correlation length.

Abstract

from arXiv · show

Diffusion Models generate data by reversing a stochastic diffusion process, progressively transforming noise into structured samples drawn from a target distribution. Recent theoretical work has shown that this backward dynamics can undergo sharp qualitative transitions, known as speciation transitions, during which trajectories become dynamically committed to data classes. Existing theoretical analyses, however, are limited to settings where classes are identifiable through first moments, such as mixtures of Gaussians with well-separated means. In this work, we develop a general theory of speciation in diffusion models that applies to arbitrary target distributions admitting well-defined classes. We formalize the notion of class structure through Bayes classification and characterize speciation times in terms of free-entropy difference between classes. This criterion recovers known results in previously studied Gaussian-mixture models, while extending to situations in which classes are not distinguishable by first moments and may instead differ through higher-order or collective features. Our framework also accommodates multiple classes and predicts the existence of successive speciation times associated with increasingly fine-grained class commitment. We illustrate the theory on two analytically tractable examples: mixtures of one-dimensional Ising models at different temperatures and mixtures of zero-mean Gaussians with distinct covariance structures. In the Ising case, we obtain explicit expressions for speciation times by mapping the problem onto a random-field Ising model and solving it via the replica method. Our results provide a unified and broadly applicable description of speciation transitions in diffusion-based generative models.

1. Introduction

Diffusion-model trajectories can undergo speciation transitions as backward denoising progressively commits them to smaller data-class subsets. This work extends existing analyses beyond first-moment-separated classes using Bayes attribution and applies the framework to multiple classes and analytically tractable mixtures.

  • Diffusion models: Diffusion Models generate samples by reversing a stochastic diffusion process from Gaussian noise toward a target probability distribution.The reverse dynamics use a score-function drift to recover the target distribution at t = 0.
  • Speciation transitions: Speciation transitions are sharp changes in backward trajectories during which class predictions fluctuate within, then commit to, progressively smaller subsets of classes.In a four-class animal example, predictions can remain within cat–dog or eagle–seagull subsets before committing further.
  • Prior theory: Earlier theory characterized an initial speciation timescale through a spectral criterion for mixtures of two well-separated Gaussians.Subsequent work extended the setting to more than two Gaussians and multiple speciation times.
  • Prior theory: Existing methods apply only when classes are spatially separated and identifiable from first moments.This limitation excludes classes distinguished through other distributional structure.
  • This work: The paper defines classes through Bayes classifiers, identifies speciation using forward-diffusion attribution error, and extends the analysis to multiple classes with successive speciation times.The criterion also covers cases whose class structure is not identifiable by first moments.
  • Examples: The framework is illustrated with one-dimensional Ising mixtures at different temperatures and zero-mean Gaussian mixtures with distinct covariance structures.For the Ising mixture, speciation times are obtained analytically by mapping to a random-field Ising model and using a replica calculation.

2. Bayes attribution and pure densities mixtures

The framework represents a target distribution as a mixture of well-distinguished components and assigns samples to components using Bayesian attribution. Proper density decompositions require asymptotically reliable attribution after finite Gaussian corruption, with separation expressed through component likelihoods.

  • Mixture model: The target distribution is modeled as a mixture of R components with non-negative weights summing to one.Each component has its own probability distribution, denoted P_r(a).
  • Bayesian attribution: Samples are assigned to the class maximizing the posterior probability P(s | a).This Bayesian attribution defines the component label for a sample drawn from the mixture.
  • Proper density decomposition: A Proper Density Decomposition requires that, after adding finite Gaussian noise, a typical sample from each component is attributed to its source component with high probability as dimension grows.For distinct components r and s, the incorrect posterior must be bounded by ε(N) = o_N(1) with high probability.
  • Large-deviation characterization: The framework models component likelihoods as P_s(a) = e^{N f_s(a,N)} and uses large-deviation dominance of the source component over alternatives to characterize proper decompositions.For a sample from component r, N f_r dominates N f_s with high probability for every s ≠ r.

large N

For pure density decompositions, component free-entropy contributions self-average at large dimension. This self-averaging distinguishes minimal proper decompositions from cases where a component can itself be decomposed.

  • Pure density decomposition: A pure density decomposition occurs when the component free-entropy density concentrates around its mean.The criterion is applied after assuming the underlying proper density decomposition.
  • Self-averaging: The free-entropy representation separates the averaged contribution from an o(N) realization-dependent disorder term.The average is taken over diffused samples originating from the relevant component.
  • Self-averaging: Self-averaging fails when a component can itself be properly decomposed, making it a signature of a minimal proper decomposition.The paper relates this notion of pure densities to earlier work on pure density decompositions.

3. General criterion for speciation

Speciation occurs when Bayesian attribution between two components becomes unreliable because their average free-entropy difference is comparable to fluctuations. The resulting criterion recovers first-moment Gaussian scaling and extends to classes indistinguishable by means.

  • General criterion: A component remains reliably identifiable when its Bayesian free-entropy score exceeds every competing component's score with high probability.This condition is expressed by Nf_ss(t) + δf_ss(x,t) ≫ Nf_rs(t) + δf_rs(x,t) for all s ≠ r.
  • General criterion: Speciation begins when the free-entropy difference between two components becomes comparable to its fluctuations.At this point, the Bayes classifier can assign finite probability to multiple classes and misattribute trajectory origins with finite probability.
  • General criterion: The criterion depends only on the pair of components whose speciation time is being predicted, regardless of other coexisting components.An arbitrary constant K controls the strictness of the no-misattribution requirement and shifts the estimated speciation time.
  • Large-time reduction: At large forward times, the average free-entropy difference can be interpreted as a Kullback–Leibler divergence between mixture components or their Gaussian approximations.The Gaussian approximation uses component means and covariances to characterize the asymptotic free-entropy difference.
  • Scaling regimes: When first moments are separated, the framework recovers the previously known speciation-time scaling.The recovered result is t_rs ∼ 1/4 log N, up to the displayed constant correction.
  • Scaling regimes: The same theory extends speciation-time predictions to classes with no first-moment separation.For large N, leading behavior is independent of target-measure details, while finite corrections depend on the specific components.

4. Two simple examples: high dimensional Gaussian data

The examples test speciation predictions through changes in reverse-diffusion potentials for Gaussian mixtures with separated means and with equal means but different variances. Both transitions are identified by the emergence of a double-well structure or vanishing curvature.

  • Examples: The Gaussian examples provide analytically tractable checks of the general criterion through explicit reverse-diffusion dynamics.The examples include both first-moment separation and separation through different isotropic variances.
  • Gaussian mixtures with different means: For separated-mean Gaussian mixtures, the speciation time scales as t ∼ 1/2 log N.The model uses balanced Gaussians with means ±m and isotropic variance σ^2, with ||m||^2 = N μ̃^2.
  • Gaussian mixtures with different means: The reverse-diffusion potential changes from quadratic at large times to a double well at t = 1/2 log N.The transition is associated with a change in curvature at the overlap coordinate r = 1.
  • Gaussian mixtures with different variances: For zero-mean Gaussian mixtures with different isotropic variances, the predicted speciation-time scaling is t ∼ 1/4 log N.This realizes the regime in which the components have identical first moments but distinct variance-dependent shell radii.
  • Gaussian mixtures with different variances: In the variance-separated model, speciation is identified by the reverse SDE potential's curvature vanishing at r = 1.The radial coordinate is r_t = ||x_t||^2/N, and the effective potential is defined from the deterministic drift.

5. Speciation times for multi-classes target distributions: 1D Ising mixtures

The framework predicts multiple speciation events in Ising mixtures by comparing free-entropy differences between components, including classes indistinguishable by first moments. Numerical and analytical attribution results exhibit the predicted successive merging structure.

  • Class structure: 1D Ising components have identical zero first moments but occupy different regions of configuration space at large system size.Their distinguishing structure is therefore collective rather than mean-based.
  • Criterion: The criterion predicts all speciation times for arbitrary numbers of Ising components, where merging is governed by correlation-length-related structure.For more than two components, this extends beyond the first transition accessible in earlier analyses.
  • Criterion: Bayes attribution reduces each component likelihood to the partition function of a 1D Ising system in the random external field supplied by the diffused sample.Speciation occurs when pairwise average free-entropy differences become comparable to their fluctuations.
  • Two-component validation: For two Ising components, free-entropy fluctuations crossing zero coincides with rising misattribution, supporting the proposed criterion at about 1% misattribution for K = 3.The example uses β1 = 0.5, β2 = 1.0, equal mixture weights, and varying numbers of spins.
  • Multiple components: For three components, predicted merging times are t1 ≃0.42 for β1 and β2 and t2 = 1.39 for β2 and β3, with block structure emerging after these times.The attribution matrices use β1 = 0.2, β2 = 0.3, β3 = 1.0 and N = 1600.
  • Multiple components: For eight hierarchically organized Ising chains, analytical and numerical attribution matrices agree, revealing successive commitment to smaller temperature subsets and finally one temperature.The comparison uses N = 1600 and places numerical matrices above analytical matrices.

6. Conclusions

The paper extends speciation theory to mixture components that are not separable by first moments, using Bayes attribution and free-entropy differences to identify transitions. It finds logarithmic merging times, finite gaps between transitions, and hierarchical attribution, with explicit Ising results validated numerically.

  • Contribution: The theory covers mixture components whose regions cannot be identified through first moments.This broadens speciation analysis beyond previously studied spatially separated mean-based classes.
  • Criterion: Speciation times are identified from Bayes component attribution during forward diffusion and differences between component free entropies.The criterion links loss of reliable attribution to free-entropy fluctuations.
  • Main result: Merging times scale logarithmically with sample-space dimension and are separated by finite time gaps, producing hierarchical attribution.The resulting hierarchy reflects progressively finer class commitment.
  • Ising example: For 1D Ising mixtures, the mean and standard deviation of the free-entropy difference decay exponentially, yielding tS ∼ 1/4 log N.Numerical experiments show excellent agreement with the predicted onset of misattribution.
  • Scope: The criterion applies to a broad range of mixture statistical-physics models when each component has finite correlation length.This is the stated scope condition for the claimed applicability.

Appendix A. Large time analysis

The large-time analysis approximates diffused mixture components by Gaussian distributions and formulates speciation through average free-entropy differences and their variance. It recovers distinct scaling regimes depending on whether classes differ in first moments.

  • Gaussian approximation: At large forward times, each diffused component is approximated by a Gaussian, making the average free-entropy difference a Gaussian Kullback–Leibler divergence.The approximation expands the exact component distribution in e^-2t and retains its mean and covariance.
  • Speciation criterion: The speciation criterion compares the average free-entropy difference with fluctuations in the empirical free-entropy difference.The variance is estimated by separating linear and quadratic contributions under the centered Gaussian approximation.
  • Scaling regimes: When first moments distinguish classes, the analysis recovers the previously established speciation-time scaling.This is the regime corresponding to nonzero first-moment separation between mixture components.
  • Scaling regimes: When first moments vanish, the speciation time scales as t_rs = 1/4 log N plus model-dependent terms.This extends speciation-time analysis to classes separated by higher-order structure rather than first moments.
  • Speciation criterion: On the logarithmic-in-N speciation timescale, the free-entropy variance is O(1/N^2), yielding an equivalent criterion based on an average difference of order 1/N.The additional 1/N factor arises from evaluating the variance at a time that grows logarithmically with N.

Appendix B.4. Scaling of speciation time

This section derives large-N scaling for speciation times and describes the Ising-mixture calculation of Bayesian attribution and free entropy. The leading logarithmic behavior is separated from finite corrections.

  • Scaling of speciation time: The scaling derivation expands the relevant ratios in small parameters of order N^-1/2 and retains terms through second order.The calculation factors one component contribution, introduces a ratio of the two terms, and applies Taylor expansions.
  • Scaling of speciation time: The resulting expression contains a leading 1/4 log N term together with corrections depending on δ and numerical constants.The displayed result gives the correction as 1/2 log δ − 1/4 log 3.
  • Ising-mixture attribution: The analysis uses component-specific inverse temperatures and weights to define Bayesian attribution probabilities along the diffused trajectory.For the Ising mixture, the component free entropy is represented through a random-field contribution proportional to e^-t Δ_t x_i.

Appendix C.2. Analytic computation of the average free entropy

The average free entropy for the Ising mixture is computed through a replica-symmetric transfer-matrix formulation. The calculation reduces the thermodynamic-limit result to temperature-dependent spin correlations.

  • Replica computation: The Ising-mixture average free entropy is evaluated with the replica trick in the Replica Symmetric approximation.Integer powers introduce replicated spins, followed by continuation to real replica number.
  • Replica computation: Replica-symmetric vectors reduce the replicated transfer matrix to a finite-dimensional subspace whose largest eigenvalue controls the calculation.The construction uses vectors indexed by the reference spin and the number of replicas carrying an up spin.
  • Transfer-matrix analysis: The transfer-matrix kernel is built from temperature-dependent factors associated with the two Ising components and their replicated eigenvectors.The leading eigenvalue is expanded around the n = 0 kernel to obtain the quantity needed for the free-entropy calculation.
  • Thermodynamic limit: For inverse temperatures β_r and β_s, the thermodynamic-limit calculation uses zero-field two-point correlations and connected correlations.The relevant correlation sums are evaluated over ordered pairs at fixed separations in a periodic chain.
  • Thermodynamic limit: The resulting correlation expression is given analytically in terms of tanh(β_r) and tanh(β_s).The expression combines each component’s squared correlation contribution with a cross-component term.

Appendix C.4. Large N

The large-N treatment combines exact Ising-mixture calculations with an asymptotic Gaussian approximation for the fluctuation side of the speciation criterion. This approximation limits direct accuracy to sufficiently large systems, while finite-size experiments supply estimates at smaller sizes.

  • Large-N approximation: For 1D Ising mixtures, the average free-entropy difference is computed analytically, but the corresponding fluctuation term lacks an analytical counterpart.The fluctuation term is therefore approximated by its large-time Gaussian asymptotic value.
  • Large-N approximation: The Gaussian approximation is valid when speciation times are sufficiently large, namely for sufficiently large N.Because U-turn experiments are expensive at large N, experiments were limited to N up to 3200 and used to estimate the fluctuation term accurately there.
  • Transfer-matrix evaluation: The exact score function and its tilted averages are evaluated using transfer-matrix methods for each value of the diffused configuration.The denominator contains standard partition functions, while the numerator contains transfer-matrix terms with an inserted spin-dependent factor.
Loading 2602.04404v1…