Source-linked AI summary
Convergence of denoising diffusion models under the manifold hypothesis
Valentin De Bortoli
TL;DR
Existing diffusion-model theory largely assumes a Lebesgue density, excluding empirical distributions and targets supported on lower-dimensional manifolds. The paper studies diffusion-model convergence in this setting and derives Wasserstein-1 bounds, showing that such models can recover low-dimensional manifold-supported target distributions. The general bounds have an exponential dependence on 1/ε, which the authors identify as potentially overly pessimistic.
Problem
Existing convergence analyses assume the target distribution has a Lebesgue density, excluding empirical measures and distributions supported on lower-dimensional manifolds.
Method
The paper analyzes a forward noising process and its approximate backward diffusion under the manifold hypothesis, deriving quantitative Wasserstein-1 convergence bounds.
Results
The results provide convergence guarantees showing that diffusion models can recover target distributions defined on low-dimensional manifolds.
Takeaways & Limitations
Diffusion models admit theoretical convergence guarantees beyond absolutely continuous targets, including distributions supported on low-dimensional manifolds.
Takeaways & Limitations
The general convergence bound has an exponential dependence on 1/ε that may be overly pessimistic; improving it requires additional Hessian conditions whose realism remains unresolved.
Abstract
from arXiv · showhide
Denoising diffusion models are a recent class of generative models exhibiting state-of-the-art performance in image and audio synthesis. Such models approximate the time-reversal of a forward noising process from a target distribution to a reference density, which is usually Gaussian. Despite their strong empirical results, the theoretical analysis of such models remains limited. In particular, all current approaches crucially assume that the target density admits a density w.r.t. the Lebesgue measure. This does not cover settings where the target distribution is supported on a lower-dimensional manifold or is given by some empirical distribution. In this paper, we bridge this gap by providing the first convergence results for diffusion models in this more general setting. In particular, we provide quantitative bounds on the Wasserstein distance of order one between the target data distribution and the generative distribution of the diffusion model.
1 Introduction
Diffusion models reverse a forward noising process and achieve state-of-the-art image and audio synthesis, but their theoretical convergence analysis has largely excluded manifold-supported and empirical target distributions. This work addresses that gap by deriving quantitative Wasserstein-1 convergence bounds under the manifold hypothesis.
- Diffusion modeling: Diffusion models add noise until reaching a Gaussian distribution, then approximate the associated backward process for generation.The backward drift requires estimating the forward logarithmic density gradient, or Stein score.
- Existing theory: Despite strong empirical results, theoretical understanding and convergence analysis of diffusion models remain limited.
- Theoretical gap: Prior analyses impose score regularity and density assumptions that fail for lower-dimensional manifold-supported or empirical target distributions.Under the manifold hypothesis, the target is supported on a lower-dimensional compact set and the score explodes as t →0.
- Related work: Pidstrigach (2022) established well-definedness and support equivalence for the continuous backward process under integrability conditions on score-estimation error.
- Contribution: This work studies convergence rates under the manifold hypothesis and derives quantitative Wasserstein-1 bounds between target and generative distributions.
2 Diffusion models for generative modeling
The model uses an Ornstein–Uhlenbeck forward noising process whose time reversal is approximated by a learned score and then discretized for implementation. Because the score can explode near t →0 under manifold or empirical targets, the analysis truncates the backward integration at T −ε.
- Forward process: The forward process is an Ornstein–Uhlenbeck process that transforms the target distribution toward a Gaussian reference distribution.The target is π, while the reference is π∞ = N(0, I_d).
- Backward process: The backward process reverses the forward trajectory and uses the gradient of the forward density's logarithm.
- Score estimation: Neural score estimators are trained by denoising score matching, approximated in practice with Monte Carlo samples.
- Discretization: The learned continuous-time backward process is discretized with stepsizes whose sum equals T to obtain an implementable algorithm.
- Truncation: Under manifold or empirical targets, score explosion near t →0 motivates truncating backward integration at T −ε and disregarding the final sample.Discrete-time diffusion models embed this truncation in their discretization scheme.
3 Main results
The paper establishes Wasserstein-1 convergence guarantees for diffusion models when data lie on compact lower-dimensional supports or are empirical measures, accommodating score explosion near zero time. The bounds separate truncation, finite-time, discretization, score-approximation, and statistical errors, while exploiting lower-dimensional structure.
- Assumptions: A1 assumes the target distribution is supported on a compact set M, encompassing lower-dimensional manifold distributions and empirical measures.The compactness assumption is motivated by finite image ranges and includes empirical distributions of the form (1/N) Σ_i δ_Xi.
- Assumptions: A3 allows the score estimator to grow as t →0 and as ∥x∥→∞, matching the explosive behavior of the true score under the manifold hypothesis.The paper contrasts this with uniform-in-time and space score bounds, and notes practical parameterizations that account for the explosion.
- Convergence bounds: Theorem 1 bounds W1(L(YK), π) by discretization and score-approximation error, finite forward integration error, and terminal-noise error.The three contributions have orders O(exp[κ/ε](M + δ1/2)/ε2), O(exp[κ/ε] exp[−T/β̄]), and O(ε1/2), respectively.
- Convergence bounds: W1(L(YK), π) →0 when T→∞, δ,M→0, and then ε→0, with δ1/2 and M entering linearly and T entering through exp[−T/β̄].The truncation parameter ε also induces an exp[κ/ε] factor in the general bound.
- Convergence bounds: The constant D0 depends only on β̄, diam(M), and d, with dimension dependence O(d) and diameter dependence O(diam(M)^4) up to logarithmic factors.For a p-dimensional hypercube, diam(M)=√p, so the diameter may reflect intrinsic rather than ambient dimension.
- Statistical guarantees and empirical measure targets: Wasserstein distance remains informative for lower-dimensional targets, unlike total variation and KL divergence, whose bounds can be vacuous when supports differ.The analysis does not quantify sample diversity, which is left for future work.
- Convergence bounds: Theorem 3 replaces Theorem 1’s exponential dependence on ε with polynomial dependence under a Hessian condition, verified for uniform distributions on hypercubes.Under suitable smoothness assumptions, the Hessian condition has strong geometric implications and implies convexity of M.
- Statistical guarantees and empirical measure targets: For empirical targets on a manifold, the expected W1 error adds a statistical term D1N^−1/(dM(M)+η) to the diffusion-model error bound.Here dM(M) denotes the Minkowski dimension of M.
4 Proof of Theorem 1
The proof decomposes the W1 error into discretization, backward-process approximation, and noising components, then bounds each using stochastic-flow and coupling arguments.
- Error decomposition: The proof controls the target error by separately bounding discretization, backward-process, and noising errors.The discretization and approximation term is treated first, followed by backward-process convergence and forward-noising error.
- Discretization analysis: The stochastic interpolation formula relates continuous backward flows to their discretized counterparts through drift differences.The proof introduces stochastic flows, their interpolation, and the drift discrepancy between the exact and discretized processes.
- Backward-process stability: Tangent-process bounds control how perturbations in initial conditions propagate through the backward dynamics and determine Wasserstein stability.The tangent process encodes local variation with respect to initialization and supports bounds on distances between backward processes.
- Backward-process stability: The proof identifies a contractive interval for the backward process and combines this stability with local error estimates.The contractive regime is controlled up to a time t⋆, after which the argument uses additional growth bounds.
- Quantitative bound: The discretization contribution is bounded by 2C0 exp[diam(M)^2(1 + β̄)/(2ε)](1 + β̄)^3(1 + log(1 + diam(M)))(M + δ^1/2)/ε^2.This is the stated bound for W1(π∞Q_tK, π∞R_K).
5 Conclusion
The paper establishes W1 convergence guarantees for diffusion models targeting low-dimensional manifolds, while identifying exponential truncation dependence and broader extensions as open issues.
- Conclusion: The results provide convergence guarantees in Wasserstein distance of order one for target distributions supported on low-dimensional manifolds.The conclusion states that diffusion models can recover target distributions defined on low-dimensional manifolds.
- Limitations: The general error bound has an exponential dependency on 1/ε, which may be overly pessimistic.The authors state that polynomial dependency is possible under additional Hessian conditions.
- Future directions: The analysis focuses on the Ornstein–Uhlenbeck forward noising process and could be extended to other forward diffusions and discretization frameworks.The authors specifically mention critically-damped diffusions and predictor-corrector schemes as possible extensions.
- Future directions: The paper leaves open analogous bounds for target distributions with full-dimensional support and tail constraints.This is identified as a challenge for extending the theoretical analysis.
- Future directions: The relationship between manifold geometry and score-function properties remains incompletely understood.Preliminary results indicate that convexity may be recoverable from score properties, but broader geometric conclusions remain unclear.
D.3 Stability and Lipschitz properties of the backward processes
This section derives moment and stability controls for backward processes, using dissipativity, Hessian bounds, tangent-process estimates, and contraction arguments.
- Moment control: Dissipativity conditions control moments of the introduced backward processes.The section uses these conditions to establish uniform moment bounds.
- Moment control: E[∥Y_k∥^2] is bounded by K = d + B(δ, M, η, diam(M))(1/A(δ, M, η, diam(M)) + δ).The bound holds under the assumptions of Lemma D.5 and the stated positivity condition on A.
- Moment control: The recursion shows that moments decrease above B/A and remain bounded near B/A below that threshold.This yields a recursive uniform bound over discretization steps.
- Contraction: For diam(M)=0, the backward diffusion converges in finite time regardless of the initialization distribution.The section states δ_xQ_T = π for every x.
- Contraction: For nonzero manifold diameter, contraction is obtained only up to a certain point, and tangent-process control bounds Wasserstein distances.The tangent process is used to control growth or contraction and thereby distances between backward processes.
F Additional comments on Theorem 1
The section discusses the validity of assumption A3 and comments on the suboptimality of Theorem 1’s bound.
- Additional comments: The authors examine whether assumption A3 is valid and discuss why Theorem 1’s bound may be suboptimal.The passage introduces these as the section’s two purposes.
F.1 Validity of A3
The section examines whether a relative-error condition on the score estimator supports A3 and illustrates explosive score and estimation-error behavior in singular targets.
- Validity of A3: A3 follows from a relative score error bound of the form ∥s(t, xt) − ∇log pt(xt)∥ ≤ Mr∥∇log pt(xt)∥.This implication uses A1, Lemma C.1, and M = 4Mr(1 + diam(M)).
- Validity of A3: For a Dirac target at zero, the score-estimation error becomes explosive as ∥x∥→+∞ and t→0.The error is evaluated at query points in a two-dimensional setting.
- Validity of A3: Figure 1 plots true-score and score-error norms across time and space, with the second spatial coordinate fixed at zero.Both panels use T = 1, with time evolution on the x-axis and first-coordinate spatial evolution on the y-axis.
- Validity of A3: For a uniform distribution on two concentric circles, Figure 2 compares target and generated samples, trajectories, and estimated-score norms.The score norm is shown over time and along the first coordinate with the second coordinate fixed at zero.
- Validity of A3: In both settings, the score network is trained with Denoising Score Matching and the ADAM optimizer.The architecture and training settings follow those used in De Bortoli et al. (2021a).
F.2 Suboptimality of the bound
This section compares the theorem’s Wasserstein bound with a naive initialization bound and shows that backward dynamics can worsen the distance under poor score estimation.
- Suboptimality of the bound: The naive bound on W1(L(Y0), π) can be smaller than the theorem’s bound, particularly for large D0.In that regime, the derived bound may initially appear vacuous because backward diffusion does not improve the Wasserstein distance.
- Suboptimality of the bound: With s = 0, the backward process remains Gaussian with law N(0, ((3 exp[2t] −1)/2) Id).This initialization approximately arises from a fully connected network with a linear last layer and no non-linearity.
- Suboptimality of the bound: For sufficiently large T, W1(L(ŶT), π) ≥ W1(L(Ŷ0), π), so backward dynamics can steer the Gaussian away from π.The explanation given is the explosive behavior of the backward Ornstein–Uhlenbeck process versus the contractive forward process.
- Suboptimality of the bound: The backward process is analyzed using practical constant, linear, and cosine noise schedules.The generalized cosine schedule is softened to be differentiable, and the schedules satisfy A2.
- Suboptimality of the bound: Euler–Maruyama analysis additionally requires the schedule s 7→βs to satisfy a Lipschitz property.This requirement concerns discretization of the approximate backward process.
H A short proof of the results of Franzese et al. (2022)
This appendix derives the Franzese et al. (2022) result by rearranging an ELBO identity and connecting score expectations to the forward-process score.
- H A short proof of the results of Franzese et al. (2022): The derivation begins by rearranging the ELBO result of Huang et al. (2021).
- H A short proof of the results of Franzese et al. (2022): Expanding the square and conditioning yields E[∇log pt|0(Xt|X0) | Xt] = ∇log pt(Xt).
- H A short proof of the results of Franzese et al. (2022): Combining the resulting identity with the ELBO expression recovers the equivalent formulation associated with Song et al. (2021a, Theorem 1).The cited theorem uses data processing, conditional KL decomposition, and Girsanov’s theorem.
I Wasserstein controls under L2 errors
This section extends Wasserstein analysis to L2 score errors by replacing A3 with weaker control, proving moment bounds, and deriving a counterpart theorem under explicit assumptions.
- I Wasserstein controls under L2 errors: The appendix replaces A3 with a weaker score-error control and develops the corresponding Wasserstein analysis.The extension targets Wasserstein distance of order one under weaker growth conditions than prior work.
- I Wasserstein controls under L2 errors: The derivation assumes Lebesgue densities and well-defined integrals, and adapting logarithmic-Sobolev-based controls is left for future work.The latter issue prevents a straightforward adaptation of Lee et al. (2022) to this Wasserstein setting.
- I Wasserstein controls under L2 errors: Theorem I.1 assumes A1, A2, A4, and A5, with T ≥ 2β̄(1 + log(1 + diam(M))), tK = T − ε, and ε, M, M/ζ, δ ≤ 1/32.Its dimensional and geometric constant is D0 = D(1 + β̄)^5(1 + d + diam(M)^4)(1 + log(1 + diam(M))).
- I Wasserstein controls under L2 errors: Under A1, A2, and A5, Lemma I.2 bounds E[∥Yk∥2] by K = d + B(δ, M, η, diam(M))(1/A(δ, M, η, diam(M)) + δ).The bound applies when the stated δ and positivity conditions hold.
- I Wasserstein controls under L2 errors: The proof controls score approximation and discretization through moment estimates and a comparison process.The argument extends a total-variation result of Lee et al. (2022) to Wasserstein distance and requires moment control under L2 error.
- I Wasserstein controls under L2 errors: The appendix also proves an improved theorem under tighter Hessian conditions, but those conditions are never satisfied on non-convex sets under appropriate smoothness assumptions.
J.1 Proof of Theorem 3
The proof develops auxiliary propositions under assumptions A1, A2, A3, A4, A6 and a truncated backward diffusion endpoint, then combines them to prove Theorem 3. Its bounds replace an exponential dependence on 1/ε with a polynomial one.
- Assumptions: Assumption A6 bounds the Hessian of log p_t by Γ/σ^2 for all t ∈ (0,T] and x_t ∈ R^d.
- Intermediate bounds: Proposition J.1 establishes an intermediate bound for the backward process up to any endpoint t_K < T under A1 and A6.
- Proof combination: The appendix explicitly contrasts an exponential dependency in 1/ε with a polynomial dependency in σ^-2Γ.
- Intermediate bounds: Proposition J.2 assumes A1–A4, A6, t_K = T − ε, and ε, δ, M ≤ 1/32 to provide the main auxiliary estimate.
- Proof combination: The proof uses bounds on T − t_K and t_K − t⋆, then combines propositions and the inequality 1 + a ≤ exp[a].
J.2 Hessian bounds for the uniform distribution
This section proves Hessian bounds for the smoothed uniform distribution on a coordinate cube by deriving coordinatewise formulas and controlling their second derivatives.
- Main bound: For the uniform distribution on [−1/2, 1/2]^p, Proposition J.4 asserts a constant Γ with ∥∇^2 log p_t(x_t)∥ ≤ Γ/σ^2.
- Derivation: The proof obtains a closed-form expression for p_t and uses the resulting diagonal structure of ∇^2 log p_t to reduce the analysis to coordinatewise second derivatives.
- Derivation: For coordinates inside the uniform support, the proof represents the relevant factor as F_t(Φ,a,b)=Φ(a+b)−Φ(a−b) and differentiates its logarithm.
- Main bound: The resulting bounds are extended across the remaining cases by symmetry and continuity, yielding a uniform constant bound on the coordinatewise second derivatives.
J.3 The role of convexity
The section connects Hessian growth of smoothed manifold-supported distributions to convexity, projection geometry, and curvature conditions. Convex manifolds satisfy the stated σ^-2 scaling, while nonconvex cases require additional geometric control.
- Projection geometry: A Chebyshev set is defined by uniqueness of nearest-point projection, and closed convex sets are exactly the Chebyshev sets.
- Curvature: The shape operator encodes local manifold geometry and appears in the curvature condition used for nonconvex manifolds.
- Theorem J.8: Theorem J.8 applies to smooth manifolds whose distribution has a smooth density with respect to Hausdorff measure.
- Convex case: For convex manifolds, the theorem provides a σ^-2-order upper scaling for the Hessian of log p_t as t approaches zero.
- Nonconvex case: The nonconvex branch assumes a point with multiple nearest projections and an additional curvature condition at those projection points.
- Scope: The authors state that the curvature condition may be relaxable, while uniform-in-space bounds needed to match A6 are left for future work.
- Scope: They conclude that nonconvex sets satisfying the stated negative-curvature restriction do not satisfy A6.