Source-linked AI summary

Stochastic Interpolants: A Unifying Framework for Flows and Diffusions

Michael S. Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden

arXiv:2303.08797v4cs.LGcond-mat.dis-nnmath.PR

TL;DR

The paper addresses the need for generative models that bridge arbitrary densities exactly in finite time while supporting both deterministic and stochastic dynamics. It develops stochastic interpolants with flexible latent-variable designs, derives their transport and Fokker–Planck formulations and learning objectives, and reports that positive diffusion empirically performs better in the studied experiments. The framework also provides likelihood estimators and formulations connected to score-based diffusion, rectifying flows, and Schrödinger bridges.

  • Problem

    Existing diffusion methods commonly map to a Gaussian and are theoretically exact only over infinite time, motivating finite-time bridges between arbitrary densities and understanding deterministic–stochastic quality differences.

  • Method

    The paper builds stochastic interpolants from endpoint samples and a latent variable, then derives transport and tunable-diffusion Fokker–Planck equations with quadratic objectives for velocities, scores, and denoisers.

  • Results

    Positive diffusion empirically performs better in the reported density-matching experiments, while designed interpolants also improve deterministic probability flow over the original interpolant.

  • Takeaways & Limitations

    The framework separates bridge construction from sampling dynamics, enabling flexible deterministic or stochastic generative models with likelihood and cross-entropy estimators.

  • Takeaways & Limitations

    Numerical investigation of the explicitly optimized Schrödinger-bridge formulation is left for future work.

Abstract

from arXiv · show

A class of generative models that unifies flow-based and diffusion-based methods is introduced. These models extend the framework proposed in Albergo and Vanden-Eijnden (2023), enabling the use of a broad class of continuous-time stochastic processes called stochastic interpolants to bridge any two probability density functions exactly in finite time. These interpolants are built by combining data from the two prescribed densities with an additional latent variable that shapes the bridge in a flexible way. The time-dependent density function of the interpolant is shown to satisfy a transport equation as well as a family of forward and backward Fokker-Planck equations with tunable diffusion coefficient. Upon consideration of the time evolution of an individual sample, this viewpoint leads to both deterministic and stochastic generative models based on probability flow equations or stochastic differential equations with an adjustable level of noise. The drift coefficients entering these models are time-dependent velocity fields characterized as the unique minimizers of simple quadratic objective functions, one of which is a new objective for the score. We show that minimization of these quadratic objectives leads to control of the likelihood for generative models built upon stochastic dynamics, while likelihood control for deterministic dynamics is more stringent. We also construct estimators for the likelihood and the cross entropy of interpolant-based generative models, and we discuss connections with other methods such as score-based diffusion models, stochastic localization, probabilistic denoising, and rectifying flows. In addition, we demonstrate that stochastic interpolants recover the Schrödinger bridge between the two target densities when explicitly optimizing over the interpolant. Finally, algorithmic aspects are discussed and the approach is illustrated on numerical examples.

1 Introduction

The paper introduces stochastic interpolants, which bridge arbitrary endpoint densities exactly in finite time while separating bridge design from deterministic or stochastic sampling. It develops the associated transport and Fokker–Planck equations, learning objectives, likelihood formulas, and connections to existing generative-modeling methods.

  • Framework: Stochastic interpolants continuously bridge samples from arbitrary densities ρ0 and ρ1, exactly matching ρ0 at t = 0 and ρ1 at t = 1.The bridge combines endpoint data with a latent variable, allowing flexible path design.
  • Framework: The interpolant density satisfies a transport equation and forward and backward Fokker–Planck equations with freely tunable diffusion coefficients.These equations generate deterministic ODE and stochastic SDE samplers sharing the same time-dependent density.
  • Learning: Velocity fields are learned as unique minimizers of quadratic objectives, including a new score objective and a denoiser objective related to the score.The framework expresses the relevant drift and score quantities through tractable conditional expectations.
  • Likelihood: Likelihood control is sufficient for SDE-based models when their drift is regressed, whereas ODE-based models additionally require Fisher-divergence minimization.The diffusion coefficient can be tuned to maximize SDE likelihood, and the paper derives likelihood and cross-entropy estimators.
  • Connections: The framework includes one-sided, mirror, and Schrödinger-bridge formulations, and connects stochastic interpolants with score-based diffusion, stochastic bridges, and rectifying flows.It avoids the generically unknown Doob h-transform for a broad class of stochastic bridges.

1. The forward Fokker-Planck equation

The forward Fokker–Planck equation is well-posed from ρ0 to ρ1, and its associated stochastic processes provide generative-model constructions with tunable diffusion.

  • The forward Fokker–Planck equation solved from ρ(0)=ρ0 reaches ρ(1)=ρ1.
  • Forward and backward Fokker–Planck equations are more robust than the transport equation to velocity approximation errors.This robustness has practical implications for generative models.
  • Time-integrated quadratic objectives allow global parameterization of the drift and score over [0,1]×R^d for numerical estimation.The objectives can be empirically estimated from samples of (x0,x1) and the resulting interpolant states.
  • The framework supports deterministic and stochastic generative models derived from transport and Fokker–Planck equations.
  • The interpolant law coincides at every time with laws generated by the associated transport and Fokker–Planck processes.

1. The solutions of the probability flow associated with the transport equation (2.9)

The probability-flow construction links transport and Fokker–Planck generative dynamics, while likelihood control differs sharply between stochastic and deterministic models.

  • The stochastic interpolant’s law is shared by deterministic probability-flow dynamics and forward or backward stochastic dynamics.Although the processes differ, their laws coincide at each time.
  • Minimizing the stated losses controls likelihood for learned Fokker–Planck generative models.
  • Matching the drift is insufficient in general to control KL divergence for deterministic transport equations.Additional control of score-related Fisher divergence is required.
  • Diffusion makes drift errors sufficient to control KL divergence between Fokker–Planck solutions by contributing a negative divergence term.This removes the need for explicit Fisher-divergence control.
  • The deterministic model does not generally obtain likelihood control from minimizing the drift objective alone.
  • The forward and backward stochastic differential equations can be used with approximate drifts to generate samples and evaluate model quality.Cross-entropy formulas provide additional evaluation tools for evolved densities.

1. The solution to the forward FPE

The forward Fokker–Planck construction supports cross-entropy estimation, empirical comparison, and practical diagnostics, but some estimators introduce uncontrolled approximation issues.

  • The resulting cross-entropy identities can test sample quality for ODE and forward or backward SDE generative models.
  • Cross-entropy formulas are available for forward and backward Fokker–Planck model densities relative to their endpoint target densities.
  • Empirical endpoint expectations enable cross-validation of drift and score approximations and comparison of transport- and Fokker–Planck-based models.
  • Taking logarithms of estimated expectations can create difficulties when divergence terms are computed with Hutchinson’s trace estimator.Jensen’s inequality removes the bias but produces upper bounds that are not sharp in general.
  • Using the learned score as a proxy for the model-density score may be useful in practice but is uncontrolled in general.

3 Instantiations and extensions

The framework instantiates stochastic interpolants as diffusive, one-sided, and related bridge constructions, deriving deterministic and stochastic generative dynamics with flexible diffusion. It also supports point-mass conditional sampling, denoiser-based models, and Schrödinger bridge formulations.

  • Diffusive interpolants: Diffusive interpolants have the same time-dependent density as Brownian-bridge constructions while enabling direct sampling from endpoint variables and Gaussian noise.The Brownian-bridge process is continuous but not time-differentiable, whereas the equivalent stochastic interpolant is directly sampled from x0, x1, and z.
  • Diffusive interpolants: The interpolant density satisfies a transport equation and forward/backward Fokker–Planck equations, yielding ODE and SDE generative models with adjustable diffusion.The diffusion coefficient can be varied, and the associated equations share the interpolant density evolution.
  • Conditional sampling: A diffusive interpolant can sample the target density from a point mass because its drift remains nonsingular at t = 0, unlike the corresponding velocity and score fields.The associated forward SDE provides a generative model from a base measure concentrated at a single x0.
  • Related constructions: The framework connects stochastic interpolants with score-based diffusion and denoising, including constructions that reparameterize diffusion models and expose the score through denoisers.The broader framework also includes a generalized Gaussian base density construction.
  • One-sided interpolants: One-sided stochastic interpolants bridge N(0, Id) and ρ1, and their score and velocity fields can be expressed through a learned denoiser.Because the score depends on the denoiser, the denoiser is the only quantity that needs to be learned in this construction.
  • Schrödinger bridges: Optimizing over the interpolant recovers the Schrödinger bridge density, with the associated optimizing velocity represented as u = ∇λ.In the zero-diffusion limit, the minimizing velocity formally reduces to the transport velocity, while solving the bridge problem numerically is left for future work.

4 Spatially linear interpolants

Spatially linear interpolants combine endpoint samples with a Gaussian latent variable, offering flexible bridges whose intermediate densities and generative trajectories can be shaped independently through γ(t) and ϵ(t). These choices can smooth densities and velocities, eliminate spurious modes, and support deterministic or stochastic generation while preserving exact endpoint transport.

  • Linear stochastic interpolants: The interpolant x_t = α(t)x_0 + β(t)x_1 + γ(t)z uses independently drawn endpoint variables and Gaussian noise, with coefficient constraints fixing the endpoint densities.The standard conditions are α(0) = β(1) = 1, α(1) = β(0) = γ(0) = γ(1) = 0, and γ(t) > 0 for interior times.
  • Factorization of the velocity field: The velocity and score decompose through conditional expectations of endpoint variables and the latent noise: b = α̇η_0 + β̇η_1 + γ̇η_z and s = −γ^-1η_z.The three conditional expectations satisfy αη_0 + βη_1 + γη_z = x, so one can be recovered from the other two.
  • Coefficient design: The constraint α^2(t) + β^2(t) + γ^2(t) = 1 preserves identity covariance throughout the interpolation when both endpoint distributions are standardized.For similarly scaled but non-identity covariances, the coefficient sum need only remain of order one rather than satisfy the constraint exactly.
  • Specific design choices: Adding γ(t)z smooths intermediate densities and suppresses duplicated complex features, making the velocity field smoother and simplifying its estimation.With γ(t) = 0, distinct endpoint mixture features can produce spurious intermediate modes; latent noise partially or completely removes them for suitable choices.
  • Specific design choices: The Gaussian encoding-decoding design maps ρ0 into standard Gaussian noise by t = 1/2 and then decodes it into ρ1 while remaining a single continuity-equation transport.Its probability-flow solution is bijective between initial and final samples, and analogous correlations remain for the forward and backward SDEs.
  • Impact of γ(t) and ϵ(t): The diffusion coefficient ϵ(t) leaves ρ(t) unchanged but controls sampling stochasticity, whereas γ(t) changes the density and can affect deterministic dynamics through the velocity field.Nonzero ϵ also provides likelihood control under approximate drift and score fields, while endpoint score contributions depend on the γ design.

5 Connections with other methods

The framework connects stochastic interpolants to score-based diffusion, denoising, and rectified flows while addressing finite-time and singularity issues. It also establishes exact transport properties and clarifies when straight-line flows do not imply optimal transport.

  • Score-based diffusion models: Stochastic interpolants connect to score-based diffusion through a time reparameterization that removes finite-time singularities.The resulting generative models operate on t ∈ [0, 1] without the bias introduced by truncating diffusion time.
  • Score-based diffusion models: Unlike naively time-changed score-based diffusion, interpolants separate density construction from sample-generation dynamics, yielding nonsingular coefficients at endpoints.This avoids the t = 0 singularity that makes the diffusion formulation difficult to solve exactly from the Gaussian endpoint.
  • Denoising methods: Iterating conditional denoising updates with infinitesimal steps produces a consistent integration scheme for the interpolant’s probability-flow ODE.The scheme links denoising formulas to the velocity field of the probability flow.
  • Rectified flows: Rectification constructs a simpler flow with the same endpoint map as the original probability flow, and linear choices can reduce trajectories to straight lines.The resulting map still pushes ρ0 onto ρ1 under the stated invertibility assumption.
  • Rectified flows: Straight-line solutions can exactly transport ρ0 to ρ1 without being the optimal transport map.Thus, straight paths are necessary but not sufficient for optimal transport.

6 Algorithmic aspects

The paper presents practical procedures for learning drift, score, and denoiser models and for sampling with ODEs or SDEs. It emphasizes endpoint stability, diffusion tuning, and denoiser-based parameterizations.

  • Practical choices: Algorithmic choices are equivalent without numerical or statistical error but can produce different practical models when those errors are present.The paper therefore treats learning and sampling choices as application-dependent design decisions.
  • Endpoint stability: Antithetic sampling keeps both conditional mean and variance finite near endpoints despite the singularity of 1/γ(t).The method is presented as necessary for stable training of objectives involving γ^-1(t).
  • Learning objectives: Learning the denoiser ηz is numerically stable across t ∈ [0, 1], but converting it to a score can make endpoint drifts singular.The paper recommends endpoint capping or tuning ϵ(t) to avoid this issue.
  • Sampling: Sampling uses learned drift and score estimates in either an ODE or an SDE, with diffusion coefficient ϵ(t) controlling stochasticity.Setting ϵ(t) = 0 gives the probability-flow ODE, while positive diffusion yields SDE sampling.
  • Denoiser parameterization: For spatially-linear one-sided interpolants, learning only the denoiser ηz is sufficient to represent the probability-flow velocity field.This enables ODE- and SDE-based models on [0, 1] using a single learned denoiser without singularity.

7 Numerical results

Numerical experiments examine diffusion, interpolant, and estimator choices on challenging densities and image data. They find benefits from stochastic sampling and denoiser-based learning, while image experiments demonstrate scalability and diversity.

  • Two-dimensional density estimation: For the checkerboard target, positive diffusion empirically performs better than deterministic sampling, with the smallest gap for the specified γ choice.The comparison uses ODE sampling and SDE sampling with ϵ = 0.5, 1.0, or 2.5.
  • Gaussian mixtures: The high-dimensional experiment maps N(0, Id) to a five-mode Gaussian mixture in dimension d = 128 using a linear interpolant.The setup evaluates four combinations of learned drift and score or denoiser functions.
  • Gaussian mixtures: The best Gaussian-mixture performance occurs when learning b and ηz together with a properly chosen ϵ > 0.The experiment compares learning b or v and s or ηz.
  • Gaussian mixtures: Small diffusion overestimates modal density and underestimates tails, whereas excessive diffusion reverses this pattern.An intermediate diffusion level gives each model its optimal performance.
  • Image generation: Image experiments on 128 × 128 Oxford flowers demonstrate scalable one-sided and mirror interpolants, while the authors defer standard benchmark evaluation to future work.SDE diffusion increases diversity from a fixed input, and nearest-neighbor comparisons show visually distinct generated images.

5 Nearest neighbors in training set

The image-generation examples use nearest-neighbor comparisons and mirror-interpolant trajectories to examine whether outputs reproduce training images or generate nearby alternatives.

  • Nearest-neighbor comparison: A generated image is compared with five nearest training-set neighbors using ℓ1 distance.The neighbors are visually distinct from the generated image, supporting the paper’s no-memorization observation.
  • Mirror interpolant: With the mirror interpolant, SDE evolution resamples a dataset flower into a proximal flower not seen in the dataset.The ODE preserves the input image at t = 1, whereas SDE noise enables new outputs from the same input.

8 Conclusion

The paper presents stochastic interpolants as a general framework for constructing deterministic and stochastic generative models that map between two densities exactly in finite time.

  • The framework provides mathematical theory and efficient algorithms for deterministic and stochastic generative models connecting two densities exactly in finite time.

Appendix A. Bridging two Gaussian mixture densities

For Gaussian endpoint densities and Gaussian-mixture settings, the interpolant yields regular densities and tractable velocity and score fields. The Gaussian case produces explicit time-dependent mean and covariance formulas and linear evolution equations.

  • For Gaussian densities, the velocity and score fields are approximately linear and grow at most linearly in x.
  • Gaussian endpoint densities produce interpolant mean m(t) = α(t)m0 + β(t)m1 and covariance C(t) = α2(t)C0 + β2(t)C1 + γ2(t)Id.
  • In the spatially linear case, the probability flow ODE and associated forward SDE become linear evolution equations.
  • The interpolant density is positive and smooth, and its velocity field has corresponding continuity and regularity properties.
  • The velocity and score fields are characterized as unique minimizers of quadratic objectives.

In addition it satisfies

The interpolant density satisfies score-based formulations that yield forward and backward Fokker–Planck equations, with corresponding stochastic dynamics solvable in the appropriate time direction.

  • The score is represented by the conditional expectation s(t, x) = −γ−1(t)E(z|xt = x).
  • The forward Fokker–Planck equation is well-posed from ρ0 and reaches ρ1 at t = 1.
  • The backward Fokker–Planck equation is well-posed from ρ1 and recovers ρ0 at t = 0.
  • The score is the minimizer of a quadratic objective derived using integration by parts.

B.3 Proofs

The proofs establish likelihood-control results for stochastic and deterministic dynamics by comparing the relevant Fokker–Planck or transport evolutions.

  • For stochastic dynamics with positive diffusion, the terminal Kullback–Leibler divergence is expressed through the corresponding Fokker–Planck evolutions.
  • Theorem 23 relates terminal-density divergence to the velocity and score objective functions for stochastic dynamics.
  • Deterministic transport solutions are represented using characteristics generated by the velocity-field ODE.
  • Forward and backward characteristic representations provide solutions from either the initial or final density condition.
  • The stochastic forward and backward models use drifts formed by adding or subtracting the score term from the velocity field.

1. The solution to the forward FPE

The forward Fokker–Planck equation is solved explicitly for an interpolant that remains fixed initially and receives Gaussian noise. Its score and drift are nonsingular near the initial time, and diffusion can later be removed.

  • The corresponding density satisfies a forward Fokker–Planck equation whose score is explicitly s(t, x) = −(x −x0)/(2at(1 −t)).
  • The score-induced drift remains nonsingular at t = 0, so the associated SDE is well-defined on [0, δ] and matches the interpolant law.
  • The construction requires an SDE initially so mass can spread from x0, but diffusion is unnecessary after δ and can be replaced by a probability-flow ODE.

B.6 Proof of Lemma 40 and Theorem 41.

The proof establishes that the constrained max–min optimization recovers the density solving the prescribed Fokker–Planck problem, with the optimizer’s velocity equal to a gradient potential. It also verifies consistency and endpoint correctness for denoising and probability-flow constructions.

  • B.6 Proof of Lemma 40 and Theorem 41.: The interpolant construction combines transformed endpoint variables with Gaussian noise while enforcing α2(t) + β2(t) + γ2(t) = 1.
  • B.6 Proof of Lemma 40 and Theorem 41.: Under the stated assumptions, every optimizer of the max–min problem produces an interpolant density equal to the solution of the constrained Fokker–Planck equation.
  • B.6 Proof of Lemma 40 and Theorem 41.: The optimizer’s velocity field is u = ∇λ, where λ is the Lagrange multiplier solving the Euler–Lagrange system.
  • Algorithmic consequences: The denoising update (5.14) is a consistent integration scheme for the probability-flow equation, including when the time grid is nonuniform.
  • Algorithmic consequences: The reconstructed probability-flow ODE transports samples from ρ0 to ρ1, and under Assumption (46) its velocity field simplifies to the stated form.
  • Algorithmic consequences: Image experiments use U-Net architectures for learning b, v, s, or η, while SDE denoising performs best when integration stops slightly before tf = 1.0.
Loading 2303.08797v4…