Source-linked AI summary
Renormalization Group Flow Matching for Scalable Local Generative Modeling
Kanta Masuki, Yuto Ashida
TL;DR
High-dimensional generative models require costly computation, while local models struggle with long-range structure. RGFM uses renormalization-group scale separation and quasi-locality to generate data across scales, reproducing long-range correlations and global structure with near-linear local computation.
Problem
Scaling diffusion and flow-based generative models to high-dimensional data remains computationally costly because expressive networks require repeated, memory-intensive evaluations.
Method
RGFM uses an exact renormalization-group flow with scale separation and quasi-locality, tracking generation through local velocity fields whose patch size grows as O(ln L).
Results
Local RGFM reproduces long-range correlations and global structures beyond its receptive field, outperforming local flow matching on one-dimensional distributions and FFHQ images.
Takeaways & Limitations
RG-guided probability flows support scalable local generation that captures long-range structure using receptive fields smaller than the full system size.
Takeaways & Limitations
At 256 × 256 resolution, local RGFM still shows inconsistencies between spatially separated facial components, suggesting weakened local predictability of global attributes.
Abstract
from arXiv · showhide
Despite their remarkable success in modeling complex data, generative models face a fundamental tradeoff. Global approaches can capture full structural coherence but suffer from high computational costs, while local models are efficient but often fail to reproduce long-range correlations and global coherence. The renormalization group (RG) bridges this gap by seamlessly connecting spatial structures across different length scales, retaining quasi-local descriptions at each step while preserving long-range correlations. We introduce renormalization group flow matching (RGFM), a generative framework that systematically structures data generation across different spatial scales. By using an exact RG flow as the probability path, RGFM progressively generates data from long- to short-wavelength structures. To reconcile scalability with global structure, we exploit two key properties of the RG: quasi-locality and scale separation. We rigorously show that the RGFM probability flow can be accurately approximated by local velocity fields acting over a spatial range $O(Λ^{-1}[\ln L+\ln(1/\varepsilon)])$ for RG wavenumber scale $Λ$, linear system size $L$, and prescribed error tolerance $\varepsilon$. This property enables local generative modeling with patches of size $O(\ln L)$ and a computational cost that scales nearly linearly with the system volume. We numerically demonstrate that local RGFM reproduces long-range correlations far beyond its receptive field in representative one-dimensional distributions, while conventional local flow matching exhibits substantial errors at long distances. On FFHQ images, RGFM yields far more coherent and higher-quality samples than local flow matching at 64x64 and 256x256. Our results establish RG-guided probability flows as a promising route toward scalable generative modeling that captures long-range structure using only local computation.
I. INTRODUCTION … A. Flow matching
The paper introduces RG flow matching (RGFM) as a coarse-to-fine generative framework that uses exact RG flows to combine local computation with long-range structure. It develops locality guarantees and experiments showing that local RGFM can preserve correlations and coherence beyond a local receptive field.
- A. Flow matching: Flow matching uses a deterministic velocity field and a flexible probability path connecting data to a simple reference distribution.Standard FM generates data by reversing an ODE flow learned between pdata and q, such as N(0, I).
- A. Background: 150–1000 V100 GPU days are required by some large-scale image-generation models, motivating methods that reduce computational cost with data dimension.Repeated neural-network evaluations during training and sampling drive the scaling challenge.
- A. Background: Local generative modeling evaluates fields from surrounding patches, yielding linear cost in data dimension for fixed patches and near-linear scaling when patch size grows logarithmically.The approach exploits spatial structure but does not automatically coordinate distant regions.
- A. Background: Fixed-neighborhood local models can miss long-range correlations and semantic features, with errors concentrated in an intermediate transition-like regime.Local approximations may remain effective near the data and Gaussian endpoints.
- A. Background: RG quasi-locality keeps each scale’s interactions concentrated over distances O(Λ^-1), while iterated scale-separated stages communicate structures across progressively longer distances.This property supports local descriptions organized by interactions and spatial derivatives.
- B. Summary of the main results: RGFM takes an exact RG flow as its probability path, progressively converting high-wavenumber fluctuations into Gaussian noise and generating data from coarse to fine scales.The terminal distribution approaches Gaussian below the infrared cutoff ΛIR ∝ 1/L.
- B. Summary of the main results: Successive lattice rescalings discard decoupled Gaussian high-wavenumber modes and keep the remaining low-wavenumber modes on reduced lattices, preventing receptive ranges from reaching system size.This construction is designed to enable scalable local training and sampling.
- B. Summary of the main results: The locality theorem bounds local RGFM approximation using buffer widths, with exponentially decaying velocity-field error and a stability bound on the resulting 2-Wasserstein distance.Experiments find that local RGFM reproduces distant Ising correlations and reconstructs latent global waveforms, unlike local FM under the same constraints.
B. Renormalization group flow matching
RGFM defines flow matching along an RG diffusion path that progressively transforms high-wavenumber modes into Gaussian noise while preserving low-wavenumber correlations. Its velocity field combines a local differential term with a generally nonlocal score-dependent term, and its endpoint is Gaussian.
- RG diffusion: High-wavenumber modes are successively transformed into Gaussian noise, proceeding from high to low wavenumbers across the scale Λ(t).At each time, the low-momentum modes retain a nontrivial effective theory while integrated-out high-momentum modes follow a Gaussian theory.
- Scale separation: The low-momentum effective distribution exactly preserves correlations in the initial data distribution, while the high-momentum modes become Gaussian under scale separation.For the smooth exponential cutoff, the stated factorization is a controlled approximation at finite Λ(t).
- Definition: RGFM is defined as flow matching associated with the RG diffusion, learning its velocity field from the corresponding conditional velocity.The model uses the standard flow-matching objective with the RG diffusion’s conditional velocity field.
- Probability path: At t = 1, the RGFM probability path converges to the Gaussian distribution pGS(ϕ), enabling generation from Gaussian samples toward data samples.The hyperparameter τ is chosen so that the endpoint is sufficiently close to pGS(ϕ).
- Velocity field: The real-space RGFM velocity contains a local term proportional to (−∇2 + m2)ϕx and a generally nonlocal term determined by the full configuration through the score.The local and nonlocal contributions arise from the explicit real-space form of the learned velocity field.
III. LOCAL GENERATIVE MODELING WITH RGFM … IV. LOCAL APPROXIMABILITY OF THE PROBABILITY FLOW IN THE RGFM
RGFM enables local generative modeling by combining RG scale separation with quasi-local probability flows, rescaling the lattice so constant-size patches can track generation across scales. Under locality and regularity assumptions, the RGFM flow admits accurate local approximations with receptive fields whose size grows only logarithmically with system size.
- III. LOCAL GENERATIVE MODELING WITH RGFM: RGFM exploits RG scale separation and locality, treating integrated-out high-wavenumber modes as independent Gaussian fluctuations while preserving a quasi-local probability flow.These properties motivate local generation along the RG-based probability path.
- A. Rescaled RGFM flow: Site decimation retains low-wavenumber modes via a DCT, reconstructs a reduced lattice, and removes short-distance degrees of freedom while preserving the effective large-scale field.For Λ(t) = ΛUV/b, the lattice size is reduced from Ld to (L/b)d.
- A. Rescaled RGFM flow: O(ln L) rescaling operations keep the score-locality length scale O(L0) in stage-specific coarse-grained lattice units, enabling constant-size patches during the flow.Without rescaling, the locality length reaches O(L) at t = 1 and would prevent efficient scaling.
- A. Rescaled RGFM flow: The probability-distribution transformation remains exactly invertible probabilistically: missing high-wavenumber modes are independently resampled from the Gaussian sector before reconstruction.This reversibility follows from RG scale separation even though the field transformation is many-to-one.
- B. Training objective and sampling procedure: Local RGFM trains one shared velocity network whose prediction on each target region depends only on that region’s patch, buffer, time, and position.The local patch comprises Ait ∪ Bit, with target and buffered-region volumes ldA and (lA + 2lB)d, respectively.
- B. Training objective and sampling procedure: Sampling alternates reverse-ODE evolution under the learned local velocity field with stochastic inversion of site-decimation transformations at successive scale transitions.Initialization samples from the Gaussian sector, while each transition lifts the configuration to a finer lattice by retaining low-momentum modes and sampling missing modes.
- IV. LOCAL APPROXIMABILITY OF THE PROBABILITY FLOW IN THE RGFM: Under physically reasonable data-locality, RG-evolution, and local-velocity regularity assumptions, the paper establishes mathematical local approximability of the RGFM probability flow.The local velocity is defined through a patchwise flow-matching objective and analyzed using locality and Lipschitz assumptions.
A. Assumptions on the data distribution and its RG flow · 1. Local and conditionally local data distributions
The paper assumes data are local or conditionally local, with locality preserved along the RG flow and latent information inferable from sufficiently large local neighborhoods. For conditionally local data, long-range dependencies are mediated by latent variables, while fixed-latent conditional actions remain uniformly local.
- A. Assumptions on the data distribution and its RG flow: The RG-flow assumptions require locality properties of the data distribution to remain valid along the flow.The framework separately introduces a Lipschitz-continuity assumption for local velocity-field approximations.
- 1. Local and conditionally local data distributions: Natural data motivate local-action models because their power spectra often follow approximate power-law behavior, E[|ϕ_k|^2] ∼ |k|^-α with α ∼ 2.This behavior is described as reminiscent of systems governed by local actions.
- 1. Local and conditionally local data distributions: A toy model shows that marginalizing a shared latent z ∈ {−1, +1} produces global structure, whereas fixed-z conditional distributions contain independent local Gaussian fluctuations.The example uses p(z) = Uniform{±1} and p(ϕ|z) = ∏_x N(ϕ_x; z, σ^2).
- 1. Local and conditionally local data distributions: The framework therefore considers either local-action distributions or distributions whose conditional action is local for every fixed latent variable z.In the conditional case, pdata(ϕ) is represented through p(z)pdata(ϕ|z), with locality parameters chosen independently of z and conditional locality scale O(L0).
- 1. Local and conditionally local data distributions: Locality is defined by exponential decay of the sensitivity of the interaction force at x to field changes at y with separation |x − y|.The decay must hold uniformly while fluctuations in an exterior region are continuously removed through the truncation interpolation.
- 1. Local and conditionally local data distributions: Conditional locality alone is insufficient for locally tracking the RGFM of the marginal distribution, so latent information must be concentrated in a local configuration.The paper formulates this requirement using conditional mutual information between z and an exterior region conditioned on a local region.
- 1. Local and conditionally local data distributions: The latent-predictability assumption requires the exterior region to provide exponentially less additional information about z as the buffer width l_B increases.Thus, the local configuration ϕ_AB becomes approximately sufficient for latent inference.
2. Properties of the RG flow for local distributions · B. Local approximation of velocity fields
The paper assumes that RG-evolved interactions remain quasi-local at scale O(Λ^-1), with latent variables remaining locally predictable for conditionally local distributions. It defines local velocity approximations through patchwise conditional fields and shows that these approximations are optimal for local-FM training under an additional regularity assumption.
- 2. Properties of the RG flow for local distributions: RG-evolved interactions are assumed local on a length scale O(Λ^-1), with constants independent of system size L and RG scale Λ.The locality assumption applies to the rescaled interaction V_Λ(ϕ) = U_Λ(√K_Λϕ).
- 2. Properties of the RG flow for local distributions: For conditionally local distributions, the locality bound for the rescaled interaction is assumed uniformly over the latent variable z.The constants γ and c can be chosen independently of z.
- 2. Properties of the RG flow for local distributions: The RG flow additionally assumes that the latent variable remains locally predictable at scale Λ^-1, measured by conditional mutual information between z and an exterior region.This requirement does not require the marginalized distribution p_Λ(ϕ) to retain a finite Markov length during the flow.
- B. Local approximation of velocity fields: A local approximation v_A,loc depends only on the field configuration in the local neighborhood ϕ_AB, and the full-domain approximation v_loc is assembled by patching local regions with buffers.The construction partitions Ω_L into equal-sized regions A_i, each surrounded by a buffer B_i of width l_B.
- B. Local approximation of velocity fields: The local approximation v_A,loc minimizes the L2(p) distance from the restricted velocity field among velocity fields depending only on local configurations.The full-domain error is defined from the local approximation errors on the patches.
- B. Local approximation of velocity fields: For standard FM and RGFM, the full-domain local approximation coincides with the minimizer of a local-FM patchwise training objective.The RGFM local velocity field is the local approximation of the rescaled velocity under the rescaled distribution.
- B. Local approximation of velocity fields: The framework assumes that the local approximation of the RGFM velocity field is Lipschitz continuous in the field configuration with a constant M.The assumption is imposed for local or conditionally local data distributions satisfying Assumption IV.1.
- B. Local approximation of velocity fields: Uniform Lipschitz regularity is expected when the interaction V_Λ avoids singular or strongly irregular dependence on the field configuration along the RG flow.The stated expectation does not follow directly from the uniform L2(p_Λ) norm bound alone.
C. Local approximability of velocity fields in the RGFM · 1. Local approximability of velocity fields in the RGFM for local data distributions
For local data distributions, RGFM velocity-field approximation errors decay exponentially with buffer width, at a length scale set by the inverse running RG scale. The analysis establishes local-region and full-domain bounds using uniform moment and velocity-field estimates.
- 1. Local approximability of velocity fields in the RGFM for local data distributions: For local data distributions, the RGFM local approximation error decays exponentially with buffer width lB, with characteristic length Λ(t)^−1.The result applies to pt = pΛ and vt = vΛ at RG wavenumber scale Λ(t).
- 1. Local approximability of velocity fields in the RGFM for local data distributions: Theorem IV.3 provides constants γ and c independent of L and t for bounding the local approximation error.The theorem is stated for local data distributions satisfying Assumption IV.1.
- 1. Local approximability of velocity fields in the RGFM for local data distributions: The RG moment ∥ϕx −⟨ϕx⟩pΛ∥L4(pΛ) is uniformly bounded by a constant M independent of Λ.This moment estimate supports the subsequent local approximation bounds.
- 1. Local approximability of velocity fields in the RGFM for local data distributions: On local regions, the approximation error is bounded using constants γ and c and a (d −1)-th degree polynomial Poly(x) independent of L and Λ.The bound applies to any local regions A and B.
- 1. Local approximability of velocity fields in the RGFM for local data distributions: The full-domain approximation bound follows by summing local-region bounds over disjoint regions partitioning ΩL.The constants γ and c and the polynomial Poly(x) remain independent of L and Λ.
- 1. Local approximability of velocity fields in the RGFM for local data distributions: The RGFM velocity field for a local data distribution has an L2(pΛ)-norm bounded by a constant γ independent of L and Λ.This estimate is combined with the full-domain approximation bound to prove Theorem IV.3.
- 1. Local approximability of velocity fields in the RGFM for local data distributions: The constants in the exponential error bound depend only on RG locality parameters and the uniform L4(pdata) moment bound.This dependence enables extension to conditionally local data distributions in the next section.
2. Local approximability of velocity fields in the RGFM for conditionally local data distributions
For conditionally local data distributions, RGFM velocity fields remain locally approximable despite latent-variable-induced nonlocal correlations, with errors controlled by exponential decay and latent-estimation costs. The unified result gives logarithmic buffer widths in system size and inverse tolerance, with conditionally local scaling α = 2d/3.
- Theorem IV.10: RGFM local approximation errors for conditionally local data distributions decay exponentially with buffer width despite latent-variable-induced nonlocal correlations.Theorem IV.10 establishes constants independent of L and t controlling this approximation error.
- Local-region bound: The local-region error combines conditional velocity-field approximation with an additional error from estimating the latent-conditioned expectation.The conditional term is controlled by RG locality, while the latent-estimation term is bounded separately for any q > 2.
- Full-domain bound: For q > 2, Corollary IV.14 bounds the full-domain local approximation error using constants independent of L and Λ.The bound follows by summing local-region inequalities over a disjoint decomposition of the full domain.
- Theorem IV.15: The unified theorem applies to local and conditionally local data distributions, with α = d/2 for local and α = 2d/3 for conditionally local distributions.The conditionally local L2d/3 scaling is a rough general bound and can be replaced by Ld/2 for translation invariant systems.
- Theorem IV.15: For any ε > 0, choosing lB,t = O(Λ(t)^−1 ln(L^α/ε)) ensures the local approximation error is at most ε.Thus, the required buffer width grows logarithmically with system size and inverse error tolerance.
D. Local approximability of probability flows in the RGFM
RGFM’s generative ODE flow remains locally approximable: suitable buffer widths make the locally generated distribution close to the target in 2-Wasserstein distance. An equivalent SDE formulation extends this result to KL divergence, with local score and velocity approximations, while avoiding the ODE’s Lipschitz requirement at continuous time.
- ODE flow: Theorem IV.16 shows that local velocity fields can approximate RGFM’s generative ODE flow with arbitrarily prescribed accuracy in 2-Wasserstein distance.The result applies to local and conditionally local data distributions under the stated assumptions.
- ODE flow: The required buffer width scales as O(Λ(t)^−1 ln(Lα/ε)), with α = d/2 for local and α = 2d/3 for conditionally local distributions.This scaling holds for any prescribed accuracy ε > 0.
- ODE flow: Theorem IV.17 bounds the 2-Wasserstein discrepancy between exact and locally approximated probability flows using the local velocity approximation error, assuming vloc,t is Lipschitz continuous.This stability result underpins the generative-flow guarantee in Theorem IV.16.
- SDE flow: The equivalent SDE can be locally approximated by replacing both the velocity field and score with vloc,t and sloc,t, yielding a KL-divergence guarantee.In RGFM, the score differs from the velocity only by a local contribution, linking their approximation errors.
- SDE flow: Theorem IV.20 obtains the same buffer-width scaling O(Λ(t)^−1 ln(Lα/ε)) for KL divergence between the target data distribution and the locally approximated generative SDE flow.Here α = d/2 for local data and α = 2d/3 for conditionally local data.
- SDE flow: Unlike the ODE result, the SDE bound does not require Lipschitz continuity to control trajectory-error propagation, though discretization can introduce practical difficulties.The SDE guarantee concerns a probability flow different from the locally approximated ODE flow, so it does not transfer directly between them.
V. NUMERICAL EXPERIMENTS
The numerical experiments compare RGFM with conventional flow matching under identical local-network constraints, first testing long-range correlations in one-dimensional distributions and then evaluating natural-image generation at 64 × 64 and 256 × 256 resolutions.
- Experimental setup: Experiments compare local RGFM and conventional FM under the same local-network constraints.The study evaluates whether each method can model data using fixed-size receptive fields.
- One-dimensional distributions: One-dimensional experiments test whether fixed-size receptive fields reproduce long-range correlations in local and conditionally local distributions.The analysis includes representative configurations and correlation functions for the one-dimensional Ising model.
- Natural-image generation: Natural-image experiments evaluate local generative modeling at resolutions of 64 × 64 and 256 × 256.These experiments are used to demonstrate the computational scalability of the approach.
A. One-dimensional local distribution: Ising model … VI. DISCUSSION
Across one-dimensional distributions and FFHQ images, local RGFM reproduces long-range correlations and global structure with local receptive fields, outperforming conventional local flow matching. The discussion attributes this advantage to RG scale separation and quasi-locality, while identifying limitations for high-resolution data and possible latent-variable extensions.
- A. One-dimensional local distribution: Ising model: The Ising demonstration embeds discrete configurations in R^L and approximates the RGFM path with a local velocity field because the action contains only nearest-neighbor interactions.The local CNN uses Nlayer = 6 layers with kernel size H = 5, giving a receptive field of 25 sites.
- A. One-dimensional local distribution: Ising model: Local-CNN RGFM reproduces Ising configurations with correlations extending far beyond its fixed receptive field, whereas local-CNN FM fails to reproduce the correct long-range structure.The comparison uses L = 1024 and correlation lengths lcorr = 32, 64, 128, and 256.
- B. One-dimensional conditionally local distribution: The conditionally local marginal has nondecaying oscillatory correlations even though the field variables are independent Gaussian variables when conditioned on the latent variable z.The experiments use L=1024 and σ=0.05 with the same local CNN architecture as the Ising experiment.
- B. One-dimensional conditionally local distribution: For the conditionally local distribution, local-CNN RGFM captures both the smooth global waveform and local Gaussian fluctuations using a fixed-size receptive field.It also reproduces the marginal distribution’s oscillatory long-range correlation function, unlike local-CNN FM.
- C. Image generation: On 64 × 64 FFHQ images, local-patch RGFM captures global facial structure under a local-neighborhood constraint, while local-patch FM struggles to construct globally coherent faces.The models use target patches of linear size 16 and receptive patches of linear size 32.
- C. Image generation: Across receptive patch sizes, local RGFM maintains a substantially lower FID, whereas increasing the receptive patch size does not substantially improve local FM.The target-patch size is fixed at lA = 16.
- C. Image generation: At 256 × 256 resolution, local-patch RGFM generates substantially more coherent images than local-patch FM, although inconsistencies remain between spatially separated facial components.Global self-attention at the first U-Net resolution requires more than 10 GB of GPU memory per image in a training batch in the stated implementation.
- VI. DISCUSSION: RGFM keeps local approximations tractable through successive site decimation, with patch sizes growing as O(ln L), and maps long-range structures onto smaller effective systems.The discussion proposes conditioning local dynamics on an inferred latent variable to address missing global information and notes possible extensions beyond exponential locality assumptions.
Appendix A: Mathematical lemmas used in the main text … 2. Proof of Proposition IV.6 in the main text
The appendices establish the mathematical tools underlying the main text and prove propositions controlling RG-field fluctuations and force truncation. The proofs combine RG locality, conditional inequalities, Minkowski’s inequality, Hölder’s inequality, and shell integration.
- Appendix A: Mathematical lemmas used in the main text: The appendix collects mathematical lemmas and proofs that support the main-text analysis.It introduces the Grönwall inequality and Minkowski integral inequality as foundational tools.
- Appendix A: Mathematical lemmas used in the main text: Grönwall’s inequality is stated and proved using an integrating factor, integration, and the initial condition a(0) = 0.The proof concludes by recovering the desired inequality.
- Appendix A: Mathematical lemmas used in the main text: Minkowski’s integral inequality is established for nonnegative functions and 1≤r<∞ by exchanging integrations and applying Hölder’s inequality with conjugate exponent r′.The r = 1 case follows directly from changing the integration order, while 1 < r < ∞ uses the conjugate-exponent argument.
- 1. Proof of Proposition IV.5 in the main text: The proof of Proposition IV.5 bounds the centered field fluctuation ∥ϕx −⟨ϕx⟩pΛ∥L4(pΛ) using the triangle inequality and conditional Jensen’s inequality.The RG diffusion representation and locality assumptions yield a bound uniform in x and Λ.
- Appendix B: Proofs in the main text: Proposition IV.6 reduces velocity-field control to bounding the functional derivative uΛ(ϕ), because vΛ(ϕ) is proportional to uΛ(ϕ) with coefficient 1/Λ2τ.The proof extends configurations outside A∪B by conditional means and constructs a local field on A.
- 2. Proof of Proposition IV.6 in the main text: The Proposition IV.6 force difference is bounded by applying Minkowski’s integral inequality, Hölder’s inequality, RG locality, and Proposition IV.5’s uniform L4 bound.These steps control the discrepancy between original and truncated forces over regions A, B, and C.
- 2. Proof of Proposition IV.6 in the main text: The remaining spatial integral is estimated by shells centered at x∈A, with polynomial factors and constants independent of L and Λ absorbed into γ′.Integrating the resulting bound over x∈A completes the proof.
3. Proof of Proposition IV.12 in the main text · 4. Proof of Proposition IV.13 in the main text
The proofs establish the local approximation bound for Proposition IV.12 using locality, velocity equality, norm inequalities, and conditional Jensen’s inequality. Proposition IV.13 is proved by partitioning configuration space and combining conditional expectations, Pinsker’s inequality, Jensen/Minkowski inequalities, and the RG conditional-mutual-information assumption.
- 3. Proof of Proposition IV.12 in the main text: Λ,loc(ϕAB|z) is local on region A because it depends only on ϕAB.This locality property is used in the proof of Eq. (92).
- 3. Proof of Proposition IV.12 in the main text: The proof of Eq. (92) applies velocity equality (85), the inequality ||a + b||2, and conditional Jensen’s inequality.These steps bound the local approximation error ΔA(pΛ, vΛ; lB).
- 3. Proof of Proposition IV.12 in the main text: The argument concludes by identifying the resulting inequality with Eq. (92).The cited passage explicitly states that this proves Eq. (92).
- 4. Proof of Proposition IV.13 in the main text: For Proposition IV.13, the configuration space C = {(ϕ, z)} is divided using an arbitrary positive constant R.This partition initiates the subsequent bound.
- 4. Proof of Proposition IV.13 in the main text: Because Ez|ϕAB[hΛ(ϕ, z)] = 0, the left-hand side of Eq. (93) is bounded using ||a + b||2 and Pinsker’s inequality.Pinsker’s inequality controls the total variation term in the bound.
- 4. Proof of Proposition IV.13 in the main text: Conditional Jensen’s inequality is applied for q > 2 and to hΛ, which depends only on ϕAB and z.The proof then uses ||a + b||2 to combine the resulting estimates.
- 4. Proof of Proposition IV.13 in the main text: Minkowski’s integral inequality combines the intermediate bounds, while the local velocity moment is controlled by Eϕ|z[|vΛ,loc,x(ϕAB|z)|q] ≤ Eϕ|z[|vΛ,x(ϕ|z)|q].The latter follows from the definition of the local approximation and conditional Jensen’s inequality.
- 4. Proof of Proposition IV.13 in the main text: Using IΛ(Z :C|AB) ≤ γe^−cΛlB, the proof derives the bound and proves inequality (93), absorbing constants into γ′ and defining c′ = ((q −2)/2q)c.The RG conditional-mutual-information assumption supplies the final decay term.
5. Proof of Theorem IV.10 in the main text
The proof reduces Theorem IV.10 to bounding an integral, then combines finite-lattice norm inequalities, conditional Jensen bounds, and kernel estimates. Using τ^-1 = O(ln L), the resulting bound together with Corollary IV.14 establishes the theorem, including q = 3.
- 5. Proof of Theorem IV.10 in the main text: Corollary IV.14 reduces the proof to establishing a bound on a specific integral.The proof explicitly begins by invoking Corollary IV.14 to identify the sufficient integral estimate.
- 5. Proof of Theorem IV.10 in the main text: Finite-lattice bounds are derived using an inequality, Hölder’s inequality, and a finite-lattice norm inequality.These steps provide the intermediate estimates needed to control the integral.
- 5. Proof of Theorem IV.10 in the main text: Conditional Jensen’s inequality twice bounds the conditionally local velocity field after expressing v_Λ,k(ϕ|z) using Lemma IV.8.The derivation uses the conditional-local representation of the velocity field.
- 5. Proof of Theorem IV.10 in the main text: Bounded functions f1(x) and f2(x), together with kernel identities for K_tk and ∂tK_tk, control the resulting terms.The bounds |f1(x)| ≤ a* and |f2(x)| ≤ b* are used explicitly.
- 5. Proof of Theorem IV.10 in the main text: For 2 < q ≤ 4, norm inequalities and Assumption IV.1 bound both terms, and τ^-1 = O(ln L) yields the final estimate.The argument uses ∥ϕ0∥_Ω^L ≤ L^d/4 ∥ϕ0∥_4 and bounds Eϕ0 by a constant M.
- 5. Proof of Theorem IV.10 in the main text: Combining the bound with Corollary IV.14 gives constants independent of L and Λ, and q = 3 recovers the bound stated in Theorem IV.10.The constants c1, c2, γ1, and γ2 are independent of L and Λ.
6. Proof of Theorem IV.17 in the main text
The proof couples the exact and local ODE flows from the same initial configuration and bounds their induced distributional discrepancy using Wasserstein distance. Flow-difference estimates, Lipschitz continuity, Grönwall’s inequality, and the local approximation error establish Theorem IV.17.
- Coupling the flows: The exact and local ODE solutions share an initial configuration, producing samples from p_t and p_t^loc, respectively.This shared initialization enables a direct map between the two evolved configurations.
- Coupling the flows: The backward exact flow followed by the forward local flow defines a coupling between p_t and p_t^loc.The coupling maps each exact-flow sample to the corresponding local-flow sample.
- Bounding flow differences: The 2-Wasserstein distance is bounded using the norm of the configuration difference δ_t = ϕ_t − ˆϕ_t.The difference evolves according to the discrepancy between the exact and local velocity fields.
- Bounding flow differences: The triangle inequality and Lipschitz continuity of v_loc,t with constant M(t) control the difference evolution, after which Grönwall’s inequality yields a time-integrated bound.The argument applies the stated norm inequality and the Lipschitz assumption before invoking Grönwall’s inequality.
- Concluding the theorem: Taking expectations and applying Minkowski’s inequality converts the trajectory estimate into a bound involving the local approximation error Δ^L_Ω(p_s, v_s; l_B,s).The proof uses ϕ_s ∼ p_s and the definition of the local approximation error to identify the final integrand.
- Concluding the theorem: Combining the derived bounds gives the claimed 2-Wasserstein estimate between p_t and p_t^loc, completing the proof of Theorem IV.17.The final step explicitly combines the preceding equations to obtain the theorem’s desired bound.
7. Proof of Proposition IV.18 in the main text … b. RGFM scale decomposition
The appendices establish equivalence between the RGFM SDE and ODE marginals, bound local approximation errors through KL divergence, and detail multiscale numerical implementation. RGFM uses scale-specific local networks, Gaussian-sector transitions, and coefficient lifting to generate across lattice resolutions.
- 7. Proof of Proposition IV.18 in the main text: The SDE (107) generates the same marginal-distribution evolution as the original ODE flow because its Fokker–Planck equation coincides with the ODE continuity equation.The proof uses pt(δ ln pt/δϕ) = δpt/δϕ.
- 8. Proof of Theorem IV.19 in the main text: Theorem IV.19 derives a time-integrated KL-divergence bound by combining continuity and Fokker–Planck equations, vanishing boundary terms, and Young’s inequality.The bound uses the initial condition DKL(p0||ˆploc 0 ) = 0.
- 9. Proof of Theorem IV.20 in the main text: O(Λ(t)−1 ln(Lα/√ε)) local patches suffice to make the RGFM local approximation error arbitrarily small for prescribed accuracy ε > 0.This conclusion follows from Theorem IV.15 and proves Theorem IV.20.
- Appendix C: Details of the numerical experiments: The numerical experiments use orthonormal DCTs and midpoint integration of learned ODE flows with torchdiffeq in PyTorch.The experiments set a=1.
- 1. One-dimensional experiments: The one-dimensional experiments use a shared local CNN architecture for RGFM and standard FM, with 6 convolutional layers, kernel size H = 5, and hidden dimension 64.The architecture includes residual blocks, group normalization, SiLU activations, and a 1×1 output convolution; optimization uses Adam with learning rate 2 × 10−4 and EMA decay 0.995.
- b. RGFM scale decomposition: RGFM begins from an original lattice of size L0 = 1024, successively decimates it, and chooses transitions after discarded modes enter the Gaussian sector.A separate copy of the same local CNN is trained on each interval between successive scale transitions.
- b. RGFM scale decomposition: The one-dimensional RGFM setup uses m = 0.02, trains for 4.5 × 105 total optimization steps, and takes approximately three hours on one NVIDIA RTX 6000 Ada GPU.The first and last scale intervals use 100,000 steps each, while five intermediate intervals use 50,000 steps each.
- b. RGFM scale decomposition: At each reverse-process transition, retained low-wavenumber DCT coefficients are lifted to a finer lattice while missing high-wavenumber coefficients are independently sampled from the Gaussian sector.The complete probability path uses 200 midpoint integration steps.
c. Dataset-specific details and evaluation … g. FID evaluation
The paper evaluates RGFM on synthetic distributions and FFHQ images using exact correlation-function references and FID, while specifying local patching, U-Net, optimization, and RG-schedule details. Image experiments use coordinate-aware local models across multiple resolutions and receptive-field configurations.
- c. Dataset-specific details and evaluation: 10,000 samples are used to compare Ising-model correlations with the exact expression in Eq. (114).
- c. Dataset-specific details and evaluation: 10,000 conditionally local samples are compared with the exact correlation function in Eq. (118), using σ = 0.05 Gaussian fluctuations.
- 2. Image-generation experiments: FFHQ images are normalized to [−1, 1] after resizing to 64 × 64, while the 256 × 256 experiment uses the first 5,000 images.
- b. Local patches and positional channels: Each scale divides images into nonoverlapping target patches of size lA, predicting their velocities from receptive patches of size lA + 2lB with boundary-aware shifts.
- b. Local patches and positional channels: Because images are not translationally invariant, the U-Net receives five channels: RGB values and normalized absolute horizontal and vertical coordinates.The network predicts three RGB velocity channels, and the loss is evaluated only within target region A.
- c. U-Net architecture and optimization: RGFM and FM share a U-Net with base width 128, multipliers, two residual blocks per resolution, and dropout 0.1.Self-attention is used at the second resolution level and middle block; input and output dimensions are 5 and 3.
- d. RG parameters: The 64 × 64 RGFM experiment uses target patches of size 16 and receptive patches of sizes 32, 32, and 16 across three successive lattice scales.The three U-Nets are trained for approximately 6 × 10^5, 2.5 × 10^5, and 1.5 × 10^5 steps, respectively.
- g. FID evaluation: 64×64 image models are evaluated with FID using 50,000 FFHQ images and clean-FID preprocessing before Inception feature statistics are computed.Images are mapped from [−1, 1] to 8-bit RGB, resized to 299 × 299, and used to form empirical feature means and covariances.