Source-linked AI summary
Learning Non-Convergent Non-Persistent Short-Run MCMC Toward Energy-Based Model
Erik Nijkamp, Mitch Hill, Song-Chun Zhu, Ying Nian Wu
TL;DR
The paper asks whether EBM learning can work with short-run MCMC that neither converges nor mixes, a setting that makes the sampler invalid for the EBM. It treats the learned dynamics as a generator or flow model and reports realistic synthesis, smooth interpolation, and faithful reconstruction, while noting limitations of the resulting estimators and maximum-likelihood assumptions.
Problem
MCMC mixing can be impractical for multimodal EBMs, while short-run MCMC is not a valid sampler of the EBM it trains.
Method
The paper repeatedly initializes fixed-length Langevin dynamics from noise and updates the EBM, interpreting the resulting dynamics as a generator or flow model.
Results
The learned short-run MCMC generates realistic images, smoothly interpolates between faces, and faithfully reconstructs observed examples.
Takeaways & Limitations
Non-convergent, non-mixing, non-persistent short-run dynamics can provide generator- or flow-like capabilities beyond traditional EBM and MCMC.
Takeaways & Limitations
The moment-matching estimator may fail to match data moments and have much lower entropy than the maximum-likelihood EBM, while maximum-likelihood optimality may not hold for misspecified models.
Abstract
from arXiv · showhide
This paper studies a curious phenomenon in learning energy-based model (EBM) using MCMC. In each learning iteration, we generate synthesized examples by running a non-convergent, non-mixing, and non-persistent short-run MCMC toward the current model, always starting from the same initial distribution such as uniform noise distribution, and always running a fixed number of MCMC steps. After generating synthesized examples, we then update the model parameters according to the maximum likelihood learning gradient, as if the synthesized examples are fair samples from the current model. We treat this non-convergent short-run MCMC as a learned generator model or a flow model. We provide arguments for treating the learned non-convergent short-run MCMC as a valid model. We show that the learned short-run MCMC is capable of generating realistic images. More interestingly, unlike traditional EBM or MCMC, the learned short-run MCMC is capable of reconstructing observed images and interpolating between images, like generator or flow models. The code can be found in the Appendix.
1 Introduction
The paper examines whether fixed-length, non-convergent short-run MCMC initialized from noise can support EBM learning despite impractical MCMC mixing, and argues that it can function as a generator or flow model.
- 1.1 Learning Energy-Based Model by MCMC Sampling: MCMC for multimodal EBMs can fail to mix, trapping chains from different starting points in separate local modes.This makes convergence impractical for complex distributions such as natural images.
- 1.1 Learning Energy-Based Model by MCMC Sampling: The proposed scheme repeatedly runs 5 to 100 Langevin steps from the same noise distribution before updating the EBM parameters.The short-run chains are non-convergent, non-mixing, and non-persistent.
- 1.2 Short-Run MCMC as Generator or Flow Model: Although short-run MCMC is not a valid EBM sampler, the learned procedure can generate realistic images.The paper frames this as learning a sampler of a learned model and investigates why the resulting images are realistic.
- 1.2 Short-Run MCMC as Generator or Flow Model: The paper argues that learned short-run MCMC matches data statistics and can act as a generator or flow model with Langevin dynamics as a noise-injected residual network.The initial image serves as a latent variable and uniform noise serves as the latent prior.
- 1.2 Short-Run MCMC as Generator or Flow Model: The learned short-run MCMC supports synthesis, image inpainting, super-resolution, style transfer, and inverse optimal control through informative initial distributions and conditional energies.These extensions are proposed as generalizations of the synthesis learning scheme.
2 Contributions and Related Work
The paper shifts attention from convergent MCMC for EBMs to learned non-convergent short-run dynamics, presenting evidence for generator-like synthesis, interpolation, and reconstruction.
- 2 Contributions and Related Work: The paper proposes treating non-convergent short-run MCMC as a valid generator or flow model rather than focusing exclusively on learning an EBM.It describes this as a conceptual shift toward energy-based dynamics.
- 2 Contributions and Related Work: Short-run MCMC interpolates smoothly between generated CelebA faces, with intermediate samples resembling realistic faces.The procedure uses interpolated latent noise and 100 MCMC steps.
- 2 Contributions and Related Work: The paper leaves understanding attractor dynamics in neuroscience to future work.This is stated as a remaining direction rather than a completed result.
- 2 Contributions and Related Work: The approach emphasizes non-persistent MCMC as more efficient and convenient than persistent MCMC while differing from contrastive divergence through noise initialization.Its theoretical analysis is related to generalized moment matching and moment matching GANs, but does not use adversarial generator learning.
3 Non-Convergent Short-Run MCMC as Generator Model
The paper treats fixed-length MCMC initialized from p0 as a learned generator qθ rather than as a sampler intended to converge to pθ, using this dynamics for generation, reconstruction, and interpolation.
- Short-run dynamics: The proposed method runs K MCMC steps from a fixed initial distribution p0 and defines qθ as the resulting marginal distribution.This abandons direct sampling of pθ while retaining pθ as the guide for the short-run dynamics.
- Learning qθ: The learning procedure updates θ using generated examples treated as independent and fair samples from qθ.The algorithm alternates between generating negative examples with short-run MCMC and updating parameters with the learning gradient.
- Training stabilization: Gaussian noise injection smooths pdata, helping the estimating equation converge when K is small and a solution may otherwise not exist.Observed examples are perturbed with εi ∼ N(0,σ^2I) before the parameter update.
- Generator interpretation: The K-step dynamics can be viewed as a K-layer noise-injected residual network with latent input z and randomness u.Because the dynamics are non-convergent and non-mixing, the output x can remain highly dependent on z, allowing z to be inferred from x.
- Interpolation and reconstruction: The learned dynamics support image interpolation by interpolating latent variables and reconstruction by optimizing z to minimize the image reconstruction loss.The paper also notes that setting u = 0 makes the dynamics deterministic and connects the formulation to attractor dynamics.
4 Understanding the Learned Short-Run MCMC
The learned short-run MCMC can be valid without sampling the converged EBM: it matches data moments while functioning as a generator-like model. Its behavior is explained geometrically through moment matching, entropy, and finite-step divergence relationships.
- 4.1 Exponential Family and Moment Matching Estimator: In the exponential-family case, short-run MCMC learning converges to a moment-matching estimator whose sufficient statistics match the data distribution.For fθ(x)=⟨θ,h(x)⟩, both the maximum-likelihood and short-run estimators match E[h(x)] to the data.
- 4.1 Exponential Family and Moment Matching Estimator: The learned q ˆθMME belongs to the moment-matching family Ω, whereas the converged EBM p ˆθMME may lie outside it and have lower entropy.The converged EBM can be farther from the uniform initialization than the maximum-likelihood projection, reducing its entropy.
- 4.1 Exponential Family and Moment Matching Estimator: Finite-step MCMC starts from p0 and stops at q ˆθMME, while the unreachable infinite-step target is p ˆθMME.The EBM guides the transition, but the learned short-run MCMC itself is retained as the valid model.
- 4.1 Exponential Family and Moment Matching Estimator: The maximum-likelihood distribution p ˆθMLE is the intersection of the EBM family Θ and moment-matching family Ω, and has maximum entropy within Ω.This follows from the stated KL Pythagorean property and the projection of pdata onto Θ.
- 4.1 Exponential Family and Moment Matching Estimator: As K increases, KL(qθ|pθ) decreases monotonically, making both q ˆθMME and p ˆθMME closer to p ˆθMLE when computational cost permits.The paper connects this monotonic decrease to smaller divergences from the maximum-likelihood distribution.
- 4.2 General ConvNet-EBM and Generalized Moment Matching Estimator: For general ConvNet energies, short-run learning uses a generalized moment-matching equation with h(x,θ)=∂fθ(x)/∂θ, rather than necessarily maximizing likelihood.The learned qθ remains valid by matching this parameter-dependent statistic to its data expectation.
5 Experimental Results
Experiments evaluate short-run MCMC on synthesis, interpolation, reconstruction, and hyperparameter sensitivity. The method produces realistic samples, supports smooth interpolation and faithful reconstruction, and responds predictably to changes in steps, noise, and model capacity.
- 5.2 Interpolation: The method produces smooth interpolations whose intermediate samples resemble realistic faces.The experiment uses K = 100 steps on CelebA and varies the interpolation coefficient ρ from 0 to 1.
- 5.3 Reconstruction: Short-run MCMC reconstructs observed images faithfully and achieves competitive per-pixel MSE on 1,000 leave-out examples.Reconstruction initializes latent variables from p0 and minimizes the least-squares loss between the observed image and the generated output.
- 5.4 Influence of Hyperparameters: Decreasing K reduces synthesis quality, while shorter chains produce colder EBMs with more dominant gradient descent dynamics.The authors describe performance degradation with small K as graceful and consider K = 100 reasonable.
- 5.4 Influence of Hyperparameters: Lower injected-noise variance improves negative-example fidelity, so σ2 should be minimized while preserving training stability.The learning rate and σ2 may be gradually decreased during training.
- 5.4 Influence of Hyperparameters: Increasing the first-layer feature-map count nf improves synthesis quality until computational resources are exhausted.The model-complexity experiment uses K = 100 and evaluates synthesis with IS and FID.
6 Conclusion
The conclusion positions short-run MCMC as complementary to, rather than a replacement for, energy-based modeling. It also identifies possible relevance to learning attractor dynamics in neuroscience.
- The authors do not advocate abandoning EBMs and instead aim ultimately to learn valid EBMs.
- The learned non-convergent short-run MCMC may help understand attractor dynamics used in neuroscience.
7 Appendix
The appendix supplies theoretical context, implementation details, network specifications, computational cost, and toy-density analyses. It connects short-run MCMC to estimating equations, residual-network dynamics, and learned low-entropy energy models.
- Theory: The appendix frames the theoretical analysis through KL divergence and generalized estimating equations leading to maximum-likelihood estimation.The supplied derivation relates the optimal estimating function to the score and Fisher information.
- Implementation: The implementation parameterizes negative energy as -fθ(x)/T and uses Langevin updates driven by derivatives of the convolutional network.The code initializes samples from uniform noise, performs K derivative-based updates, and trains with the difference between observed and synthesized energies.
- Computational Cost: A 64×64 CelebA run with K = 100 and nf = 64 requires 16 hours for 100,000 parameter updates on four Titan Xp GPUs.Each short-run MCMC iteration requires K derivatives of the convolutional energy network.
- Toy Examples in 1D and 2D: Training with K1 = 100 and sampling with K2 > K1 causes over-saturation, showing that training and sampling step counts need not coincide.The appendix interprets short-run MCMC as a variable-depth residual network and explicitly varies K2 at sampling time.
- Toy Examples in 1D and 2D: Toy experiments find that MCMC-sample density closely matches the true density, while the learned EBM has much lower entropy or temperature.The learned energy captures the true density’s modes but has a much larger scale and concentrates density near the global energy minimum.