Source-linked AI summary
An exact mapping between the Variational Renormalization Group and Deep Learning
Pankaj Mehta, David J. Schwab
TL;DR
Deep learning’s theoretical underpinnings remain an open question. This paper establishes a one-to-one mapping between RBM-based deep neural networks and variational renormalization group, finding self-organized coarse-graining reminiscent of Kadanoff block renormalization.
Problem
Deep learning’s success raises natural questions about its theoretical underpinnings.
Method
The paper constructs a one-to-one mapping between variational renormalization group and RBM-based deep neural networks, illustrating it with one- and two-dimensional Ising models.
Results
The analyzed deep neural networks self-organize to implement a coarse-graining procedure reminiscent of Kadanoff block renormalization.
Takeaways & Limitations
The mapping suggests that deep learning may implement a generalized renormalization-group-like scheme for learning relevant features from data.
Takeaways & Limitations
Applying the mapping to improve real-space renormalization remains a possible route because variational RG has often been limited by inadequate approximations.
Abstract
from arXiv · showhide
Deep learning is a broad set of techniques that uses multiple layers of representation to automatically learn relevant features directly from structured data. Recently, such techniques have yielded record-breaking results on a diverse set of difficult machine learning tasks in computer vision, speech recognition, and natural language processing. Despite the enormous success of deep learning, relatively little is understood theoretically about why these techniques are so successful at feature learning and compression. Here, we show that deep learning is intimately related to one of the most important and successful techniques in theoretical physics, the renormalization group (RG). RG is an iterative coarse-graining scheme that allows for the extraction of relevant features (i.e. operators) as a physical system is examined at different length scales. We construct an exact mapping from the variational renormalization group, first introduced by Kadanoff, and deep learning architectures based on Restricted Boltzmann Machines (RBMs). We illustrate these ideas using the nearest-neighbor Ising Model in one and two-dimensions. Our results suggests that deep learning algorithms may be employing a generalized RG-like scheme to learn relevant features from data.
I. OVERVIEW OF VARIATIONAL RG
Variational RG coarse-grains a spin system by introducing hidden spins that represent larger-scale degrees of freedom. Its variational parameters are chosen to preserve long-distance behavior by minimizing the free-energy difference.
- Coarse-graining: RG introduces M < N hidden spins to represent coarse-grained degrees of freedom with short-scale fluctuations averaged out.Coarse-graining typically increases a characteristic length scale, such as lattice spacing.
- Coarse-graining: Interactions among visible spins induce interactions among hidden spins, defining a new Hamiltonian and a mapping from {K} to {˜K}.The exact mapping depends on the chosen RG scheme.
- Variational transformation: Kadanoff’s variational RG couples visible and hidden spins through Tλ({vi}, {hj}) before marginalizing over visible spins.The resulting hidden-spin Hamiltonian provides the coarse-grained description.
- Block spin example: Block spin renormalization groups adjacent spins into effective block variables and iterates this construction across larger length scales.The lattice spacing doubles after each iteration.
- Variational transformation: The variational parameters λ are selected to minimize the free-energy difference ∆F between the physical and coarse-grained systems.Exact preservation of long-distance observables motivates this optimization, but the exact condition is generally unattainable.
II. RBMS AND DEEP NEURAL NETWORKS
RBMs model binary data with visible and hidden spins, while stacked RBMs form deep architectures by reusing each hidden layer as the next layer’s input. Training uses variational distribution matching.
- RG connection: The variational RG procedure has a natural interpretation as a deep learning scheme based on RBMs.The paper develops this interpretation as an exact mapping.
- RBM representation: RBMs introduce hidden spin variables coupled to visible units to model the probability distribution of binary data.Visible spins can encode pixels, while the data distribution captures statistical structure across examples.
- RBM representation: The visible-hidden interactions are parameterized through an energy function and define joint, visible-marginal, and hidden-marginal distributions.The parameters include hidden biases, visible biases, and interaction weights.
- Training: RBM parameters are chosen by minimizing the Kullback-Leibler divergence between the true data distribution and the variational visible distribution.This is the unsupervised-learning objective used for the model.
- Deep architecture: Stacked RBMs create a DNN by treating one RBM’s trained hidden layer as the visible layer for the next RBM.Hidden-layer activities responding to data samples become training data for subsequent layers.
III. MAPPING VARIATIONAL RG TO DEEP LEARNING
The paper establishes a one-to-one correspondence between variational RG and RBM-based DNNs. Under this mapping, RG operators act as conditional probabilities and coarse-grained Hamiltonians describe hidden-layer distributions.
- Exact correspondence: The equation relating H[{vi}] and T({vi}, {hj}) defines a one-to-one mapping between variational RG and RBM-based DNNs.H[{vi}] encodes the data probability distribution P({vi}).
- Exact correspondence: The coarse-grained RG Hamiltonian HRG λ [{hj}] also describes the hidden spins in the RBM.Equivalently, the hidden-spin marginal distribution has Boltzmann form with this Hamiltonian.
- Probability interpretation: The operator Tλ({vi}, {hj}) can be interpreted as a variational approximation to the conditional probability of hidden spins given visible spins.This provides a probability-theoretic interpretation of variational RG.
- Exact correspondence: When the RG transformation is exact, the variational Hamiltonian equals the true Hamiltonian describing the data.The condition is TrhjeTλ({vi},{hj}) = 1.
- Approximation schemes: Variational RG and machine learning use distinct variational approximation schemes for coarse graining, and the correspondence extends beyond a specific energy form to any Boltzmann Machine.The distinction is between Hamiltonian/free-energy approximations and KL-divergence minimization.
- One-dimensional Ising model: For the 1-D Ising model, decimation halves the spins, doubles lattice spacing, and relates successive couplings through the square of the hyperbolic tangent.The same transformation is represented by successive layers in the deep architecture.
IV. EXAMPLES
The paper illustrates the RG–deep-learning correspondence analytically for the one-dimensional Ising model and numerically for the two-dimensional nearest-neighbor Ising model.
- Examples: The one-dimensional Ising model provides an exactly solvable example, while the two-dimensional model is explored numerically with an RBM-based deep architecture.These examples are used to examine the mapping in concrete settings.
A. One dimensional Ising Model
The one-dimensional Ising model provides an exact setting where RG decimation corresponds to a deep architecture that successively coarse-grains spin variables. This construction is interpretable but omits half of the visible spins.
- The one-dimensional Ising model consists of binary spins on a lattice, with neighboring alignment favored by a ferromagnetic coupling J.
- RG decimation marginalizes every other spin, doubles the lattice spacing, and produces a new effective interaction J(1).
- After n successive transformations, the coupling flows toward J = 0 as n approaches infinity.
- The deep architecture implements repeated decimation because each layer’s hidden spins correspond to spins remaining after marginalizing the layer below.
- The simple architecture is easy to interpret and construct, but contains no information about half of the visible spins that do not couple to the hidden layer.
B. Two dimensional Ising Model
The paper applies RBM-based deep learning to coarse-grain samples from the two-dimensional Ising model near its critical point. Training produces an emergent local block-spin structure whose receptive fields grow across layers and whose compressed representations preserve macroscopic features qualitatively.
- The two-dimensional Ising model has a phase transition at J/(kBT ) = 0.4352, where its correlation length diverges and near-critical coarse-graining is productive.
- 20, 000 samples from a periodic 40 × 40 model at J = 0.408 trained a four-layer RBM-based network with 1600, 400, 100, and 25 spins.
- Each hidden spin couples locally to a block below, with similarly sized blocks whose characteristic size increases at higher layers.
- The local block-spin structure emerges from training, suggesting that the DNN self-organizes to implement block-spin renormalization.
- With only 25 spins in the top layer, reconstructions qualitatively reproduce macroscopic features of individual samples at a compression ratio of 64.
V. DISCUSSION
The paper connects RBM-based deep neural networks with variational renormalization, illustrating the mapping on one- and two-dimensional Ising models. It argues that deep networks may learn important features through a generalized RG-like coarse-graining procedure, while also identifying opportunities for deep learning to improve real-space renormalization.
- V. DISCUSSION: A one-to-one mapping links RBM-based deep neural networks with the variational renormalization group.The paper analytically constructs a DNN for the 1D Ising model and numerically examines the 2D Ising model.
- V. DISCUSSION: The mapped networks self-organize into a coarse-graining procedure reminiscent of Kadanoff block renormalization.This behavior was observed across the paper’s analytical 1D and numerical 2D Ising-model examples.
- V. DISCUSSION: RG explains how coarse-graining extracts relevant features while integrating out short-distance fluctuations and diminishing irrelevant operators.Relevant operators become increasingly important at larger scales, whereas irrelevant operators have diminishing effects on large-scale properties.
- V. DISCUSSION: Tensor-network and matrix-product-state techniques motivate importing low-redundancy ideas such as entanglement entropy and disentanglers into deep learning.The paper presents this as an open question rather than an established result.
- V. DISCUSSION: Deep-learning techniques may help address variational RG’s difficulty in making good approximations for complicated physical systems.The paper frames this as a possible route for overcoming limitations of real-space renormalization techniques.
Appendix A: Learning Deep Architecture for the Two-dimensional Ising Model
The appendix describes the unsupervised training setup for stacked RBMs on the two-dimensional Ising model, including optimization, sampling, batching, and regularization choices.
- Appendix A: Learning Deep Architecture for the Two-dimensional Ising Model: Stacked RBMs were trained only in the unsupervised learning phase on 40,000 two-dimensional Ising-model samples with J = 0.408.The individual RBMs used contrastive divergence for 200 epochs.
- Appendix A: Learning Deep Architecture for the Two-dimensional Ising Model: Training used momentum 0.5 and mini-batches of size 100.These settings were applied during contrastive-divergence training.
- Appendix A: Learning Deep Architecture for the Two-dimensional Ising Model: L1 regularization was implemented with strength 2 × 10^-4.
Appendix B: Visualizing Effective Receptive Fields
The effective receptive field visualizes which visible-layer spins influence each hidden spin, using recursively composed weight matrices across network layers.
- Appendix B: Visualizing Effective Receptive Fields: Each column of the effective receptive field matrix encodes the visible-layer receptive field of one hidden spin.The matrix for layer l is denoted r^(l), and the visible layer is l = 0.
- Appendix B: Visualizing Effective Receptive Fields: The receptive field is computed by setting r^(1) = W^(1) and recursively applying r^(l) = r^(l−1)W^(l) for l > 1.The weight matrices contain the interlayer weights w_ij.
- Appendix B: Visualizing Effective Receptive Fields: The resulting effective receptive field measures how strongly a hidden spin influences spins in the visible layer.