Source-linked AI summary

Sampling using $SU(N)$ gauge equivariant flows

Denis Boyda, Gurtej Kanwar, Sébastien Racanière, Danilo Jimenez Rezende, Michael S. Albergo, Kyle Cranmer, Daniel C. Hackett, Phiala E. Shanahan

arXiv:2008.05456v2hep-latcs.LGstat.ML

TL;DR

The paper addresses efficient sampling for SU(N) lattice gauge theories while preserving gauge symmetry by construction. It introduces matrix-conjugation-equivariant flows for SU(N) variables, applies them to single-variable distributions and two-dimensional SU(2) and SU(3) lattice gauge theories, and reports correct observables with predicted statistical scaling.

  • Problem

    Flow-based samplers for lattice gauge theories need to incorporate gauge symmetries directly rather than relying on models to learn them approximately.

  • Method

    The paper constructs invertible SU(N) transformations equivariant under matrix conjugation, usable as kernels in gauge-equivariant coupling layers.

  • Results

    The models achieve ESSs greater than 73% on all reported SU(3) distributions at β = 9.

  • Takeaways & Limitations

    The construction enables gauge-symmetric flow-based samplers for single SU(N) variables and two-dimensional SU(2) and SU(3) lattice gauge configurations.

  • Takeaways & Limitations

    The proof-of-principle implementation does not optimize architecture or training for expressivity or efficiency.

Abstract

from arXiv · show

We develop a flow-based sampling algorithm for $SU(N)$ lattice gauge theories that is gauge-invariant by construction. Our key contribution is constructing a class of flows on an $SU(N)$ variable (or on a $U(N)$ variable by a simple alternative) that respect matrix conjugation symmetry. We apply this technique to sample distributions of single $SU(N)$ variables and to construct flow-based samplers for $SU(2)$ and $SU(3)$ lattice gauge theory in two dimensions.

I. INTRODUCTION

The paper develops flow-based samplers for lattice gauge theories, using invertible transformations to approximate target distributions while enabling unbiased observable estimates. Its central extension is a gauge-equivariant construction for SU(N) and U(N) variables based on matrix-conjugation-equivariant kernels.

  • I. INTRODUCTION: The paper develops flow-based models that use invertible transformations to approximate target distributions and enable efficient sampling of lattice field theories.The model combines an easily sampled prior with an invertible function having a tractable Jacobian.
  • I. INTRODUCTION: The contribution is a universal unitary-group kernel construction: an eigenvalue transformation equivariant under permutations is equivariant under matrix conjugation.The construction applies to SU(N), with a similar adaptation for U(N), and connects to the maximal torus and Weyl group.
  • I. INTRODUCTION: Flow-based samplers can produce unbiased observable estimates through reweighting or a Metropolis accept/reject step, despite approximating the target density.A flow-based Markov chain uses model samples as independent proposals and guarantees asymptotic exactness through Metropolis correction.
  • I. INTRODUCTION: Training optimizes a modified KL divergence using samples from the model and stochastic gradients of log q(U) + S(U), without requiring training data from existing samplers.The unknown constant log Z is removed because it does not affect gradients or the minimizer.
  • I. INTRODUCTION: Reweighting quality is measured by an effective sample size normalized so that ESS = 1 for a perfect model.Reweighting is computationally efficient when observable evaluation is inexpensive relative to drawing samples.

B. Symmetries in flow models

The section explains how flow models can preserve lattice gauge symmetries by combining invariant priors with equivariant transformations. The approach imposes gauge and translational symmetry simultaneously without gauge-fixing away degrees of freedom.

  • B. Symmetries in flow models: The model uses a gauge-valued site field to transform each link as Ω(x)U_μ(x)Ω†(x + μ̂), while preserving the action under gauge transformations.The gauge symmetry group is one of the internal symmetries considered alongside translations and hypercubic transformations.
  • B. Symmetries in flow models: Explicitly imposing symmetries restricts the variational family to symmetry-respecting maps and can reduce variance or correlations caused by symmetry breaking.For gauge symmetry, the corresponding distribution factorizes into uniform pure-gauge directions and nontrivial gauge-invariant directions.
  • B. Symmetries in flow models: Gauge fixing can factor out pure-gauge degrees of freedom, but maximal-tree gauge fixing is not translationally invariant.The paper instead acts on all SU(N) gauge-field degrees of freedom while imposing gauge and translational symmetries together.
  • B. Symmetries in flow models: Gauge equivariance follows by sampling an exactly invariant prior and applying an invertible transformation whose coupling layers individually commute with the symmetry action.The construction handles spacetime symmetries through equivariant networks and symmetric frozen/updated-link decompositions, while internal symmetries require invariant context functions and appropriate updates.

C. Gauge equivariance

Gauge-equivariant coupling layers transform open loops through matrix-conjugation-equivariant kernels, preserving the gauge transformation law of the output configuration. The paper specializes this construction to SU(N) using eigenvalue-based kernels and plaquette updates.

  • C. Gauge equivariance: The coupling-layer construction changes variables to open loops, applies a matrix-conjugation-equivariant kernel, and changes variables back to links.
  • C. Gauge equivariance: In lattice applications, plaquettes are transformed while traces of unmodified plaquettes provide gauge-invariant inputs to the context functions.
  • C. Gauge equivariance: Invertibility requires active and passive loop updates to remain disjoint, while changing loop subsets across layers allows all links to be updated.
  • C. Gauge equivariance: The novel contribution is a matrix-conjugation-equivariant transformation for SU(N), with a straightforward U(N) adaptation, extending the earlier U(1) construction.
  • C. Gauge equivariance: A kernel is an invertible single-group-variable map satisfying h(XUX^-1) = Xh(U)X^-1, which preserves gauge equivariance when applied to open loops.
  • C. Gauge equivariance: The SU(N) construction acts on eigenvalues while remaining permutation equivariant, thereby preserving conjugation equivariance and manipulating gauge-invariant spectral distributions.

A. Target densities

The paper defines conjugation-invariant target densities on SU(N) through eigenvalue spectra and constructs permutation-equivariant spectral flows, illustrated explicitly for SU(2).

  • The target family is invariant under matrix conjugation and therefore depends only on the spectrum, with coefficients controlling shape and β controlling scale.
  • The c(0) coefficients exactly match the marginal distribution of open plaquettes in two-dimensional lattice gauge theory, while additional coefficient sets probe related single-peaked densities.These coefficient sets are intended to represent local densities relevant to two-dimensional lattice gauge theory.
  • Combining a uniform Haar prior with one permutation-equivariant kernel imposes matrix-conjugation symmetry exactly and tests the transformations’ ability to reproduce target densities.Model quality is assessed using effective sample size and density plots.
  • Eigenvalue-space plots use Lebesgue measure, whereas the full SU(N) model uses Haar measure and includes conjugacy-class volume in the induced eigenvalue measure.
  • B. Flows on SU(2): For SU(2), eigenvalue permutations act as θ → −θ, so the flow applies an invertible finite-interval transformation on the canonical cell [0, π] and extends it symmetrically.The construction negates negative angles, applies the interval flow, then restores the original sign when needed.
  • B. Flows on SU(2): The SU(2) flow precisely reproduces distribution peaks and exact mirror symmetry, with deviations only in tails below roughly 10^-4.All models reached ESS above 97% for every coefficient set, and the symmetry is enforced by construction.

C. Flows on SU(3)

The SU(3) construction maps eigenvalue phases into a canonical cell, applies invertible spline flows there, and restores permutation equivariance across all six cells.

  • SU(3) eigenvalues are represented by two angular variables, producing six cells related by permutations of the three eigenvalues.
  • The algorithm enumerates eigenvalue-phase permutations, selects the ordering satisfying the canonical condition, shifts the connected domain, and applies an invertible flow on the triangular cell.The inverse permutation restores the final eigenvalue phases after flowing in canonical coordinates.
  • The canonical cell is a simplex in the maximal torus, whose boundaries occur where eigenvalues become degenerate; its chosen SU(3) region is shown in orange.
  • An alternative SU(3) construction averages over all permutations, but its N! cost makes it unsuitable for large N.
  • The SU(3) flow accurately reproduces target peaks and six-fold permutation symmetry, achieving ESS greater than 73% on all tested distributions.At β = 9, deviations occur only in tails below roughly 10^-3; the plaquette marginal achieves a particularly high ESS.
  • D. Flows on SU(N): The general SU(N) procedure extracts and sorts eigenvalue angles, maps them into a canonical cell, applies spline transformations, and reverses the recorded permutation afterward.Sorting avoids checking all N! permutations explicitly for large N, unlike the SU(3) construction.
  • D. Flows on SU(N): For larger N, the flow preserves the canonical simplex by mapping it to an open box, applying a boundary-preserving flow, and mapping back.
  • D. Flows on SU(N): For N = 9, multimodal coefficient sets perform worse, while flows trained on c(0) achieve ESS above 90% for every N from 10 through 100.All model distributions retained exact permutation invariance.

IV. APPLICATION TO SU(2) AND SU(3) LATTICE GAUGE THEORY IN 2D

The authors apply matrix-conjugation-equivariant kernels to flow-based samplers for two-dimensional SU(2) and SU(3) lattice gauge theory on periodic lattices.

  • Gauge-equivariant kernels acting on plaquette loops are used to construct flow-based samplers for two-dimensional lattice gauge theory.
  • The study investigates SU(2) and SU(3) models on 16 × 16 periodic lattices at approximately equivalent ’t Hooft couplings.
  • The application evaluates sampler exactness and whether model symmetries are exactly built in or approximately learned after training.

A. Model architecture and training

The models use gauge-equivariant coupling layers and spectral flows to sample SU(2) and SU(3) lattice gauge theories, with training aided by smaller-volume transfer. In two dimensions, the resulting ensembles reproduce analytical observables and expected statistical-error scaling.

  • Model architecture: The architecture uses 48 gauge-equivariant coupling layers acting on plaquettes, with alternating rotations and translations so every link is updated after eight layers.The models use a uniform prior and convolutional context networks acting on gauge-invariant quantities.
  • Model architecture: For SU(2), the spectral flow is permutation equivariant and uses a four-knot spline on θ ∈ [0, π], with knot positions conditioned on neighboring plaquette invariants.The conditioning networks have 32 channels in each of two hidden layers.
  • Model architecture: For SU(3), the spectral flow maps eigenvalues in the canonical triangular cell to an open box and applies a two-step conditioned spline transformation.The horizontal coordinate is transformed first, followed by the vertical coordinate conditioned on the new horizontal coordinate.
  • Training strategy: Transferring an SU(3) model trained on 8×8 to 16×16 reaches similar model quality in far fewer training steps than random initialization.The advantage follows from learning local correlation structure on smaller volumes, with exponentially small finite-volume corrections in the relevant regime.
  • Training and sampling: Flow-based Markov chains use independent model proposals followed by Metropolis accept/reject steps, preserving exactness in the infinite-statistics limit.Training uses batches of 3072, Adam optimization, a modified KL-divergence estimate, and ESS monitoring; final ESS values are reported in Table V.
  • Validation: Flow-based SU(2) and SU(3) samples are statistically consistent with analytical observables, while Re W11 errors follow the expected 1/√n scaling.The observable checks include Wilson loops, Polyakov-loop quantities, and Polyakov-loop two-point functions; the SU(3) Re W11 scaling is illustrated at β = 6.

C. Symmetries

The models impose several exact internal and spacetime symmetries, while hypercubic symmetry is learned and residual translations can be symmetrized post hoc. Explicit model symmetries are generally preferable because brute-force averaging improved ESS only about twofold at 16-fold cost.

  • C. Symmetries: Explicit gauge, center, conjugation, and large-subgroup translational symmetries are built into the flow-based models to reduce sampling inefficiencies.Model symmetry breaking increases reweighting variance or Markov-chain correlations, whereas exact symmetries are preserved in proposals.
  • C. Symmetries: Gauge transformations leave both the flow-based and true actions exactly invariant across 32 SU(3) configurations at β = 6.The invariance was measured along smoothly varied gauge transformations.
  • C. Symmetries: Translations preserve lines of constant effective action, although residual fluctuations remain across translations because the tiled pattern breaks Z4 × Z4 symmetry.The spatial structure of these fluctuations is shown separately.
  • C. Symmetries: Hypercubic symmetry is learned rather than imposed directly, yielding approximate invariance under all 8 rotations and reflections in two dimensions.The group has only 8 elements in two dimensions, making learning it practical in this implementation.
  • C. Symmetries: An additional 16-fold translation average increased ESS by roughly two but cost a factor of 16, making post hoc residual-symmetry averaging counterproductive.The comparable histogram widths indicate only O(1) improvement in the spread of reweighting factors.

Appendix A: Proof that equivariance under matrix conjugation can be represented as equivariance under eigenvalue permutation

The appendix proves that conjugation-equivariant diffeomorphisms of SU(N) or U(N) are characterized by Weyl-equivariant diffeomorphisms on a maximal torus. For diagonal matrices, this reduces conjugation equivariance to permutation equivariance of eigenvalues.

  • Appendix A: A conjugation-equivariant diffeomorphism on G restricts to a Weyl-equivariant diffeomorphism on a maximal torus T.For SU(N) and U(N), T is the subgroup of diagonal matrices and the Weyl group permutes diagonal elements.
  • Appendix A: The key lemma shows that any matrix commuting with D also commutes with f(D), enabling consistency across different diagonalizations of the same group element.Equal eigenvalues force corresponding entries of f(D) to coincide, preserving the relevant eigenspaces.
  • Appendix A: Conversely, a Weyl-equivariant function on T extends to a conjugation-equivariant function on G, and invertibility on T implies invertibility on G.The extension is well-defined because different diagonal decompositions are related through the normalizer and the centralizer argument.
  • Appendix A: For SU(2), the maximal torus is U(1), so any bijection of U(1) respecting its two-element Weyl action gives an equivariant bijection of SU(2).The nontrivial Weyl element exchanges the two eigenvalues.

Appendix B: Details of permutation equivariance of SU(N) spectral flows

The appendix constructs an (N − 1)-simplex in the maximal torus whose Weyl-group orbit tiles the torus into cells of regular points. Its boundary maps to repeated eigenvalues, while its interior maps to regular group elements.

  • Appendix B: The vertices y1, …, yN define an (N − 1)-simplex Ψ whose exponential image is a cell C in the maximal torus.The cell is a connected region of regular points bounded by eigenvalue degeneracies.
  • Appendix B: Boundary points of Ψ exponentiate to matrices with repeated eigenvalues, because at least one barycentric coordinate vanishes.The γN = 0 case produces an angular separation of 2π, while γk = 0 for k < N collapses adjacent angles.
  • Appendix B: The simplex is obtained by barycentric subdivision and edge-length adjustment so its Weyl-group orbit tiles the torus without holes or overlaps.The tiling constraint requires vertex-coordinate differences to lie in {0, 2π}, giving αk = k.
  • Appendix B: Interior points of Ψ have all γk > 0 and exponentiate to regular points with no eigenvalue pair equivalent modulo 2π.Thus Ψ separates regular points from the degeneracy boundaries in the maximal torus.
  • Appendix B: The resulting vertices and translated copies form a Bravais lattice, with the canonical cell identified as the fundamental simplex and its Weyl orbit as the root polytope.This connects the construction to standard lattice-theoretic fundamental domains.

3. Proof that Algorithm 1 projects into Ψ

The appendix proves that Algorithm 1 always returns canonical angles inside the simplex Ψ. It does so by explicitly expressing the output as a convex combination of Ψ’s vertices.

  • 3. Proof that Algorithm 1 projects into Ψ: Algorithm 1 outputs canonical angles θcanon that lie in the convex hull Ψ of the vertices yk.The proof constructs the barycentric weights γk for the convex combination.
  • 3. Proof that Algorithm 1 projects into Ψ: The constructed weights satisfy γk ≥ 0 and sum to one, which establishes membership in Ψ.The proof concludes by identifying the convex-combination result with θcanon.

4. Full algorithm

The full algorithm constructs an equivariant SU(N) coupling layer whose output and log-det-Jacobian respect matrix conjugation symmetry, while differentiating through unitary diagonalization.

  • The algorithm accumulates all log-det-Jacobians before returning the transformed variable and total LDJ.
  • The coupling layer returns U′ equivariant under SU(N) matrix conjugations and an LDJ invariant under those conjugations.
  • The LDJ contains Haar-measure density terms for eigenvalues, while the forward and backward Jacobian factors for ζ cancel.
  • Unitary diagonalization defines eigenvalues w and eigenvectors P through U = P D P†, with gradients propagated from w and P back to U.
  • The diagonalization Jacobian is fixed by differentiating the decomposition and selecting valid components despite phase ambiguities in eigenvectors.

Appendix D: Conjugation equivariant maps on SU(3) via averaging

The appendix builds SU(3) conjugation-equivariant maps by averaging Lie-algebra diffeomorphisms over the Weyl group and identifies conditions ensuring they descend to torus diffeomorphisms.

  • Averaging a suitable Lie-algebra map produces a Weyl-group-equivariant map that can descend to the torus and preserve diffeomorphic structure.
  • The construction requires periodicity-compatible shifts satisfying Equation (D3), and its simple single-circle form performed worse than the main-body networks.
  • For maps of the form h̃(x,y) = (f(x), f(y)), invertibility follows because the Jacobian determinant is proportional to sums of pairwise products of f′ values.
  • If f is connected to the identity by a path of diffeomorphisms, the descended map is a Weyl-equivariant diffeomorphism of T.
  • Any circle diffeomorphism from Ref. [15], including mixtures of NCPs, Möbius maps, or splines, defines an equivariant diffeomorphism of SU(3).

Appendix E: The case of U(N)

For U(N), the determinant constraint is absent, enabling a simpler permutation-equivariant torus flow; a U(3) test rapidly converged with high effective sample size.

  • U(N) flows can be built directly on the N-dimensional torus by mapping it to R^N, stacking permutation-equivariant layers, and projecting back.
  • U(3) experiments at β = 1, 5, 9 achieved effective sample sizes above 95% in every case.
Loading 2008.05456v2…