Source-linked AI summary

Approaching the Thermodynamic Limit with Neural-Network Quantum States

Luciano Loris Viteritti, Riccardo Rende, Subir Sachdev, Giuseppe Carleo

arXiv:2602.02665v1cond-mat.str-elcond-mat.dis-nnquant-ph

TL;DR

Large frustrated quantum systems remain difficult to simulate accurately enough for controlled thermodynamic-limit extrapolation. The paper introduces a distance-aware Spatial Attention bias for Transformer-based Neural-Network Quantum States, which stabilizes large-scale optimization and yields state-of-the-art results, including M0 = 0.148(1) on triangular-lattice clusters up to 42 × 42 sites. These calculations also expose strongly non-local sign structure and generalize to the J1-J2 model on a 20 × 20 square lattice.

  • Problem

    Strongly correlated quantum systems require very large lattices for controlled finite-size scaling, but frustrated two-dimensional systems remain difficult for existing numerical methods.

  • Method

    The paper introduces Spatial Attention, a Hamiltonian-agnostic distance-dependent kernel with a learned inverse length scale inside Transformer-based Neural-Network Quantum States.

  • Results

    The approach reaches state-of-the-art variational results on triangular-lattice clusters up to 42 × 42 sites, with thermodynamic-limit magnetization M0 = 0.148(1).

  • Takeaways & Limitations

    Large-system analysis indicates a highly non-local sign structure in the triangular-lattice Heisenberg antiferromagnet and supports an intrinsic sign problem.

  • Takeaways & Limitations

    Further scaling is primarily constrained by available computational resources, rather than by conceptual or methodological limitations.

Abstract

from arXiv · show

Accessing the thermodynamic-limit properties of strongly correlated quantum matter requires simulations on very large lattices, a regime that remains challenging for numerical methods, especially in frustrated two-dimensional systems. We introduce the Spatial Attention mechanism, a minimal and physically interpretable inductive bias for Neural-Network Quantum States, implemented as a single learned length scale within the Transformer architecture. This bias stabilizes large-scale optimization and enables access to thermodynamic-limit physics through highly accurate simulations on unprecedented system sizes within the Variational Monte Carlo framework. Applied to the spin-$\tfrac12$ triangular-lattice Heisenberg antiferromagnet, our approach achieves state-of-the-art results on clusters of up to $42\times42$ sites. The ability to simulate such large systems allows controlled finite-size scaling of energies and order parameters, enabling the extraction of experimentally relevant quantities such as spin-wave velocities and uniform susceptibilities. In turn, we find extrapolated thermodynamic limit energies systematically better than those obtained with tensor-network approaches such as iPEPS. The resulting magnetization is strongly renormalized, $M_0=0.148(1)$ (about $30\%$ of the classical value), revealing that less accurate variational states systematically overestimate magnetic order. Analysis of the optimized wave function further suggests an intrinsically non-local sign structure, indicating that the sign problem cannot be removed by local basis transformations. We finally demonstrate the generality of the method by obtaining state-of-the-art energies for a $J_1$-$J_2$ Heisenberg model on a $20\times20$ square lattice, outperforming Residual Convolutional Neural Networks.

I. INTRODUCTION

The paper introduces Spatial Attention, a distance-aware inductive bias for Transformer-based Neural-Network Quantum States, to stabilize optimization on large frustrated lattices. Applied to triangular- and square-lattice Heisenberg models, it enables large-scale simulations, controlled finite-size scaling, and state-of-the-art variational results.

  • Motivation and method: Spatial Attention reweights Transformer attention with a learned distance-dependent kernel, providing a Hamiltonian-agnostic geometric bias inspired by the cluster property.The modification preserves global connectivity and can represent multiple correlation scales through independent length scales across attention heads.
  • Results: The method enables controlled finite-size scaling of energy and magnetization, including extraction of low-energy spin-excitation velocities and susceptibilities.These quantities are presented as experimentally relevant outputs of the large-system calculations.
  • Results: The optimized states show an exponentially decreasing sign overlap with system size, indicating highly non-local sign structure while stopping short of a no-go theorem for all local basis transformations.The analysis supports an intrinsic sign problem in the triangular-lattice Heisenberg antiferromagnet.
  • Results: 20 × 20 square lattice: the approach obtains state-of-the-art energies for the J1-J2 Heisenberg model compared with existing variational methods.This result supports generality beyond the triangular-lattice benchmark.
  • Motivation and method: Large-scale optimization becomes more stable because the architecture explicitly encodes the spatial attenuation that is otherwise difficult to learn from random initialization.The authors describe this improvement as especially important in the large-system regime.

III. HEISENBERG TRIANGULAR LATTICE

The triangular-lattice Heisenberg study evaluates spatial-attention ViT wave functions on large periodic clusters, achieving accurate energies and controlled scaling of energy and magnetization. The results include state-of-the-art variational energies and a strongly renormalized thermodynamic-limit magnetization.

  • The model uses periodic L × L clusters with L a multiple of 3, enforcing compatibility with three-sublattice 120° Néel order.
  • The ViT achieves lower variational energies than prior Gutzwiller-projected and recurrent-neural-network states across accessible sizes, improving relative energy accuracy by more than an order of magnitude.
  • Finite-size energies follow linear scaling in 1/L^3 for L ≥24, enabling reliable extrapolation to the thermodynamic limit and state-of-the-art results below iPEPS estimates up to L = 42.
  • Relative energy error and V-score remain essentially stable as system size grows, despite an approximately fixed parameter count of P ≈4.5×10^5.
  • The extrapolated magnetization is M0 = 0.148(1), the smallest available estimate and substantially below the linear spin-wave prediction of 0.239.

B. Finite-Size Scaling

Large periodic-lattice wave functions enable controlled finite-size scaling of energy, magnetization, and experimentally relevant response parameters. Fitting the large-L data yields renormalized spin-wave and susceptibility coefficients that differ from simpler analytical estimates.

  • Finite-size scaling: Large periodic lattices permit finite-size scaling analyses unavailable to methods restricted to the thermodynamic limit.Accurate variational states on periodic clusters provide access to energies and order parameters across system sizes.
  • Low-energy description: The low-energy parameterization uses orthogonal unit vectors n1 and n2, represented through a unit-length complex spinor z.The constraints are |n1|2 = |n2|2 = 1 and n1 · n2 = 0.
  • Low-energy description: Spin-wave expansion around the ordered state produces a harmonic theory with three modes governed by spin-wave velocities and uniform susceptibilities.The three real fluctuation fields πℓ enter the quadratic expansion.
  • Extrapolated parameters: The fitted coefficients are c̄ ≈ 3.88 and cχ ≈ 0.061, differing from spin-wave estimates c̄SW ≈ 3.14 and cχSW ≈ 0.086.The comparison also includes the Schwinger-boson estimate c̄SB ≈ 3.15.

C. Non-Local Sign Structure

The triangular-lattice ground-state sign structure is tested against local and extended prescriptions using mean sign overlap. Agreement is nearly exact on a small cluster but decays exponentially with system size, indicating increasingly important non-local sign information.

  • Motivation: The sign problem asks whether a basis transformation can make the frustrated ground state non-negative, as the Marshall–Peierls rule does on bipartite lattices.No analogous prescription is known for the triangular lattice.
  • Candidate sign rules: The Huse–Elser prescription combines a classical three-sublattice phase factor with an additional three-body contribution controlled by β.Its sign prescription is ΦHE(σ) = ℜ{T3body(σ)Tclass(σ)}.
  • Small-system validation: For L = 6, the ViT achieves mean sign overlap ⟨s⟩ = 0.99999707 with the exact ground state, while Huse–Elser reaches approximately 0.96 near β ≈ 0.27.The overlap is estimated by sampling from |Ψθ(σ)|2.
  • System-size dependence: Mean sign overlap decays exponentially with the number of sites for both classical and Huse–Elser sign rules.The learned sign structure increasingly differs from these extended local prescriptions as the system grows.
  • Interpretation: The observed decay provides quantitative evidence for a highly non-local sign structure, while not proving a no-go theorem for every local basis transformation.The authors connect this behavior to an intrinsic sign problem in the triangular-lattice antiferromagnet.

IV. J1-J2 HEISENBERG SQUARE LATTICE

Spatially informed Transformer wave functions generalize beyond the triangular lattice to the frustrated J1-J2 square-lattice model. At J2/J1 = 0.5 on a 20 × 20 cluster, they achieve state-of-the-art variational energies and surpass a ResNet2 benchmark.

  • Generality: The Spatial Attention improvement is based on geometric information and is not specific to triangular lattices or nearest-neighbor Hamiltonians.The authors therefore test it on another frustrated model.
  • Model and regime: The benchmark is the spin-1/2 J1-J2 Heisenberg model on a 20 × 20 periodic square lattice at J2/J1 = 0.5.This strongly frustrated point has been associated with a possible quantum spin liquid and provides a stringent variational test.
  • Energy results: A zero-variance extrapolation gives energy −0.49684(1), while full point-group symmetry yields variational energy −0.496732(1).The reported symmetry sequence is no symmetry, translational symmetry, then full C4v symmetry.
  • Comparison: The ViT energies surpass those of the deep ResNet2 Ansatz augmented with one Lanczos step.The translationally symmetric ViT already lies below both the ResNet2 result and its zero-variance extrapolation.
  • Discussion: The broader method stabilizes large-scale optimization and makes controlled finite-size scaling feasible in frustrated two-dimensional systems.The discussion connects this capability to estimates of energies, magnetization, spin-wave velocities, and susceptibilities.

B. Architecture and Optimization Details

The work combines a Vision Transformer wave function with symmetry restoration and variance extrapolation, while analyzing its phase structure across system sizes. The architecture uses fixed patch-based embeddings and a complex-valued variational state.

  • Architecture: The Vision Transformer uses 12 attention heads, 8 layers, embedding dimension 72, and approximately 4.5 × 10^5 trainable parameters.The same parameter count is used across system sizes from L = 18 to L = 42.
  • Architecture: Spin configurations are divided into non-overlapping patches, with patch size b = 3 for the triangular lattice and b = 4 for the square lattice.The triangular-lattice choice is compatible with 120° Néel antiferromagnetic order.
  • Optimization: Symmetry projection lowers energy and variance, enabling a controlled zero-variance extrapolation through unprojected, translationally projected, and full point-group projected states.For the triangular lattice, the full point-group projection uses the C6v group and supports extrapolation from L = 18 to L = 42.
  • Optimization: The variance per site remains stable as system size increases, indicating that the variational Ansatz is size consistent.
  • Wave-function structure: The complex-valued wave function represents sign structure in the complex plane, while optimized phases remain sharply concentrated around 0 and π across system sizes.The phase peaks broaden mildly with increasing system size.

E. Comparison of Standard and Spatial Factored Attention at initialization

At initialization, standard factored attention couples a central patch uniformly to all patches, whereas spatial attention imposes distance-dependent weighting through a learnable length scale.

  • Standard attention: Standard factored attention assigns comparable weights to all patches, producing no spatial structure in the central patch representation.
  • Spatial attention: Spatial attention suppresses contributions from distant patches by introducing a learnable length scale.

A. Notation

The notation embeds a finite L × L torus in an infinite lattice and defines its reciprocal geometry through primitive and reciprocal lattice vectors. The reciprocal metric determines vector norms and the unit-cell area.

  • Real-space lattice: A finite torus with primitive vectors a1 and a2 has supercell vectors A1 = La1 and A2 = La2, with lattice points R = n1A1 + n2A2.
  • Reciprocal lattice: Reciprocal vectors satisfy Ai · Bj = 2π δij, and allowed reciprocal vectors are G = m1B1 + m2B2.
  • Reciprocal lattice: Primitive reciprocal vectors bi are defined by Bi = 2π/L bi and ai · bj = δij.
  • Reciprocal geometry: The reciprocal-space metric gij = bi · bj determines the norm of G through the quadratic form m^T g m.
  • Reciprocal geometry: The reciprocal-cell area obeys det g = 1/|a1 × a2|, so the torus volume can be written as V = L^2/√det g.

B. Ground-State Energy

The ground-state energy analysis derives finite-size corrections on a torus and identifies a universal coefficient governed by reciprocal-space geometry. The resulting triangular- and square-lattice coefficients agree with reported reference values.

  • Finite-size scaling: The torus energy density is organized as a non-universal term plus a leading finite-size correction proportional to C1/L^3.The coefficient C1 is universal and depends on the torus modular parameter rather than the underlying lattice.
  • Derivation: Poisson summation and the Fourier transform of |k| are used to obtain the discrete finite-size correction.
  • Triangular lattice: For the triangular lattice, the finite-size coefficient is expressed as C1 = −S△/(2π√det g).
  • Triangular lattice: The triangular-lattice geometric constant is S△ = 6ζ(...), with the associated numerical factor reported as 11.03417573.
  • Square lattice: For the square lattice, the calculation gives C1 = −S□/(2π) = −1.437745, in agreement with prior references.

C. Ground-State Magnetization

This section derives the universal finite-size coefficient controlling ground-state magnetization on a torus. Two regularization procedures agree on the triangular-lattice value.

  • Finite-size correction: C2 is a universal number determined only by the torus modular parameter.
  • Finite-size correction: The magnetization finite-size correction is formulated using the difference between a discrete reciprocal-lattice sum and its continuum integral.Although both terms diverge separately, their difference is finite.
  • Exponential regularization: C2 ≈ 0.670583 is obtained by extrapolating the regulated sum-integral difference to ε → 0.An exponential regulator is introduced through the dimensionless parameter ε.
  • Zeta-function regularization: Zeta-function regularization independently reproduces the exponential-cutoff result for the triangular lattice.Analytic continuation extracts the finite part without explicitly subtracting divergent contributions.
  • Square-lattice comparison: For the square lattice, the corresponding coefficient is C2 = 0.62074644.

II. NUMERICAL RESULTS FROM LINEAR SPIN-WAVE THEORY

This section applies linear spin-wave theory to the spin-1/2 triangular-lattice Heisenberg antiferromagnet. It specifies spin-wave velocities, their 1/S renormalization, uniform susceptibilities, and coefficients entering finite-size corrections.

  • Model and setup: Linear spin-wave theory is evaluated for the spin-1/2 triangular-lattice Heisenberg antiferromagnet with S = 1/2, J = 1, and a = 1.
  • Spin-wave velocities: The theory yields two independent spin-wave velocities for the triangular-lattice model.
  • Spin-wave velocities: Including 1/S corrections gives renormalized spin-wave velocities.
  • Susceptibilities and finite-size coefficients: Uniform susceptibilities are used to define the masses entering the nonlinear σ-model.The masses satisfy m1 = 4χ∥ and m2 = m3 = 4χ⊥.
  • Susceptibilities and finite-size coefficients: The spin-wave velocities and susceptibilities determine coefficients governing finite-size scaling of the ground-state energy and magnetization.
Loading 2602.02665v1…