Source-linked AI summary
Empowering deep neural quantum states through efficient optimization
Ao Chen, Markus Heyl
TL;DR
Accurate ground-state calculations for complex two-dimensional quantum systems remain difficult because existing NQS optimization methods do not scale to modern deep networks. The paper introduces MinSR, a compressed stochastic-reconfiguration method, and demonstrates accurate deep-NQS energies and numerical evidence for a gapless QSL.
Problem
Existing NQS optimization relies on stochastic reconfiguration whose cost impedes training deep, large-parameter networks for complex two-dimensional quantum systems.
Method
MinSR uses a compressed matrix with the same non-zero eigenvalues as the quantum metric, producing an SR-equivalent optimization with lower cost.
Results
MinSR trains deep networks with more than one million parameters and yields best variational energies for frustrated J1-J2 benchmarks, including E/N = −0.4976921(4).
Takeaways & Limitations
The results provide numerical evidence for a gapless QSL at the maximally frustrated square-lattice point and show that large-scale deep NQSs can approximate quantum many-body complexity.
Takeaways & Limitations
The authors identify future applications to fermionic systems and quantum chemistry, where suitable variational-wavefunction designs are still needed to lower cost and increase accuracy.
Abstract
from arXiv · showhide
Computing the ground state of interacting quantum matter is a long-standing challenge, especially for complex two-dimensional systems. Recent developments have highlighted the potential of neural quantum states to solve the quantum many-body problem by encoding the many-body wavefunction into artificial neural networks. However, this method has faced the critical limitation that existing optimization algorithms are not suitable for training modern large-scale deep network architectures. Here, we introduce a minimum-step stochastic-reconfiguration optimization algorithm, which allows us to train deep neural quantum states with up to $10^6$ parameters. We demonstrate our method for paradigmatic frustrated spin-1/2 models on square and triangular lattices, for which our trained deep networks approach machine precision and yield improved variational energies compared to existing results. Equipped with our optimization algorithm, we find numerical evidence for gapless quantum-spin-liquid phases in the considered models, an open question to date. We present a method that captures the emergent complexity in quantum many-body problems through the expressive power of large-scale artificial neural networks.
Results
MinSR reformulates stochastic reconfiguration using a compressed matrix, reducing optimization cost while preserving the traditional SR solution. Deep NQSs trained with MinSR achieve highly accurate energies for square-lattice Heisenberg models and provide numerical evidence for a gapless QSL at the maximally frustrated point.
- Minimum-step stochastic reconfiguration: The NQS represents the many-body wavefunction with a neural network whose parameters are optimized to minimize variational energy.In VMC, the network maps spin configurations to wavefunction components, defining the variational state.
- Minimum-step stochastic reconfiguration: MinSR uses a compressed matrix with the same non-zero eigenvalues as the quantum metric, making it equivalent to SR while reducing the matrix size.The compressed formulation is especially useful when the number of parameters greatly exceeds the number of samples.
- Benchmark models: 10^-7 relative variational-energy error is reached for the 10 × 10 non-frustrated Heisenberg model, outperforming existing results.The deep NQS result uses E/N = −0.67155260(3), with EGS/N = −0.67155267(5) as the reference.
- Benchmark models: E/N = −0.4976921(4) is obtained for the 10 × 10 frustrated J1-J2 model, outperforming existing numerical results.The result uses networks including a 64-layer ResNet and a model with more than one million parameters.
- Energy gaps of a QSL: Δ = 0.00(3) in the thermodynamic limit supports a gapless QSL at the maximally frustrated square-lattice J1-J2 point.The finite rescaled gap Δ × L further corroborates a vanishing gap, and the authors describe the result as strong numerical evidence.
Discussion
The work positions deep neural quantum states as another approach to approximating quantum many-body complexity through large-scale neural networks. It identifies extensions to fermionic systems, quantum chemistry, tensor networks, and broader machine-learning settings as future directions.
- Deep neural quantum states provide another approach to approximating quantum many-body complexity through the expressive power of large-scale neural networks.
- MinSR may extend beyond neural quantum states to other variational wavefunctions, including tensor networks, enabling more complex ansätze.The authors describe this as a future application of MinSR as a general optimization method in variational Monte Carlo.
- Future applications include fermionic systems such as the Hubbard model and ab initio quantum chemistry, where traditional methods have limited accuracy in strongly interacting regimes.
- A broader machine-learning application would require a suitable optimization space and an equation analogous to the paper’s variational formulation.Reinforcement learning is proposed as one possible setting for a MinSR-like natural policy gradient.
Methods
The methods derive MinSR as a least-squares minimum-norm alternative to stochastic reconfiguration and stabilize its numerical solution with pseudo-inverses. The approach combines deep residual neural quantum states with physical sign structures, symmetry projections, momentum encoding, and finite-size error estimation.
- MinSR optimization: MinSR selects the minimum-norm parameter update among solutions with minimum residual error, reducing higher-order effects, overfitting, and instability.The derivation assumes an underdetermined system with Ns < Np and uses a least-squares minimum-norm condition.
- MinSR optimization: The MinSR and stochastic-reconfiguration updates are both equivalent to the pseudo-inverse solution when Ns < Np.This establishes MinSR as a natural alternative to SR in the underdetermined regime.
- Numerical stabilization: Numerical MinSR uses diagonalization and a cutoff-based pseudo-inverse, with a soft cutoff to avoid abrupt changes as eigenvalues cross the threshold.The typical relative and absolute cutoffs are rpinv = 10^-12 and apinv = 0.
- Complex neural networks: The complex-network formulation treats complex observables and residuals while retaining real parameters, yielding a corresponding MinSR equation for non-holomorphic networks.The resulting non-holomorphic SR solution agrees with the widely used formulation.
- Neural quantum states: Two ResNet designs represent neural quantum states, with ResNet2 generally providing better accuracy and stability and supporting non-zero momentum for low-lying excited states.ResNet1 performs better for transfer learning from smaller to larger lattices.
- Physical constraints: Model-specific sign structures and symmetry projections are applied to improve accuracy and target suitable quantum-number sectors.The square-lattice Marshall sign rule is exact for the non-frustrated Heisenberg model but approximate near J2/J1 ≈ 0.5; triangular-lattice states use a 120°-order sign structure.
- Error estimation: The extrapolation assumes the error state changes little across training attempts, while an approximately constant (E − Eg)/σ2 ratio is used to estimate large-lattice errors from smaller systems.This empirical procedure is intended to reduce error and computational time.
Data availability
The research does not rely on external datasets, and the reported figures and neural-network weights are publicly available through Zenodo.
- Data availability: The study uses no external datasets, while figures and obtained neural-network weights are available via Zenodo.The passage identifies the repository by DOI URL 10.5281/zenodo.7657551.
Funding
Funding for the research was provided through open-access support from Universität Augsburg.
- Funding: Open-access funding was provided by Universität Augsburg.
Additional information
The supplementary materials compare optimization performance, wavefunction accuracy, extrapolation procedures, and variational energies for square and triangular J1-J2 systems.
- Optimization performance: Extended Data Fig. 1 evaluates optimization time, residual error, and training curves across sample sizes, parameter counts, and optimizers.The measurements use the 10 × 10 square Heisenberg J1-J2 model, ResNet architectures, and A100 GPUs.
- Wavefunction accuracy: Extended Data Fig. 2 compares MinSR-trained neural quantum state amplitudes with exact-diagonalization amplitudes on a 6 × 6 square lattice.The comparison uses a 64-layer ResNet1 with 146320 parameters and reports infidelity in the inset.
- Extrapolation: Zero-variance extrapolation is applied to square-lattice J1-J2 data at J2/J1 = 0.5 using multiple lattice sizes, network sizes, and a Lanczos step.For L = 20, the slope is estimated from the average of slopes measured at L = 10, 12, and 16 because direct fitting is inaccurate.
- Extrapolation: Triangular-lattice extrapolation at J2/J1 = 0.125 uses complex-valued ResNet2 networks with 34944 and 139008 parameters and a Lanczos step.The plotted sectors are S = 0, k = Γ and S = 1, k = K.
- Variational energies: The supplementary energy reference for the square J1-J2 model is −0.497715(9) from zero-variance extrapolation.Extended Data Table 1 concerns variational ground-state energies on a 10 × 10 square lattice with periodic boundary conditions at J2/J1 = 0.5.