Source-linked AI summary
Machine Learning Phases of Strongly Correlated Fermions
Kelvin Ch'ng, Juan Carrasquilla, Roger G. Melko, Ehsan Khatami
TL;DR
The paper asks whether neural networks can classify finite-temperature phases of strongly correlated fermions despite the complications of quantum Monte Carlo data. It trains a 3D CNN on Hubbard-model auxiliary-field configurations, then transfers the half-filled model to doped systems. The network captures the half-filled magnetic phase behavior and predicts that magnetic instability persists to at least 5% doping.
Problem
Classifying phases in quantum many-body systems is challenging because quantum Monte Carlo configurations include imaginary time and quantum fluctuations can obscure ordered-phase patterns.
Method
A 3D convolutional neural network is trained on auxiliary spin configurations generated by determinantal quantum Monte Carlo for the Hubbard model.
Results
The trained network captures the half-filled Néel-temperature trend across interaction strengths and predicts that the magnetic instability remains nonzero away from half filling.
Takeaways & Limitations
Transfer learning from half filling predicts that the magnetic phase extends to at least 5% doping in the studied region.
Abstract
from arXiv · showhide
Machine learning offers an unprecedented perspective for the problem of classifying phases in condensed matter physics. We employ neural-network machine learning techniques to distinguish finite-temperature phases of the strongly correlated fermions on cubic lattices. We show that a three dimensional convolutional network trained on auxiliary field configurations produced by quantum Monte Carlo simulations of the Hubbard model can correctly predict the magnetic phase diagram of the model at the average density of one (half filling). We then use the network, trained at half filling, to explore the trend in the transition temperature as the system is doped away from half filling. This transfer learning approach predicts that the instability to the magnetic phase extends to at least 5% doping in this region. Our results pave the way for other machine learning applications in correlated quantum many-body systems.
I. INTRODUCTION
Machine learning is being applied to classify phases of matter, but quantum systems are harder because imaginary-time structure and quantum fluctuations complicate the input configurations. This work uses a 3D CNN with Hubbard-model quantum Monte Carlo data to classify finite-temperature phases and estimate transition temperatures.
- Neural networks can classify intricate labeled data by adjusting connection parameters and biases during training.
- Machine-learning methods have located critical temperatures in classical Ising-type models, reaching up to 99% accuracy for 2D Ising models.
- Quantum systems are more difficult because finite-temperature quantum Monte Carlo data include an imaginary-time dimension and quantum fluctuations can obscure ordered-phase patterns.
- The network architecture uses four-dimensional auxiliary-field inputs, three spatial dimensions of size 4, and an imaginary-time dimension of size L = 200.The hidden feature-map counts are n(2) = 32, n(3) = 16, and n(4) = 8, with dropout regularization at rate 0.5.
- A 3D CNN applied to Hubbard-model auxiliary configurations can classify finite-temperature phases of strongly correlated fermions and estimate transition temperatures on relatively small lattices.
II. MODEL
The paper studies the particle-hole-invariant Fermi-Hubbard model on three-dimensional cubic lattices. At half filling, it undergoes an antiferromagnetic Néel transition whose temperature varies nonmonotonically with interaction strength.
- The Fermi-Hubbard Hamiltonian contains nearest-neighbor hopping, onsite Coulomb interaction U, and chemical potential µ.
- The chemical potential µ = 0 corresponds to half filling, defined here as average density n = 1.
- The model is studied on three-dimensional cubic lattices with hopping t = 1 as the energy unit.
- At half filling, the 3D model has a finite-temperature transition to an antiferromagnetic Néel phase for every U > 0.
- The Néel temperature rises rapidly at weak coupling but decreases at large U, where the model is effectively described by an antiferromagnetic Heisenberg model with exchange proportional to 1/U.
III. METHOD
The method trains convolutional neural networks on auxiliary fields generated by determinantal quantum Monte Carlo and uses them to map magnetic phase boundaries. Spatial locality is processed by convolutions, while imaginary-time slices serve as filter channels.
- Determinantal quantum Monte Carlo reduces Hubbard-model observables to stochastic averages over discrete auxiliary fields in space and imaginary time.
- Training uses configurations sampled at temperatures around one or two critical points, after which the network is tested across parameters to map the associated phase boundary.
- The 3D CNN applies convolutions to the cubic lattice’s three spatial dimensions and treats imaginary-time slices as separate filter channels.
- The architecture uses three or four hidden feature-extraction layers followed by a fully connected layer and an output layer.The number of neurons is optimized for larger systems using Monte Carlo optimization.
IV. RESULTS
The neural network predicts the half-filled Hubbard model’s magnetic transition across interaction strengths and is then transferred to doped systems. Results show that joint weak- and strong-coupling training is important, while sign-problem treatment materially affects the inferred phase boundary.
- Half filling: Approximately 80 000 DQMC-labeled configurations train CNNs to distinguish ordered from high-temperature spin configurations at U = 5 and 16.Training uses configurations sampled around the Néel transition for finite cubic systems.
- Half filling: Six independently initialized CNNs provide averaged transition-temperature estimates, with larger error bars for N = 83 because that network has more parameters.The error bars are standard errors across the six classifications.
- Half filling: Transition temperatures at other interaction strengths are estimated where the ordered-phase neuron output crosses 0.5, using polynomial fits when classifications are noisy.The classifications use 90 000 configurations for N = 43 and 10 000 for N = 83.
- Half filling: Joint training with U = 5 and U = 16 is crucial: single-regime training overestimates intermediate- and strong-coupling TN or underestimates TN below U = 16.The combined data encode competing weak- and strong-coupling signatures, improving estimates for other U.
- Doping away from half filling: The transferred U = 9 network finds a nonzero transition temperature away from half filling, with the instability expected to decrease rapidly as density decreases.The phase diagram uses the neuron-output 0.5 crossing to estimate critical chemical potentials and Néel temperatures.
- Doping away from half filling: Ignoring the DQMC sign changes the inferred boundary: the critical chemical potential is −2.0 with sign treatment versus −2.6 without it.Sign omission also produces a smaller average density at fixed chemical potential.
- Doping away from half filling: At T = 0.32 and critical µ = −2, the density is less than 0.95, consistent with conventional results.The density begins deviating from unity around µ = −1.0, where the Mott gap is not yet fully developed at accessible temperatures.
V. CONCLUSION
Neural-network methods predict finite-temperature magnetic order in the three-dimensional Fermi-Hubbard model. Training at half filling captures the interaction dependence of the Néel temperature and indicates that magnetic instability persists near commensurate filling, while sign treatment affects the results.
- V. CONCLUSION: A 3D CNN trained on DQMC auxiliary spin configurations captures the half-filled model’s Néel-temperature trend across interaction strengths.Training samples systems in weak- and strong-coupling regimes and classifies configurations at other interactions.
- V. CONCLUSION: The half-filled U = 9 network predicts that magnetic instability persists when the system is doped near commensurate filling.The conclusion reports persistence in the region close to density one.
- V. CONCLUSION: Excluding the sign from neuron-output expectation values produces different results in the doped system.The sign treatment is therefore part of the reported conclusion about the transferred classification.
APPENDIX A: DETERMINANT QUANTUM MONTE CARLO
Determinantal quantum Monte Carlo rewrites the Hubbard model using auxiliary fields, enabling Monte Carlo estimation of observables while introducing determinant-sign constraints and a direct link to fermionic spin correlations.
- DQMC discretizes inverse temperature into imaginary-time slices and separates kinetic and interaction terms before applying the Hubbard–Stratonovich transformation.The transformation introduces auxiliary spin variables in a D+1-dimensional space, with D spatial dimensions and one imaginary-time dimension.
- Integrating out the fermionic degrees of freedom leaves a sum over auxiliary-spin configurations weighted by products of two fermion-species determinants.The matrix Mσ depends on the auxiliary-spin configuration and its spatial and imaginary-time indices.
- Observables are estimated by importance sampling, using the determinant product to accept or reject local auxiliary-field updates.The simulations use the Metropolis algorithm as an example of this sampling procedure.
- At half filling, symmetry gives the two determinants the same sign, whereas away from half filling the sign problem must be accounted for in observable estimates.The expectation value divides the sign-weighted average by the average sign, and jackknife resampling provides error bars.
- The chosen decoupling scheme makes auxiliary-field correlations suitable proxies for fermionic spin correlations, while other decouplings could target charge or pairing phases.The paper does not explore those alternative phase-detection possibilities.
1. Case of N = 43
For N = 43, the paper constructs volumetric feature maps from auxiliary-field inputs and feeds progressively extracted features through a regularized classifier with a decaying learning rate.
- CNN architecture: Each imaginary-time slice is encoded as a separate filter channel in a 3D CNN processing the cubic-lattice spatial dimensions.The input blocks are connected to hidden volumetric feature maps for feature extraction.
- CNN architecture: A shared 2 × 2 × 2 kernel sweeps each input block, adds a shared bias, and applies ReLU to construct volumetric feature maps.The kernel labels the imaginary-time input block and the hidden feature-map block; stride one completes the map.
- CNN architecture: Zero padding preserves the spatial volume during convolution, including by expanding input cubes to 5 × 5 × 5 so outputs remain 4 × 4 × 4.This limits depletion of border information in deeper layers.
- CNN architecture: Two additional ReLU convolutional layers use 16 and 8 volumetric feature maps, followed by an eight-neuron classification layer and two softmax outputs.The final feature-extraction neurons are fully connected to the classification layer.
- Training: Dropout deactivates half of the eight fully connected neurons during training to mitigate overfitting and encourage robust features.The dropout rate is 0.5.
- Training: The network uses cross-entropy loss and an exponentially decaying learning rate, with initial learning rate η0 = 10^-3 and decay rate λ = 0.925.The learning rate begins rapidly and slows as the model approaches convergence.
2. Case of N = 83
For N = 83, the study modifies the convolution and architecture choices for computational manageability and uses Monte Carlo hyperparameter sampling rather than hand selection.
- For N = 83, convolutions omit zero padding to keep optimization computationally manageable, and the network uses four hidden feature-extraction layers.
- The learning-rate decay is increased to λ = 0.875 to reduce the risk of overshooting the optimal model.
- 20 Monte Carlo search iterations over feature-map and classification-neuron counts replace handpicked hyperparameters, which yielded no learning.The search constrains both quantities to [8,128] and selects n(2) = 54, n(3) = 26, n(4) = 14, n(5) = 18, and n(6) = 64.
APPENDIX C: TRAINING THE CNN
The CNN training procedure monitors training and testing accuracy to detect overfitting, saves improving models under explicit criteria, and reports testing accuracies for alternative training setups.
- Training procedure: 85% of labeled configurations are used for training and 15% for testing, with batches of 200 processed during each complete training epoch.Testing accuracy gauges overfitting when training accuracy continues improving but testing accuracy stalls.
- Training procedure: Training stops when overfitting is detected from the divergence between improving training accuracy and stalled testing accuracy.
- Model selection: Saved models require a training–testing accuracy difference below 2.5%, training accuracy above testing accuracy, and improved current testing accuracy.The last saved locations are marked by green circles in Fig. 6.
- Results: 92.7% and 93.2% testing accuracy are reported for N = 43 and N = 83 under simultaneous U = 5 and U = 16 training.The corresponding single-interaction results are 91.9% for U = 5 and 94.3% for U = 16.
- Results: Figure 6 compares training accuracy versus epochs for simultaneous U = 5 and 16 training, U = 5 alone, and U = 16 alone.
APPENDIX D: WHAT DOES THE MACHINE LEARN?
The CNN’s learned local features can be analyzed as overlaps with the 256 binary ordering patterns on a 2×2×2 receptive cube. The dominant features correspond to patterns with the largest overlap, with the perfect AFM pattern providing a reference.
- 256 binary ordering patterns exhaust the possible 1 and -1 configurations on the eight-site 2×2×2 receptive cube.The input configurations are binary, so the local feature space contains 2^8 = 256 distinct patterns.
- The analysis focuses on the first feature-extraction layer, whose weights indicate the local patterns sought by the CNN.Each ordering pattern is convolved with the learned weights to measure its correlation with the extracted features.
- The overlap function compares each binary ordering pattern with the learned weights while ignoring biases and taking absolute values to account for spin inversion symmetry.Biases contribute only a uniform shift and do not encode local correlations.
- Patterns with the largest γx values are treated as the CNN’s dominant features, with γx playing a role analogous to an order parameter for each pattern.For the perfect AFM pattern, the corresponding overlap is analogous to a staggered magnetization of the learned weights.
- Figure 8 compares overlap histograms for U = 5 and U = 16, marking the perfect AFM pattern’s overlap with a vertical dashed red line.The distributions show which local ordering patterns occupy the large-γx tail and are therefore dominant.