Source-linked AI summary
Large Electron Model: A Universal Ground State Predictor
Timothy Zaklama, Max Geier, Liang Fu
TL;DR
The paper addresses the challenge of predicting interacting-electron wavefunctions across physical settings without system-specific labeled solutions. It introduces a single unsupervised, parameter-conditioned universal fermionic network optimized variationally, and demonstrates accurate generalization across unseen interaction strengths and particle numbers, including systems with up to N = 50 electrons.
Problem
Reliable many-electron prediction requires methods beyond DFT, which is not variational, does not produce many-body ground-state wavefunctions, and fails in strongly correlated materials.
Method
LEM is a single parameter-sharing network built from a universal fermionic wavefunction representation and conditioned on Hamiltonian parameters and particle number.
Results
The model accurately predicts ground-state wavefunctions, energies, and charge densities across unseen interaction strengths and particle-number sectors, with energy predictions extending to N = 50 electrons.
Takeaways & Limitations
The results support a reusable variational foundation-model approach for interacting electrons in the continuum and material-discovery calculations.
Takeaways & Limitations
The compact single-determinant Slater–Jastrow baseline remains limited by its restricted determinant backbone and cannot represent the full multiconfigurational mixing needed in strongly correlated regimes.
Abstract
from arXiv · showhide
We introduce Large Electron Model, a single neural network model that produces variational wavefunctions of interacting electrons over the entire Hamiltonian parameter manifold. Our model employs the Fermi Sets architecture, a universal representation of many-body fermionic wavefunctions, which is further conditioned on Hamiltonian parameter and particle number. For interacting electrons in a two-dimensional harmonic potential, a single trained model accurately predicts the ground state wavefunction while generalizing across unseen coupling strengths and particle-number sectors, producing both accurate real-space charge densities and ground state energies, even up to $50$ particles. Our results establish a foundation model method for material discovery that is grounded in the variational principle, while accurately treating strong electron correlation beyond the capacity of density functional theory.
INTRODUCTION
Large Electron Model addresses the need for reusable, variational prediction of interacting-electron wavefunctions across Hamiltonian settings. It conditions a universal fermionic neural network on physical parameters and learns the mapping through shared, unsupervised energy minimization.
- Motivation: Accurate prediction of interacting-electron properties requires solving many-electron systems, while DFT is not variational, does not produce many-body ground-state wavefunctions, and fails in strongly correlated materials.The paper motivates a wavefunction-level approach beyond system-specific electronic-structure calculations.
- Motivation: Existing neural-network variational Monte Carlo methods typically optimize a new network from scratch for each Hamiltonian and parameter choice, preventing computation reuse across systems or sizes.The stated gap is the absence of a shared foundation-model solver across a physical-parameter manifold.
- Approach: Large Electron Model is a single parameter-sharing network conditioned on Hamiltonian parameters, such as interaction strength and particle number, that predicts ground-state wavefunctions for unseen Hamiltonians.The model is designed to represent the manifold of wavefunctions with shared parameters.
- Training: The model is trained fully unsupervised by variational energy minimization, using Monte Carlo estimates of energy expectations over a fixed ensemble of Hamiltonian parameters.Gradients of the accumulated energy update the shared network weights.
- Approach: LEM uses a universal Fermi-wavefunction ansatz that combines an antisymmetric Slater-determinant component with a symmetric attention-based correlation component.This architecture targets fermionic sign structures and electron correlations within one representation.
- Results: On interacting electrons in a two-dimensional harmonic trap, one trained model predicts accurate wavefunctions, energies, and charge densities across unseen interaction strengths and particle-number sectors, scaling to N = 50.The reported results establish wavefunction-level generalization across changing Hilbert-space sectors and interaction regimes.
APPROACH
The approach uses a single conditional neural wavefunction to predict ground states across continuum Hamiltonians, extending a universal Fermi Sets ansatz with parameter dependence. Its architecture combines antisymmetric Slater determinants with symmetric, attention-based correlation modeling and permutation-invariant outputs.
- Conditional foundation model: LEM maps Hamiltonian parameters to ground-state wavefunctions for continuum interacting-fermion systems using one shared conditional network.The model is optimized once and reused across a manifold of physical settings.
- Fermi Sets ansatz: Fermi Sets factorizes fermionic wavefunctions into symmetric functions and antisymmetric cores, with parameter dependence added for foundation-model prediction.The factorization preserves fermionic antisymmetry while allowing configuration-dependent symmetric prefactors.
- Symmetric component: Geometric particle features are combined with spin, embedded, processed by a transformer, and conditioned on the Hamiltonian parameter space.Reference centers define the geometric features used to represent particles.
- Symmetric component: Transformer outputs are pooled over particles to produce permutation-invariant representations, which an MLP converts into symmetric coefficients.Permutation-equivariant weights can increase expressivity.
- Antisymmetric component: The antisymmetric factor is implemented with Slater determinants built from learned single-particle orbitals.Antisymmetry is carried by the determinant component of the ansatz.
PHYSICAL SYSTEM
The study evaluates the foundation-model framework on spin-polarized interacting electrons confined in a two-dimensional isotropic harmonic potential with Coulomb repulsion. The Hamiltonian is parameterized by particle number, confinement frequency, and interaction strength, and energies are evaluated through variational Monte Carlo.
- Physical system: The physical system contains N spin-polarized interacting electrons in d = 2, confined by an isotropic harmonic potential and coupled through Coulomb repulsion.This is the quantum-dot problem used to demonstrate the framework.
- Hamiltonian parameters: The Hamiltonian parameter vector is Λ = (N, ω, λ), where ω controls confinement curvature and λ controls interaction strength.The calculations set ω = 1.
- Hamiltonian parameters: For ω ≠ 1, solutions at ω = 1 are rescaled using the dimensionless coupling λ/ω^1/4 and overall energy scale √ω.This provides the stated mapping from the fixed-frequency calculations to other confinement frequencies.
- Evaluation: The variational Monte Carlo local energy is constructed from the predicted wavefunction and enters the energy-minimization objective.The wavefunction is written as Ψθ(R, s; Λ).
- Evaluation: LEM optimizes one universal Fermi network across a target Hamiltonian manifold and predicts ground-state wavefunctions throughout the family.The workflow targets the family {H(Λ)} over the chosen parameter set.
RESULTS
LEM predicts ground-state wavefunctions, energies, and observables across unseen interaction strengths and particle-number sectors from one optimized model. It reproduces symmetry and correlation features, including zero-shot angular structures, and remains accurate up to N = 50.
- Ground-state prediction: A single model predicts ground-state wavefunctions across 0 ≤ λ ≤ 10.0 and every integer 5 ≤ N ≤ 11 after optimization on only selected λ and N values.The model is evaluated against high-quality variational wavefunctions, DMC, and FermiNet-based NN-VMC benchmarks.
- Benchmarking: From a single multi-system optimization, LEM beats DMC energies for all tested cases while providing an explicit wavefunction and direct access to equal-time observables.Individually optimized Fermi Sets reach even lower energies, but LEM requires no problem-specific optimization or supervision.
- Wavefunction structure: The predicted densities reproduce closed-shell rotational symmetry, interaction-driven radial restructuring, and correct angular-momentum patterns in open-shell systems.For N = 9, the network zero-shot predicts the twofold structure despite never being optimized for that angular structure.
- Ground-state prediction: The model produces accurate out-of-sample energies for unseen interaction strengths and particle numbers across the (λ, N) grid.The same learned parameters generalize across distinct Hilbert-space sectors.
- Large-system generalization: At N = 50, the model accurately predicts unseen coupling strengths without reoptimization, and most comparisons for 6.0 < λ < 8.0 outperform fresh single-system runs.Multi-system optimization is described as an inductive bias that can stabilize optimization and improve generalization.
- Scope: The parameter-conditioned Fermi Sets architecture is presented as applicable beyond quantum dots to general electron systems specified by a Hamiltonian family and conditioning parameters.The authors identify solids with electron–nuclei interactions as a possible extension.
CONCLUSION
The paper introduces a parameter-sharing foundation model that predicts many-electron ground states across systems and Hamiltonian parameters. It generalizes across unseen interaction strengths and particle numbers while producing accurate wavefunctions, variational energies, and charge densities.
- LEM maps system and Hamiltonian parameters directly to accurate many-body ground states through a single parameter-sharing neural network.
- The model generalizes at the wavefunction level to unseen interaction strengths and completely unseen particle numbers, including sectors with changing Hilbert-space dimensions.
- LEM predicts ground-state energies for up to N = 50 electrons and out-of-sample energies that outperform fresh single-system runs.
- Because it is fully unsupervised and requires only the Hamiltonian, the method offers a unified route to reusable continuum-electron foundation models.
Supplementary Material for Large Electron Model: A Universal Ground State
The supplementary material identifies the authors and their institutional affiliation.
- The paper is authored by Timothy Zaklama, Max Geier, and Liang Fu.
- All authors are affiliated with the Department of Physics at the Massachusetts Institute of Technology in Cambridge, Massachusetts.
Table of benchmarked energies
The benchmark table compares ground-state energies from the foundation model with neural Slater–Jastrow and Hartree–Fock approaches at N = 10. The accompanying discussion reports that the foundation model outperforms these standard variational benchmarks using one training run.
- Table II compares ground-state energies at N = 10 across NN–SJ backflow, NN–SJ, Hartree–Fock, and the Fermi Sets foundation model.
- The baseline energies come from calculations trained from scratch on single systems, using the same network size where applicable.
- The table’s training column records whether the Fermi Sets network encountered each parameter value during training.
- A single foundation-model training run produces energies for seen and unseen systems that outperform the standard variational Monte Carlo techniques in the comparison.
Neural Network Slater–Jastrow and Hartree–Fock comparison
The comparison shows how density patterns differ between Hartree–Fock and neural Slater–Jastrow baselines as interaction strength increases. Hartree–Fock develops broken-symmetry localization, while correlation-aware ansätze remain more physically structured but retain determinant-based limitations.
- Symmetry interpretation: For circular-dot eigenstates in a single angular-momentum sector, the one-body density is rotationally symmetric; angular modulation signals sector mixing or symmetry breaking.
- Density comparison: At λ = 1.0, 3.0, and 7.0, Figure S1 compares real-space densities from NN Slater–Jastrow and Hartree–Fock calculations for N = 10.
- Hartree–Fock limitation: At λ = 7.0, Hartree–Fock produces a strongly non-uniform, irregular density because its mean-field orbitals localize and mix angular-momentum sectors.
- Slater–Jastrow structure: The neural Jastrow reshapes amplitudes and captures higher-order correlations, but its positive factor leaves the determinant’s antisymmetric sign structure and nodes unchanged.
- Ansatz limitations: Backflow can improve nodal structure, yet the Slater–Jastrow family remains limited by its restricted determinant backbone and cannot capture essential full multiconfigurational mixing in strong correlation.
- Foundation-model comparison: The Fermi Sets model mixes antisymmetric components through a symmetric, configuration- and parameter-dependent mechanism, producing densities that track ground-state structure across interaction strengths and particle-number sectors.
Weak coupling density and angular momentum for the open-shell case
The open-shell N = 7, λ = 1.0 case exposes sensitivity to training coverage near weak-coupling angular-momentum degeneracies. Adding a weak-coupling training point corrects the inferred sector and yields the expected L = ±3 character.
- Physical setting: The N = 7 spin-polarized dot tests inference where the seventh electron occupies degenerate orbitals in an open shell.The first six electrons form a closed shell with L = 0, while the seventh enters the next harmonic-oscillator shell.
- Physical setting: The weakly interacting ground state is expected to select the |L| = 3 sector, while interactions lift the |L| = 3 and |L| = 1 degeneracy.The selection follows the modified Hund second rule for circular quantum dots.
- Training-distribution diagnostic: At unseen N = 7 and λ = 1.0, the original training distribution produced a twofold density pattern associated with the wrong open-shell sector.The original model was trained on N ∈ {6, 8, 10} and λ ∈ {0.0, 1.0, 2.0, 5.0, 8.0, 10.0}.
- Training-distribution diagnostic: Adding one weak-coupling training point corrected the angular-momentum sector and produced the correct L = ±3 character at the unseen N = 7 system.The comparison identifies the earlier result as an inference error caused by insufficiently resolved weak-coupling training data, not a physical transition or ansatz limitation.
- Angular-momentum validation: Sector projections show dominant paired L = +3 and L = −3 weights for N = 7, paired L = +1 and L = −1 weights for N = 9, and L = 0 for N = 10.These projections provide a wavefunction-level diagnostic beyond visual inspection of densities.
- Model diagnostic: The model learns a parameter- and particle-number-dependent map to variational wavefunctions, but accuracy depends on resolving the physical regimes represented in training.Weak-coupling open-shell states are demanding because small energetic splittings determine the angular-momentum sector; local energy variance can quantify uncertainty because exact eigenstates have zero variance.
- Model diagnostic: Because the output is an explicit many-electron wavefunction, it supports equal-time observables beyond energies and densities, including pair correlations and structure factors.The same sector-weight procedure can also be applied to conditional densities and other probes of correlated electronic structure.
Optimization stability across system size
Co-training is stable for the small-particle-number model but noisier for large-N tasks with widely separated energy scales. Loss reweighting, task-scale normalization, regularization, or increased capacity can reduce this optimization imbalance.
- Observed stability: The small-N co-trained model converges markedly more stably than the large-N model despite training on more systems.Figure S4 compares 18 small-N parameter groups with the N = 50 co-training run using mean variational energy and energy variance.
- Observed stability: Per-parameter curves show smooth convergence with mild transient noise for small-N co-training, whereas large-N co-training has stronger fluctuations at lower λ.The noise ordering differs from single-task training.
- Optimization mechanism: The larger low-λ fluctuations arise from simultaneous optimization across tasks with widely separated energy scales, not from greater physical difficulty at weak coupling.Averaging the objective lets higher-energy groups set a larger effective shared-parameter step scale.
- Mitigation: Loss balancing can reduce the imbalance by reweighting groups or normalizing them by energy or variance scales.These strategies aim to make task gradients comparable across parameter groups.
- Mitigation: Additional mitigation options include weak variance regularization for low-λ groups or increased model capacity when resources allow.The stated goals are suppressing excess variance and reducing competition between tasks.
Extended data
Extended visualizations expose the model’s full many-body wavefunction through conditional probability, phase, and pair correlations. They reveal nodal and exchange-correlation structure that one-body densities alone cannot provide.
- Wavefunction visualization: Conditional wavefunction slices fix N − 1 electron coordinates at a representative high-probability configuration and scan the remaining coordinate across the plane.This directly visualizes Ψθ rather than only a reduced derivative quantity.
- Wavefunction visualization: The visualization separates conditional probability density, phase, and pair correlation into left, middle, and right columns.Figure S6 compares these quantities for N = 10 and N = 7 at λ = 8.0.
- Phase structure: Phase-domain boundaries coincide with suppressed amplitude and identify nodal contours, exposing sign and phase structure inaccessible from one-body densities alone.The nodal domains demonstrate that the model captures fermionic antisymmetry.
- Pair correlations: For the N = 10, λ = 8 full-shell case, pair correlations show a short-distance exchange-correlation hole and an approximately isotropic correlation ring.This is consistent with the rotationally symmetric density expected for a filled shell.