Source-linked AI summary
Phases in a class of associative memories via hidden neurons
Toshihiro Ota, Masato Taki
TL;DR
The paper seeks a common architecture for understanding what fixes storage scale across polynomial and exponential associative-memory regimes. It studies the bipartite class H and finds distinct crosstalk statistics and phase structures, with visible and hidden Lagrangians governing stability and storage scale respectively.
Problem
Polynomial and exponential storage regimes use different analytical tools, leaving no common architecture for studying what fixes storage scale and how crosstalk statistics change.
Method
The paper analyzes the bipartite class H using replica methods at polynomial load and a copy representation that maps exponential-load thermodynamics onto counting over pattern assignments.
Results
Polynomial load yields replica-symmetric phase diagrams and capacities, while exponential load exhibits paramagnetic, condensed, and frozen phases with quantized attention reassignments destabilizing retrieval under heating.
Takeaways & Limitations
Retrieval separates into visible-layer stability and hidden-layer storage-scale roles, while crosstalk is central-limit-like at polynomial load and large-deviation-based at exponential load.
Takeaways & Limitations
Models A and C rely on the replica-symmetric ansatz, and Models A, C, and B use an adiabatic hidden-sector limit that excludes finite β/τh phase diagrams.
Abstract
from arXiv · showhide
Associative memory in the Hopfield network is attractor dynamics in a disordered many-body system, and higher-order and exponential extensions turn its retrieval update into softmax attention. The polynomial and exponential regimes have been analyzed by different methods, with no common architecture in which to ask what fixes the storage scale. In this paper we study the bipartite architecture of Krotov and Hopfield, which we call the class $H$, whose model is fixed by a Lagrangian for each layer, taking the hidden neurons as the order parameter of retrieval. At polynomial load the replica method yields the replica-symmetric phase diagrams and closed-form capacities, and the crosstalk moment is common to Ising and spherical visible neurons, so their differences come from the visible entropy. With a softmax hidden layer the load is exponential, and a copy representation maps the thermodynamics onto random-energy-model counting, with paramagnetic, condensed, and frozen phases. Heating destabilizes retrieval by quantized reassignments of attention, and typical Gaussian patterns remain metastable at every load. The regimes differ in their crosstalk statistics, central-limit at polynomial load and large-deviation at exponential load, and the class $H$ splits retrieval into two roles, the visible Lagrangian fixing stability and the hidden one the storage scale, two axes that may also guide the design of new Lagrangians.
I. INTRODUCTION
The paper introduces class H as a common bipartite architecture for comparing polynomial and exponential associative-memory regimes. Its results connect crosstalk statistics and storage scale to separate visible and hidden Lagrangians.
- Motivation: Class H provides one architecture for asking how the storage scale and crosstalk statistics change across polynomial and exponential regimes.It uses visible and hidden layers coupled by pairwise interactions, with each layer specified by a Lagrangian.
- Polynomial load: Replica analysis gives phase diagrams and closed-form zero-temperature capacities for Models A and C at polynomial load.The crosstalk moment is shared by Ising and spherical visible neurons, while their retrieval phases differ through visible entropy.
- Exponential load: With a softmax hidden layer, exponential load maps through a copy representation to random-energy-model counting with paramagnetic, condensed, and frozen phases.Heating destabilizes retrieval through quantized attention reassignments rather than smooth overlap erosion.
- Crosstalk regimes: Polynomial load has central-limit crosstalk largely insensitive to pattern ensembles, whereas exponential load has ensemble-dependent large-deviation crosstalk.Rare large-norm Gaussian patterns can lie below typical retrieval in free energy, leaving typical memories metastable rather than equilibrated.
- Two Lagrangian roles: The visible Lagrangian fixes retrieval stability through visible entropy, while the hidden Lagrangian fixes storage scale and disorder statistics.The framework places Hopfield, dense associative memory, and attention on two common design axes.
II. PRELIMINARIES
This section introduces class H as the framework used for the paper’s statistical-mechanical analysis. Technical details and the broader class-Hk construction are deferred to the appendices.
- Class H is overviewed as the basis for the paper’s subsequent statistical-mechanical analysis, with broader class-Hk details provided in Appendix A.
A. Overview of the class H
Class H consists of visible and hidden neurons in a bipartite network, with each layer governed by a Lagrangian. Hidden neurons encode retrieval through their responses to visible-pattern overlaps.
- Class H uses visible and hidden neurons with symmetric cross-layer interactions and no intralayer connections.The layer sizes are Nv and Nh, and the interaction matrices are transposes of one another.
- The visible and hidden Lagrangians determine the neurons’ activation functions and define the associated energy function.The energy combines the Legendre transforms of the two Lagrangians with the cross-layer coupling.
- Positive-semidefinite Lagrangian Hessians make the energy decrease along trajectories, and a lower-bounded energy guarantees convergence to fixed-point attractors.These attractors are identified with stored memories, so convergence corresponds to retrieval.
- In the adiabatic limit, hidden neurons instantaneously follow the visible configuration and measure overlaps between visible activations and stored patterns.The hidden states therefore act as feature detectors and as the retrieval order parameter.
B. Partition function
The partition function provides the statistical-mechanical formulation of class H, but its precise definition depends on the chosen Lagrangians and integration measure. In the adiabatic limit, hidden fluctuations are suppressed and the hidden integral localizes at a saddle point.
- The class-H partition function is formal because some Lagrangians create flat energy directions and divergent integrals.For degree-one homogeneous visible Lagrangians, the energy can be independent of the visible scale, so the domain and measure must be specified model by model.
- The adiabatic limit β/τh →∞ suppresses hidden thermal fluctuations and permits saddle-point evaluation of the hidden integral.The integral localizes at the stationary hidden configuration, yielding the reduced partition function.
- The leading adiabatic expression suppresses a Gaussian prefactor involving the hidden-Lagrangian Hessian, whose role is examined separately for Model B.
- The hidden neurons are treated as the retrieval order parameter in the model analyses that follow.
III. MODELS A AND C
Models A and C form a class-H family specified by visible ℓp Lagrangians and a common higher-order hidden Lagrangian. The replica analysis integrates out visible neurons and uses hidden neurons as retrieval order parameters.
- III. MODELS A AND C: The energy depends on visible neurons through their activation g because the bare visible-neuron contribution vanishes.This makes the activation, rather than the overall scale of v, the relevant visible degree of freedom.
- III. MODELS A AND C: The visible partition function is restricted to a sphere because the energy is independent of the overall scale of v.For p = 2 the sphere has radius √Nv, while p = 1 reduces to a sum over Ising configurations.
- III. MODELS A AND C: Models A and C correspond to p = 1 and p = 2, respectively, within the ℓp visible-neuron family.The visible activation is normalized on a sphere, with p = 1 giving Ising spins and p = 2 giving spherical neurons.
- III. MODELS A AND C: For positive even k, the replica method first integrates out visible neurons, then averages over random patterns, focusing on replica-symmetric solutions.The hidden neurons serve as the order parameters for memory retrieval.
A. Model A
Model A is analyzed with replica-symmetric equations for polynomial-load retrieval, yielding closed-form capacities and a phase diagram. Its crosstalk is central-limit-like for k > 2, but higher moments cause factorial capacity loss and increasingly narrow retrieval regions as k grows.
- A. Model A: For k > 2, non-condensed overlaps behave as bare Gaussian variables because their Onsager feedback vanishes in the thermodynamic limit.At k = 2, the feedback is marginal and must be resummed, producing the denominator 1 − βk(1 − q).
- A. Model A: The crosstalk moment is the 2(k − 1)-th Gaussian moment, whose factorial growth drives the collapse of capacity as k increases.The k-body signal sharpens retrieval through mk−1, while rare large crosstalk fluctuations increasingly penalize additional stored patterns.
- A. Model A: αc ≃0.138 for k = 2, αc ≃1.32 × 10−2 for k = 4, and αc ≃1.67 × 10−4 for k = 6.For k = 2, the critical overlap is mc ≃0.97; for k > 2, the capacity has a closed form based on the Gaussian moment (2k − 3)!!.
- A. Model A: At zero load, critical temperature decreases only slowly with k, while zero-temperature capacity collapses factorially, making the retrieval boundary nearly vertical for large k.The resulting asymmetry leaves Model A robust to thermal noise at αk = 0 but fragile to loading at αk > 0.
- A. Model A: Retrieval is globally stable below TM(αk), metastable between TM(αk) and TR(αk), and absent above the retrieval spinodal.For k > 2, the first-order boundary is reentrant because heating melts the m = 0 background while retrieval remains effectively frozen.
- A. Model A: For k > 2, the first-order boundary exceeds its zero-temperature load by about 20% for k = 4 and 54% for k = 6.This reentrance reflects the different entropic responses of the spin-glass background and retrieval state.
B. Model C
Model C uses spherical visible neurons within the class H architecture, preserving the crosstalk statistics of Model A while changing retrieval through visible entropy. Its retrieval capacity collapses at k = 2, remains lower than Model A for k > 2, and lacks reentrant phase boundaries.
- Model definition: Model C replaces Ising visible neurons with continuous variables constrained to a sphere, while retaining the same hidden-layer architecture and crosstalk term.The visible trace is a Gaussian integral on the sphere, and the pattern ensemble is rotationally invariant.
- Crosstalk and entropy: The crosstalk moment is universal across Models A and C, so their phase differences arise from the visible entropy.The spherical visible state competes through the entropy of the sphere rather than changing the non-condensed overlap statistics.
- Capacity: For k > 2, Model C has closed-form capacities, including αc = 4/405 ≃9.88 × 10−3 for k = 4 and αc ≃8.67 × 10−5 for k = 6.The corresponding critical overlaps are lower than Model A’s, reflecting a retrieval-quality trade-off from soft spherical spins.
- Phase diagram: The k > 2 phase diagrams retain Model A’s topology but have uniformly smaller retrieval regions and no reentrant boundaries.Visible confinement entropy suppresses reentrance by causing the spherical retrieval state to melt with the m = 0 background at low temperature.
- Capacity: For k = 2, Model C has zero capacity: retrieval survives only at αk = 0 and T ≤1.The spherical quadratic model is marginal, with αc = 0 and no defined steepness in Table II.
IV. MODEL B
Model B is the softmax-hidden member of class H, where retrieval is represented by attention weights over stored patterns. Its thermodynamics reduces to a competition between attention entropy and Gram-matrix energy, requiring exponential load for a genuine balance.
- Model definition: Model B uses linear visible activation and softmax hidden activation, with the hidden weights serving as the retrieval order parameter.Retrieving pattern µ corresponds to concentrating the attention vector f at the simplex vertex eµ.
- Parameters: The attention sharpness λ controls the effective interaction scale, while β remains the genuine inverse temperature.With τv = 1 and τh = λ, the effective energy matches the modern Hopfield-network form.
- Effective thermodynamics: After integrating out the quadratic visible sector, Model B’s thermodynamics depends on attention weights through Gram-matrix energy and attention entropy.The positive-semidefinite term f^T Gf favors concentration, whereas entropy favors delocalized attention.
- Storage scale: The natural load regime is exponential because energy is extensive in the number of visible neurons, while attention entropy scales as log Nh.A genuine competition requires log Nh ∼ Nv.
- Method: The disorder analysis uses large deviations of pattern norms and overlaps, so Model B is treated by direct counting rather than the replica method.The relevant random-matrix object is the Gram matrix of Gaussian patterns.
A. Capacity at zero temperature
At zero temperature, Model B retrieval is governed by vertex condensation and extreme-value statistics over exponentially many patterns. Capacity is set by whether attention leakage from competing patterns vanishes, with distinct soft-attention and frozen-competitor failure regimes.
- Phase structure: At exponential load, the Gibbs measure spreads over exponentially many nearly pure vertex states rather than interior attention distributions.The resulting counting entropy provides the competition that determines typical-pattern retrieval capacity.
- Retrieval criterion: Zero-temperature retrieval is a fully concentrated attention state, and its capacity is determined by whether the attention leak vanishes.The retrieval state lies near the retrieved pattern, while the leak is controlled by the largest competing pattern contributions.
- Failure regimes: For λ < 1, retrieval fails through aggregate crosstalk from exponentially many weakly correlated patterns, whereas for λ > 1 it is limited by the single most aligned competitor.The two regimes correspond respectively to an interior large-deviation saddle and a frozen boundary maximum.
- Capacity: The two capacity branches match continuously at λ = 1, while no attention sharpness overcomes a competitor as aligned as the signal itself.The Gaussian-pattern capacity therefore has a ceiling in the frozen regime.
- Pattern ensemble: The capacity ceiling is ensemble-dependent: it disappears for spherical patterns, whose capacity continues to grow logarithmically with λ.This sensitivity distinguishes exponential-load behavior from the polynomial regime.
- Equilibrium status: Typical Gaussian retrieval is metastable at every exponential load because atypically large-norm patterns have lower energy and become the equilibrium states.The typical-pattern capacity is therefore a spinodal rather than an equilibrium threshold.
1. Copy representation
The copy representation rewrites the softmax-hidden thermodynamics as a discrete counting problem over pattern selections. Its multinomial entropy makes the measure explicit and exposes retrieval as alignment of all copies on one pattern.
- Copy construction: At integer n = β/λ, exponentiating the logarithm produces n copies, each selecting one stored pattern and interacting through their Gram matrix.The copy construction provides a discrete representation of the adiabatic partition function.
- Counting measure: Coarse-graining copy configurations by empirical pattern frequencies turns their multiplicity into multinomial counting entropy.Stirling’s formula converts the multinomial coefficient into the functional used in the f representation.
- Counting measure: The counting measure is essential at exponential load because the simplex dimension is itself exponential, making continuum volume factors uncontrolled.The model-derived counting measure avoids these uncontrolled prefactors.
- Physical interpretation: Retrieval is the fully aligned copy configuration, while a single copy defecting to a competitor carries attention quantum 1/n = λT.As T approaches zero, the number of copies diverges and aligned configurations reproduce vertex condensation.
- Relation to replica methods: The copy formulation parallels replicas while keeping a physical integer number of copies and interpreting the frozen branch through random-energy-model counting.Its symmetric order parameter emerges as the exact maximizer within copy classes, rather than as a replica ansatz.
B. Phase diagram at finite temperature
At exponential load, Model B exhibits paramagnetic, condensed, and frozen phases whose boundaries follow random-energy-model counting and norm fluctuations. Retrieval remains metastable under quantized attention defections, with finite-temperature stability split between annealed and frozen branches.
- Phase structure: The phase boundaries arise by comparing free energies and by counting exponentially many patterns subject to large-deviation constraints on their Gram matrices.The counting problem has annealed and frozen branches, with the latter pinned to the most favorable patterns available.
- Phase structure: The equilibrium phase diagram contains paramagnetic, condensed, and frozen sectors, with first-order paramagnetic–condensed boundaries and a continuous freezing transition.For λ < 1, the paramagnetic and condensed phases meet at a triple point; for λ ≥ 1, the paramagnetic phase is absent.
- Retrieval stability: Heating destabilizes retrieval through discrete attention defections, because attention weights move on the lattice p = 1 − kλT rather than continuously.At T = 1/λ, one defection empties the condensate because the attention quantum reaches the full attention weight.
- Retrieval stability: Finite-temperature retrieval stability follows annealed and frozen branches that switch near λ^2 ≈ 2α and reduce to the zero-temperature capacity as T → 0.The frozen branch controls stability when optimal-overlap destinations cease to exist, sustaining a narrow metastable strip up to T = 1/λ.
- Construction: The phase-diagram construction extends from integer copy numbers to arbitrary real temperatures without changing the sector decomposition or stability formulas.The copy number is n = β/λ, and generalized binomial expansion supplies the real-temperature continuation.
- Retrieval stability: For Gaussian patterns, typical retrieval is never the equilibrium phase, but it remains metastable throughout the retrieval region with escape times exponentially large in N_v.The frozen phase lies below pure retrieval because positive norm fluctuations favor maximum-norm patterns.
V. DISCUSSION
The class-H analysis unifies associative-memory models through hidden-neuron order parameters and separates retrieval stability from storage-scale mechanisms. It also identifies unresolved replica-symmetry and finite-hidden-temperature corrections, while motivating broader Lagrangian design.
- V. DISCUSSION: The class-H framework treats Models A–C as positions on axes set by visible entropy and hidden-sector storage behavior, rather than separate theories.The discussion connects these axes to possible new Lagrangians and deeper hierarchical class-H_k models.
- V. DISCUSSION: Replica-symmetric phase diagrams and closed-form capacities are obtained for Models A and C at polynomial load, while Model B maps to REM counting with paramagnetic, condensed, and frozen phases.The Model B hidden sector reduces to attention weights, and typical retrieval is metastable.
- V. DISCUSSION: The main limitations are the untested AT stability of the RS saddle points and the restriction of k > 2 and Model B analyses to the adiabatic limit β/τh →∞.Future work includes RSB phase diagrams, hidden-sector fluctuation stability, and testing frozen-phase attention statistics.
- 1. A family of associative memories that generates Models A and C: Positive homogeneity makes the activation degree-zero and the energy constant along rays, producing a radial zero mode whose divergent volume factors out of observables.The visible integral is therefore reduced to ray space or an appropriate sphere constraint.
- 1. A family of associative memories that generates Models A and C: The visible model family is classified by a convex body K, with the activation on ∂K selected by the direction of the visible input.The hypercube yields Ising activations, while the Euclidean ball yields spherical spins.
a. Replica setup and order parameters
The replica setup isolates one condensed pattern and represents the remaining overlaps as crosstalk noise. Replica symmetry then reduces the calculation to order parameters describing retrieval and visible-spin correlations.
- a. Replica setup and order parameters: Replica averaging introduces overlap matrices and conjugate variables, followed by the replica-symmetric ansatz for diagonal and off-diagonal entries.The ansatz uses common diagonal and off-diagonal order parameters across replicas.
- a. Replica setup and order parameters: The retrieval ansatz condenses only the first pattern and collects all non-condensed contributions into crosstalk noise.The non-condensed terms become Gaussian at large visible-neuron number by the central limit theorem.
- a. Replica setup and order parameters: The replicated partition function factorizes over non-condensed modes, allowing Gaussian representations of replica-correlated noise and subsequent single-mode integrations.Stationarity equations are extracted from the resulting factorized expression.
- a. Replica setup and order parameters: The conjugate variable r̂ becomes the Edwards–Anderson order parameter q of the visible Ising spins.Its interpretation follows because the relevant integrand is the squared thermal average of a visible spin.
b. The case k = 2: Gaussian hidden sector and the AGS theory
For k = 2, the hidden-sector integral is Gaussian and yields the standard Hopfield replica structure. The resulting free energy reproduces the AGS equations without an adiabatic approximation.
- b. The case k = 2: Gaussian hidden sector and the AGS theory: For k = 2, the hidden-sector single-mode integral is Gaussian, so the replica calculation closes through the RS free energy.The diagonal and off-diagonal replica contributions combine into the noise entropic term and the q-dependent interaction.
- b. The case k = 2: Gaussian hidden sector and the AGS theory: The RS ansatz diagonalizes the conjugate overlap matrix into one non-degenerate mode and n − 1 degenerate replica-difference modes.Expanding these eigenvalue contributions to first order in n produces the replicated thermodynamic expression.
- b. The case k = 2: Gaussian hidden sector and the AGS theory: The resulting equations of state reproduce the AGS theory of the Hopfield model near saturation.For k = 2, the hidden variables enter quadratically and their Gaussian integration is exact at any β/τh.
c. The case k > 2: breakdown of the hidden-sector integral and the adiabatic reduction
For k > 2, the direct hidden-sector integral loses an extensive saddle-point structure and the central-limit scaling fails. The analysis therefore defines the models through adiabatic elimination of hidden neurons.
- c. The case k > 2: breakdown of the hidden-sector integral and the adiabatic reduction: For k > 2, no rescaling balances the hidden-sector exponents at a nontrivial central-limit scale, so the integral is dominated by the bare confinement scale.At c = −1/2 the integrand becomes flat, whereas c = −1/k controls the dominant scale.
- c. The case k > 2: breakdown of the hidden-sector integral and the adiabatic reduction: The typical non-condensed hidden-mode amplitude scales as |m^μ| ∼ N_v^−1/k, contradicting the assumed O(1) covariance and causing divergent crosstalk variance.Continuous hidden variables therefore cannot sustain load N_h = α_k N_v^(k−1) at finite β_k.
- c. The case k > 2: breakdown of the hidden-sector integral and the adiabatic reduction: The conjugate-variable contribution becomes superextensive for k > 2, eliminating an extensive saddle-point structure.This failure cannot be removed by taking a time-scale limit at fixed β_k because the equilibrium measure depends only on β_k.
- c. The case k > 2: breakdown of the hidden-sector integral and the adiabatic reduction: Models A and C are consequently defined for k > 2 by adiabatic reduction, replacing hidden neurons with stationary values conditioned on the visible configuration.The reduction is motivated dynamically by hidden neurons tracking visible states when τ_v ≫ τ_h.
d. The case k > 2: noise sector from the cumulant expansion
For k > 2, cumulant corrections become subextensive, crosstalk noise reduces to bare Gaussian moments, and the spherical-model retrieval analysis exposes marginality at k = 2 as a contrasting case.
- Cumulant expansion: For k > 2, the third and higher cumulants are subextensive, so non-condensed overlaps remain bare Gaussian variables with r = M_k(q).The feedback vanishes in the thermodynamic limit, unlike the O(1) resummation required for k = 2.
- Spherical marginality: At k = 2, retrieval lies at the boundary βk(1 − q) = 1, where the condensed direction is flat and non-condensed fluctuations diverge.The same marginality appears as loss of normalizability in the noise sector and as a zero mode of the quadratic spherical energy.
- Noise sector: The crosstalk moment is identical for Models A and C, while the k = 2 spherical model is marginal and stores nothing without higher-order interactions.This agreement is confirmed both by the replica calculation and by an independent feedback argument.
- Exponential load: At exponential load, maximal-norm Gaussian patterns dominate the quenched free energy, producing freezing and lowering equilibrium energy below typical retrieval.Their extreme-value enhancement raises the equilibrium boundary above the naive α = λ/2 comparison.
- Exponential load: The zero-temperature capacity follows different branches: α < 1/2 on the frozen branch and α < λ − λ^2/2 on the annealed branch, with the applicable branch set by λ ≷ 1.The two conditions combine into a capacity continuous at λ = 1.
- Copy representation: The copy representation shows that hidden-sector fluctuations shift the exponent and pattern terms, but their consistent treatment at exponential load remains open.The correction is nonperturbative because the shifted components scale as √N_h when N_h = e^(αN_v).
3. Finite-temperature analysis in the copy representation
The copy representation identifies paramagnetic, condensed, and frozen branches and shows that finite-temperature retrieval stability is controlled by discrete attention reassignments. Its real-temperature continuation preserves the one-quantum spinodal and connects retrieval to bulk phases through quantized nucleation ladders.
- Bulk branches: The copy construction yields paramagnetic, condensed, and frozen branches, with the condensed and frozen phases joining continuously at the freezing line.The condensed entropy vanishes on the freezing line, where both branches share the same rate and the frozen state corresponds to complete retrieval of the maximal-norm pattern.
- Bulk branches: The annealed existence condition is β∥f∥2 < 1, linking attention concentration to divergence of the visible Gaussian integral.For symmetric attention over M patterns, ∥f∥2 = 1/M; the corresponding conditions are λ < 1 for the paramagnetic branch and β < 1 for the condensed branch.
- Real-temperature continuation: The one-quantum criterion is exact for arbitrary real temperature, and the replica-copy construction extends formulas beyond the integer temperature points.The real-temperature derivation also resolves the fluctuation determinant and hidden-sector zero mode while establishing unique analytic continuation.
- Real-temperature continuation: The real-temperature construction makes the retrieval surface continuous in overlap but quantized in attention, with sectors located at p_k = 1 − kλT.Maximizing over the overlap recovers the retrieval branch, while the discreteness of finite-temperature physics resides in the attention coordinate.
- Stability and nucleation: Finite-temperature retrieval stability is set by the one-copy spinodal, with the minimum of the annealed and frozen criteria governing the capacity at every real temperature.Co-defection channels do not destabilize retrieval first; they instead provide nucleation paths toward condensed phases.
- Stability and nucleation: Distinct-defection ladders terminate at the paramagnetic branch, whereas co-defection ladders terminate at condensed or frozen branches, closing retrieval onto the bulk phases.At noninteger temperatures, dual expansions connect the ladder endpoints to the same bulk values without a gap.