Source-linked AI summary
Large Associative Memory Problem in Neurobiology and Machine Learning
Dmitry Krotov, John Hopfield
TL;DR
The paper addresses the storage limits of conventional associative memories and the many-body synapses used by large-capacity models. It introduces hidden neurons in a continuous energy-based network with pairwise connections, and shows that the increased capacity comes from added hidden neurons and synapses while the framework recovers established models.
Problem
Conventional associative memories have storage limits, while Dense Associative Memories use many-body interactions that are not microscopic biological synapses.
Method
The paper introduces continuous feature and hidden neurons connected through a bipartite, two-body synaptic network with an energy function and selected activation functions.
Results
The framework has large memory storage capacity with two-body synaptic connections and recovers Dense Associative Memory limits, including exponential capacity for F(x) = exp(x).
Takeaways & Limitations
The increased capacity is attributed to unfolding the effective theory by adding hidden neurons and synapses, while each synapse retains the information capacity of conventional synapses.
Takeaways & Limitations
The capacity is bounded by the number of hidden neurons, Nmem ≤ Nh, and the biological-plausibility claim is defined specifically as the absence of many-body synapses.
Abstract
from arXiv · showhide
Dense Associative Memories or modern Hopfield networks permit storage and reliable retrieval of an exponentially large (in the dimension of feature space) number of memories. At the same time, their naive implementation is non-biological, since it seemingly requires the existence of many-body synaptic junctions between the neurons. We show that these models are effective descriptions of a more microscopic (written in terms of biological degrees of freedom) theory that has additional (hidden) neurons and only requires two-body interactions between them. For this reason our proposed microscopic theory is a valid model of large associative memory with a degree of biological plausibility. The dynamics of our network and its reduced dimensional equivalent both minimize energy (Lyapunov) functions. When certain dynamical variables (hidden neurons) are integrated out from our microscopic theory, one can recover many of the models that were previously discussed in the literature, e.g. the model presented in "Hopfield Networks is All You Need" paper. We also provide an alternative derivation of the energy function and the update rule proposed in the aforementioned paper and clarify the relationships between various models of this class.
1 INTRODUCTION
Associative-memory models must retrieve many related items from partial cues, but conventional networks have storage limits and Dense Associative Memories use non-biological many-body interactions. The paper asks whether hidden circuitry can preserve large capacity while using only two-body synapses.
- Associative memory: Associative memory retrieves the remaining items of a memory when given a sufficiently large subset of its items.The paper frames this ability as linking many sets of otherwise unrelated items.
- Storage limits: Conventional feature-only systems store at most roughly Nf unrelated memories because their synapses contain only limited information.The paper reports a classical Hopfield limit below approximately 0.14Nf memories even with precise synapses.
- Many-body interactions: Dense Associative Memories increase storage capacity but use higher-order interaction tensors that cannot represent ordinary two-cell biological synapses.For a cubic interaction, the coupling tensor has three indices rather than the two indices of a synaptic connection.
- Research question: The paper asks whether hidden neurons and suitable interactions can provide capacity significantly larger than Nf while remaining describable using only two-body synapses.This question targets the hidden circuitry underlying the effective large-capacity models.
- Paper approach: The authors extend prior Dense Associative Memory work to continuous states and time, add complex hidden neurons, and identify limits that recover earlier associative-memory models.The proposed system has Nf + Nh variables and includes models A, B, and C.
2 MATHEMATICAL FORMULATION
The proposed continuous-time network uses two neuron groups connected as a bipartite graph, with an energy function composed of feature, hidden, and interaction terms. Its construction supports pairwise synapses and monotonic energy decrease under stated conditions.
- Network architecture: The network uses feature and hidden neurons connected only across groups, forming a bipartite architecture without within-group synapses.The hidden-neuron outputs and feature-neuron outputs can use nonlinear activations or contrastive normalization, depending on the model.
- Dynamics: The model is formulated as a continuous-time dynamical system whose updates depend on inputs from other neurons and each neuron's own decay state.The Lagrangian functions are chosen so their derivatives correspond to neuron outputs and support the energy formulation.
- Energy decrease: The energy decreases monotonically along dynamical trajectories when the specified dynamical equations and energy construction are used.This establishes a Lyapunov-function interpretation for the network dynamics.
- Convergence: A bounded energy, for example obtained with bounded feature activations such as tanh or sigmoid, leads the dynamics to a fixed point at a local energy minimum.Positive-semidefinite Hessian conditions provide the relevant bound in the stated formulation.
- Energy construction: The energy function separates feature-only, hidden-only, and cross-group interaction terms, with coupling parameters ξµi representing feature–memory synapses.This structure yields conventional recurrent-neural-network equations in which neurons collect weighted outputs from the other group.
3 EFFECTIVE THEORY FOR FEATURE NEURONS
Integrating out fast hidden neurons reduces the general two-layer theory to established associative-memory models, including Dense Associative Memory and modern Hopfield networks, while preserving energy-based dynamics. The resulting limits also expose capacity bounds, attention-like updates, and divisive normalization.
- Effective reductions: Integrating out hidden neurons recovers classical Hopfield, Dense Associative Memory, and modern Hopfield models from the general theory.The modern Hopfield update rule has the mathematical structure of dot-product attention and is used in Transformer networks.
- Model A: With F(x) = x^n, the Dense Associative Memory limit stores N_mem ∼ N_f^(n−1), while F(x) = exp(x) gives exponential storage capacity.For n = 2, the model reduces further to the classical Hopfield network.
- Model A: Capacity estimates assume unlimited hidden neurons and are additionally bounded by N_mem ≤ N_h.Thus, hidden-neuron count constrains the realizable capacity even when the feature-space estimate is larger.
- Model A: For additive models, positive-definite Hessians are equivalent to monotonically increasing feature- and memory-neuron activation functions.These conditions support the energy-based analysis of the dynamics.
- Model B: In the fast-hidden-neuron limit τ_h → 0, contrastively normalized models reduce to the energy function studied by Ramsauer et al. up to additive constants.The derivation assumes zero input currents and sets inverse temperature β to one.
- Model B: The corresponding effective update is exactly the Ramsauer et al. rule, whose one-step form is equivalent to dot-product attention and is used in Transformers.The paper presents this as a continuous-time counterpart before taking finite differences.
- Model C: Model C introduces spherical normalization in the feature layer and yields an activation function implementing canonical divisive normalization.Divisive normalization is also reported as beneficial in deep CNNs and RNNs for image classification and language modeling.
4 A FEW EXAMPLES OF LARGE ASSOCIATIVE MEMORY PROBLEMS
The paper illustrates large associative-memory needs in AI and biology, where the number of desired memories can greatly exceed the feature-space dimension. Examples include image and immune-repertoire classification, color perception, and hippocampal memory systems.
- Pattern memorization: Small grayscale images can require far more memories than standard Hopfield networks can store.For 64×64 images, standard associative memory stores approximately 573 patterns, while the Kuzushiji-Kanji dataset contains over 140,000 characters.
- Immune repertoire classification: Immune-repertoire classification involves more than 10,000 sequences embedded in only 32 dimensions.The task therefore requires storage capacity much larger than the feature-space dimensionality.
- Cortical-hippocampal system: The hippocampus is discussed as a candidate biological substrate for associative memory, including CA3 recurrent networks and possible CA1–entorhinal mappings.The paper identifies place cells and other hippocampal responses as potentially related to feature and memory neurons.
- Color representation: Color vision presents a large-memory problem because three cone dimensions support many named color sensations.The paper treats color representations as continuous feature inputs with Nf = 3.
5 DISCUSSION AND CONCLUSIONS
The paper argues that large associative-memory models can be recast as two-body, biologically interpretable systems by adding hidden neurons. It also frames the general formulation as a foundation for relating existing models and developing new recurrent architectures.
- 5 DISCUSSION AND CONCLUSIONS: The proposed dynamical system combines large memory capacity with an explicit description using only two-body synaptic connections.The authors compare its biological plausibility with conventional continuous Hopfield networks and its psychological plausibility with their larger memory capacity.
- 5 DISCUSSION AND CONCLUSIONS: Adding hidden neurons increases storage capacity by unfolding the effective theory and adding synapses.The paper attributes the increase to more synapses, each retaining the same information capacity as conventional synapses.
- 5 DISCUSSION AND CONCLUSIONS: The general formulation provides a conceptually grounded derivation of associative-memory models and their relationships.The authors propose that it may support development of new recurrent neural-network architectures.
APPENDIX A
The appendix derives the time evolution of the proposed energy function under the network dynamics. It concludes that the energy decreases along trajectories under a positive-semidefinite Hessian condition.
- APPENDIX A: The appendix differentiates the energy with respect to feature-neuron and hidden-neuron activities.The input currents are treated as time-independent during the calculation.
- APPENDIX A: The dynamical equations convert the derivative terms into activity time derivatives, completing the Lyapunov-decrease proof.The result holds for arbitrary feature- and memory-neuron time constants when the corresponding Hessians are positive semi-definite.
APPENDIX B. THE LIMIT OF STANDARD CONTINUOUS HOPFIELD NETWORKS.
The appendix shows how classical continuous Hopfield networks arise as a special limit of the general theory. It establishes equivalence between apparently different energy expressions through Legendre-transform identities.
- APPENDIX B. THE LIMIT OF STANDARD CONTINUOUS HOPFIELD NETWORKS.: Classical continuous Hopfield dynamics are obtained from the general theory for graded-response neurons.The formulation uses an activation function and its inverse.
- APPENDIX B. THE LIMIT OF STANDARD CONTINUOUS HOPFIELD NETWORKS.: The model is classified as a special limit of models A with a particular choice of Lagrangian functions.Those Lagrangians determine the corresponding activation functions.
- APPENDIX B. THE LIMIT OF STANDARD CONTINUOUS HOPFIELD NETWORKS.: Integrating out hidden neurons reduces the general dynamics to equations on the feature neurons and an effective energy.The appendix presents this reduction as the route from the general theory to the classical formulation.
- APPENDIX B. THE LIMIT OF STANDARD CONTINUOUS HOPFIELD NETWORKS.: The apparently different third energy terms are equivalent because derivatives of a function and its Legendre transform are inverse functions.Differentiating both expressions gives vi g(vi)′, so they differ only by an additive constant.