Source-linked AI summary
Why Neurons Have Thousands of Synapses, A Theory of Sequence Memory in Neocortex
Jeff Hawkins, Subutai Ahmad
TL;DR
The paper addresses how neocortical neurons integrate thousands of synapses and how this supports large-scale sequence behavior. It models active dendrites as multiple pattern detectors and builds a predictive sequence-memory network, finding robust recognition and sequence learning under sparse activity. The authors propose that related sequence-memory mechanisms may operate throughout neocortex, while acknowledging unresolved biological interactions and implementation requirements.
Problem
It is unclear how neocortical neurons integrate thousands of synapses, especially distal synapses, and what network behavior this enables.
Method
The paper models active dendrites as nonlinear pattern detectors and uses proximal, basal, and apical synapses for feedforward responses, contextual predictions, and top-down expectations.
Results
The models recognize hundreds of patterns robustly and learn predictive time-based sequences with online learning, multiple predictions, and resistance to noise and variation.
Takeaways & Limitations
The authors propose sequence memory as a common property of neocortical tissue, with cortical layers implementing variations of a shared algorithm.
Takeaways & Limitations
The physiological interaction between apical and basal dendrites remains an area of ongoing research, and the model does not fully resolve it.
Abstract
from arXiv · showhide
Neocortical neurons have thousands of excitatory synapses. It is a mystery how neurons integrate the input from so many synapses and what kind of large-scale network behavior this enables. It has been previously proposed that non-linear properties of dendrites enable neurons to recognize multiple patterns. In this paper we extend this idea by showing that a neuron with several thousand synapses arranged along active dendrites can learn to accurately and robustly recognize hundreds of unique patterns of cellular activity, even in the presence of large amounts of noise and pattern variation. We then propose a neuron model where some of the patterns recognized by a neuron lead to action potentials and define the classic receptive field of the neuron, whereas the majority of the patterns recognized by a neuron act as predictions by slightly depolarizing the neuron without immediately generating an action potential. We then present a network model based on neurons with these properties and show that the network learns a robust model of time-based sequences. Given the similarity of excitatory neurons throughout the neocortex and the importance of sequence memory in inference and behavior, we propose that this form of sequence memory is a universal property of neocortical tissue. We further propose that cellular layers in the neocortex implement variations of the same sequence memory algorithm to achieve different aspects of inference and behavior. The neuron and network models we introduce are robust over a wide range of parameters as long as the network uses a sparse distributed code of cellular activations. The sequence capacity of the network scales linearly with the number of synapses on each neuron. Thus neurons need thousands of synapses to learn the many temporal patterns in sensory stimuli and motor sequences.
1. Introduction
The paper develops a theory explaining how neocortical neurons use thousands of synapses and active dendrites to recognize patterns and support sequence memory. It proposes that these mechanisms are broadly applicable across neocortical tissue.
- Motivation: Neocortical neurons have thousands of excitatory synapses, most of them distal, whose functional contribution is difficult to explain using conventional neuron models.Active dendrites are presented as processing elements that may resolve this integration problem.
- Contribution: The theory shows that pyramidal neurons with active dendrites can recognize hundreds of unique cellular-activity patterns despite substantial noise and variability.Recognition remains reliable when overall neural activity is sparse.
- Contribution: The model assigns proximal inputs to feedforward responses, while distal patterns depolarize neurons as predictions without directly causing action potentials.This separates classic receptive-field responses from predictive cellular states.
- Contribution: A network of these neurons learns and recalls sequences by using predictive depolarization to bias future sparse activity.The network is designed to learn time-based sequences continuously and robustly.
- Implications: Because excitatory neurons and sequence memory are widespread across neocortex, the paper proposes sequence memory as a unifying principle of neocortical function.The authors further propose that cortical layers implement variations of a common sequence-memory algorithm.
2.1. Neurons Recognize Multiple Patterns
Active dendrites let a neuron act as multiple nonlinear pattern detectors rather than as a single point neuron. Synaptic location separates feedforward receptive-field responses from contextual and top-down predictions.
- Multiple pattern recognition: Active dendrites turn neighboring synapses into nonlinear pattern detectors, so a neuron’s thousands of synapses can recognize multiple patterns.Coincident activation of approximately eight to twenty nearby synapses can generate an NMDA dendritic spike.
- Multiple pattern recognition: Sparse activity enables robust pattern recognition by allowing dendritic segments to classify population patterns from a small synaptic sample.The model defines recognition as at least θ active matches among s synapses in a sparse population.
- Multiple pattern recognition: 300 patterns are approximately supported by 6,000 synapses when 20 synapses are allocated to each pattern.This is a rough capacity estimate for a neuron with active dendrites.
- Synaptic integration zones: Most recognized patterns depolarize the neuron without directly producing an action potential, allowing them to serve as predictive states.This predictive role differs from proximal inputs that define the cell’s feedforward receptive field.
- Synaptic integration zones: Basal dendrites learn preceding activity patterns, and their subthreshold depolarization causes correctly predicted cells to fire earlier and inhibit neighbors.The proposed mechanism makes anticipated inputs more sparse.
- Synaptic integration zones: Apical dendrites use recognized patterns to establish top-down expectations through depolarization that typically remains below somatic action-potential threshold.The interaction among apical spikes, basal spikes, and somatic action potentials remains under study.
2.2. Networks of Neurons Learn Sequences
The paper proposes a cortical sequence-memory network in which local dendritic learning stores transitions and predictions across sparse cellular representations. Simulations show online adaptation, high-order and simultaneous predictions, and robustness, while some biological mechanisms remain unresolved.
- 2.2. Networks of Neurons Learn Sequences: The proposed fundamental neocortical operation is learning and recalling sequences of patterns, with cortical layers implementing variations of a common algorithm.The paper presents a basic sequence-memory algorithm without detailing all layer-specific variations.
- Network requirements: The network is designed for continuous online learning, high-order predictions, multiple simultaneous predictions, local learning rules, and robustness during streaming data.These properties are required to occur simultaneously rather than independently.
- Sequence learning: Basal synapses learn transitions between feedforward patterns by depolarizing cells that are predicted to become active in the next input.Feedforward input activates cells, whereas basal input generates predictions.
- Sequence learning: Ambiguous subsequences can produce multiple simultaneous predictions because different cells within a mini-column learn different temporal contexts.Sparse predictions allow several candidate continuations without confusion.
- Top-down expectation: Top-down apical input can predict multiple elements of a sequence simultaneously, complementing basal prediction of the next input.The physiological interaction between apical and basal dendrites remains an active research area.
3. Simulation Results
The HTM network learns high-order temporal predictions online from sparse inputs, reaches the task’s maximum accuracy, and remains robust to substantial neuron loss.
- Network configuration: The network used 2048 mini-columns with 32 neurons each; every neuron had 128 basal dendritic segments with up to 40 synapses per segment.The simulation omitted apical synapses to focus on sequence-memory properties.
- Input and task: Inputs activated 40 of 2048 mini-columns, and the network learned predictions from transitions combining random elements with six-element sequences.The sequences required high-order temporal context for disambiguation and best prediction accuracy.
- Learning and prediction: 50% maximum average prediction accuracy was achieved by the high-order HTM network, compared with about 33% for the first-order network.The 50% ceiling was defined by the task design and was only achievable with high-order representations.
- Learning and prediction: After the embedded sequences changed, accuracy initially dropped but recovered as continual learning acquired the new high-order patterns.This demonstrates online adaptation rather than fixed-sequence recall.
- Robustness: At up to about 40% cell death, performance showed minimal impact; with greater loss, performance initially declined and then recovered through relearning.The reported robustness is attributed to dendritic segments forming more synapses than necessary to generate an NMDA spike.
4. Discussion
The discussion presents active-dendrite neurons as reliable pattern recognizers whose predictive interactions support sparse, robust sequence memory. It also outlines capacity, biological predictions, implementation boundaries, and unresolved mechanisms for cortical and hippocampal applications.
- Model neuron: Active dendrites and thousands of synapses let model neurons recognize hundreds of unique patterns despite noise and variation.Proximal synapses define feedforward receptive fields, while basal and apical synapses depolarize cells as predictions.
- Network behavior: A network of these neurons learns predictive models of data streams using contextual basal inputs and feedback-related apical inputs.The model operates with sparse activity and can learn continuously, use variable context, and make multiple simultaneous predictions.
- Capacity: Network capacity measures stored transitions rather than complete sequences and scales linearly with cells per column and basal patterns recognized per neuron.With 2% active columns, 32 cells per column, and 200 basal patterns per cell, the paper estimates approximately 320,000 stored transitions.
- Experimental predictions: The theory makes testable predictions about activity sparsening during predictable streams, temporal context, and dendritic plasticity.It predicts higher, vertically correlated activity for unanticipated inputs and localized plasticity after depolarization followed shortly by a back action potential.
- Limitations and extensions: Important scope boundaries include unresolved mini-column excitation and inhibition, costly silent-synapse extensions for rapid hippocampal learning, and incomplete cortical-layer modeling.The authors also relate the mechanism to HMMs and spiking sequence models while identifying cortical motor-sequence integration as ongoing work.
- Generalization: Sparse distributed representations support generalization because dendritic segments can treat novel but semantically related patterns as similar.The system may therefore generate novel predictions based on analogy across different sequences.
5. Materials and Methods
The HTM sequence-memory network represents cells with active, predictive, and inactive states, and updates dendritic segments and synapses through local rules. Feedforward activation selects winning columns, while distal segments detect context and learning reinforces or decays synapses based on prediction outcomes.
- Network representation: Each cell can be active, predictive (depolarized), or non-active, and maintains distal segments containing synapses to other cells.Segments store permanence values; synapses above a connection threshold are treated as connected.
- Cell-state computation: At each time step, inhibitory selection chooses k winning columns that best match the feedforward input.All cells within a mini-column share the same feedforward receptive field.
- Cell-state computation: A cell in a winning column becomes active if it was previously predictive; otherwise, all cells in that column become active.The predictive state for the current time step is then computed from dendritic segment activity.
- Cell-state computation: A dendritic segment becomes active when more than θ connected synapses have active presynaptic cells, depolarizing its cell when at least one segment is active.Here, θ is the NMDA spiking threshold and ∘ denotes element-wise multiplication.
- Learning rules: Correct predictions reinforce the segment responsible for depolarization, while unpredicted transitions select the segment closest to activation for future context.The learning rule rewards synapses from active presynaptic cells and punishes inactive ones.
- Learning rules: The model decreases all permanence values slightly, increases values for active presynaptic cells more strongly, and applies small decay to mistakenly active segments.The simulations use N = 2048, M = 32, and k = 40, with segments created as needed when no existing segment matches.
S1 Text. Chance of Error When Recognizing Large Patterns with a Few Synapses
The supplement analyzes false matches when dendritic segments recognize sparse activity patterns through subsampled synapses. It reports that larger sampling, added synapses, and appropriate thresholds support reliable recognition under noise and when multiple patterns share one segment.
- Chance of error: The false-match probability depends on population size, active-cell count, synapse count, and the NMDA spike threshold.The formulation treats nonlinear dendritic segments as classifiers that subsample a larger active-cell population.
- Chance of error: Increasing the sampling size rapidly reduces the chance of error, so relatively few synapses can provide reliable pattern matching.This is the main result summarized by Table A.
- Noise robustness: Setting s = 2θ provides immunity to 50% noise while retaining a low probability of false matches.The chance of error decreases rapidly as θ increases, and a small number of synapses remains sufficient even with noise.
- Mixed-pattern recognition: A segment can recognize m independent mixed patterns while remaining robust to 50% noise when s = 2mθ.Higher accuracy for larger m is possible with a slightly higher threshold.
- Algorithm comparison: The supplement presents the proposed HTM model alongside HMM and LSTM in a comparison of common sequence-memory algorithms.The comparison notes that LSTMs require a globally computed error signal and backpropagation despite local weight updates.