Source-linked AI summary
A sparse coding model with synaptically local plasticity and spiking neurons can account for the diverse shapes of V1 simple cell receptive fields
Joel Zylberberg, Jason Timothy Murphy, Michael Robert DeWeese
TL;DR
It was unknown whether sparse codes for natural images could be learned with biologically realistic, synaptically local plasticity rules. The paper develops a spiking network using such rules and shows that it accounts for the diversity of V1 simple-cell receptive fields, while mathematically relating local learning to sparseness and decorrelation. The authors also identify emergent network properties relevant to comparisons with visual cortex.
Problem
It was unknown whether sparse coding could be achieved under biological architectural constraints using plasticity rules that are local to synapses.
Method
The paper develops a biophysically motivated spiking network that learns from natural images using only synaptically local rules, with homeostasis and lateral inhibition maintaining sparse, independent activity.
Results
The network accounts for the observed diversity of V1 simple-cell receptive-field shapes and mathematically approximates an optimal linear generative model under firing-rate and temporal-correlation constraints.
Takeaways & Limitations
The work provides a demonstration that synaptically local plasticity can support sparse coding of natural images while reproducing diverse V1 receptive fields.
Takeaways & Limitations
The model's inhibitory connection strengths may not be biologically realistic, and the distribution of inhibitory strengths in cortex remains unknown.
Abstract
from arXiv · showhide
Sparse coding algorithms trained on natural images can accurately predict the features that excite visual cortical neurons, but it is not known whether such codes can be learned using biologically realistic plasticity rules. We have developed a biophysically motivated spiking network, relying solely on synaptically local information, that can predict the full diversity of V1 simple cell receptive field shapes when trained on natural images. This represents the first demonstration that sparse coding principles, operating within the constraints imposed by cortical architecture, can successfully reproduce these receptive fields. We further prove, mathematically, that sparseness and decorrelation are the key ingredients that allow for synaptically local plasticity rules to optimize a cooperative, linear generative image model formed by the neural representation. Finally, we discuss several interesting emergent properties of our network, with the intent of bridging the gap between theoretical and experimental studies of visual cortex.
I. INTRODUCTION
Sparse coding offers a candidate principle for visual cortical processing, but prior models did not establish that biologically realistic, synaptically local mechanisms could learn it. This work presents a spiking network that addresses that gap and reproduces the diversity of V1 simple-cell receptive fields.
- Motivation: Sparse coding has been proposed as an efficient neural code, but cortical activity includes sparse, dense, and mixed response patterns.The absolute standard for assessing cortical sparseness is also unclear.
- Prior models: Earlier natural-image sparse-coding models reproduced some V1 receptive-field features, but agreement with measured simple-cell fields was not perfect.The SSC network later learned small unoriented features, Gabor-like oriented features, and elongated edge detectors.
- Biological gap: Prior sparse-coding algorithms required learning rules with access to receptive-field information from many distant neurons, unlike synaptically local cortical plasticity.Existing models also commonly used non-spiking computational units, whereas cortical information is transmitted through discrete spikes.
- Biological gap: No previous spiking image-processing network had been shown to learn the full diversity of V1 receptive-field shapes using local plasticity rules.The unresolved question was whether sparse coding could be achieved under biological architectural constraints.
- Present work: The proposed spiking network uses synaptically local learning, with homeostasis and lateral inhibition maintaining sparse and independent activity during training.The model is presented as a biologically inspired variation of a Foldiak network.
- Present work: The authors mathematically show that the network approximates an optimal linear generative model under constraints on average firing rates and temporal correlations.They report the first demonstration that local plasticity can account for diverse V1 simple-cell receptive fields and derive its relationship to sparseness and decorrelation.
II. RESULTS
SAILnet is a biophysically inspired spiking network that learns sparse representations of natural images using synaptically local plasticity. Its learned receptive fields reproduce the diversity of V1 simple-cell shapes.
- Network and learning: SAILnet uses spiking leaky integrate-and-fire neurons whose outputs are binary spikes generated by threshold crossing.The network represents image inputs through time-dependent internal variables and discrete spike outputs.
- Network and learning: The model learns sparse, weakly correlated activity while updating feed-forward, inhibitory, and threshold parameters with local information.Synaptic updates depend on pre- and postsynaptic activity or the unit’s own firing rate, rather than global network activity.
- Network and learning: Linear decoding from SAILnet activity approximately recovers the input stimulus despite the network’s nonlinear spike-based encoding.The paper attributes this property to the learning rules.
- Network and learning: The local learning rule approximates a prior non-local rule when neuronal activity is highly sparse and uncorrelated, averaged across images.This provides the mathematical link between local plasticity and optimization of the cooperative linear generative model.
- Experimental setup: The model was trained with p = 0.05 on 16 × 16 patches from whitened natural images using a 1536-unit, six-times-overcomplete network.The overcomplete architecture reflects the larger number of V1 neurons relative to its inputs, while larger networks were limited by O(N^2) parameters.
- Receptive fields: SAILnet receptive fields show the diversity observed in macaque V1, including unoriented, oriented Gabor-like, and elongated feature classes.The model accounts for these shapes using only synaptically local rules applied to natural images.
natural images
Although learning targets the same average firing rate for every unit, SAILnet produces broad firing-rate distributions when probed with natural images. Stimulus contrast changes the distribution’s qualitative form.
- Natural-image responses: A broad, approximately lognormal distribution of mean firing rates emerges when the trained network is probed with natural-image stimuli.The distribution is measured after training with learning turned off.
- Contrast dependence: Lower-contrast images produce a monotonic decreasing firing-rate distribution that is fit similarly by lognormal or exponential functions.The same learned network parameters are used for both contrast conditions.
- Contrast dependence: The contrast-dependent shift follows from slower charging and fewer spikes within the fixed presentation period for reduced pixel values.The low-rate tail is effectively truncated because negative firing rates are impossible.
- Measurement effects: Fifty thousand probe images are sufficient to estimate the underlying firing-rate distribution because its variance reaches an asymptote after approximately 25,000–30,000 presentations.The authors conclude that finite sample-size effects do not determine the observed distributions.
- Learning effects: Non-zero learning updates broaden firing-rate distributions because threshold jumps can push units above or below their target firing rate.This variation persists after learning converges because parameters continue fluctuating around their average values.
- Learning effects: Changing learning rates alters distribution variance but preserves the qualitative contrast-dependent pattern.Training-set responses remain non-monotonic and approximately lognormal, whereas low-contrast responses remain monotonic and exponential/lognormal.
Pairs of SAILnet units have small firing rate correlations.
SAILnet produces near-zero pairwise firing-rate correlations, matching a qualitative feature observed experimentally. The model’s correlation distribution is narrower than the experimental distribution.
- Correlation structure: Correlations between SAILnet unit spike counts across 30,000 natural images tend to be near zero.This agrees qualitatively with experimental observations.
- Correlation structure: The model’s correlation-coefficient distribution has smaller variance than the experimental data.The authors note that experimental correlations show a larger spread.
- Correlation structure: Larger learning-rate updates produce a larger variance in the simulated correlation-coefficient distribution.Thus, plasticity step size affects the spread of measured correlations.
- Correlation structure: The correlation distribution appears truncated on the left because never-coactive neurons have a lower bound on their possible correlation.With p = 0.05, this bound is not far below zero.
model
SAILnet learns connectivity patterns alongside sensory representations, including approximately lognormal inhibitory strengths and stronger inhibition between neurons with overlapping receptive fields.
- Connection-strength distributions: Inhibitory connection strengths learned from natural images are approximately lognormally distributed, with a Gaussian fit to their logarithms explaining 98% of the variance.The model nonetheless shows systematic deviations from the fit, especially in the low-strength tail.
- Experimental predictions: The model predicts approximately lognormal inhibitory functional connections between excitatory V1 simple cells.This extends connectivity predictions beyond the excitatory connections measured experimentally.
- Experimental predictions: The model does not specify whether inhibitory-strength variation arises from dendritic or axonal synapses of inhibitory interneurons.This leaves the anatomical source of the predicted variability unresolved.
- Receptive-field connectivity: Inhibitory strength correlates with receptive-field overlap: neurons with substantially overlapping fields tend to inhibit one another strongly.Shared feed-forward input increases the need for mutual inhibition to maintain uncorrelated activity.
- Receptive-field connectivity: This overlap-dependent inhibition is learned naturally by SAILnet in response to natural stimuli, rather than being imposed as in the LCA algorithm.The comparison concerns how the connectivity structure arises in the models.
- Experimental predictions: The predicted connectivity patterns are directly testable experimentally, but measuring functional interactions mediated by two or more synapses may be difficult.The practical challenge concerns inhibitory pathways between excitatory simple cells.
III. DISCUSSION
The discussion presents synaptically local plasticity as sufficient for sparse coding that accounts for diverse V1 simple-cell receptive-field shapes, while identifying biological simplifications and testable predictions. It also highlights dynamic-stimulus behavior emerging from the network’s finite update time.
- Contribution: Synaptically local plasticity is presented as sufficient for learning sparse codes that account for diverse V1 simple-cell receptive-field shapes.The authors contrast this result with prior models that used nonlocal rules or did not demonstrate the diversity of observed receptive fields.
- Contribution: The learning rule updates connection strengths using only the number of arriving presynaptic spikes and postsynaptic spikes.This local rule is contrasted with a nonlocal rule requiring inhibitory synapses to track receptive-field changes.
- Limitations: The model alternates between brief inference and learning periods, and the authors state that this timing may not be biologically realistic.They note uncertainty about how cortical neurons would know when inference ends and learning begins, while suggesting a possible link to saccade onset.
- Limitations: Additional scope limitations include continuous-valued inputs despite spiking retinal inputs, direct inhibitory connections instead of interneuron populations, omitted spike-timing-dependent plasticity, and no intrinsic activity noise.The authors describe ongoing work incorporating timing-dependent learning for time-varying natural movies.
- Emergent properties: Finite integration time produces hysteresis: previous frames influence how the network processes and represents the current frame.The authors suggest this may stabilize image representations relative to models such as ICA or sparsenet.
- Biological correspondence: SAILnet combines spiking neurons, sparse activity, largely uncorrelated responses, inhibitory lateral connections, and an overcomplete representation.These properties are described as qualitative features of visual cortex captured by the simplified model.
- Experimental implications: The model supports falsifiable experimental predictions about interneuronal connectivity and cortical population activity.The authors propose that these predictions could help identify coding principles in visual cortex.
IV. METHODS
The methods implement SAILnet as a leaky integrate-and-fire network whose neurons integrate image-driven and lateral synaptic currents, spike at threshold, reset, and learn through synaptically local updates. Simulations run the membrane dynamics for a fixed interval after each image while slowly adapting thresholds.
- Neuron dynamics: Each SAILnet neuron follows leaky integrate-and-fire dynamics with an internal variable analogous to membrane voltage and a brief binary output spike.The internal variable represents capacitor voltage in an RC-circuit model, while the output is 1 briefly after threshold crossing and 0 otherwise.
- Synaptic inputs: Image inputs and spikes from other neurons modify each neuron’s internal variable through feed-forward weights Q_ik and lateral strengths W_im.These weights determine how much each pixel value or incoming spike changes the neuron’s internal state.
- Neuron dynamics: The membrane variable obeys du_i(t)/dt + u_i(t) = I_input(t), integrating the synaptic input current over time.The model treats the RC time constant as one unit of time.
- Simulation: For each input image, the dynamics run for five RC time units using a 0.1-unit integration step, with all neurons initialized at u_i(t = 0) = 0.The simulation numerically integrates the differential equation in discrete time.
- Spike generation: When u_i(t*) exceeds the neuron-specific threshold θ_i, the neuron emits y_i(t* + 1) = 1 and its output returns to zero afterward unless it spikes again.After spiking, the internal variable returns to its resting value of 0 and can be recharged.
- Learning and adaptation: SAILnet learning rules are framed as gradient descent, while thresholds adapt slowly relative to inference and synaptic weights remain approximately constant during inference.This separates fast neural dynamics from slower parameter adaptation.
constrained optimization problem
SAILnet formulates learning as constrained optimization that minimizes reconstruction error while enforcing sparse, weakly correlated neural activity. Under these constraints, its local weight updates approximately solve the same error-minimization problem as non-local sparse-coding methods and support linear input reconstruction.
- Constraints and optimization: SAILnet uses Lagrange multipliers to minimize reconstruction error while approximately enforcing fixed average firing rates and minimal correlations.Variable thresholds and inhibitory connections implement the corresponding constraints in the network.
- Constraints and optimization: The objective terms enforcing sparseness and decorrelation are critical because feed-forward weight changes otherwise alter firing rates and inter-neuron correlations.Without forces returning the network to the constraint surface, reconstruction-error minimization alone would not preserve these properties.
- Synaptic locality: The non-local correlation term is biologically problematic because synaptic updates should use only information available locally at each synapse.A synapse can access presynaptic activity, postsynaptic activity, and its own strength, but not other neurons’ receptive fields or activities.
- Synaptic locality: Sparse, uncorrelated activity makes the non-local update approximately equivalent to Oja’s synaptically local Hebbian rule.This lets SAILnet approximately solve the same error-minimization problem as previous non-local sparse-coding algorithms.
- Linear representation: Despite nonlinear spike generation, SAILnet supports a linear decoding in which the input is approximately reconstructed from neural activity and feed-forward weights.Spike-triggered averages are proportional to the feed-forward weights when probe and training stimulus statistics match.
Training SAILnet
SAILnet is trained in batches of natural-image stimuli by averaging activity-dependent updates and maintaining inhibitory connections as nonpositive. Training reaches a dynamic equilibrium after roughly 10^7 presentations, while simulations continue for about 2 × 10^8 presentations.
- Training procedure: Training presents batches of 100 zero-mean, unit-standard-deviation images and counts each neuron’s spikes separately for each image.Batch-wise averaging enables matrix operations that speed computation of network updates.
- Training procedure: Negative inhibitory connection values are reset to zero after each update, preventing those connections from becoming excitatory.This preserves the model’s inhibitory-connection convention during training.
- Learning-rate schedule: The relative learning rates are chosen so β is much smaller than α or γ, maintaining sparse and uncorrelated activity as feed-forward weights change.The rates are selected based on the requirement that network activity remain sparse and uncorrelated.
Fitting SAILnet RFs to Gabor functions
The study fits SAILnet receptive fields with two-dimensional Gabor functions and applies quality-control cuts to retain interpretable fits. These restrictions leave 299 SAILnet receptive fields for subsequent analysis, while recognizing that some receptive-field shapes are not Gabor-like.
- Gabor fitting: A Gabor receptive-field model combines a two-dimensional Gaussian envelope with a sinusoid.Amplitude, center, orientation, phase, spatial extents, and frequency are represented by the fitted parameters.
- Gabor fitting: The fitting procedure chooses Gabor parameters by unconstrained optimization that minimizes mean squared error ||G − RF||^2.The resulting parameters are used to characterize each receptive field’s shape.
- Quality control: Receptive fields with relative fit error ||G−RF||^2/||RF||^2 > 0.5 are excluded from the SAILnet analysis.This threshold imposes a mild minimum signal-to-noise requirement.
- Quality control: Fits are also excluded when the Gabor center lies outside the 16 × 16 pixel patch or within one envelope standard deviation of its edge.The restriction avoids poorly constrained or biased shape estimates caused by truncation at patch boundaries.
- Quality control: 299 SAILnet receptive fields remain after the quality-control cuts, whereas 116 of 250 macaque receptive fields pass the gentler 0.8 fit-error threshold.The macaque threshold is gentler because measured macaque fields contain sizable zero-support regions and noise.
- Scope of the fit: Gabor functions cannot accurately describe every receptive-field shape, including center-surround fields better represented by differences of Gaussians.The authors leave selection of a broader receptive-field function family for future work.
Competing financial interests
SAILnet uses spiking neurons and synaptically local learning to represent whitened natural-image patches, producing diverse V1-like receptive fields and several experimentally relevant response properties.
- Network and neuron model: SAILnet uses leaky integrate-and-fire neurons whose membrane-like internal variable integrates input current, spikes at threshold, and resets after firing.The model represents membrane dynamics with an RC-circuit analogy and simulates the voltage equation in discrete time.
- Image representation: A linear decoder reconstructs whitened image patches from SAILnet activity, despite representing each 256-parameter patch with 75 binary spikes on average.The reconstruction resembles the original but is not identical because of the severe compression ratio.
- Receptive fields: SAILnet receptive fields match the diversity of macaque V1 simple-cell fields, including unoriented features, Gabor-like wavelets, and elongated edge detectors.Gabor-fit parameters show that SAILnet receptive fields span the macaque data space and account for the observed shapes.
- Network statistics: Pairwise spike-count correlations are near zero, while inhibitory connectivity is stronger between neurons with significantly overlapping receptive fields.Cells with negatively overlapping receptive fields tend not to have inhibitory connections.