Source-linked AI summary

Searching for collective behavior in a network of real neurons

Gašper Tkačik, Olivier Marre, Dario Amodei, Elad Schneidman, William Bialek, Michael J Berry

arXiv:1306.3061v1q-bio.NCcond-mat.stat-mechphysics.bio-ph

TL;DR

The paper asks whether minimally structured maximum entropy models can describe collective spiking states in large retinal populations. It applies these models to repeated natural movies, testing increasingly rich constraints on neural activity. The resulting models accurately describe the network, with pairwise interactions supplemented by population synchrony for larger groups and with consequences for entropy, collective modes, inhomogeneity, and predictability.

  • Problem

    Large neural networks have many possible spiking states, making it unclear whether limited measurements can support accurate, minimally structured models of their collective activity.

  • Method

    The authors construct maximum entropy distributions constrained by selected neural statistics, extending pairwise models with a global synchrony constraint when needed.

  • Results

    The models accurately describe whole-network states: pairwise models break down at N ≥40, whereas K-pairwise models recover good performance by matching the synchrony distribution.

  • Takeaways & Limitations

    Collective retinal activity can be represented with relatively simple data-constrained models, while retaining substantial structure in the neural state distribution.

Abstract

from arXiv · show

Maximum entropy models are the least structured probability distributions that exactly reproduce a chosen set of statistics measured in an interacting network. Here we use this principle to construct probabilistic models which describe the correlated spiking activity of populations of up to 120 neurons in the salamander retina as it responds to natural movies. Already in groups as small as 10 neurons, interactions between spikes can no longer be regarded as small perturbations in an otherwise independent system; for 40 or more neurons pairwise interactions need to be supplemented by a global interaction that controls the distribution of synchrony in the population. Here we show that such "K-pairwise" models--being systematic extensions of the previously used pairwise Ising models--provide an excellent account of the data. We explore the properties of the neural vocabulary by: 1) estimating its entropy, which constrains the population's capacity to represent visual information; 2) classifying activity patterns into a small set of metastable collective modes; 3) showing that the neural codeword ensembles are extremely inhomogenous; 4) demonstrating that the state of individual neurons is highly predictable from the rest of the population, allowing the capacity for error correction.

I. INTRODUCTION

The paper applies maximum entropy modeling to correlated spiking in large salamander-retina populations responding to natural movies. It finds that collective activity can be modeled accurately and has consequences for neural coding.

  • I. INTRODUCTION: Maximum entropy models retain only structure required to reproduce measured network statistics, avoiding hypothesized microscopic dynamics.The approach seeks minimally structured probability distributions consistent with experimental observations.
  • I. INTRODUCTION: For larger neural populations, models must address collective behavior that cannot be inferred safely from small groups of approximately 10 neurons.The introduction motivates testing population size, homogeneity, and collective effects directly.
  • I. INTRODUCTION: The study examines neural entropy, metastable collective modes, inhomogeneous codeword ensembles, and predictability of individual neurons from the population.These analyses address the structure and coding implications of the neural vocabulary.
  • I. INTRODUCTION: New multi-electrode methods enable recordings from N = 100 −200 ganglion cells in a densely interconnected retinal patch during natural movies.The experiments target large populations rather than the small groups studied previously.
  • I. INTRODUCTION: The resulting models provide an accurate description of whole-network states and show that those states are collective.The authors connect this collective structure to implications for visual information encoding.

II. MAXIMUM ENTROPY

Maximum entropy modeling constructs the least structured distribution matching selected neural statistics. The paper builds independent, pairwise, and K-pairwise models by constraining firing rates, pairwise correlations, and population synchrony.

  • II. MAXIMUM ENTROPY: Maximum entropy models match measured expectation values while otherwise imposing as little structure as possible.Their Boltzmann form is mathematically equivalent to a statistical-mechanics representation, not an assumption of thermal equilibrium.
  • II. MAXIMUM ENTROPY: The framework has 2^N possible states but aims to capture their essential structure with far fewer measured functions fµ({σi}).Choosing constraints is therefore the central modeling decision rather than a fixed activity model.
  • II. MAXIMUM ENTROPY: Independent models constrain each neuron’s mean spike probability, represented by ⟨σi⟩ and equivalent to its mean firing rate.The binary variable σi distinguishes spiking from silence in each time bin.
  • II. MAXIMUM ENTROPY: Pairwise models add constraints on correlations between every pair of neurons, producing effective couplings Jij between cells.Positive and negative couplings correspond to effective tendencies toward joint firing or separation.
  • II. MAXIMUM ENTROPY: K-pairwise models additionally constrain PN(K), the probability that K of N neurons spike synchronously in one time slice.This global synchrony distribution is independent of neuron identity and captures population-level activity.

III. CAN WE LEARN THE MODEL?

The authors learn maximum entropy models from large salamander-retina recordings and test reconstruction accuracy and overfitting. The models reproduce measured statistics, while pairwise models eventually require a global synchrony constraint for larger populations.

  • III. CAN WE LEARN THE MODEL?: 160 retinal neurons were selected from a roughly 200-cell patch recorded during repeated naturalistic movies.The experiment uses multi-electrode arrays to sample a large fraction of neurons in the patch.
  • III. CAN WE LEARN THE MODEL?: Pairwise models reproduce measured mean firing rates and correlations within experimental errors.The reconstructed coupling constants are widespread and approximately balanced between positive and negative values.
  • III. CAN WE LEARN THE MODEL?: N(N+1)/2 = 5050 measured quantities are estimated for a 100-neuron pairwise model, creating a substantial overfitting concern.The concern is whether finite-data errors across the correlation matrix could be fitted by the parameters.
  • III. CAN WE LEARN THE MODEL?: Training-versus-testing likelihoods remain consistent as network size expands from N = 10 to N = 120, with no evident overfitting.Randomly held-out movie repeats provide the testing data.

IV. DO THE MODELS WORK?

Maximum-entropy models were tested against collective spiking statistics that cannot be directly enumerated in large networks. Pairwise models capture much of the structure in small populations, while K-pairwise models extend accurate predictions to larger populations by constraining global synchrony.

  • Synchrony: ~10 neurons already show substantial departures from independent spiking, while at N = 40 synchronous firing of K = 10 neurons is ~10^4 times more probable than independence predicts.The departure from independence becomes even larger at N = 100.
  • Pairwise-model limits: At N ≥ 40, pairwise models begin to break down, including nearly threefold errors for the highly probable silent state at N = 100.Despite these errors, pairwise models remain far closer to the data than independent models.
  • K-pairwise extension: Adding the observed synchrony distribution P_N(K) defines a K-pairwise model with only ~N additional parameters beyond the ~N^2/2 pairwise parameters.The global potential V(K) regulates population activity without specifying the identities of neurons within triplets.
  • Higher-order tests: K-pairwise models correct systematic three-point-correlation errors that pairwise models make, including overestimated positive and missed negative correlations.This improvement arises even though the fitted global potential V_N(K) is numerically small.
  • Energy distributions: The model and data have similar energy distributions across more than 90% of observed states, including high-energy tails near probabilities of 10^-1.At E ≈ 25, individual states have predicted probability e^-25 ≈ 10^-11, yet the aggregate energy distribution remains quantitatively matched.

V. WHAT DO THE MODELS TEACH US?

The models reproduce mean firing rates, pairwise correlations, and the distribution of summed activity in populations of at least 100 neurons. This supports a quantitative statistical-mechanical description of the network’s distribution of spiking and silence.

  • V. WHAT DO THE MODELS TEACH US?: K-pairwise models match mean spike probabilities, pairwise correlations, and summed-activity distributions, with parameters well determined for groups of N = 100 neurons or more.Figures 7–9 indicate accurate description of the many spiking-and-silence states adopted by the network.

A. Basins of attraction

The K-pairwise energy landscape organizes retinal activity into metastable basins whose number grows rapidly with network size. These basins recur across stimulus repeats and subnetworks, indicating reproducible collective states despite variable individual responses.

  • Energy landscape: 48±2% of interacting neuron triplets are frustrated, contributing to a complex energy landscape with many local minima.Competing pairwise interactions prevent all terms from being minimized simultaneously.
  • Energy landscape: Descending observed activity patterns on the model energy function identifies locally stable metastable states and partitions all 2^N patterns into their basins of attraction.Each basin is represented by the metastable state reached through downhill single-neuron flips.
  • Scaling: Beyond N > 30, metastable states proliferate with growth faster than linear and show no sign of saturation.The counted states may underestimate the total because only basins accessible from experimentally observed patterns are included.
  • Collective modes: Patterns within the same basin are substantially more correlated than patterns assigned to different basins.The energy valleys therefore cluster activity patterns without requiring an arbitrary similarity metric.
  • Reproducibility: When the same movie is repeated, the retina revisits the same attraction basin with very high probability, while transitions between valleys take approximately 2.5Δτ.The same metastable states and corresponding valleys are identifiable across different population subsets.

B. Entropy

Entropy quantifies the effective neural vocabulary and bounds information transmission. For these correlated retinal populations, multi-information grows faster than linearly, so entropy per neuron continues decreasing through N = 120.

  • Interpretation: Maximum entropy models provide an upper bound on the true entropy, S[P({σi})] ≤ S[P({fµ})({σi})].This bound follows because the constrained maximum-entropy distribution has the least structure consistent with the measured statistics.
  • Estimation: Three entropy-estimation methods agree to better than 1% for groups of N = 120 neurons.The methods exploit Monte Carlo sampling and statistical-mechanics relationships rather than direct enumeration of all states.
  • Scaling: Multi-information initially grows quadratically with N and remains faster than linear, so entropy per neuron has not reached extensive scaling at N = 120.The statistical structure reflects pairwise interactions and the K-spike constraint.

C. Coincidences and surprises

Correlated retinal activity makes exact population responses recur far more often than independent models predict. As population size grows, coincidence probabilities approach exponential decay with a much smaller rate than the entropy per neuron suggests.

  • Scaling: −(ln Pc)/N = 0.0127 ± 0.0005 in the large-N limit, an order of magnitude below the estimated entropy per neuron.The real networks therefore approach a thermodynamic limit with unusually frequent recurrence of combinatorial patterns.
  • Inhomogeneity: Correlations make recurrence of exact spike-and-silence patterns thousands of times more likely than expected from independent neurons.The effect exceeds the reduction in entropy alone and indicates an extremely inhomogeneous state distribution.
  • Model comparison: Independent neurons predict exponential coincidence decay with N, but real networks lie orders of magnitude away from that prediction.The K-pairwise model captures the observed coincidence probabilities precisely.

D. Redundancy and predictability

Although individual-neuron responses vary across repeated natural movies, the population state predicts each neuron’s time-varying spike probability without using the visual stimulus. Prediction quality improves as more neurons are included and can reach approximately 0.8 correlation.

  • Prediction method: Conditional distributions P(σi|{σj≠i}) convert the states of other neurons into predictions of one neuron’s spike probability over time.The predicted probability can be plotted as a PSTH and evaluated against the experimental PSTH.
  • Prediction accuracy: Including more neurons makes the collective prediction increasingly resemble the real PSTH, including near-zero spike-probability epochs.Both average and single-trial prediction quality continue to improve around N ∼ 100.
  • Prediction accuracy: Correlation coefficients between predicted and experimental PSTHs reach approximately 0.8 on average, with higher values for some cells.Predictions are somewhat more variable for lower-rate cells.
  • Redundancy: Predicting individual neurons from the network rather than the visual input indicates substantial redundancy in the population representation.Downstream cells receiving only a fraction of the neurons can access the population’s full information about this visual patch.

VI. DISCUSSION

The paper models collective neural activity as a structured probability distribution and tests how well minimally structured maximum entropy models capture populations of retinal neurons. These models reveal strong, heterogeneous, and highly predictable population dynamics, while their parameters should not be interpreted as literal mechanisms or network dynamics.

  • Modeling collective behavior: Collective behavior is defined by non-factorizable structure in the distribution of population spiking states, which the paper models directly despite limited data and a very large state space.The authors treat model construction as a first step toward exploring collective activity, rather than as the final explanation of neural function.
  • Model evaluation: Pairwise models break down at N ≥40, but adding the probability of K synchronous spikes restores good performance without greatly increasing model complexity.The global synchrony constraint is well sampled and supplements pairwise correlations in the larger networks.
  • Model evaluation: Most observed states are individually rare, yet the model correctly predicts their aggregate statistical weight through cumulative energy or state-probability distributions.Agreement with experiment extends far into the low-probability tail, where rare states collectively contribute measurably.
  • Population structure: The models do not simplify through weak, sparse, low-rank, or low-dimensional interactions: correlations are full rank, interactions are widespread, and their effects are strong.The authors therefore conclude that the network is not close to a collection of non-interacting neurons.
  • Functional implications: Population activity clusters into energy-defined basins, and individual neurons are highly predictable from the others without using the visual stimulus.This supports the possibility of collective modes and error correction or pattern completion in the neural population.
  • Population structure: With N = 120 neurons, the entropy corresponds to significant occupancy of roughly one million distinct spike-and-silence combinations, showing that simple models can capture substantial distributional structure.This motivates theoretical approaches for managing the complexity of recordings from increasingly large neural populations.
  • Population structure: The neural state distribution is extremely inhomogeneous: repeated states are far more likely than under independence, while entropy remains close to the independent-neuron value.Extrapolating to roughly 250 neurons gives an estimated one-percent probability that two randomly chosen states are identical.
  • Functional implications: The maximum entropy construction is unsupervised and uses spike-train structure alone, although temporal correlations remain an additional structure that could be modeled.This makes the identified structure available in principle to a brain that receives only the retinal population activity.

Appendix A: Experimental methods

The experiments recorded tiger salamander retinal ganglion cells responding to repeated naturalistic movie clips. Activity was sampled from neuron subgroups using 20 ms binary spike-or-silence states across hundreds of stimulus repeats.

  • Electrophysiology: Recordings used tiger salamander retinal ganglion cells responding to naturalistic movie clips after the isolated retina was transferred into oxygenated Ringer’s medium.The preparation followed institutional animal-care standards and was optimized for recording stability.
  • Stimulus display: The stimulus was a 19 s grayscale movie of swimming fish and water plants, repeated 297 times at 30 frames per second.The display used standard optics and gamma correction.
  • Data preparation: Thirty subgroups were randomly selected at each size from N = 10 through 120 among 160 sorted cells, yielding 360 analyzed neuron groups.Time was discretized into 20 ms bins, and each neuron was coded as spiking or silent in each bin.
  • Data preparation: The binary state representation was σi(t) = +1 for at least one spike and −1 for silence; bins with multiple spikes comprised approximately 0.5% of time bins and were coded as spiking.The mean probability of non-silence was approximately 3.1%, and the experiment provided T = 283,041 samples per subgroup.

Appendix B: Learning maximum entropy models from data

The appendix describes learning maximum entropy models from measured constraints in large neural populations, validating the estimates, and characterizing how learned interactions change with network size.

  • Learning procedure: The modified L1-regularized learning procedure accepts arbitrary constraint functions and learns Hamiltonian parameters sequentially to greedily optimize a log-likelihood bound.The procedure extends earlier learning methods beyond single and pairwise marginals.
  • Interaction structure: As network size increases, the distribution of pairwise couplings becomes symmetric while remaining broad, and its standard deviation declines only slightly.The typical coupling does not decline as approximately 1/N, implying a thermodynamic limit unlike conventional spin-glass models.
  • Data and validation: The models were trained on 277 stimulus repeats and validated on 20 withheld repeats, using firing rates, covariances, and the k-spike distribution as constraints.The constraints were the only input to the learning algorithm.
  • Data and validation: 110·10^3 effective independent samples represented about 37% of the approximately 300·10^3 total pattern samples because samples within repeats were stimulus-dependent.The effective sample size was estimated by comparing error scaling in the original and time-shuffled data.
  • Data and validation: The largest models had fewer than 8·10^3 constrained statistics estimated from at least 15 times as many effective samples, and training-versus-testing likelihoods showed no overfitting.The overfitting check compared log likelihoods on training and testing data across 360 subgroups.

Appendix C: Exploring the energy landscape

Metastable states are identified by performing energy-lowering single-spin flips until no individual flip can further reduce the configuration’s energy.

  • Metastable-state search: Starting from an observed pattern, spins are tested in increasing index order and a flip is retained only when it lowers the energy.The final configuration is recorded as a metastable state when no spin can be flipped to reduce energy.
  • Metastable-state search: The set of metastable states can depend on the descent procedure, especially when some states are visited in different orders.The procedure’s ordering choice can therefore affect which metastable states are found.

Appendix D: Computing the entropy and partition function of the maximum entropy distributions

The appendix evaluates several entropy and partition-function estimation methods, finding strong agreement among statistical-mechanics-based approaches while identifying sample-counting limits for larger networks.

  • Estimation strategy: Entropy estimation is difficult because direct frequency counting can fail when the number of possible states is very large relative to the available samples.The appendix therefore compares multiple estimators within the maximum entropy framework.
  • Estimation strategy: For N = 10 and 20 neurons, exact enumeration agrees with heat-capacity integration, while Wang–Landau sampling also agrees with the heat-capacity method.Exact enumeration is possible because all 2^N network states can be evaluated at these sizes.
  • Partition function: K-pairwise models permit direct measurement of the partition function from the all-silent probability, followed by entropy calculation from a single Monte Carlo sampling run.The resulting estimates agree with heat-capacity integration and Wang–Landau estimates to better than 1%.
  • Sample-based estimates: For N < 50, NSB entropy estimates are reliable and largely insensitive to sample size, whereas larger networks show sample-size dependence and significant disagreement with heat-capacity integration.Model and real-data NSB estimates nevertheless agree with one another up to N = 120.

Appendix E: Are real networks in the perturbative regime?

The appendix tests whether neural correlations can be treated perturbatively and finds that this approximation increasingly fails as network size grows, consistent with collective behavior.

  • Perturbative regime: Weak pairwise correlations can still have a non-perturbative effect when they are widespread across the network.Thus, two-neuron independence can be a good approximation even when larger populations require collective modeling.
  • Perturbative regime: Perturbation theory is expected to converge at finite N in principle, but convergence may become slower for larger networks, signaling collective behavior.The appendix uses this convergence behavior to assess whether interactions remain weak.
  • Perturbative regime: For larger networks, perturbative estimates of Jij increasingly deviate from the exact maximum entropy couplings, with deviations exceeding the useful regime beyond N > 10.The comparison uses the relationship between pairwise correlations and interactions and evaluates exact versus perturbative couplings.
  • Entropy-estimation context: Figure S3 compares NSB entropy estimates from finite model samples with true entropy across network sizes and shows negligible bias for N ≤ 40 but substantial underestimation for larger networks.The figure also compares model-derived and experimental sample estimates, with network size encoded by color.
Loading 1306.3061v1…