Source-linked AI summary

Predictive information in a sensory population

Stephanie E. Palmer, Olivier Marre, Michael J. Berry, William Bialek

arXiv:1307.0225v1q-bio.NC

TL;DR

The paper asks how early sensory populations can represent information about future inputs efficiently. It measures predictive information in retinal ganglion-cell groups, compares it with the limit imposed by input statistics, and examines downstream compression. Every observed cell participates in a near-bound group, while predictor neurons can compress retinal predictive information into output spikes or silence and show feature selectivity.

  • Problem

    The paper addresses how neural populations can retain information about future sensory inputs while compressing information about the past.

  • Method

    The study measures information between retinal activity and future visual inputs, compares neural representations with an information-bottleneck bound, and tests downstream predictor neurons encoding predictions in one output bit.

  • Results

    Every monitored retinal cell participates in a group operating close to the predictive-information bound, and downstream predictor neurons can encode compressed retinal predictions while exhibiting feature selectivity.

  • Takeaways & Limitations

    Efficient representation of predictive information is proposed as a candidate principle for neural computation at successive stages of processing.

Abstract

from arXiv · show

Guiding behavior requires the brain to make predictions about future sensory inputs. Here we show that efficient predictive computation starts at the earliest stages of the visual system. We estimate how much information groups of retinal ganglion cells carry about the future state of their visual inputs, and show that every cell we can observe participates in a group of cells for which this predictive information is close to the physical limit set by the statistical structure of the inputs themselves. Groups of cells in the retina also carry information about the future state of their own activity, and we show that this information can be compressed further and encoded by downstream predictor neurons, which then exhibit interesting feature selectivity. Efficient representation of predictive information is a candidate principle that can be applied at each stage of neural computation.

I. INTRODUCTION

Prediction is a general computational problem because sensory data guide behavior only insofar as they reveal future states. The paper asks whether sensory systems retain limited past information that is maximally predictive of the future.

  • Sensory information guides actions by providing information about the future state of the world.
  • Prediction spans tasks from extrapolating moving-object trajectories to learning abstract rules about unfolding events.
  • Because representing and transmitting information has costs, sensory codes may preserve only past features that are maximally informative about the future.
  • The paper also considers whether successive neural-processing stages predict future patterns of neural activity.

II. CODING FOR THE POSITION OF A SINGLE VISUAL OBJECT

The study measures how retinal ganglion-cell activity represents the past and future position of a moving bar. Responses are encoded as binary population words, whose information varies with temporal delay and group size.

  • A moving bar with partially predictable trajectories provides the visual stimulus for recordings from salamander retinal ganglion cells.
  • Within 1/60-second windows, each ganglion cell is represented as spiking or silent, and N cells form a binary word wt.
  • The analysis estimates mutual information between neural words and the bar’s position at past or future times, using the distributions of words, positions, and conditional positions.
  • Information is normalized by the mean spikes generated by each cell group and averaged across groups in bits per spike.
  • At approximately 80 ms in the past, retinal responses are most informative and larger groups show declining information per spike from redundancy.
  • Predictive information extends into the future, with decreasing redundancy and hints of synergistic coding far ahead.

III. BOUNDS ON PREDICTABILITY

Predictive information is limited by the statistical structure and stochasticity of sensory inputs, creating an information-bottleneck bound. Efficient representations retain the minimum past information needed to achieve a chosen predictive power.

  • Unobserved causal factors make sensory evolution irreducibly stochastic, so even a complete record of the past cannot yield perfect predictions.
  • Predictive information Ipred(T) = I(Xpast; Xfuture) is the finite number of bits that the past provides about the future.
  • A compressed representation Z contains past information Ipast = I(Z; Xpast) and predictive information Ifuture = I(Z; Xfuture).
  • For a target Ifuture, the representation must capture at least I∗past(Ifuture) bits about the past; conversely, limited Ipast imposes a maximum Ifuture.
  • The information-bottleneck plane separates accessible from impossible combinations of past and future information, and efficiency means approaching its boundary.
  • For the modeled trajectories, exact position and velocity specify the future completely but require infinite information, motivating error-limited representations.

IV. DIRECT MEASURES OF PREDICTIVE INFORMATION

The paper directly measures how retinal ganglion-cell responses predict future sensory inputs and compares that information with a theoretical efficiency bound. Measured groups operate close to this bound, including groups containing every observed cell.

  • Direct measurement: The authors measure predictive information by generating independent stimulus trajectories that converge onto common futures.They synthesize one hundred independent pasts for each of thirty futures.
  • Direct measurement: Responses are initially independent of the common future, then become modulated by features that predict it as convergence approaches.This pattern is visible in single-cell spike probabilities and supports estimating information for groups of 1–7 neurons.
  • Efficiency bound: 0.11 bits about the past defines a theoretical maximum of 0.097 bits about the future for the five-cell group analyzed.The group captured 0.78 bits/spike, and its predictive ratio was 0.98±0.39, within error bars of optimality.
  • Efficiency bound: Predictive power decays with forecasting distance in a way that follows the theoretical sensory-input limit within error bars.The comparison is made for the five-cell group as the prediction begins progressively farther ahead of the current time.
  • Population result: Every one of the 53 monitored neurons participates in at least one group operating close to the predictive-information bound.This remains true for larger groups until the finite dataset prevents effective sampling of the relevant distributions.

V. PREDICTING THE FUTURE STATE OF THE RETINA

The authors also ask how well retinal activity predicts its own future, using model-independent conditional distributions of neural response words. Predictive information is stronger and longer-lasting for naturalistic stimuli and larger cell groups than for simpler or random stimuli.

  • Retinal prediction: Predicting future visual input from the brain’s perspective is equivalent to predicting future retinal output.The retina is the only source of visual-stimulus information available to the central nervous system.
  • Retinal prediction: The conditional distribution P(w_t+∆t|w_t) provides a model-independent measure of how current retinal activity predicts future activity.The approach does not depend strongly on the complexity of the visual inputs.
  • Population scaling: Larger cell groups carry predictive information for hundreds of milliseconds, with maximum predictive information averaging above 1 bit/spike across sampled groups.Smaller groups lack long-term predictive power and provide roughly half the short-term information per spike of larger groups.
  • Stimulus dependence: Naturalistic movies produce the strongest and longest-ranged predictions, single-object motion produces intermediate results, and random checkerboards lose predictability within a few frames.The comparison links retinal predictability to the statistical structure of the sensory inputs.
  • Stimulus dependence: These results suggest that predicting future retinal activity can highlight activity patterns that are especially informative about the visual world.The paper states that this possibility is borne out by subsequent analysis.

VI. PREDICTOR NEURONS?

Downstream predictor neurons compress the predictive information in retinal population activity into a single spike-or-silence output. These stimulus-independent optimizations preserve predictive content while producing biologically plausible computations and revealing motion-selective features.

  • Compression into one bit: A single output bit can preserve almost all predictive information carried by four retinal ganglion cells.The output is optimized by grouping input words according to how predictive they are of future population activity.
  • Compression into one bit: The optimal grouping generalizes across cell groups and is well approximated by a perceptron thresholding a weighted sum of inputs.Predictor neurons need not fire at anomalously high rates, suggesting biological realizability.
  • Emergent stimulus selectivity: Stimulus-independent predictor neurons also carry information about the visual input, despite being optimized only to predict future retinal activity.Their visual information increases with the predictive information they capture.
  • Emergent stimulus selectivity: Predictive optimization selects motion features such as constant speed and long epochs of constant velocity, while refining position estimates relative to individual inputs.Optimizing for farther-future prediction shifts the time of sharpest stimulus discrimination closer to the downstream spike, compensating for latency.
  • Emergent stimulus selectivity: Motion estimation is efficient because, in an inertial world, it represents the future state of the visual scene.The result connects the emergence of motion computation to predictive information rather than motion representation alone.

VII. DISCUSSION

The discussion frames predictive information as an alternative to redundancy reduction or total-information maximization for understanding neural coding. The authors argue that this principle can extend from retinal coding to broader neural computation and prediction problems.

  • Predictive information as a coding principle: Predictive-information optimization differs from classical principles based on reducing redundancy or maximizing total information transmission.The paper suggests these candidate principles can be distinguished experimentally.
  • Broader scope: The authors propose efficient predictive representation as a principle applicable at every layer of neural processing, not only retinal visual coding.They present prediction as a unified problem spanning trajectory extrapolation, rewards, action outcomes, and learned rules.
  • Predictive information as a coding principle: Figure 5 evaluates how efficiently a single output neuron captures the predictive information of four-cell retinal groups and compares all mappings with perceptron rules.Panel b summarizes 150 groups; the average maximum efficiency is y = .82, against y = 1 for perfect capture.
  • Predictive information as a coding principle: The optimized predictor neurons generate reasonable output firing rates and have computational structures that can be learned by biologically plausible rules.Their prediction rules are found without reference to the stimulus, yet the neurons also efficiently transmit sensory information.
  • Predictive information as a coding principle: Predictive optimization lets the nervous system identify features of a retina’s combinatorial code that are informative about the visual world without external calibration.The paper links this result to downstream feature selectivity emerging from prediction.

Methods

The study recorded retinal activity under controlled visual stimuli, generated stochastic moving-bar trajectories, and estimated predictive information using compression and extrapolation methods.

  • Data collection: Retinal voltages were recorded from larval tiger salamander retina using dense 252-electrode arrays.The tissue was perfused while monitor images were projected onto the photoreceptor layer.
  • Stimulus presentation: Naturalistic and moving-bar movies were refreshed at 60 fps, while randomly flickering checkerboards were refreshed at 30 fps.The moving bar was 11 pixels wide and black against a grey background.
  • Stimulus generation: The moving-bar trajectory followed a stochastic spring-bound Brownian-motion process with specified position and velocity updates.The dynamics were slightly overdamped, and the time step matched the display refresh interval.
  • Stimulus generation: Common-future trajectories were constructed by joining multiple distinct pasts to a shared future from matching endpoint positions.The procedure began with a 10-million-step trajectory and selected 52-step segments.
  • Information estimation: Mutual information estimates were extrapolated to infinite sample size from bootstrap subsamples, with shuffled-data checks for reliability.Reliable estimates required shuffled information within error or 0.02 bits/spike of zero.
  • Information bottleneck: The information-bottleneck analysis optimized compression of past activity while retaining information about future stimuli.The bottleneck problem defines the optimal predictive-information tradeoff for each compression amount.

Supplementary Information

Supplementary analyses connect predictive information to stimulus coding and reveal feature selectivity in optimized predictor neurons.

  • Stimulus coding: Capturing more predictive information in four-cell retinal activity enabled downstream binary rules to convey greater stimulus information.The analysis supports the conclusion that better local predictions lead to better stimulus coding.
Loading 1307.0225v1…