Source-linked AI summary

Recurrence is required to capture the representational dynamics of the human visual system

Tim C Kietzmann, Courtney J Spoerer, Lynn Sörensen, Radoslaw M Cichy, Olaf Hauk, Nikolaus Kriegeskorte

arXiv:1903.05946v2q-bio.NCcs.CVcs.LG

TL;DR

The paper addresses limited understanding of rapid, time-varying computations in the human ventral stream, which are often studied with feedforward models. It measures and models multi-region representational dynamics using time-resolved MEG, Granger causality, and matched feedforward and recurrent neural networks. Recurrent models better captured the cortical dynamics, while cooling experiments supported roles for lateral and top-down connections.

  • Problem

    Rapid representational computations across ventral-stream regions are not well understood because prior analyses commonly rely on time-averaged data and feedforward models.

  • Method

    The study models MEG representational-dynamics movies across three ventral-stream regions with parameter-matched feedforward and recurrent networks, alongside Granger-causality and virtual-cooling analyses.

  • Results

    Recurrent networks significantly outperformed ramping feedforward models, matching trajectory correlations of 0.95, 0.93, and 0.97 across V1-3, V4t/LO, and IT/PHC.

  • Takeaways & Limitations

    The analyses consistently indicate that human ventral-stream dynamics arise from recurrent message passing involving lateral and top-down connections.

  • Takeaways & Limitations

    The Granger-causality model did not include common input to source and target regions.

Abstract

from arXiv · show

The human visual system is an intricate network of brain regions that enables us to recognize the world around us. Despite its abundant lateral and feedback connections, object processing is commonly viewed and studied as a feedforward process. Here, we measure and model the rapid representational dynamics across multiple stages of the human ventral stream using time-resolved brain imaging and deep learning. We observe substantial representational transformations during the first 300 ms of processing within and across ventral-stream regions. Categorical divisions emerge in sequence, cascading forward and in reverse across regions, and Granger causality analysis suggests bidirectional information flow between regions. Finally, recurrent deep neural network models clearly outperform parameter-matched feedforward models in terms of their ability to capture the multi-region cortical dynamics. Targeted virtual cooling experiments on the recurrent deep network models further substantiate the importance of their lateral and top-down connections. These results establish that recurrent models are required to understand information processing in the human ventral stream.

Significance Statement

Time-resolved analyses reveal dynamic, recurrent transformations within and across ventral-stream regions, and recurrent neural networks better capture these dynamics than matched feedforward models.

  • Representational dynamics: Animacy peaks first in IT/PHC around 160 ms, then in V4t/LO around 260 ms, while resurfacing less strongly in IT/PHC.The effect is absent around 200 ms before reappearing in V4t/LO.
  • Representational dynamics: Representational distinctions emerge sequentially within regions and cascade forward and backward across ventral-stream areas.Low-level features, faces, and animacy show staggered temporal profiles, including a reverse animacy cascade.
  • Interregional information flow: Granger causality indicates significant feedforward and feedback information flow between successive ventral-stream regions.Feedforward effects emerge around 70 ms, whereas feedback effects emerge more gradually and peak at region-specific later latencies.
  • Neural network modelling: Recurrent DNNs matched ventral-stream average-distance trajectories with correlations of 0.95, 0.93, and 0.97 for V1-3, V4t/LO, and IT/PHC, respectively.They significantly outperformed ramping feedforward models for all ROIs and cross-validation splits at p<0.0001.
  • Recurrent connectivity: Cooling experiments show that lateral and top-down connections contribute to object categorization and ventral-stream dynamic predictions.Lower-layer performance was especially affected by cooling top-down connections, while later layers were more robust to top-down cooling.

MEG data acquisition, pre-processing and source reconstruction

MEG recordings captured visual responses from 16 participants, with source reconstruction analyzed for 15 participants across three ventral-stream ROI groups. The experiment used 92 object stimuli and multimodal anatomical procedures to define spatially distinct regions.

  • Data collection: MEG data were recorded from 16 right-handed participants, with source reconstruction performed for 15 participants who also underwent structural and functional MRI.Participants had normal or corrected-to-normal vision and provided informed consent.
  • Stimuli: Participants viewed 92 different objects, enabling comparisons across modalities, species, and recording sites.The stimulus set had also been used for human fMRI, perceptual judgments, macaque recordings, and deep-network studies.
  • Experimental design: Each MEG session presented objects briefly on a grey background while participants detected randomly occurring paper-clip targets; target trials were excluded.Stimuli lasted 500 ms, with trial onset asynchronies of 1.5 or 2 s.
  • Pre-processing: MEG signals were acquired from 306 channels at 1 kHz, filtered, cleaned with spatiotemporal methods, downsampled to 500 Hz, and baseline-corrected.The preprocessing pipeline included bandpass filtering from 0.03 to 330 Hz.
  • Source reconstruction: ROI masks were defined on individual cortical surfaces and projected into functional volumes using FreeSurfer.The anatomical source space used participant-specific structural T1 scans and boundary-element models.
  • Source reconstruction: Three spatially distinct ROI groups covered early V1-3, intermediate V4t/LO1-3, and downstream IT/PHC visual areas.The ROIs were made comparatively large to maximize signal-to-noise ratio while limiting cross-talk.

MEG Representational dynamics analysis

Time-resolved representational similarity analysis converted multivariate activity into RDM movies, allowing representational transformations and inter-ROI information transfer to be modeled across time. Hierarchical GLMs and Granger causality analyses quantified model contributions and directional predictive relationships.

  • Representational similarity analysis: RDM movies represented time-varying distances between all pairs of experimental conditions in each ROI.Distances used correlation distance, defined as 1-Pearson correlation, producing one RDM at each time point.
  • RDM modelling: Hierarchical GLMs combined external computational and categorical predictors to explain observed condition-specific representational distances.Predictors could be non-orthogonal, so their contributions were assessed within the combined model.
  • RDM modelling: Four main and 10 control predictors were standardized before GLM entry, with main predictors including animate, face, low-level GIST, and real-world-size geometries.Nonnegative least squares estimated the weights for the linear combination.
  • Statistical analysis: Unique variance was tested against pre-stimulus baseline increases, with a non-parametric cluster test controlling multiple comparisons across time.The first 600 ms of stimulus processing were included, while statistical comparisons used unsmoothed signals.
  • Inter-ROI information transfer: Granger causality tested whether past RDMs from a source ROI improved prediction of a target ROI beyond the target ROI’s own past.The analysis used a hierarchical GLM with nonnegative least-squares fitting.
  • Noise ceiling: Signal-noise ceilings were estimated per ROI and time point using leave-one-participant-out lower bounds and group-average upper bounds.The upper bound was overfit because each participant contributed to the group average used for prediction.

FMRI data acquisition and analyses

fMRI data were collected across repeated runs and analyzed in ROIs aligned with the MEG regions. Activation-pattern distances formed fMRI RDMs, which were predicted from recurrent-network representations using cross-validated temporal combinations.

  • Data acquisition: fMRI data were collected from 15 participants, who completed 10 to 14 runs each with stimulus and null trials.During null trials, participants reported brief fixation-cross luminance changes by button press.
  • ROI representational geometry: fMRI ROIs were aligned with the MEG ROIs, and stimulus-pattern distances were computed using 1-Pearson correlation.Activation patterns were extracted for all possible stimulus pairs using t-values.
  • Network prediction: MEG-fitted recurrent convolutional networks predicted temporally smooth fMRI representational similarities by linearly combining time points from ROI-matched network layers.Layer selection matched each fMRI ROI to its corresponding MEG ROI.
  • Cross-validation: Cross-validated nonnegative least squares fit time-point weights on N-1 participants and evaluated predictions on the left-out participant.Prediction accuracy was measured by correlating the upper triangles of the predicted and observed RDMs.

Neural network models

The study compared parameter-matched feedforward and recurrent convolutional networks trained to predict time-varying ventral-stream representations. Networks were evaluated on held-out stimuli and held-out MEG data to assess their ability to capture human representational dynamics.

  • Architectures: Feedforward bottom-up networks were compared with recurrent BLT networks containing bottom-up, lateral, and top-down connections.The architectures had approximately the same number of parameters.
  • Architectures: Feedforward networks were allowed to ramp up unit activity over time so they could exhibit non-trivial dynamics.This design provided a dynamic feedforward comparison rather than only a static feedforward baseline.
  • Training: Networks used representational distance learning to predict ventral-stream representational dynamics up to 300 ms after stimulus onset.Training used 141,000 natural images spanning 61 categories derived from the experimental stimulus set.
  • Training: Training images were augmented through random cropping, vertical flipping, brightness, saturation, and contrast changes, then resized to 96×96 pixels.Crops covered at least one third of the image area and used aspect ratios from 0.9 to 1.1.
  • Evaluation: To reduce overfitting, evaluation used the 92 experimental images, which were independent and visually dissimilar from the natural training images.Networks were also tested on MEG data held out from model fitting using two-fold cross-validation.

Architectural overview

The networks use six convolutional layers followed by a linear readout, with pooling reducing spatial dimensions between layers. Feedforward and recurrent architectures are approximately parameter-matched, using larger kernels in feedforward models to compensate for recurrent connections.

  • Architectural overview: Six convolutional layers precede a linear readout, with 1×1 strides and padding that preserves each convolution’s spatial dimensions.Before each convolution except the first, 2×2 max pooling reduces height and width by half.
  • Architectural overview: Feedforward B and recurrent BLT networks retain the same number of units and layers while approximately matching parameter counts.Because lateral and top-down connections increase BLT parameters, B uses larger kernels for compensation.
  • Architectural overview: 3.0 million parameters occur in BK9, 4.3 million in BK11, and 4.0 million in BLT.BK9 and BK11 are the two closest feedforward parameter matches to BLT.

Recurrent convolutional layers

Recurrent convolutional layers represent activations over time and combine bottom-up, lateral, and top-down inputs. Feedforward layers remove recurrence but can retain temporal ramping through nonnegative self-connections.

  • Recurrent convolutional layers: Each recurrent convolutional layer represents a three-dimensional activation array indexed by time step and layer, with height, width, and feature dimensions.The network input is defined as the activation at the initial layer.
  • Recurrent convolutional layers: Conventional feedforward networks reduce recurrent convolutional layers to standard convolutional layers by removing recurrent connections.Their units are inactive before feedforward input arrives at the corresponding layer.
  • Recurrent convolutional layers: Feedforward B layers use shared nonnegative self-connections to let unit activity ramp up over time.Setting the self-connection parameter ω_n to zero recovers conventional feedforward models.
  • Recurrent convolutional layers: BLT layers add lateral and top-down convolutions to the bottom-up computation.Top-down outputs are nearest-neighbor up-sampled so their spatial dimensions match bottom-up and lateral outputs.
  • Recurrent convolutional layers: In the final BLT layer, top-down input comes from the readout through a fully connected connection rather than a convolution.Elsewhere, top-down connections use convolutions.

Readout layer

The readout converts final-layer features into category outputs after global average pooling and receives recurrent input from its previous state. This recurrent input allows categorization responses to persist without continuous bottom-up input.

  • Readout layer: A linear readout produces one output for each category on which the network is trained.The readout follows the network’s convolutional layers.
  • Readout layer: Global average pooling averages the final layer over spatial dimensions, producing a vector whose length equals the final layer’s number of features.This pooled vector is supplied to the readout.
  • Readout layer: The readout receives lateral input from its previous time step, allowing categorization responses to continue without continuous bottom-up input.In B networks this input uses self-connections; in BLT networks readout units use fully connected lateral weights.

Training

The networks were trained with two objectives: learning representational distances and classifying objects.

  • Training: The networks were trained using representational distance learning and object classification.

Representational distance learning

Representational distance learning trains selected network layers to match time-resolved representational dynamics across three ventral-stream regions while retaining category classification. The objective combines representational, categorization, and L2-regularization terms, with time-variance normalization and scheduled loss weighting.

  • Representational distance learning: RDL matches representational dynamics across network layers 2, 4, and 6 to V1-3, V4t/LO, and IT/PHC, respectively.The training images came from the RDL61 set, whose categorical structure matches the experimental stimuli.
  • Representational distance learning: Distances are averaged within stimulus categories, producing a 61×61 representational dissimilarity matrix for each time point.For example, distances among 12 face images are averaged into one category-level estimate.
  • Representational distance learning: Mini-batches contain pairs of images from different categories, whose layerwise correlation distances are compared with the corresponding empirical RDM distances.
  • Representational distance learning: RDM error is normalized by empirical RDM variance at each time step so that high-variance time points do not dominate optimization.This makes each time point contribute independently of RDM variance or noise level.
  • Representational distance learning: The overall loss combines RDL and categorization objectives with L2 regularization, weighting their contributions through separate coefficients.The regularization coefficient is λ = 10⁻⁶, and the objective averages the component losses over mini-batches.
  • Representational distance learning: The categorization coefficient starts at 10 and decays tenfold every 10,000 mini-batches until reaching 10⁻², while Adam performs optimization.Training terminates after 4 million mini-batches; Adam uses learning rate α = 10⁻⁴.

Virtual cooling studies

Virtual cooling is implemented by independently dropping lateral and top-down convolution outputs at different keep probabilities. The resulting activations are evaluated for both human ventral-stream prediction and object categorization.

  • Virtual cooling studies: Dropout at different keep probabilities targets lateral and top-down connections to emulate cortical cooling studies.
  • Virtual cooling studies: The cooled networks are tested for predicting human ventral-stream representational dynamics and performing object categorization.Mean network or layer activity is corrected before evaluation.

Model fitting for off-the-shelf architectures

Off-the-shelf AlexNet and VGG16 feedforward networks are evaluated as candidate models by selecting layers that best predict MEG data. The selected layers differ across ventral-stream regions and validation splits.

  • Model fitting for off-the-shelf architectures: Feedforward DNNs provide static layer outputs but can still be informative for predicting dynamic neural responses.
  • Model fitting for off-the-shelf architectures: AlexNet and VGG16 are probed with the RDL61 image set, using activation vectors from candidate layers to select the best layer for each ROI.
  • Model fitting for off-the-shelf architectures: AlexNet selects L5, L2, and L2 for V1-3, V4t/LO, and IT, while VGG16 selections vary across validation splits for V1-3 and use deeper layers for V4t/LO and IT/PHC.For VGG16, layers 5 and 12 are selected for the two V1-3 splits, and layer 13 for V4t/LO and IT/PHC.
Loading 1903.05946v2…