Source-linked AI summary

Neural Encoding and Decoding with Deep Learning for Dynamic Natural Vision

Haiguang Wen, Junxing Shi, Yizhen Zhang, Kun-Han Lu, Jiayue Cao, Zhongming Liu

arXiv:1608.03425v2q-bio.NCq-bio.QM

TL;DR

The paper addresses how the brain represents dynamic natural vision and whether activity can be decoded to reconstruct and categorize what a person sees. It applies CNN-based encoding and decoding models to fMRI responses during natural movies, finding coverage across visual cortex and support for voxel visualization, visual reconstruction, and semantic decoding.

  • Problem

    Existing neural encoding and decoding studies focused mainly on static or artificial stimuli, limiting evidence about representations underlying dynamic natural vision.

  • Method

    The study uses CNN-based encoding and decoding models to relate fMRI responses during natural movies to visual and semantic feature representations.

  • Results

    The CNN explained fMRI responses across nearly the entire visual cortex, visualized voxel representations and category selectivity, and supported direct visual reconstruction of natural movies.

  • Takeaways & Limitations

    The findings support using deep learning as a computable model for understanding and decoding cortical representations during dynamic natural vision.

  • Takeaways & Limitations

    Because fMRI resolution limited the decoded activity, visual reconstructions were blurry, lacked texture and color details, and may poorly resolve cluttered scenes.

Abstract

from arXiv · show

Convolutional neural network (CNN) driven by image recognition has been shown to be able to explain cortical responses to static pictures at ventral-stream areas. Here, we further showed that such CNN could reliably predict and decode functional magnetic resonance imaging data from humans watching natural movies, despite its lack of any mechanism to account for temporal dynamics or feedback processing. Using separate data, encoding and decoding models were developed and evaluated for describing the bi-directional relationships be-tween the CNN and the brain. Through the encoding models, the CNN-predicted areas covered not only the ventral stream, but also the dorsal stream, albe-it to a lesser degree; single-voxel response was visualized as the specific pixel pattern that drove the response, revealing the distinct representation of individual cortical location; cortical activation was synthesized from natural images with high-throughput to map category representation, con-trast, and selectivity. Through the decoding models, fMRI signals were directly decoded to estimate the feature representations in both visual and semantic spaces, for direct visual reconstruction and seman-tic categorization, respectively. These results cor-roborate, generalize, and extend previous findings, and highlight the value of using deep learning, as an all-in-one model of the visual cortex, to understand and decode natural vision.

Introduction

The study asks how the brain represents dynamic visual information and whether activity can be decoded to reconstruct and categorize what a person sees. It addresses these questions with deep-learning models of distributed cortical activity and reports encoding and decoding results across visual and semantic representations.

  • Research questions: The study investigates dynamic visual representation and direct decoding of brain activity for reconstructing and categorizing viewed content.It seeks an alternative strategy capable of embracing the complexity of natural vision and distributed cortical activity.
  • Approach: Deep learning provides hierarchical feature representations spanning low-level visual properties through high-level objects and actions.These representations motivate using a deep convolutional neural network to model cortical visual processing.
  • Dataset: The study acquired 11.5 hours of fMRI data from each of three subjects watching 972 diverse video clips.The dataset was independent of, larger than, and more broadly covered than datasets in prior studies.
  • Encoding results: CNN-based encoding explained significant fMRI variance across nearly the entire visual cortex, including ventral and dorsal streams, with weaker effects dorsally.The models were trained and tested with distinct data to describe relationships between cortical activity and CNN representations.
  • Encoding results: Voxel-wise encoding visualized distinct single-voxel representations and revealed category representation and selectivity.These analyses used CNN-based voxel-wise encoding models to characterize representations at individual cortical locations.
  • Decoding results: CNN-based decoding supported direct visual reconstruction of natural movies and direct semantic categorization.Reconstructions highlighted foreground objects but had blurry details and missing colors, while categorization used the CNN’s embedded semantic space.

Materials and Methods

The study used repeated natural-movie viewing in three healthy volunteers, combined with 3T fMRI acquisition and CNN-based visual feature extraction. AlexNet features, a 15-category classifier, and deconvolutional projections supported voxel-wise encoding and visualization analyses.

  • Participants and stimuli: Three healthy volunteers with normal vision watched diverse natural-movie clips, with separate 2.4-hour training and 40-minute testing movies presented repeatedly.The training movie was watched twice and the testing movie ten times; testing clips differed from training clips.
  • Data Acquisition and Preprocessing: Functional MRI data were acquired at 3T with 3.5 mm isotropic spatial resolution and 2 s temporal resolution.The acquisition used gradient-recalled echo-planar imaging with 38 interleaved axial slices.
  • Convolutional Neural Network (CNN): The pretrained eight-layer AlexNet extracted hierarchical visual features from every movie frame, producing activation time series for individual CNN units.AlexNet comprised five convolutional layers followed by three fully connected layers.
  • Convolutional Neural Network (CNN): The original output categories were reduced to 15 movie-relevant classes, and a new softmax classifier was trained while retaining all lower CNN layers.The classifier used approximately 20,500 labeled training images and approximately 3,500 separate testing images.
  • Deconvolutional neural network (De-CNN): A deconvolutional neural network approximately reversed CNN operations through top-down projections from selected units toward input pixel space.The procedure involved unpooling, rectification, and filtering onto lower layers until reaching the input representation.
  • Encoding models: Voxel-wise encoding models related visual inputs to voxel responses and optimized pixel-level representations, using L1 regularization to favor sparse visual-feature coding.The visualized representation was defined as an optimal gradient pattern reflecting pixel-wise influence on a voxel’s response.

Results

CNN representations aligned with hierarchical and categorical organization across ventral and dorsal visual pathways, while encoding and decoding models predicted, visualized, and reconstructed natural-movie-related cortical activity. Decoding captured salient object structure and motion, and fMRI responses showed significant cross-subject reproducibility.

  • Functional representations: Voxel response visualizations revealed location- and function-specific representations, including retinotopy, foreground objects, motion or action, body parts, facial features, scenes, and houses.FFA voxels responded selectively to human and animal faces, whereas PPA voxels represented backgrounds, scenes, or houses.
  • CNN–brain correspondence: CNN layers mapped progressively higher-level features onto cortical areas from striate to extrastriate cortex across both ventral and dorsal streams, consistently across subjects.Low- to high-level representations were associated with progressively abstract visual information along the visual pathways.
  • CNN–brain correspondence: CNN categorical units correlated with corresponding high-order visual areas, including face-related representations in FFA, OFA, and pSTS-FA, with stronger correlations in the right hemisphere.Face-unit and FFA responses also peaked on movie frames containing human faces; related representations included indoor scenes, land animals, cars, and birds.
  • Encoding models: Encoding models predicted partially but significantly widespread fMRI responses to unseen natural movies, with ROI prediction accuracy ranging from 0.4 to 0.6 in visual cortex.The evaluation used five 8-minute testing movies containing clips unseen during model training, while a control model estimated potentially explainable variance.
  • Decoding models: Decoded first-layer feature maps correlated with actual CNN feature maps at r=0.30±0.04, enabling reconstructions that captured salient objects’ location, shape, and motion but missed color.Less salient objects and backgrounds were poorly reproduced.
  • Cross-subject generalization: Inter-subject fMRI responses were significantly reproducible for 82% of locations within visual cortex, supporting cross-subject neural encoding and decoding.Reproducibility was significant at p<0.01 using a t-test with Bonferroni correction.

Discussion

The study extends CNN-based modeling to widespread cortical responses during natural movie viewing, using encoding and decoding models to link brain activity with visual and semantic representations. These models support voxel-level visualization, high-throughput analysis of natural vision, direct movie reconstruction, and semantic categorization.

  • Overall contribution: CNN-based encoding and decoding models generalized deep-learning analyses to widespread fMRI responses during naturalistic movie viewing and established a feedforward account across visual-cortex processing levels.The models were trained with hours of fMRI data collected during movie viewing.
  • Cortical coverage: Model-predictable voxels covered nearly the entire visual cortex, including dorsal-stream activity despite the CNN’s lack of recurrent or feedback connections.The broader coverage exceeded predictions from Gabor or motion filters, manually defined categorical features, and models trained on limited static pictures.
  • Voxel representations: Voxel-wise encoding visualized the pixel patterns driving individual cortical responses and revealed increasingly complex, category-selective, and complementary representations across downstream locations.The method illustrated distinct representations at locations such as FFA and PPA.
  • High-throughput analysis: Encoding models generalized to novel visual stimuli, enabling high-throughput prediction and analysis of cortical responses to many natural pictures or videos beyond practical fMRI scanning.The trained models function as a computational workbench for studying neural representations of natural vision.
  • Decoding: Decoding directly reconstructed natural movies without candidate-image comparison and estimated CNN visual and semantic representations for visual reconstruction and semantic categorization.The method used features learned from natural images rather than a predefined candidate-image set; reconstruction in this study relied on decoded 1st-layer features.
  • Semantic interpretation: Semantic decoding translated CNN representations into human-defined categorical labels through a separate classifier, enabling textual interpretation of brain activity without redefining the semantic space.The CNN semantic space emerges progressively from lower-level visual features and supports object recognition across category definitions.
Loading 1608.03425v2…