Source-linked AI summary
Generic decoding of seen and imagined objects using hierarchical visual features
Tomoyasu Horikawa, Yukiyasu Kamitani
TL;DR
Existing fMRI decoders generally predict only categories seen during training, so this paper uses hierarchical visual features to decode arbitrary seen and imagined objects. The approach predicts features from fMRI and identifies categories beyond decoder training, including imagined objects.
Problem
Prior fMRI decoding was limited to classes included in decoder training, whereas the study sought to predict arbitrary object categories.
Method
The approach decodes hierarchical visual feature representations from fMRI signals in human visual cortex to predict object categories.
Results
Predicted features identified both seen and imagined object categories that were not used for decoder training.
Takeaways & Limitations
Hierarchical visual feature representations support brain-based retrieval of arbitrary seen and imagined object categories.
Takeaways & Limitations
The earlier decoding approach was unable to predict classes not used during training.
Abstract
from arXiv · showhide
Object recognition is a key function in both human and machine vision. While recent studies have achieved fMRI decoding of seen and imagined contents, the prediction is limited to training examples. We present a decoding approach for arbitrary objects, using the machine vision principle that an object category is represented by a set of features rendered invariant through hierarchical processing. We show that visual features including those from a convolutional neural network can be predicted from fMRI patterns and that greater accuracy is achieved for low/high-level features with lower/higher-level visual areas, respectively. Predicted features are used to identify seen/imagined object categories (extending beyond decoder training) from a set of computed features for numerous object images. Furthermore, the decoding of imagined objects reveals progressive recruitment of higher to lower visual representations. Our results demonstrate a homology between human and machine vision and its utility for brain-based information retrieval.
Introduction
The study introduces generic object decoding, predicting hierarchical visual features from fMRI to identify arbitrary seen and imagined object categories beyond decoder training. It links feature complexity to visual cortical hierarchy and finds that mid-level features are especially useful for category identification.
- Contribution: Generic object decoding predicts visual features from fMRI signals and identifies arbitrary seen and imagined object categories beyond those used for decoder training.The approach represents numerous object categories in a shared feature space, including 15,372 ImageNet categories.
- Feature decoding: Predicted feature values positively correlated with true values across feature–ROI combinations, supporting feature decoding from fMRI signals.Category-average features also extended feature decoding to seen and imagined objects.
- Hierarchical feature decoding: Higher-order features were better predicted from higher visual areas, whereas lower-order features were better predicted from lower visual areas.The interaction between feature layer and ROI was significant (ANOVA, P < 0.01).
- Imagery decoding: Imagery-induced brain activity showed progressive recruitment of hierarchical visual representations, with higher CNN layers and higher ROIs tending to peak earlier than lower ones.This temporal ordering was not found with stimulus-induced brain activity.
- Category identification: Mid-level features, especially CNN5–6, were most useful for identifying both seen and imagined object categories.Mid-level features decoded from higher ROIs were the most useful, and semantically related categories were often highly ranked when exact identification failed.
Supplementary figures … 7. Prediction of category-average features from stimulus- and imagery-induced brain activity by category-average feature decoders
The supplementary material reports distributions of correlation coefficients between predicted and category-average feature values for seen and imagined conditions.
- 7. Prediction of category-average features from stimulus- and imagery-induced brain activity by category-average feature decoders: Correlation-coefficient distributions compare predicted with category-average feature values under seen and imagined conditions.The analysis concerns predictions generated by category-average feature decoders from stimulus- and imagery-induced brain activity.
9. Time course of feature prediction from imagery-induced brain activity for individual CNN layers … 16. Identification accuracy as a function of the number of feature units
The supplied passage identifies results obtained by image feature decoders within the section on identification accuracy across feature types, layers, and ROIs.
- 11. Identification accuracy for all combinations of feature types/layers and ROIs: Identification-accuracy results are reported for combinations of feature types, layers, and ROIs obtained by image feature decoders.
17. Identification accuracy with true image feature values (generic object recognition, GOR)
Generic object recognition accuracy was evaluated using true image features, equivalent to perfectly predicted features from brain activity. Identification was tested across extensive candidate sets, with CNN features showing high performance and CNN8 slightly below CNN7.
- Evaluation: GOR accuracy corresponds to identification using visual features perfectly predicted from brain activity by image feature decoders.Identification accuracy was evaluated for each visual feature type/layer.
- Evaluation: Identification from two categories compared each of 50 test object categories with one of 15,322 candidate categories.Error bars represented 95% CI across 50 test categories, with a 50% chance level.
- Evaluation: Identification from 100 categories used 100 randomly selected candidate sets for each of the 50 test categories.Correct-identification percentages were averaged across candidate sets, with a 1% chance level.
- Results: GOR identification was slightly poorer with CNN8 than with CNN7.The high accuracy of original CNN features in object recognition was given as one reason for high CNN-feature accuracy in generic decoding.