Source-linked AI summary

Deep Neural Networks Rival the Representation of Primate IT Cortex for Core Visual Object Recognition

Charles F. Cadieu, Ha Hong, Daniel L. K. Yamins, Nicolas Pinto, Diego Ardila, Ethan A. Solomon, Najib J. Majaj, James J. DiCarlo

arXiv:1406.3284v1q-bio.NCcs.NE

TL;DR

The paper asks whether DNN representations rival the primate brain's IT cortex for core visual object recognition despite experimental and computational comparison constraints. It directly compares matched neural and model representations using kernel analysis and related measures, finding that the latest DNNs rival IT performance and also align with IT responses and representational similarity. The authors caution that the study's macaque measurements and human-performance comparison leave important limitations.

  • Problem

    It remains unclear whether DNN representations rival IT cortex for object recognition, because fair comparison must account for neural noise, recording scale, trials, classifier complexity, and training examples.

  • Method

    The study directly compares IT and DNN representations using matched measurements, an extension of kernel analysis, encoding-model predictions, and representational similarity.

  • Results

    Latest DNNs rival IT cortex in representational performance, and high-performing models also show strong IT representational similarity and IT multi-unit response prediction.

  • Takeaways & Limitations

    The latest DNNs cannot be ruled out on representational-performance grounds as relying on computational mechanisms similar to the primate visual system.

  • Takeaways & Limitations

    The study's macaque measurements are constrained by brief viewing, the behavioral paradigm, and mapping neural recordings to neural features; human performance was not measured with the same cross-validated procedure as models.

Abstract

from arXiv · show

The primate visual system achieves remarkable visual object recognition performance even in brief presentations and under changes to object exemplar, geometric transformations, and background variation (a.k.a. core visual object recognition). This remarkable performance is mediated by the representation formed in inferior temporal (IT) cortex. In parallel, recent advances in machine learning have led to ever higher performing models of object recognition using artificial deep neural networks (DNNs). It remains unclear, however, whether the representational performance of DNNs rivals that of the brain. To accurately produce such a comparison, a major difficulty has been a unifying metric that accounts for experimental limitations such as the amount of noise, the number of neural recording sites, and the number trials, and computational limitations such as the complexity of the decoding classifier and the number of classifier training examples. In this work we perform a direct comparison that corrects for these experimental limitations and computational considerations. As part of our methodology, we propose an extension of "kernel analysis" that measures the generalization accuracy as a function of representational complexity. Our evaluations show that, unlike previous bio-inspired models, the latest DNNs rival the representational performance of IT cortex on this visual object recognition task. Furthermore, we show that models that perform well on measures of representational performance also perform well on measures of representational similarity to IT and on measures of predicting individual IT multi-unit responses. Whether these DNNs rely on computational mechanisms similar to the primate visual system is yet to be determined, but, unlike all previous bio-inspired models, that possibility cannot be ruled out merely on representational performance grounds.

Author Summary

Primate vision recognizes object categories rapidly despite substantial visual variation, while recent DNNs reach performance equal to IT cortex on this task.

  • Latest artificial deep neural networks achieve performance equal to IT cortex for recognizing object categories across substantial visual variation.The tested variation includes object exemplar, position, pose, scale, and background.
  • The study compares neural features from macaque IT cortex with features derived from the latest deep neural networks across thousands of images.
  • Primate object recognition remains effective during brief presentations and changes in exemplar, position, pose, scale, and background.

Introduction

Primate vision supports rapid, robust object recognition through representations formed along the ventral stream, especially in IT cortex. The paper develops a fair comparison between IT and DNN representations by correcting experimental limitations and evaluating accuracy across representational complexity.

  • Motivation: Primate object recognition remains highly accurate at low latency despite brief presentations and changes in exemplars, transformations, and backgrounds.Humans and macaques can solve such tasks with presentation times shorter than 100 ms.
  • Motivation: The ventral stream extends from V1 through V2 and V4 to IT cortex, where visual representations are selective for object identity and tolerant to nuisance variation.
  • Methodological contribution: The paper addresses shortcomings in prior comparisons by explicitly accounting for experimental noise, recorded neural sites, stimulus presentations, and other limitations.The authors report that these corrections can dramatically affect comparison results.
  • Methodological contribution: A novel extension of kernel analysis measures representation accuracy as a function of decision-boundary complexity rather than relying on fixed-complexity classifiers.This identifies representations that achieve high accuracy at a given task complexity.
  • Experimental design: The study evaluates class-level object representations using a dataset of 1960 images, compared with prior datasets containing 150 and 96 images.The larger image set samples greater stimulus variation relevant to assessing IT representational performance.
  • Study aim: The paper establishes human performance on brief object-categorization trials and compares latest DNNs with previous bio-inspired models and IT cortex representations.

Results

The study evaluates core visual object category recognition using controlled image variation and compares neural and model representations under matched sampling and noise. Recent DNNs, especially Zeiler & Fergus 2013, rival or exceed IT representations, while earlier biologically inspired models perform near chance or marginally above it.

  • Task: The task tests object category recognition across exemplar, position, scale, pose, rotation, and unique-background variation in 1960 images.Categories included Cars, Fruits, Animals, Planes, Chairs, Tables, and Faces.
  • Method: Kernel analysis measures precision as a function of decision-boundary complexity, revealing representation quality beyond a fixed classifier complexity.The analysis uses inverse regularization, 1/λ, as the complexity measure.
  • Model comparison: V1-like and V2-like models perform near chance across complexity, while HMAX performs only marginally better and recent DNNs perform substantially better.The task is difficult for simpler and earlier biologically inspired representations under the tested visual variation.
  • Multi-unit comparison: After matching sampling and neural noise, IT multi-unit performance is only matched by the Zeiler & Fergus 2013 representation across the full complexity range.The comparison uses 80 samples for multi-unit representations and adds matched noise to model representations.
  • Single-unit comparison: After single-unit matching, IT performs better than HMO, slightly worse than Krizhevsky et al. 2012, and worse than Zeiler & Fergus 2013.Higher single-unit noise and fewer trials produce lower precision-versus-complexity curves than the multi-unit analysis.
  • Robustness: Across sample sizes, Zeiler & Fergus 2013 rivals IT multi-unit performance, while Krizhevsky et al. 2012 and Zeiler & Fergus 2013 surpass IT single-unit performance.The area-under-the-curve relationship remains robust across the measured sampling range.

Discussion

The latest DNNs rival IT cortex in representational performance, but several dataset, measurement, and biological-comparison limitations remain. Encoding and decoding analyses provide complementary evidence, while the computational mechanisms underlying the similarity remain unresolved.

  • Results: The latest DNNs rival IT cortex on rapid object category recognition after corrections for sampling, noise, and trial limitations.Zeiler & Fergus 2013 matched IT multi-unit performance, while Krizhevsky et al. 2012 and Zeiler & Fergus 2013 surpassed IT single-unit performance.
  • Method: Kernel analysis measures precision as a function of classifier complexity, favoring representations that support accurate classification with simple decision functions.Such representations can predict class labels from few examples, whereas poor representations require complex classifiers and more training examples.
  • Limitations: The controlled image set omits contextual effects and variations such as lighting, texture, natural deformation, and occlusion.The authors describe these omissions as targets for future datasets and neural measurements.
  • Limitations: Macaque measurements were limited by brief viewing, passive viewing, and unresolved mapping from recordings to neural features.The authors identify viewing time, behavioral paradigm, and neural-feature mapping as issues for determining ultimate representational performance.
  • Limitations: Human and DNN overall task performance were not directly compared, although the authors infer that human performance exceeds current DNN performance.The inference is based on human mean accuracy of 85% versus DNN test-set accuracy of 77%, alongside differences in exposure to labeled training images.
  • Method: Encoding and decoding measures are complementary: Zeiler & Fergus 2013 rivals IT decoding performance but fails to capture over 40% of explainable variance in the IT neural sample.Kernel analysis and linear-SVM analyses decode class labels, whereas IT-response prediction and representational-similarity analyses encode neural variation from image-derived measurements.
  • Related work: High-performing DNNs use concepts extending to early primate visual-system models, including successive simple-complex-like layers and max-pooling.The models also use convolution or weight sharing, and some use backpropagation.
  • Limitations: The tested seven categories are only a small fraction of the 1000 classes used to train the successful DNNs, and the class correspondence is unclear.The effects of non-relevant training classes and the role of ecologically relevant categories remain unresolved, as does comparability between biological development and 15M labeled training images.

Methods

The study followed institutional oversight requirements for both animal experiments and human behavioral measurements.

  • Ethics: Animal procedures complied with NIH laboratory-animal guidelines and received approval from MIT’s Committee on Animal Care.The animal protocol was 0111-003-014.
  • Ethics: Human behavioral measurements received approval from MIT’s Committee on the Use of Humans as Experimental Subjects.The approval number was 0812003043.

Image dataset generation

The image dataset used synthetic 3-D object renderings with controlled object variation and randomly selected backgrounds.

  • Generation: Synthetic object images were generated from 3-D models rendered with POV-Ray.The workflow converted purchased 3-D models to POV-Ray format and enabled arbitrary numbers of objects with controlled identity-preserving transformations.
  • Backgrounds: Each projected object image was combined with a randomly chosen background, and no two images shared the same background.Some backgrounds were accidentally correlated with object identity, but most were uncorrelated and provided no identity information.
  • Image formation: A circular aperture with radial fall-off was applied to create each final image.

Neural data collection

Neural data were collected from V4 and IT in two macaques using chronic multi-electrode arrays, then normalized into multi-unit and single-unit representations.

  • Recording: Recordings sampled 128 visually driven sites in one monkey and 168 in another across V4 and IT.The recordings included 58 IT and 70 V4 sites in one animal, and 110 IT and 58 V4 sites in the other.
  • Multi-unit representation: Multi-unit representations used firing-rate vectors computed from spikes recorded 70–170 ms after image onset.The normalization process also subtracted background firing measured during gray-screen presentation.
  • Single-unit representation: Single-unit representations were created by spike-sorting multi-unit recordings before applying a similar normalization process.Affinity propagation and a published method isolated 160 IT and 95 V4 single-units; 40 units with highest consistency were selected separately for IT and V4.

Kernel analysis methodology

Kernel analysis evaluates representations by plotting generalization precision across prediction-function complexity. It uses regularized kernel regression and leave-one-out error, with complexity defined as 1/λ.

  • Kernel analysis: Kernel analysis measures representation quality by evaluating precision as prediction-function complexity increases.It uses regularized kernel regression to assess how well the task can be solved in the feature space.
  • Kernel analysis: The regularization parameter λ controls function complexity, while precision is defined as 1−looe(λ), where looe is leave-one-out generalization error.The plotted curve is precision against complexity 1/λ.
  • Computation: The procedure maps images to feature representations, constructs a Gaussian kernel matrix, and solves regularized least squares over a range of regularization values.The Gaussian kernel uses scale parameter σ and is applied to extracted feature representations.
  • Error estimation: Leave-one-out errors are computed from the regularized kernel solution and averaged as mean squared error to estimate generalization error.The error can be computed efficiently from an eigendecomposition of the kernel matrix.
  • Model selection: The kernel width σ is optimized separately at each complexity value, and the resulting precision-complexity curves can plateau at high complexity.The procedure minimizes looe over σ for each λ before plotting precision against 1/λ.
  • Evaluation: The evaluation uses 1960 images from seven object categories and repeats curve estimation ten times using 80% image subsamples with replacement.The categories are Animals, Cars, Chairs, Faces, Fruits, Planes, and Tables.

Machine representations

The study evaluates DNN and biologically inspired visual representations, including convolutional networks, V1-like and V2-like models, and HMAX. The models differ in training procedure, architecture, dimensionality, and biological motivation.

  • Deep neural networks: The evaluated representations include three recent convolutional DNNs that had successively surpassed state-of-the-art ImageNet performance.The DNNs examined are identified through references [25] and Yamins et al.
  • Biologically inspired models: The V1-like model computes locally normalized, thresholded Gabor wavelets across orientation and frequency, forming an 86400-dimensional baseline representation.It is intended as a first-order account of primary visual cortex.
  • Biologically inspired models: The V2-like model combines Gabor outputs nonlinearly and averages them within receptive-field windows, producing a 24316-dimensional representation.The model is intended to capture functional properties associated with visual area V2.
  • Biologically inspired models: HMAX is a two-layer biologically inspired hierarchical model using sparse localized features and has a 4096-dimensional representation.It had performed relatively well on previous invariant object-recognition measures and explained some ventral-stream responses.
  • Deep neural networks: The HMO model is a four-layer convolutional network developed by hierarchical modular optimization on a screening task with objects placed on randomly selected backgrounds.Its representation has 1250 top-level outputs and is evaluated on a task with different objects and backgrounds.
  • Deep neural networks: SuperVision is a supervised convolutional network trained on ImageNet data, while the Zeiler–Fergus model is an eight-layer supervised network trained on LSVRC-2012.SuperVision uses 4096 penultimate-layer features; the Zeiler–Fergus representation is also described through learned convolutional features.

Experimental noise matched model

The noise-matched model procedure adjusts artificial representations to match noise observed in neural measurements. It estimates rate-dependent neural noise, scales model variance, and adds signal-dependent noise.

  • Noise matching: Model representations receive matched noise because neural representational performance is limited by observed measurement noise.Noise is added to models rather than fully removed from neural representations.
  • Neural noise model: The neural noise model assumes approximately Poisson spike-count variation and estimates a rate-dependent relationship between response mean and variance.Responses are first normalized so their variance equals 1.
  • Neural noise model: The fitted mean–variance coefficients are a = 0.14 and b = 0.92 for multi-unit sites, versus a = 0.76 and b = 0.71 for single-unit sites.Coefficients are averaged across neural sites to produce one mean–variance relationship for each measurement type.
  • Noise estimation: The procedure separates empirical response variation into underlying signal and noise contributions, estimating signal standard error jointly across neural sites.Joint estimation can improve the noise-model parameters when sites share similar noise characteristics.
  • Model correction: Scaling model variance makes the variance after noise addition approximately equal to the variance observed in the neural sample.The procedure then adds signal-dependent noise to the scaled model representation.
  • Validation: Using the empirical estimate of σ²_noise produces nearly identical results to the jointly estimated noise model.The resulting model variance was empirically verified against the neural signal variance.

Linear-SVM methodology

The supplied passage only states that supporting information exists and does not describe a linear-SVM methodology.

  • Linear-SVM methodology: The supplied material provides no substantive description of the linear-SVM methodology.

Predicting IT multi-unit sites from model representations

The study uses ridge-regression generalized linear models to predict IT multi-unit responses from model representations, evaluating predictions on held-out data.

  • Ridge-regression generalized linear models predict each IT multi-unit response from model representations.Models are estimated on 80% of the data and evaluated on the remaining 20%, repeated across 10 randomizations.

Representational similarity analysis

Representational similarity is measured by comparing object-level representational dissimilarity matrices, using Spearman rank correlation between their corresponding entries.

  • Feature vectors are computed for each object by averaging representation vectors across image variations.
  • Each representation produces a 49x49 representational dissimilarity matrix because the task contains 49 unique objects.
  • Similarity between two representational dissimilarity matrices is measured by Spearman rank correlation over upper-triangular, non-diagonal elements.The same 20% image split is used across RDM calculations, while images used for encoding models are kept separate.
  • The additional IT multi-unit fit provides an evaluation metric distinct from image-level explained-variance prediction of IT multi-units.For the “+ IT-fit” representations, IT multi-unit predictions are evaluated using an object-level measure.

Supporting Information

The supporting analyses justify the 100 ms neural-recording regime, describe classifier and correction procedures, compare IT recording types, and contextualize model speed and energy use.

  • Human performance on our task as a function of presentation time: Mean human accuracy reaches 92.8% at 2 seconds, while 50 ms presentations reach 82% of the 2000 ms performance.Performance is robust to shorter presentations but declines toward chance as presentation time continues to decrease.
  • Noise and sampling corrections: Noise-model corrections reduce measured performance relative to empirically observed noise, conservatively penalizing noise-matched model representations.Figure S1 varies experimental repetitions and noise-model trials for both multi-unit and single-unit IT samples.
  • Human performance on our task as a function of presentation time: 92% of 2000 ms human performance is reached at 100 ms, supporting this presentation time for neural recordings.At 100 ms, performance is also close to that at 200 ms, while enabling nearly double the data-collection throughput.
  • Linear-SVM methodology: The linear-SVM evaluates generalization from training images to held-out testing images using a simple linear decision boundary.The detailed procedure trains on 1568 images and tests on 392, with regularization selected by 3-fold cross-validation.
  • Comparing IT multi-unit and single-unit representations: With six fixed trials, multi-unit recordings surpass single-unit recordings by recording site and become comparable to unscreened single-units after neuron-count adjustment.The adjustment estimates approximately 4–5 single-units per multi-unit and multiplies multi-unit counts by 5.0.
  • Processing time and energy consumption of computational models: 65 ms per image is reported for a DNN processing batches of 128 images, comparable to the 100 ms presentation and IT response-latency regimes.The estimate excludes phototransduction and computer-system communication latencies.
  • Processing time and energy consumption of computational models: Current DNN implementations are estimated to be around 400 times less energy efficient than the macaque ventral stream.The estimate compares GPUs operating at 200–350 W under load with an estimated 0.5 W macaque ventral stream.
Loading 1406.3284v1…