Source-linked AI summary

Recurrent computations for visual pattern completion

Hanlin Tang, Martin Schrimpf, Bill Lotter, Charlotte Moerman, Ana Paredes, Josue Ortega Caro, Walter Hardesty, David Cox, Gabriel Kreiman

arXiv:1706.02240v2q-bio.NCcs.AIcs.CVcs.LG

TL;DR

How the brain recognizes objects from partial information remains an important question. The paper combines behavioral, physiological, and computational evidence, finding that recurrent computations support recognition when feed-forward models fail on partial objects.

  • Problem

    The study addresses how the brain completes patterns and interprets partial visual information.

  • Method

    The authors combine behavioral, physiological, and computational analyses, including feed-forward models augmented with top-level recurrent connections.

  • Results

    Feed-forward architectures lacked robust recognition of partially visible objects, whereas adding recurrent connections significantly improved recognition across average and object-level analyses.

  • Takeaways & Limitations

    The findings support recurrent computations as a plausible mechanism for recognizing objects from partial information.

  • Takeaways & Limitations

    Because infinitely many bottom-up models exist, failure of the tested architectures does not show that every bottom-up architecture lacks robustness.

Abstract

from arXiv · show

Making inferences from partial information constitutes a critical aspect of cognition. During visual perception, pattern completion enables recognition of poorly visible or occluded objects. We combined psychophysics, physiology and computational models to test the hypothesis that pattern completion is implemented by recurrent computations and present three pieces of evidence that are consistent with this hypothesis. First, subjects robustly recognized objects even when rendered <15% visible, but recognition was largely impaired when processing was interrupted by backward masking. Second, invasive physiological responses along the human ventral cortex exhibited visually selective responses to partially visible objects that were delayed compared to whole objects, suggesting the need for additional computations. These physiological delays were correlated with the effects of backward masking. Third, state-of-the-art feed-forward computational architectures were not robust to partial visibility. However, recognition performance was recovered when the model was augmented with attractor-based recurrent connectivity. These results provide a strong argument of plausibility for the role of recurrent computations in making visual inferences from partial information.

Significance Statement

The study investigates how humans recognize heavily occluded objects, combining behavioral, neurophysiological, and computational evidence to argue that recurrent computations support pattern completion. Humans remain robust under limited visibility, whereas feed-forward models are impaired, and interruption by masking exposes the need for additional processing.

  • Motivation: Pattern completion enables recognition of partially visible objects despite infinitely many possible contours joining their parts.This ability is described as a central property of intelligence and depends on integrating spatial and temporal information.
  • Behavioral evidence: Humans robustly recognize objects from limited information, but recognition rapidly deteriorates when computations are interrupted by a noise mask.The study evaluates partially visible objects using psychophysics, neurophysiology, and computational modeling.
  • Neurophysiological evidence: Backward-masking effects correlate image by image with increased latency in intracranial field potentials along the ventral visual stream.Partially visible objects also take more time to recognize behaviorally and physiologically than fully visible objects.
  • Computational account: Modern feed-forward convolutional hierarchical models are not robust to occlusion, whereas adding recurrent attractor dynamics yields a proof-of-concept model capturing human pattern completion.The proposed recurrent computations include within-layer and top-down mechanisms for spatial and temporal integration.

Results · Robust recognition of partially visible objects. Subjects performed a recognition

Subjects recognized partially visible and occluded objects robustly, including novel shapes, but backward masking substantially disrupted recognition across visibility levels and SOAs. The masking effect was selective for partial information, with a significant SOA-by-masking interaction.

  • Robust recognition of partially visible objects. Subjects performed a recognition: Longer SOAs produced a small but significant improvement in recognition of partially visible objects.Performance correlated with SOA at Pearson r = 0.56, p<0.001.
  • Robust recognition of partially visible objects. Subjects performed a recognition: Recognition of heavily occluded objects was also robust, and occlusion improved performance relative to partially visible objects.The improvement was significant at p<10-4 by Chi-squared test.
  • Robust recognition of partially visible objects. Subjects performed a recognition: Visual categorization of novel shapes likewise remained robust under limited visibility.This extended robustness beyond familiar object images and categories.
  • Robust recognition of partially visible objects. Subjects performed a recognition: Backward masking severely impaired recognition of partially visible objects, whereas whole-object performance was affected only at the shortest SOAs.The contrast was reported at 100% visibility with p<0.01 by two-sided t-test.
  • Robust recognition of partially visible objects. Subjects performed a recognition: Backward masking also reduced recognition of occluded and novel objects across a broad range of SOA values and visibility levels.Effects were significant for occluded objects at p<0.001 and novel objects at p<0.0001 by two-way ANOVA.

Images more susceptible to backward masking elicited longer neural delays

Partially visible objects elicited visually selective neural responses that were significantly delayed relative to whole objects. Across preferred-category images, longer neural response latencies were associated with stronger impairment from backward masking.

  • Neural response delays: Partially visible objects elicited visually selective neural signals, but their responses were significantly delayed relative to whole objects.Different renderings of the same object produced a wide latency distribution because visible features varied across trials.
  • Neural response delays: 248 ms versus 206 ms: peak response timing varied across two example images of the same partially visible object.The peak voltage occurred at 206 ms for the first image and 248 ms for the last image in Fig. 2C.
  • Neural response delays: Pearson r = 0.37, p = 0.004: masking effects correlated with neural response latency for preferred-category images after accounting for image difficulty and recording site.The masking index was not correlated with neural response latency for non-preferred-category images (p = 0.33).
  • Neural response delays: Images producing longer neural response latencies were associated with stronger effects of interrupting computations via backward masking.This association persisted despite noisy single-trial latency measures, different participant groups, and across-subject variability in masking indices.

Discussion

Converging behavioral, physiological, and computational evidence supports recurrent computations as a mechanism for recognizing partially visible objects. Backward masking disrupts these computations, while recurrent model connections improve recognition and reproduce observed behavioral and neural delays.

  • Behavioral evidence: 10-20% visibility still supported recognition, including for novel objects, demonstrating robust pattern completion under severe occlusion.Subjects categorized completely novel objects under low visibility despite never having seen those objects or similar ones before.
  • Behavioral and physiological evidence: 25ms ≤ SOA ≤ 100 ms backward masking impaired recognition of briefly presented partial images.The disruptive effect was correlated with neural delays measured along the ventral visual stream.
  • Computational evidence: Bottom-up architectures trained on whole objects failed to recognize partially visible objects robustly, whereas adding top-level recurrent connections improved performance.The recurrent extension also correlated with human recognition at the object-by-object level and accounted for backward-masking effects.
  • Computational mechanism: Recurrent dynamics brought partial-object representations closer to whole-object representations and matched behavioral and physiological temporal lags.These dynamics evolved over time and were interrupted by a backward mask presented near the image.
  • Physiological evidence: 50 ms physiological delays during partial-object recognition provided time for recurrent connections to contribute beyond whole-object processing.Partial-object representations were delayed relative to whole objects, and masking significantly impaired recognition by interrupting additional computations.

Methods

The Methods combined psychophysical experiments with neurophysiological analyses of partially visible, occluded, novel, and matched stimuli. Behavioral experiments involved 106 volunteers, while neural latency was measured from intracranial field-potential responses.

  • Psychophysics: 106 volunteers (62 female, ages 18-34) participated in the behavioral experiments.The experiments included partially visible objects rendered through bubbles and three variations involving occluded objects, novel objects, and stimuli matched to a previous neurophysiological experiment.
  • Neurophysiology experiments: Neurophysiological intracranial field-potential data in Figs. 2 and 3 were taken from reference (14).Neural latency for each image was defined as the peak-response time in the intracranial field potential and calculated in single trials.

Computational Models. We tested state-of-the-art feed-forward vision models,

The models focused on AlexNet with ImageNet-pretrained weights and tested a recurrent extension with all-to-all connections at its top feature layer. The recurrent weights were based on a Hopfield attractor network and defined using whole-object information.

  • Computational Models: The study focused on AlexNet with weights pre-trained on ImageNet, while evaluating other models in supplementary analyses.AlexNet was the primary feed-forward architecture examined.
  • Computational Models: The recurrent model added all-to-all recurrent connections to AlexNet’s top feature layer.This recurrent neural network was introduced as a proof-of-principle model.
  • Computational Models: The recurrent weights were defined using whole-object information and a Hopfield attractor network implementation.The implementation used MATLAB’s Hopfield attractor network.

Figure Legends

A forced-choice categorization task varied stimulus visibility, exposure duration, masking, and object condition. Masking significantly degraded recognition of partial objects compared with unmasked trials (p<0.001).

  • Task design: Twenty-one subjects completed forced-choice categorization trials with stimulus onset asynchronies (SOAs) from 25 to 150 ms.After 500 ms of fixation, stimuli were followed by either a gray screen or a 500 ms noise mask.
  • Stimulus conditions: Stimuli were presented whole, partially visible, or occluded, with a separate variation using novel objects.The figure legend identifies these as the Whole, Partial, occluded, and novel-object conditions.
  • Behavioral measures: Behavioral performance was plotted against visibility for unmasked and masked trials, with colors indicating SOA and error bars showing SEM.Chance performance was 20%, bin size was 2.5%, and the x-axis was discontinuous to report 100% visibility.
  • Additional conditions: Performance was also plotted against SOA for occluded stimuli, with chance set to 25%, and for novel-object stimuli.The occluded-stimulus analysis refers to the conditions shown in D, while the novel-object analysis refers to E.

on an image-by-image basis

Image-by-image analyses linked partial-object neural responses to backward-masking effects and evaluated whether feed-forward models could generalize from whole to partial objects. Neural latency and masking strength were significantly correlated across partial exemplars.

  • Physiology: Intracranial field potentials showed stronger responses to faces than other whole-object categories in the left fusiform gyrus.Responses were averaged across five whole-object categories during 150 ms presentations without masking.
  • Masking–latency relationship: 0.37 Pearson r linked each partial image’s masking index to its neural response latency, with P = 0.004.The correlation was significant by linear regression and permutation test across preferred-category partial objects from two electrodes.
  • Computational models: Feed-forward computational models were compared with humans by training an SVM on whole-object features and testing it on partial-object features.Training and test objects were nonoverlapping, and human performance was measured at 150 ms SOA for the same images.

time, and was impaired by backward masking

Adding attractor-based recurrence improved AlexNet’s recognition of partially visible objects, approaching human performance and moving representations toward whole-image categories. The recurrent model also became more human-like over time, with masking performance improving at longer SOAs.

  • Recurrent model performance: RNNh significantly improved recognition over the original fc7 layer and approached human performance across visibility levels.The comparison used human and original fc7 curves; error bars represented SEM.
  • Recurrent model performance: Over recurrent time, partial-object representations approached the correct categories formed by whole images.This temporal evolution was visualized with t-SNE.
  • Recurrent model performance: Over time, RNNh’s object classifications became more human-like.Model–human correlations were computed separately by category and averaged; error bars denoted S.D. across categories.
  • Backward masking: Performance improved with increasing SOA when the same backward mask used in psychophysics was fed to RNNh.The model was tested with the mask at different SOA values using 5-way cross-validation; error bars denoted SEM.

Funding

The work was supported by a FITweltweit fellowship, a DAAD fellowship, and grants from NSF and NIH.

  • Funding included a fellowship from the FITweltweit programme of the German Academic Exchange Service (DAAD).
  • Additional support came from NSF STC award CCF-123121 and NIH award R01EY026025.

Data availability statement

All data and code used in the study will be made available upon publication through the lab’s website and GitHub repository.

  • Data availability statement: Upon publication, the lab will release image databases, behavioral and physiological measurements, and computational algorithms through its website and GitHub repository.The release includes both data and code.

Figure 2

The supplied figure passages indicate that backward masking impaired recognition across partial and occluded stimuli, while recurrent connectivity improved model performance under low visibility. Purely feed-forward models were impaired in low-visibility conditions, whereas recurrent models supported robust recognition, including for novel objects.

  • Computational models: Adding recurrent connectivity improved performance in both AlexNet- and VGG16-based models.The VGG16 recurrent implementation used 4096 units from the fc1 layer.
  • Psychophysics: Backward masking consistently impaired categorization of both partial and occluded stimuli.The effect was observed across experiments using 16 exemplars from four categories.
  • Computational models: Purely feed-forward models were impaired under low-visibility conditions.The comparison included AlexNet, VGG, ResNet, and Inception variants pretrained on ImageNet 2012.
  • Novel objects: Recurrent and feed-forward models showed performance patterns for novel objects similar to those for known object categories.The supplemental comparison included feed-forward models, the RNNh model, and the temporal evolution of RNNh feature representations.
Loading 1706.02240v2…