Source-linked AI summary
Primate vision reveals a missing principle for robust dynamic AI
Matteo Dunnhofer, Christian Micheloni, Kohitij Kar
TL;DR
The paper asks how visual systems combine object appearance and motion while remaining robust when appearance changes. It compares human perception and macaque IT with diverse image- and video-based models using controlled dynamic stimuli and appearance disruption. Predictive world models best combine behavioral generalization with neural correspondence, but no model reproduces IT’s temporal transformation toward appearance-invariant motion coding.
Problem
Naturalistic video performance and existing neural comparisons do not establish whether models learn appearance-robust motion representations or merely exploit correlated visual features.
Method
The study compares human perception, macaque IT activity, and image-, video-, segmentation-, optic-flow-, and predictive-world-model representations on dynamic object tasks with appearance disruption.
Results
Predictive world models combined strong cross-appearance behavioral generalization with the highest correspondence to macaque IT among tested artificial systems, while other video models generally generalized poorly.
Takeaways & Limitations
Robust dynamic vision requires progressive integration of appearance and motion into object representations, with predictive learning a promising route toward this target.
Takeaways & Limitations
The analyses used controlled videos focused on object identity, motion direction, and speed, and examined only IT rather than the broader dynamic vision system.
Abstract
from arXiv · showhide
How does an intelligent visual system combine what objects look like with how they move while remaining robust as appearance changes? We addressed this question by comparing human perception and neural activity in macaque inferior temporal cortex with representations from image- and video-based neural networks spanning recognition, segmentation, optic-flow processing and predictive world modeling. Temporal integration improved object representations, but most video recognition models generalized poorly when appearance was disrupted while motion structure was preserved. Humans and macaque IT remained robust. Notably, predictive world models combined strong cross-appearance generalization with the closest correspondence to IT, outperforming other video-modeling approaches in neural fidelity. Yet no model reproduced the cortical transformation from early appearance-dominated responses toward later appearance-invariant motion coding. These results identify progressive integration of motion into object representations as a principle of robust dynamic vision and implicate predictive learning as a promising route toward realizing this computation in artificial systems.
Authors
Matteo Dunnhofer, Christian Micheloni, and Kohitij Kar are the paper’s authors.
- The paper is authored by Matteo Dunnhofer, Christian Micheloni, and Kohitij Kar.
Affiliation
The authors are affiliated with York University in Toronto and the University of Udine in Udine, Italy.
- York University’s Department of Biology and Centre for Vision Research is located in Toronto, Canada.
- The University of Udine’s Department of Mathematics, Computer Science, and Physics is located in Udine, Italy.
Introduction
The paper asks how visual systems combine object appearance and motion over time while remaining robust to changing appearance. It compares biological vision with multiple artificial approaches using controlled dynamic stimuli and appearance disruption, identifying progressive appearance–motion integration and predictive learning as key directions for dynamic AI.
- Motivation: Dynamic vision must maintain stable object representations while extracting motion from continuously changing visual input.
- Motivation: High video-benchmark performance can reflect appearance–motion correlations rather than motion representations that generalize when appearance changes.
- Approach: The study compares humans, macaque IT, image- and video-recognition networks, segmentation systems, optic-flow models, and predictive world models.
- Findings: Temporal integration improves access to identity and motion, but most video models generalize poorly across appearance disruption while humans and macaque IT remain robust.
- Findings: Predictive world models show the strongest correspondence with IT among tested video-modeling approaches, although no model reproduces IT’s temporal shift toward appearance-invariant motion coding.
- Approach: Image-based networks process frames independently, whereas video networks jointly process multiple frames through spatiotemporal computations.
Humans accurately discriminate moving objects
Humans robustly extract object identity, motion direction, and speed from brief videos, whereas conventional video models gain decodability from temporal integration without reliably generalizing motion across appearance changes. Optic-flow and predictive world models improve cross-appearance robustness, while predictive models best align with IT representations, although no model reproduces IT’s temporal transformation.
- Human behavioral benchmarks: Humans achieved near-ceiling identity accuracy and robust motion-direction and speed judgments from 300 ms videos.Mean accuracies were 0.94 ± 0.02 SD for identity, 0.91 ± 0.07 SD for motion direction, and 0.82 ± 0.06 SD for speed.
- Temporal integration improves decoding: Video ANNs outperformed image-based ANNs on dynamic object recognition, with mean accuracy 0.73 ± 0.07 SD versus 0.66 ± 0.07 SD.The population difference was Δ mean accuracy = 0.07, p = 0.04.
- Temporal integration improves decoding: Temporally varying input increased access to identity- and motion-related information, but naturalistic decoding did not establish appearance-robust motion representations.Because appearance and motion covary in natural videos, models were tested on appearance-free stimuli preserving motion structure.
- Cross-appearance robustness: Humans maintained robust motion judgments after recognizable appearance was removed, establishing a stringent cross-appearance benchmark.Appearance-based performance was mean accuracy = 0.92 ± 0.04 SD, approximately equal to performance with appearance-free stimuli.
- Cross-appearance robustness: Image- and video-recognition ANNs decoded motion well from naturalistic videos but transferred markedly less successfully to appearance-free videos.Naturalistic decoding reached 0.78 ± 0.09 SD for video ANNs and 0.71 ± 0.05 SD for image ANNs; image models fell to chance-level cross-appearance performance.
- Robust model strategies: Optic-flow and predictive world models approached human cross-appearance performance, with predictive models reaching 0.85 ± 0.04 SD on appearance-free videos.World models achieved 0.92 ± 0.02 SD on naturalistic videos, while optic-flow models achieved 0.91 ± 0.01 SD and 0.79 ± 0.01 SD, respectively.
- Biological robustness: IT contained motion-direction information that generalized across appearance disruption, providing a biological target for artificial representations.IT decoding exceeded chance for both held-out naturalistic and appearance-free responses.
- Neural correspondence: Spatiotemporal models aligned better with late IT than spatial models, while predictive world models showed the strongest IT correspondence among tested artificial systems.Early CKA was similar for spatial and spatiotemporal models, but spatiotemporal representations improved for late IT; predictive models reached spatial mean CKA = 0.30 ± 0.02 SD and spatiotemporal mean CKA = 0.36 ± 0.0 SD.
Ethics Approval
The study received ethics approval for both human-participant recruitment and non-human-primate procedures, with consent and animal-care guidelines documented.
- Human Experiments: Human participant research was approved by the York University Ethics Review Committee's Human Participant Review Subcommittee.Participants were recruited through Amazon Mechanical Turk and gave consent before proceeding.
- Human Experiments: Participants provided consent on the first page before taking part in the study.The compensation rate was $15 CAD/hour.
- Non-human Primate Experiments: Non-human-primate procedures followed NIH, MIT, and Canadian Council on Animal Care guidelines and received York University Animal Care Committee approval.The passage states that all data were collected and animal procedures performed in accordance with these requirements.
Human subjects
The study collected online behavioral data from consenting adults recruited through MTurk, with compensation, quality control, and reliability checks specified.
- 107 human participants were recruited through Amazon Mechanical Turk and compensated at $15 CAD per hour.
- Participants provided consent before proceeding and received no additional training before the task.
- Participants with accuracy below 0.6 during the first 30 trials were excluded, compared with a chance level of 0.5.
- Video-level behavioral reliability increased with additional repetitions, reaching approximately 0.8 and approaching asymptote after about 60 repetitions.
Non-human primates
Non-human primate data came from five adult male rhesus monkeys under approved procedures consistent with institutional and animal-care guidelines.
- The study used 5 adult male rhesus monkeys (Macaca mulatta).
- Animal procedures and data collection followed NIH, MIT, and Canadian Council on Animal Care guidelines and received York University Animal Care Committee approval.
Visual Stimuli
The study used rendered single-object videos and controlled motion manipulations, including appearance-free stimuli generated by advecting noise with estimated optic flow and incoherent stimuli created by frame permutation.
- Stimulus construction: High-quality single-object videos used 2D projections of 3D models rendered against random backgrounds across 10 object categories.
- Stimulus construction: Objects were translated according to selected speeds and eight motion directions using custom MATLAB scripts.
- Appearance-free videos: Appearance-free videos used random noise frames advected over time by optic flow estimated from the original videos.
- Motion controls: Incoherent-motion videos randomly permuted 18 frames, preserving individual-frame content while disrupting smooth trajectories and consistent motion.
Behavioral Tasks
Human participants viewed brief videos of moving objects and reported object identity, motion direction, or motion speed. Although stimulus sets contained ten categories and eight directions, each trial used a two-alternative forced-choice decision.
- Behavioral procedure: Participants viewed 300-ms videos of moving objects after fixation and reported attributes through an interactive interface.Trials began with 100 ms of fixation, followed by the video presentation.
- Tasks: The behavioral battery tested object recognition, motion-direction discrimination, and motion-speed discrimination.Recognition and direction used two-way choices; speed required selecting the faster of two videos.
- Task structure: Ten object categories and eight motion directions defined the stimulus sets, while each trial required a binary choice between two candidates.The complete stimulus space was larger than the two-alternative decision made on any individual trial.
- Performance metric: Behavioral performance was summarized with the video-level discriminability metric I1, averaging pairwise accuracy against distractor classes.Each trial compared the ground-truth class with one distractor, and accuracies were averaged across distractors.
Non-human primate electrophysiology
Macaque IT activity was recorded during controlled passive viewing with implanted Utah arrays and monitored fixation. Analyses pooled IT sites while applying visual-response and reliability criteria to selected site-level analyses.
- Electrophysiology: Multiple 10x10 Utah arrays recorded simultaneously from 96 electrodes per array implanted in macaque inferior temporal cortex.Electrodes were 1.5 mm long and spaced 400 µm apart.
- Eye tracking: Monkeys were trained to fixate a central 0.2° target within a ±2° window, with calibration repeated when eye-position drift occurred.Eye movements were monitored with an EyeLink 1000 system.
- Viewing protocol: During passive viewing, monkeys fixated for 100 ms, viewed a 300-ms video, received a fluid reward, and waited through a 500-ms inter-trial interval.Videos were displayed on a 1,920 × 1,080 monitor positioned 42.5 cm from the animal.
- Signal processing: Neural activity was sampled continuously at 20 kHz after 0.1-Hz-to-10-kHz band-pass filtering, primarily as pooled multiunit activity across the IT axis.Specific recording locations were not included in the analyses, treating sites as random samples from pooled IT.
- Site selection: Population analyses used all recorded IT sites, whereas model alignment and neuron-specific correlations retained visually responsive, reliable sites.The reliability threshold required video rank-order response reliability greater than 0.4.
- Reliability: Individual-site reliability was estimated with split-half internal consistency across random, non-overlapping trial subsets.Mean stimulus responses were compared across repeated splits to assess response-profile stability.
Neural Data Analysis and Statistics
The study analyzed time-resolved decoding, representational stability, and appearance-versus-motion sensitivity in IT. Noise correction and reliability normalization were used to compare neural responses across time and stimulus conditions.
- Decoding: Motion direction was decoded independently in 30-ms bins with eight-class LDA and 10-fold cross-validation, including transfer across appearance conditions.Cross-condition transfer tested whether motion information generalized when object appearance changed.
- Decoding metric: Neural decoding used the video-level discriminability metric I1, comparing predicted ground-truth probabilities with each alternative class.Pairwise probability ratios were averaged across alternatives.
- Temporal dynamics: Temporal representation stability was measured by Pearson correlations between each neuron's stimulus-wise response vectors separated by 30 or 60 ms.Firing rates were first averaged across trials within each time bin.
- Noise correction: Temporal correlations were normalized by split-half reliability estimates corrected with the Spearman–Brown formula.This correction addressed noise and finite trial variability at each neuron and time point.
- Population statistics: Population temporal change was summarized from median correlations across neurons over early and late windows extending to 390 ms after stimulus onset.The change measure contrasted correlations at 30-ms and 60-ms offsets from the same response origin.
- Appearance and motion: Three video classes separated naturalistic appearance from coherent motion, incoherent motion, and appearance-free coherent motion.These conditions enabled separate tests of appearance-related and motion-related response sensitivity.
- Appearance and motion: Appearance- and motion-related correlations were noise-corrected using geometric means of Spearman–Brown-adjusted split-half reliabilities.The resulting estimates quantified individual-neuron encoding of appearance-specific or motion-specific information over time.
- Population statistics: Time-resolved population trajectories used median corrected correlations across neurons, with variability summarized by median absolute deviations.Histograms additionally characterized neuron-level correlation distributions at representative early and late times.
ANN Model Analysis
The analysis compared pretrained image- and video-based networks spanning recognition, segmentation, optic flow, and predictive modeling with macaque IT representations. Models were evaluated using linear decoding, cross-condition transfer, temporal analyses, and noise-normalized CKA alignment.
- Models: The benchmark included 15 video ANNs and 14 image-based architectures, all evaluated as publicly available pretrained models.The video set covered recognition, segmentation, optic-flow, and predictive world-modeling approaches.
- Feature extraction: Image models processed video frames independently, while video models extracted features from current and preceding frames held in temporal buffers.Recurrent image models were treated as feedforward frame-wise systems for this analysis.
- Feature extraction: When a temporal buffer exceeded the available history, the first frame was repeated to fill the remaining positions.This established the input sequence used for early video frames.
- Feature extraction: Frame-wise-output networks such as SAM2, MatNet, and FusionSeg supplied features at each frame for analysis.These models were included alongside fixed-buffer video architectures.
- Model characterization: Optic-flow fields between consecutive frames were computed with RAFT, and each video model's IT-like layer was selected by maximum explained variance on HVM640 images.This layer-selection procedure used neural predictivity to macaque responses.
- Temporal controls: Spatial non-time-integrative controls were created by removing temporal signals, including repeated-frame buffers and appearance-independent optic-flow processing.These controls supported comparisons between spatial and spatiotemporal representations.
- Representation processing: All model representations were compressed to 1,000 features using Sparse Random Projection with 10 repetitions.The same compression procedure was applied across model analyses.
- Model decoding: Linear decoding used eight-class LDA for motion direction and ten-class LDA for object identity, while speed used PLS regression with 10-fold cross-validation.Object-speed performance was evaluated from continuous scalar predictions.
Statistical Analyses
The study used correlation-based analyses for continuous variables and selected paired or unpaired group tests according to distributional properties. Statistical families were kept separate, with exact p-values and test statistics reported without global multiple-comparison correction.
- Continuous relationships were assessed using Spearman rank or Pearson correlation coefficients.
- Normality was evaluated with the Shapiro–Wilk test before choosing parametric or nonparametric group comparisons.
- Paired t-tests were used for paired comparisons, whereas independent-samples t-tests were used for unpaired comparisons when normality assumptions held.
- Conceptually independent analyses were evaluated within separate statistical families rather than using global multiple-comparison corrections.
Supplementary Figures
The supplementary figures document the models, decoding analyses, and appearance–motion comparisons used to relate artificial systems to macaque IT responses. Across these analyses, IT shows a later shift toward motion information that the ANN models do not reproduce, while predictive world models remain robust when appearance is removed.
- Behavioral decoding: Figure S1 compares object-identity, motion-direction, and speed decoding across spatial and spatiotemporal video ANNs.
- IT decoding: IT motion decoding was trained on naturalistic videos and tested on held-out naturalistic and paired appearance-free videos, directly probing cross-appearance generalization.
- Model robustness: Optic-flow models retain motion decoding through their motion branch, whereas appearance-only streams fall to chance when recognizable appearance is removed; predictive world models remain accurate.
- Temporal factor analysis: Figure S4 tracks the time-resolved evolution of appearance and motion factors in IT neurons and ANN representations.
- Temporal transformation: Early appearance–motion differences were IT: 0.85, image ANNs: 0.84, and video ANNs: 0.70; later, IT reduced its appearance bias while ANN models retained a stronger static-appearance bias.
- Model specifications: The supplementary model table lists ImageNet-trained image networks, their IT-like layers, and video-model training configurations.