Source-linked AI summary
Actions in the Eye: Dynamic Gaze Datasets and Learnt Saliency Models for Visual Recognition
Stefan Mathe, Cristian Sminchisescu
TL;DR
The paper addresses the gap between computer-vision action-recognition pipelines and human visual attention data for dynamic, task-controlled video. It collects and analyzes large-scale human fixations, develops fixation-based saliency and recognition systems, and reports accurate fixation prediction and state-of-the-art results in some difficult benchmarks.
Problem
Large-scale datasets linking human eye movements to dynamic visual recognition tasks are lacking, while prior gaze and saliency studies largely use static or free-viewing conditions.
Method
The paper collects task-controlled eye movements for Hollywood-2 and UCF Sports, introduces spatial and sequential consistency measures, and trains fixation-based saliency and action-recognition systems.
Results
Human fixation patterns are remarkably stable, fixation predictors are accurate, and automatic saliency-guided systems achieve state-of-the-art results in some hard action-recognition benchmarks.
Takeaways & Limitations
Human eye movements can provide usable supervision for saliency prediction and can be integrated into end-to-end automatic visual action-recognition systems.
Takeaways & Limitations
The action labels rely on existing computer-vision datasets and labeling, although weakly supervised learning may better map persistent structure to higher-level semantic labels.
Abstract
from arXiv · showhide
Systems based on bag-of-words models from image features collected at maxima of sparse interest point operators have been used successfully for both computer visual object and action recognition tasks. While the sparse, interest-point based approach to recognition is not inconsistent with visual processing in biological systems that operate in `saccade and fixate' regimes, the methodology and emphasis in the human and the computer vision communities remains sharply distinct. Here, we make three contributions aiming to bridge this gap. First, we complement existing state-of-the art large scale dynamic computer vision annotated datasets like Hollywood-2 and UCF Sports with human eye movements collected under the ecological constraints of the visual action recognition task. To our knowledge these are the first large human eye tracking datasets to be collected and made publicly available for video, vision.imar.ro/eyetracking (497,107 frames, each viewed by 16 subjects), unique in terms of their (a) large scale and computer vision relevance, (b) dynamic, video stimuli, (c) task control, as opposed to free-viewing. Second, we introduce novel sequential consistency and alignment measures, which underline the remarkable stability of patterns of visual search among subjects. Third, we leverage the significant amount of collected data in order to pursue studies and build automatic, end-to-end trainable computer vision systems based on human eye movements. Our studies not only shed light on the differences between computer vision spatio-temporal interest point image sampling strategies and the human fixations, as well as their impact for visual recognition performance, but also demonstrate that human fixations can be accurately predicted, and when used in an end-to-end automatic system, leveraging some of the advanced computer vision practice, can lead to state of the art results.
1 INTRODUCTION
The paper bridges human visual attention and computer action recognition by collecting task-controlled gaze data, measuring fixation consistency, and using human fixations to guide recognition systems.
- Motivation: Computer action-recognition systems commonly use bag-of-words descriptors extracted at maxima of sparse spatio-temporal interest-point operators.These operators target nontrivial local space-time structure such as corners.
- Motivation: The paper investigates whether computational insights from human vision and fixation-guided sampling can improve sparse computer-vision recognition systems.It compares human fixations with interest-point operators and evaluates fixation-derived operators for action classification.
- Contributions: Large-scale gaze recordings from Hollywood-2 and UCF Sports add task-controlled human eye movements to dynamic action-recognition datasets.The recordings are publicly available and complement existing computer-vision annotations.
- Contributions: Human fixation patterns show remarkable spatial and sequential consistency across subjects, while task has less influence on dynamic fixations within the studied actions.The paper introduces consistency models and video-adapted relevance measures for these analyses.
- Contributions: Learned saliency detectors use human fixation data and visual features to predict fixations, with performance evaluated using average precision and spatial Kullback-Leibler measures.The resulting predictors are incorporated into automatic action-recognition systems.
2 RELATED WORK
Prior saliency and gaze datasets largely emphasize static, free-viewing settings, whereas this work targets learned saliency and gaze analysis for dynamic, task-controlled action recognition.
- Saliency models: Saliency models may be pre-specified or learned from eye-tracking data.Learned models can use low-level or combined low-, mid-, and high-level features.
- Saliency models: Existing saliency research has mainly addressed static images, with comparatively little attention to dynamic video.Video extensions of static bottom-up models incorporate motion and flicker channels and are generally pre-specified.
- Gaze datasets: Prior learned fixation detectors include work trained on free-viewing fixation data rather than a specific recognition task.This distinction motivates task-controlled dynamic gaze collection for action recognition.
- Gaze datasets: Most human gaze datasets contain at most a few hundred images or videos and are usually collected under free viewing.The paper contrasts this setting with its large-scale, dynamic, task-controlled data.
- Action recognition: Computer-vision action-recognition systems commonly use interest-point detectors and complex features within bag-of-visual-words pipelines.Other approaches include random-field models and structured-output SVMs, but the study focuses on bag-of-words spatio-temporal pipelines.
3 LARGE SCALE HUMAN EYE MOVEMENT DATA COLLECTION IN VIDEO
The study adds eye-tracking annotations to two challenging action datasets using controlled recordings from active and free-viewing participants under calibrated laboratory conditions.
- Datasets: Eye recordings were collected for Hollywood-2 and UCF Sports as additional annotations for large-scale video action-recognition datasets.Hollywood-2 contains 12 action classes from 69 movies, while UCF Sports contains 150 videos across 9 sports actions.
- Datasets: Hollywood-2 provides 823 training and 884 test sequences across 12 real-world action classes.The dataset is described as one of the largest and most challenging available for real-world actions.
- Participants: Sixteen volunteers participated: 12 performed the action-recognition task and 4 viewed the videos without a specific task.The participants were 9 male and 7 female volunteers aged 21–41.
- Recording setup: Gaze was recorded binocularly from the right eye at 500 Hz using a tower-mounted eye tracker, with the head stabilized 60 cm from the display.The display resolution was 1280 × 1024 pixels.
- Recording setup: Calibration used 13 screen locations and was restarted when estimated error exceeded 0.75° of visual angle.Validation was performed at four locations and again at the end of each block.
- Protocol: Active participants fixated the screen center before each video, identified its action from multiple-choice options, and submitted answers after viewing.Playback proceeded automatically after the required central fixation.
4 SPATIAL AND SEQUENTIAL CONSISTENCY
Human fixations are highly consistent across viewers, both spatially and in temporal ordering, while task-related comparisons and class-level effects require qualification. The analysis also relates this consistency to action-recognition viewing behavior and semantic AOIs.
- 4.1 Action Recognition by Humans: Near-perfect human action recognition on Hollywood-2 supports the use of gaze data collected during a visual action-recognition task.Errors that occur are typically between semantically related actions, while entirely mislabeling a video is almost nonexistent.
- 4.2 Spatial Consistency Among Subjects: 94.8% and 93.2% AUC indicate strong spatial inter-subject agreement on Hollywood-2 and UCF Sports, respectively.Cross-stimulus controls are lower at 72.3% and 69.2%, reflecting stimulus-related effects in addition to shared fixation patterns.
- 4.2 Spatial Consistency Among Subjects: Spatial consistency remains strong across action classes, although cross-stimulus control varies substantially, especially in UCF Sports.The authors conjecture that filming styles and directors’ presentation choices contribute to this class-level variation.
- 4.3 The Influence of Task on Eye Movements: None of the free-viewing subjects’ fixation patterns deviates significantly from the active group across 1000 sampled frames.This comparison does not establish task discriminability because it contrasts action recognition with no task rather than multiple specific tasks.
- 4.4 Sequential Consistency Among Subjects: AOIs cluster fixations into semantically meaningful regions, enabling scanpaths to be represented as ordered sequences of object-centered symbols.The example includes car parts, people, and carried objects, with arrows representing saccades between AOIs.
- 4.4 Sequential Consistency Among Subjects: 70% average AOI transition probability versus 13% for a random baseline, and 71% aligned AOI symbols versus 51%, reveal strong sequential consistency.The analysis represents scanpaths as discrete AOI sequences and evaluates them with Markov dynamics and temporal alignment.
5 SEMANTICS OF FIXATED STIMULUS PATTERNS
The authors build visual vocabularies from image patches fixated during Hollywood-2 action videos and find that these patches capture recurring, semantically meaningful action-related regions.
- Protocol: Fixated patches are represented with HoG descriptors and clustered into 500 visual-word groups.Each patch spans 1° from fixation center; 500 clusters balance semantic under- and over-segmentation.
- Visualization: Fig. 5 displays high-probability vocabulary patches, with each row showing the five most probable patches from one cluster across different videos.The displayed entries are sampled from vocabularies built by clustering fixated image regions for several Hollywood-2 action classes.
- Findings: Fixated regions include actors, manipulated objects, and some surrounding context, indicating semantically meaningful structure across action classes.Examples include people eating or running, dishes, telephones, car doors, vegetation, and street signs.
- Findings: Fixations usually land on objects or object parts rather than unstructured image regions, while subjects generally avoid object boundaries.The resulting vocabularies capture recurring aspects of the observed actions, although some clusters are not semantically homogeneous because HoG descriptors are limited.
- Implication: The authors suggest that repeatable fixation on people and objects could support computer-based action recognition.This proposed use motivates the recognition experiments that follow.
6 EVALUATION PIPELINE
The evaluation pipeline converts video locations selected by an interest point operator into descriptors, visual-word histograms, and kernel-based action classifiers.
- Pipeline: Each model uses an interest point operator, descriptor extraction, bag-of-visual-words quantization, and an action classifier.The same processing pipeline is used for all computer action recognition models.
- Interest Point Operator: Interest point operators take a video as input and output spatio-temporal coordinates with associated spatial and, usually, temporal scales.The experiments include both computer-vision and biologically derived operators; the 2D fixation operator lacks temporal scale.
- Descriptors: HoG and MBH descriptors are extracted at returned locations across seven grid configurations, producing 14 classification features.MBH is computed from optical flow, and the same grid configurations are used throughout the experiments.
- Visual Dictionaries/Second Order Pooling: Descriptors are clustered by k-means into 4000-word vocabularies, and each video is represented by an L1-normalized visual-word histogram.Clustering uses 500,000 randomly sampled descriptors for computational reasons; second-order pooling is also evaluated later.
- Classifiers: The classifier combines 14 RBF-χ2 kernel matrices with Multiple Kernel Learning and trains one-vs-all classifiers for each action label.Kernel and regularization parameters are selected by cross-validation and grid search.
- Evaluation: Recognition is evaluated using average precision on the test set for each action class.Average precision is identified as the standard metric for Hollywood-2 action recognition.
7 HUMAN FIXATION STUDIES
The fixation studies compare human gaze with computer-vision interest points and test fixation- and saliency-derived sampling within a common recognition pipeline.
- 7 HUMAN FIXATION STUDIES: The experiments evaluate action-recognition operators derived from ground-truth human fixations using the pipeline introduced in Section 6.The study asks whether information in fixated regions can aid automated action recognition.
- 7.1 Human vs. Computer Vision Operators: Human fixations correlate weakly with classical interest-point locations: only approximately 6% of spatio-temporal Harris corners are fixated.The probability varies little across actions, consistent with subjects generally avoiding object boundaries.
- 7.1 Human vs. Computer Vision Operators: Using human fixations directly as interest points does not improve recognition performance over the Harris-based operator.Table 3 compares Harris corners with fixation operators using one point per fixated frame or one point per fixation.
- 7.2 Impact of Human Saliency Maps for Computer Visual Action Recognition: Because the entire foveated area may not be informative, the study also samples more finely within fixated regions to target possible covert-attention subregions.The experiment is motivated by the possibility that relevant action information lies in only part of the fovea.
- 7.2 Impact of Human Saliency Maps for Computer Visual Action Recognition: The saliency operator blurs fixation-count maps with a Gaussian filter, combines the resulting fixation probability with a uniform distribution, and samples locations from the mixture.The mixture is psal = (1 − α)pfix + αpunif; sampled locations receive random spatio-temporal scales from [2, 8].
- 7.2 Impact of Human Saliency Maps for Computer Visual Action Recognition: Ground-truth saliency sampling significantly outperforms both Harris and uniform sampling at equal interest-point sparsity rates.The result indicates that fixation surface structure, without temporal ordering, can boost contemporary action-recognition methods and descriptors.
8 SALIENCY MAP PREDICTION
The paper evaluates saliency predictors for video by combining static, motion, and fixation-trained features, using both ranking-based AUC and distribution-sensitive KL divergence. A HoG-MBH detector trained on human fixations performs best under KL divergence, while combining predictors improves AUC.
- Evaluation measures: AUC ranks saliency maps by their ability to separate fixated pixels from non-fixated pixels, whereas spatial KL divergence compares predicted and ground-truth fixation distributions.AUC does not ensure that normalized predicted probabilities match the ground-truth distribution; lower KL indicates closer distributional agreement.
- Saliency predictors: The predictor set includes static low-, mid-, and high-level features, five motion or space-time feature maps, and fixation-trained HoG-MBH detectors.Motion features include optical-flow magnitude, flow boundaries, flow bimodality, and spatio-temporal Harris cornerness.
- Action recognition application: Using interest points sampled from predicted or ground-truth saliency maps significantly improves Hollywood-2 action recognition over the state-of-the-art baseline, unlike uniform sampling.The comparison uses the same number of interest points per frame as the Harris spatio-temporal corner baseline.
- Fixation-trained detector: The HoG-MBH detector uses static HoG and motion MBH descriptors centered at human fixations and produces saliency maps through sliding-window detection.The detector is trained to exploit semantically meaningful structure often present in fixated regions.
- Detector results: 76.7% average precision is achieved by the combined HoG-MBH detector, compared with 73.4% for HoG alone and 70.5% for MBH alone.The result is measured on 106 random locations, evenly divided between fixated and non-fixated examples.
- Evaluation results: Under KL divergence, the HoG-MBH detector performs best, while under AUC, fusing predicted maps and static and dynamic features gives the highest results.The ranking changes between metrics because AUC optimizes pixel-level classification, whereas KL evaluates spatial probability distributions.
9 AUTOMATIC VISUAL ACTION RECOGNITION
The section evaluates saliency-driven interest-point sampling for action recognition, comparing predicted and ground-truth human saliency with uniform and central-bias baselines across two video datasets. It also examines prediction quality, computational variability, and combinations with dense trajectories.
- Saliency prediction: HoG-MBH detector maps most closely approximate ground-truth saliency spatially, motivating their use for predicted-map interest-point sampling.The detector is selected using the KL divergence criterion; Figure 6 reports that its maps are closest to ground truth.
- Action recognition evaluation: Saliency-based pipelines using predicted or ground-truth maps outperform uniform interest-point sampling on both Hollywood-2 and UCF Sports.Central bias approximates human saliency but does not match the recognition performance of the saliency-map pipelines.
- Action recognition evaluation: Predicted and ground-truth saliency pipelines have similar recognition performance, with a slight predicted-map advantage on Hollywood-2 and a ground-truth advantage on UCF Sports.The comparison is reported for the corresponding results in Tables 4d,e and 5.
- Action recognition evaluation: Combining sparse saliency-sampled descriptors with dense-trajectory kernels using MKL goes beyond state-of-the-art performance.The combined representation uses 18 kernels: 14 sparse saliency-sampled feature kernels and 4 dense-trajectory kernels.
- Variability analysis: The main bag-of-visual-words results use a single random seed because repeated training is impractical with approximately 28 million Hollywood-2 interest points.The large number of random samples is expected to reduce variance, with somewhat higher variability anticipated for smaller UCF Sports.
- Variability analysis: Less than 0.8% standard deviation was observed across 10 random seeds for every UCF Sports pipeline using second-order pooling.The faster second-order-pooling experiments confirmed the trends of the bag-of-visual-words pipeline, although recognition results were somewhat lower.
10 CONCLUSIONS
The paper contributes large-scale human eye-tracking datasets for video action recognition, consistency measures for fixation patterns, and saliency-based recognition systems. It reports accurate trainable saliency operators whose predictions support state-of-the-art action-recognition results on difficult benchmarks.
- Contributions: The authors release comprehensive human eye-tracking annotations for Hollywood-2 and UCF Sports and quantify spatial and sequential fixation consistency.The datasets address video action recognition under task-controlled conditions, while the consistency models compare subjects, videos, and actions.
- Contributions: The paper proposes trainable saliency operators based on human fixations and evaluates them in end-to-end visual action-recognition systems.The analysis focuses on computer-vision interest-point operators and descriptors.
- Conclusions: Automatic saliency predictors achieve state-of-the-art results in some of the field’s hardest action-recognition benchmarks.This conclusion concerns end-to-end systems using predicted human-saliency information.