Source-linked AI summary

CAT2000: A Large Scale Fixation Dataset for Boosting Saliency Research

Ali Borji, Laurent Itti

arXiv:1505.03581v1cs.CV

TL;DR

Saliency models have been evaluated on small, biased fixation datasets, raising concerns about the breadth of their apparent performance. The paper addresses this gap by collecting controlled eye movements from 120 observers viewing 4,000 images across 20 categories. CAT2000 reveals category-dependent fixation statistics, observer consistency, and model performance, and is positioned for benchmarking and behavioral studies.

  • Problem

    Saliency models have performed well on fixation prediction, but existing evaluations use small, biased datasets with limited stimulus variety and center-bias.

  • Method

    The paper records controlled eye movements from 120 observers viewing 4,000 images across 20 diverse categories, then analyzes dataset properties and evaluates four saliency models.

  • Results

    The dataset shows category-dependent center-bias, saccade rates, inter-observer consistency, and model performance, with models performing best on Sketch and poorly on Line drawings.

  • Takeaways & Limitations

    CAT2000 provides a resource for saliency-model benchmarking, behavioral studies, and investigating semantic cues that guide gaze during free viewing.

Abstract

from arXiv · show

Saliency modeling has been an active research area in computer vision for about two decades. Existing state of the art models perform very well in predicting where people look in natural scenes. There is, however, the risk that these models may have been overfitting themselves to available small scale biased datasets, thus trapping the progress in a local minimum. To gain a deeper insight regarding current issues in saliency modeling and to better gauge progress, we recorded eye movements of 120 observers while they freely viewed a large number of naturalistic and artificial images. Our stimuli includes 4000 images; 200 from each of 20 categories covering different types of scenes such as Cartoons, Art, Objects, Low resolution images, Indoor, Outdoor, Jumbled, Random, and Line drawings. We analyze some basic properties of this dataset and compare some successful models. We believe that our dataset opens new challenges for the next generation of saliency models and helps conduct behavioral studies on bottom-up visual attention.

1. Introduction & motivation

Saliency models have achieved strong fixation prediction, but their evaluation has relied on small, biased datasets with limited stimulus variety and center-bias. CAT2000 addresses these shortcomings by systematically collecting a large-scale fixation dataset across diverse image categories.

  • Research gap: Recent saliency models predict fixations nearly as well as the human inter-observer model, but have been evaluated on biased fixation datasets.Existing datasets contain few scenes, few observers, limited stimulus variety, and frequently center-positioned objects.
  • Research gap: Webcam- and crowdsourcing-based collection makes eye-tracking accuracy, calibration, viewing conditions, and observer characteristics difficult to control.These factors include observer distance, field of view, mood, age, intelligence, and concentration.
  • Dataset response: CAT2000 systematically collects a large-scale fixation dataset across several image categories to address dataset-bias challenges in saliency modeling.The motivation is to better understand human information selection and scene perception using controlled eye-movement recordings.

2. CAT2000 dataset

CAT2000 contains 4,000 high-resolution images spanning 20 categories designed to probe varied attentional cues. Eye movements from 120 observers were collected in controlled sessions, with every image viewed by 24 observers.

  • 2.1. Stimuli: The dataset includes 200 Caltech256 object categories, 200 psychological patterns, texture defects, random-viewpoint images, random-location satellite views, and sketches of 200 objects.These constructions provide stimulus sets for object, bottom-up, center-bias, and scene-content analyses.
  • 2.2. Observers: 120 observers participated, comprising 40 male and 80 female undergraduates with normal or corrected-to-normal vision.Observers were naive to the experiment’s purpose and had not previously seen the stimuli; the protocol received USC IRB approval.
  • 2.2. Observers: 4,000 images were partitioned into five 800-image cohorts, with each observer viewing one cohort across four sessions.Each session showed 200 images, lasted about 25 minutes, and included 5 minutes of rest; the eye tracker was recalibrated each session.
  • 2.2. Observers: Each image was viewed by 24 observers through 24 passes, with five observers assigned to each pass.Images were shuffled while ensuring each 200-image section contained 10 images from every category.
  • 2.3. Eye tracking procedure: Each trial presented a fixation cross, a target image for 5 seconds, and a gray screen for 3 seconds while observers freely looked around.Viewing distance was 106 cm, and scenes subtended approximately 45.5° × 31° of visual angle.
  • 2.3. Eye tracking procedure: Eye movements were recorded with a 1000 Hz infrared EyeLink tracker using five-point calibration and a spatial resolution below 0.5°.Stimuli were displayed at 60 Hz and 1920 × 1080 resolution with a chin rest stabilizing head movements.

3. Dataset statistics & model comparison

CAT2000 contains substantial viewing data with category-dependent center-bias, saccade rates, and inter-observer consistency. Four established saliency models perform unevenly across categories, with particular difficulty on several complex or top-down scenes.

  • Dataset statistics: 24,148,768 saccades were recorded over 240 hours, and center-bias varied substantially across categories.Action, Objects, and Sketch may reflect photographer bias, whereas Random, Jumbled, Satellite, and Social distribute content more broadly.
  • Dataset statistics: The median was around 20 saccades per image during 5 seconds of viewing, with variance about 6 saccades.Low-Resolution, Noisy, Sketch, and Pattern had fewer saccades than Social, Jumbled, Affective, and Cartoon.
  • Observer consistency: Inter-observer consistency was highest for Sketch, Low Resolution, Affective, and Black & White, and lowest for Jumbled, Satellite, Indoor, Cartoon, and Inverted.Consistency was measured by predicting one observer’s fixations from the smoothed fixation map of the other 23 observers using NSS.
  • Observer consistency: High center-bias usually coincided with higher inter-observer consistency.This relationship links category-level fixation concentration with agreement among observers.
  • Model comparison: Four popular models—ITTI, HouCVPR, GBVS, and AWS—performed best on Sketch but poorly on Line drawings.The authors attribute this contrast partly to centered objects in Sketch versus content spread across Line drawings.
  • Model comparison: Social, Satellite, Jumbled, Cartoon, and Inverted were also difficult categories for the evaluated models.The paper discusses top-down social cues, observer boredom, block borders, and upright-image biases as possible category-specific factors.

4. Discussion & conclusion

The paper introduces CAT2000 as a 4,000-image eye-movement dataset spanning diverse categories. It supports model benchmarking, behavioral studies, and investigation of semantic cues guiding gaze.

  • Contribution: CAT2000 contains 4,000 images from a variety of categories and is intended for large-scale eye-movement research.The paper presents this dataset as the central contribution.
  • Applications: The dataset can support model benchmarking, behavioral studies, and investigation of semantic cues guiding gaze during free viewing.The conclusion characterizes the reported analysis as only an initial exploration of these uses.
  • Public release: The public release separates train and test images, with fixation data from 18 observers shared for training and 6 observers held out.Test-image fixations from all 24 observers are held out, enabling evaluation on new observers or unseen images.
Loading 1505.03581v1…