Source-linked AI summary

Gaze360: Physically Unconstrained Gaze Estimation in the Wild

Petr Kellnhofer, Adria Recasens, Simon Stent, Wojciech Matusik, Antonio Torralba

arXiv:1910.10088v1cs.CV

TL;DR

Gaze360 addresses the shortage of large, diverse annotated data for robust 3D gaze estimation in unconstrained scenes. It introduces a scalable collection setup and a temporal model that estimates gaze with uncertainty, then evaluates transfer and real-world attention estimation. The resulting dataset covers 238 subjects across varied indoor and outdoor conditions and supports cross-dataset evaluation and supermarket attention prediction.

  • Problem

    Robust in-the-wild 3D gaze estimation lacks sufficiently large and diverse annotated training data, although gaze is a useful behavioral cue.

  • Method

    Gaze360 combines scalable mobile collection of annotated 3D gaze data with a temporal appearance-based model using quantile regression to estimate gaze uncertainty.

  • Results

    238 subjects were recorded across indoor and outdoor locations, and the dataset and model were evaluated through cross-dataset comparisons, domain adaptation, and unseen-video and supermarket applications.

  • Takeaways & Limitations

    The dataset and model support gaze-based analysis across varied natural scenes, including estimating customer attention in a supermarket.

Abstract

from arXiv · show

Understanding where people are looking is an informative social cue. In this work, we present Gaze360, a large-scale gaze-tracking dataset and method for robust 3D gaze estimation in unconstrained images. Our dataset consists of 238 subjects in indoor and outdoor environments with labelled 3D gaze across a wide range of head poses and distances. It is the largest publicly available dataset of its kind by both subject and variety, made possible by a simple and efficient collection method. Our proposed 3D gaze model extends existing models to include temporal information and to directly output an estimate of gaze uncertainty. We demonstrate the benefits of our model via an ablation study, and show its generalization performance via a cross-dataset evaluation against other recent gaze benchmark datasets. We furthermore propose a simple self-supervised approach to improve cross-dataset domain adaptation. Finally, we demonstrate an application of our model for estimating customer attention in a supermarket setting. Our dataset and models are available at http://gaze360.csail.mit.edu .

1. Introduction

Gaze360 addresses the lack of large, diverse annotated data that limits robust gaze estimation by introducing a scalable dataset-collection method and a model for unconstrained 3D gaze. The work evaluates the dataset and model across domains and demonstrates applications in unseen imagery and customer-attention estimation.

  • Motivation: Gaze estimation trails related human-modeling tasks because sufficiently large, diverse annotated training data are difficult to collect outside the lab.The paper identifies data scarcity and collection difficulty as the primary reasons for the performance gap.
  • Contributions: The authors introduce a methodology for efficiently collecting annotated 3D gaze data in arbitrary environments.The collection method is designed to address the inflexibility and slowness of existing acquisition procedures.
  • Contributions: Gaze360 records video from 238 subjects across indoor and outdoor conditions, making it the largest 3D gaze dataset by subject and variety.The dataset contribution is paired with careful evaluation of its error and characteristics.
  • Contributions: The final model uses multi-frame input and pinball-loss quantile regression to estimate gaze direction and uncertainty.Temporal information helps resolve single-frame ambiguities, while quantile regression provides an error estimate.
  • Evaluation and application: The paper evaluates cross-dataset generalization, introduces self-supervised domain adaptation, and applies Gaze360 to customer-attention estimation in a supermarket.These evaluations test transfer to other datasets and an application setting.

2. Related Work

Prior gaze datasets and models are often physically constrained, lab-based, frontal, or dependent on detected facial features. Gaze360 instead targets varied natural scenes with mobile collection and a model designed for occlusion, temporal evidence, and uncertainty.

  • Datasets: Many existing gaze datasets use static or smartphone-integrated cameras, providing control and accuracy but limited illumination and motion diversity.These constraints reduce their suitability for general applications.
  • Datasets: Natural-scene gaze data should cover broad head and eyeball orientations, including oblique views, rather than constraining subjects to frontal poses.The paper emphasizes coverage of the full range despite increased eye occlusion at large head yaw.
  • Gaze360 collection: Gaze360 uses a mobile multi-camera setup and a free-moving target to collect many subjects across varied lighting, scale, motion blur, and gaze directions.This design more closely approximates robot and surveillance-camera domains than fixed laboratory setups.
  • Gaze360 collection: Compared with a target-free approach requiring gaze-tracking glasses, motion-capture cameras, and semantic in-painting, Gaze360 uses a simpler setup and scales to 238 rather than 15 subjects.The comparison concerns acquisition complexity and subject count.
  • Gaze models: Appearance-based models learn image-to-gaze mappings, while Gaze360 avoids eye or face detectors to improve robustness when required features are partially occluded.Geometric models instead rely on physical assumptions and can be sensitive to occlusion and lighting interference.
  • Gaze models: Gaze360 predicts a gaze direction and uncertainty under partial or full eye occlusion, and aggregates additional frames to recover features visible only intermittently.Its uncertainty estimate is learned through quantile regression with a pinball loss.

3. Dataset collection method

The dataset collection method combines a central 360° camera with a moving AprilTag-marked target board, enabling simultaneous recording of multiple subjects across broad gaze and pose variations. Geometry from camera rays, subject positions, and target pose yields gaze labels in a camera-independent eye coordinate system.

  • Motivation: Existing 3D gaze datasets are typically slow, inflexible, indoor, and small because precise target positioning and constant gaze verification are difficult to relocate or parallelize.The paper motivates a scalable alternative that captures diverse subjects, illumination, head poses, and gaze directions.
  • Setup: A central Ladybug5 360° panoramic camera and a moving rigid AprilTag-marked board allow simultaneous recording of multiple subjects in arbitrary environments.The camera uses synchronized overlapping units, while subjects continuously fixate on a cross attached to the board.
  • Subject positioning: AlphaPose detects head keypoints and feet, and camera rays plus measured camera height estimate subject eye positions in spherical coordinates.The training collection assumes a flat ground plane, although the paper states this is not restrictive at test time.
  • Target positioning: An AprilTag provides the target board’s 3D pose, while the adjacent cross serves as the gaze-fixation target used to locate the target point.Known camera calibration and board geometry support the target-point computation.
  • Gaze direction: The gaze vector is computed from the target point to the eye and converted into a local eye coordinate system whose forward direction aligns with that vector.This convention makes direct camera gaze equal to g = [0, 0, −1] independently of subject position.
  • Collection protocol: A looping board trajectory covers gaze directions around subjects, while vertical motion and an inner path induce larger pitch variation.The board is kept approximately frontoparallel to reduce pose-estimation error.
  • Collection protocol: Alternating “move” and “freeze” instructions vary relative head and eye poses by allowing natural body orientation or restricting movement to the eyes.The protocol broadens the combinations of head and eyeball orientations captured during recording.

4. Gaze360 dataset summary

Gaze360 provides diverse 3D gaze annotations from 238 subjects across indoor and outdoor settings, covering broad gaze directions and head poses. Its acquisition procedure supports large-scale, continuous video collection with measured annotations.

  • Dataset characteristics: The dataset combines 3D gaze annotations, wide gaze and head-pose ranges, indoor and outdoor environments, subject diversity, and short continuous videos at 8 Hz.Gaze360 is surpassed in subject count only by GazeCapture, which provides 2D gaze with a narrower range.
  • Summary statistics: 238 subjects were collected across 5 indoor and 2 outdoor locations over 9 sessions, yielding 129K training, 17K validation, and 26K test images.The dataset includes 58% female and 42% male subjects, with visually diverse ages and ethnicities.
  • Data distribution: Gaze360 covers the full 360° horizontal gaze range and reaches approximately ±140° gaze yaw when one eye remains visible.Vertical coverage is limited by the marker’s achievable elevation, and rear-region sampling is less dense because subjects can occlude the target board.
  • Data distribution: Figure 4 compares logarithmic joint distributions of gaze yaw and pitch across TabletGaze, MPIIFaceGaze, iTracker, and Gaze360 using a Mollweide projection of the unit sphere.The projection visualizes the full spherical gaze domain.
  • Dataset samples: Figure 5 illustrates variation in environment, illumination, age, sex, ethnicity, head pose, and gaze direction using full-body and close-up head crops.Yellow arrows indicate measured ground-truth gaze.
  • Annotation validation: A control experiment using an additional front-facing test camera and AprilTag-based measurement was conducted to validate the accuracy of the gaze annotations.The test camera was mounted above the participant’s right eye, with background AprilTags registering the camera views.

5. Gaze360 model

Gaze360 models gaze from video sequences and predicts both gaze direction and uncertainty. The architecture combines convolutional frame features with bidirectional temporal modeling, spherical-coordinate regression, and pinball-based quantile estimation, with self-supervised adaptation for new domains.

  • Temporal model: The model uses a 7-frame sequence to predict gaze for the central frame, exploiting gaze’s natural temporal continuity.Bidirectional LSTMs use both past and future inputs in the sequence.
  • Architecture: Each head crop is processed by a convolutional backbone, then bidirectional two-layer LSTMs produce features for gaze prediction and error-quantile estimation.The backbone outputs 256-dimensional features before temporal aggregation.
  • Gaze representation: Gaze direction is regressed in spherical coordinates relative to the camera view, with pole singularities assigned to rare strictly vertical directions.This representation is intended to be naturally interpretable for the task.
  • Uncertainty motivation: Because regression outputs cannot use softmax magnitude as confidence, the model explicitly estimates error bounds for difficult views such as partial eye occlusion.The paper motivates uncertainty estimation for sideways views and cases where glasses frames obscure an eye.
  • Error quantile estimation: The model predicts the expected gaze direction and 10% and 90% error quantiles with a pinball loss, producing an 80% error cone in one forward pass.The method assumes isotropy in spherical coordinates, which is less accurate near large pitch angles but simplifies interpretation for most observed directions.
  • Error quantile estimation: The output f(I) = (θ, φ, σ) represents expected spherical gaze angles and a quantile offset, with θ ± σ and φ ± σ defining the 10% and 90% quantiles.The pinball loss drives the angular predictions toward ground truth and σ toward the quantile threshold.
  • Self-supervised adaptation: For domain adaptation, the model is fine-tuned with labeled Gaze360 and unlabeled target images using a feature-domain discriminator and a left-right symmetry consistency loss.The discriminator operates on backbone features, while the symmetry loss compares predictions from original and horizontally flipped images.

6. Experimental Analysis

Gaze360 models are evaluated against simple baselines, temporal variants, and uncertainty-estimation methods on within- and cross-dataset tests. Pinball-based temporal models provide the recommended accuracy–uncertainty trade-off, while cross-dataset adaptation improves transfer further.

  • Within-dataset evaluation: All gaze models outperform mean-gaze and head-pose baselines, which cannot capture the dataset’s rich eye-movement variation.
  • Within-dataset evaluation: Pinball loss generally yields the lowest error and the strongest correlation between predicted uncertainty and actual prediction error, using one forward pass.
  • Within-dataset evaluation: Figure 7 compares ground-truth and predicted gaze, angular errors, and uncertainty, including failure cases where the model is overconfident.
  • Within-dataset evaluation: Temporal modeling substantially improves gaze accuracy over a single-frame static model; TRN and LSTM perform similarly, with Pinball LSTM recommended.
  • Within-dataset evaluation: Accuracy decreases as gaze yaw increases, while the model transitions toward head-pose estimation for rear views and reports higher uncertainty.
  • Cross-dataset evaluation: Cross-dataset testing is harder than within-domain evaluation; Gaze360-trained models perform best consistently, and self-supervised adaptation improves performance on every evaluated dataset.

7. Tracking gaze in the wild

Gaze360 is tested on uncurated online imagery and a supermarket attention task. The model transfers without fine-tuning and can identify looked-at shelf objects, with accuracy depending on viewpoint and occlusion.

  • Prediction in unconstrained environments: The model performs well on unseen YouTube images and videos without additional training or fine-tuning.
  • Estimating attention in a supermarket: In a supermarket shelf task, the system predicts which object a customer is looking at and produces a heatmap of customer attention.
  • Estimating attention in a supermarket: 51% accuracy is achieved from a shelf-side camera, increasing to 68% with a smartphone camera embedded in the shelf.
  • Estimating attention in a supermarket: Bottom-shelf objects have the highest error rate because downward gaze nearly fully occludes the eyes.

8. Conclusion

The paper introduces a scalable collection method, a large diverse dataset, and a temporal gaze model with uncertainty estimation. Evaluations show value for cross-dataset transfer and unconstrained imagery, supporting broader use of gaze for vision-based behavior understanding.

  • The collection approach produces a large, diverse dataset suitable for deep learning of 3D gaze from images and video.
  • The proposed temporal appearance-based model uses a loss function that estimates gaze-error quantiles.
  • Cross-dataset comparisons and application to unseen YouTube imagery demonstrate the dataset’s and model’s value beyond constrained settings.
Loading 1910.10088v1…