Source-linked AI summary
An Eye-Tracking Dataset for Viewing Distance Categories in Real-World Scenarios
Dohwa Kim, Yejin Choi, Seungbok Lee, Chi Yoon Jeong, Eunji Park
TL;DR
Existing gaze-depth datasets often rely on constrained laboratory, static, or VR settings that do not capture natural distance changes. GazeDepth addresses this gap with multimodal wearable eye-tracking recordings across fixed and variable real-world viewing scenarios, and its analyses show distinguishable gaze characteristics across viewing-distance categories.
Problem
Prior gaze-depth studies mainly used constrained settings with fixed heads, stationary targets, or VR, limiting coverage of natural viewing-distance changes.
Method
GazeDepth collected multimodal wearable eye-tracking data from 19 participants across baseline, fixed-distance, and variable-distance tasks spanning near (33 cm), middle (50 cm), and far (300 cm) conditions.
Results
Statistical analyses found significant differences across near, middle, and far conditions, and classification models demonstrated distinguishable gaze characteristics for the three categories.
Takeaways & Limitations
The dataset supports studying distance-dependent gaze behavior and evaluating gaze-based distance inference in realistic scenarios.
Takeaways & Limitations
Outdoor annotations had many unassigned samples and lower agreement, while fixed-distance activities may have introduced task-related confounds and variation in gaze direction.
Abstract
from arXiv · showhide
Estimating viewing distance from gaze behavior is essential for understanding user intent and enabling distance-aware interactive systems. However, most existing eye-tracking datasets have been collected in constrained settings, such as laboratory environments or static tasks. Consequently, they only partially capture viewing behaviors in real-world situations where viewing distance changes with natural head and body movements. We introduce GazeDepth, an eye-tracking dataset collected from 19 participants using a wearable tracker during tasks reflecting real-world scenarios. GazeDepth includes fixed-distance viewing scenarios with constant observer-target distances at near (33 cm), middle (50 cm), and far (300 cm), as well as variable-distance viewing scenarios in which participants shift gaze among targets at different depths in indoor and outdoor environments. The dataset provides synchronized gaze data, pupil size, 3D eye-vectors, and head-motion signals, along with distance labels. Statistical analyses showed that distance-related gaze features, such as vergence angle and estimated viewing distance, differed consistently across viewing-distance categories. In addition, classification models trained on GazeDepth further demonstrated that the dataset captures gaze characteristics that distinguish the three viewing-distance categories, supporting gaze-based distance inference and distance-aware interaction in realistic scenarios.
Background & Summary
GazeDepth addresses the limited realism of prior gaze-depth studies by combining controlled and dynamic viewing-distance scenarios. Its multimodal recordings and analyses show that gaze behavior differs across near, middle, and far categories.
- Motivation: Prior gaze-depth studies often used fixed heads, stationary targets, or VR, limiting coverage of natural distance changes and real-world viewing contexts.Two-dimensional gaze positions can also leave target depth ambiguous when objects share similar viewing directions.
- Dataset design: GazeDepth combines baseline measurements with fixed-distance and variable-distance scenarios designed to capture gaze behavior across realistic viewing conditions.Fixed distances were near (33 cm), middle (50 cm), and far (300 cm), while variable tasks involved natural gaze shifts across depths.
- Dataset contents: The dataset records fixation, saccade, and blink events, pupil size, 3D gaze direction, eye position vectors, and head IMU data from 19 participants.It contains 22.9 hours of recordings and 213,881 fixation events.
- Results: Statistical analyses found significant differences in distance-related gaze indicators across near, middle, and far conditions.The analyzed indicators included pupil size and vergence angle.
- Results: Machine-learning classification further showed that GazeDepth supports subject-independent discrimination of the three viewing-distance categories.The classification analysis evaluated six standard classifiers using gaze features.
Methods
The study protocol was designed to measure distance-related gaze differences while retaining realistic visual behavior. It combined baseline measurements with fixed- and variable-distance tasks.
- Protocol design: The protocol targeted measurable differences in gaze characteristics and transitions across distances while providing scenarios resembling real-world viewing.These choices were intended to capture gaze dynamics underlying natural visual attention and gaze shifts.
- Task structure: Data collection included baseline measurements and five main tasks organized into fixed-distance and variable-distance scenarios.The baseline condition was relaxed and intended to capture individual differences in gaze characteristics.
- Fixed-distance viewing: Fixed-distance scenarios kept target distance constant at near, middle, and far positions while using realistic activities such as book reading and monitor viewing.These tasks were intended to compare gaze characteristics across distance categories in realistic contexts.
- Ethics: The study received institutional ethics approval, and participants provided written consent before data collection.The approved consent process covered procedures, collected data, protection protocols, and compensation.
<Baseline Measurement>
The experimental protocol began with baseline measurement before the distance-viewing tasks. Baseline activities included dot fixation, surface viewing, and multi-object viewing.
- Protocol phase: Baseline measurement was the first phase of the protocol, preceding fixed-distance and variable-distance viewing tasks.The full protocol comprised three phases: baseline measurement, fixed-distance viewing, and variable-distance viewing.
- Baseline activities: Baseline measurement included dot-fixating, surface-viewing, and multi-object viewing activities.These activities established reference gaze patterns before the main distance-viewing scenarios.
- Participant rights: Participants were informed that they could withdraw from the study at any time after providing written consent.The consent process was reviewed and approved by the university ethics center.
Participants
The study recruited adult participants without reported ophthalmic or relevant mobility and fixation difficulties. Participants used a glasses-type eye tracker during everyday visual tasks lasting up to two hours.
- Recruitment: The study recruited 19 adults through university and local information-sharing platforms.The sample included 9 participants aged 20–39 and 10 aged 40–59, with mean age 39.21 years (SD = 11.83).
- Eligibility: Eligibility required no ophthalmic disorders and no difficulty with object fixation or walking during data collection.Participants had corrected visual acuity of 1.0 or higher or used supported corrective lenses.
- Data-collection conditions: Participants wore a glasses-type eye tracker and were told that sessions could last up to 2 hours.They were also asked to remove makeup around the eyes and informed about possible mild physical discomfort.
- Experimental setup: Figure 2 depicts the experimental setup, including the eye-tracking device, participant positioning, and task workflow.The passage identifies these setup elements but does not provide additional participant-specific measurements.
Data Collection Setup
Indoor data collection used controlled classrooms, fixed target distances, and a wearable eye tracker connected to a smartphone. The tracker recorded gaze events that were automatically processed and stored in Pupil Cloud.
- Indoor tasks were conducted in two classrooms under controlled illumination, with target-object distances fixed for baseline and fixed-distance scenarios.
- Participants wore a Pupil Labs Neon eye tracker connected to a smartphone, with corrective lenses attached when necessary.The setup was checked to ensure the tracker did not obstruct the participant’s field of view.
- Gaze events such as fixations and saccades were automatically detected, labeled, and stored in Pupil Cloud.
Data Acquisition Procedure
Data collection proceeded from onboarding and baseline measurements through fixed- and variable-distance tasks, combining controlled viewing conditions with realistic gaze transitions. Participants completed reading, computer, screen-based, board-game, and outdoor exploration activities, followed by a post-survey.
- Data Collection Procedure: Data collection comprised onboarding, baseline measurement, fixed-distance scenarios, and variable-distance scenarios, with two-minute breaks between task sessions.Participants removed and reattached the eye tracker when a location change was required.
- Calibration: Calibration and offset correction were applied per task when feasible, but outdoor walking used the Neon tracker’s native gaze-estimation pipeline without an additional initial calibration step.
- Baseline Measurements: Baseline measurements included dot fixation, surface viewing, and multi-object viewing to characterize individual gaze behavior and oculomotor stability.The dot-fixating task used near (33 cm), middle (50 cm), and far (300 cm) positions.
- Fixed-distance Viewing Scenarios: Fixed-distance tasks included natural book reading at approximately 25–35 cm, computer tasks on a monitor 50 cm away, and target identification on a screen 300 cm away.The computer tasks involved account creation and instructed web browsing, while the far-distance task used images and videos.
- Variable-distance Viewing Scenarios: Variable-distance tasks required gaze shifts among near, middle, and far targets during an indoor board game and an outdoor route-based exploration task.The indoor task involved manipulation objects, an opponent, and an instruction screen; outdoor exploration covered approximately 550 meters and lasted at least 10 minutes.
- Post-survey: A post-survey assessed concentration, frequently fixated objects, and perceived similarity between the experimental and real environments.
Viewing Distance Annotation
GazeDepth assigned viewing-distance labels to variable-distance data using scene-based automatic annotation, then evaluated the labels against manual fixation annotations. Indoor labeling used spatial zones, whereas outdoor labeling matched gaze coordinates to distance-specific object detections.
- Purpose: Automatic labeling was introduced to support modeling and evaluation of viewing-distance categories in variable-distance tasks.
- Indoor Automatic Annotation: Indoor annotation divided each frame into near, middle, and far zones and assigned the zone containing the gaze coordinate as the label.The zones were defined using recurring scene objects including the plate, opponent, and table, with camera-distortion and head-pose correction based on the table mask.
- Indoor Exception Handling: 81.36% of indoor annotation frames contained both reference objects, while lady-only detections accounted for 17.73%; plate-only and no-object cases were each below 1%.A plate-position estimation method was developed to address the relatively frequent lady-only case.
- Indoor Annotation Results: Indoor frames were labeled near 59.30%, middle 16.59%, far 21.45%, and unassigned 2.66% on average.
- Outdoor Automatic Annotation: Outdoor annotation matched gaze coordinates to bounding boxes of predefined near-, middle-, and far-distance objects, selecting the smaller box when detections overlapped.
- Outdoor Annotation Results: Outdoor frames were labeled near 30.15%, middle 10.97%, far 24.66%, and unassigned 34.21% on average.The higher unassigned proportion reflected frequent looks toward task-irrelevant objects during walking.
- Manual Validation: Manual validation used fixation events from six indoor and six outdoor participants, with three annotators independently labeling each selected task.Annotators labeled the central frame of each fixation and propagated the label across the fixation interval using near, middle, far, none, and skip categories.
Data Records
GazeDepth organizes task-specific recordings, metadata, synchronized gaze and motion signals, viewing-distance labels, and questionnaire data into a structured directory. Files can be linked through recording and section identifiers, while videos and sensor streams use specified sampling formats.
- Directory Structure: The dataset directory contains Baseline, Fixed-distance_viewing, Variable-distance_viewing, and TaskResult subfolders organized by task category.Supplementary Information describes the detailed file formats, data fields, and recording exceptions.
- Recording Files: Each recording folder contains a front-facing MP4 scene video, two metadata JSON files, and eight CSV data files.A sections.csv file is also provided within each task’s timeseries-data and scene-video folder.
- Sensor Records: CSV files contain timestamps, task events, gaze coordinates, fixation, saccade, and blink events, binocular 3D eye states, pupil sizes, and IMU-based head-motion data.Records across files are linked using recording and section identifiers.
- Sampling Formats: Scene videos were recorded at 60 fps and 1600 × 1200 pixels, while gaze and IMU signals were sampled at 200 Hz.
- Event Detection: Fixation and saccade events were automatically detected in Pupil Cloud v7.2 and stored in dedicated CSV files.The fixation detector extends the I-VT algorithm and includes compensatory eye movements during head or body motion.
- Screen Coordinates: Screen-based task folders include transformed gaze coordinates that estimate participants’ points of gaze on recorded screens.
- Distance Labels: Variable-distance folders provide participant-specific indoor and outdoor files containing automatically generated near-, middle-, and far-distance labels aligned with scene data.
- Questionnaire and Task Results: XLSX files provide pre- and post-survey responses, task instructions, and subtask results for selected fixed- and variable-distance tasks.
Technical Validation
Technical validation found stable gaze recordings, reasonably accurate automated labels, and distance-related gaze features that consistently separated near, middle, and far conditions. Classification retained meaningful performance across tasks, while participant-specific vergence variability indicates a need for calibration or adaptation.
- Gaze Data Quality: 0.18% of timestamp intervals exceeded 5.5 ms, and average participant-level data loss was 0.18% (SD = 0.06), indicating stable near-200 Hz recording.The highest participant-level data loss rate was 0.37%.
- Label Validation: Automated labels matched majority-vote annotations with accuracy of 0.905 (SD = 0.07) indoors and 0.822 (SD = 0.05) outdoors.F1-scores were 0.881 (SD = 0.05) indoors and 0.81 (SD = 0.05) outdoors.
- Statistical Analysis: Both vergence angle and geometry-based viewing-distance estimates differed significantly across near, middle, and far conditions, with all pairwise comparisons significant at p < .01 in every setting.Normality was assessed before selecting repeated-measures ANOVA or Friedman tests with paired post hoc comparisons.
- Distance-Dependent Vergence Characteristics: Vergence adjusted after viewing-distance transitions within a few hundred milliseconds, supporting the use of a 0.5-s analysis window.Transitions involved a saccade and vergence adjustment toward the new distance category.
- Limitations: Participant-specific vergence variability and decision boundaries suggest that absolute-vergence classifiers may require initial user calibration.Calibration would establish user-specific distance distributions and decision boundaries.
- Viewing Distance Classification: 0.966 was the highest fixed-distance accuracy for LDA with a 10 s window, while RF reached 0.827 indoors and 0.763 outdoors with a 0.5 s window.These results were interpreted as offline validation because session statistics were available for feature normalization.
- Ablation Study: Excluding optical-axis y features yielded mean accuracy of 0.785 (SD = 0.07), while left-eye-only and right-eye-only features reached 0.802 (SD = 0.04).Classification therefore did not rely predominantly on vertical gaze information or binocular information alone.
- Cross-Task Generalization: Inter-task accuracies were 0.750 for fixed-distance, 0.601 for indoor, and 0.555 for outdoor testing, remaining above majority-class baselines despite an 18.6 % point mean decrease.The results indicate that models retained viewing-distance-related information beyond task-specific characteristics.
Usage Notes
GazeDepth supports benchmarking distance-related gaze models and potential applications in interaction and vision training, but users should account for categorical labels, outdoor annotation challenges, and task-related confounds.
- Benchmarking: GazeDepth can benchmark distance classification, human behavior classification, and intention recognition in real-world environments.The outdoor route-finding task also supports benchmarking visual attention, spatial cognition, and environmental exploration during navigation.
- Interactive systems: Distance estimation models could adapt AR/MR and context-aware assistance to whether users inspect nearby objects, explore workspaces, or monitor farther surroundings.Distance information could also support finer-grained activity classification.
- Vision training: GazeDepth could support vision-training validation by measuring stable focus shifts and delayed transitions across target distances.The proposed uses include vergence-training metrics, Brock string exercises, and monitoring near-distance focusing difficulties.
- Usage boundaries: Variable-distance outdoor annotations had lower agreement, while categorical labels are less suitable for estimating exact continuous viewing distances.Future annotation could use frames before and after fixation events to provide more temporal context in complex outdoor scenes.
- Usage boundaries: Different activities across fixed-distance conditions may introduce task-related confounds, although distance information remained informative after reducing gaze-direction cues.Optical-axis y features may reflect posture and target placement, and participant factors may affect generalizability and recording quality.
Code Availability
The paper provides processing code and describes dataset-resource documentation, participant and labeling tables, automatic-labeling cases, and classification-performance reporting.
- Code: Processing code for automated annotation and anonymization is available through the project’s GitHub repository.The repository documents instructions for downloading the required model.
- Dataset documentation: Table 1 summarizes eye-tracking datasets by conditions, distance-estimation use, participants, data types, devices, and setting.Viewing distance is encoded as continuous range R or discrete distance D.
- Dataset documentation: Table 2 reports participant demographics and ophthalmic characteristics for N = 19 participants.
- Annotation: Table 3 defines automatic-labeling cases from reference-object detection status and reports frame counts and proportions before and after plate-position estimation.
- Evaluation: Table 4 reports Near/Middle/Far classification accuracy and F1-score across three tasks, repeated runs, and window-size settings as Mean (SD).
Author Contributions Statement
D.K. and Y.C. contributed equally, with the authors distributing responsibilities across study design, acquisition, annotation, analysis, writing, review, and oversight.
- Study design and acquisition: D.K. and E.P. conceptualized the study and designed the research, while D.K. contributed to experimental sessions and data acquisition.
- Annotation: D.K. and Y.C. handled data annotation and annotation-tool development, with D.K. focusing on outdoor work and anonymization and Y.C. on indoor work.
- Verification and analysis: D.K., Y.C., and S.L. verified and edited the dataset for anonymization, statistical analysis, and modeling, while Y.C. focused on annotation validation.
- Oversight: E.P. oversaw research planning and execution, including mentoring and funding acquisition, while all authors verified the results and reviewed the final manuscript.
- Writing and review: D.K. and Y.C. wrote the original draft with support from S.L., and the broader author team reviewed and edited the manuscript.