Source-linked AI summary
Drift Calibration in Geometric Eye Tracking Systems
Jiaqi Liu, Zixuan Wang, Yuhong Zhang, Dingkang Liang, Jane Hanqi Li, Tzyy-Ping Jung, Gert Cauwenberghs
TL;DR
Residual calibration error makes gaze-based behavioral and interactive systems difficult to evaluate consistently. The paper builds a calibration-focused benchmark with disjoint fitting and test grids, compares correction families, and adds a neural refiner. Mean error falls from 1.53° to 1.03° with the best classical composite and 0.96° with the refiner, while lower residual error is associated with better closed-loop performance.
Problem
Different devices, target layouts, error definitions, and evaluation splits make post-vendor gaze-correction methods difficult to compare.
Method
The paper releases 163 trials from 12 participants with disjoint 18-point fitting and 32-point test grids, evaluates global, local, and composite corrections, and combines ranked calibrator predictions with a neural refiner.
Results
1.03° mean error was achieved by the strongest classical composite and 0.96° by the neural refiner, while lower residual error was associated with higher closed-loop task performance.
Takeaways & Limitations
The benchmark provides a reproducible basis for evaluating post-vendor gaze correction and its relationship to closed-loop interaction performance.
Takeaways & Limitations
Evaluation was limited to one device and a controlled within-population setting, so cross-user and cross-device validation remains future work.
Abstract
from arXiv · showhide
Geometric eye trackers can provide the spatial accuracy required for gaze-based interaction and multimodal studies, but their measurements remain sensitive to residual session-specific calibration error. Research on correcting this error is difficult to compare because methods are typically evaluated with different devices, target layouts, and error definitions. We present a calibration-focused dataset containing 163 trials from 12 participants, with separate 18-point fitting and 32-point test grids, and use it to evaluate global, local, and composite correction functions under a common spatial-extrapolation protocol. We further introduce a lightweight neural refiner that combines ranked predictions from complementary calibrators. On this controlled dataset, post-vendor correction reduces the mean angular error from $1.53^\circ$ to $1.03^\circ$ with the strongest classical composite and to $0.96^\circ$ with the refiner. In a closed-loop gaze task, lower residual error is associated with higher performance across four online correction conditions. These results provide a reproducible data-quality benchmark for using gaze as a behavioral signal in interactive modeling.
1 Introduction
Gaze is valuable for behavioral and interactive systems, but residual spatial offsets can mislabel behavior or trigger incorrect actions. The paper addresses inconsistent comparisons of post-vendor correction methods with a common protocol and a neural refiner.
- Residual spatial offsets can change behavioral labels or actions in gaze-based systems.Gaze quality must be judged relative to the size and spacing of task-relevant units.
- Appearance-based and geometric gaze estimators use different sensors and benchmarks, making reported angular errors difficult to compare.
- Calibration studies conflate initial mapping, post-vendor correction, implicit updating, and appearance-model personalization.These settings differ in devices, layouts, error definitions, and evaluation splits.
- The paper introduces a common empirical basis, an open 163-trial dataset, systematic calibrator comparisons, and a lightweight neural refiner.The refiner combines predictions from seven complementary calibrators and reaches 0.96° mean error.
2 A Calibration-Oriented Dataset
The dataset is designed to measure residual drift after vendor calibration and to test spatial generalization rather than recovery of fitted calibration points. It combines disjoint grids, repeated trials, and complete fixation samples.
- The dataset contains 163 trials from 12 participants using a monocular head-mounted tracker at a fixed 70 cm viewing distance.Each session used a 24-inch 1920 × 1080 display after standard five-point vendor calibration.
- Each trial uses an 18-point fitting grid and a disjoint 32-point test grid, preventing evaluation at fitted locations.This makes the evaluation a test of spatial generalization.
- The acquisition design discards the first 2 s of each 6 s target presentation because over 95% of targets stabilized within 1.83 s.
- The dataset retains full fixation sample clouds, separating within-fixation dispersion from centroid offset.Dispersion reflects noise and precision, whereas centroid offset reflects drift and accuracy.
3 Combining Classical Calibrators with Neural Refinement
The correction pipeline predicts a session-specific drift field at unseen screen positions by comparing global, local, and composite calibrators, then combining their ranked predictions with a neural refiner.
- Tracker estimates and known target locations define displacement observations that calibrators extrapolate to unseen positions.The corrected gaze is obtained by adding the predicted displacement to the raw gaze position.
- Similarity and polynomial models impose global smoothness, while kernel methods allow more local variation.
- Composite models remove a global similarity component before modeling the remaining residual with a flexible calibrator.This probes the tradeoff between stable sparse-data extrapolation and position-dependent flexibility.
- The refiner ranks calibrator predictions by fitting-grid residual error and combines them with query position in a multilayer perceptron.Training perturbs hypothesis channels with Gaussian noise to improve robustness to unreliable base calibrators.
4 Results
Post-vendor correction reduces residual gaze error, with the neural refiner outperforming the strongest classical composite. Improvements are especially valuable under sparse calibration, and lower residual error is associated with better closed-loop task performance.
- 1.03° mean error was achieved by Similarity+RBF, versus 1.53° for raw tracker output.The best classical composite reduced mean error by 32.3%.
- 0.96° mean error was achieved by the neural refiner, improving over the best classical composite by 7.0%.The refiner also reduced standard deviation by 21.0% relative to that composite.
- 42.3 px error with all seven hypotheses improved on 44.5 px with one hypothesis and 43.1 px with three.The full model improved by 4.9% over one hypothesis and 1.9% over three.
- With six calibration points, the refiner reached 46.7 px versus 51.4 px for the strongest classical composite, a 9.1% reduction.Hypothesis perturbation contributed a 10.9% gain in the six-point condition.
- Across four online correction conditions, task score increased with calibration quality and was negatively associated with residual angular error.The reported trend indicates that differences near one degree remain relevant to real-time interaction.
5 Conclusion
The benchmark shows that residual post-vendor gaze error can be reduced substantially with classical composite correction and a lightweight neural refiner, while lower residual error corresponds to better closed-loop task performance. The evaluation is controlled but limited to one device and a young, controlled population, motivating broader validation.
- Conclusion: 0.96° mean error was achieved by the seven-hypothesis refiner, versus 1.03° for the best classical composite and 1.53° for raw output.The refiner further improved mean error by 7.0% over the best classical method and reduced error variability by 21.0%.
- Conclusion: A 37.2% reduction from raw output was achieved by the best correction pipeline, with the strongest classical composite reducing mean error to 1.03°.The corresponding raw and corrected errors were 1.53° (67.3 px) and 1.03° (45.5 px).
- Conclusion: Six-point error fell by 9.1% over the strongest classical composite, while perturbation contributed a 10.9% gain.The refiner was particularly effective when calibration was sparse.
- Conclusion: Lower residual error coincided with higher game score across four closed-loop gaze-control correction conditions.The reported relationship connects offline calibration quality with online task performance without establishing causation.
- Conclusion: The reported gains require validation beyond 12 young adults, one device, and controlled laboratory conditions before application to other populations, devices, or deployments.The authors release a benchmark intended to support future cross-user and cross-device validation.
B.2 Full affine transform
The calibration families span global, local, and composite models for predicting spatial drift from sparse correspondences. Their trade-offs involve flexibility, extrapolation behavior, interpolation, uncertainty, and variance under limited calibration data.
- Full affine transform: The full affine model adds anisotropic scale and shear to translation, using six parameters.It represents the correction as C_aff(p) = Ap + t − p.
- Polynomial regression: With N = 18 and d4 = 15, the order-4 polynomial fit has only three more observations than coefficients per output, increasing estimation variance.Polynomial models regress drift components independently per output axis with ridge regularization.
- Radial basis functions: RBF interpolation represents the drift field as weighted kernels centered on calibration targets, with weights fit to observed drifts.The unsmoothed form interpolates calibration points exactly, while smoothed variants need not.
- Radial basis functions: RBF variants are evaluated away from the calibration grid because flat systems can be poorly conditioned and extrapolation depends on target geometry.The study evaluates smoothing values 0, 1, and 2 for the multiquadric kernel.
- Gaussian process regression: Gaussian-process regression permits nonzero training residuals, can reduce sensitivity to fixational noise, and provides position-dependent uncertainty.The study does not use posterior variance as an input to the refiner.
- Piecewise-affine interpolation: Piecewise-affine interpolation is continuous but not differentiable across triangle edges and is applied only inside the convex hull.Outside the hull, the implementation returns the target associated with the nearest control point.
B.8 Composite calibrators
Composite calibration first fits a global anchor and then models its residual, combining stable global structure with flexible local correction. On this dataset, similarity+RBF outperforms similarity and other reported single-family baselines, although the result does not establish why.
- B.8 Composite calibrators: Composite calibration first fits a global anchor, then models the residual with a flexible calibrator.For similarity+RBF and similarity+GPR, the residual model is fitted at the raw gaze coordinates.
- B.8 Composite calibrators: Similarity+RBF has lower error than similarity and the other reported single-family baselines in Table 1.
- B.8 Composite calibrators: The observed advantage is consistent with residual modeling but does not establish why composite calibration performs better.
C Observed Drift Fields
Figure 5 depicts structured drift as raw gaze estimates displaced from targets, while recovered fields extrapolate differently beyond calibration points. Cross-study calibration results are not directly comparable because apparatus and error definitions differ.
- C Observed Drift Fields: Raw gaze estimates are displaced from targets, and the fixation cloud marks sample dispersion around the observed gaze.
- C Observed Drift Fields: Four recovered fields extrapolate differently beyond the calibration points.
- C Observed Drift Fields: Calibration-related values from different studies should not be compared directly because devices, target layouts, and error definitions differ.
E Acquisition Protocol
Each session separates an 18-point fitting phase from a 32-point test phase, with optional closed-loop evaluation. Targets appear sequentially for 6 seconds, and gaze from the first 2 seconds is discarded.
- E Acquisition Protocol: Each session uses an 18-point fitting grid and a disjoint 32-point test grid before any optional closed-loop task.Targets are presented at disjoint screen positions.
- E Acquisition Protocol: Targets illuminate sequentially in random order for 6 s each.
- E Acquisition Protocol: The protocol discards the first 2 s and records gaze during the remaining 4 s.
- E Acquisition Protocol: The closed-loop task applies different calibration functions to the live gaze stream while scoring fixation duration inside illuminated circles.
F Fixation Acquisition Time
Fixation acquisition is generally rapid: over 95% of targets are acquired within 1.83 seconds, supporting the protocol’s 2-second discard window.
- F Fixation Acquisition Time: Over 95% of targets are acquired within 1.83 s.The estimate comes from more than 2,000 collected points using a stable region defined around each target’s mean fixation position.
- F Fixation Acquisition Time: The acquisition-time result supports discarding the first 2 s of each target presentation.
G Refiner Ablations
The refiner ablations examine how training perturbations and the number of ranked inputs affect refinement performance.
- Table 3 evaluates hypothesis perturbation during refiner training across calibration densities.
- Table 3 compares how many ranked hypotheses are supplied as inputs to the refiner.