Source-linked AI summary
Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans
Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Ulas Bagci, Alessandro Bruno
TL;DR
Whether MLLMs search through human-matched foveated input as humans do remains unclear, limiting their interpretation as models of human vision. This paper compares fixation-driven MLLM and human search across decision, target acquisition, and gaze process, finding that matched outcomes coexist with non-human scanpaths.
Problem
It remains unclear whether MLLMs given human-matched foveated input reproduce human search across decisions, target acquisition, and gaze processes.
Method
The study drives three general-purpose MLLMs fixation by fixation through human-acuity-calibrated foveated views and compares their decisions, target acquisition, and scanpaths with humans.
Results
All three models match or exceed humans on decisions and target acquisition, but share low-entropy, large-amplitude, self-consistent scanpaths that remain non-human across degradation regimes.
Takeaways & Limitations
Answer-alignment and saliency scores cannot certify human-like vision, making zero-shot MLLMs suitable for outcome and spatial questions but not temporal, process-level ones.
Takeaways & Limitations
The study evaluates general-purpose models under a fixed foveation constraint and does not assess search-trained agentic, pointing-native, or frontier closed-source models.
Abstract
from arXiv · showhide
Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal large language models (MLLMs), given the same foveated input, search as humans do bears on their use as models of human vision and on attention-alignment scores. We compare three general-purpose MLLMs with human eye-movement scanpaths on goal-directed search (COCO-Search18), driving each model fixation by fixation through an identical, human-matched foveated view and assessing it along three axes: the decision of target presence, the efficiency of reaching the target, and the gaze process itself. The axes dissociate. On the decision and on target acquisition the models match or exceed humans, detecting present targets near ceiling and reaching them on the first saccade more often than people do. The gaze process is not human. Under the human-matched condition, all three share one signature: low-entropy, large-amplitude, self-consistent scanpaths that agree with themselves far more closely than two humans agree with each other. That is consistent with a single-pass, non-serial architecture rather than a limit of acuity. Matched retinal input reproduces where humans look but not how the looking unfolds in time, and no degradation regime recovers human-like search at human-like success. The gap sits on a process axis that answer-alignment and saliency metrics do not measure. Because they miss it, such metrics cannot certify human-like vision, and zero-shot models suit outcome and spatial questions but not temporal, process-level ones.
1 Introduction · 2 Related work
The paper tests whether human-matched foveated input makes MLLM visual search human-like, separating target decisions, target acquisition, and gaze process. Models match or exceed humans on outcomes but produce non-human, low-entropy and highly self-consistent scanpaths, challenging attention-alignment as evidence of human-like vision.
- 1 Introduction: Human goal-directed search is serial because peripheral detections require foveal fixation for confirmation.The fovea resolves fine detail only within approximately 1–2° of gaze, making search a sequence of saccades and fixations.
- 1 Introduction: Matched foveated input provides a falsifiable test of whether MLLMs can serve as human observers and whether attention overlap measures human-like vision.The model sees sharp detail at gaze, degraded peripheral detail, and must re-fixate to displace the aperture.
- 1 Introduction: The study separates search into decision, target finding, and gaze axes, which can dissociate rather than yield a single measure of human-likeness.The comparison uses COCO-Search18 and drives each model fixation by fixation through an identical human-matched foveated view.
- 1 Introduction: 0.97/0.97/0.80 first-saccade target fixation versus 0.49 for humans shows that models exceed people in early target acquisition while maintaining comparable eventual success.Present-target detection is also near ceiling, and models require no more fixations than humans.
- 1 Introduction: 0.84/0.91/0.71 cross-seed ScanMatch versus the 0.53 inter-observer ceiling indicates that models agree with themselves far more than humans agree with one another.Their gaze signature is low-entropy, large-amplitude, and highly self-consistent, despite matched retinal input.
- 2 Related work: Human-search research spans scanpath prediction, inverse reinforcement learning, transformers, stochastic generators, and ideal-observer or Bayesian searchers using COCO-Search18.Classical accounts combine parallel peripheral evaluation with serial focal inspection and explain search asymmetries through natural-image statistics.
- 2 Related work: Foveation is established both as a model of human visual representation and as a search mechanism, with the Geisler–Perry acuity falloff supplying the human-matched renderer.Prior work includes emergent foveated perception, central–peripheral scene recognition, variable-resolution representations, and foveal detectors trained for search.
- 2 Related work: Prior MLLM studies report dissociations between attention similarity, task performance, spatial guidance, and human-like processing, while this work tests general-purpose models under fixed foveation without search-specific training.Agentic systems instead optimize where to look through LLM guidance, zooming, reinforcement learning, or embodied search.
3 Method · 4 Results
The study evaluates three foveated MLLMs on human-matched visual search, separating decision accuracy, target acquisition, and gaze dynamics. Models match or exceed human outcomes but share a non-human, highly consistent gaze process, and no degradation level restores human-like search without sacrificing success.
- 3 Method: The comparison bracket included sharp, human-acuity-calibrated geisler–perry, synthetic gaussian gist-k, and crop conditions, with only geisler–perry human-matched.At the study’s viewing geometry, geisler–perry appeared mild and nearly indistinguishable from sharp.
- 3 Method: Each episode began at a forced central fixation; the model received a gaze-centered rendered glimpse and returned LOOK, FOUND, or ABSENT until termination.The loop imposed no search policy, and the glimpse cap never forced termination.
- 3 Method: The evaluation used 285 COCO-Search18 scenes, five seeds, and nine foveation conditions for each of three MLLMs, yielding 38,475 completed episodes.The human reference comprised ten observers per scene, with 141 target-present and 144 target-absent validation scenes.
- 4.1 Decision and finding: models match or exceed humans: At the human-matched GP condition, one-shot existence accuracy was 0.99/0.81, 0.99/0.81, and 0.97/0.80 for target-present/target-absent scenes across Qwen, GLM, and Gemma.Agentic present/absent decisions had d′ 3.84/4.23/3.14 against the human 2.91, while absent scenes were declared absent at 0.93/0.91/0.96.
- 4.1 Decision and finding: models match or exceed humans: At GP, first-saccade target fixation was 0.97/0.97/0.80 for Qwen/GLM/Gemma, exceeding the human 0.49.Finding and outcome measures matched or exceeded human values on existence-passed trials.
- 4.2 Gaze dynamics: a shared, non-human signature: At GP, all three models showed lower gaze entropy and larger saccade amplitudes than humans, with Cliff’s δ entropy −.67/−.65/−.27 and amplitude +.50/+.61/+.23.Their scanpaths were highly self-consistent, occupying a region disjoint from the human reference; entropy and amplitude effects had 95% intervals excluding zero for all models.
- 4.3 No regime recovers human-like search; a single-pass account: Synthetic degradation did not produce human-like search: when TFP@1 approached the human rate, eventual success had already collapsed, with Qwen at k=32 reaching 0.42/0.71.At intermediate k16 and k24, gaze partly converged while target acquisition declined, so convergence was not evidence of human-like search.
- 4.4 Generality and robustness: The outcome–gaze dissociation held across three model families and difficulty strata, while Qwen’s high cross-seed determinism remained above the inter-observer ceiling even at temperature 1.0.Cross-seed and human inter-observer agreement are not identical constructs because no matched human intra-observer baseline exists.
5 Discussion
The models reproduce answer outcomes and partially match human saliency maps, but diverge on every temporal gaze measure. Thus, zero-shot MLLMs serve as surrogates for outcome and spatial questions, while their non-serial behavior provides a null model for isolating seriality.
- Metrics and surrogacy: Density correlations of 0.58/0.63/0.50 indicate partial saliency-map matching, while the models diverge on every temporal gaze measure.Answer-alignment and saliency scores capture outcomes or time-collapsed maps, making high values necessary but insufficient for human-like process alignment.
- Metrics and surrogacy: Zero-shot MLLMs are adequate surrogates for outcome and spatial questions but not for process and temporal ones.This limitation reflects the class of models rather than model selection alone.
- A non-human searcher as a null model: A system achieving human-or-better outcomes without a serial bottleneck provides a useful null model for isolating seriality from detection and spatial priors.The null model preserves reproduced detection and spatial behavior while removing the serial constraint.
6 Limitations and scope · 7 Conclusion · Supplementary Material
The study finds that foveated MLLMs match or exceed humans on search outcomes while exhibiting non-human gaze dynamics. Its conclusions are limited to general-purpose models and supported by supplementary methodological and per-model analyses.
- 6 Limitations and scope: The study evaluates general-purpose models under a fixed foveation constraint, excluding search-trained agentic, pointing-native, and frontier closed-source models.Cross-model trends are descriptive and confounded by architecture, recipe, and sparsity.
- 6 Limitations and scope: The gaze signature is measured under one prompt whose memory clause may shape refixation, and fixation durations are available only for humans.These design choices limit interpretation of the process comparison.
- 7 Conclusion: Three MLLMs match or exceed humans on target-presence decisions and target acquisition under human-calibrated foveation.At legible conditions, all three share low-entropy, large-amplitude, highly self-consistent gaze; no degradation regime restores human-like search with human-like success.
- 7 Conclusion: Matched retinal input reproduces where humans look but not how looking unfolds, consistent with a single-pass reader carrying a human-like spatial prior.When the signature converges under degradation, models are failing to resolve the scene rather than searching.
- 7 Conclusion: Answer-alignment and saliency scores cannot certify process-level correspondence because they evaluate outcomes or time-collapsed maps.A searcher achieving human-or-better outcomes without a serial bottleneck provides a null model for interpreting such metrics.
- Supplementary Material: The supplementary material specifies the stimuli, foveation model, search procedure, behavioural metrics, statistical methods, and complete per-model results.It organizes these materials across Secs. S1–S6, covering decision, finding, and gaze.
S1 Stimuli and the foveation model · S2 Search procedure
The study uses COCO-Search18 scenes and human scanpaths with deterministic foveation rendered at each model gaze point. Models search fixation by fixation from a forced central fixation using explicit look, found, or absent directives across four visual conditions.
- S1 Stimuli and the foveation model: COCO-Search18 supplies 1680×1050 px scenes, cued-object presence judgments, and human scanpaths for trials defined as scene–target pairs.Scenes subtend ∼54°×35° of visual angle at an angular resolution of ρ ≈30 px/deg.
- S1 Stimuli and the foveation model: Visual angle is computed throughout as θ = d/ρ for a pixel distance d.
- S1 Stimuli and the foveation model: Foveation is imposed by a deterministic renderer centered on the model’s current gaze point g.The renderer follows the Geisler–Perry approach, relating resolvable spatial frequency to retinal eccentricity e.
- S1 Stimuli and the foveation model: A Gaussian image pyramid selects a local level at each pixel so the image cutoff matches the eye’s, producing a sharp centre and smooth peripheral falloff.
- S1 Stimuli and the foveation model: The experiment brackets vision with sharp no-foveation, human-matched Geisler–Perry, synthetic Gaussian gist-k, and fovea-only crop conditions.The human-matched GP condition uses ρ=30 and a viewing distance of 0.6 m; gist-k varies k ∈{8, 16, 24, 32, 48, 128}.
- S1 Stimuli and the foveation model: The renderer is identical across models, while crop retains only a fovea-only disc of radius ∼2.5° and removes the periphery.
- S2 Search procedure: Each episode starts at a forced central fixation, after which the model receives the current rendered view while retaining earlier glimpses in context.
- S2 Search procedure: At every step, the model must return exactly one directive: look(x, y), found(x, y), or absent.look continues search at a new gaze point; found terminates with a present decision at (x, y), and absent terminates with an absent decision.
S3 Metrics and statistical methodology
The study separates decision, target-acquisition, and gaze-process outcomes using defined scanpath, fixation, similarity, density-map, and statistical measures. Effects are quantified against humans with Cliff’s δ, mixed-effects models, and a joint PCA of gaze statistics.
- Scanpath and task metrics: Target hits are defined by fixation within the dilated target box, with first-hit index h(s) recording when the target is first fixated.The default tolerance is τ=1°=ρ pixels; h(s)=∞ when the target is never fixated.
- Scanpath and task metrics: Existence accuracy, yes-bias, and signal-detection measures quantify present/absent decisions from correct answers, false positives, hit rate, and false-alarm rate.The human reference uses recorded gamepad present/absent responses.
- Scanpath and task metrics: Target-fixation probability reports reaching the target after each saccade, including TFP@1 for first-saccade targeting and TFP-end=TFP15 for eventual success.The estimator is applied identically to models and humans, across strata and hit-tolerance analyses; NumFix is the per-group median fixation count.
- Gaze-process metrics: Gaze-process statistics include saccade amplitude and direction, turning angle, scanpath length, center bias, convex-hull coverage, gaze entropy, and refixation rate.Statistics are summarized by per-group medians and effect sizes against the human distribution; refixation measures landings in previously visited grid cells.
- Agreement metrics: ScanMatch quantifies pairwise scanpath similarity on a 14×9 grid using Needleman–Wunsch alignment, with the human–human inter-observer ceiling reported as 0.53.The method is retained for comparability despite grid quantization discarding foveal scale.
- Statistical methodology: Human comparisons use Cliff’s δ and crossed-random-intercept linear mixed-effects models, while PCA jointly analyzes five gaze statistics after dropping collinear coverage.The sign convention is agent−human, |δ|>0.33 is non-trivial, and a 95% confidence interval excluding zero is significant.
S4 Decision axis
All three models detect present targets near ceiling and show human-or-better sensitivity in the present/absent decision. Their comparable false-positive bias indicates that later search divergences are not attributable to detection failure.
- Decision axis: 0.99/0.99/0.97 target-present existence accuracy places all three models near ceiling.The passage reports these values for the three models in that order.
- Decision axis: 0.19/0.19/0.20 false-positive bias on absent scenes is comparable across models.This comparability supports matched existence-passed sets across models.
- Decision axis: 3.84/4.23/3.14 model d′ exceeds the human 2.91 reference in the agentic present/absent decision.The passage characterizes the decision axis as human-or-better for every model.
- Decision axis: Later search divergences cannot be attributed to detection failure because models detect targets near ceiling and their existence-passed sets are comparable.This conclusion follows from the reported target-present accuracy and absent-scene false-positive bias.
S5 Finding axis · S6 Gaze axis
Under human-matched foveation, models acquire targets at least as successfully as humans but follow a distinct gaze process. Their scanpaths are spatially concentrated, larger-amplitude, and more self-consistent than human scanpaths.
- S5 Finding axis: S5 Finding axis: Under human-matched foveation, models fixate present targets on the first saccade more often than humans: TFP@1 0.97/0.97/0.80 versus 0.49.The values are ordered Qwen / GLM / Gemma against the human reference.
- S5 Finding axis: S5 Finding axis: Eventual target acquisition is comparable to humans, with TFP-end 0.98/0.98/0.93 versus 0.93.The models’ advantage is therefore in first-saccade efficiency rather than eventual accuracy.
- S5 Finding axis: S5 Finding axis: Models use no more fixations than humans, with median NumFix 2/2/3 versus 3.The values are ordered Qwen / GLM / Gemma against the human reference.
- S6 Gaze axis: S6 Gaze axis: All models show lower gaze entropy than humans, with Cliff’s δ −.67/−.65/−.27, indicating spatially concentrated sampling.Gemma-4-E4B’s effect is roughly half those of the two reasoning-tuned models.
- S6 Gaze axis: S6 Gaze axis: All models show larger saccade amplitudes than humans, with Cliff’s δ +.50/+.61/+.23, indicating direct jumps to the target.The magnitude varies across models, with Gemma-4-E4B’s effect roughly half those of the two reasoning-tuned models.
- S6 Gaze axis: S6 Gaze axis: Cross-seed agent↔agent ScanMatch exceeds the 0.53 inter-observer ceiling for every model, reaching 0.84/0.91/0.71.The values are ordered Qwen / GLM / Gemma.
- S6 Gaze axis: S6 Gaze axis: Models reach targets on the first saccade, whereas human target-fixation probability accrues over several fixations.This pattern appears under both sharp and human-matched GP conditions in the cumulative target-fixation curves.
- S6 Gaze axis: S6 Gaze axis: Agent↔human similarity remains at or below the 0.53 ceiling, so no model is more similar to humans than two humans are to each other.Gemma-4-E4B shows the lowest self-consistency among the models.
S7 Multivariate structure of the gaze signature · S8 Mechanism of the divergence
The gaze signature separates models from humans along interpretable multivariate axes, while matched retinal input reproduces spatial relevance but not temporal search organization. Degradation does not recover human-like search: making first-saccade targeting human-like simultaneously reduces eventual success.
- S7 Multivariate structure of the gaze signature: PCA reduces five gaze statistics to two axes—exploration extent and saccade amplitude versus refixation—and every model occupies a region disjoint from humans except under target-losing degradation.The human reference is approached only when degradation is severe enough that the target is no longer found.
- S8 Mechanism of the divergence: The human-matched condition is behaviourally indistinguishable from no foveation for every model and statistic, including first-saccade targeting 0.97 under GP versus 0.97 under sharp.This supports a single-pass architecture: a parallel vision encoder resolves a legible frame and saccades directly to the target.
- S8 Mechanism of the divergence: Human medians in the intrinsic signature are gaze entropy 1.58 bits, saccade amplitude 8.57◦, refixation 0.00, scanpath length 19.4◦, and center bias 10.0◦.Table S3 compares each model’s distribution with humans using Cliff’s δ, with |δ| > 0.33 marked in bold.
- S8 Mechanism of the divergence: The models’ spatial prior is approximately human: center bias is statistically indistinguishable from humans, and fixation-density correlations are CC 0.58/0.63/0.50.Their divergence concerns the temporal organization of fixation order, amplitude, and determinism.
- S8 Mechanism of the divergence: At human-rate first-saccade targeting of 0.49, degradation requires k=32 for Qwen3.5-35B-A3B and k=16 for Gemma-4-E4B, while TFP-end falls to 0.71 and 0.75.No degradation regime is simultaneously human-like in first-saccade targeting and eventual success.
- S8 Mechanism of the divergence: At human-matched input, reasoning-tuned models concentrate gaze with low entropy and large saccades; Gemma-4-E4B is closest to humans on entropy, amplitude, and scanpath length.Qwen3.5-35B-A3B is closest on refixation.
- S8 Mechanism of the divergence: As peripheral evidence vanishes, refixation first rises through revisiting, while severe degradation increases target-absent declared-absent responses as models default to “absent.”The absolute refixation level is partly shaped by the prompt’s glimpse-memory instruction; the trend across degradation is the failure signature.
S9 Robustness and cross-model variation · S10 Data completeness
Robustness checks preserve the models’ first-saccade and gaze-distribution effects across target difficulty, tolerance, and scene/rater controls, while cross-model differences remain descriptive. The complete dataset and recovery procedures support analysis of all recorded scanpaths without exclusions.
- S9 Robustness and cross-model variation: First-saccade advantage holds in every eccentricity × size stratum, including the hardest targets.The result is not attributable to easy targets.
- S9 Robustness and cross-model variation: Mixed-effects controls for scene and rater retain nonzero gaze-entropy, saccade-amplitude and first-saccade effects for all three models.All corresponding confidence intervals exclude zero in Table S9.
- S9 Robustness and cross-model variation: Gemma-4-E4B has the smallest effects on five of seven gaze statistics, making it the most human-like model overall.The statistics are entropy, saccade amplitude, scanpath length, center bias, and coverage.
- S9 Robustness and cross-model variation: Models’ saccade-direction and turning-angle structure departs from the human reference, consistent with entropy and amplitude effects.The figure documents directional and meander differences under the human-matched condition.
- S9 Robustness and cross-model variation: Cross-model ordering is descriptive because architecture, training recipe, and mixture-of-experts sparsity covary across only three observations.High cross-seed determinism also remains above the inter-observer ceiling through temperature 1.0 on Qwen3.5-35B-A3B.
- S10 Data completeness: All three models completed 285 scenes, 5 seeds and nine conditions, with no trials excluded from final analysis.GLM-4.6V-Flash detections were re-collected after server interruption, and Gemma-4-E4B terminal decisions were canonically normalised before parsing.
- S10 Data completeness: Logged turn sequences reproduced recorded gaze paths up to decisions, while residual truncated episodes were re-collected under the identical protocol.All quantities use recorded scanpaths and the single definition set in Sec. S3.
S11 Implications and scope
The models align with human outcomes and approximate spatial allocation but diverge in temporal gaze dynamics, limiting their use as human-vision surrogates for process-level studies. Under the study’s fixed foveation constraint, they also provide a useful parallel one-pass null model for isolating effects of serial search.
- Evaluation metrics: Answer-alignment scores and single-shot saliency overlap miss the temporal axis on which model gaze diverges from humans.These metrics evaluate outcomes or time-collapsed spatial maps rather than the gaze process.
- Scope: Human-matched foveation changes model behaviour negligibly, so matched retinal input does not alter the study’s central process-level conclusions.The figure compares statistics under sharp and human-matched conditions, with points on the identity line.
- Suitability as human-vision surrogates: Zero-shot multimodal models suit outcome and spatial-allocation studies but not process or temporal-dynamics studies.The latter include scanpath prediction, fixation counts and amplitudes, stopping behaviour, and evidence-accumulation time courses.
- A null model for serial search: A parallel one-pass system achieving human-or-better outcomes serves as a null model for isolating behaviours attributable to seriality.The comparison separates serial-search effects from detection ability and spatial priors, which the null model already reproduces.
- Scope: The study examined general-purpose models under fixed foveation but excluded search-trained agents, tool-use policies, pointing-native or frontier closed-source models, and explicit semantic-guidance probes.It also did not run per-model temperature sweeps beyond the anchor model.
S12 Prompt
The prompt casts the model as a single eye searching a photograph through human-like foveated vision. At each turn, it uses the current gaze view and search history to move, report PRESENT when the target is visible, or report ABSENT when confident it is missing.
- Foveated search: The model sees sharply only at its current gaze point, with peripheral blur increasing with distance, and must move its gaze to inspect another region.This implements a foveated, sequential search interface modeled on human peripheral vision.
- Search history: Each turn provides the image from the current gaze point plus earlier gaze locations and observations, enabling history-based search and avoidance of repeated checks.The history is explicitly used to choose the next location and avoid re-checking the same spots.
- Decision task: The task requires a PRESENT/ABSENT decision about whether the target object appears in the image.A PRESENT decision is made when the target is clearly visible at the gaze point; ABSENT requires confidence that it is nowhere in the image.
- Turn-level actions: On every turn, the model must either move its gaze to a new point to continue searching or issue a final target-presence directive.The prompt constrains each response to one action and requires the reply to end with exactly one directive form.