Source-linked AI summary
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
Rohit Sinha, Aditya Kanade, Sai Srinivas Kancheti, Vineeth N Balasubramanian, Tanuja Ganu
TL;DR
MLLMs show strong progress on perceptual vision tasks, but their visuospatial cognitive reasoning remains insufficiently characterized. Mind’s Eye introduces an eight-task ART benchmark that separates abstraction, relation, and transformation while diagnosing reasoning errors. Humans achieve 80% accuracy, whereas top MLLMs remain below 50%, with analyses showing that models often localize relevant visual information without reliably reasoning over it.
Problem
Existing evaluations do not isolate visuospatial transformation and often leave unclear whether models reason from images or exploit linguistic shortcuts.
Method
Mind’s Eye evaluates MLLMs with eight cognitively grounded tasks organized under Abstraction, Relation, and Transformation, using diagnostic distractors and controlled visual inputs.
Results
Humans average 80% accuracy while top MLLMs remain below 50%, and attention analyses show that models can localize relevant regions but fail to reason over them reliably.
Takeaways & Limitations
The benchmark exposes limited visuospatial reasoning and task-dependent prompting effects, with structured guidance helping Abstraction but impairing Transformation.
Takeaways & Limitations
The benchmark uses multiple-choice scoring and 2D renderings with controlled 3D implications; fully 3D inputs and interactions remain future work.
Abstract
from arXiv · showhide
Multimodal large language models (MLLMs) have achieved impressive progress on vision language benchmarks, yet their capacity for visual cognitive and visuospatial reasoning remains less understood. We introduce "Mind's Eye", a multiple-choice benchmark of eight visuo-cognitive tasks inspired by classic human intelligence tests and organized under a novel "A-R-T" taxonomy: Abstraction, Relation, and Transformation. The tasks probe core processes of fluid intelligence such as pattern induction, analogical relation mapping, and mental transformation. We evaluate a diverse suite of closed-source and open-source MLLMs and compare their performance with human participants. Humans achieve 80% accuracy, while top performing MLLMs remain below 50%. Error analysis reveals failures in: (i) visual attention allocation, (ii) internal perceptual manipulation, and (iii) weak abstraction of underlying visual concepts. Our findings suggest that current MLLMs exhibit limited visuospatial reasoning capabilities, when compared with human participants, highlighting the need for more cognitively grounded evaluation frameworks.
1 Introduction
Mind’s Eye addresses gaps in evaluating visuospatial transformation and separates visual reasoning from linguistic shortcuts through cognitively grounded tasks. Its evaluation finds substantial human–MLLM differences and diagnoses attention, manipulation, and abstraction failures.
- Existing evaluations often conflate visual evidence with linguistic priors and do not isolate visuospatial transformation such as rotation, folding, or recomposition.
- Mind’s Eye introduces eight tasks organized by Abstraction, Relation, and Transformation to probe pattern induction, analogical mapping, and mental shape manipulation.
- The benchmark isolates visuospatial reasoning from world knowledge and linguistic priors while using diagnostic distractors to identify specific error types.
- Humans average 80% accuracy across tasks, whereas the best evaluated models remain below 50%.
- Diagnostic analyses reveal attention misalignment, difficulty-invariant failures, and reasoning-trace errors.
2 Related Work
Prior benchmarks cover broad perception, compositional reasoning, or isolated cognitive skills but provide limited control over internal visuospatial simulation. Mind’s Eye is presented as satisfying a broader set of diagnostic and psychometric design criteria.
- General-purpose benchmarks measure visual question answering, OCR, and mathematical reasoning breadth but lack parametric control for studying visuospatial understanding.
- Cognitive benchmarks emphasize rule discovery, analogy, spatial completion, or perception grounding rather than stepwise geometric simulation.
- Mind’s Eye bridges the under-specification of internal simulation by probing whether models can perform internal transformations.
- Mind’s Eye is described as the first benchmark in this space to satisfy all six criteria: psychometric taxonomy, established assessments, diagnostic distractors, no knowledge reliance, parametric control, and scalability.
3 Mind’s Eye: The Benchmark
Mind’s Eye operationalizes fluid visual reasoning through the ART taxonomy and an eight-task, programmatically generated benchmark. Its design isolates cognitive operations, controls difficulty, and uses diagnostic answer choices for fine-grained analysis.
- Conceptual foundations: ART decomposes fluid visual reasoning into Abstraction, Relation, and Transformation: pattern induction, relational correspondence, and mental spatial manipulation.
- Conceptual foundations: Abstraction induces latent structure, Relation maps correspondences across configurations, and Transformation mentally simulates rotation, folding, or composition.
- Design principles: Tasks require visual-structure reasoning rather than world knowledge, use distractors tied to specific errors, and follow factorial stimulus-generation designs.
- Task suite: The eight tasks are VRA and HPE for Abstraction; DSC, VCS, and SS for Relation; and MT, PF, and MC for Transformation.
- Benchmark scale: 800 items are included, with 100 items per task balanced across the three ART dimensions.
- Human evaluation: Human evaluation recruited 30 participants, who completed calibrated multiple-choice samples from every task.
4 Experiments and Results
The experiments evaluate diverse closed- and open-source MLLMs on standardized Mind’s Eye inputs and find broad weaknesses in visuospatial reasoning, especially when perception must support transformation or temporal integration.
- Experimental setup: The study evaluates recent proprietary and open-source MLLMs using identical visual inputs and standardized textual prompts.
- Prompting strategies: Four prompting strategies are tested: Chain-of-Thought, Meta-Task Framing, Step-by-Step Instruction, and Hint-based prompting.
- Main results: MLLMs generally underperform humans and often fail to integrate recognized 3D arrangements or object correspondences into consistent reasoning.
- Main results: Models particularly struggle with temporal sequences and tracking visual elements across transformations.
- Main results: Mental Composition performance depends on visual resemblance to a cube, declining when correct inference requires folding a shape into a nontrivial 3D structure.
5 Analysis and Discussion
Analysis shows that models can localize relevant visual regions, but reliable reasoning over those regions remains limited. Prompting and difficulty produce dimension-dependent patterns: scaffolding helps abstraction yet harms transformation, while model performance stays comparatively flat across difficulty.
- Attention Alignment: OAScorrect correlates with accuracy, but even high-attention models remain below human performance, showing that localization is necessary but insufficient.Across 200 items, the point-biserial correlation is rpb = 0.34, p < 0.001; highest-attention accuracy remains below human performance (>80%).
- Attention Alignment: Relative attention strengthens the association between correct-option focus and accuracy without resolving the broader reasoning deficit.The correlation increases from rpb = 0.34 to rpb = 0.41, while correct-prediction attention preference grows from Cohen’s d = 1.15 to d = 1.41.
- Interpretation: The central bottleneck is reasoning over correctly localized information, not simply directing attention to relevant regions.Attention to relevant regions correlates with accuracy, but localization alone does not produce reliable reasoning.
- Prompt Stability: Prompting effects depend on ART dimension: abstraction benefits from structured guidance, while transformation consistently degrades under alternative prompts.Meta-task and Step-by-step prompting improve abstraction by approximately +1.3 pts, whereas Hint prompting lowers transformation performance by approximately −0.9 pts.
- Difficulty Effects: Models fail across difficulty levels, suggesting missing foundational visuocognitive operations rather than difficulty-specific weakness.Human accuracy declines with difficulty, whereas model performance remains flat; this pattern appears across all three ART dimensions.
6 Conclusions
Mind’s Eye evaluates MLLMs’ visual intelligence through an ART-organized visuocognitive benchmark. Its results show a persistent human-model gap and task-dependent prompting effects, while diagnostic analyses identify limited grounding and internal simulation as central failure patterns.
- Benchmark: Mind’s Eye is a visuocognitive benchmark organized around Abstraction, Relation, and Transformation.The three axes are inspired by Carroll’s three-stratum theory.
- Results: Non-expert humans achieve 80% mean accuracy, while top MLLMs remain below 50%.
- Results: Prompting strategies produce task-dependent but modest improvements without changing error profiles.
- Failure Modes: MLLMs rely heavily on perceptual cues, with limited coupling between textual reasoning and visual evidence.
- Implications: The benchmark’s ART-aligned parametric design exposes specific failure modes and suggests directions involving grounded attention, spatial working memory, and transformation-aware representations.
Limitations
Mind’s Eye’s evaluation is bounded by its multiple-choice format, controlled 2D renderings, and specific human baseline population. The authors also identify broader validity, interpretation, and misuse limits for synthetic cognitive-style benchmarking.
- Evaluation scope: The multiple-choice format improves reliability and objectivity of comparison but leaves open-ended generation as an untested source of insights.The benchmark focuses on multiple-choice scoring, while open-ended generation may reveal different properties.
- Evaluation scope: The tasks use 2D renderings with controlled 3D implications, so fully 3D inputs and interactions remain outside the current scope.The authors identify fully 3D inputs and interactions as future work.
- Human baseline: The human baseline comprises non-expert adults in a single-language setting, and cross-lingual or expert cohorts may change absolute performance levels.The authors hypothesize that relative gaps are likely to remain, based on their observations.
- Validity: Synthetic controlled items may not transfer to natural images, while the tasks remain proxies for broader visuocognition and cannot eliminate all heuristics.The authors release generators to enable domain shifts and use bootstrap confidence intervals and mixed-effects models to quantify uncertainty.
- Interpretation: Behavioral success on narrowly specified tasks should not be treated as mechanistic equivalence to human cognition.The authors therefore distinguish construct-level claims about what is measured from implementation claims about how models compute.
- Broader impact: Leaderboards and cognitive-style tests could encourage narrow optimization or inappropriate gatekeeping in safety-critical, hiring, or educational settings.The authors prohibit using the benchmark for human evaluation or selection and caution against ranking individuals or groups.
A Additional Results
Additional results show that model scale and prompting produce task-dependent effects rather than reliably resolving visuospatial reasoning failures. Attention and reasoning analyses indicate persistent weaknesses in visual grounding, internal simulation, and concept abstraction.
- Model scale: Medium-scale models can outperform larger counterparts, while scaling helps some transformation tasks but leaves abstraction and compositional reasoning difficult.Qwen-2.5-VL 32B outperforms smaller variants on several transformation-heavy tasks, yet smaller models remain competitive on abstraction-oriented tasks and Mental Composition accuracy stays low.
- Prompting: Prompting effects vary by task: structured guidance benefits some abstraction and relation tasks, whereas transformation tasks consistently deteriorate relative to CoT.Mental Composition drops by approximately 0.8–1.4 points across prompting strategies, while Hierarchical Pattern Equivalence gains approximately 1.8 points under meta-task prompting.
- Visual grounding: Less than 20% of normalized attention mass reaches explicitly referenced visual objects, and mean Region Aligned Attention is 0.18.These results indicate that fluent reasoning traces are often not grounded in the corresponding visual regions.
- Visual grounding: Attention alignment predicts performance modestly: rpb = 0.34, but even the highest-attention quartile reaches only 35.7% accuracy.Correct answers receive significantly more attention than distractors (Cohen’s d = 1.15), whereas selected incorrect and correct options do not differ significantly when models err (p = 0.40).
- Failure analysis: Reasoning traces reveal persistent misperception, prompt-insensitive reasoning, and failure to perform final multi-step disambiguation against visually similar distractors.Models may encode a shape incorrectly, retrieve domain-based interpretations, and choose a plausible distractor without completing the required reasoning sequence.
B.9 Thought Anchors CoT Annotation Analysis
The analysis categorizes chain-of-thought sentences into reasoning stages and compares model reasoning with ground-truth transformations and answers. It finds near-chance mental-transformation performance, axis-specific biases, and frequent failures to bind correct reasoning to the final visual option.
- Thought Anchor Annotation: Six reasoning stages distinguish comprehension, planning, option analysis, answer emission, self-checking, and unknown statements.The categorization is used to identify where failures occur in the reasoning process.
- Thought Anchor Annotation: Ground-truth transformation and answer annotations enable separate evaluation of predicted options and reasoning alignment.
- Mental Transformation Results: 32.2% overall accuracy places Qwen-7B only marginally above four-option chance and below human performance above 80%.
- Mental Transformation Results: 28.2% Y-axis, 32.4% X-axis, and 35.7% Z-axis accuracy reveal strong axis-specific variation inconsistent with robust human visuospatial reasoning.The analysis attributes this pattern to superficial 2D heuristics rather than flexible 3D representations.
- Mental Transformation Results: 61.1% of cases contained reasoning that correctly described the transformation but produced an incorrect or misaligned final option.This pattern is characterized as systematic mis-binding between linguistic reasoning and visual candidates.
- Prompt Analysis: Optimized prompts produced gains below 0.10 on evaluated tasks, but persistent error patterns indicate that prompt wording does not explain the core limitations.Reported gains included +0.08, +0.08, +0.07, and +0.06 on four tasks.
- Benchmark Context: Mind’s Eye procedurally generates eight visuospatial task families with explicit answer, violation-type, and difficulty metadata.Examples include visual conceptual slippage, visual relation abstraction, mental composition, and paper folding.
D Benchmark Design
Mind’s Eye is designed as a controlled, psychometrically structured benchmark covering Abstraction, Relation, and Transformation. Its generators vary task and nuisance factors, use diagnostic distractors, calibrate difficulty with human data, and standardize model evaluation.
- Stimulus Generation: Factorial generators independently vary structural and nuisance factors to reduce annotation artifacts and superficial shortcuts.Examples include fold-sequence length, transformation-chain length, rendering styles, color schemes, and layout jitter.
- Difficulty Calibration: Human-consensus calibration assigns 32% of items to Easy, 45% to Medium, and 23% to Hard difficulty levels.Each item was independently evaluated by five participants from a cohort of 30 adults.
- Diagnostic Design: Distractors are keyed to confounds such as mirrored transformations, fold-parity errors, swapped correspondences, and superficial feature matches.This design supports error analysis beyond binary correctness.
- Benchmark Construction: The main suite contains 800 items, with 100 items per subtask unless otherwise stated; option order and answer keys are uniformly randomized.
- Construct Coverage: The Q-matrix maps benchmark tasks to latent Abstraction, Relation, and Transformation skills and supports multi-trait psychometric modeling.It serves as an explicit construct-coverage blueprint for the benchmark.
- Model Evaluation: Evaluation uses identical visual inputs and standardized prompts, followed by expert-LLM answer extraction and mapping to task-specific labels.Gemma-3 serves as the judging model for parsing free-form answers.
- Prompting Strategies: Four prompting strategies are evaluated to test model sensitivity to instructions and elicit cognitive reasoning rather than shallow pattern matching.The strategies include Hint-Based, Elimination-Based, Meta-Task, and Step-by-Step prompting.
F Human Evaluation Protocol
The human evaluation establishes a standardized reference for Mind’s Eye performance and item difficulty. Thirty adults complete randomized eight-task batteries under controlled conditions, producing data for model comparison and psychometric analysis.
- Participants: 30 adults aged 20–40 participated, including 17 male and 13 female participants without prior expertise in the benchmark tasks.
- Procedure: Each participant completed all eight task families in randomized order, with randomized items and standardized instructions.Responses were collected digitally through an image-based multiple-choice interface.
- Procedure: Approximately 60 minutes of testing yielded a human-response dataset for model reference, item-difficulty calibration, and discrimination analysis.
- Prompting Evaluation: Performance was evaluated under Hint-Based, Elimination-Based, Meta-Task, and Step-by-Step prompting strategies.Detailed results are reported in separate tables for each strategy.
I Qualitative CoT Analysis
Qualitative analysis of GPT-4o reasoning traces finds coherent-looking explanations that remain largely surface-level and perceptually driven. Across representative tasks, the traces rely on heuristic visual cues rather than systematic cognitive reasoning.
- Qualitative Findings: GPT-4o often produces syntactically coherent explanations and sometimes correct answers, but its reasoning remains largely surface-level and perceptually driven.
- Qualitative Findings: The traces rely on heuristic visual cues rather than systematic cognitive reasoning.
- Theoretical Alignment: Figure 25 aligns Fluid Intelligence constructs from the CHC framework with the Abstraction, Relation, and Transformation taxonomy.
K Difficulty Analysis
Humans show graded sensitivity to difficulty, whereas MLLMs perform substantially worse and remain relatively flat across difficulty levels. This divergence is strongest in Transformation and Abstraction tasks, with prompting effects varying by cognitive dimension.
- Human Difficulty Sensitivity: Human accuracy declines from 0.85-0.95 on Easy items to 0.55-0.65 on Medium and 0.10-0.25 on Hard items.These ranges correspond to the benchmark’s annotator-based difficulty calibration.
- Model–Human Performance: 80% human accuracy contrasts with model accuracy typically ranging from 0.2-0.5 across tasks, with the gap consistent across all eight subtasks.The deficit is particularly pronounced in Transformation and Abstraction dimensions.
- Difficulty Curves: Model performance varies by only 0.02-0.08 accuracy points across Easy, Medium, and Hard conditions, unlike the progressive human decline.The flat curve is presented as evidence that models do not perform the core cognitive operations underlying genuine spatial reasoning.
- Model Comparisons: Closed-source models outperform open-source models across tasks and difficulty levels, but both remain below humans and show similarly flat difficulty curves.The advantage is largest in Transformation, where closed-source models reach 0.30-0.45 accuracy versus 0.25-0.35 for open-source models.
- Interpretation: The human–model divergence suggests that current MLLMs use fundamentally different mechanisms and may require architectural innovations for perceptual transformation and cognitive simulation.The conclusion contrasts genuine human visuospatial reasoning with shallow model heuristics that fail uniformly across difficulty levels.
- Prompting Effects: Prompting effects depend on cognitive dimension: structured scaffolding benefits Abstraction but consistently impairs Transformation performance.The reported pattern suggests prompting supports rule derivation more readily than procedural visuospatial operations.
- Error Analysis: Attention analysis shows that models can localize relevant answer regions but do not reliably reason over the information they identify.A Mental Transformation attention map illustrates misplaced model attention relative to expected regions.
- Reasoning Trace Analysis: Reasoning traces attribute errors to superficial heuristics, including color matching, failure to recognize cube nets, and failure to track holes during unfolding.These analyses cover Mental Transformation, Mental Composition, and Paper Folding tasks.