Source-linked AI summary

Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance

Jason Qiu, Zachary Meurer, Xavier Thomas, Deepti Ghadiyaram

arXiv:2604.01848v4cs.CV

TL;DR

The paper investigates whether VLMs use robust geometric reasoning or rely on semantic familiarity when recognizing transformed images. It evaluates rotation, scaling, and identity matching across domains with different semantic richness, finding sharp degradation on sparse and unfamiliar content across models and prompting strategies. The results indicate a gap between semantic recognition and invariant geometric understanding, while the benchmark remains limited to 2D transformations and does not measure downstream application impact.

  • Problem

    The paper addresses whether VLMs possess robust geometric invariance and equivariance rather than relying on semantic familiarity when identifying transformed content.

  • Method

    The study evaluates six VLMs on rotation, scaling, and identity matching using image pairs drawn from visual domains spanning sparse sketches and scripts to photographs and abstract art.

  • Results

    Performance degrades sharply on sketches, symbolic characters, and unfamiliar scripts, with failures persisting across architectures, model scale, and prompting strategies.

  • Takeaways & Limitations

    Current VLMs rely on semantic anchors and lack a fundamental invariant grasp of geometry across the tested transformations and visual domains.

  • Takeaways & Limitations

    The benchmark evaluates only 2D transformations and does not directly measure effects on downstream tasks such as visual grounding or embodied robotics.

Abstract

from arXiv · show

This work investigates the fundamental fragility of state-of-the-art Vision-Language Models (VLMs) under basic geometric transformations. While modern VLMs excel at semantic tasks such as recognizing objects in canonical orientations and describing complex scenes, they exhibit systematic failures at a more fundamental level: lack of robust spatial invariance and equivariance required to reliably determine object identity under simple rotations, scaling, and identity transformations. We demonstrate this limitation through a systematic evaluation across diverse visual domains, including symbolic sketches, natural photographs, and abstract art. Performance drops sharply as semantic content becomes sparse, and this behavior is observed across architectures, model capacities, and prompting strategies. Overall, our results reveal a systematic gap between semantic understanding and spatial reasoning in current VLMs, highlighting the need for stronger geometric grounding in future multimodal systems.

1 Introduction

The paper asks whether VLMs can switch from semantic recognition to geometric reasoning when identifying transformed content. Across visual domains and transformations, it finds sharp, broadly persistent failures as semantic cues become sparse.

  • Evaluation and research question: The study tests rotation, scaling, and identity matching across familiar and unfamiliar scripts, sketches, photographs, and other visual domains.The evaluation uses semantic granularity as a stress test for geometric reasoning.
  • Core findings: VLM performance collapses on semantically sparse content under basic geometric transformations, despite strong performance on richer familiar inputs.This pattern appears across symbolic sketches, handwritten scripts, cartoons, and photographs.
  • Core findings: Models perform better when asked whether two images contain the same object than when asked whether one is a rotated variant of the other.The comparison indicates reliance on object labels without a corresponding grasp of underlying geometry.
  • Core findings: 92.67% to 76.49%: Gemini-2.5-Pro accuracy drops from photos to symbolic sketches for rotation, while scale accuracy falls from 99.81% to 82.56%.These values illustrate the degradation associated with sparse semantic content.
  • Core findings: Rotation is consistently the most challenging transformation, even when identity and scale performance is stronger.The finding is reported across the evaluated models and domains.
  • Robustness of the finding: Failures persist across architectures, model capacities, and prompting strategies, indicating that model scale or prompt design does not explain the limitation.The paper characterizes the limitation as fundamental to current VLM behavior.

2 Related Work

Prior work documents VLM weaknesses in visual reasoning and transformation-related tasks, but commonly emphasizes natural images or isolated failure modes. This paper positions its contribution as a systematic comparison across visual domains and semantic richness.

  • Prior visual-reasoning research: Earlier studies report VLM failures in object counting, recognition, visual analogies, orientation, depth estimation, and spatial correspondence.These works motivate broader investigation of visual reasoning limitations.
  • Transformation invariance: Unlike convolutional networks, ViT backbones lack inherent transformation equivariance, and emergent invariance does not necessarily produce robust transformation reasoning.The related work distinguishes representational invariance from reliable task-level reasoning.
  • Cross-domain evaluation: Prior visual-reasoning benchmarks primarily focus on natural images rather than systematically comparing photos, cartoons, sketches, and other domains.The paper uses cross-domain evaluation to address this coverage gap.

3 Studying VLM’s invariance equivariance dilemma

The study tests whether VLMs recognize identical content across rotation, scaling, and identity transformations, rather than relying on semantic familiarity. Across scripts and visual domains, models show strong semantic recognition but fragile geometric reasoning, especially when semantic cues are sparse.

  • Evaluation setup: The evaluation defines transformation equivariance as recognizing the same underlying content while identifying whether rotation, scaling, or identity was applied.Each instance pairs an image I with either a transformed version of I or a transformed different image.
  • Model capacity: Increasing model capacity yields modest accuracy gains but does not solve geometric brittleness, with sketch TPR remaining 14.50% for Qwen2.5-VL-32B and 5.67% for Qwen3-VL-30B.Larger models improve more on photos than sketches, so scale alone is insufficient for robust geometric reasoning.
  • Prompt and task formulation: Near-perfect object recognition and identification contrast with substantially lower rotation recognition, exposing a gap between semantic recognition and geometric reasoning.This comparison uses distinct task formulations on the same transformed image pairs.
  • Rotation invariance: Models fail across datasets, capacities, and prompts when asked to recognize rotated variants of the same image.Rotation recognition shows consistently low TPR and a strong bias toward predicting “No,” producing near-random accuracy despite high TNR.
  • Identity and scale invariance: Identity and scale performance is strong for familiar scripts but deteriorates for unfamiliar Omniglot scripts, including Qwen2.5-VL-7B’s 62.78% TPR on identical Omniglot pairs.For scale, familiar Times New Roman exceeds 98% across models, while Omniglot ranges from 74.21% to 82.56% accuracy.
  • Reasoning behavior: Reasoning traces show semantic-first processing for familiar characters and more low-level geometric analysis for unfamiliar characters.For a familiar “C,” Gemini-2.5-Pro identifies the letter before confirming rotation; for unfamiliar Gujarati, it analyzes structural geometry.
  • Prompting interventions: In-context learning can raise TPR but often lowers TNR, while smaller models may gain little or remain at zero TPR.For Qwen2.5-VL-32B on Malayalam, two examples raise TPR from 6.38% to 51.06% while TNR falls from 97.87% to 65.96%.

4 Conclusion

The study finds that VLM performance degrades sharply on sparse and unfamiliar visual inputs across transformations, with failures persisting across architectures, model scale, and prompting strategies. It concludes that current VLMs rely on semantic anchors rather than an invariant grasp of geometry, while noting that the benchmark covers only 2D transformations and does not measure downstream impacts.

  • Across three transformations, four visual domains, and six models, performance is robust on semantically rich inputs but degrades sharply on sketches, symbolic characters, and unfamiliar scripts.
  • These failures persist regardless of model architecture, scale, or prompting strategy and are only partially mitigated by in-context learning or structured visual prompts.
  • The findings indicate that current VLMs lack a fundamental invariant grasp of geometry and instead rely on semantic anchors for spatial tasks.
  • The benchmark focuses exclusively on 2D transformations, does not evaluate 3D rotations, and does not directly measure effects on downstream applications such as visual grounding or embodied robotics.

A.1 Dataset overview

The evaluation combines sparse character datasets, semantically rich object images, and controlled geometric transformations. Feature representations are extracted from multiple vision encoders for cosine-similarity analysis under rotation.

  • Datasets: Omniglot contains handwritten binary characters from 50 diverse scripts, while Times New Roman and Handwritten English provide familiar English character datasets.
  • Datasets: PACS contains seven object categories across Photograph, Art, Cartoon, and Sketch domains, supplying semantically rich images with varied texture, color, and abstraction.
  • Transformations: Rotations are centered, and non-orthogonal rotations introduce empty regions filled with white padding; scaling resizes characters by s ∈ {0.1, 0.3, 0.5, 0.9} while preserving image dimensions.
  • Feature extraction: Cosine similarity uses global representations from CLS tokens, attention pooling, averaged vision tokens, or spatially averaged diffusion features, depending on the encoder.
  • Feature extraction: Across encoders, feature similarity decreases as rotation angle increases, with DINOv2 dropping steepest while SigLIP and Qwen2.5-VL-7B maintain relatively higher similarity.

D Suspect 4: Interaction with the Language Decoder

The language-decoder analysis compares feature-level rotational similarity with task-level rotation recognition. Despite similar representations for transformed inputs, downstream accuracy remains near chance and is shaped by prediction bias.

  • Accuracy remains near chance for most Idefics2 scripts despite consistently high TPR and substantially lower TNR, indicating a strong “YES” prediction bias.
  • Vision encoders maintain high cosine similarity between characters and rotated counterparts, yet MLLMs consistently fail to recognize the same rotation.
  • Idefics2 uses the same frozen SigLIP-SO400M-384 encoder evaluated in the feature analysis, enabling a controlled comparison between encoder representations and task performance.
  • The mismatch shows that high representation similarity alone is insufficient for reliable downstream reasoning about transformations.

E PACS Performance for the Scale Invariance Task

Scale-invariance performance is high on natural-image domains and familiar scripts but drops on Omniglot. This pattern appears for both Qwen2.5-VL and Gemini-2.5-Pro, with Gemini approaching perfect performance on familiar inputs.

  • Qwen2.5-VL performs highly on Art Painting, Cartoon, Photo, Sketch, Times New Roman, and Handwritten English, but drops on Omniglot.
  • Gemini-2.5-Pro achieves near-perfect scale-invariance performance on natural-image domains and familiar scripts but drops on Omniglot.

F Perimetric Complexity Analysis

Perimetric complexity does not explain the performance gap across scripts: structural intricacy has only a weak relationship with scale-invariance accuracy.

  • r = −0.18 correlation between perimetric complexity and Qwen2.5-VL-7B scale-invariance accuracy is weak.Perimetric complexity is defined as P^2/A, where P is character perimeter and A is occupied area.

G Additional Scale Invariance Analysis

Additional analyses show that geometric transformation failures persist across datasets, prompting strategies, and model components, while familiar scripts remain substantially easier than unfamiliar ones.

  • Near-perfect scale accuracy holds for Times New Roman and Handwritten English across scales, but Omniglot performance remains substantially lower.
  • Rotation information is linearly decodable from encoder features, yet downstream VLMs achieve only 50.72% and 53.55% on Omniglot rotation.Encoders exceed the 11.1% random-chance angle-classification baseline, while Qwen2.5-VL-7B and Idefics2 downstream performance collapses.
  • Forced-choice prompting raises Qwen2.5-VL-32B TPR by +65–75% but reduces Omniglot TNR from 94.93% to 39.11%.The trade-off indicates response-bias mitigation rather than stable geometric understanding.
  • Gemini-2.5-Pro remains stable across prompt variants, while unfamiliar scripts and sparse visual domains produce lower transformation-recognition performance.
  • Humans achieve 100% accuracy across scripts, with only marginally longer response times for unfamiliar scripts.Reported times include 2.91s for Malayalam and 2.78s for Armenian versus 2.28s for Times New Roman.

I.8 Statistical Robustness and Confidence Intervals

Bootstrap confidence intervals are narrow, supporting the statistical robustness of the reported performance gaps, which recur across additional open-source model families.

  • 95% bootstrap confidence intervals based on 10,000 resamples are narrow across rotation and scale evaluations.
  • Gemini-2.5-Pro rotation accuracy is 89.32% on Times New Roman versus 76.90% on Omniglot, with narrow confidence intervals.
  • Molmo2-8B and InternVL3.5-8B reproduce semantic-richness degradation, with scale-task TPR dropping from 50.96% to 12.52% and from 96.15% to 29.22%.

I.10 The Role of Data Domain and Semantic Grounding

Semantic-domain fine-tuning improves geometric transformation recognition without geometric supervision, supporting a role for semantic grounding in accessing visual representations.

  • Semantic VQA fine-tuning on 3,729 PACS sketch images excludes rotation examples, scale information, and geometric supervision.The vision encoder remains frozen while the language component is fine-tuned on semantic descriptions.
  • Rotation TPR improves from 12.72% to 81.50% (+68.78%) after semantic-only fine-tuning, despite zero rotation or scale examples.The reported improvement occurs without modifying the visual encoder or adding geometric supervision.
  • The authors interpret the result as evidence that frozen vision encoders preserve geometric features while language decoders struggle to access them without semantic grounding.
  • Concurrent work similarly reports improved visual correspondence after fine-tuning VLMs with arbitrary semantic labels for novel shapes.
Loading 2604.01848v4…