Source-linked AI summary

VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?

Minkyu Kim, Sangheon Lee, Dongmin Park

arXiv:2603.07888v1cs.CVcs.AIcs.LG

TL;DR

VLM-SubtleBench addresses the limited evaluation of subtle comparative reasoning between visually similar images. It introduces a diverse benchmark and controlled analyses, finding systematic model–human gaps, particularly for spatial, temporal, and viewpoint reasoning, while prompting strategies provide only limited improvements.

  • Problem

    Existing comparative VLM benchmarks emphasize salient differences and natural images, leaving subtle reasoning across diverse real-world domains insufficiently evaluated.

  • Method

    The paper constructs VLM-SubtleBench with paired image-question-answer examples spanning ten difference types and six domains, then evaluates proprietary and open-source VLMs using prompting and synthetic controls.

  • Results

    VLMs retain large human-performance gaps, with the best model trailing humans by over 30 percentage points on spatial, temporal, and viewpoint reasoning; simple prompting yields limited improvements.

  • Takeaways & Limitations

    VLM-SubtleBench serves as a benchmark and diagnostic tool for developing models capable of subtle comparative reasoning across domains.

Abstract

from arXiv · show

The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for vision-language models (VLMs) have recently emerged, they primarily focus on images with large, salient differences and fail to capture the nuanced reasoning required for real-world applications. In this work, we introduce VLM-SubtleBench, a benchmark designed to evaluate VLMs on subtle comparative reasoning. Our benchmark covers ten difference types - Attribute, State, Emotion, Temporal, Spatial, Existence, Quantity, Quality, Viewpoint, and Action - and curate paired question-image sets reflecting these fine-grained variations. Unlike prior benchmarks restricted to natural image datasets, our benchmark spans diverse domains, including industrial, aerial, and medical imagery. Through extensive evaluation of both proprietary and open-source VLMs, we reveal systematic gaps between model and human performance across difference types and domains, and provide controlled analyses highlighting where VLMs' reasoning sharply deteriorates. Together, our benchmark and findings establish a foundation for advancing VLMs toward human-level comparative reasoning.

1 INTRODUCTION

VLM-SubtleBench addresses the lack of challenging, diverse benchmarks for subtle comparative reasoning across visually similar inputs. It evaluates fine-grained differences and finds substantial gaps between advanced VLMs and humans.

  • Subtle visual differences support tasks spanning micro-expression recognition, industrial anomaly detection, satellite-image assessment, and medical disease-stage distinction.
  • Prior comparative benchmarks mainly use dissimilar natural scenes, making salient differences easy for state-of-the-art VLMs such as GPT-4o.
  • VLM-SubtleBench contains 13K image-pair, question, and answer triplets across 10 difference types and six image domains.
  • GPT-5 and Gemini-2.5-pro still show significant human-performance gaps, especially for spatial, temporal, and viewpoint reasoning.
  • The studies examine model performance, test-time prompting, and synthetic image pairs to identify where comparative reasoning deteriorates.
  • Even the best model trails humans by over 30 percentage points on spatial, temporal, and viewpoint reasoning, while simple prompting strategies yield limited improvements.

2 RELATED WORK

Prior work evaluates multi-image reasoning, domain-specific visual differences, and image-difference captioning, but does not specifically target subtle change identification between similar image pairs.

  • Multi-image benchmarks cover verification, compositional matching, perception, relational reasoning, retrieval, and scene understanding across multiple images.
  • These benchmarks do not specifically evaluate comparative reasoning that identifies what changed between two similar images.
  • Specialized datasets study visual differences in surveillance, synthetic, bird, and medical imagery.
  • Image-difference captioning datasets target object replacement, removal, color, texture, spatial, style, and text changes.

3 VLM-SUBTLEBENCH

VLM-SubtleBench is constructed to test subtle comparative reasoning across diverse domains and a comprehensive taxonomy of difference types. Its paired examples combine curated real data, annotations, synthetic edits, and controlled motion or object changes.

  • Scope: The benchmark spans six domains: natural, game, industry, aerial, synthetic, and medical imagery.
  • Scope: Its taxonomy contains ten difference types, extending prior categorization with Viewpoint and Action.
  • Scope: The categories range from object attributes, states, emotions, and quality to temporal, spatial, existence, quantity, viewpoint, and action changes.
  • Dataset curation: Question-answer pairs are generated from paired images using rule-based, annotation-driven, and model-assisted methods.
  • Dataset curation: Attribute and state examples use industrial damage data, natural images, medical pairs, and synthetic primitives to encode graded visual changes.
  • Dataset curation: Temporal and spatial examples derive from video frames with timestamp or motion annotations, requiring models to infer event order or transformations.
  • Dataset curation: Existence, quantity, quality, viewpoint, and action examples use aerial, synthetic, video, crowd, industrial, and camera-motion sources with human or automated annotations.

4 EXPERIMENT

The experiments evaluate open-source and proprietary VLMs on subtle comparative reasoning, testing model performance, prompting strategies, synthetic difficulty factors, fine-tuning, and downstream transfer. Results show persistent weaknesses in spatial, temporal, and viewpoint reasoning, limited prompting gains, strong sensitivity to controlled visual factors, and better transfer from VLM-SubtleBench than from MLLM-CompBench.

  • Benchmark results: GPT-5-thinking ranks first in 7 of 10 difference types and achieves the highest average accuracy among proprietary models.
  • Benchmark results: GPT-5-thinking achieves 93.1% accuracy on emotion, whereas models perform weakly on temporal, spatial, and viewpoint differences.These weaker categories require common-sense reasoning and spatial understanding.
  • Benchmark results: 43.0% LLM-as-a-judge accuracy for GPT-5-thinking captioning remains below ground-truth captions, while Qwen2.5-VL-72B scores 24.3.Large open-source models can match proprietary systems on CSS but perform substantially worse under LLM-as-a-judge evaluation.
  • Prompting strategies: Reasoning steps improve performance in 9 of 10 domains, while concatenating two images degrades accuracy in 9 of 10 domains.Highlighting helps modestly but declines when brightness or image-quality variation causes localization failures; overlap and subtraction help mainly in layout-dependent settings.
  • Controlled synthetic evaluation: Synthetic controls reveal that attribute accuracy exceeds 70% at roughly 25% brightness shifts, while existence accuracy falls below 60% beyond 32 objects.Quantity approaches the 50% random baseline in scenes containing ten or more objects, and viewpoint performance requires camera translations around 160 pixels.
  • Fine-tuning and transfer: Fine-tuning improves all difference types, but spatial and temporal gains are more modest, while VLM-SubtleBench produces larger downstream gains than MLLM-CompBench.The transfer studies evaluate industrial anomaly detection and aerial surveillance tasks.

5 DISCUSSION

VLM-SubtleBench evaluates subtle comparative reasoning across ten difference types and six domains, where humans succeed but current VLMs struggle. It also serves as a diagnostic tool for understanding systematic gaps and failure modes.

  • VLM-SubtleBench evaluates subtle image-pair changes across ten difference types and six domains.
  • The benchmark can diagnose perceptual capabilities needed by agents in games, robotics, navigation, and graphical interfaces.

A BENCHMARK CURATION DETAIL

The benchmark curates subtle comparative pairs from annotated, edited, and synthetic sources across industrial, natural, medical, video, aerial, and synthetic settings. Annotation procedures select or verify pairs and generate comparative questions for distinct difference types.

  • The test split is organized by difference type, domain, and data source.
  • Attribute: Industrial attribute pairs compare same-type MVTEC-AD anomalies with different defect sizes, while COCO edits object colors into similar alternatives.
  • Attribute: Medical pairs reformulate chest-X-ray comparisons into visually perceptible questions, while synthetic scenes vary one shape’s brightness or size.
  • State, Emotion, Temporal, and Spatial: Video sources provide state, emotion, spatial, and temporal pairs through frame sampling, annotation, and selection of action-relevant or ordered frames.
  • Temporal and Spatial: Temporal pairs use frame order that annotators judge irreversible, whereas spatial pairs include both reversible and non-reversible movements.
  • Existence and Quantity: Aerial and synthetic sources construct existence pairs through small-area changes or object removal/addition, while MVTEC-LOCO supplies quantity-defect pairs.

B.3 CAPTION EVALUATION METRICS

Caption evaluation combines semantic similarity with model-based matching judgments, while overlap and subtraction images provide visual inputs for comparison.

  • Caption-level semantic agreement is measured using cosine similarity between Sentence-BERT embeddings of reference and generated captions.
  • Visual input construction: Overlap images blend two aligned inputs equally, producing a composite that merges both images visually.
  • Visual input construction: Subtraction images highlight maximal pixel-wise change between aligned inputs for visual comparison.
  • GPT-4o judges whether each candidate caption matches its reference, with accuracy defined as the fraction of “Yes” judgments.

B.4 HUMAN EVALUATION

Human evaluation samples 1,000 examples across the ten difference types, while model training and transferability experiments use specified hardware, splits, and standardized formatting. The evaluation materials include highlighted change regions and benchmark-wide model comparisons.

  • Human evaluation: Human evaluation samples 1,000 examples, with 100 examples per difference type and source allocation proportional to each source’s contribution.
  • Training setup: Training uses four NVIDIA A100 GPUs, an effective batch size of 32, three epochs, and joint optimization of all model parameters.
  • Transferability study: The transferability study fine-tunes Qwen-2.5-VL-7B on matched validation subsets and standardizes task formats across datasets.
  • Visual inputs: Highlight images dim the background while preserving original appearance inside selected boxes and marking their boundaries in green.
  • Results reporting: Table 8 reports model-wise accuracy on VLM-SubtleBench, MLLM-CompBench, MMAD, and QAG-360K.

C.1 CORRELATION ANALYSIS WITH DOWNSTREAM BENCHMARKS

The analysis evaluates pure language-based comparison by generating image descriptions and answering questions without the original images. Performance declines for emotion and existence, while quality remains nearly unchanged.

  • Prompt design: Emotion prompts ask models to describe expressed emotion and rate its intensity from 1–10.
  • Prompt design: Existence prompts require listing all visible objects and their approximate locations.
  • Prompt design: Quality prompts ask models to assess image quality on a 1–10 scale using blur, noise, exposure, compression, and related issues.
  • Method: The pure language-based method generates descriptions for each image, then performs VQA using only those descriptions.The original images are withheld during question answering.
  • Results: Emotion and existence performance is lower, whereas quality performance remains nearly identical to the comparison baseline.The authors suggest explicit 1–10 rating prompts may provide a reliable intermediate representation for quality comparison.

C.3 DOMAIN-WISE PERFORMANCE ANALYSIS

Domain-wise results show that o3 and GPT-5-thinking lead most categories, with especially large advantages in synthetic and medical imagery but weaker performance on aerial imagery.

  • Overall performance: o3 and GPT-5-thinking achieve the highest scores in most categories among proprietary models.
  • Domain differences: Both models show a substantial performance gap over other models in the synthetic and medical domains.
  • Domain differences: Both models produce comparatively weaker results on the aerial domain.

C.4 EXTENDED COLOR-SENSITIVITY ANALYSIS

The extended analysis tests color sensitivity under controlled synthetic variation. GPT-4o struggles particularly with green and magenta hues, while brightness and cross-factor interactions show limited effects.

  • Analysis design: The synthetic-control analysis varies color, hue shift, brightness, object size, count, and translation across five representative colors.The colors include two green tones and blue, red, and magenta; hue shifts are measured as ∆E in OKLAB space.
  • Color sensitivity: GPT-4o has significantly lower accuracy distinguishing green hues than red or blue tones.
  • Color sensitivity: Magenta causes the most severe degradation, approaching 0%, revealing a systematic color-specific weakness.
  • Controlled effects: Brightness variation produces no notable color-dependent gap, while color interactions with size, count, and viewpoint are minimal.The analysis indicates that hue discrimination contributes little to the reasoning objective in these cross-factor tasks.

C.7 PERFORMANCE OF PROPRIETARY MODELS BY DATA SOURCE

Across data sources, proprietary models perform strongest on natural and industrial imagery and weaker on synthetic and aerial settings. o3 and GPT-5-thinking lead most categories, especially in synthetic and medical domains.

  • Model comparison: o3 and GPT-5-thinking achieve the highest scores in most categories across data sources.
  • Domain differences: Their performance advantage over other models is substantial in synthetic and medical domains.
  • Domain differences: Models generally perform strongest on natural and industrial imagery, while performance degrades in synthetic and aerial settings.
Loading 2603.07888v1…