Source-linked AI summary
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
Dingming Li, Hongxing Li, Zixuan Wang, Yuchen Yan, Hang Zhang, Siqi Chen, Guiyang Hou, Shengpei Jiang, Wenqi Zhang, Yongliang Shen, Weiming Lu, Yueting Zhuang
TL;DR
VLMs remain limited in cross-viewpoint spatial reasoning, especially when adopting another entity’s perspective. The paper introduces ViewSpatial-Bench and an automated 3D annotation pipeline, then trains MVSM with multi-perspective data. MVSM achieves a reported 46.24% improvement over baselines, while the benchmark’s human-perspective annotation still requires manual labeling.
Problem
VLMs perform adequately from camera-centered frames but struggle to reason about spatial relationships from alternative entity perspectives, a limitation relevant to embodied interaction.
Method
The paper introduces ViewSpatial-Bench, an automated 3D spatial annotation pipeline, and multi-perspective training data for evaluating and fine-tuning viewpoint-aware VLMs.
Results
46.24% improvement over baselines was reported for the Multi-View Spatial Model across ViewSpatial-Bench tasks.
Takeaways & Limitations
The work establishes a benchmark and trained model for spatially intelligent VLMs aligned with human-perspective reasoning in embodied environments.
Takeaways & Limitations
Human-perspective relative-direction annotation requires manual labeling, creating scaling constraints and potential annotator bias.
Abstract
from arXiv · showhide
Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and reasoning about visual content, but significant challenges persist in tasks requiring cross-viewpoint understanding and spatial reasoning. We identify a critical limitation: current VLMs excel primarily at egocentric spatial reasoning (from the camera's perspective) but fail to generalize to allocentric viewpoints when required to adopt another entity's spatial frame of reference. We introduce ViewSpatial-Bench, the first comprehensive benchmark designed specifically for multi-viewpoint spatial localization recognition evaluation across five distinct task types, supported by an automated 3D annotation pipeline that generates precise directional labels. Comprehensive evaluation of diverse VLMs on ViewSpatial-Bench reveals a significant performance disparity: models demonstrate reasonable performance on camera-perspective tasks but exhibit reduced accuracy when reasoning from a human viewpoint. By fine-tuning VLMs on our multi-perspective spatial dataset, we achieve an overall performance improvement of 46.24% across tasks, highlighting the efficacy of our approach. Our work establishes a crucial benchmark for spatial intelligence in embodied AI systems and provides empirical evidence that modeling 3D spatial relationships enhances VLMs' corresponding spatial comprehension capabilities.
1 Introduction
Current VLMs handle camera-centered spatial judgments better than cross-viewpoint reasoning, motivating ViewSpatial-Bench and spatially annotated training data for multi-perspective localization.
- Current VLMs perform adequately on egocentric spatial judgments but struggle with spatial relationships from alternative entity perspectives.
- Perspective-taking is critical for human-machine interaction, spatial navigation, and multi-agent collaboration, especially in 3D environments involving depth, occlusion, and camera pose.
- Sparse three-dimensional spatial annotations and predominantly static, single-view task designs may underlie VLM deficiencies in cross-viewpoint spatial understanding.
- ViewSpatial-Bench evaluates spatial localization from camera and human perspectives across five task types using an automated 3D orientation annotation pipeline.
- 46.24% improvement over baselines was achieved by the Multi-View Spatial Model trained on the resulting multi-viewpoint VQA data.
2 Related Works
VLM spatial reasoning remains largely camera-centric, while existing benchmarks mostly emphasize single-perspective understanding and provide limited multi-task perspective transformation assessment.
- VLM spatial understanding remains fundamentally limited for spatial relationships, object localization, and embodied interaction reasoning.
- Most existing benchmarks evaluate single-perspective spatial understanding in two-dimensional images or compositional spatial queries.
- 3DSRBench and SPHERE address cross-viewpoint spatial understanding but remain insufficient in multi-task comprehensiveness and perspective transformation depth.
3 ViewSpatial-Bench
ViewSpatial-Bench combines ScanNet and MS-CoCo data with automated spatial processing and manual verification to evaluate five localization tasks across camera and human perspectives.
- Construction Pipeline: The construction pipeline creates metadata, extracts spatial relations, generates templated questions, filters data, and manually verifies annotations.
- Task Design: Camera-perspective tasks assess object relative direction and object view orientation from the image or camera frame.
- Task Design: Human-perspective tasks require adopting a character’s viewpoint for relative direction, object gaze orientation, or simulated scene-relative direction.
- Data Sources: The benchmark uses ScanNet for accurate 3D coordinates and MS-CoCo for diverse images containing human subjects and annotated keypoints.
- Labeling and QA: Raw coordinates and orientation angles are discretized into standardized directional labels, while distractor rules control ambiguity and question difficulty.
- Dataset Overview: 5,712 samples across 1,338 unique scenes form a benchmark combining automated 3D annotation with manual verification.
4 Multi-View Spatial Model
The Multi-View Spatial Model uses spatially annotated image-text data and multi-perspective fine-tuning to improve viewpoint-aware spatial reasoning.
- Training Data: Approximately 43K diverse spatially annotated examples were generated to train the Multi-View Spatial Model.
- Evaluation Scope: ViewSpatial-Bench provides a comprehensive camera- and person-viewpoint evaluation framework with broad directional categories and spatial query targets.
- Training Data: Training data uses consistent natural-language templates with standardized directional classifications for spatial relationships.
- Fine-Tuning: Multi-Perspective Fine-Tuning explicitly trains the model to reason from different observational viewpoints and develop a unified representation of 3D spatial relations.
5 Experiments
Experiments show that current VLMs struggle with perspective-dependent spatial localization, while perspective-aware training substantially improves performance and generalization across benchmarks and embodied scenarios.
- 5 Experiments: 34.98% for GPT-4o and 32.56% for Gemini-2.0-Flash barely exceed the 26.33% random-chance baseline on spatial localization.
- 5 Experiments: Camera-perspective accuracy averages 33.2%, below the 35.7% average for human-viewpoint reasoning across most VLMs.
- 5 Experiments: Perspective and task interact asymmetrically: human-perspective Object View Orientation reaches 42.6%, compared with 36.9% for Relative Direction.
- 5 Experiments: MVSM gains 46.24% absolute performance over Qwen2.5-VL (3B), with consistent improvements across all task categories.
- 5.3.1 Transfer Learning Performance: MVSM outperforms its backbone on VSI-Bench Object Relative Direction and Route Planning, including a +9.54% Route Planning gain.
- 5.3.1 Transfer Learning Performance: Figure 4 contrasts MVSM’s correct perspective-taking answers with GPT-4o’s viewpoint errors on representative VSI-App examples.
- 5.3.1 Transfer Learning Performance: On VSI-App’s 50 indoor and outdoor scenes, MVSM improves by +20.00% indoors and +4.00% outdoors, revealing a domain gap.
- 5.3.1 Transfer Learning Performance: Without perspective-aware training, models alternate between human and camera frames within responses, whereas MVSM consistently follows the specified perspective.
6 Conclusions
The paper introduces ViewSpatial-Bench and an automated spatial annotation pipeline to study and improve multi-perspective spatial localization in VLMs. MVSM achieves substantial gains on benchmark tasks and shows generalization to embodied interaction scenarios.
- 6 Conclusions: ViewSpatial-Bench evaluates VLM multi-perspective spatial localization across five distinct task types.
- 6 Conclusions: An automated spatial annotation pipeline and large-scale multi-perspective dataset support training MVSM for improved spatial performance.
- 6 Conclusions: MVSM achieves substantial overall improvements on ViewSpatial-Bench and demonstrates generalization on VSI-Bench and VSI-App embodied-interaction evaluations.
A Limitations
ViewSpatial-Bench has limitations in annotation scalability, environmental coverage, and temporal scope. These constraints bound automation, generalization beyond indoor scenes, and evaluation of dynamic spatial reasoning.
- Annotation Challenges for Human-Perspective Tasks: Human-perspective relative-direction annotation requires manual labeling because natural-image coordinates and environmental contexts prevent full automation.This introduces scaling constraints and potential annotator biases.
- Domain Constraints in Environmental Coverage: Camera-perspective relative-direction tasks use exclusively indoor ScanNet environments, potentially limiting generalizability to outdoor settings.Outdoor scenes differ in spatial scale, object density, and visual characteristics, and transfer experiments indicate a substantial indoor–outdoor domain gap.
- Static vs. Dynamic Spatial Reasoning: ViewSpatial-Bench evaluates static spatial orientation comprehension rather than dynamic scenarios involving moving objects or observers.The authors identify temporal sequences and motion-based reasoning as future extensions for embodied-AI evaluation.
- Future Work: The authors frame these limitations as directions for future research building on ViewSpatial-Bench while addressing its current constraints.The stated directions include reducing reliance on manual annotation, extending environmental coverage, and incorporating temporal reasoning.
B.1 Dataset Collection and Unification
The dataset combines curated imagery, metadata-derived spatial annotations, and template-based question generation. Its construction includes automated processing alongside filtering and human verification.
- ScanNet Data Collection: ScanNet data undergoes three-stage frame sampling: extracting all frames, sampling every 10th frame, then selecting minimal consecutive frames capturing complete scenes.This strategy is intended to optimize benchmark data quality while retaining comprehensive scene information.
- MS-CoCo Data Collection: MS-CoCo samples are filtered for biologically salient objects and consistent gaze–head orientation before automatic orientation processing.Objects must occupy at least 20% of the image area, and samples with substantially divergent gaze and head directions are removed.
- QA Pair Generation: Object metadata and angle annotations populate predefined question templates, with computed angles serving as ground-truth answers for multiple-choice questions.The templates instantiate object names into spatial questions across the dataset’s task construction process.
- QA Pair Generation: The question templates include formulations that ask for an object’s location from a position facing or viewing another object.These variants operationalize perspective-dependent spatial queries using object placeholders such as {object1}, {object2}, and {object3}.
- QA Pair Generation: Table 4 records the prompt templates used to generate spatial reasoning questions and pair them with metadata-derived directional multiple-choice answers.Object names are inserted into the templates before answer pairing.
B.2 Data Statiscs
ViewSpatial-Bench emphasizes humans and objects while covering diverse directional language and common everyday entities. The benchmark also presents model-response examples across question types.
- Object Categories: The dataset’s two major object-category groups are humans and objects, matching its camera- and human-perspective localization design.This category structure is shown through wordcloud analysis and aligns with the benchmark’s dual perspective targets.
- Spatial Prepositions: The benchmark includes comprehensive directional terminology, with primary directions more frequent than compound directions.Examples of primary terms include “front,” “right,” and “left,” while compound terms include “front-left,” “back-right,” and “above-left.”
- Object Frequencies: Common everyday entities, including chairs, tables, sofas, and desks, are well represented in the top-20 object distribution.This distribution is described as supporting practical relevance for embodied-AI spatial reasoning.
- Model Responses: Figures 7–9 provide response examples from different models across multiple ViewSpatial-Bench question types.These examples illustrate model answers rather than introducing another dataset statistic.
B.4 VSI-App Dataset Construction
VSI-App is a curated out-of-distribution benchmark for testing human-perspective spatial reasoning in realistic interaction settings. Its evaluation uses repeated option-order tests and voting to reduce ordering effects.
- Scene Selection: VSI-App contains 200 professionally sourced scene images, evenly divided between indoor and outdoor environments.Two professional annotators selected the images for a dataset evaluating multi-view spatial models under out-of-distribution conditions.
- Scene Selection: Selected scenes must contain rich 3D structure, identifiable human viewpoint references, and explicit human–object spatial relationships supporting interaction.The criteria are intended to simulate complex real-world human–computer interaction environments.
- Question Annotation: Annotators create questions about object positions from a human’s first-person perspective and about navigation from the human’s position to targets.The annotation process conducts in-depth spatial analysis rather than relying solely on template-based question generation.
- Evaluation Purpose: VSI-App evaluates whether multi-view spatial models generalize to realistic human–computer interaction questions using multiple-choice testing.The dataset is designed to assess human-perspective spatial reasoning, generalization, and practical utility.
- Evaluation Details: Evaluation settings include zero-shot testing, while VSI-Bench testing uses lmms-eval with batch size 1 and a maximum of 32 frames.Open-source models use default Transformers generation settings, and each model is evaluated five times for consistency.
- Evaluation Procedure: Each question is tested with five option orderings and five independent runs, then the most frequent answer is selected by voting.This repeated-testing strategy reduces the potential impact of answer-option ordering on predictions.
C.3 Analysis experiment
The analysis tests whether MVSM’s gains reflect genuine spatial understanding rather than shortcut learning and whether its training procedure generalizes across model backbones. It also presents ViewSpatial-Bench examples and compares models across camera- and person-perspective spatial tasks.
- Training Robustness: Table 6 analyzes MVSM training robustness across question formats and model architectures, with MC denoting Multiple Choice and DA denoting Direct Answer.The table organizes the robustness analysis around both format and backbone comparisons.
- Training Format and Shortcut Learning Analysis: The shortcut-learning analysis uses approximately 43K multi-perspective samples spanning all five task categories and converts training formats to reduce option-elimination strategies.The experiments compare training formats on the same multi-perspective spatial dataset.
- Training Format and Shortcut Learning Analysis: 82.09% overall accuracy for multiple-choice training versus 79.34% for direct-answer training, with the small difference supporting gains from spatial understanding rather than option-structure shortcuts.Both formats produced substantial improvements over baseline under identical experimental settings.
- Multi-Backbone Generalization Validation: Qwen2.5-VL(3B) improved from 35.85% to 82.09% (+46.24%), InternVL-2B from 34.98% to 76.45% (+41.47%), and Qwen2.5-VL(7B) from 36.85% to 83.01% (+46.16%).These results are reported across three representative backbones using the same dataset and specified training configurations.
- Multi-Backbone Generalization Validation: The reported gains are consistent across different model families and parameter scales, supporting broader applicability of the training methodology.The comparison covers Qwen2.5-VL at 3B and 7B scales and InternVL-2B.
- ViewSpatial-Bench Examples: Figure 7 compares Qwen2.5-VL(3B), GPT-4o, and MVSM on five spatial reasoning tasks from camera and person perspectives.Figures 7–9 provide ViewSpatial-Bench examples across the benchmark’s visual materials.