Source-linked AI summary
MetaphorVU: Towards Metaphorical Video Understanding
Zhuoqun Li, Boxi Cao, Guiping Jiang, Fangrui Lv, Ruotong Pan, Jianan Wang, Xiangyu Wu, Hongyu Lin, Yaojie Lu, Yong Du, Ruyin Jia, Liyan, Tingting Gao, Han Li, Xianpei Han, Le Sun
TL;DR
MetaphorVU addresses the lack of systematic evaluation for metaphorical video understanding, a high-order capability important for interpreting complex visual meanings. It introduces MetaphorVU-Bench and a knowledge-graph-based inference-time framework, MetaphorBoost; experiments show current MLLMs remain well below humans, while MetaphorBoost consistently improves performance.
Problem
Existing MLLM research largely studies literal video tasks, leaving systematic evidence about metaphorical video understanding and its high-order cross-domain mapping capability limited.
Method
The paper constructs MetaphorVU-Bench and a metaphorical knowledge graph, then uses MetaphorBoost for inference-time mapping augmentation.
Results
Current MLLMs struggle with metaphorical video understanding: leading models average around 64, nearly 20 points below humans, while MetaphorBoost consistently improves performance.
Takeaways & Limitations
Defective cross-domain mapping is the primary identified failure factor, and knowledge-graph augmentation provides a promising direction for improving MLLM metaphorical video understanding.
Abstract
from arXiv · showhide
Metaphorical videos are prevalent across various real-world scenarios to convey complex ideas, and understanding them typically requires high-order cognitive capabilities. The lack of systematic studies on metaphorical video understanding not only constrains the real-world applicability of MLLMs but also impedes the thorough assessment of their high-order cognitive capabilities. To bridge this gap, we propose MetaphorVU-Bench, the first systematic and comprehensive benchmark dedicated to metaphorical video understanding. Through experiments, we find current MLLMs struggle with accurate metaphorical video understanding, lagging far behind human level, primarily due to defective cross-domain mapping. Motivated by this finding, we construct a metaphor knowledge graph as mapping augmentation and propose MetaphorBoost, an inference-time enhancement framework achieving consistent performance improvement. Our benchmark, analysis, and method provide useful insights and a foundation for future research on advancing MLLMs.
1. Introduction
Metaphorical videos convey complex ideas through high-order cross-domain mapping, but existing MLLM research has focused mainly on literal video understanding. MetaphorVU-Bench evaluates this capability systematically, revealing substantial performance gaps and motivating knowledge-graph-based MetaphorBoost.
- Motivation: Metaphorical videos convey complex ideas in real-world settings and require mapping visual elements to underlying concepts.Human interpretation can connect visual elements such as tailcoat pigs, a banquet, and cats under a table with social concepts and implicit criticism or sympathy.
- Research gap: Existing MLLM video research largely emphasizes literal tasks such as object recognition and event description rather than metaphorical understanding.This leaves the transformation from perceived visual signals to deeper semantics insufficiently studied.
- Benchmark: MetaphorVU-Bench is the first comprehensive benchmark for metaphorical video understanding, combining a systematic taxonomy, real-world curation, and rigorous human annotation.Its construction includes eight metaphor types, multi-stage filtering of billions of candidates yielding 860 videos, and strict cross-validation.
- Findings: Current MLLMs struggle with accurate metaphorical video understanding, with leading models averaging around 64 and nearly 20 points below human performance.Evaluation covers 11 representative close-source and open-source MLLMs.
- Findings: More than 80% of MLLM failures arise from defective cross-domain mapping rather than recognition errors.This finding motivates improving the links between visual elements and underlying concepts.
- MetaphorBoost: MetaphorBoost augments inference-time metaphor interpretation by querying a metaphorical knowledge graph for interconnected references.The framework uses recognized content to retrieve knowledge that supports cross-domain mapping and precise interpretations.
2. MetaphorVU-Bench
MetaphorVU-Bench provides a systematic benchmark for metaphorical video understanding through an eight-type taxonomy, real-world video curation, manual annotation, and quality control. Its evaluation framework exposes current MLLMs’ limited metaphor-understanding capability and compares them with human-written interpretations.
- 2.1. Video Metaphor Taxonomy: The taxonomy defines eight video-metaphor types grounded in multimodal metaphor theory, providing a principled basis for systematic evaluation.Examples include body language, atmosphere language, cultural symbolism, naturalistic symbolism, causal montage, and analogical montage.
- 2.2. Benchmark Construction: MetaphorVU-Bench is built from diverse real-world short videos selected through multi-stage filtration and manually annotated with structured metaphor interpretations.The construction process uses large-scale candidate filtering, model-assisted verification, and human selection and annotation.
- 2.2. Benchmark Construction: The benchmark covers diverse video topics and uses suitable video durations and interpretation formats to support real-world metaphorical video-understanding evaluation.Its statistics summarize sample count, average video duration, and average token count of golden interpretations.
- 2.2. Benchmark Construction: Manual annotations specify which visual elements convey which implicit meanings, while cross-validation and multiple reviewers improve reliability and reduce individual errors.The benchmark also removes speech and subtitles so annotation and evaluation rely solely on visual information.
- 2.2. Benchmark Construction: Table 2 compares MLLM results with human-written interpretations on a 100-instance upper-bound sample and reports limited MLLM capability and weak gains from existing reasoning-enhanced methods.The table positions the proposed method as more effective than those existing reasoning-enhanced approaches.
- 2.3. Evaluation Task and Metric: The evaluation formulates metaphor understanding as mapping an input video and title to a thinking process and a structured metaphor interpretation.The system is expected to recognize visual elements, link them to underlying concepts, and explain which elements convey which implicit meanings.
3. MetaphorVU Evaluation
The evaluation compares diverse MLLMs and reasoning-enhanced methods on metaphorical video understanding, finding substantial gaps from humans and limited gains from methods designed for literal video tasks. Error and type analyses identify defective cross-domain mapping as the central weakness.
- Evaluation Setup: The evaluation includes close-source and open-source MLLMs of various scales alongside reasoning-enhanced methods.Baselines include post-training and inference-time scaling approaches, plus metaphor-specific prompt engineering and few-shot examples.
- Evaluation Setup: The experiments use official APIs or repositories and specified post-training weights or inference-time strategies to improve evaluation consistency.Close-source models use official APIs, while open-source models are deployed with vLLM.
- Overall Results: 52.0 average score for Qwen3-VL-8B-Thinking remains far below the human score of 83.4, while Gemini-3-Pro reaches 63.8 but still trails humans.Gemini-3-Pro shows the strongest overall baseline performance, yet remains below human level.
- Overall Results: Inference-time scaling methods developed for recognition and event description provide marginal gains or can degrade metaphorical video understanding.LTR and ViTCoT degrade the Qwen3-VL-8B-Thinking base model, while VideoRFT and Vision-R1 yield only marginal improvements over Qwen2.5-VL-Instruct.
- Detailed Analysis: Most identified deficiencies involve missing, superficial, or improper cross-domain mapping rather than incorrect visual recognition.The analysis concludes that improving links from visual elements to underlying concepts is key to improving performance.
- Detailed Analysis: MLLMs perform worse on metaphor types with richer visual elements and greater cross-domain mapping requirements.The pattern across metaphor types supports mapping augmentation as a core improvement direction.
4. MetaphorBoost
MetaphorBoost augments MLLMs’ cross-domain mapping with a metaphorical knowledge graph queried at inference time, consistently improving metaphorical video understanding. Ablations attribute its gains to external, structured, metaphor-oriented knowledge and reductions in mapping errors.
- Ineffective cross-domain mapping is identified as the primary factor limiting current MLLMs’ metaphorical video understanding.
- Metaphorical Knowledge Graph: MetaphorBoost constructs a metaphorical knowledge graph as an external scaffold for inference-time mapping augmentation.The graph contains 54,687 nodes and 200,268 edges.
- Inference-time Mapping Augmentation: MetaphorBoost extracts visual keywords, retrieves top target concepts within a maximum of h graph hops, and supplies them as references for interpretation.The retained targets are those linked to the most identified keywords.
- Effectiveness: MetaphorBoost consistently improves MLLMs, raising average scores from 33.8 to 37.9, 52.0 to 55.9, and 63.8 to 66.1 across three base models.The reported gains surpass previous post-training or inference-time scaling methods where stated, and 66.1 achieves a state-of-the-art score.
- Ablation Analysis: Ablations show that external knowledge, graph structure, and metaphor-oriented knowledge each provide more effective augmentation than their respective alternatives.The comparisons are against querying the MLLM itself, raw textual datasets, and ConceptNet6.
- Error Analysis: MetaphorBoost reduces missing, superficial, and improper mapping, which is associated with improved metaphorical video interpretation.A case study reports that the framework mitigates all three deficiencies.
5. Related Work
Prior metaphor-understanding research centers on text and images, while video metaphor remains comparatively scarce. Related deep-semantic video benchmarks assess broader reasoning abilities, whereas MetaphorVU focuses specifically on systematic metaphorical video understanding.
- Metaphor Understanding: Textual metaphor research detects metaphor and identifies source and target domains through relationships between tokens.
- Metaphor Understanding: Image metaphor research builds datasets such as internet-meme collections and explores multimodal fusion to improve performance.
- Metaphor Understanding: Video metaphor research is relatively scarce, with recent datasets largely constructed from advertising videos.Compared with text and images, videos are temporal and can convey richer information and complex metaphor.
- Deep-semantic Video Understanding: Deep-semantic video understanding studies extend beyond object recognition and event description to prediction, domain-knowledge problem solving, and inference of underlying event logic.
- Deep-semantic Video Understanding: MMR-V evaluates broad implicit video reasoning, whereas MetaphorVU provides a dedicated taxonomy and benchmark for fine-grained metaphorical video analysis.The two works are described as complementary for evaluating high-order cognitive capabilities.
6. Conclusion
The paper introduces MetaphorVU-Bench and MetaphorBoost to address metaphorical video understanding, finding that current MLLMs struggle primarily because of defective cross-domain mapping. The proposed knowledge-graph-based augmentation consistently improves performance and offers a direction for further research.
- MetaphorVU-Bench is presented as the first systematic benchmark for comprehensive evaluation of metaphorical video understanding.
- Experiments show that current MLLMs struggle with accurate metaphorical video understanding, primarily due to defective cross-domain mapping.
- MetaphorBoost uses a metaphorical knowledge graph to augment mapping at inference time and consistently improve MLLM performance.
- The benchmark, analysis, and method provide insights and a foundation for future research on advancing MLLMs.
Impact Statement
The paper describes benchmark construction and data curation procedures for evaluating metaphorical video understanding. It also states that the work advances machine learning without identifying specific societal consequences requiring emphasis.
- Impact Statement: The impact statement identifies potential societal consequences but says none require specific highlighting.
- Taxonomy: The benchmark’s taxonomy is designed from multimodal metaphor theory and extensions in the video field to support principled evaluation.
- Taxonomy: The taxonomy distinguishes metaphorical realizations involving visual arrangement, symbolic signs, montage, and narrative-based performance.
- Data Curation: Candidate videos are filtered using video introductions, ASR results, audience comments, LLM analysis, MLLM verification, and final human review.
- Data Curation: The resulting curated collection contains 860 videos with definite metaphorical logic, with metaphor types annotated while balancing sample counts as much as possible.
C. Prompt for Evaluation and LLM Judge, and Consistency Experiments
The evaluation prompt guides MLLMs from recognizing visual contents through mapping them to external concepts and implicit meanings before producing metaphor interpretations. Because outputs are free-form text, DeepSeek-V3.2 serves as an LLM judge using prompts designed for this evaluation setting.
- C.1. Prompt for Evaluation: MLLMs first recognize the visual contents appearing in a video.
- C.1. Prompt for Evaluation: The evaluation process establishes mappings from visual contents to external concepts.
- C.1. Prompt for Evaluation: MLLMs then unveil implicit meanings during their thinking process.
- C.1. Prompt for Evaluation: The final output identifies which visual contents convey which implicit meanings.
- C.2. Prompt for LLM Judge: MetaphorVU-Bench uses free-form text for video metaphor interpretations.
- C.2. Prompt for LLM Judge: Rule-based metrics are difficult to align with human evaluation habits for these outputs.
- C.2. Prompt for LLM Judge: The evaluation follows metrics used in previous free-form question-answering studies.
- C.2. Prompt for LLM Judge: DeepSeek-V3.2 is used as the LLM judge, with its detailed evaluation prompt provided separately.
C.3. Consistency Experiments for LLM Judge
The paper validates its LLM-based evaluation by comparing judge scores with human scores, while describing the knowledge-graph construction and inference-time mapping procedure used by MetaphorBoost.
- C.3. Consistency Experiments for LLM Judge: Human annotators scored 100 randomly sampled model-generated metaphor interpretations using the same evaluation guidelines as the LLM judge.
- C.3. Consistency Experiments for LLM Judge: Human and LLM judge scores show a Pearson correlation coefficient of 0.85 with p-value 3e-20.
- C.3. Consistency Experiments for LLM Judge: The reported correlation indicates a strong positive and statistically significant relationship between human and LLM judge scores.
- Metaphorical Knowledge Graph: MetaphorBoost constructs a metaphorical knowledge graph from textual metaphorical datasets containing metaphorical concept pairs.
- Metaphorical Knowledge Graph: GPT-5 translates portions of originally Chinese data into English to support the knowledge graph’s universality.
- Metaphorical Knowledge Graph: DeepSeek-V3.2 extracts metaphorical concept pairs from each text, using the extracted pairs as knowledge-graph nodes.
- Inference-Time Mapping Augmentation: During inference, MetaphorBoost identifies visual elements in the video and outputs them as a keyword list.
- Inference-Time Mapping Augmentation: It queries the metaphorical knowledge graph with the identified visual elements, uses retrieved concepts as augmentation, and generates the final metaphor interpretation.
F. Details of Reasoning-based Baselines
The baselines span post-training, inference-time reasoning, prompting, and few-shot adaptation. Their relevance varies because several target explicit, low-level, logical, mathematical, or domain-specific understanding rather than metaphorical cross-domain mapping.
- Reasoning-Enhanced Methods: Reasoning-enhanced baselines improve base-model reasoning through post-training or inference-time scaling.
- Reasoning-Enhanced Methods: VideoRFT combines supervised fine-tuning with chain-of-thought annotations and reinforcement learning using a semantic-consistency reward.
- Reasoning-Enhanced Methods: Vision-R1 uses a 200K multimodal CoT dataset and Progressive Thinking Suppression Training to refine multimodal reasoning.
- Reasoning-Enhanced Methods: ReAd-R targets implicit marketing logic and persuasive strategies in advertisement videos, but its domain-specific training limits broader generalizability.
- Reasoning-Enhanced Methods: LTR recursively decomposes complex video questions into manageable parts and performs bottom-up reasoning in a language-centric logical tree.
- Reasoning-Enhanced Methods: ViTCoT interleaves visual and textual information so models can re-examine visual content during reasoning, while focusing on explicit content reasoning.
- Prompting and Few-Shot Methods: Prompt Engineering encourages cross-domain mapping from visual contents to implicit meanings through carefully designed prompts.
- Prompting and Few-Shot Methods: Few-shot Example supplies annotated demonstrations showing how explicit visual contents project onto abstract concepts.
G. Experiments about Query Strategy and Hyperparameters
MetaphorBoost queries the metaphorical knowledge graph using bounded multi-hop retrieval and retains highly connected target nodes. Experiments show that structured, connection-based selection and the default hyperparameters outperform tested alternatives on average.
- Query Strategy: MetaphorBoost uses a maximum of h=2 hops and retains the Top-z=10 target nodes most associated with the query keywords.
- Query Strategy: The query strategy is designed to exploit the knowledge graph’s support for multi-hop and structured reasoning.
- Query Strategy: Average performance decreases when results are retained randomly instead of selecting those with the most common connections to query keywords.
- Query Strategy: The authors interpret the result as evidence that structured federated querying provides low-noise augmentation.
- Hyperparameter Analysis: Variants using h=1 or z=5 have lower average scores than the default settings, despite fluctuations across subsets.
H. More Examples of MetaphorVU-Bench
This section provides additional examples of eight metaphor types and documents the prompts and guidelines used to filter, annotate, evaluate, and analyze metaphorical videos. The workflow includes extracting visual elements, identifying explicit–implicit concept pairs, and scoring interpretation completeness and correctness.
- Additional metaphor examples: Eight metaphor types are illustrated: Body Language, Atmosphere Language, Cultural Symbol, Naturalistic Symbol, Causal Montage, Analogical Montage, Surreal Narrative, and Performative Narrative.The examples note that videos may contain multiple metaphor types, while figures show a dominant type for convenient illustration.
- Mapping deficiencies: MetaphorBoost is described as mitigating missing, superficial, and improper mapping deficiencies that collectively contribute to poor metaphorical video interpretation.The figure presents mapping augmentation as improving MLLM performance on metaphorical video understanding.
- Benchmark evaluation: The benchmark evaluates whether a VLM interprets implicit ideas expressed through a short video's presented content.The evaluation prompt focuses on the video's actual content rather than titles or embedded text.
- Data construction and analysis: Additional prompts support metaphor interpretation, evaluation, LLM and human filtration, and interpretation of video comments as evidence of metaphorical expression.The filtration task can determine whether a video likely uses metaphor based solely on its comment section.
- Benchmark evaluation: The evaluation uses a strict binary score requiring complete, error-free coverage and a loose 0-to-10 score requiring core correctness.The strict score is awarded only when all major metaphorical meanings are covered without significant errors or contradictions.
- Metaphor analysis workflow: The analysis workflow extracts key visual content elements and metaphorical association pairs linking explicit concepts to the abstract concepts they imply.The prompts separately identify video elements and semantic relations between explicit and implicit concepts.