Source-linked AI summary
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
Chaoyou Fu, Haozhi Yuan, Yuhao Dong, Yi-Fan Zhang, Yunhang Shen, Xiaoxing Hu, Xueying Li, Jinsen Su, Chengwu Long, Xiaoyao Xie, Yongkang Xie, Xiawu Zheng, Xue Yang, Haoyu Cao, Yunsheng Wu, Ziwei Liu, Xing Sun, Caifeng Shan, Ran He
TL;DR
Video-MME-v2 targets the gap between saturated benchmark scores and trustworthy video understanding by introducing a progressive hierarchy and group-based nonlinear evaluation. Built with extensive human quality control, it finds large model–human differences, hierarchical performance bottlenecks, and dependence on textual cues.
Problem
Existing video benchmarks lack comprehensive capability hierarchies and often rely on per-question accuracy, limiting evaluation of robust and faithful video understanding.
Method
Video-MME-v2 combines three progressive capability levels with consistency- and coherence-based question groups scored using a nonlinear evaluation framework and controlled human annotation.
Results
Human experts achieved a Non-Lin Score of 90.7 versus 49.4 for Gemini-3-Pro, while errors in visual aggregation and temporal modeling constrained higher-level reasoning.
Takeaways & Limitations
The benchmark exposes consistency, reasoning-coherence, hierarchical-understanding, and text-dependence weaknesses in current video MLLMs.
Takeaways & Limitations
Scores on Action & Motion and Physical World Reasoning remain below 30 even for state-of-the-art models.
Abstract
from arXiv · showhide
With the rapid advancement of video understanding, existing benchmarks are becoming increasingly saturated, exposing a critical discrepancy between inflated leaderboard scores and real-world model capabilities. To address this widening gap, we introduce Video-MME-v2, a comprehensive benchmark designed to rigorously evaluate the robustness and faithfulness of video understanding. To systematically evaluate model capabilities, we design a \textbf{progressive tri-level hierarchy} that incrementally increases the complexity of video comprehension, ranging from multi-point visual information aggregation, to temporal dynamics modeling, and ultimately to complex multimodal reasoning. Besides, in contrast to conventional per-question accuracy, we propose a \textbf{group-based non-linear evaluation} strategy that enforces both consistency across related queries and coherence in multi-step reasoning. It penalizes fragmented or guess-based correctness and assigns credit only to answers supported by valid reasoning. To guarantee data quality, Video-MME-v2 is constructed through a rigorously controlled human annotation pipeline, involving 12 annotators and 50 independent reviewers. Backed by \textbf{3,300 human-hours} and up to \textbf{5 rounds} of quality assurance, Video-MME-v2 aims to serve as one of the most authoritative video benchmarks. Extensive experiments reveal a substantial gap between current best model Gemini-3-Pro and human experts, and uncover a clear hierarchical bottleneck where errors in visual information aggregation and temporal modeling propagate to limit high-level reasoning. We further find that thinking-based reasoning is highly dependent on textual cues, improving performance with subtitles but sometimes degrading it in purely visual settings. By exposing these limitations, Video-MME-v2 establishes a demanding new testbed for the development of next-generation video MLLMs.
1 Introduction
Video-MME-v2 addresses incomplete and unreliable video-MLLM evaluation with a progressive hierarchy and group-based assessment. Its controlled construction and experiments expose substantial gaps in model reliability, hierarchical understanding, and reasoning coherence.
- Existing benchmarks lack a comprehensive capability hierarchy and rely heavily on isolated per-question accuracy, limiting holistic assessment of video MLLMs.
- Video-MME-v2 evaluates video understanding through three progressive levels: information aggregation, temporal dynamics modeling, and complex reasoning.
- Its group-based strategy measures capability consistency across related tasks and reasoning coherence across temporally and causally linked questions.
- 12 annotators and 50 reviewers contributed more than 3,300 human-hours to curate 800 videos and 3,200 questions through multi-stage quality control.
- Human experts scored 90.7, whereas Gemini-3-Pro reached 49.4, and accumulated lower-level errors limited higher-level reasoning.
- Thinking modes improve performance with subtitles but can regress without textual cues, while group-based nonlinear scoring reveals inconsistency hidden by per-question accuracy.
2 Related Work
Video-MLLM research has progressed from frame-based comprehension toward complex reasoning, while benchmarks have diversified across specialized and general capabilities. However, existing evaluations remain fragmented and often insufficiently comprehensive.
- Recent video MLLMs increasingly address complex reasoning through reinforcement learning and tool use, extending earlier frame-sequence approaches.
- Existing benchmarks cover specialized abilities such as action understanding and long-video comprehension, alongside broader but relatively basic general-video evaluations.
- Figure 1 contrasts models using group-based nonlinear rankings while showing average accuracy only as a reference measure.
3 Benchmark Design
Video-MME-v2 organizes video comprehension into progressive capability levels and evaluates both breadth of competence and depth of reasoning. Its nonlinear group metrics reward consistent, coherent performance rather than isolated correct answers.
- Benchmark Structure: The benchmark contains three hierarchical levels, 12 sub-categories, and over 30 task types spanning aggregation, temporal modeling, and complex reasoning.
- Capability Hierarchy: Level 1 tests visual recognition, cross-modal consistency, and basic counting; Level 2 tests motion, ordering, and causal reasoning.
- Capability Hierarchy: Level 3 evaluates narrative understanding, social dynamics, and physical-world reasoning through professional knowledge and multi-hop inference.
- Group Design: Question groups model relationships among related queries to assess both capability consistency and multi-step reasoning coherence.
- Metrics: The evaluation combines conventional per-question accuracy with nonlinear group scores that emphasize consistency and reasoning coherence.
- Metrics: Each group score depends on the joint correctness pattern of its four questions rather than treating answers independently.
- Metrics: Consistency groups score four related answers as (N /4)^2, penalizing isolated guesses and rewarding performance across multiple capability facets.
4 Dataset Construction and Annotation
Video-MME-v2 combines recency-oriented, diverse video curation with group-based question design and extensive human, model, and independent quality control. Its construction targets leakage, guessing shortcuts, ambiguity, and weak distractors while scaling to 800 videos with four questions each.
- Dataset scale and quality control: 3,300 human-hours produced 800 videos, each paired with 4 questions and 8 answer options, through a pipeline involving 12 annotators and 50 reviewers.The process also includes recency-oriented curation and explicit decontamination.
- Video curation: Over 80% of videos were published in 2025 or later, including nearly 40% published after October 2025, to mitigate pretraining leakage.The collection process specifically targeted newly released internet content.
- Video curation: Four top-level domains and 31 fine-grained subcategories provide balanced diversity across subject matter and visual styles.The domains are Sports & Competition, Lifestyle & Entertainment, Art & Literature, and Knowledge & Education.
- Video curation: Mean and median view counts of 4.83 million and 355 thousand support filtering toward higher-exposure, human-curated content.The dataset excludes low-exposure videos; 84.3% exceed 10,000 views and 94.4% exceed 1,000 views.
- Video curation: Manual screening excludes classic films, television works, and flagship influencer content to reduce memorization effects.The stated goal is for performance to reflect perception and reasoning rather than retrieval of training-time memories.
- Question and option design: Questions and answers lengthen from Q1 to Q4 while option word counts remain consistent, supporting progressively deeper reasoning without length-based guessing.The group-based construction uses later questions as more difficult links in logical chains, and eight options reduce random-guessing probability to 12.5%.
- Annotation and verification: Three cross-review rounds, independent blind testing, text-only baseline checks, and correction-reverification address ambiguity, language-only shortcuts, and revision quality.Frontier-model adversarial testing and manually refined adversarial distractors further test fine-grained discrimination.
5 Experiments
Experiments show a substantial reliability gap between models and humans, with performance degrading across the benchmark’s hierarchy. Results also reveal that robustness depends on capability completeness, model scale, group consistency, and access to textual or audio cues.
- Benchmark Results: 90.7 versus 49.4: human experts substantially outperform Gemini-3-Pro on the group-based nonlinear score.The reported gap is 41.3 points.
- Hierarchical Bottlenecks: Level 1 to Level 3 performance declines monotonically across models, as lower-level errors propagate into temporal modeling and complex reasoning.The authors identify perception and temporal grounding as prerequisite components for reliable multi-hop reasoning and global consistency.
- Model Comparisons: 49.4 versus 39.1: Gemini-3-Pro outperforms Qwen3.5-397B Think, while open-source models generally show larger drops without subtitles or audio.The comparison suggests greater dependence on textual or auxiliary multimodal cues among current open-source models.
- Audio and Text Cues: 49.4 from 38.2 (+11.2) for Gemini-3-Pro and 38.6 from 29.9 (+8.7) for MiMo-v2-Omni when native audio is available.For omni models, the with-subtitle/audio setting includes raw audio, which provides complementary semantic and paralinguistic information.
- Model Scale: 31.4: Qwen3.5-27B-Think exceeds Qwen3-VL-235B-Instruct at 25.0, showing that smaller models can remain competitive with strong post-training and reasoning alignment.The results indicate that model quality factors can sometimes outweigh sheer parameter scale.
- Group-Based Evaluation: 66.1% versus 49.4: Gemini-3-Pro’s average accuracy exceeds its Non-Lin Score, showing that correlated-question consistency is weaker than isolated correctness suggests.The group-based score emphasizes consistency across related queries rather than isolated correct predictions.
6 Conclusion
Video-MME-v2 introduces a rigorous benchmark combining progressive evaluation levels with group-based nonlinear scoring to assess video MLLM robustness and faithfulness.
- Video-MME-v2 evaluates video MLLMs through a progressive multi-level hierarchy spanning diverse video understanding tasks.
- Its group-based framework probes capability robustness through consistency groups and reasoning faithfulness through coherence groups.
- Non-linear scoring is applied within each group to sharpen evaluation beyond conventional per-question assessment.
- Experiments demonstrate the benchmark’s challenging nature and analyze model capabilities and reasoning behaviors.
Contributors and Acknowledgements
This section lists the paper’s contributors and identifies senior leads.
- Chaoyou Fu, Haozhi Yuan, Yuhao Dong, and Yi-Fan Zhang are marked as core contributors with equal contributions.
- Xing Sun, Caifeng Shan, and Ran He are identified as senior leads.