Source-linked AI summary

MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao, Yining Li, Dahua Lin, Kai Chen

arXiv:2406.14515v3cs.CVcs.MM

TL;DR

Existing VideoQA benchmarks inadequately cover long-form content, fine-grained abilities, and temporal comprehension, while open-ended answer scoring can be unreliable. MMBench-Video addresses these gaps with diverse YouTube videos, human-created free-form questions, a hierarchical capability taxonomy, and GPT-4 evaluation, revealing substantial performance limitations in current Video-LLMs.

  • Problem

    Existing VideoQA benchmarks are often short, limited to basic capabilities, and evaluated with GPT-3.5-based scoring that has lower accuracy and stability.

  • Method

    MMBench-Video uses long-form YouTube videos, free-form human-created QA pairs covering 26 fine-grained capabilities, temporal-question curation, and GPT-4-based scoring.

  • Results

    Current Video-LLMs substantially underperform proprietary LVLMs and some open-source LVLMs on spatial and temporal understanding.

  • Takeaways & Limitations

    MMBench-Video provides a quantitative benchmark for comparing LVLM video understanding across diverse capabilities and temporal demands.

  • Takeaways & Limitations

    The evaluation covers a curated model selection, uses GPT-4 as judge, and limits video clips to 6 minutes because existing Video-LLMs have limited capabilities.

Abstract

from arXiv · show

The advent of large vision-language models (LVLMs) has spurred research into their applications in multi-modal contexts, particularly in video understanding. Traditional VideoQA benchmarks, despite providing quantitative metrics, often fail to encompass the full spectrum of video content and inadequately assess models' temporal comprehension. To address these limitations, we introduce MMBench-Video, a quantitative benchmark designed to rigorously evaluate LVLMs' proficiency in video understanding. MMBench-Video incorporates lengthy videos from YouTube and employs free-form questions, mirroring practical use cases. The benchmark is meticulously crafted to probe the models' temporal reasoning skills, with all questions human-annotated according to a carefully constructed ability taxonomy. We employ GPT-4 for automated assessment, demonstrating superior accuracy and robustness over earlier LLM-based evaluations. Utilizing MMBench-Video, we have conducted comprehensive evaluations that include both proprietary and open-source LVLMs for images and videos. MMBench-Video stands as a valuable resource for the research community, facilitating improved evaluation of LVLMs and catalyzing progress in the field of video understanding. The evalutation code of MMBench-Video will be integrated into VLMEvalKit: https://github.com/open-compass/VLMEvalKit.

1 Introduction

MMBench-Video addresses shortcomings in existing VideoQA benchmarks by evaluating long-form video understanding across diverse fine-grained capabilities with free-form questions. GPT-4-based scoring supports broad LVLM evaluation, revealing substantial limitations in current Video-LLMs.

  • Motivation: Existing video understanding benchmarks often focus on specific tasks and short videos, limiting coverage of real-world, zero-shot, contextual, emotional, and linguistic understanding.The paper identifies short-video coverage, limited capabilities, and biased open-ended-answer evaluation as central gaps.
  • Benchmark: MMBench-Video contains approximately 600 YouTube videos, roughly 2,000 volunteer-created QA pairs, and 26 fine-grained capabilities.The videos span 16 major categories, range from 30 seconds to 6 minutes, and prioritize temporally indispensable questions.
  • Evaluation: GPT-4-based evaluation improves the accuracy, consistency, and human-judgment alignment of automated scoring for variable-length, free-form answers.The evaluator emphasizes semantic similarity while overlooking minor differences in language organization.
  • Evaluation: Comprehensive evaluations compare open-source Video-LLMs with open-source and proprietary LVLMs across diverse capabilities.Performance rankings are used to compare models and identify their limitations.
  • Findings: Current Video-LLMs substantially underperform proprietary LVLMs and some open-source image LVLMs, with gaps in both spatial and temporal understanding.The paper reports this pattern on MMBench-Video and image VQA benchmarks.

2 Related Work

Prior LVLM and VideoQA research established multimodal modeling and free-form evaluation, but existing VideoQA benchmarks remain constrained in domain coverage, answer format, temporal scope, and evaluation reliability.

  • LVLMs: LVLMs connect pretrained vision and language components through mechanisms such as gated cross-attention, Querying Transformers, and instruction tuning.Representative systems include Flamingo, BLIP, and LLaVA.
  • VideoQA benchmarks: VideoQA benchmarks span domains including movies, television, video games, synthetic scenarios, and egocentric videos.These datasets commonly evaluate models trained on their respective training sets with concise answers.
  • Evaluation: Video-ChatGPT uses GPT-3.5 to score free-form VideoQA responses, but the approach has limited accuracy, stability, and alignment with human preferences.The evaluated benchmarks primarily cover concept existence, object relationships, and activity recognition.
  • MMBench-Video: MMBench-Video organizes video understanding into three hierarchical ability levels containing 26 leaf abilities.Its taxonomy includes perception, reasoning, hallucination, commonsense reasoning, and temporal reasoning dimensions.

3 MMBench-Video

MMBench-Video is constructed as a diverse, long-form, multi-shot benchmark with hierarchical capability coverage, temporally relevant free-form questions, and GPT-4 adjudication. Its statistics and controlled comparisons indicate stronger temporal demands than existing VideoQA datasets.

  • Video collection: MMBench-Video collects YouTube videos across 16 major categories and excludes videos shorter than 30 seconds from collection.Question-answer pairs are derived from clips no longer than 6 minutes to balance duration and task complexity.
  • Capability taxonomy: The benchmark uses a three-level capability taxonomy with Perception and Reasoning at the top level and 26 fine-grained leaf abilities.Additional video-specific dimensions include Hallucination, Commonsense Reasoning, and Temporal Reasoning.
  • Question composition: Questions prioritize temporal reasoning and reduce static questions that can be answered from nearly any video frame.Static questions remain for necessary coarse capabilities such as Video Style and Video Topic.
  • Question composition: Human-created questions are cross-validated and filtered with an LVLM-based mechanism to remove a portion of static questions.Each video has multiple independent questions targeting one or more leaf capabilities, and questions are free-form.
  • Evaluation paradigm: GPT-4 assigns 0-to-3 scores based on content similarity between model outputs and ground-truth answers.Experiments report strong consistency and alignment with human assessments.
  • Dataset statistics: MMBench-Video reaches a maximum of 210 shots and substantially exceeds other benchmarks in average shot count.Its shot-number distribution is long-tailed, supporting its long-form, multi-shot design.
  • Temporal indispensability: GPT-4o retains 47.8% of its 8-frame efficacy with one frame on MMBench-Video, versus over 75% on previous VideoQA datasets.Its normalized one-frame score on MMBench-Video is 26.0%, indicating greater temporal indispensability.

4 Experiment

Evaluations compare open-source and proprietary LVLMs and Video-LLMs on MMBench-Video, including frame-count and subtitle-based settings. Results reveal substantial gaps between model groups and show that GPT-based judging and speech context affect evaluation outcomes.

  • Open-Source Video-LLMs: Open-source Video-LLMs perform comparably poorly on MMBench-Video, with LLaVA-NeXT-Video reaching only 1.14 out of 3.VideoChat2’s 18% advantage over Video-ChatGPT on MSVD-QA narrows to 6% on MMBench-Video.
  • Open-Source LVLMs for Images: VILA-1.5-40B achieves the highest open-source image-LVLM score, reaching 1.61 with 14 input frames.Except for Qwen-VL-Chat, open-source image LVLMs improve substantially when using 8 frames instead of one.
  • Proprietary LVLMs for Images: GPT-4o achieves 1.86 with 16 frames, outperforming the best open-source video LLM by 63% and the best open-source image LVLM by 16%.Increasing sampled frames improves perception because adjacent frames support interpretation and reveal previously missed content.
  • Incorporating Speech: Adding YouTube-generated subtitles improves GPT-4o’s performance, but richer speech context also increases hallucination risk and requires balancing information density with redundancy.The evaluation uses subtitles as supplementary prompt context on the subset with available subtitles.

5 Conclusion

MMBench-Video is presented as a long-form, multi-shot VideoQA benchmark for evaluating LVLM video understanding across diverse topics and fine-grained capabilities. Its evaluations identify significant spatial and temporal understanding limitations in existing Video-LLMs.

  • Conclusion: MMBench-Video evaluates LVLM understanding of long-form, multi-shot videos spanning diverse topics and fine-grained capabilities.The benchmark is designed specifically for video content understanding.
  • Conclusion: Extensive evaluations identify significant spatial and temporal understanding limitations among existing Video-LLMs.

A OpenSource Datasets and Codes

The MMBench-Video dataset and evaluation code are publicly released, with additional analysis examining speech information across the full benchmark. Speech remains unavailable for approximately half of the videos.

  • Datasets and Codes: The full MMBench-Video dataset is available on HuggingFace, and evaluation code is released through VLMEvalKit.Test results are published in the OpenVLM Video Leaderboard, and a DataSheet is provided.
  • Datasets and Codes: The dataset is open-sourced under the Attribution-NonCommercial 4.0 International license.
  • Speech Across the Full Benchmark: Approximately half of MMBench-Video videos lack parseable video title tracks, yet subtitles still significantly improve performance across the full dataset.The analysis reports a 200% boost for GPT-4o with subtitles when tested without visual input.
  • Speech Across the Full Benchmark: Table 8 evaluates GPT-4o’s improvement from incorporating YouTube-generated subtitles over the whole dataset.

B.2 Detailed Analysis of L-2 Capability

Detailed analyses examine hallucination, frame-count effects, temporal reasoning, and performance across video lengths and shot counts. Model behavior is especially sensitive to shot transitions and the availability of diverse temporal information.

  • Hallucination: Hallucination is the most significant L-2 perceptual limitation for all evaluated Video-LLMs, unlike state-of-the-art proprietary LVLMs.Video-LLMs tend to answer questions about nonexistent visual content rather than dismissing uncertain queries.
  • Frame Count: Increasing frame counts improves perceptual capabilities more than reasoning capabilities, while proprietary LVLMs perform better in logical, commonsense, and hallucination-related capabilities.
  • Temporal Reasoning: Idefics2-8B-[1f] outperforms all Video-LLMs on temporal reasoning despite receiving only one image.The finding suggests that Video-LLMs are not effectively leveraging diverse temporal information.
  • Video Length and Shots: Figure 6 plots model scores against video shot count and length for InternVL-Chat-V1.5, GPT-4o, and Video-LLaVA under specified frame-sampling settings.
  • Video Length and Shots: Video length and shot count affect model scores, with longer or multi-shot videos generally requiring more visual content.
  • Video Length and Shots: GPT-4o’s score falls to 75% of its original value above 50 shots, indicating stronger sensitivity to shot count than video length.Frequent shot transitions make videos more difficult for the model to comprehend.

C.1 Temporal Dispensable Data Filtering with LVLM

MMBench-Video filters questions for temporal indispensability by testing whether single random frames can answer them, then manually verifies the retained set.

  • GPT-4v evaluates each question using four inferences from independently sampled random frames.GPT-4 scores the responses, and the average score determines temporal relevance.
  • Questions with high average single-frame scores are filtered out to preserve temporally necessary VideoQA instances.

C.2 Question Type Analysis

MMBench-Video broadens question types beyond conventional interrogatives and deliberately balances their distribution, unlike the skewed distributions of MSVD and MSRVTT.

  • Motivation: The benchmark’s question set was curated because existing VideoQA benchmarks cover a limited range of question types.
  • Question type coverage: MMBench-Video includes ‘why’, ‘which’, ‘is / are’, and ‘does / do’ alongside ‘what’, ‘who’, ‘how’, ‘when’, and ‘where’.The expanded interrogative set is intended to better match natural human dialogue.
  • Distribution comparison: MSVD and MSRVTT concentrate over 60% of questions in ‘what’, while ‘when’ and ‘where’ together comprise only 1%.
  • Distribution comparison: MMBench-Video keeps ‘what’ most prevalent but distributes the remaining interrogative forms more evenly.

E Limitations and Broader Impacts

The benchmark evaluates long-form multi-shot VideoQA but remains bounded by model coverage, judge-model choice, and a six-minute video limit. Its small scale may also leave topics and fine-grained capabilities uncovered.

  • Limitations: The evaluation covers representative open-source and proprietary VLMs rather than all recent high-performing models.The authors attribute this scope to budget constraints.
  • Limitations: GPT-4 serves as the judge, while the feasibility of state-of-the-art open-source LLM judges remains for future study.
  • Limitations: Videos are capped at 6 minutes because existing Video-LLMs have limited capabilities, rather than extending to tens of minutes or hours.
  • Broader Impacts: MMBench-Video’s small scale may omit video topics and fine-grained capabilities, limiting representativeness for specific tasks or scenarios.
  • Broader Impacts: The benchmark provides insights for model optimization and reports limited spatial and temporal understanding in evaluated video-LLMs.

F Datasheet for Datasets.

The datasheet describes MMBench-Video as a publicly released, evaluation-only dataset of YouTube videos paired with human-authored questions and answers targeting fine-grained video understanding.

  • Composition: MMBench-Video contains 609 videos and 1,998 question-answer pairs collected from public YouTube sources.
  • Composition: Each instance includes a 16.9-second-to-6-minute video, question, answer, video category, and targeted fine-grained capability.YouTube subtitles may also be included when applicable.
  • Usage: The dataset is designed only for evaluation and has no recommended training, development, or validation splits.
  • Collection Process: Authors and undergraduate volunteers manually propose questions and corresponding answers based on the videos.The participants are reported to receive a fair wage.
  • Quality and preprocessing: Videos are human-selected, manually quality-checked for ethical violations, and some raw videos are trimmed according to annotations.
  • Distribution: The benchmark is publicly released as a HuggingFace Dataset under a CC BY-NC 4.0 license.

Checklist

The checklist records how the study addresses reporting, asset, participant, and ethics requirements. It also notes that evaluation code was not yet publicly available and that limitations were discussed in the supplementary materials.

  • The authors state that the abstract and introduction accurately reflect the benchmark’s contributions and scope.
  • Limitations are reported as discussed in the supplementary materials.
  • The new dataset and instructions are provided in the paper and supplement, while evaluation code and results were not yet released publicly.
  • Existing datasets are cited with licensing information, although the exact license for TGIF-QA could not be found.
  • The data are described as publicly available, free for research use, and without personally identifiable information or offensive content.
  • Participants received an estimated hourly wage of around 5 US dollars, with around 2000 US dollars spent on compensation.
Loading 2406.14515v3…