Source-linked AI summary
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Caifeng Shan, Ran He, Xing Sun
TL;DR
MLLM evaluation has focused largely on static images, leaving comprehensive assessment of sequential video understanding underexplored. Video-MME addresses this gap with a manually annotated benchmark covering diverse domains, durations, and modalities, and finds Gemini 1.5 Pro leading commercial models while longer multimodal sequences remain challenging.
Problem
Existing MLLM development and evaluation primarily focus on static visual data, leaving sequential video understanding insufficiently assessed.
Method
Video-MME manually curates 900 diverse videos and annotates 2,700 multiple-choice questions spanning domains, durations, subtitles, audio, and temporal reasoning.
Results
Gemini 1.5 Pro is the best-performing commercial model, achieving 75% average accuracy, while VILA-1.5 reaches 59% overall accuracy.
Takeaways & Limitations
Video-MME’s findings underscore the need for improved handling of longer multimodal data in MLLMs.
Takeaways & Limitations
Long-video understanding remains constrained by restricted input frames and limited availability of datasets focused on complex temporal reasoning.
Abstract
from arXiv · showhide
In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image understanding. The potential of MLLMs in processing sequential visual data is still insufficiently explored, highlighting the absence of a comprehensive, high-quality assessment of their performance. In this paper, we introduce Video-MME, the first-ever full-spectrum, Multi-Modal Evaluation benchmark of MLLMs in Video analysis. Our work distinguishes from existing benchmarks through four key features: 1) Diversity in video types, spanning 6 primary visual domains with 30 subfields to ensure broad scenario generalizability; 2) Duration in temporal dimension, encompassing both short-, medium-, and long-term videos, ranging from 11 seconds to 1 hour, for robust contextual dynamics; 3) Breadth in data modalities, integrating multi-modal inputs besides video frames, including subtitles and audios, to unveil the all-round capabilities of MLLMs; 4) Quality in annotations, utilizing rigorous manual labeling by expert annotators to facilitate precise and reliable model assessment. 900 videos with a total of 254 hours are manually selected and annotated by repeatedly viewing all the video content, resulting in 2,700 question-answer pairs. With Video-MME, we extensively evaluate various state-of-the-art MLLMs, including GPT-4 series and Gemini 1.5 Pro, as well as open-source image models like InternVL-Chat-V1.5 and video models like LLaVA-NeXT-Video. Our experiments reveal that Gemini 1.5 Pro is the best-performing commercial model, significantly outperforming the open-source models. Our dataset along with these findings underscores the need for further improvements in handling longer sequences and multi-modal data. Project Page: https://video-mme.github.io
1. Introduction
Video-MME addresses the limited evaluation of MLLMs on sequential visual data by providing a manually curated benchmark spanning diverse domains, durations, and modalities. Its evaluation finds strong commercial-model performance but persistent weaknesses on longer videos and multimodal inputs.
- Video-MME evaluates MLLMs on sequential visual data, addressing benchmarks’ predominant focus on static visual understanding.
- 75% average accuracy makes Gemini 1.5 Pro the highest-performing commercial model, while VILA-1.5 reaches 59% overall accuracy.Open-source models show substantial gaps relative to commercial models.
- MLLM performance generally declines as video length increases, identifying longer-sequence processing as a critical bottleneck.The paper discusses architectural context extension and temporal-reasoning training data as potential improvement directions.
2. Related Work
Prior work has advanced MLLM architectures and image benchmarks, while video-specific models increasingly incorporate temporal modeling and localization. Video-MME extends this trajectory by targeting comprehensive evaluation of video understanding.
- MLLMs typically combine a vision encoder, modality alignment module, and LLM backbone to process multimodal inputs.Representative components include CLIP or SigLIP encoders and LLaMA or Vicuna language backbones.
- Video-LLM architectures encode frames into vision tokens and use temporal modules or Q-Formers to model and compress video information.Video-LLaMA and VideoChat2 exemplify this approach.
- Video benchmarks complement architectural advances by evaluating perception, cognition, scientific understanding, mathematical reasoning, and multidisciplinary capabilities.
3. Video-MME
Video-MME constructs a diverse, manually reviewed benchmark spanning video types, durations, modalities, and temporal difficulty. Its statistics and examples show standardized QA formatting, increasingly information-dense subtitles, and questions requiring broad or multimodal video comprehension.
- Dataset Construction: The dataset construction combines video collection, multiple-choice QA annotation, and rigorous manual quality review.Annotators view each video fully, create three questions with four options, and independently review QA clarity, answerability, and correctness.
- Dataset Construction: 900 videos span 6 domains and 30 fine-grained categories, supporting broad coverage across video scenarios.The domains include Knowledge, Film & Television, Sports Competition, Life Record, and Multilingual; the supplied passage lists five domain names while the benchmark description identifies six.
- Dataset Construction: Gemini 1.5 Pro achieves less than 15% accuracy on text-only questions after filtering, indicating that retained QA pairs require video content as a critical clue.Questions answerable from text alone are excluded or returned for revision.
- Dataset Statistics: Question, option, and answer lengths remain consistent across video lengths, while subtitle volume rises from 198.6 words for short videos to 6.5K words for long videos.The option distribution is near-uniform at 25.1/27.2/25.3/22.4% for A/B/C/D.
- Dataset Statistics: Video-MME includes 6 key domains, 30 subfields, a full spectrum of video lengths, and varied question types for evaluating temporal understanding.Qualitative cases require integrating frames with audio or subtitles, arithmetic reasoning, and information distributed across videos up to 30 minutes.
- Dataset Statistics: 26s, 164.7s, and 890.7s are the median certificate lengths for short, medium, and long videos, respectively.Certificate length measures the total duration of the minimum necessary and sufficient sub-clips needed to convince a human verifier that an annotation is correct.
4. Experiments
Video-MME evaluates commercial, open-source video, and image MLLMs across tasks, modalities, and video durations. Results show strong commercial-model performance but substantial challenges for open-source models, long videos, and temporal reasoning.
- Evaluation Settings: Four commercial, nine open-source video, and three image-based MLLMs are evaluated on Video-MME.The benchmark supports both video and multi-image evaluation settings.
- Quantitative Results: 75% accuracy makes Gemini 1.5 Pro the best-performing commercial model, exceeding GPT-4V by 15.1% and GPT-4o by 3.1%.These results use video frames alone.
- Quantitative Results: Image-based Qwen-VL-Max and InternVL-Chat-V1.5 achieve performance comparable to LLaVA-NeXT-Video, supporting Video-MME's applicability to both image and video MLLMs.The findings also indicate that image understanding is foundational to video understanding.
- Modality Analysis: +16.7% and +12.5% accuracy improvements occur on long multilingual videos when Gemini 1.5 Pro adds subtitles and audios, respectively, over frames alone.Subtitles provide greater assistance than audios across video durations, while multilingual performance may depend on subtitle quality.
- Duration Analysis: Both commercial and open-source models decline as video duration increases, partly because longer videos contain harder reasoning questions and sparser frame information.Fixed frame counts in open-source models reduce information density as video length grows; additional modalities can supplement missing information.
5. Discussions
The discussion identifies long-context modeling and complex temporal-reasoning data as major directions for improving MLLM video understanding. It links performance degradation on long videos to context and training-data constraints.
- Improving Long Context Modeling Capabilities of MLLMs: Performance declines as video duration increases, making long-context modeling a significant challenge for MLLMs.Restricted input frames can bottleneck open-source models' ability to represent long videos.
- Improving Long Context Modeling Capabilities of MLLMs: Architectural and infrastructural context extension, adaptive key-frame selection, and video-token compression are proposed to improve long-sequence processing.Examples include ring attention, training-free context extension, and temporal Q-Former architectures.
- Building Datasets with Complex Temporal Understanding: Instruction-tuning datasets for complex temporal reasoning remain limited relative to datasets for text and images.Long-tailed data distributions make acquisition difficult, motivating human-in-the-loop annotation and automatic synthesis.
- Building Datasets with Complex Temporal Understanding: Developing richer temporal-reasoning datasets would provide training supervision for robust video understanding and help exploit architectural advances.The discussion emphasizes that both model design and suitable training data are needed for progress.
6. Conclusion
Video-MME evaluates video understanding with diverse video types, durations, modalities, and expert-annotated questions. Its evaluation highlights the need to improve MLLMs' handling of longer multimodal data.
- Conclusion: Video-MME combines diverse video types, temporal durations, data modalities, and high-quality expert-annotated question-answer pairs.The benchmark is designed specifically to evaluate MLLMs on video understanding tasks.
- Conclusion: The evaluation underscores the need for further advances in handling longer multimodal data.The authors hope the benchmark will inspire research improving MLLM video-understanding capabilities.
7. Detailed Experimental Settings
The detailed settings define the evaluated model groups, frame and subtitle handling, input format, and exact-match accuracy computation. These procedures standardize multimodal video evaluation across models.
- Models: The evaluation covers commercial, open-source video, and image-based MLLMs using their official configurations where applicable.Image models are included to test generalization to multi-image inputs.
- Frame Extraction: Gemini 1.5 Pro samples one frame per second for short and medium videos and one frame every two seconds for long videos.Other models follow their respective official frame-extraction guidelines.
- Subtitle Utilization: Subtitle-enabled evaluation selects subtitles matching the timestamps of sampled frames to synchronize textual and visual inputs.For example, subtitles corresponding to ten sampled frames are selected.
- Evaluation: The evaluation input consists of complete video frames plus optional complete subtitles or audios and a multiple-choice question prompt.Models respond with the letter of the selected answer when the standardized prompt is used.
- Evaluation: Accuracy is computed by extracting the model's answer with regular expressions and directly comparing it with the ground-truth answer.The procedure does not rely on external judges such as ChatGPT.
8. Additional Analysis
Video-MME’s additional analysis tests multimodal inputs and long-range temporal reasoning, revealing both modality-dependent gains and sharp differences in models’ ability to track events across videos.
- Qualitative Evaluation: LLaVA-NeXT-Video and Gemini 1.5 Pro correctly track a target individual’s events across an entire video, unlike several open-source models that associate the person with nearby events.The cases jointly probe OCR, attribute perception, object recognition, and long-range temporal reasoning.
- Qualitative Evaluation: The highlighted cases show that benchmark questions remain challenging because successful video analysis requires both perception and reasoning over temporal context.Models may identify visual or subtitle information yet still fail when reasoning across the video is required.
- Additional Modalities: Subtitles and audio improve Gemini 1.5 Pro’s video understanding across categories, but the magnitude of improvement varies by domain.Figure 4 compares frames alone with frames plus subtitles and frames plus audio across 30 subcategories.