Source-linked AI summary
MLVU: Benchmarking Multi-task Long Video Understanding
Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Zhengyang Liang, Shitao Xiao, Minghao Qin, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, Zheng Liu
TL;DR
Existing benchmarks inadequately evaluate long-video understanding because they use short videos, limited genres and tasks, and designs not tailored to long-video information. MLVU addresses these gaps with flexible video lengths, diverse genres, and diversified LVU-oriented tasks, while evaluation of 23 MLLMs shows that long-video understanding remains challenging, especially for fine-grained tasks.
Problem
Existing video benchmarks largely use short videos, limited genres or tasks, and evaluation designs that inadequately reflect long-video understanding.
Method
MLVU benchmarks long-video understanding using videos from 3 minutes to 2 hours, diverse genres, and holistic, single-detail, and multi-detail evaluation tasks.
Results
Evaluation of 23 MLLMs shows that existing methods struggle with most tasks and fine-grained information from entire videos; GPT-4o leads with an M-Avg of 54.5% in multiple-choice tasks.
Takeaways & Limitations
MLVU provides a comprehensive analysis of MLLMs’ long-video understanding and indicates that context length, image understanding ability, and LLM backbones are influential factors for future advances.
Abstract
from arXiv · showhide
The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in video types and evaluation tasks, and the inappropriateness for evaluating LVU performances. To address the above problems, we propose a new benchmark called MLVU (Multi-task Long Video Understanding Benchmark) for the comprehensive and in-depth evaluation of LVU. MLVU presents the following critical values: \textit{1)} The substantial and flexible extension of video lengths, which enables the benchmark to evaluate LVU performance across a wide range of durations. \textit{2)} The inclusion of various video genres, e.g., movies, surveillance footage, egocentric videos, cartoons, game videos, etc., which reflects the models' LVU performances in different scenarios. \textit{3)} The development of diversified evaluation tasks, which enables a comprehensive examination of MLLMs' key abilities in long-video understanding. The empirical study with 23 latest MLLMs reveals significant room for improvement in today's technique, as all existing methods struggle with most of the evaluation tasks and exhibit severe performance degradation when handling longer videos. Additionally, it suggests that factors such as context length, image-understanding ability, and the choice of LLM backbone can play critical roles in future advancements. We anticipate that MLVU will advance the research of long video understanding by providing a comprehensive and in-depth analysis of MLLMs.
1. Introduction
MLVU addresses shortcomings in long-video understanding evaluation by extending video durations, broadening video and task diversity, and tailoring tasks to long-video information. Evaluation of 23 MLLMs shows persistent difficulty, especially on fine-grained whole-video tasks.
- Existing benchmarks often use videos lasting only a few seconds, limiting their ability to reflect long-video understanding.
- Prior benchmarks lack diversity in video genres and evaluation tasks, often focusing on one video type or task.
- MLVU spans videos from 3 minutes to 2 hours, averages about 15 minutes, and supports evaluation across incrementally longer clips.
- MLVU includes varied genres and nine evaluation tasks covering multiple task formats and global or local video information.
- 23 MLLMs were evaluated, and all methods struggled with fine-grained tasks requiring information from entire videos, including counting, ordering, and summarization.
2. Related Work
Related benchmarks establish broad multimodal and video-understanding evaluation, but long-video benchmarks remain fragmented across video types, task dimensions, or context-length strategies. MLVU is motivated by the need for comprehensive evaluation of long-video capabilities.
- MLLMs combine LLM backbones with visual encoders and adapters, with video-oriented systems extending this framework through video instruction data and specialized adapters.
- Long-video approaches use memory components, extended context lengths, or selective frame and clip retrieval to process longer inputs.
- The related-work landscape motivates a benchmark combining diverse long videos with carefully designed tasks for comprehensive LVU evaluation.
- Existing video benchmarks cover specialized areas such as temporal perception, action understanding, classification, reasoning, and captioning, particularly for short videos.
- Specialized benchmarks such as EgoSchema focus on a single aspect or video source rather than comprehensive long-video understanding.
3. MLVU: Multi-task Long Video Understanding Benchmark
MLVU is a 3,102-question benchmark spanning diverse long videos, multiple duration levels, and nine LVU-oriented tasks. Its tasks assess holistic, single-detail, and multi-detail understanding through global, critical-plot, or multiple-plot information.
- 3.1. Overview: MLVU contains 3,102 questions across nine categories, divided into 2,593 development and 509 test questions.
- Diversified Video Categories: The benchmark includes movies, documentaries, television series, egocentric videos, life records, sports, tutorials, surveillance footage, animated series, and game videos.
- Substantial Extension of Video Length: Videos range from 3 minutes to more than 2 hours and are partitioned into incremental segments for evaluation at different lengths.
- Diversified Evaluation Tasks: Nine tasks align with reasoning, captioning, recognition, perception, and summarization while requiring in-depth understanding of long videos.
- Task Categories: Holistic, single-detail, and multi-detail LVU tasks require global video information, one critical plot, or multiple plots, respectively.
- Holistic LVU: Video Summarization uses 257 narrative-rich videos with manually annotated summaries evaluated by GPT-4 against the annotations.
- Single-Detail LVU: NQA inserts a short needle clip into a long background video and tests whether the model can locate and answer questions about that clip.
4. Experiments and Analysis
Experiments on 23 MLLMs show that GPT-4o leads overall, but current systems remain weak on fine-grained, multi-detail, and longer-video understanding. Performance varies by task, video length, clip position, probe count, and model design factors.
- Main Results: 23 MLLMs were evaluated across holistic, single-detail, and multi-detail LVU tasks using the MLVU benchmark.Table 2 reports individual task results alongside M-Avg and G-Avg aggregates.
- Main Results: GPT-4o leads the benchmark with an M-Avg of 54.5% and a G-Avg of 5.87.These averages cover multiple-choice and generation tasks, respectively.
- Main Results: 42.9% is GPT-4o’s accuracy on NQA, while existing methods also struggle with ego reasoning, action ordering, and action counting.The results indicate that long-video understanding remains difficult even for the strongest evaluated model.
- Main Results: Holistic tasks such as topic reasoning and anomaly recognition differentiate models more strongly than other tasks.GPT-4o and several leading open-source models perform accurately, whereas many other MLLMs fail to produce meaningful results.
- Main Results: Multi-detail tasks show catastrophic degradation relative to single-detail tasks, with most methods failing action ordering, action counting, and summarization.GPT-4o and Video-XL are exceptions on action ordering and counting.
- Further Analysis: Performance declines as videos grow from 180s to 600s, and Video-LLaMA-2 approaches random performance at 10 minutes.Figure 3 measures average accuracy across NQA, ER, PQA, AC, and AO.
- Further Analysis: Long-video models such as LongVA and Video-XL remain robust across referring-clip positions, unlike short-video models.The analysis covers ego reasoning and plot question-answering across four clip-position intervals.
- Further Analysis: Action-count performance significantly declines as the number of probes increases across all models.The probes correspond to the multiple details that must be comprehended and processed simultaneously.
5. Conclusion
MLVU is introduced as a benchmark designed to assess long-video understanding comprehensively through longer videos, diverse genres, and diversified evaluation tasks. Experiments show that LVU remains technically challenging for state-of-the-art MLLMs, motivating joint attention to several influencing factors.
- 5. Conclusion: MLVU extends video lengths, includes varied video genres, and develops diversified LVU-oriented evaluation tasks.These innovations aim to support comprehensive and in-depth analysis of MLLMs’ long-video understanding performance.
- 5. Conclusion: LVU remains a technically challenging problem for today’s state-of-the-art MLLMs.The conclusion identifies this as the central empirical finding of the MLVU study.
- 5. Conclusion: Future advancements may require joint optimization of context length, image understanding ability, and LLM backbones.The paper anticipates that MLVU will facilitate further research in long-video understanding.
Supplementary Material
The supplementary material provides additional evaluation results, dataset-construction details, duration-ladder specifications, annotation information, baseline procedures, retrieval experiments, and visualized examples.
- Supplementary Material: The supplementary material includes evaluation results on the MLVU development set.
- Supplementary Material: It documents the Universal Long Video Collection and the MLVU Time-Ladder.
- Supplementary Material: It also covers annotations, baselines, evaluation procedures, retrieval-augmented generation, and visualized MLVU examples.
B. Evaluation Results on MLVU Dev Set
The MLVU dev-set materials describe how long videos from a diverse collection are transformed into benchmark videos and questions. They also distinguish the dev and test evaluation formats and explain how one source video can support multiple tasks and clip granularities.
- B. Evaluation Results on MLVU Dev Set: The MLVU dev set uses four-option multiple-choice questions, whereas the test set uses six options, making the test set more challenging and discriminative.
- B. Evaluation Results on MLVU Dev Set: The Universal Long Video Collection contains 986 long videos spanning movies, documentaries, games, surveillance, egocentric footage, cartoons, television, tutorials, sports, and life records.The collection includes 168 movies, 60 documentaries, 65 game videos, 239 surveillance videos, 100 egocentric videos, 72 cartoons, 92 TV series, 60 tutorials, 60 sports videos, and 70 life records.
D. Details of the MLVU Time-Ladder
MLVU’s Time-Ladder and annotation procedures vary video duration and task-specific segment references to study long-video difficulty. The materials also describe the task distribution, task categories, and annotation sources across the benchmark.
- D. Details of the MLVU Time-Ladder: The MLVU Time-Ladder contains videos of 3, 6, and 10 minutes to investigate how duration affects LVU task difficulty.Segment-level annotation enables video-length adjustment without requiring additional human annotators.
- D. Details of the MLVU Time-Ladder: Video summarization annotations follow the initial 3- and 6-minute segments, while PQA and SSC annotations identify answer-containing segments.Ego reasoning uses answer intervals already present in Ego4D, and synthetic tasks include NQA, AO, and AC.
- D. Details of the MLVU Time-Ladder: MLVU contains 3,102 questions divided into 2,593 development questions and 509 test questions.Table 1 provides the detailed distribution across tasks.
- D. Details of the MLVU Time-Ladder: The benchmark organizes tasks into holistic, single-detail, and multi-detail LVU categories.Table 2 lists TR, AR, and VS as holistic tasks; NQA, ER, PQA, and SSC as single-detail tasks; and AO and AC as multi-detail tasks.
F.3. Video Summarization (VS).
The Video Summarization task uses long clips and requires chronological summaries that avoid character names while preserving identifiable roles or attributes. The broader MLVU annotation process combines manual, semi-automated, and synthetic data construction across complementary tasks.
- Video Summarization: VS annotations use pronouns and distinctive attributes because most existing MLLMs cannot reliably process audio or subtitles or identify specific characters.This guideline is intended to reduce dependence on unavailable multimodal signals.
- Needle Question-Answering: NQA uses GPT-4 and WebVid captions to generate question-answer pairs about randomly inserted needle clips within longer background videos.The needle clips are selected from WebVid, and their captions are supplied to GPT-4.
- Plot Question-Answering: PQA questions target intricate plot details and combine perception with reasoning while avoiding character names and objective hints.Questions should identify unique events rather than vague actions that could occur repeatedly.
- Video Summarization: VS asks annotators to summarize key events in chronological order for video clips lasting 3 to 15 minutes.Annotations avoid specific character names and instead use pronouns plus attributes or roles.
- Action Tasks: Action Order and Action Count use synthetically generated videos, questions, and answers, with probe-selection and review strategies intended to improve evaluation-data validity.Action Order uses uncommon actions and reviewed background videos to reduce accidental matches.
G. Details of Baselines and the Evaluation Process
The baseline suite covers image- and video-based MLLMs, including proprietary systems with multi-image APIs. Evaluation templates differ by task type to support option extraction for multiple choice without adding interventions to generation tasks.
- Baselines: The baselines include Otter-I, LLaVA-1.6, InternVL, Claude-3-Opus, and Qwen-VL-Max for multi-image inference.Available models are evaluated with maximum input frames determined by LLM context length.
- Baselines: The baseline selection accounts for the fact that most image-based MLLMs lack multi-image inference capabilities.The selected image models have official multi-image implementations, while Claude-3-Opus and Qwen-VL-Max provide APIs.
- Evaluation Process: MLVU uses separate templates for Multiple-Choice and Generation tasks, with option prediction guidance added only to the Multiple-Choice template.Generation tasks instead use fixed-question guidance without additional interventions.
G.3. Evaluation Metrics
MLVU evaluates multiple-choice answers by exact option matching and generation outputs with GPT-4-based criteria tailored to captioning and summarization. Its prompts target plot events and specific, non-identifying details.
- Evaluation Metrics: Multiple-Choice performance is measured by absolute accuracy through matching the predicted option with the ground truth.
- Evaluation Metrics: GPT-4 ranks generated answers against provided answers for generation-task evaluation.
- Evaluation Metrics: Sub-Scene Captioning is evaluated with Accuracy and Relevance, while Video Summary is evaluated with Completeness and Reliability.
- Plot Question-Answering: Plot-question prompts require specific factual or inferential details, avoid character names, and identify unique scenes within the long video.
- Plot Question-Answering: The PQA word-cloud comparison contrasts TVQA’s character-specific questions and MovieChat’s less diverse questions with PQA’s intended diversity and usability.
H. Explorations of Video Retrieval Augmented Generation
The zero-shot video RAG strategy retrieves clips relevant to a question before inference, improving detail-oriented reasoning but offering limited support for global understanding and incurring substantial overhead.
- Findings: RAG benefits detail-oriented single-detail reasoning because retrieval focuses models on clips containing answer-related cues.
- Findings: RAG has limited capability on multi-detail reasoning and holistic understanding tasks requiring global perception and knowledge aggregation.
- RAG Pipeline: RAG divides each long video into N clips of C frames, extracts clip embeddings with LanguageBind, and retrieves the top K clips using text-video similarity.The selected clips are concatenated for question answering and configured around 16-second intervals for models limited to 16 frames.
- Limitations: The retrieval process takes more than one minute, motivating more efficient approaches for long-video understanding.