Source-linked AI summary
MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
Yilun Zhao, Lujing Xie, Haowei Zhang, Guo Gan, Yitao Long, Zhiyuan Hu, Tongyan Hu, Weiyuan Chen, Chuhan Li, Junyang Song, Zhijian Xu, Chengye Wang, Weifeng Pan, Ziyao Shangguan, Xiangru Tang, Zhenwen Liang, Yixin Liu, Chen Zhao, Arman Cohan
TL;DR
Expert-level reasoning over specialized-domain videos remains under-evaluated despite foundation models’ broader multimodal capabilities. MMVU addresses this gap with a 3,000-example benchmark built through expert, textbook-guided annotation and evaluates 32 frontier models, finding that o1 performs best while models remain below human expertise.
Problem
Specialized-domain video benchmarks rarely evaluate whether foundation models can combine visual understanding with expert-level knowledge and reasoning.
Method
MMVU contains 3,000 expert-annotated questions over 1,529 videos spanning 27 subjects and four disciplines, constructed through a textbook-guided annotation pipeline.
Results
o1 achieves the highest performance among 32 tested models, but GPT-4o reaches 66.7% versus 86.8% for human experts in the open-book setting.
Takeaways & Limitations
MMVU supports evaluation of specialized-domain video reasoning with expert-validated visual understanding, annotated solutions, and reasoning analyses.
Takeaways & Limitations
A 20% error category reflects models’ difficulty integrating domain knowledge with visual perception in specialized videos.
Abstract
from arXiv · showhide
We introduce MMVU, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. MMVU includes 3,000 expert-annotated questions spanning 27 subjects across four core disciplines: Science, Healthcare, Humanities & Social Sciences, and Engineering. Compared to prior benchmarks, MMVU features three key advancements. First, it challenges models to apply domain-specific knowledge and perform expert-level reasoning to analyze specialized-domain videos, moving beyond the basic visual perception typically assessed in current video benchmarks. Second, each example is annotated by human experts from scratch. We implement strict data quality controls to ensure the high quality of the dataset. Finally, each example is enriched with expert-annotated reasoning rationals and relevant domain knowledge, facilitating in-depth analysis. We conduct an extensive evaluation of 32 frontier multimodal foundation models on MMVU. The latest System-2-capable models, o1 and Gemini 2.0 Flash Thinking, achieve the highest performance among the tested models. However, they still fall short of matching human expertise. Through in-depth error analyses and case studies, we offer actionable insights for future advancements in expert-level, knowledge-intensive video understanding for specialized domains.
1 INTRODUCTION
MMVU addresses the gap in evaluating expert-level reasoning over specialized-domain videos, which convey temporal dynamics, procedural knowledge, and complex interactions. It introduces a 3,000-example benchmark built through expert annotation and evaluates 32 frontier multimodal models, finding persistent distance from human expertise.
- Motivation: Specialized-domain video benchmarks remain limited despite videos conveying temporal dynamics, procedural knowledge, and complex interactions.
- Benchmark: MMVU contains 3,000 expert-annotated QA examples spanning 1,529 videos, 27 subjects, and four disciplines.The disciplines are Science, Healthcare, Humanities & Social Sciences, and Engineering.
- Benchmark: Experts use textbooks to identify concepts, source relevant videos, create questions requiring domain knowledge, and provide reasoning rationales and domain knowledge.The benchmark also applies thorough data quality controls.
- Evaluation: 32 frontier multimodal models were evaluated, with o1 performing best while other models remained noticeably below human-level capability.The evaluation covers models from 17 organizations.
2 RELATED WORK
Prior benchmarks assess specialized expertise mainly through text and images, while video benchmarks largely emphasize general-purpose comprehension. MMVU targets knowledge-intensive expert reasoning in videos, where visual perception and domain expertise are jointly required.
- Video Understanding Benchmark: Existing video benchmarks primarily evaluate general-purpose tasks such as action recognition, captioning, grounding, and temporal reasoning.
- Research Gap: Specialized-domain video understanding requires both visual perception and domain-specific expertise, particularly in healthcare, engineering, and science.
- Multi-discipline Evaluation Benchmark: Earlier multi-discipline benchmarks established expert-reasoning evaluation mainly for textual domains, with newer work extending assessment to other modalities.
3 MMVU BENCHMARK
MMVU is constructed as a broad, expert-level video benchmark using textbook-guided concept selection, vision-intensive licensed videos, expert question and rationale annotation, and multiple quality controls. Its 3,000 examples are split into validation and hidden test subsets, with human performance measured under three resource conditions.
- MMVU Benchmark: MMVU targets breadth, depth, true visual understanding, and fine-grained evaluation through textbook-guided annotation and expert validation.Each example includes expert-annotated solutions and requisite knowledge for analysis.
- Preliminary Setup: A user study with 133 college and graduate students guides subject selection across diverse disciplines.Participants curated QA examples requiring expert-level video understanding in subjects relevant to their fields.
- Preliminary Setup: At least two relevant experts are assigned per subject, with 67 annotators participating in the annotation process.
- Textbook-Guided QA Example Annotation: Annotators identify textbook concepts requiring dynamic visual representation, then select Creative Commons videos that are vision-intensive, exclude audio, and contain minimal on-screen text.
- Textbook-Guided QA Example Annotation: Each selected clip receives two or three expert-level questions with relevant timestamps, plus domain knowledge and a step-by-step reasoning rationale.Questions may be multiple-choice or open-ended.
- Data Quality Control: Human review verifies annotation accuracy, while time-based compensation discourages rushed work during time-intensive curation.Annotating one example averages 20 minutes and 17 seconds, and validation averages 4 minutes and 12 seconds.
- Data Statistics: 3,000 examples are randomly divided into 1,000 validation examples and a hidden 2,000-example test set reserved for standard evaluation.
- Human Performance: Human accuracy averages 49.7% closed-book, 86.8% open-book, and 95.3% after oracle revision.
4 EXPERIMENTS
MMVU evaluates frontier multimodal foundation models on expert-level video understanding using CoT prompts, direct answering, and human error analysis. Results show strong variation across models and persistent gaps from human expertise, while System-2-capable models and CoT often improve performance.
- Main Findings: GPT-4o achieves 66.7% accuracy with CoT prompting, compared with 86.8% for human experts in the open-book setting.The result demonstrates a substantial model–human performance gap on MMVU.
- Main Findings: Open-source models generally lag behind proprietary models, although Qwen2-VL-72B and DeepSeek-VL2 exceed human benchmarks in closed-book settings.These models approach the performance of leading proprietary models.
- Main Findings: CoT reasoning generally improves performance over direct answering, but Claude 3.5 Sonnet gains 11.0% while GPT-4o shows only marginal improvement.The effect of CoT varies across foundation models on MMVU.
- Main Findings: o1 and Gemini 2.0 Flash Thinking achieve the top two MMVU results, indicating advantages for long CoT and increased test-time compute.The finding highlights the effectiveness of System-2 thinking for expert-level video reasoning.
- Qualitative Analysis: Human error analysis identifies visual perception errors at 18%, domain-knowledge errors in visual perception at 20%, and domain-knowledge errors in reasoning at 27%.Other reported categories include textual-information reliance at 20%, logical reasoning errors at 6%, and other errors at 9%.
5 CONCLUSION
MMVU assesses expert-level, knowledge-intensive reasoning on specialized-domain videos through a high-quality, multidisciplinary benchmark and evaluates frontier multimodal models against human expertise. The study finds that o1 approaches human expert proficiency, while a notable gap remains for other models.
- 5 CONCLUSION: MMVU evaluates expert-level, knowledge-intensive reasoning on specialized-domain videos using a multidisciplinary benchmark.The benchmark is constructed through a textbook-guided annotation pipeline intended to capture breadth of knowledge and depth of reasoning.
- 5 CONCLUSION: Each MMVU example is annotated from scratch by human experts and subjected to expert review for annotation accuracy.The construction process includes validation by authors or top-performing annotators against benchmark guidelines.
- 5 CONCLUSION: 32 frontier multimodal foundation models were evaluated, with o1 achieving the highest performance and approaching human expert-level proficiency.The conclusion also reports a remaining performance gap between other evaluated models and human experts.
- 5 CONCLUSION: MMVU’s construction draws on textbooks spanning Science, Engineering, Healthcare, and Humanities and Social Science disciplines.The cited tables list textbooks and corresponding example numbers for each discipline.
- 5 CONCLUSION: The benchmark uses a dedicated annotation and validation interface to support quality-controlled example construction.The interface covers video collection, multiple-choice and open-ended annotation, and validation, including revision or discarding of low-quality examples.
A.5 DATA ANNOTATION AND VALIDATION PAYMENT
MMVU’s annotation and validation process lasted three months and compensated annotators according to time spent rather than examples completed.
- A.5 DATA ANNOTATION AND VALIDATION PAYMENT: The three-month annotation and validation process paid annotators based on time spent rather than completed examples.This payment scheme was intended to prevent rushing, particularly because suitable Creative Commons videos were limited in some subjects.
B EXPERIMENT SETUP
The experiments evaluate multimodal models under standardized inference settings using CoT reasoning and direct-answer prompts for multiple-choice and open-ended questions. Evaluation prompts assess final-answer accuracy, including semantic equivalence for open-ended responses.
- B EXPERIMENT SETUP: All experiments use temperature 1.0 and a maximum output length of 1024 tokens, except Gemini-2-Flash-Thinking, which uses 8192 tokens.The longer limit accommodates Gemini-2-Flash-Thinking’s long CoT reasoning mechanism.
- B EXPERIMENT SETUP: The study compares CoT reasoning with direct answering for both multiple-choice and open-ended questions.The corresponding prompts are shown for each question type, and Figure 16 compares their validation-set performance.
- B EXPERIMENT SETUP: Multiple-choice and open-ended responses are evaluated by extracting the final answer and comparing it with the ground truth.For open-ended QA, an answer is correct when it demonstrates the same technique or meaning, rather than matching word-for-word.
C.2 ERROR CASE ANALYSIS: VISUAL PERCEPTION ERROR
This section presents visual examples associated with thermodynamics, electromagnetism, and art error cases, including a multiple-choice thermodynamics question.
- C.2 ERROR CASE ANALYSIS: VISUAL PERCEPTION ERROR: The thermodynamics example asks which process is shown in an animation, with five compression or expansion options.
- C.2 ERROR CASE ANALYSIS: VISUAL PERCEPTION ERROR: Figure 17 presents an error case involving Thermodynamics.
- C.2 ERROR CASE ANALYSIS: VISUAL PERCEPTION ERROR: Figure 18 presents an error case involving Electromagnetism.
- C.2 ERROR CASE ANALYSIS: VISUAL PERCEPTION ERROR: Figure 19 presents an error case involving Art.
C.3 ERROR CASE ANALYSIS: MISUSE OR LACK DOMAIN KNOWLEDGE IN VISUAL PERCEPTION
This section presents error cases from Computer Science, Electrical Engineering, and Pharmacy.
- Figure 20 presents an error case from Computer Science.
- Figure 21 presents an error case from Electrical Engineering.
- Figure 22 presents an error case from Pharmacy.
C.4 ERROR CASE ANALYSIS: MISUSE OR LACK DOMAIN KNOWLEDGE IN REASONING
This section presents reasoning-related error cases alongside questions spanning algorithms, biology, chemistry, and related domains.
- The examples include questions identifying a demonstrated common algorithm and a biological process.
- Figure 20 presents an error case from Computer Science.
- Figure 24 presents an error case from Biology.
- The chemistry example asks for the approximate precipitate mass produced by a reaction involving 2.24 liters of gas at standard temperature and pressure.
- Figure 25 presents an error case from Chemistry.
C.5 ERROR CASE ANALYSIS: HEAVY RELIANCE ON TEXTUAL INFORMATION
This section presents error cases from Healthcare, Humanities & Social Sciences, and Engineering, including questions about physiological adjustment, risk handling, and mechanical transformation.
- The healthcare example asks which physiological system or activity is most important to adjust during a phenomenon.
- Figure 27 presents an error case from Clinical Medicine.
- The management example asks for a method to handle risks associated with a company shown in the video.
- Figure 28 presents an error case from Management.
- The mechanical-engineering example asks which mechanical transformation is not represented in the video.