Source-linked AI summary
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon
TL;DR
Existing benchmarks mostly test perceptual recognition, not whether MLLMs can infer artistic intent. MuseBench evaluates intent-level audiovisual arts understanding with expert-derived questions, finding that the best model reaches 48.29% accuracy versus 87.18% for human experts.
Problem
Existing benchmarks largely test what happens in scenes, leaving whether MLLMs can infer the intent and artistic significance behind creative decisions underexplored.
Method
MuseBench uses video essays and a four-phase human-in-the-loop pipeline to construct 4,016 expert-validated questions across four audiovisual art categories.
Results
48.29% accuracy is achieved by the best of 28 zero-shot MLLMs, versus 87.18% for human experts, with models consistently weakest on game arts.
Takeaways & Limitations
Audiovisual-arts reasoning remains far from saturated, motivating richer artistic and cultural supervision beyond generic video understanding.
Takeaways & Limitations
The benchmark covers four art categories and relies primarily on video essays, which may not capture the full diversity of artistic expression across forms and languages.
Abstract
from arXiv · showhide
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing, or grief conveyed through silence and lingering close-ups). True artistic understanding extends beyond recognizing what is depicted to reasoning about why it is expressed through particular creative choices. Despite the strong progress of multimodal large language models (MLLMs), this critical aspect of artistic understanding remains underexplored, as existing benchmarks largely measure perceptual recognition while overlooking reasoning about creative intent. To address this gap, we introduce Musebench, a comprehensive benchmark designed to evaluate MLLMs on nuanced artistic understanding. It comprises 4,016 questions spanning cinematic arts, static visual arts, stage performing arts, and game arts, distilled from over 10K candidate video essays that pair professional commentary with visual demonstration. To capture the open-ended nature of artistic analysis at scale, the benchmark combines single-select and variable-option multi-select questions. All questions are generated and refined through a four-phase iterative pipeline combining shortcut filtering, adversarial distractors, and expert validation. Comprehensive zero-shot evaluation of 28 state-of-the-art MLLMs reveals that even the best-performing model achieves only 48.29% accuracy, substantially below human expert performance of 87.18%, exposing a significant gap in current models' creative domain expertise.
1 Introduction
MUSEBENCH evaluates whether MLLMs can infer artistic intent from audiovisual form rather than merely recognize depicted content. It combines a scalable expert-knowledge pipeline with heterogeneous evaluation, revealing a substantial gap between models and human experts.
- Motivation: Audiovisual-art understanding requires interpreting why creators use particular visual, auditory, and narrative choices, not merely recognizing what appears on screen.Examples include linking symmetric framing and warm lighting, or prolonged silence, to emotional intent.
- Motivation: Existing video benchmarks primarily test scene content with a single correct option, leaving creative-intent inference underexplored.The paper contrasts factual recognition with reasoning about the intent behind creative decisions.
- Data Construction: The four-phase pipeline transforms over 10,000 video essays into visually grounded questions while using adversarial distractors to reduce transcript-based shortcuts.Distractors use technical misread, over-simplification, factual error, and conceptual confusion strategies.
- Evaluation: 48.29% accuracy is achieved by the best model versus 87.18% for human experts in zero-shot evaluation of 28 state-of-the-art MLLMs.The analysis additionally finds persistent weakness on game arts, recovery of only the most salient multi-select option, and little benefit from adaptive key-frame selection.
- Benchmark Design: MUSEBENCH covers four art categories and 11 sub-domains, combining single-select and variable-option multi-select questions to capture interpretive plurality.Its questions target audiovisual arts expertise across cinema, visual arts, stage performance, and interactive media.
2 Related Work
Prior work advances video understanding through efficient long-video processing, agentic retrieval, temporal reasoning, multimodal breadth, and domain knowledge. MuseBench is positioned within audiovisual perception research while targeting the interpretive expertise required by audiovisual arts.
- Multimodal Large Language Models for Video Understanding: MLLM video-understanding research addresses efficient processing through sparse token memory, visual summarization tokens, and native sparse attention for long contexts.Another line frames video understanding as agentic retrieval using tree search and interleaved reasoning with temporal grounding.
- Benchmarks for Video Understanding: Video-understanding benchmarks have progressed from short-clip QA to story-level and temporal-reasoning frameworks.Recent benchmarks also expand multimodal breadth, long-video scale, and domain knowledge involving expert lectures and STEM reasoning.
- Benchmarks for Video Understanding: Audiovisual perception benchmarks include AV-Odyssey Bench, which probes fine-grained contrasts such as pitch and loudness.MuseBench extends this inquiry toward audiovisual arts expertise and interpretive reasoning involving cinematographic technique, compositional principles, and performance craft.
3 MuseBench Construction
MUSEBENCH is constructed from expert-narrated video essays using a four-category taxonomy, complementary question formats, and an iterative generation and validation pipeline. The resulting benchmark contains 4,016 expert-validated questions distilled from over 10,000 candidate video essays.
- 3.1 Video Essays: Video essays provide expert commentary aligned with visual or auditory evidence for analyzing not only what artistic techniques are used, but why they produce particular effects.The format combines expert-narration density with narration-to-evidence alignment.
- 3.2 Taxonomy: The capability taxonomy spans four art categories—Cinematic Arts, Static Visual Arts, Stage Performing Arts, and Game Arts—and 11 sub-domains.The taxonomy guides data collection and reporting for comprehensive coverage across audiovisual arts.
- 3.3 Construction Pipeline: The construction pipeline partitions videos into 10-second intervals, captions each segment, and generates candidate question-answer pairs through four successive phases.Retained videos are separated into narrator transcripts for question construction and narrator-removed 10-second audiovisual clips for model evaluation.
- 3.2 Question Formats: MUSEBENCH combines single-select questions with 4–8 options and one correct answer, and multi-select questions containing 2–4 correct answers to capture set-valued analytical judgment.Single-select questions probe discrete recognition, whereas multi-select questions represent multiple valid perspectives.
- 3.3 Quality Control: Around 9% of generated question-answer pairs were flagged as incorrect across eight failure modes, while every retained item was manually verified before assessment.The iterative review loop updates prompts with exclusion rules and domain-specific constraints.
- 3.4 Benchmark Statistics: 4,016 expert-validated questions were distilled from over 10,000 candidate video essays across four art categories and 11 sub-domains.Each retained video contributes 3–5 questions, and each question offers 4–8 options.
4 Experiments
Zero-shot evaluation of 28 MLLMs shows that audiovisual-arts reasoning remains substantially unsaturated, with category-specific weaknesses and no decisive advantage from video specialization or adaptive frame selection. Complementary analyses expose multi-select underprediction, modality effects, and a strong first-position bias in open-source models.
- Evaluation setup: 28 MLLMs are evaluated zero-shot across proprietary, open-source general-purpose, open-source video-specific, and dynamic key-frame-selection tiers.Inputs use narrator-removed evidence clips sampled at 1 fps, or each model’s maximum supported rate.
- Leaderboard findings: No MLLM approaches saturation; proprietary systems lead but remain below experts, while video-specialized models offer no decisive advantage.The reported bottleneck is artistic vocabulary, cultural priors, and grounded inference rather than formatting, evaluation noise, or temporal localization.
- Category competence: Game arts are a shared weakness, while models show divergent competence profiles across cinematic, static visual, and stage performing arts.Models competitive on the other categories drop markedly on game footage across tiers and formats.
- Key-frame selection: 14.42–20.51 ACC is the range for all five adaptive key-frame models, which trail the strongest video-specific models by 7–13 points.Adaptive selection does not unlock further headroom because the bottleneck lies in artistic vocabulary and cultural priors rather than locating salient frames.
- Multi-select behavior: On multi-select questions, F1 exceeds EM and precision exceeds recall for most models, indicating underprediction of valid alternatives rather than distractor over-selection.The precision–recall gap widens for mid-tier systems.
5 Conclusion
MUSEBENCH is introduced as a comprehensive benchmark for audiovisual arts understanding, built from expert knowledge in video essays and validated through iterative human review. It evaluates 28 MLLMs using 4,016 questions spanning four art categories.
- MUSEBENCH contains 4,016 expert-validated questions spanning cinematic, static visual, stage performing, and game arts.The benchmark uses video essays as a scalable source of expert knowledge.
- The benchmark is constructed through a four-phase pipeline with iterative human-in-the-loop quality review.
- MUSEBENCH supports comprehensive zero-shot evaluation of 28 multimodal large language models.
Appendices … G Additional Examples of MUSEBENCH
The appendices organize benchmark details, construction procedures, and detailed evaluation metrics for MUSEBENCH. They cover benchmark definitions and statistics, a multi-phase construction pipeline with quality review and human evaluation, and separate single-select and multi-select metrics.
- B Benchmark Details: Benchmark details cover comparisons with existing benchmarks, category and sub-domain definitions, dataset statistics, and a vocabulary overview.These materials appear in Sections B.1–B.4 of the benchmark details appendix.
- Appendices: The appendix structure places benchmark details before construction details and detailed evaluation metrics.Sections B, C, and D are listed consecutively in the appendix contents.
- C Construction Details: Construction details span source curation, transcription, clip segmentation, clip captioning, question generation, and distractor generation.These procedures are listed as Sections C.1–C.6.
- C Construction Details: The construction appendix also documents a quality review loop, failure taxonomy, and human evaluation details.These topics are covered in Sections C.7–C.9.
- D Detailed Evaluation Metrics: Single-select evaluation uses chance-adjusted accuracy as its detailed metric.This metric is specified in Section D.1.
- D Detailed Evaluation Metrics: Multi-select evaluation reports precision, recall, and F1.These metrics are specified in Section D.2.
A Limitation and Future Work · B Benchmark Details
MuseBench currently covers four art categories and primarily uses video essays, whose expert commentary may not fully represent artistic diversity across forms and languages. Future work will broaden the benchmark’s coverage, sources, and evaluation format.
- A Limitation and Future Work: MuseBench currently focuses on four art categories and uses video essays as its primary data source.Video essays provide rich expert commentary but remain the benchmark’s principal source format.
- A Limitation and Future Work: Video essays may not fully capture the diversity of raw creative works, and their availability varies across art forms and languages.These limitations affect both representational breadth and source coverage.
- A Limitation and Future Work: Future work will expand to additional art forms, incorporate more diverse multilingual sources, and develop open-ended evaluation beyond multiple-choice.Examples of additional art forms include music and architecture.
B.1 Comparison with Existing Benchmarks · B.2 Category and Sub-domain Definitions
MUSEBENCH distinguishes itself from existing video benchmarks by jointly requiring open-domain audiovisual evidence, domain expertise, subtitles and audio, and visual dependency. Its 4,016 questions span four art categories and 11 sub-domains covering cinematic, static visual, stage performing, and game arts.
- B.1 Comparison with Existing Benchmarks: Existing benchmarks largely use short clips of everyday activities or open-domain factual QA without coupling subtitles and audio, domain expertise, or controlled visual dependency.Video-MMMU adds multi-level difficulty to lecture videos but remains an expert-oriented effort with a different target scope.
- B.1 Comparison with Existing Benchmarks: MUSEBENCH is the only compared benchmark satisfying Open Domain, Sub.&Aud., Domain Expert, and Visual Dep. simultaneously.Its hierarchical taxonomy supplies domain expertise, while narrator-removed clips and timestamped expert commentary enforce visual dependency and audio use.
- B.2 Category and Sub-domain Definitions: 4,016 questions are organized into four top-level art categories and 11 fine-grained sub-domains grounded in creative disciplines and expert video-essay topics.The taxonomy is hierarchical and designed to represent topics actively analyzed by professional commentators.
- B.2 Category and Sub-domain Definitions: Cinematic Arts examines how filmmakers translate dramatic intent into image and sound through Cinematography, Mise-en-scène, Editing and Pacing, and Sound.These sub-domains cover visual construction, on-screen staging, shot-to-shot rhythm, and diegetic or non-diegetic audio.
- B.2 Category and Sub-domain Definitions: Static Visual Arts replaces temporal structure with compositional and material reasoning across Fine Art and Photography.The category covers tangible artifacts and still images, including technique, composition, exposure, framing, and documentary intent.
- B.2 Category and Sub-domain Definitions: Stage Performing Arts analyzes live meaning through performer presence, staged space, and time-bound audience reception.Its sub-domains are Performance, Stage Design, and Theatrical Lighting, covering delivery, spatial organization, and attention or mood shaping.
- B.2 Category and Sub-domain Definitions: Game Arts addresses audiovisual craft shaped by interactive systems and player agency through CG and Interactive Visuals.These sub-domains cover rendered game aesthetics alongside player-conditioned visual elements such as level design, HUDs, navigation cues, camera, and animation.
B.3 Dataset Statistics … C.6 Phase D: Distractor Generation
MUSEBENCH is characterized through global, source, and question-format statistics, with vocabulary centered on artistic analysis. Its construction pipeline curates expert video essays, transcribes and segments them, captions clips, generates questions, and adds distractors through iterative design.
- B.3 Dataset Statistics: Table 4 consolidates MUSEBENCH’s global scope, source data, and question-format statistics.The supplied passage identifies Table 4 as covering global properties and construction parameters.
- B.4 Benchmark Vocabulary Overview: Dominant vocabulary terms include emotional, analysis, composition, color, spatial, narrative, character, and audience.Vocabulary frequencies span question text, answer options, and core intents after removing function words and generic analytical fillers.
- C Construction Details: The appendix traces construction from source curation through transcription, category inference, 10 s clip segmentation, captioning, question generation, and distractor synthesis.These per-video phases expand the main text’s construction pipeline and human-in-the-loop quality review.
- C.1 Source Curation: Source curation uses four LLM-controlled stages to discover long-form video essays where experts analyze artistic artifacts while showing their visual content.The stages include category-aware keyword generation, public-web search, relevance judgment, and final human vetting; variant expansion is used when keywords exhaust results.
- C.2 Transcription: Whisper-Large-v3 transcribes each retained video into a JSON record whose transcript, timestamped segments, and metadata support downstream processing.Timestamped segments feed clip-level captioning, the full transcript feeds question generation, and metadata supports segmentation constraints.
- C.3 Phase A: Clip Segmentation: 10 s non-overlapping clips provide uniform temporal granularity, while a 1,800 s maximum duration bounds each source’s clips and QA pairs.Stable clip indices are reused by captioning, clip matching, and each question’s relevant_clips field.
- C.4 Phase B: Clip Captioning: Keye-VL-1.5 captions each 10 s clip at one frame per second using aligned narration, describing color, composition, motion, and scene context.These captions are construction artifacts for question generation and review and are not exposed to evaluated models.
- C.5 Phase C: Question Generation; C.6 Phase D: Distractor Generation: The QA generator creates 3 to 5 candidate pairs per video, with about 30% multi_select and the remainder single_select, before Phase D adds 3 to 7 plausible distractors.Multi-select questions contain 2 to 4 independent correct answer points; distractors use seven runtime strategies, including scope error, temporal confusion, and partial truth.
C.7 Quality Review Loop · C.8 Failure Taxonomy
MuseBench uses category-specific, expert-reviewed prompt revisions to detect and retire failures in generated QA pairs. An audit of 5,000 early pairs identified eight tagged failures across four severity tiers and seven systemic issues requiring prompt-level refinement, with none of the tagged failures remaining in the released benchmark.
- C.7 Quality Review Loop: The review loop independently generates pilot QA batches for each art category, has domain experts tag failures, consolidates new types, and rewrites prompts with added rules.Rules include hard content constraints and schema-level requirements.
- C.7 Quality Review Loop: Four representative failure dimensions are demonstrated through bad cases, added prompt rules, and regenerated versions of the same items.The dimensions are narrator-dependent answerability, ambiguous stems, weak or incorrect distractors, and misaligned clip references.
- C.7 Quality Review Loop: Narrator-dependent questions were rejected because answers relying on transcripts cannot be recovered from the evaluated visual or audible evidence.The added R3 rule requires every stem to be answerable from evidence inside the clip.
- C.7 Quality Review Loop: Ambiguous stems were addressed by requiring concrete visible or audible elements and rejecting under-specified prompts with multiple equally valid answers.Rules R2 and F4 target grounding in audiovisual evidence and specificity.
- C.7 Quality Review Loop: Weak distractors were addressed by requiring each option to use a distinct strategy from a seven-strategy taxonomy rather than lexical paraphrases.The motivating example contained four near-equivalent descriptions of editing rhythm.
- C.7 Quality Review Loop: In Game Arts, three prompt revisions moved from abstract stems without committed keys to named-game decision focus and finally visible-element anchoring without proper nouns.The final revision was designed to prevent title-recognition shortcuts and force use of visual evidence.
- C.8 Failure Taxonomy: 5,000 early benchmark QA pairs were audited round by round, surfacing eight failure tags across four severity tiers, each paired with a prompt-level rule.CRITICAL and HIGH issues were eliminated by replacing affected items, while MEDIUM and LOW issues used strengthened generation, distractor prompts, and programmatic post-hoc alignment.
- C.8 Failure Taxonomy: Seven systemic issues required full prompt-level refinement, including low option discriminability, academic language, imprecise intent targeting, indistinct distractors, text-vision misclassification, inconsistent visual-evidence enforcement, and embedded option problems.These issues resisted remediation through item swapping alone.
C.9 Human Evaluation Details · D Detailed Evaluation Metrics
Human evaluation used a self-contained interface and four domain experts to rate benchmark items across four 0–5 Likert dimensions. The rubric assessed holistic quality, visual necessity, mechanistic technique–effect–intent reasoning, and answer integrity, with category-specific calibration before scoring.
- C.9 Human Evaluation Details: A self-contained web interface let domain experts play clips, read questions and options, and rate each item on four Likert dimensions.The interface supported the human evaluation reported in Section 3.4.
- C.9 Human Evaluation Details: Four domain experts were assigned by formal training to the four arts categories, with two independent raters per category whose scores were averaged.Two experts trained in Drama, Film and Literature rated Stage Performing Arts and Cinematic Arts; two trained in Computational Media and Arts rated Game Arts and Static Visual Arts.
- C.9 Human Evaluation Details: 0–5 Likert scores measured holistic quality, visual necessity, mechanistic trace, and answer integrity for every item.Each dimension had a one-line definition and anchor descriptions shown beside the rating slider.
- C.9 Human Evaluation Details: Holistic Quality judged whether an item exemplified a high-quality benchmark question, ranging from publication-ready to fundamentally flawed.Its anchors were 5 = publication-ready, 3 = acceptable with minor edits, and 1 = fundamentally flawed.
- C.9 Human Evaluation Details: Visual Necessity assessed whether answering required watching the visual content rather than reading text alone.The highest anchor required a specific on-screen cue unavailable without watching, while the lowest had no audiovisual anchor.
- C.9 Human Evaluation Details: Mechanistic Trace evaluated whether the correct option linked technique → effect → intent through a specific craft object and causal explanation.The lowest anchor represented aesthetic labels without a mechanism.
- C.9 Human Evaluation Details: Answer Integrity evaluated whether the option set contained a uniquely supported correct answer and diverse distractors, or a closed correct subset for multi-select.Lower scores reflected weak, wrong, duplicated, or missing answer options.
- C.9 Human Evaluation Details: Raters completed a calibration walkthrough on three pre-selected pairs per category and reserved anchors 0 and 5 for unambiguous cases.Calibration occurred before scoring began.
D.1 Single-Select Evaluation: Chance-Adjusted Accuracy · D.2 Multi-Select Evaluation: Precision, Recall, and F1
MuseBench evaluates single-select answers with chance-adjusted accuracy to normalize varying random-guess baselines, and evaluates multi-select answers with set-based precision, recall, and F1. These metrics improve comparability and diagnose incorrect distractor selection and missed valid perspectives beyond exact-match accuracy.
- D.1 Single-Select Evaluation: Chance-Adjusted Accuracy: Chance-adjusted accuracy normalizes single-select performance across questions with different numbers of answer options and random-guess baselines.Raw accuracy is not directly comparable when option counts vary; CAA measures performance relative to random guessing.
- D.1 Single-Select Evaluation: Chance-Adjusted Accuracy: CAA assigns 1 to correct predictions, has expected score 0 under uniform random guessing, and becomes negative for worse-than-chance performance.The score is defined using each question’s random-guess accuracy c_i and correctness indicator a_i.
- D.1 Single-Select Evaluation: Chance-Adjusted Accuracy: Averaging CAA over N_single single-select questions yields the reported single-select score.The evaluation aggregates question-level chance-adjusted scores across the single-select set.
- D.1 Single-Select Evaluation: Chance-Adjusted Accuracy: CAA makes results more comparable across option counts because identical raw accuracy on questions with more options indicates stronger performance.The metric accounts for the increased difficulty implied by larger answer spaces.
- D.2 Multi-Select Evaluation: Precision, Recall, and F1: Multi-select evaluation supplements exact-match accuracy with precision, recall, and F1 because exact matching treats near-correct and completely incorrect predictions alike.These metrics evaluate overlap between the predicted answer set and the ground-truth correct-option set.
- D.2 Multi-Select Evaluation: Precision, Recall, and F1: For each multi-select question, TP counts selected correct options, FP counts selected incorrect options, and FN counts missed correct options.The counts are defined as TP_j = |Ŷ_j ∩ Y_j|, FP_j = |Ŷ_j \ Y_j|, and FN_j = |Y_j \ Ŷ_j|.
- D.2 Multi-Select Evaluation: Precision, Recall, and F1: Precision, recall, and F1 are computed per question and can be macro-averaged across questions or micro-averaged after aggregating counts.Macro averaging averages per-question scores, whereas micro averaging first combines TP, FP, and FN across questions.
- D.2 Multi-Select Evaluation: Precision, Recall, and F1: Set-based metrics diagnose incorrect distractor selection through precision, missed valid analytical perspectives through recall, and their balance through F1.Compared with exact-match accuracy, they provide a more informative characterization of model behavior.
E Complete Experimental Results · F Broader Impact · G Additional Examples of MUSEBENCH
The complete results report zero-shot evaluation procedures and category-level performance patterns, while the broader-impact discussion positions MUSEBENCH as a probe for intent-level artistic reasoning and the examples demonstrate its inferential coverage across four art domains.
- E Complete Experimental Results: All open-source MLLMs use a single internal NVIDIA A800 cluster, while proprietary systems are queried through official APIs in one zero-shot pass over the full test set.The proprietary systems listed include GPT-5.4, Claude-4.6-Opus, and Gemini-2.5-Pro.
- E Complete Experimental Results: 57.77% Stage ACC and 50.20% Cinematic ACC make Stage Performing Arts and Cinematic Arts the most accessible categories for Claude-4.6-Opus.For GPT-5.4, the corresponding values are 51.68% and 47.58%; for Doubao-Seed-1.8-Pro, they are 50.56% and 48.38%.
- E Complete Experimental Results: Game Arts trails by 15 to 25 points for nearly every system, while category-level ACC rankings remain almost identical across rows.The passage identifies Stage Performing Arts and Cinematic Arts as most accessible and Game Arts as the trailing category.
- F Broader Impact: MUSEBENCH targets intent-level reasoning about audiovisual artistic expression across cinematic, static visual, stage performing, and game arts.The benchmark exposes a gap between expert artistic understanding and current models.
- F Broader Impact: The benchmark can guide training corpora and instruction tuning toward richer cultural and stylistic supervision and serve as a standardized probe for arts education tools and accessibility.The supplied passage also identifies the expert-model understanding gap as a motivation for these uses.
- G Additional Examples of MUSEBENCH: Additional samples present eight uniformly sampled frames, a multiple-choice question, all answer options, the highlighted correct option or options, and metadata for question type, option count, and sub-domain.Across categories, the pairs demand inferential reasoning grounded in on-screen evidence.
- G Additional Examples of MUSEBENCH: Game Arts examples test boss design, level pacing, narrative cinematography, and audiovisual signaling to probe design intent behind visual choices.The examples are presented as analytical readings of game elements rather than perceptual identification alone.
- G Additional Examples of MUSEBENCH: Cinematic, Stage Performing, and Static Visual Arts examples cover shot composition and editing, staged coordination, and compositional, color, technical, and aesthetic analysis.These examples respectively emphasize dramatic or compositional effects, jointly constructed meaning in live performance, and discrimination among visually similar artistic choices.