Source-linked AI summary
MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
Yujie Wei, Yujin Han, Zhekai Chen, Yongming Li, Kaixun Jiang, Zhihang Liu, Quanhao Li, Zhiwu Qing, Xiang Wang, Zhen Xing, Ruihang Chu, Lingyi Hong, Yefei He, Junjie Zhou, Junqiu Yu, Yang Shi, Difan Zou, Kai Zhu, Shiwei Zhang, Yingya Zhang, Yu Liu, Xihui Liu, Hongming Shan
TL;DR
MSAV models lack comprehensive, reliable evaluation across diverse cinematic audio-video tasks. MSAVBench addresses this with broad coverage and adaptive hybrid assessment, finding persistent weaknesses in director-level control and fine-grained audio-visual synchronization.
Problem
Existing MSAV benchmarks lack comprehensive audio evaluation, diverse challenging scenarios, and robust adaptive pipelines for systematically assessing modern models.
Method
MSAVBench combines four-dimensional benchmark coverage with adaptive shot correction, instance-wise rubrics, and tool-grounded evidence extraction.
Results
Evaluation of 19 systems finds persistent closed–open-source gaps and continued difficulty with director-level control, cinematic structure, and fine-grained audio-visual alignment.
Takeaways & Limitations
Modular or agentic pipelines may narrow the open-source gap, while complex MSAV generation likely requires unified audio-video architectures.
Takeaways & Limitations
Some evaluation components rely on multimodal foundation-model judges, increasing the cost of large-scale evaluations.
Abstract
from arXiv · showhide
Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However, evaluating such frontier models remains a fundamental challenge. Existing benchmarks are limited in scope and data diversity, and rely on rigid evaluation pipelines, preventing systematic and reliable assessment of modern MSAV models. To bridge these gaps, we introduce MSAVBench, the first comprehensive benchmark and adaptive hybrid evaluation framework for multi-shot audio-video generation. Our benchmark spans four key dimensions, video, audio, shot, and reference, covering diverse task settings, varying shot counts of up to 15, and challenging non-realistic scenarios. Our evaluation framework improves robustness through an adaptive self-correction mechanism for shot segmentation, instance-wise rubrics for subjective metrics, and tool-grounded evidence extraction for complex judgments. Furthermore, MSAVBench achieves high alignment with human judgments, reaching a Spearman rank correlation of 91.5%. Our systematic evaluation of 19 state-of-the-art closed- and open-source models shows that current systems still struggle with director-level control and fine-grained audio-visual synchronization, while modular or agentic generation pipelines offer a promising path toward narrowing the gap between open- and closed-source models. The benchmark data and evaluation code are publicly available at https://github.com/ali-vilab/MSAVBench.
1 Introduction
MSAVBench addresses the limited coverage and rigidity of existing MSAV evaluation by combining broad, challenging benchmark dimensions with an adaptive hybrid framework. Its evaluation of 19 state-of-the-art models identifies persistent gaps in open-source performance, director-level control, and fine-grained audio-visual synchronization.
- Motivation: Existing benchmarks inadequately evaluate MSAV because they typically isolate single-shot, silent, audio, or multi-shot video aspects and lack diverse, complex scenarios.They overlook cinematic language, counterfactual content, and systematic assessment of task adaptability.
- Benchmark: MSAVBench spans video, audio, shot, and reference dimensions across diverse prompts, shot counts up to 15, subject counts, and non-realistic scenarios.The benchmark is presented as the first comprehensive benchmark for multi-shot audio-video generation.
- Findings: 19 state-of-the-art closed- and open-source models reveal a persistent performance gap, although modular or agentic pipelines may narrow it.The analysis identifies modular and agentic generation pipelines as promising for open-source systems.
- Findings: Current models remain far from reliable director-level generation, struggling with cinematic control, structural consistency, and fine-grained audio-visual synchronization.These findings highlight limitations in controlling complex MSAV narratives despite recent progress.
- Evaluation framework: Its adaptive hybrid evaluation framework improves robustness through dynamic shot-boundary correction, instance-wise rubrics, and tool-grounded evidence extraction.The evaluation suite combines specialized expert models, rubric-based scoring, and tool-grounded assessment across global, cross-shot, intra-shot, and reference levels.
2 Related Work
Prior work has progressed from single-shot visual generation and evaluation toward multi-shot audio-video systems, but existing methods and benchmarks remain limited in narrative structure, audio assessment, and prompt complexity.
- Audio-video generation models: Current video generative models mainly target single-shot synthesis, which is insufficient for multi-scene narratives requiring synchronized audio.Frontier closed-source systems have more recently explored multi-shot audio-video generation, alongside emerging open-source efforts.
- Audio-video evaluation benchmarks: Early benchmarks primarily assess single-shot visual quality, while later multi-shot benchmarks add story structure and cross-shot consistency but remain largely video-centric.The passage identifies VBench, Video-Bench, and AesVideo-Bench as early examples, and cites later multi-shot benchmarks as extending evaluation scope.
- Audio-video evaluation benchmarks: Audio-video benchmarks assess audio quality and audio-visual alignment, yet mostly use single-shot or weakly structured prompts.This limits their coverage of complex multi-shot audio-video generation settings.
3 MSAVBench
MSAVBench is designed around broad diversity and escalating complexity across video, audio, shot, and reference dimensions. Its benchmark construction combines expert taxonomy, automated prompt generation, human refinement, and reference-media curation to support systematic evaluation of cinematic, multilingual, and reference-conditioned generation.
- Data design: MSAVBench spans four dimensions—video, audio, shot, and reference—to cover varied visual content, sound and language, cinematic controls, and identity or timbre preservation.Shot conditions include professional cinematic language, while reference conditions extend beyond standard text conditioning to characters, scenes, and audio.
- Data design: The benchmark probes complexity through realistic and non-realistic subjects and scenes, including fictional worlds and counterfactual compositions.These combinations test faithful prompt adherence without mode collapse or failure on complex scenarios.
- Benchmark construction: Six experts refined 2200 generated scripts into 286 prompts containing 2198 individual shots after removing redundancy, incoherence, unnatural transitions, and hallucinations.The pipeline first samples 2200 seed quadruples, generates global-to-shot scripts, and then applies expert review and refinement.
- Benchmark composition: Its cinematic controls include 5 shot scales, 5 camera angles, varied camera movements and lighting, and multiple cross-shot transitions.These requirements target rigorous assessment of cinematic generation capabilities across shots.
- Benchmark composition: The benchmark scales shot counts from 2 to 15, averages 7.7 shots per prompt, and includes multi-subject compositions in 32.2% of prompts.Some scenarios require five or more simultaneous subjects and cross-combine realistic and non-realistic subjects and scenes.
- Evaluation framework: MSAVBench organizes evaluation into four hierarchical levels with 20 metrics covering global narrative and audio-visual alignment, cross-shot consistency, and other generation properties.The framework combines iterative shot self-correction with expert models, rubric-based VLM scoring, and tool-grounded agentic scoring.
4 Experiments
Experiments benchmark 19 representative closed- and open-source video generators, revealing persistent gaps in structural control, long-horizon consistency, and fine-grained audio-visual alignment. Modular or agentic pipelines show promise, while MSAVBench’s rubric- and tool-grounded evaluation remains robust across VLM judges.
- Experimental setup: 19 representative video generators are benchmarked across closed-source commercial systems and open-source pipelines on MSAVBench.The evaluated families include commercial systems, reference-conditioned models, and several open-source pipeline categories.
- Overall findings: Commercial systems consistently dominate, whereas modular image-plus-audio-video pipelines can rival closed systems and may enable cost-effective agentic generation.Native open-source multi-shot audio-video models remain absent because of data scarcity and prohibitive computational costs.
- Overall findings: Open-source models lag in director-level structural control, especially layout alignment and camera control, acting more like passive pixel renderers than controllable storytellers.The gap is pronounced in complex spatial and cinematic compliance despite less severe differences in basic audio-visual fidelity.
- Overall findings: Fine-grained joint audio-visual alignment remains unsolved across lip-speech synchronization, sound attribution, audio-visual synchronization, and multi-talker timbre consistency.The passage attributes these weaknesses to difficulty coupling phoneme-level audio with dynamic visual content across shots.
- Quantitative analysis: From 1–4 to 11–15 shots, Kling-V3-T2V drops 3.5%, while LongLive+HunyuanFoley drops 24.5% and Wan2.2+HunyuanFoley drops 11.7%.Performance declines for all models as shot counts increase, with open-source pipelines degrading more severely; non-realistic prompts also reduce overall scores across methods.
- Evaluation robustness: Rubric- and tool-grounded evaluation remains stable across VLM backbones, with narrative coherence dropping only from 0.850 to 0.820 and outperforming direct VLM scoring.This supports robustness to the specific VLM judge choice and the reliability of MSAVBench’s metric design.
5 Conclusion · Appendix · A More Data Details on MSAVBench
MSAVBench is presented as the first multi-shot audio-video generation benchmark with an adaptive hybrid evaluation framework. It covers multiple data dimensions and challenging scenarios while evaluating 19 state-of-the-art systems, highlighting the potential of modular and agentic open-source pipelines.
- 5 Conclusion: MSAVBench is presented as the first benchmark for multi-shot audio-video generation.
- 5 Conclusion: Its adaptive hybrid evaluation framework is designed for reliable assessment.
- 5 Conclusion: The benchmark covers video, audio, shot, and reference dimensions.
- 5 Conclusion: It includes comprehensive data coverage and challenging generation scenarios.
- 5 Conclusion: Agentic shot self-correction and stratified scoring support reliable evaluation.
- 5 Conclusion: 19 state-of-the-art systems were evaluated, with modular and agentic open-source pipelines showing potential to narrow the gap.
A.1 Data Design Details … A.2.2 LLM Prompt Templates
MSAVBench designs prompts across four orthogonal dimensions and constructs them through expert-curated taxonomies, GPT-5.4 script generation, and cinematic prompt enhancement. The resulting templates enforce structured metadata, shot-level continuity, diversity, safety, and coherent multi-shot captions.
- A.1 Data Design Details: Every prompt spans Video, Audio, Shot, and Reference dimensions, with realistic and non-realistic reality classes for subjects and scenes.The non-realistic class includes coherent fictional and counterfactual sub-types.
- A.1 Data Design Details: The Video dimension covers eight genres, six visual styles, four subject classes, and realistic or non-realistic scene types.Genres include Action, Narrative, Tutorial, Singing & Music performance, Multi-person Dialogue, Science / Game, Advertising, and Nature.
- A.1 Data Design Details: The Audio dimension specifies six content classes, seven emotions, and six spoken languages, while Shot annotations define scale, angle, motion, transition, and lighting.Reference assets include 68 subject images, 65 paired audio clips, and 32 scene images assigned across 96 prompts.
- A.2.1 Expert-Curated Sub-Category Vocabulary: Stage 1 uses an eight-genre seed taxonomy whose released vocabulary contains 144 fine-grained sub-categories.The taxonomy includes representative categories such as martial-arts duel, detective reasoning, aurora, and glacier collapse.
- A.2 Data Construction Details: The data-construction pipeline samples theme, subject, scene, and style quadruples, generates structured multi-shot scripts, rewrites them into global-to-shot prompts, and obtains expert review.GPT-5.4 performs initial synthesis, while a Prompt-Enhancement model adds explicit cinematic language.
- A.2.2 LLM Prompt Templates: Stage 2 uses GPT-5.4 initial-prompt and Prompt-Enhancement templates to produce metadata-rich scripts and cinematic prompts for downstream generators.The initial template accepts dimension constraints and a target shot count; the PE template rewrites the script into the downstream format.
A.3 Data Analysis Details · B More Evaluation Suite Details on MSAVBench · B.1 Metric Definitions, Tools and Score Mapping
MSAVBench releases a 286-prompt, 2,198-shot suite with diverse visual, acoustic, linguistic, cinematic, reference, and multi-level complexity distributions. Its evaluation framework defines 20 metrics across Story, Cross-Shot, Intra-Shot, and Reference levels, mapping tool- or judge-derived outputs to [0, 1].
- A.3 Data Analysis Details: The released benchmark contains 286 prompts and 2,198 shots, with high-level and cinematic-language distributions reported in Figures 2 and 6.Cinematic-language distributions cover shot scale, camera angle, transition, and tone × saturation.
- A.3 Data Analysis Details: The suite balances eight video genres while spanning human, animal, inanimate, and fictional subjects across realistic and non-realistic scenes.Genre shares include Action 16.4%, Tutorial 16.4%, Narrative 15.7%, Singing & Music 16.1%, Multi-person Dialogue 15.7%, Science / Game 8.4%, Advertising 8.4%, and Nature 2.8%.
- A.3 Data Analysis Details: Audio content is dominated by speech and human-made environmental sounds, while per-shot emotional colour spans seven categories and spoken content covers six languages.The supplied distributions report speech at 28.7%, human-made environmental sounds at 20.3%, and joy as the most common emotional category at 42.5%.
- A.3 Data Analysis Details: Cinematic coverage includes five major shot scales, five major shot angles, four major camera-motion types, four major transition types, and five lighting types.The most common reported categories are eye-level angles at 59.2%, hard-cut transitions at 66.9%, and push-pull camera motion at 44.6%.
- A.3 Data Analysis Details: The reference subset contains 68 subject reference images, 65 paired reference audio clips, and 32 scene reference images assigned across 96 prompts.References span realistic and anime subjects, five age buckets, multiple ethnicities, six languages, and indoor and outdoor environments.
- A.3 Data Analysis Details: Task complexity ranges from 2 to 15 shots per prompt, averages 7.7 shots, and includes multi-subject composition and combinations of realistic and non-realistic subjects and scenes.The supplied distribution reports 32.2% of prompts requiring multi-subject composition, with over 10% demanding at least five simultaneous subjects.
- B.1 Metric Definitions, Tools and Score Mapping: The MSAVBench evaluation suite organizes 20 metrics into four levels: Story, Cross-Shot, Intra-Shot, and Reference.For each metric, the framework specifies its target, tool or judge, computation, and mapping of raw output to a score in [0, 1].
B.1.1 Story-Level Metrics
Story-level evaluation measures narrative coherence, visual quality, audio-visual synchronization, lip-speech synchronization, and sound attribution using rubric-based judges and specialized audiovisual tools.
- B.1.1 Story-Level Metrics: Narrative coherence assesses event ordering, causal validity, and completeness across uniformly sampled full-video frames, scoring the proportion of positive binary judge answers.It uses a rubric-based VLM judge, Qwen 3.5.
- B.1.1 Story-Level Metrics: Visual quality measures whether prompt-specified visual attributes are realized, using prompt-instantiated multiple-choice questions scored by average answer accuracy.Each prompt slot is converted into an MCQ and evaluated by a rubric VLM judge.
- B.1.1 Story-Level Metrics: Audio-visual synchronization measures whole-video temporal alignment between visual events and sound using DeSync’s predicted global audio-video offset.The raw offset ∆t is mapped to [0, 1] by max(0, 1 −|∆t|/2.0 s).
- B.1.1 Story-Level Metrics: Lip-speech synchronization measures dialogue-shot lip-sync quality by averaging matched speaking-segment confidence across the video.It combines active-speaker localization, speaker diarization [46], and StableSyncNett [34].
- B.1.1 Story-Level Metrics: Sound attribution measures whether speech aligns temporally with the correct visible speaker through cross-modal speaker matching and mean temporal overlap ratio.It uses visual active-speaker detection and audio diarization [46].
B.1.2 Cross-Shot-Level Metrics … C Additional Experimental Details
The benchmark evaluates cross-shot, intra-shot, and reference-level properties with specialized models, rubric judges, and tool-grounded agents. It normalizes and aggregates 20 metrics into 11 dimensions, applying a shot-completion penalty to account for missing shots.
- B.1.2 Cross-Shot-Level Metrics: Cross-shot visual consistency covers layout, subject identity, background, style, illumination, and colour using adjacent-shot checks, embeddings, and rubric or agentic judges.Final scores use average pass rates or mean clipped cosine similarities, depending on the metric.
- B.1.2 Cross-Shot-Level Metrics: Cross-shot audio consistency evaluates music continuity through embeddings, BPM, and beat alignment, and voice timbre through speaker embeddings compared across dialogue-bearing shots.Music consistency is a weighted sum in [0, 1], while voice timbre consistency is the mean clipped cosine similarity.
- B.1.3 Intra-Shot-Level Metrics: Intra-shot metrics assess captioned layout and hand actions, camera adherence, audio production quality, rendered text accuracy, and script transcription.They use agentic or rubric VLM judges, Audiobox-Aesthetic, PP-OCRv5, and FireRedASR2-LLM or Whisper-large-v3, with deterministic score mappings.
- B.1.4 Reference-Level Metrics: Reference-level metrics measure subject identity and appearance fidelity, plus speaker timbre fidelity, by comparing generated embeddings with reference image or voice embeddings.Both final scores are mean clipped cosine similarities using the corresponding cross-shot embedding pipelines.
- B.1.5 Overall Score Aggregation: Five visual consistency metrics merge into Visual Quality and four dialogue-related audio metrics merge into Multi-Speaker Dialogue Audio, producing 11 final dimensions.The merger prevents overlapping atomic metrics from overweighting the same capability during aggregation.
- B.1.5 Overall Score Aggregation: All dimensions are mapped to [0, 1], averaged, and multiplied by a shot-completion penalty equal to valid generated shots divided by the specified shot count.The paper reports strong alignment between this aggregation design and human judgments.
- B.2 Stratified Scoring Paradigms: 10 metrics use specialized expert models or deterministic signal-processing pipelines, including synchronization, attribution, consistency, quality, text, transcription, and fidelity measures.These metrics avoid VLM-based reasoning and include audio-visual sync., lip-speech sync., sound attribution, style consistency, music consistency, voice timbre consistency, audio quality, text rendering accuracy, ASR transcription (WER), and voice fidelity.
- B.2 Stratified Scoring Paradigms: Five metrics use instance-wise rubric-based scoring and five use tool-grounded agentic scoring, with pass rates or localized perception evidence supporting their judgments.Rubric metrics include narrative coherence, visual quality, illumination consistency, colour consistency, and camera parameter adherence; agentic metrics include layout, subject, background, and fidelity measures.
C.1 Implementation … D.1 Experts for Benchmark Construction
The framework combines distributed perception tools, model-specific VLM judges, and cached intermediate results to support efficient multi-shot audio-video evaluation. Benchmark construction uses six qualified AIGC and audio-video researchers with multi-expert review and majority-vote resolution.
- C.1 Implementation: Perception tools run as independent FastAPI micro-services on 8×A100 hosts.This infrastructure supports the evaluation framework’s distributed tool deployment.
- C.1 Implementation: GPT-5.4 generates and enhances prompts, while Gemini 3.1 Pro judges audio and Qwen3.5 judges visual content.The framework assigns different models to prompt processing, audio-related judgments, and visual-related judgments.
- C.2 Cost-Efficient Evaluation: Tool outputs are cached at the case level and reused across metrics whenever possible.Caching avoids repeating the same perception computations for multiple metrics.
- C.2 Cost-Efficient Evaluation: Many metrics use specialized expert models or deterministic pipelines instead of VLM judges, substantially reducing evaluation cost.The framework balances evaluation accuracy with computational cost by limiting VLM calls.
- D Human Expert Annotation: Benchmark construction relies on domain experts for taxonomy design and prompt curation.The supplied passage identifies these activities as stages in the data-construction pipeline.
- D.1 Experts for Benchmark Construction: Six full-time AIGC and audio-video researchers, each with a relevant graduate degree, perform taxonomy design and prompt curation.Their fields include computer vision, multimedia, and audio signal processing.
- D.1 Experts for Benchmark Construction: Each PE-rewritten prompt receives review from at least two experts during Stage 3.The review process covers prompt filtering or refinement.
- D.1 Experts for Benchmark Construction: Disagreements are escalated to a third senior expert and resolved by majority vote.This provides a specified adjudication procedure for filtering or refinement disagreements.
D.2 Evaluation Experts and Pairwise Annotation Protocol · D.3 Annotation Interface · E Ethics, Privacy, and Licensing
MSAVBench uses experienced experts and anonymized pairwise comparisons to assess overall and fine-grained video quality. Its interface supports rubric-based ranking, while its prompt and reference materials follow privacy and licensing safeguards.
- D.2 Evaluation Experts and Pairwise Annotation Protocol: Annotators are full-time AIGC researchers and aesthetic-quality annotators with prior aesthetic-quality annotation experience.
- D.2 Evaluation Experts and Pairwise Annotation Protocol: 30 experts compare 16 video-generation models through 1,200 system-level pairwise judgments, with each annotator labeling 40 video pairs.
- D.2 Evaluation Experts and Pairwise Annotation Protocol: 10 experts perform fine-grained comparisons across narrative coherence, cross-shot layout consistency, and intra-shot layout-text alignment.
- D.2 Evaluation Experts and Pairwise Annotation Protocol: Anonymized videos appear in random order under unified rubrics, with outcomes recorded as “A wins,” “B wins,” or “both good / both bad.”Ties count as 0.5 for each method when computing win rates, and rankings are compared with automatic metrics using Spearman’s ρ.
- D.3 Annotation Interface: A custom web interface presents two candidate videos with their prompt and relevant metadata for criterion-specific preference selection.The resulting pairwise preferences are aggregated into system-level rankings; Figure 7 illustrates the interface.
- E Ethics, Privacy, and Licensing: MSAVBench prompts are synthetically generated from expert-designed taxonomies, reviewed by domain experts, and exclude personal data, identifiable individuals, sensitive content, and real proper names.
- E Ethics, Privacy, and Licensing: Reference images and audio clips come from published benchmarks with open redistribution terms, while generated evaluation videos are not redistributed.The release will include the prompt set, legally shareable reference assets, and the evaluation framework.
F Limitations
MSAVBench’s agentic evaluation pipeline may incur substantial costs because some components rely on multimodal foundation models as judges. However, its alignment with human judgment remains strong when using a smaller open-source model, indicating robustness to the VLM backbone choice.
- Evaluation pipeline limitations: Multimodal foundation-model judges can introduce additional cost in large-scale evaluations.This cost arises from components of the agentic evaluation pipeline that rely on multimodal foundation models as judges.
- Evaluation pipeline limitations: The framework remains well aligned with human judgment even when instantiated with a smaller open-source model.This result suggests the evaluation method is robust to the choice of VLM backbone.