Source-linked AI summary
VABench: A Comprehensive Benchmark for Audio-Video Generation
Daili Hua, Xizhi Wang, Bohan Zeng, Xinyi Huang, Hao Liang, Junbo Niu, Xinlong Chen, Quanqing Xu, Wentao Zhang
TL;DR
Existing benchmarks provide strong visual evaluation but lack comprehensive, convincing assessment of synchronous audio-video generation. VABench addresses this gap with a multidimensional benchmark spanning generation tasks, content categories, cross-modal properties, synchronization, and stereo audio, and its analysis highlights the difficulty of balancing semantics, synchronization, and realism.
Problem
Existing benchmarks lack systematic evaluation for synchronous audio-video generation, including cross-modal coupling, higher-order coherence, and stereo spatial-audio properties.
Method
VABench evaluates T2AV, I2AV, and stereo audio-video generation across seven content categories using 15 fine-grained metrics and dedicated spatial-audio tests.
Results
VABench provides automated, multidimensional, and human-aligned assessment while revealing that semantic consistency, synchronization, and realism are difficult to balance simultaneously.
Takeaways & Limitations
VABench offers a reliable and interpretable framework for assessing synchronous audio-video generation and informing more coherent, perceptually grounded systems.
Takeaways & Limitations
Manual inspection found demographic bias in generated human subjects, with models tending toward different demographic appearances.
Abstract
from arXiv · showhide
Recent advances in video generation have been remarkable, enabling models to produce visually compelling videos with synchronized audio. While existing video generation benchmarks provide comprehensive metrics for visual quality, they lack convincing evaluations for audio-video generation, especially for models aiming to generate synchronized audio-video outputs. To address this gap, we introduce VABench, a comprehensive and multi-dimensional benchmark framework designed to systematically evaluate the capabilities of synchronous audio-video generation. VABench encompasses three primary task types: text-to-audio-video (T2AV), image-to-audio-video (I2AV), and stereo audio-video generation. It further establishes two major evaluation modules covering 15 dimensions. These dimensions specifically assess pairwise similarities (text-video, text-audio, video-audio), audio-video synchronization, lip-speech consistency, and carefully curated audio and video question-answering (QA) pairs, among others. Furthermore, VABench covers seven major content categories: animals, human sounds, music, environmental sounds, synchronous physical sounds, complex scenes, and virtual worlds. We provide a systematic analysis and visualization of the evaluation results, aiming to establish a new standard for assessing video generation models with synchronous audio capabilities and to promote the comprehensive advancement of the field.
1. Introduction
VABench addresses the lack of systematic evaluation for synchronous audio-video generation with a comprehensive benchmark spanning diverse tasks, content categories, metrics, and stereo-audio properties.
- 1. Introduction: Existing benchmarks lack comprehensive coverage of cross-modal coupling, physical plausibility, emotional expressiveness, and spatial acoustic properties in joint audio-video generation.Examples include motion-induced Doppler effects, audiovisual emotion, and coordination between background music and visual rhythm.
- 1. Introduction: VABench evaluates audio-video generation through T2AV and I2AV tasks across seven content categories, including animals, music, physical sounds, complex scenes, and virtual worlds.The benchmark filters test samples with human workers and large language models while adjusting their distribution.
- 1. Introduction: 15 fine-grained metrics assess audio-visual synchronization, lip-speech consistency, cross-modal similarity, and other multimodal quality dimensions.The metrics comprise eight expert-model-based measures and seven multimodal-LLM-based measures.
- 1. Introduction: VABench adds dedicated stereo dual-channel evaluation for spatial auditory perception and sound-field rendering, properties overlooked by existing benchmarks.This extends evaluation beyond perceptual coherence toward spatial audio realism.
2. Related Works
Related work has advanced visual and audio-video generation, but evaluation remains insufficient for reference-free, scalable, and holistic assessment of joint audio-visual synthesis.
- 2. Related Works: Recent audio-video generation methods include unified models and video-to-audio modules whose performance depends substantially on video-to-audio quality.MMAudio and Kling-Foley improve semantic controllability and event alignment through joint video-text conditioning.
- 2. Related Works: Visual benchmarks such as VBench and VBench2.0 measure frame quality, temporal consistency, and physical plausibility but lack mechanisms for cross-modal consistency.Earlier metrics such as IS and FVD were replaced by multidimensional visual frameworks.
- 2. Related Works: Existing joint audio-visual benchmarks remain underdeveloped because many require real audio-video references, limiting their suitability for text-to-audio-video generation.T2AV instead requires reference-free evaluation of consistency among text, video, and audio.
- 2. Related Works: Current evaluations are constrained by manual assessment and limited quantitative coverage of higher-order physical or emotional coherence.These limitations reduce scalability and leave important multimodal couplings insufficiently measured.
3. VABench
VABench combines T2AV, I2AV, and stereophonic evaluation across seven content categories with a multi-track framework spanning expert metrics and MLLM-based assessment.
- 3.1. Tasks and Data Categories: VABench evaluates T2AV, I2AV, and stereophonic audio generation across a taxonomy spanning basic sources, physical interactions, complex semantics, and non-realistic content.The benchmark’s categories are grounded in human auditory perception and include seven major content groups.
- 3.2. Data Generation and Collection: The benchmark dataset uses dual T2AV and I2AV pipelines with LLM- and VLM-generated prompts and QA pairs followed by human verification.The pipelines produce 778 T2AV and 521 I2AV samples while checking semantic accuracy, observability, categorization, and physical or commonsense constraints.
- 3.3. Evaluation Metrics: VABench combines expert-model evaluation of perceptual quality with MLLM-based evaluation of complex audio-visual semantics.The framework is designed to pair specialized precision with holistic assessment.
- 3.3.1. Expert Model-based Evaluation: Its expert metrics cover uni-modal quality, cross-modal semantic alignment, and temporal synchronization using measures such as text-video, text-audio, and audio-visual alignment.Audio quality includes speech clarity, speech quality and naturalness, and audio aesthetics; cross-modal metrics use specialized models including ViCLIP and CLAP.
- 3.3.2. MLLM-based Evaluation: MLLM-based evaluation uses coarse macro scores for alignment and realism alongside fine-grained QA accuracy averaged across samples.Each sample receives detail-oriented questions, and the fine-grained score averages the proportion of satisfied requirements across the test set.
- 3.3.3. Stereophonic Analysis: Stereophonic analysis evaluates spatial imaging quality and signal integrity through human checks and nine acoustic metrics.The metrics include stereo width, ITD and ILD stability, envelope correlation, and transient synchronization.
4. Experiment
Experiments compare end-to-end and decoupled systems across T2AV, I2AV, stereo audio, and fine-grained category evaluations. Results show complementary strengths, but integrated AV models generally perform better on synchronization, semantics, and complex audio-visual behavior.
- Text to Audio-Video Generation: Veo3 leads overall T2AV performance, Wan2.5 achieves the strongest synchronization, and Sora2 offers stronger realism despite weaker audio aesthetics and synchronization.The results indicate that semantic consistency, synchronization, and realism remain difficult to optimize simultaneously.
- Main Results: Figure 5 compares key video frames and audio waveforms across I2AV, Stereo, and T2AV, while Figure 6 evaluates QA performance across seven audio categories.These visualizations provide qualitative cross-task comparisons and fine-grained category-level evaluation.
- Main Results: Integrated AV models generally outperform V+A systems in T2AV and I2AV, while decoupled systems remain viable alternatives.The reported advantage is attributed to end-to-end joint training and its ability to capture cross-modal synergies.
- Image to Audio-Video Generation: Seedance+MMAudio achieves the best major I2AV results, while image conditioning narrows model gaps and can let V+A systems surpass AV models on alignment.AV models retain a clearer advantage in T2AV, whereas I2AV produces more balanced scores across systems.
- Multi-Categories Analysis: Across seven categories, systems perform best on Virtual Worlds and struggle most with Complex Scenes and Human Sounds, where AV models show their largest advantages.Veo3 is strongest overall in AQA, while category-specific leaders include Sora2 for Human Sounds and Wan2.5 for Virtual Worlds or Music-related performance.
- Stereo Evaluation: Stereo evaluation finds no model reliably generating text-prompted stereo separation, although Veo3 and Sora2 sometimes produce localized spatial cues.Veo3 occasionally aligns moving spatial sources with visual motion, while Sora2 can produce distinct left-right vocal tracks in multi-speaker scenes.
5. Conclusion
VABench evaluates synchronous audio-video generation across T2AV, I2AV, and stereo tasks using automated, multidimensional, and human-aligned measures. The benchmark exposes the difficulty of balancing semantics, synchronization, and realism while supporting more interpretable evaluation.
- Conclusion: VABench provides comprehensive evaluation of synchronous audio-video generation across T2AV, I2AV, and stereo tasks.Its automated, multidimensional, and human-aligned design is intended to make performance assessment reliable and interpretable.
6. Additional evaluation metrics
Supplementary metrics reinforce broad differences among audio-video models while revealing distinct model specializations across categories and evaluation dimensions.
- SpeechClarity and Artistry: AV models achieve the best overall SpeechClarity performance, while Kling+MMAudio reaches state-of-the-art Artistry performance among the compared systems.SpeechClarity uses DNSMOS, and Artistry uses Qwen2.5 Omni 7B.
- Model specializations: Veo3 shows strong acoustic-visual detail control, Sora2 excels in Physical plausibility and event consistency, and Kling leads visual style and fidelity.Kling+MMAudio also surpasses Sora2 on Text-Audio Align and Audio-Visual Align while remaining strong in subjective dimensions.
- Audio QA: Veo3 leads Animals but is weaker in Environmental Sounds and Complex Scenes, whereas Sora2 provides the most consistently balanced category performance.The models therefore exhibit complementary strengths rather than a uniform ranking across audio QA categories.
- Model specializations: MMAudio is strongest in Environment and Virtual categories, ThinkSound specializes in Music, and Wan2.2 and Wan2.5 remain weaker in Physical and Music.Among visual-only models, Kling performs strongly in Human Sounds and Virtual, while Seedance slightly exceeds it in Environment.
- Cross-category comparison: AV models consistently occupy the top positions across categories, with their largest advantage over V+A models in Human Sounds and Virtual Worlds.The reported pattern attributes this advantage to integrated spatial, material, rhythmic, and emotional audio cues.
- Cross-category comparison: Audio cues improve visual dynamism and narrative coherence in abstract or surreal Virtual scenarios, extending their role beyond temporal synchronization.The analysis links these cues to rhythm, energy distribution, and emotional tone.
7. Audio-Video Generation Models in Evaluation
The evaluation uses each video model’s default generation configuration and records attributes such as duration and frame rate for comparison.
- Evaluation settings: Experiments follow each video generation model’s default configuration parameters, including output duration and frame rate.These settings are summarized in Table 4.
- Evaluation settings: Sora2 generates 10-second videos at 30 FPS by default, while several other evaluated models generate 5–8-second videos at 24 FPS.The latter group includes Veo3-fast, Wan2.5 Preview, Seedance-1.0-Lite, Wan2.2-TI2V, and Kling2.5 Turbo.
8. Detail Analysis of Different Tasks
Across T2AV and I2AV, category strengths remain broadly consistent, while image conditioning stabilizes realism and alignment but constrains artistic variation.
- Task comparison: The section compares T2AV and I2AV results using Tables 5 and 6 to identify shared strengths and the effects of image-conditioned input.The comparison focuses on common category patterns and task-specific outcome differences.
- Common Strengths and Core Challenges: Music is the strongest category across both T2AV and I2AV tasks, indicating robust handling of structured melodic content.This pattern is reported across most evaluation metrics.
- Impact of Image Input: I2AV image conditioning reduces cross-model variance and raises minimum performance in Alignment, Visual Realism, and Audio Realism.The supplied analysis describes this as a stabilization effect relative to T2AV.
- Impact of Image Input: Artistry scores are highest for T2AV Virtual content but converge around 3.8–4.0 across I2AV categories, suggesting reduced creative variation under image conditioning.The analysis associates this convergence with fidelity to the visual reference.
9. Qualitative Analysis
Qualitative analyses test physical timing, Doppler behavior, and stereo spatialization, finding partial success in synchronization but substantial weaknesses in semantic spatial control.
- Doppler effect: Veo3 most clearly reproduces the Doppler effect, with a smooth descending frequency trajectory, while Wan2.5 captures attenuation but a weaker shift.The analysis uses spectrograms of airplane-video outputs to assess approach-and-recede acoustics.
- Lightning and thunder: For lightning scenes, the evaluated systems place thunder after visible lightning or sustain it after the flash, consistent with light arriving before sound.The analysis examines Veo3, Wan2.5, and Kling+MMAudio using video and spectrogram evidence.
- Stereophonic analysis: Veo3 creates channel-dependent energy shifts and perceptible depth, but primarily through energy panning rather than semantic separation of waves and seagulls.Its high cross-correlation leaves source localization ambiguous.
- Stereophonic analysis: Sora2 produces nearly identical channels, effectively mono audio in dual-channel form despite synchronization, while Wan2.5 reports a 0.9998 channel correlation.Both outputs lack the requested left-right source layout and stereophonic width.
- Stereophonic analysis: All evaluated models show room for improvement in semantic stereo generation, especially in localizing requested sound sources across channels.The coastal test specifies waves on the left and seagulls on the right.
10. Special samples Analysis
Special samples show Veo3 reproducing Doppler and spatial audio dynamics, while Sora2 translates emotional conflict into alternating stereo channels; manual inspection also reveals demographic representation bias.
- Veo3 stereo and physical audio: Veo3 reproduces Doppler trajectories, channel dominance, and spatial alignment with moving vehicles, supporting physical and stereophonic consistency.Its spectrogram shows a Doppler arc and high-frequency tire-friction bursts, while waveform analysis aligns channel balance with visual approach and recession.
- Sora2 emotional stereo rendering: Sora2 alternates left-right whisper tracks and dual-channel bursts to render an internal conflict as a structured emotional stereo scene.The alternating spatial separation and emotional tones represent temptation and conscience as an adversarial dialogue.
- Demographic bias: Manual inspection finds demographic representation biases, with Veo3 tending toward Caucasian features and Seedance toward Asian appearances.The authors hypothesize that these tendencies correlate with model geography and private training-data distributions.
11. MLLM Based Evaluation Cases
The MLLM evaluation uses specialized prompts, five-point rubrics, and timestamped JSON judgments to assess narrative, visual, and micro-level audio-video qualities. Example cases show strong QA performance for Wan2.5 and Veo3, contrasting with weaker Seedance+MMAudio results.
- Evaluation workflow: The MLLM workflow casts each evaluation as a specialized expert task with a five-point rubric and constrained JSON output.Prompts target dimensions such as narrative-emotional synergy and visual realism, requiring timestamped reasons tied to observed moments or anomalies.
- Micro-evaluation examples: Seedance+MMAudio’s T2AV example preserves footsteps and bottle clinks but omits the distant bus rumble and clear bird chirping.Its QA answers also report an overall soundscape that includes undescribed traffic noise and footsteps.
- Micro-evaluation examples: Wan2.5 and Veo3 each receive an Overall Score of 0.8 in the I2AV example, whereas Seedance+MMAudio receives 0.4 in the T2AV example.The example QA responses support the higher I2AV scores and identify multiple missing or present sound elements in the T2AV comparison.
- Micro-evaluation examples: Wan2.5’s I2AV example is supported by correctly reproduced powerful music and deep bass spanning the venue.The cited QA answer directly confirms the requested musical and spatial sound characteristics.