Source-linked AI summary

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Ziwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yang, Zhen Xing, Yuqing Yang, Qi Dai, Lili Qiu, Chong Luo

arXiv:2604.08540v1cs.CVcs.AIcs.CL

TL;DR

T2AV evaluation remains fragmented because existing benchmarks often isolate modalities or use coarse similarity, despite realistic prompts requiring joint correctness. AVGen-Bench introduces a task-driven benchmark with 11 real-world categories and a hybrid specialist-model/MLLM evaluation suite. Results show strong general audio-visual aesthetics alongside weak fine-grained semantic control, while pitch evaluation exhibits a documented floor-effect limitation.

  • Problem

    Existing T2AV benchmarks often assess audio and video separately or use coarse embedding metrics, limiting verification of fine-grained joint correctness.

  • Method

    AVGen-Bench uses realistic prompts across 11 categories and combines specialist models with MLLMs to evaluate perceptual quality, alignment, and semantic constraints.

  • Results

    The evaluation reveals strong general audio-visual aesthetics but weak fine-grained semantic control, including failures in text rendering, pitch accuracy, and physical plausibility.

  • Takeaways & Limitations

    Current T2AV systems require finer-grained supervision to progress beyond coarse alignment toward physically grounded world models.

  • Takeaways & Limitations

    Pitch Accuracy has a lower Pearson correlation of 0.5544 because extremely poor pitch control creates a floor effect and narrow low-score range.

Abstract

from arXiv · show

Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embedding similarity, failing to capture the fine-grained joint correctness required by realistic prompts. We introduce AVGen-Bench, a task-driven benchmark for T2AV generation featuring high-quality prompts across 11 real-world categories. To support comprehensive assessment, we propose a multi-granular evaluation framework that combines lightweight specialist models with Multimodal Large Language Models (MLLMs), enabling evaluation from perceptual quality to fine-grained semantic controllability. Our evaluation reveals a pronounced gap between strong audio-visual aesthetics and weak semantic reliability, including persistent failures in text rendering, speech coherence, physical reasoning, and a universal breakdown in musical pitch control. Code and benchmark resources are available at http://aka.ms/avgenbench.

1. Introduction

AVGen-Bench addresses fragmented T2AV evaluation with realistic task-driven prompts and a joint, fine-grained framework. It combines specialist models and MLLMs to assess perceptual quality, cross-modal behavior, and semantic control.

  • Evaluation Framework: AVGen-Bench differs from prior benchmarks through joint audio-visual evaluation, metrics across 10 dimensions, and rich prompts with high token counts.These design choices target rigorous assessment of realistic, complex generation requests.
  • Evaluation Gap: Existing benchmarks often evaluate audio and video separately, leaving joint fine-grained T2AV correctness insufficiently assessed.Real prompts interleave visual and acoustic requirements and expose failures in speech, sound-event alignment, lip motion, pitch, and physical logic.
  • Benchmark: AVGen-Bench introduces realistic prompts spanning 11 daily-life categories and professional, creator-economy, and world-simulation scenarios.Its task-centric design evaluates whether models accomplish user intentions rather than only achieving perceptual quality.
  • Evaluation Framework: The benchmark jointly evaluates uni-modal quality, audio-visual consistency, and fine-grained semantic alignment using specialist models and MLLMs.Specialists provide signal-level measurements, while MLLMs support semantic reasoning and holistic intent verification.
  • Findings: The framework systematically diagnoses weaknesses that coarse or modality-isolated evaluations may miss.The paper highlights a gap between strong audio-visual aesthetics and weak fine-grained semantic control, especially in text, speech, and physical reasoning.

2. Related Works

Prior T2AV research advances synchronized generation, but benchmark protocols remain limited in modality coverage, prompt complexity, and fine-grained verification. AVGen-Bench is positioned against these gaps with broader audio-aware and semantic evaluation.

  • Benchmark Comparison: AVGen-Bench reports the highest average prompt complexity and comprehensive metrics covering all audio modalities in its comparison with existing benchmarks.The comparison is summarized in Table 1.
  • Joint Audio-Video Generation: Recent T2AV systems pursue unified audio-video generation through dual-stream, flow-matching, hybrid, and conditional architectures.Examples include Sora 2, Veo 3.1, Wan 2.6, Kling 2.6, Ovi, JavisDiT, LTX-2, and MAViD.
  • Benchmark Comparison: The benchmark literature spans high-fidelity generation models and increasingly unified evaluation efforts, but fine-grained interpretable verification remains an open need.AVGen-Bench is introduced as a response to the limitations of coarse matching and modality-separated protocols.
  • Benchmark Limitations: Prior benchmark protocols commonly isolate visual and acoustic modalities, limiting assessment of joint generation quality.Visual benchmarks focus on video, while audio benchmarks evaluate sound separately and may rely heavily on subjective human evaluation.
  • Benchmark Limitations: Embedding-based metrics such as CLIP and CLAP support general matching but cannot verify specific musical notes or precise synchronization.These black-box metrics may therefore miss fine-grained hallucinations and semantic failures.

3. AVGen-Bench

AVGen-Bench is built from intent-first, human-reviewed tasks across realistic domains and evaluated with hybrid specialist–MLLM pipelines. Its modules cover perceptual quality, cross-modal alignment, and targeted semantic constraints.

  • Task-Driven Prompt Curation: The benchmark uses an intent-first taxonomy and a Human-in-the-Loop pipeline to curate prompts from realistic user scenarios.GPT-5.2 generates candidates from scenario definitions, followed by manual review for complexity and clarity.
  • Task-Driven Prompt Curation: The dataset contains 235 curated tasks across 3 domains and 11 sub-categories, averaging 1.6 shots per prompt.Speech appears in 44% of samples and environmental sound effects in 88%.
  • Task-Driven Prompt Curation: Prompt curation is decoupled from evaluation metrics so tasks derive from genuine user needs and can be extended to new domains.This avoids reverse-engineering prompts around available detectors.
  • Task-Driven Prompt Curation: Creator-economy prompts inject musical scales and chord constraints to test whether audio frequencies match visual finger positions.This targets precise audio-visual alignment rather than generic music generation.
  • Multi-Granular Evaluation: The evaluation suite combines lightweight specialist models with MLLMs across uni-modal aesthetics, cross-modal alignment, and text-to-media consistency.Specialists target low-level signal fidelity, while MLLMs support high-level semantic reasoning.
  • Fine-Grained Evaluation Modules: Fine-grained modules evaluate scene text, facial consistency, pitch accuracy, speech intelligibility, physical plausibility, and holistic semantic alignment.The workflows combine tools such as OCR, InsightFace, DBSCAN, Audio-to-MIDI, ASR, kinematics, and constraint decomposition.

4. Experiment

AVGen-Bench reveals that current T2AV systems combine strong perceptual aesthetics with persistent failures in fine-grained visual, audio, physical, and semantic control. Automated evaluation generally aligns with expert judgment and remains stable across repeated runs, although pitch evaluation is weaker.

  • Basic Uni-modal Quality: Visual quality is consistently high, with Seedance-1.5 Pro reaching 0.970 and Veo 3.1 reaching 0.960.Top-scoring models produce professional lighting, composition, and cinematic aesthetics.
  • Basic Uni-modal Quality: Audio production quality trails visual quality, while Seedance-1.5 Pro reaches a PQ score of 7.48.Higher PQ corresponds to crisper sound, whereas lower scores are associated with background noise or signal artifacts.
  • Basic Cross-modal Alignment: AV synchronization remains imperfect, with mean absolute offsets of 0.2s–0.44s and Lip Sync errors from 2.0 to over 5 frames.Offsets above 2 frames can disrupt the perceptual illusion of a talking head.
  • Fine-grained Visual: Text rendering degrades with longer or smaller text, while incidental text fails universally through glyph collapse or graffiti-like scribbles.Short, explicitly prompted text in dominant regions is more reliable than lengthy or incidental text.
  • Fine-grained Visual: Kling-V2.6 reaches only 57.33 in facial consistency, with identity drift after shot changes and severe degradation in multi-face scenes.Other models score around 48–54, while crowds produce distorted features and flickering.
  • Fine-grained Audio: All models score below 12/100 for pitch accuracy, generating random notes despite prompts specifying scales or chord sequences.Models can reproduce instrument timbre but not explicit musical pitch control.
  • Fine-grained Audio: Veo 3.1 Quality and Fast achieve speech scores of 96.09 and 94.53, but open-source systems exhibit hallucinated or truncated speech.Incidental speech can become gibberish, while long or complex verbatim dialogue may omit words or end prematurely.
  • Physical Plausibility: Most models fail the 4.0 passing threshold for low-level kinematic plausibility and often misrepresent high-level physical phenomena.For sodium dropped into water, models commonly show sinking or color changes instead of floating with correct dynamics.

5. Conclusion

AVGen-Bench finds a sharp dichotomy in T2AV generation: systems produce strong general audio-visual aesthetics but weak fine-grained semantic control. The authors argue that coarse alignment training is insufficient for precise pitch, text, and physical reasoning.

  • 5. Conclusion: Current state-of-the-art T2AV models create cinematic audio-visual content but fail significantly at fine-grained semantic control.Low performance is especially evident in precise pitch, text rendering, and physical-logic tasks.
  • 5. Conclusion: The findings suggest that future training should prioritize finer-grained supervision and physically grounded world models.The stated goal is to move beyond probabilistic texture generation.

A. Additional Qualitative Results

Additional qualitative examples organize the paper’s observed failures into text rendering, consistency and speech, and physical and semantic logic categories. Figure 7 specifically illustrates prompted-text and incidental-text rendering failures.

  • A. Additional Qualitative Results: The qualitative appendix groups failures into text rendering, consistency and speech, and physical and semantic logic categories.These categories correspond to Figures 7, 8, and 9.
  • Text Rendering Failures: Figure 7 shows glyph collapse, layout errors, and illegible or misplaced prompted text, alongside gibberish hallucinations for incidental background text.The examples include strings such as “Your customers are talking” and “EIGHTY-SEVEN SECONDS.”

B. Human Evaluation Protocols and Interfaces

The human-evaluation infrastructure uses a unified Gradio annotation platform and selects pairwise or pointwise protocols according to the evaluation dimension.

  • B. Human Evaluation Protocols and Interfaces: A unified Gradio platform supports the paper’s reproducible meta-evaluation.The annotation strategy chooses Pairwise or Pointwise protocols based on the evaluation dimension.

B.1. Hybrid Annotation Strategy

AVGen-Bench uses hybrid annotation protocols tailored to each evaluation dimension. Pairwise comparison handles relative subjective quality, while pointwise scoring captures absolute correctness in text rendering.

  • Pairwise Comparison: Speech Quality and Holistic Semantic Alignment are evaluated through blind pairwise A/B comparisons with randomized presentation and a Tie option.Annotators compare two anonymized videos side by side and vote for the superior model or select Tie.
  • Pairwise Comparison: Side-by-side comparison is used because relative judgments such as which voice sounds more natural are cognitively easier and more consistent.
  • Pairwise Comparison: The pairwise interface displays required speech lines, forcing verification of verbatim adherence to prompt constraints.
  • Pointwise Scoring: Text Rendering is assessed pointwise because pairwise comparison can mark two equally poor outputs as a Tie, obscuring absolute failure.Absolute scoring captures whether text is legible and correctly spelled.
  • Pointwise Scoring: Text quality uses a three-point rubric: Good for fully legible and correct, OK for minor artifacts but legible, and Poor for illegible, hallucinated, or missing text.
Loading 2604.08540v1…