Source-linked AI summary
PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation
Yuchen Sun, Qian Yang, Jun Wang, Detai Xin, Guoqiao Yu, Guanglu Wan, Qi Jia
TL;DR
T2AV evaluation has not adequately diagnosed audio across content types and audiovisual grounding conditions. PRISM-Bench addresses this gap with a factorized benchmark and blind, reference-based judging, finding uneven progress: perceptual audio quality has improved, but audiovisual grounding remains difficult, especially for on-screen audio.
Problem
Existing T2AV benchmarks do not jointly diagnose audio content type and sound-source visibility, limiting insight into model behavior under specific audio conditions.
Method
PRISM-Bench evaluates Speech, Music, and Sound across On-screen and Off-screen conditions using 900 human-verified samples, four dimensions, 35 criteria, and blind side-by-side judging against ground-truth references.
Results
Evaluation shows uneven progress: Seedance 2.0 leads across subsets, while audiovisual grounding remains weaker than perceptual audio quality and frontier proprietary systems outperform open-source models.
Takeaways & Limitations
Factorized evaluation reveals grounding-sensitive weaknesses that aggregate progress in acoustic realism and prompt adherence does not eliminate.
Abstract
from arXiv · showhide
Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio modality. Existing benchmarks either treat audio as an auxiliary component of video quality or assess it in isolation from audiovisual grounding, making it difficult to diagnose where current systems truly succeed or fail in audio generation. We present PRISM-Bench, the first audio-centric diagnostic benchmark for T2AV generation. Built from a rigorously curated dataset of 900 human-verified samples, PRISM-Bench factorizes audio evaluation along two orthogonal axes: audio type (Speech, Music, and Sound) and sound-source visibility (On-screen vs. Off-screen). It evaluates generated content across four perceptual dimensions (Audio-Visual Coherence, Audio Quality, Audio Expressiveness, and Prompt Following) with 35 fine-grained criteria. To ensure reliable assessment, we adopt an enhanced MLLM-as-a-Judge protocol based on blind, side-by-side comparison against ground-truth references, demonstrating strong alignment (over 70% mean agreement) with human raters. Our evaluation of recent T2AV systems highlights a significant performance gap between frontier and open-source models. Furthermore, we demonstrate that current generation paradigms overfit to perceptual fidelity while struggling with complex grounding and control tasks, particularly in generating music and synchronized On-screen audio.
1 Introduction
PRISM-Bench addresses a gap in T2AV evaluation by jointly diagnosing audio type and sound-source visibility rather than relying on aggregate audio-video scores. It combines a curated benchmark, fine-grained perceptual criteria, and blinded reference-based judging to expose model-specific strengths and weaknesses.
- Existing benchmarks organize evaluation around tasks, scenarios, broad dimensions, or physical failures, without jointly stratifying audio content type and source visibility.
- PRISM-Bench factorizes evaluation into Speech, Music, and Sound across On-screen and Off-screen conditions, using 900 human-verified samples, four perceptual dimensions, and 35 criteria.
- The benchmark uses blinded, randomized side-by-side comparisons between generated and ground-truth samples, with each candidate evaluated independently.
- Seedance 2.0 ranks first across On-screen, Off-screen, and Mixed subsets, while factorized scores reveal weaknesses hidden by its aggregate lead.
- On-screen Audio-Visual Coherence and Off-screen Sound Prompt Following trail the model’s other dimensions, while frontier proprietary systems retain an advantage over open-source models.
- The findings indicate that acoustic realism and prompt adherence alone do not resolve fine-grained audiovisual grounding and sound-event control.
2 Related Work
T2AV generation has shifted toward unified audiovisual models, while evaluation has expanded but remains insufficiently factorized. Existing benchmarks do not jointly diagnose sound type and source visibility, limiting insight into model behavior across grounding conditions.
- T2AV systems increasingly use unified or tightly coupled generation to improve temporal alignment and semantic consistency between visual events and audio.
- Recent systems include both proprietary and publicly released generators, making systematic comparison more feasible.
- Prior benchmarks address audiovisual alignment, signal-level quality, physical grounding, multi-speaker scenarios, and blinded human comparisons, but not the joint audio-type and visibility taxonomy.
- Existing T2AV benchmarks therefore provide limited insight into behavior across Speech, Music, and Sound under On-screen and Off-screen conditions.
3 PRISM-Bench
PRISM-Bench is built as an audio-centric, factorized benchmark with a curated dataset and a visibility-aware evaluation framework. Its pipeline combines scenario-aware curation, multi-stage annotation, human verification, and tag-conditioned blind judging across audio types and grounding conditions.
- 3.1 Overview: The benchmark organizes audio type and sound-source visibility into On-screen, Off-screen, and Mixed settings for fine-grained diagnosis.
- 3.2 Dataset Curation Pipeline: The dataset pipeline uses scenario-aware sampling, multimodal and audio-only captioning, cross-source fusion, and benchmark-level human correction.
- 3.2 Dataset Curation Pipeline: All 900 samples undergo human verification of synchronization, speech, paralinguistic cues, music, environmental sounds, audio type, and source visibility; the benchmark contains three sets of 300 pairs.
- 3.2 Dataset Curation Pipeline: On-screen samples generally receive longer captions than Off-screen samples because they require descriptions of audible content, visible sources, and grounding relations.
- 3.3 Evaluation Dimensions: Evaluation activates type- and visibility-specific criteria for Audio-Visual Coherence, Audio Quality, Audio Expressiveness, and Prompt Following.
- 3.4 Evaluation Framework: Generated and ground-truth candidates are presented under matched conditions without revealing which is which, reducing presentation bias in controlled comparisons.
- 3.4 Evaluation Framework: The two-stage judge evaluates coherence and quality from merged clips, then expressiveness and prompt following using human-verified captions as references.
4 Experiments
PRISM-Bench evaluates recent T2AV systems across visibility conditions, audio types, and perceptual dimensions, revealing strong proprietary-model performance alongside persistent grounding weaknesses. Its MLLM-based evaluation shows over 70% mean agreement with human judgments, supporting scalable assessment.
- Overall results: Seedance 2.0 leads every subset with Final scores of 97.73, 78.11, and 99.88, while the strongest open-source results are 49.50, 40.16, and 64.38.The proprietary-model lead over the strongest open model is at least 35 points on every subset.
- Evaluation setup: Table 1 reports GT-calibrated means across On-screen, Off-screen, and Mixed subsets, broken down by four dimensions and three audio types.Final aggregates sum applicable dimension totals, while AV Coherence is unavailable for Off-screen evaluation.
- Dimension-wise results: Seedance 2.0 exceeds ground truth in Audio Quality and Prompt Following across all subsets, but On-screen AV Coherence remains 20.91 versus the ground-truth 20.93.Kling-v3-omni and Veo 3.1 show the same internal weakness, with lower On-screen AV Coherence than Audio Quality.
- Audio-type results: Seedance 2.0’s On-screen subtype totals are nearly balanced across Speech, Music, and Sound, but Music AV Coherence remains lower than Sound.Its On-screen totals are 32.62 for Speech, 32.54 for Music, and 32.58 for Sound; Music AV Coherence is 6.46 versus 7.81 for Sound.
- Reference interpretation: The calibrated ground-truth reference is an anchor rather than a hard upper bound, as Seedance 2.0 exceeds its Final score on On-screen, Off-screen, and Mixed.The corresponding Seedance 2.0 and ground-truth values are 78.11 versus 71.10 on Off-screen and 99.88 versus 93.69 on Mixed.
- Human alignment: Mean agreement between the automated evaluator and human raters exceeds 70% across dimensions, with PF highest at 77.6% and AV lowest at 67.9%.Agreement is higher for clearly separated model pairs at 76.7%–77.7% and lower for the closer LTX-2 versus Ovi comparison at 63.4%.
5 Conclusion
PRISM-Bench provides an audio-centric, visibility-aware diagnostic framework for T2AV generation. Experiments show that audio quality has improved substantially, while audiovisual grounding—especially on-screen music—remains difficult.
- Benchmark contribution: PRISM-Bench organizes evaluation by Speech, Music, and Sound and by On-screen versus Off-screen source visibility.It uses around 900 human-verified samples and a side-by-side MLLM-as-a-Judge protocol.
- Main conclusion: Experiments show uneven progress across dimensions and audio conditions, with substantial perceptual audio-quality gains but persistent audiovisual-grounding challenges.The conclusion identifies on-screen music as an especially difficult case.
A Ground-Truth Reliability Analysis
The ground-truth row is a context-averaged descriptive reference rather than a fixed upper bound. Reliability analysis therefore asks whether pairing-induced reference variation is small relative to differences among candidate models.
- Ground-truth interpretation: The single ground-truth row should be interpreted as a context-averaged descriptive reference because the same clip is judged repeatedly across candidate pairings.Some reference-side fluctuation across pairing contexts is expected.
- Reliability criterion: The reliability question is whether ground-truth fluctuation remains small relative to real performance differences among candidate models.The analysis uses GT Std, d_bias, and d_s as reliability measures.
A.1 Definitions and Subset-Level Summary
The evaluation quantifies contextual drift and inter-model separation across On-screen, Off-screen, and Mixed subsets. Model signal exceeds contextual bias in every subset, with On-screen showing the greatest variability.
- Definitions: The framework defines GT-induced contextual drift and inter-model performance gaps using standardized effect sizes with pooled score variance.These quantities support subset-level comparison of evaluator bias and model separation.
- Subset-level results: 0.521, 0.641, and 0.814 are the average model signals for On-screen, Off-screen, and Mixed, respectively.The values indicate at least medium model-level separation across all three subsets.
- Subset-level results: Off-screen combines d_bias = 0.152 with d_signal = 0.641, while Mixed combines d_bias = 0.233 with d_signal = 0.814.Both subsets show relatively small contextual drift compared with model-level separation.
- Subset-level results: On-screen is the most variable subset, Off-screen the most stable, and Mixed intermediate.The ordering is interpreted through raw ground-truth fluctuations, which indicate dispersion rather than the benchmark’s sole decision criterion.
A.2 Dimension-Wise Stability Breakdown
Dimension-wise instability is concentrated in synchronization-sensitive Audio-Visual Coherence rather than distributed uniformly across the rubric. On-screen speech and music are especially context-sensitive, while Off-screen dimensions are generally more stable.
- Synchronization-sensitive dimensions: On-screen speech and music Audio-Visual Coherence have the highest reported standard deviations, at 2.145 and 2.314.These dimensions require resolving lip-sync, performer–sound consistency, or event-level onset alignment.
- Subset comparison: Off-screen removes Audio-Visual Coherence and has several remaining entries below 1.0, including speech Audio Quality at 0.801 and music Audio Quality at 0.889.Off-screen Audio Expressiveness for speech is also 0.881.
- Interpretation: Higher On-screen fluctuation is driven primarily by synchronization-sensitive dimensions rather than uniform instability across the rubric.The instability is concentrated in visible grounding and synchronization components.
A.3 Interpretation of Bias, Signal, and SNR
Benchmark validity depends on limited contextual drift alongside at least moderate model-level separation. Off-screen and Mixed provide the clearest margins, while On-screen remains informative despite harder synchronization conditions.
- Validity criterion: The benchmark’s main validity criterion requires contextual drift to remain limited while model-level separation remains at least moderate.Under this criterion, PRISM-Bench remains reliable across all three subsets.
- Subset interpretation: Off-screen and Mixed provide the clearest validity margins because they combine relatively small d_bias with strong d_signal.On-screen is more conservative because its elevated bias is concentrated in the visible grounding branch.
- SNR interpretation: On-screen has average SNR = 1.85 but average d_signal = 0.521, so SNR is not treated as the sole decision criterion.The ratio can be depressed by harder contextual conditions or near-tie comparisons while meaningful model separation remains.
B Judge Model Ablation
The judge ablation evaluates whether low contextual bias also preserves sensitivity to meaningful model differences. Gemini-3.1 Pro is preferred because it balances small bias with stronger inter-model separation, whereas overly conservative judges compress score differences.
- Experimental setup: The controlled Mixed-set ablation isolates judge behavior using identical prompts, blind two-candidate inputs, and output schemas.The Mixed subset jointly includes on-screen and off-screen conditions.
- Bias–signal trade-off: Qwen3-Omni has the smallest d_bias but only Eval SNR = 1.11 because its d_signal is similarly small.The result illustrates that minimizing bias alone can reduce discriminative power.
- Caveat: The authors caution that the Qwen3-Omni result may reflect misalignment between the blind pairwise protocol and its training input format or supervision signals.They interpret it as evidence for balancing low bias with sensitivity, not as a general weakness of the model.
- Bias–signal trade-off: Gemini-3.1 Flash and Gemini-3.0 Flash reduce contextual drift to around 0.07–0.08 but also reduce d_signal to 0.273 and 0.367.These judges may compress meaningful score differences when compared models are not extremely far apart.
- Final selection: Gemini-3.1 Pro achieves d_bias = 0.141 and the largest tested average inter-model separation, d_signal = 0.429.This balance is preferred over minimizing each reliability statistic independently.
- Final selection: Under Gemini-3.1 Pro, the strongest model contrast reaches d_signal = 0.639 while corresponding average d_bias remains well below it.The pattern combines non-compressed model separation with small contextual drift.
C Evaluation Protocol Details
PRISM-Bench uses a shared blind, two-stage judging scaffold while routing active evaluation dimensions according to sound-source visibility. The protocol merges stage-specific scores into one auditable record for each candidate.
- Shared judging scaffold: All subsets use a common blind side-by-side scaffold with shared instructions, preprocessing, scoring scale, and response schema.The scaffold is instantiated through two stage-specific prompt slots.
- Visibility-aware routing: On-screen evaluation scores AV Coherence and Audio Quality in Stage A, then Expressiveness and Prompt Following with caption conditioning in Stage B.The protocol treats visible-source synchronization and source-grounded correspondence as valid targets for this subset.
- Visibility-aware routing: Mixed evaluation routes AV Coherence only to on-screen audio, while Audio Quality covers all active audio types regardless of visibility.Stage B evaluates Expressiveness and Prompt Following under the same visibility-aware request.
- Input and identity handling: Each judge call presents two blinded candidate videos separately, using a 0.5-second identifier frame only to recover candidate identity.The identifier frame is explicitly ignored during scoring.
- Score aggregation: The final record deterministically merges AV Coherence and Audio Quality from Stage A with Expressiveness and Prompt Following from Stage B.The system verifies that all required dimensions are present and does not perform a third scoring pass.
- Score aggregation: Figure 7 illustrates one sample from caption and activated tags through stage-wise score objects to the final merged record.The case study is intended to make the evaluation outcome directly auditable.
D Human Alignment
The human-alignment study mirrors the automated protocol with blinded ground-truth-versus-model comparisons and staged scoring across the four evaluation dimensions. Sample analyses show that semantic role assignment can remain correct even when synthetic audio, missing ambience, or lip-sync failures reduce quality and grounding.
- Human annotation protocol: Three human annotators independently assign absolute integer scores from 0–10, with majority vote producing each final human label.Rare unanimous disagreement is resolved through brief reconciliation.
- Observed alignment cases: On-screen and off-screen speech can remain semantically well grounded while speech sounds synthetic and the off-screen ambience misses continuous insect chirping.The sample preserves role assignment and broad mouth-motion alignment, but audio realism and ambient completeness remain weaker.
- Observed alignment cases: Other samples show that severe on-screen lip-sync mismatch, truncated off-screen speech, and sparse off-screen sound jointly weaken coherence and prompt recovery.These failures affect both source grounding and the completeness of the intended audio scene.
- Study design: The alignment study spans Sora 2, LTX-2, and Ovi, selected to cover different model families and performance levels.This design examines alignment across a broader range of generation quality than a single performance tier.
- Human annotation protocol: Human raters mirror the automated two-stage design: Stage 1 scores AV Coherence and Audio Quality, while Stage 2 scores Expressiveness and Prompt Following after reading the prompt.AV Coherence is restricted to on-screen audio, whereas Audio Quality covers speech, music, and sound regardless of meaning.
- Study design: The interface presents an English prompt and two blinded side-by-side videos while organizing the four dimensions into stage-specific scoring panels.It also records subset and active-tag metadata and provides a discard option for invalid cases.