Source-linked AI summary
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
Tengfei Liu, Yang Shi, Xuanyu Zhu, Jiafu Tang, Liu Yang, Qixun Wang, Zhuoran Zhang, Yuqi Tang, Fengxiang Wang, Yuhao Dong, Xinlong Chen, Bozhou Li, Bohan Zeng, Yue Ding, Xiaohan Zhang, Jialu Chen, Haotian Wang, Yuanxing Zhang, Pengfei Wan, Leye Wang
TL;DR
Existing benchmarks rarely unify minute-scale audio-visual evaluation across text, image, and video inputs. LongAV-Compass addresses this gap with a 284-case benchmark and diagnostic framework, showing that long-form generation requires multiple capabilities to hold jointly over extended durations.
Problem
Existing benchmarks focus on short-form generation or only part of the task space, limiting unified minute-scale evaluation across text, image, and video inputs.
Method
LongAV-Compass combines taxonomy-guided test construction with event-level annotations and a unified framework assessing over 20 dimensions across three conditioning modalities.
Results
Current systems cannot be characterized by a single overall score; strong performance requires event completion, temporal continuity, visual quality, semantic alignment, and audio-visual synchronization jointly.
Takeaways & Limitations
LongAV-Compass provides a diagnostic basis for analyzing capabilities and failure modes in minute-scale audio-visual generation across diverse input modalities.
Abstract
from arXiv · showhide
Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to short-form settings. Existing benchmarks primarily focus on 5--10 second text-conditioned generation and rarely support unified evaluation across text, image, and video conditioning modalities. Moreover, they provide limited insight into how identity consistency, narrative coherence, and audio-visual alignment degrade over extended temporal horizons. To bridge this gap, we introduce LongAV-Compass, a systematic benchmark for minute-long audio-visual generation. LongAV-Compass contains 284 curated test cases spanning text-to-audio-video (T2AV), image-to-audio-video (I2AV), and video-to-audio-video (V2AV), organized by application scenario and generation complexity. The benchmark combines taxonomy-guided benchmark construction with a unified evaluation framework that integrates MLLM-assisted assessment with complementary perceptual and multimodal metrics, including DINO-v2, ArcFace, CLIP, and ImageBind. The framework evaluates more than 20 fine-grained dimensions covering within-segment quality, cross-segment consistency, global narrative coherence, semantic alignment, and audio-visual synchronization. Through experiments on 11 representative models together with human-alignment validation, LongAV-Compass provides a diagnostic testbed for analyzing the limitations of current systems in sustaining coherent, semantically aligned, and temporally consistent minute-scale audio-visual generation across diverse input modalities.
1. Introduction
LongAV-Compass addresses the lack of unified, minute-scale evaluation for audio-visual generation across text, image, and video inputs. It combines a curated benchmark with hierarchical, multidimensional assessment designed to diagnose long-range quality, consistency, coherence, alignment, and synchronization failures.
- Motivation: Existing benchmarks remain short-form and fragmented across input conditions, limiting evidence about minute-long coherence, scene transitions, and synchronization decay.They typically assess local visual quality or coarse semantic alignment rather than long-range behavior.
- Benchmark: 284 curated test cases unify T2AV, I2AV, and V2AV across application scenarios and generation complexities.The dataset includes 128 T2AV, 115 I2AV, and 41 V2AV examples, covering Vlog, Content-Creator, Performance Ads, and Brand Ads.
- Evaluation framework: The framework evaluates more than 20 dimensions spanning within-segment quality, cross-segment consistency, narrative coherence, audio quality, synchronization, and conditioned semantic alignment.It decomposes long-video assessment into complementary perspectives rather than relying on a flat leaderboard.
- Evaluation framework: An MLLM-centered protocol is complemented by DINO-v2, CLIP, and other multimodal metrics, with human-alignment validation and evaluation of 11 representative systems.The hybrid design supports diagnostics for segment quality, subject consistency, script following, semantic alignment, input anchoring or continuation, and audio-visual synchronization.
2. Related Work
Prior work has developed complementary benchmarks for short-form video, synchronized audio-visual generation, and story-level evaluation, but LongAV-Compass targets minute-long audio-visual generation across T2AV, I2AV, and V2AV with unified diagnosis of long-range consistency and cross-modal alignment.
- Short-form video evaluation: Short-form benchmarks standardize evaluation of visual quality, motion realism, semantic alignment, and prompt following, but primarily target short text-conditioned clips.Examples include VBench, EvalCrafter, and FETV.
- Audio-visual generation evaluation: Audio-visual benchmarks and generation systems extend evaluation beyond visual quality to synchronized, temporally aligned, and semantically aligned sound generation.VABench covers multiple task types, while T2AV-Compass provides a unified protocol for text-to-audio-video systems.
- Story-level evaluation: StoryBench and MSVBench introduce temporally structured assessment of event sequences, hierarchical scripts, story coherence, and cross-shot consistency, but focus on text-conditioned story visualization.These benchmarks represent progress toward long-horizon generative evaluation without targeting minute-long audio-visual generation.
- LongAV-Compass: LongAV-Compass evaluates minute-long audio-visual generation across T2AV, I2AV, and V2AV using taxonomy-guided coverage and a unified framework.Its diagnostic focus includes long-range consistency, event-level continuity, and cross-modal alignment as duration and structure increase.
- LongAV-Compass: The benchmark spans four application scenarios and multiple complexity levels, enabling analysis by content domain and generation difficulty.Complexity levels are denoted L1–L4; S, RI, and RV denote script, reference image, and reference video.
3. LongAV-Compass
LongAV-Compass provides a unified benchmark for minute-scale audio-visual generation across T2AV, I2AV, and V2AV, while preserving modality-specific conditioning requirements. It organizes cases by application scenario and generation complexity and evaluates event-, segment-, full-video-, audio-, synchronization-, and reference-image quality.
- Unified task design: LongAV-Compass unifies evaluation across T2AV, I2AV, and V2AV instead of treating conditioning modalities as independent benchmarks.T2AV uses structured event scripts, I2AV uses reference images and event scripts, and V2AV uses reference videos with continuation scripts.
- Benchmark organization: The benchmark uses a two-dimensional taxonomy spanning four application scenarios—Vlog, Content-Creator, Performance Ads, and Brand Ads—and generation complexity.The scenarios cover creator-oriented content, platform-oriented promotional content, and large-scale brand marketing.
- Task construction: 128 T2AV, 115 I2AV, and 41 V2AV cases cover structured-script generation, reference-image conditioning, and reference-video continuation.The V2AV cases use 10–15 second reference clips followed by 45–50 second continuation scripts.
- Annotation schema: Each case combines a global description with a temporally aligned event sequence, supplemented by modality-specific fields such as identity constraints or continuation descriptions.The global description conditions generation, while the event sequence supports event-level evaluation and fine-grained diagnosis.
- Evaluation framework: Six shared video metrics assess event fulfillment, segment quality, long-range continuity, transition stability, holistic presentation, and text-video alignment across multiple temporal levels.Three audio metrics additionally cover temporal alignment, event-level audio quality, and long-range soundtrack coherence; audio-video synchronization is scored on a 1–5 scale.
- Task-specific evaluation: I2AV evaluation adds first-frame image anchoring (IV1) and image alignment (ImgAlign) to measure reference-image preservation initially and over time.ImgAlign computes CLIP image-image similarity between the reference image and sampled generated frames.
4. Experiments
Experiments evaluate 11 representative systems on minute-long T2AV, I2AV, and V2AV generation using a unified diagnostic framework and fairness-controlled protocols. Results show that proprietary models generally lead, while event fulfillment, continuity, and holistic presentation distinguish long-form quality better than isolated alignment metrics.
- Experimental Setup: 11 representative systems are evaluated across proprietary, open-source, and agent-based categories.The proprietary group includes Seedance 2.0, Kling 3.0, and Veo 3.1; the open-source group includes seven systems, with VideoDirectorGPT as an agent-based baseline.
- Experimental Setup: Outputs target at least 60 seconds, typically 60–120 seconds, while preserving native generation configurations whenever possible.Models without native audio use the shared video-only protocol, whereas models with native audio also undergo audio-visual evaluation.
- Evaluation Framework: The unified framework combines event-aligned segment evaluation, full-video assessment, and task-specific reference checks across complementary diagnostic dimensions.The framework separately measures event fulfillment, segment-level quality, long-range consistency, global presentation, semantic alignment, and audio-visual synchronization rather than collapsing them into one score.
- Main Results: Seedance 2.0 is the most consistent model, while LTX 2.3 leads TVAlign and HunyuanVideo 1.5-I2V leads transition score but trails proprietary leaders on broader capabilities.These results indicate that embedding-level alignment or smooth local transitions alone do not suffice for minute-long script realization.
- Scenario-Level Behavior: Proprietary models maintain clear advantages across Brand Ads, Performance Ads, Content-Creator, and Vlog scenarios.Open-source and agent-based methods show more pronounced weaknesses under scenario-specific requirements.
- Failure Analysis: Performance Ads is the most challenging scenario, with degradation driven primarily by event fulfillment and long-form continuity drops.Recurring failures include missing product operations, broken demonstration sequences, inconsistent causal outcomes, and unstable narrative pacing.
5. Conclusion
LongAV-Compass is introduced as a unified benchmark for minute-scale audio-visual generation across text, image, and video inputs. It combines taxonomy-guided test construction with task-aligned diagnostic evaluation to assess structured long-form scenarios beyond short clips.
- LongAV-Compass unifies minute-scale audio-visual generation evaluation across text, image, and video inputs.
- Taxonomy-guided test construction and task-aligned diagnostic evaluation move assessment beyond short clips toward structured long-form scenarios.
- Current generation systems cannot be adequately characterized by a single overall score.
A. Data Construction Details · A.1. Prompt Design Templates
LongAV-Compass constructs structured long-form scripts through multiple Gemini 3.1 Pro pipelines, with tailored prompts for T2AV, I2AV, and V2AV tracks. These templates encode event structure, audiovisual requirements, identity or consistency constraints, and modality-specific conditioning.
- A. Data Construction Details · A.1. Prompt Design Templates: Multiple construction pipelines use Gemini 3.1 Pro with tailored prompts to build structured long-form benchmark scripts.The paper details prompt designs for each construction track.
- A.1. Prompt Design Templates: T2AV real-video prompts decompose source videos into coherent events containing 2–4 shots, audiovisual descriptions, completion flags, identity tracking, and physical constraints.The structured script is generated from a provided source video.
- A.1. Prompt Design Templates: Across tracks, scripts organize generation around structured events while incorporating modality-specific inputs and explicit audiovisual or consistency requirements.The requirements range from identity tracking and physical constraints to reference-image anchoring and separate video/audio prompts.
- A.1. Prompt Design Templates: T2AV LLM-template generation lets designers specify scenario, L1–L4 complexity, and language before Gemini 3.1 Pro creates scripts with three fixed QA questions per event.The questions test subject presence, core action occurrence, and key visual detail correctness.
- A.1. Prompt Design Templates: I2AV construction first extracts an image prior covering subjects, objects, composition, lighting, motion potential, and consistency constraints, then generates an anchored event script.The extracted prior is fed into a second prompt to preserve visual anchoring to the reference image.
- A.1. Prompt Design Templates: V2AV cases provide a 10–15 s reference video, which Gemini 3.1 Pro uses to generate a continuation script for the remaining 45–50 s.The output separates video and audio prompt fields that concatenate event descriptions for downstream generation models.
A.2. Dataset Statistics
LongAV-Compass comprises 284 samples spanning three tasks, four application scenarios, and four complexity levels. Its splits are balanced by application scenario, dominated by L2–L3 complexity, and structured with differing event and shot distributions across modalities.
- Dataset composition: 284 samples span three tasks, four application scenarios, and four complexity levels (L1–L4).The tasks are T2AV, I2AV, and V2AV.
- Application scenarios: The four application scenarios—Performance Ads, Content-Creator, Brand Ads, and Vlog—are distributed in a balanced manner.
- Complexity levels: L2–L3 dominate across all tasks, while V2AV contains no L1 samples because of video-continuation complexity.
- Event and shot statistics: T2AV and I2AV have similar event counts (avg 6.9–7.0, range 2–18) and shot counts (avg 16.5–17.3), whereas V2AV averages 5.7 events, ranging from 3–7.
B. Evaluation Framework Details · B.1. Evaluation Dimensions and Scoring Rubric · C. Case Studies
The framework evaluates generated audio-video content across nine shared dimensions using decomposed sub-items and a 1–5 MOS rubric, with additional task-specific measures for I2AV. Case studies illustrate annotation quality, generation challenges, and evaluation behavior through representative benchmark examples.
- B.1. Evaluation Dimensions and Scoring Rubric: The framework assesses generated videos along six shared video dimensions and generated audio along three audio dimensions.Each dimension is decomposed into sub-items.
- B.1. Evaluation Dimensions and Scoring Rubric: All MOS dimensions use a 1–5 scale ranging from failed or severe defects to excellent quality.The anchors are 1=failed/- severe defects, 2=poor with major issues, 3=acceptable with visible issues, 4=good with minor issues, and 5=excellent.
- B.1. Evaluation Dimensions and Scoring Rubric: The six video evaluation dimensions and their constituent sub-items are detailed in Table 8.The table provides the framework’s video-dimension specification.
- B.1. Evaluation Dimensions and Scoring Rubric: The three audio evaluation dimensions are listed in Table 9.This table specifies the framework’s audio-dimension coverage.
- B.1. Evaluation Dimensions and Scoring Rubric: For I2AV, First-frame Anchoring uses Gemini 1–5 MOS to assess whether the generated opening preserves the reference image.This task-specific dimension is denoted IV1.
- B.1. Evaluation Dimensions and Scoring Rubric: For I2AV, Image Alignment uses CLIP ViT-L/14 cosine similarity between the reference image and uniformly sampled event frames.Scores use a trimmed mean per event followed by a duration-weighted average.
- C. Case Studies: The case studies present representative benchmark cases illustrating annotation quality, generation challenges, and evaluation behavior.Figures 16, 17, and 18 show keyframe visualizations for the three cases.
C.1. T2AV Case: Performance Ads Product Demonstration (L4) · C.2. V2AV Case: Content-Creator Short Film Continuation (L4) · C.3. I2AV Case: Product Lifestyle Image to Video (L4)
The three L4 cases test minute-scale audio-visual generation across product advertising, narrative continuation, and image-conditioned lifestyle video. Together, they stress multi-actor coordination, narrative and identity continuity, product and environment preservation, and cross-event visual consistency.
- C.1. T2AV Case: Performance Ads Product Demonstration (L4): The T2AV case is a 60-second skincare advertisement with two actors, four events, and rich product-interaction sequences requiring multi-actor coordination and role consistency.Its four events span recommendation, water-burst texture demonstration, product trial, and a final product hero shot with brand reveal.
- C.1. T2AV Case: Performance Ads Product Demonstration (L4): The T2AV case tests distinct-role interaction, facial-expression transitions, texture demonstration, and brand-packaging consistency across shots.The QA checklist also verifies the two women, product handoff, and bright indoor window setting.
- C.2. V2AV Case: Content-Creator Short Film Continuation (L4): The V2AV case continues a reference video of a man running through a train station into a dramatic encounter, fantasy montage, and bittersweet ending.The continuation includes collision and paper scatter, eye contact, a six-subscene romantic montage, the woman boarding, the man remaining alone, and title cards.
- C.2. V2AV Case: Content-Creator Short Film Continuation (L4): The V2AV case requires preserving both subjects’ facial and physical features across scenes, including the montage, while maintaining platform and prop constraints.The man must not board the train, and scattered papers and the open briefcase must remain on the platform at the end.
- C.2. V2AV Case: Content-Creator Short Film Continuation (L4): The V2AV generation challenges combine reference-continuation consistency, rapid montage identity preservation, emotional transitions from surprise to love to loss, and cross-event prop consistency.These demands make the case particularly challenging across style, subjects, emotions, and recurring papers.
- C.3. I2AV Case: Product Lifestyle Image to Video (L4): The I2AV case transforms a product flat-lay photo of an Apple Watch into a 60-second performance advertisement while preserving the watch and desktop environment.The reference contains a silver aluminum watch with a gray-white woven band, notebook, stylus, and iPod on a warmly lit wooden desk.
- C.3. I2AV Case: Product Lifestyle Image to Video (L4): The I2AV script has five events covering notification, wrist placement, health-tracking close-up, watch-face montage, and return to the desk with logo and CTA.Its challenges include preserving reference style and objects, transitioning naturally from still image to video, animating the watch UI, and restoring the original composition.
C.4. Challenging Cases
The challenging cases stress long-horizon generation through dense event sequences, multi-actor emotional narratives, and continuous-motion cinematography. Across these settings, models exhibit identity drift, event collapse, transition artifacts, and product inconsistency, with identity often failing after event 6–7 or 30–40 seconds.
- High Event Count: 18-event nail-art reviews require consistent hands, product colors, and backgrounds across rapid transitions while following a sequential procedure.This T2AV performance-ad case represents the high-event-count challenge.
- Multi-Actor Drama with Emotional Arcs: Multi-actor dramas require persistent costumes and facial features across 13 events, diverse camera angles, and anger, grief, and confrontation.Models typically fail identity preservation after event 6–7 or reduce emotional expression to a neutral default.
- Aerial Cinematography with Continuous Motion: Aerial brand advertisements challenge models to preserve landscape geometry across continuous motion while integrating text overlays and aerial-to-ground transitions.The case contains 15 events and combines drone footage, continuous camera movement, text overlays, and perspective changes.
- Common Failure Patterns: Common failures include identity drift after 30–40 seconds, skipped or merged later events, transition artifacts, and inconsistent branded products.Identity drift particularly affects hair and clothing, while products may change shape, color, or labeling between shots.
D. Generation Protocol Details … D.3. Output Processing
The generation protocol builds task-specific prompts from structured annotations, adapts them to each model’s native conditioning format, and standardizes outputs for evaluation. It supports T2AV, I2AV, and V2AV through modality-specific inputs, continuation instructions, event-aware prompts, and canonical temporal processing.
- D. Generation Protocol Details: Task-specific prompts are constructed from structured annotations while preserving event order, conditioning semantics, and audio expectations.Prompts are adapted to each model’s native format across end-to-end, pipeline, and agent-based systems.
- D.1. Prompt Construction for Generation: T2AV uses the global description, optionally the structured event list, and semicolon-separated event audio expectations for separate audio prompts.The video prompt is the global description, while the audio prompt concatenates event1.audio expectation, event2.audio expectation, and subsequent fields.
- D.1. Prompt Construction for Generation: I2AV combines a reference image and global description, optionally including the full event list with time ranges and modality-specific prompts.The reference image, video prompt, and semicolon-separated audio expectations define the conditioning inputs.
- D.3. Output Processing: Audio-bearing event clips are reextracted from full video.mp4 to ensure audio continuity during evaluation.This processing step is specified for audio evaluation.
- D.1. Prompt Construction for Generation: V2AV provides a 10–15 s reference video clip with continuation prompts for subsequent visual events, audio expectations, and applicable speech.The video and audio instructions explicitly continue the scene and sound after the reference video.
- D.2. Model-Specific Adaptations: End-to-end models receive one full text prompt with native image/video conditioning, pipeline models separate video and audio stages, and agent-based models receive structured event instructions.Pipeline audio modules receive the generated video plus the audio prompt, while agents orchestrate from the event list.
- D.3. Output Processing: Generated videos are saved as full video.mp4 files, segmented at canonical event boundaries, and supplemented with 2 s boundary clips for transition evaluation.Boundary clips span 2 s before and after each event boundary.