Source-linked AI summary
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang, Yiting He, Lizhuo Shao, Weiming Zhu, Tengfei Liu, Yang Shi, Jialu Chen, Yuanxing Zhang, Huaxiong Li
TL;DR
Existing benchmarks provide limited insight into whether MR2AV systems can interpret, bind, and compose multiple references coherently. MultiRef-Compass introduces a controlled benchmark and hybrid evaluation framework, and experiments show substantial limitations across multiple evaluation dimensions.
Problem
Existing benchmarks rarely evaluate whether MR2AV systems can jointly interpret, bind, and compose multiple references while maintaining multimodal consistency.
Method
MultiRef-Compass constructs 350 controlled samples and evaluates MR2AV systems with automatic metrics plus rejudging-enhanced MLLM judging across four dimensions.
Results
Experiments on eight representative systems reveal substantial limitations across multiple MR2AV evaluation dimensions.
Takeaways & Limitations
MultiRef-Compass provides a practical diagnostic testbed that distinguishes model capabilities and exposes failure modes absent from existing benchmarks.
Takeaways & Limitations
The benchmark standardizes samples to three reference images, excludes substantially larger reference sets, and its automatic tools have limitations for cartoons and dynamic faces.
Abstract
from arXiv · showhide
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these settings, MR2AV requires models to jointly reason over multiple references while generating synchronized visual and audio content. Models must not only preserve each reference faithfully but also correctly bind and compose multiple referenced entities into coherent audio-visual events. To address this gap, we introduce MultiRef-Compass, a unified benchmark for MR2AV generation. It comprises $350$ carefully curated samples constructed through a scalable and controllable asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. To provide interpretable assessment, MultiRef-Compass defines an evaluation protocol with four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following, using 14 sub-metrics. MultiRef-Compass integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition. Extensive experiments on eight representative MR2AV systems reveal substantial room for improvement across multiple evaluation dimensions, underscoring the need for a comprehensive benchmark and positioning MultiRef-Compass as a foundation for future MR2AV research.
Introduction
MultiRef-Compass addresses the limited evaluation of multi-reference-to-audio-video generation, where existing benchmarks do not jointly assess multiple-reference understanding, binding, multimodal consistency, and composition. It introduces a 350-sample benchmark with controlled asset construction and a hybrid, auditable evaluation framework spanning four dimensions.
- Introduction: Existing text-to-audio-video and reference-to-audio-video benchmarks mainly assess instruction following, audio-visual alignment, or single-reference consistency, leaving multi-reference grounding insufficiently evaluated.Current protocols provide limited insight into whether models can interpret and bind multiple references correctly.
- Introduction: Unified MR2AV evaluation must test cross-reference understanding, correct reference binding, multimodal consistency, and integration of multiple references.The benchmark targets models’ ability to understand, bind, and integrate references while maintaining multimodal consistency.
- Introduction: 350 curated samples form MultiRef-Compass, constructed through a scalable, controllable, taxonomy-driven asset-composition pipeline covering multi-view subject preservation, multi-entity binding, and human-object-scene composition.The pipeline supports controlled and reproducible composition of multimodal assets.
- Introduction: The hybrid framework combines automatic metrics with rejudging-enhanced MLLM-as-a-Judge evaluation across Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following.Its omni-reference schema with metric routing supports image, video, and audio references alongside textual instructions, enabling scalable and auditable assessment.
Related Work
Image-to-video research has progressed from single-image animation and foundational reference control toward richer spatiotemporal guidance and heterogeneous multi-reference conditioning. Existing benchmarks assess general quality, prompt alignment, or single-reference consistency but rarely evaluate multi-view identity versus distinct-entity composition.
- I2V methods initially addressed subject customization, pose control, identity consistency, and localized motion for reference-based animation.
- Later systems added compositional spatiotemporal conditioning, trajectory guidance, and multiple heterogeneous references to specify target generation elements.
- General-purpose benchmarks mainly measure generation quality and prompt alignment, while specialized benchmarks evaluate multiple generation tasks or subject and reference consistency.
- Existing benchmarks rarely distinguish same-identity multi-view references from compositions involving distinct entities.
MultiRef-Compass Dataset
MultiRef-Compass is a structured MR2AV benchmark built by curating and recombining multimodal assets into reference-conditioned samples. It covers controlled multi-view, heterogeneous-composition, and multi-entity scenarios through structured prompts and a 350-sample scale.
- Construction pipeline: The dataset construction pipeline comprises asset-pack curation, board-specific asset composition, and structured prompt generation.Prompts separately describe video, audio, and speech before integration into a unified generation prompt, with sample-level metadata supporting evaluation routing.
- Asset curation: Asset packs cover scenes, videos, voices, subjects, and multimodal references sourced primarily from royalty-free repositories and manually verified for consistency and quality.Subject packs include multi-view identities extracted from videos or 3–6 auxiliary references generated from a primary image, varying viewpoint, expression, and pose.
- Board-specific composition: The main benchmark uses three boards: B1 for same-subject multi-view conditioning, B2 for subject-object-scene compositions, and B3 for multi-entity scenarios across subject packs.These configurations target reference preservation and binding under controlled recombination of assets.
- Dataset scale: 350 samples comprise 300 samples across three main boards and 50 challenge-board samples, each paired with a structured prompt.This scale balances benchmark coverage with evaluation cost for comparisons across proprietary and open-source models.
- Evaluation coverage: The benchmark supports controlled evaluation of multi-reference preservation and binding, fine-grained instruction following, and coordinated audio-visual generation.Its evaluation framework reports diagnostic scores across Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following.
Evaluation Framework
MultiRef-Compass combines automatic metrics with rejudging-enhanced MLLM evaluation across four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following. Results show complementary, modality-specific strengths and no system consistently dominates all dimensions.
- Evaluation Protocol: The framework evaluates every generated video across Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following using automatic metrics and rejudging-enhanced MLLM judgments.MLLM metrics use a 1–5 MOS scale, while rejudging verifies potentially unreliable judgments.
- Evaluation Protocol: A rule-based router activates only applicable metrics, such as lip-sync checks for dialogue or monologue samples, while automatic metrics provide objective evidence when reliable tools exist.The protocol also supports checklist-style and rubric-based MLLM evaluation.
- Basic Quality: Basic Quality measures visual, audio, and anatomical plausibility through DOVER++, Audiobox, and MLLM checklist judgments of humans, objects, and backgrounds.Anatomical Quality additionally evaluates biological constraints, rigid-body rigidity, and 3D depth coherence.
- Reference Consistency: Kling 3.0 achieves the highest EF, whereas Seedance 2.0 leads BC and DP, showing that embedding similarity and fine-grained reference judgments capture complementary capabilities.High embedding-based similarity does not always imply the best fine-grained preservation.
- Audio-Visual Consistency and Instruction Following: Gemini-Omni leads SLS, Seedance 2.0 leads ESM, and HappyHouse 1.1 leads SC, while instruction-following strengths vary across visual, audio, speech, and temporal tasks.Seedance 2.0 is strongest visually, Kling 3.0 in audio and speech compliance, and Wan 2.7 in temporal compliance.
- Overall Findings and Implications: No evaluated system performs consistently strongly across all four dimensions; reference consistency declines with more complex compositions, and anatomical failures remain widespread.The benchmark separates capabilities to expose failures hidden by aggregate quality scores.
Limitations
MultiRef-Compass is a controlled diagnostic benchmark with limited coverage of domains, styles, modalities, and larger reference sets. Its evaluation also faces limitations in automatic evidence tools, synchronization metrics, MLLM judging, and system reproducibility.
- Benchmark scope: The benchmark does not cover every creative domain, cultural style, or reference modality.Its controlled diagnostic design prioritizes comparability over exhaustive coverage.
- Benchmark scope: Each sample uses three reference images, leaving substantially larger reference sets outside the current scope.This standardization supports comparability across models with different conditioning interfaces.
- Evaluation limitations: Automatic evidence tools have limitations for cartoon subjects and faces undergoing large pose or ex...
- Evaluation limitations: More robust synchronization metrics are needed for dynamic faces and complex multi-speaker scenes.MLLM-based judging also incurs additional cost and may inherit model-specific biases.
- Reproducibility: Proprietary systems and APIs may change over time, so model versions and access dates are reported for reproducibility.
Conclusion · Appendix · A MR2AV Definition
MultiRef-Compass is introduced as a comprehensive benchmark for multi-reference-to-audio-video generation, combining controllable sample construction with interpretable evaluation across four dimensions. MR2AV is defined as reference-conditioned synthesis that jointly interprets multimodal references and composes their information rather than directly editing source content.
- Conclusion: MultiRef-Compass evaluates multi-reference-to-audio-video generation with a unified benchmark.It targets generation conditioned on multiple references and textual instructions.
- Conclusion: An asset-reuse pipeline constructs diverse multi-reference evaluation samples.The pipeline supports benchmark sample construction at scale.
- Conclusion: The evaluation combines automatic metrics with a rejudging-enhanced MLLM-as-a-Judge protocol.This combination is used to assess MR2AV systems.
- Conclusion: The benchmark evaluates systems along four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following.These dimensions define the benchmark’s evaluation scope.
- A MR2AV Definition: MR2AV synthesizes new audio-video content from multiple references and a textual instruction.References may include images, videos, and audio clips whose information should be preserved or reflected in the generated content.
- A MR2AV Definition: Multi-reference generation requires jointly interpreting and composing information from multiple reference sources.The references may depict different views of the content to be generated.
- A MR2AV Definition: MR2AV is a reference-conditioned generation task rather than an editing task.Reference videos can provide subject identity, appearance, or motion cues, while reference audio can provide voice identity, timbre, or acoustic characteristics.
- A MR2AV Definition: References provide information to preserve or reflect, rather than source content to be directly edited.This distinction applies to both reference videos and reference audio clips.
B Dataset Construction Details … B.4 Prompt Schema and Sample Records
MultiRef-Compass is constructed through reusable, metadata-guided asset packs and board-specific recombination to support realistic, controlled, and diagnostic MR2AV evaluation. Its boards target subject preservation, heterogeneous composition, multi-entity binding, and optional audio-video reference challenges, with prompts generated through a structured, reviewed schema.
- B Dataset Construction Details: The construction pipeline combines realistic multi-reference R2AV workflows with controlled, reproducible, and diagnostic evaluation through asset collection, board construction, and structured prompt generation.The benchmark is designed to remain practical while supporting reproducible assessment.
- B.1 Overview and Design Rationale: Reusable subject, object, scene, video, and audio packs are recombined under board-specific rules, reducing redundant annotation and clarifying sources of difficulty.Each pack is verified once for quality, identity consistency, and usage constraints.
- B.2 Asset-Pack Collection: Subject packs contain multiple same-identity images collected from Pexels videos or augmented from a canonical image with Nano Banana.The provided passage describes complementary construction paths for real-world subjects and single-image cases.
- B.2 Asset-Pack Collection: Object, scene, audio, and video packs provide heterogeneous references spanning entities, environments, voice timbres, and dynamic subject information.Metadata records category, style, attributes, and usage constraints to automatically match compatible packs.
- B.3 Board Construction: Board 1 tests whether multiple views of one identity become a coherent dynamic subject without generating duplicate subjects.References may vary in viewpoint, expression, pose, or speaking state.
- B.3 Board Construction: Board 2 evaluates preservation and coherent placement of subject, object, and scene references across heterogeneous compositions without explicit entity-role binding.It checks for ignored objects, altered scenes, or dominance by one reference type.
- B.3 Board Construction: Board 3 assigns distinct roles to multiple sampled entities and directly tests binding through actions, object interactions, speech, and spatial placement.Failures include identity swaps, entity merging, and attribute transfer.
- B.4 Prompt Schema and Sample Records: Prompts use difficulty-dependent fields for visual grounding, temporal actions, audio events, speech, and integrated instructions, then undergo LLM drafting and manual refinement.Review removes unreasonable requirements, inconsistent reference use, and unsafe violent or sexual scenarios.
C Board4 Analysis · D Case Analyses
Board 4 extends the benchmark with audio-video reference inputs, creating more complex multimodal interactions while broadly preserving earlier trends. Case-level and board-level analyses further illustrate protocol behavior and rejudging corrections.
- C Board4 Analysis: Board 4 adds audio-video reference inputs to the core sample types, increasing multimodal interaction challenges beyond Boards 1–3.Its evaluation covers the same broad benchmark structure while introducing audio-video reference conditioning.
- C Board4 Analysis: Seedance 2.0 achieves higher visual technical quality and clearly stronger anatomical quality, whereas Kling 3.0 obtains slightly higher audio technical quality.The result suggests stronger structural stability for Seedance 2.0 and higher audio technical quality for Kling 3.0.
- C Board4 Analysis: The structured prompt schema specifies fields activated by board and sample requirements, defining how MultiRef-Compass configurations are organized.Table 5 documents the schema used to activate requirements across boards.
- C Board4 Analysis: The additional Board 4 leaderboard reports automatic and MLLM-based metrics, including TS for voice timbre similarity between generated speech and the provided audio reference.This supplements the main-paper result tables with selected-model comparisons.
- C Board4 Analysis: Seedance 2.0 is slightly stronger on visual and audio task following, while Kling 3.0 performs better on speech content accuracy and temporal order following.These results show that task-following strengths differ across visual, audio, speech-content, and temporal-order dimensions.
- C Board4 Analysis: Overall, Seedance 2.0 is more balanced across reference preservation, structural quality, and audio-reference use, while Kling 3.0 remains competitive on selected instruction and source-attribution dimensions.Board 4 therefore preserves the main benchmark’s broad performance pattern.
- D Case Analyses: The case analyses provide representative metric- and board-level views of protocol behavior under different board settings.They complement metric construction details and examine how rejudging corrects inconsistent checklist judgments.
D.1 Entity Fidelity Analysis
Entity Fidelity declines as reference complexity increases, with the largest degradation occurring in human consistency. Stronger systems remain comparatively stable across boards, but preserving interacting human entities remains the dominant challenge.
- Entity Fidelity across boards: Board 1 achieves the highest average Entity Fidelity score at 0.6930, compared with 0.5893 on Board 2 and 0.5748 on Board 3.The decline is mainly driven by the Human branch, whose average EF decreases from 0.6930 to 0.5522 and 0.5409 across the three boards.
- Branch-level analysis: Human consistency drives the cross-board decline, while Object and Background scores remain relatively stable around the reported level.As multi-reference composition and entity interaction become more complex, maintaining human consistency becomes the dominant challenge for current R2AV systems.
- Model-level analysis: Kling and Seedance remain relatively stable across boards and rank among the top models on all three board totals.Other models show branch-level weaknesses involving human consistency, object preservation, or background preservation.
D.2 Speech-Lip Synchronization Analysis
The analysis shows that speech-lip synchronization evaluation becomes less stable in complex generation settings, while speech-content accuracy is broadly balanced across English and Chinese but varies substantially by model and language. SCA therefore assesses language, lexical content, dialogue assignment, and utterance structure beyond speech synthesis and synchronization alone.
- Speech-Lip Synchronization: 84 valid videos remain for Gemini-Omni, compared with 47 for Kling and Seedance 2.0 under the filtering procedure.The pre-filter removes large head motion, invisible speaking faces, off-screen speech, and insufficient mouth visibility; residual motion can still affect SyncNet-based scores.
- Speech-Lip Synchronization: More complex generation settings retain fewer valid samples, making VSLS scores more sensitive to sample selection.Board 1 retains the most valid samples after stable-frontal and temporal-offset filtering, whereas Boards 2 and 3 retain substantially fewer across models.
- Speech Content Accuracy: 87.8% English strict accuracy and 88.0% Chinese strict accuracy indicate nearly identical aggregate speech-content accuracy across languages.Model-level results nevertheless reveal clear language-specific behaviors.
- Speech Content Accuracy: Kling reaches 90.6% English and 97.6% Chinese accuracy, HappyHouse 1.1 reaches 89.9% and 100.0%, and Seedance 2.0 reaches 93.5% and 91.2%, respectively.Remaining errors include extra words, repetitions, paraphrases, omissions, language-control errors, and content-control errors; Gemini-Omni has more Chinese-side errors, including English substitution and semantic reversal.
- Speech Content Accuracy: SCA evaluates requested language, lexical content, dialogue assignment, and prompt-specified utterance structure, so it should be assessed separately from synchronization and audio quality.A model may produce fluent, synchronized speech while failing to reproduce the required spoken content.
D.3 Rejudging Cases Analysis
Rejudging corrects overly literal initial judgments by recognizing semantically valid actions that satisfy checklist intent, even when the exact wording is not explicit.
- Representative Rejudging Cases: The rejudge stage corrects failures where picking up, carrying, or placing @Image2 changes its position despite no explicit “adjusting” action being identified.This addresses overly strict initial judgments for the checklist question “Does someone adjust the position of @Image2?”
D.4 Stability of MLLM-Dependent Evaluation
Repeated MLLM-dependent evaluation on 60 sampled examples shows stable outcomes, with relative standard deviations below 1% across all metrics and consistently small item-level MAEs.
- Evaluation Protocol: Five repeated evaluations of 60 randomly sampled examples used identical generated outputs and evaluation settings, including MLLM-assisted EF and SLS processing.EF applies MLLM-estimated paste-naturalness calibration, while SLS uses MLLM-based pre-screening for reliable lip-sync measurement.
- Stability Results: Below 1% relative standard deviations were observed across all metrics, alongside consistently small item-level MAEs.The results indicate that evaluation outcomes are not artifacts of judge stochasticity.
- Stability Results: Repeated evaluation preserved reliable overall model comparisons under the same evaluation pipeline.This conclusion follows from the low metric variability and small item-level errors.
E Detailed Metric Constructions and Prompts
This section specifies how MultiRef-Compass constructs its evaluation metrics, including checklist generation, MLLM-as-a-Judge prompts, and automatic-metric implementation details. The instruction-following prompts operationalize audio, speech, and temporal evaluation through fixed, reusable checklists.
- E Detailed Metric Constructions and Prompts: The section details metric-specific checklists, MLLM-as-a-Judge prompts, and automatic-metric implementation details.Prompt boxes show core evaluation components while omitting output-format and serialization specifications.
- E Detailed Metric Constructions and Prompts: Prompt boxes present core evaluation-prompt components while omitting output-format specifications and implementation-specific serialization details.
- E.4 Instruction Following: The checklist-generation prompt converts each video-generation prompt into a fixed checklist for Audio Task Following, Speech Content Accuracy, and Temporal Order Following.The checklist is reused to evaluate generated videos from different models and is designed to remain independent of any individual generation.
- E.4 Instruction Following: The audio and speech checklist extracts the expected speaker, exact quoted text, and language, using zh, en, or unknown labels.Vague requirements such as “speaks” or “talks” do not generate checklist items when no exact text is provided; absent exact text is marked not applicable.
- E.4 Instruction Following: The evaluation checklist is designed for consistent comparison across multiple generated videos from different models.Its reusable structure supports instruction-following assessment independently of any single generated video or model.
- E.4 Instruction Following: Temporal Order Following evaluates explicit temporal structure by checking shot or scene segmentation, required shots or scenes, and their correct order.These questions are generated when the prompt explicitly describes multiple shots or scenes.