Source-linked AI summary
T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu, Zihan Xu, Zhenguo Li, Xihui Liu
TL;DR
Compositional T2V generation lacks dedicated evaluation despite requiring coordinated objects, attributes, actions, motions, and temporal dynamics. The paper introduces T2V-CompBench with category-specific metrics and finds that current models remain highly challenged by compositional generation.
Problem
Existing T2V benchmarks largely neglect the ability to compose multiple objects, attributes, actions, motions, and temporal dynamics in videos.
Method
T2V-CompBench provides seven compositional categories over 1400 prompts, using MLLM-, detection-, and tracking-based metrics validated against human evaluations.
Results
Compositional text-to-video generation is highly challenging for current models across the benchmark’s categories.
Takeaways & Limitations
The benchmark and systematic model analysis establish a basis for studying and improving compositionality in text-to-video generation.
Takeaways & Limitations
The benchmark lacks a unified metric across categories and primarily evaluates videos within 2 to 5 seconds, with fixed-frame sampling potentially insufficient for longer videos.
Abstract
from arXiv · showhide
Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this important ability for evaluation. In this work, we conduct the first systematic study on compositional text-to-video generation. We propose T2V-CompBench, the first benchmark tailored for compositional text-to-video generation. T2V-CompBench encompasses diverse aspects of compositionality, including consistent attribute binding, dynamic attribute binding, spatial relationships, motion binding, action binding, object interactions, and generative numeracy. We further carefully design evaluation metrics of multimodal large language model (MLLM)-based, detection-based, and tracking-based metrics, which can better reflect the compositional text-to-video generation quality of seven proposed categories with 1400 text prompts. The effectiveness of the proposed metrics is verified by correlation with human evaluations. We also benchmark various text-to-video generative models and conduct in-depth analysis across different models and various compositional categories. We find that compositional text-to-video generation is highly challenging for current models, and we hope our attempt could shed light on future research in this direction.
1. Introduction
Compositional text-to-video generation requires combining multiple objects, attributes, actions, and temporal dynamics, but existing T2V benchmarks largely overlook this capability. T2V-CompBench addresses the gap with a seven-category benchmark and category-specific metrics.
- Generating complex, dynamic scenes from fine-grained prompts while accurately composing multiple objects, attributes, and motions remains challenging.
- Existing T2V work often uses simple prompts, while video benchmarks primarily assess video quality rather than compositionality.
- T2V-CompBench contains seven compositional categories with 200 prompts each, covering objects, attributes, quantities, actions, interactions, and spatio-temporal dynamics.
- The benchmark uses MLLM-based metrics for several semantic categories, detection-based metrics for spatial relationships and numeracy, and tracking-based metrics for motion binding.
- The study benchmarks varied T2V models and validates its evaluation metrics by correlating their scores with human evaluations.
2. Related Work
Prior work evaluates T2V quality, alignment, or selected temporal phenomena, while compositional T2V generation remains insufficiently benchmarked. T2V-CompBench extends compositional evaluation from images to videos with temporal-aware metrics and human-correlation validation.
- Compositional T2I benchmarks study attributes, relationships, and complex compositions, whereas video generation additionally requires spatio-temporal reasoning.
- T2V models are commonly evaluated with video-quality and text-video-alignment metrics, but these metrics are unsuitable for complex compositional prompts.
- Existing T2V benchmarks cover broad quality dimensions, controllable attributes, prompt complexity, or time-lapse generation, but mostly use single-object prompts.
- T2V-CompBench is presented as the first benchmark for compositional text-to-video generation, with tailored metrics validated through human correlation studies.
3. Benchmark Construction
The benchmark constructs seven compositional prompt categories spanning spatial and temporal relations, uses real-user vocabulary analysis, and generates verified prompts with controlled complexity and active verbs.
- Prompt Categories: The seven categories cover attribute binding, dynamic attribute changes, spatial relationships, motion, action binding, object interactions, and generative numeracy.
- Prompt Categories: Generative numeracy varies object counts within quantity groups and combines quantities for multiple object types.
- Prompt Categories: Each category contains 200 prompts, with Figure 3 organizing their subgroups and counts.
- Prompt Suite Generation: Vocabulary is grounded in 1.67 million unique VidProM prompts, emphasizing frequent, individually boxable objects to support composition and evaluation.
- Prompt Suite Generation: GPT-4 generates prompts and evaluation metadata; humans verify them, while every prompt includes at least one active verb to discourage static videos.
- Prompt Suite Statistics: Benchmark prompts average 3.6 nouns, 1.4 verbs, and 10.4 words, emphasizing multiple concepts, temporal dynamics, and moderate length.
4. Evaluation Metrics
The evaluation framework adapts metric design to compositional T2V’s spatial and temporal complexity, combining frame sampling with MLLM-, detection-, and tracking-based evaluators.
- Because T2V videos contain many frames and complex dynamics, evaluation samples 6 frames for MLLM metrics, 16 for detection metrics, and 8 FPS for tracking.
- MLLM-based Evaluation Metrics: Grid-LLaVA uses six uniformly sampled frames, chain-of-thought, and disentangled questions to assess semantic video-text alignment while reducing hallucinations.
- Detection-based Evaluation Metrics: Detection-based metrics use GroundingDINO boxes and rule-based geometry to score spatial relationships and object numeracy across frames.
- Tracking-based Evaluation Metrics: Motion-binding evaluation separates foreground and background motion because camera movement can obscure an object’s true direction.
- Tracking-based Evaluation Metrics: GroundingSAM masks foreground and background regions, while DOT tracks points whose average-vector difference estimates the object’s actual movement.
5. Experiments
The experiments evaluate 23 T2V models and compare compositionality metrics against human judgments. Results show strong variation across architectures and categories, with dynamic and relational composition remaining difficult.
- Evaluated Models: 23 T2V models—17 open-source and 6 commercial—are evaluated on T2V-CompBench.
- Evaluation: The study compares conventional metrics with MLLM-, detection-, and tracking-based metrics, then correlates automatic scores with human evaluations.
- Metric Analysis: Grid-LLaVA outperforms LLaVA in action binding and object interactions because it analyzes multiple frames simultaneously.
- Quantitative Results: T2V-Turbo-V2 and CogVideoX-5B lead their respective open-source architecture groups, while PixVerse-V3 performs best overall among commercial models.
- Qualitative Results: Dynamic attribute binding is the most challenging category, followed by spatial relationships, motion binding, and numeracy.
- Qualitative Results: Object interactions, action binding, and consistent attribute binding are easier overall but still produce static interactions, incorrect actions, or misbound attributes.
6. Conclusion
The paper introduces T2V-CompBench as the first systematic benchmark for compositional text-to-video generation. Its metrics are validated against human evaluation, and experiments show that current models still struggle with compositionality.
- T2V-CompBench contains 1400 prompts across seven compositional categories for systematic text-to-video evaluation.
- The benchmark provides category-specific metrics whose effectiveness is validated through correlation with human evaluations.
- Benchmarking models with different architectures shows that compositional text-to-video generation remains highly challenging for current systems.
Appendices
The appendices describe vocabulary construction, prompt composition, word distributions, and metric stability analyses. These analyses support a 1400-prompt evaluation design intended to balance stable scores with computational cost.
- A.1. Vocabulary Construction: WordNet groups real-user prompt words into multi-level noun and verb classes used to select benchmark vocabulary.
- A.1. Vocabulary Construction: The vocabulary includes 260 nouns, 200 verbs, and 100 attributes, with attributes selected from color, shape, texture, and related categories.
- A.1. Vocabulary Construction: LLM-generated prompts use high-frequency nouns, verbs, attributes, and templates specified in the appendix tables.
- A.2. Prompt Analysis: The 1400 prompts concentrate nouns in artifact, object, and person classes, while verbs emphasize travel, movement, and change.
- A.4. Stability of 1400 Prompts and Videos: G-Dino stabilizes at approximately 125 videos for spatial relationships and numeracy, while Grid-LLaVA stabilizes around 150 videos for three categories.
- A.4. Stability of 1400 Prompts and Videos: The benchmark limits each category to 200 videos, producing 1400 prompts and videos to avoid excessive computational cost.
B. Implementation Details
The evaluation follows official default T2V implementations, records generated-video specifications, and uses MLLM procedures designed to reduce hallucinations and improve query reliability.
- T2V models are evaluated using their official default implementations, with resolution, total frames, FPS, and duration reported for generated videos.
- MLLM-based metrics face hallucination risks, including misidentifying visual content or assigning unmatched grades.
- The evaluation first asks the MLLM to describe visual content without revealing the subsequent question, reducing influence from the query.
- Evaluation aspects are split into parallel or sequential queries so grading remains sufficiently differentiated without overwhelming the MLLM.
C.2. How to obtain reliable and reproducible results
The benchmark analyzes score reliability and compositional performance across attribute, spatial, motion, action, and interaction subdimensions using proposed metrics and model comparisons.
- How to obtain reliable and reproducible results: Increasing the MLLM evaluation sample size raises mean and median correlations with human scores and generally stabilizes model rankings.Scores are averaged over randomly sampled experiments, with repeated sampling producing correlations for sample sizes of 2 and 3.
- Consistent Attribute Binding: Color is easiest in consistent attribute binding, texture follows, and shape is most challenging.T2V-Turbo-V2 accurately represents color-object binding in the cited visualization and also shows noticeable object movement.
- Dynamic Attribute Binding: PixVerse-V3 achieves the highest dynamic attribute binding score at 0.0687, with only 9 of 200 videos showing meaningful attribute transitions.Only 31 videos contain relevant elements or transitions, indicating that changing attributes remain difficult for current models.
- Spatial Relationships: For spatial relationships, “Coexist” measures joint object generation, “Acc.” measures relationship accuracy, and “Acc.Score” measures the average score of correct relationships.Both “Acc.” and “Acc.Score” must be high for accurate spatial relationships.
- Spatial Relationships: LVD ranks highest in spatial “Acc.” and “Acc.Score,” while Vico and Mochi perform best among open-source models for generating multiple objects.The results are interpreted as evidence of LVD’s strong layout-planning capability.
- Motion Binding: Gen-3 achieves 42% motion-direction accuracy, while PixVerse-V3 and CogVideoX-5B generate substantial object motion.Motion Level captures object displacement but not movement direction.
- Action Binding: Uncommon action prompts are more challenging than common prompts, especially when they require animals to perform anthropomorphic actions.
- Object Interactions: Physical object interactions are harder than social interactions because they require understanding physical laws.The examples contrast an inaccurate interaction process with one that captures both progression and outcome.
D.7. Generative Numeracy
Generative numeracy declines as prompts request more objects, with commercial models generally outperforming open-source models across quantity groups.
- As specified object quantity increases, the average generative numeracy score tends to decrease.The relationship is shown across quantity groups for videos generated from single-object-class prompts.
- Commercial models generally outperform open-source models, with PixVerse-V3 achieving the highest scores across almost all quantity groups.Among DiT-based models, Open-Sora 1.2 has the best overall performance, while diffusion U-Net-based models show comparable numeracy results.
- Open-Sora 1.2 correctly represents object quantity in an example despite producing a video with an unrealistic style.
E. Human Evaluation
Human evaluations use category-specific AMT interfaces with explicit instructions, examples, rationales, and quality checks; the paper also identifies broader societal risks for future evaluation.
- Human Evaluation: AMT interfaces assess each category’s targeted dimension and present video-text pairs with rating options from 5 to 1.The consistent attribute binding interface focuses annotators on objects and attributes.
- Annotation Instruction: Annotators receive category-specific criteria, examples, score rationales, and concise reminders accompanying each rating option.
- Strategies for Ensuring Quality: Quality control includes interface warnings and random sampling of 20% of each worker’s completed tasks for rejection of evident instruction failures.These measures are intended to improve human-evaluation reliability and accuracy.
- Societal Impacts: The benchmark does not yet evaluate unbiased composition, although the paper identifies risks involving misinformation, deepfakes, stereotypes, and exclusion.The authors plan to incorporate unbiased composition as a future evaluation dimension.
G. Limitations and Future Work
The benchmark’s future work is constrained by the absence of a unified evaluation metric and by its focus on videos lasting 2 to 5 seconds. Fixed-frame sampling may be insufficient for longer videos, whose evaluation is left for future work.
- Limitations: The benchmark lacks a unified evaluation metric that covers all compositional categories.The authors identify this limitation as a challenge for developing better and larger multimodal LLMs or video understanding models.
- Limitations: The benchmark evaluates videos within a 2 to 5 second duration range.This duration scope defines the setting targeted by the current evaluation.
- Future Work: For categories other than motion binding, evaluation samples a fixed number of frames, which may be insufficient for videos longer than 5 seconds.The authors leave evaluation of long videos for future work.
- Future Work: In motion binding, longer videos may produce greater object displacement and better performance, complicating evaluation across video durations.The authors specifically identify long-video evaluation as an open direction.