Source-linked AI summary

VBench: Comprehensive Benchmark Suite for Video Generative Models

Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, Ziwei Liu

arXiv:2311.17982v1cs.CV

TL;DR

Video-generation evaluation remains difficult because existing metrics may not match human judgment and do not fully reveal model-specific strengths and weaknesses. VBench addresses this gap with a hierarchical, disentangled benchmark spanning 16 dimensions, human-preference validation, and analyses across abilities and content types. The authors report close alignment with human perception and position the benchmark as a source of detailed evaluation insights.

  • Problem

    Existing video-generation metrics are inconsistent with human judgment, while real-video quality-assessment methods overlook artifacts and other challenges specific to generated videos.

  • Method

    VBench evaluates video generation through 16 hierarchical, disentangled dimensions organized under Video Quality and Video-Condition Consistency, supported by tailored methods and human preference annotations.

  • Results

    VBench evaluations highly correlate with human preferences across every fine-grained evaluation dimension and provide insights into model abilities across content types.

  • Takeaways & Limitations

    VBench offers detailed feedback on model strengths and weaknesses that can inform evaluation, training, and architectural or data choices for video generation.

  • Takeaways & Limitations

    VBench currently covers limited open-sourced text-to-video models, does not yet extend to additional tasks such as image-to-video, and does not assess safety or equality dimensions.

Abstract

from arXiv · show

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align with human perceptions; 2) An ideal evaluation system should provide insights to inform future developments of video generation. To this end, we present VBench, a comprehensive benchmark suite that dissects "video generation quality" into specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods. VBench has three appealing properties: 1) Comprehensive Dimensions: VBench comprises 16 dimensions in video generation (e.g., subject identity inconsistency, motion smoothness, temporal flickering, and spatial relationship, etc). The evaluation metrics with fine-grained levels reveal individual models' strengths and weaknesses. 2) Human Alignment: We also provide a dataset of human preference annotations to validate our benchmarks' alignment with human perception, for each evaluation dimension respectively. 3) Valuable Insights: We look into current models' ability across various evaluation dimensions, and various content types. We also investigate the gaps between video and image generation models. We will open-source VBench, including all prompts, evaluation methods, generated videos, and human preference annotations, and also include more video generation models in VBench to drive forward the field of video generation.

1. Introduction

VBench addresses the need for video-generation evaluation that reflects human perception while exposing models’ specific strengths and weaknesses. It decomposes quality into detailed dimensions, validates alignment with human preferences, and provides insights across abilities and content types.

  • Existing video-generation metrics can conflict with human judgment, motivating evaluation methods tailored to synthesized-video challenges.
  • VBench decomposes video-generation quality hierarchically and disentangled into dimensions that isolate individual aspects of quality.The framework separates Video Quality from Video-Condition Consistency and further subdivides each into granular criteria.
  • Human preference annotations show that VBench’s fine-grained evaluation methods closely align with human perception.The annotations can also support instruction tuning of evaluation models.
  • VBench provides detailed feedback on model strengths and weaknesses across ability aspects, informing video-generation analysis and development.
  • VBench is open-sourced with its evaluation dimensions, methods, prompts, generated videos, and human preference annotations.

2. Related Works

Related work builds on rapid progress in diffusion-based image synthesis and extends it to video generation. Video models now use multiple guidance modalities beyond text.

  • Diffusion models have achieved significant progress in image synthesis and enabled many video-generation approaches.
  • Many recent diffusion-based video-generation works are text-to-video models.
  • Video-generation systems also support image-to-video, video-to-video, and control-map guidance such as pose, depth, and sketch.

3. VBench Suite

VBench decomposes video generation quality into 16 disentangled dimensions, organized around video quality and video-condition consistency. It pairs dimension-specific evaluation methods with curated prompts and human preference annotations.

  • Evaluation Dimension Suite: VBench evaluates video generation through 16 disentangled dimensions grouped under Video Quality and Video-Condition Consistency.Video Quality asks whether the video looks good independently of the prompt, while Video-Condition Consistency asks whether it matches the requested condition.
  • Video Quality: Video Quality separates temporal quality from frame-wise quality, covering consistency, dynamics, aesthetics, and imaging distortion.Temporal dimensions include subject and background consistency, temporal flickering, motion smoothness, and dynamic degree; frame-wise dimensions include aesthetic and imaging quality.
  • Video-Condition Consistency: Video-Condition Consistency separates semantics from style and evaluates object classes, multiple objects, human actions, colors, and spatial relationships.The suite uses specialized evaluators such as GRiT and UMT for several semantic criteria.
  • Prompt Suite: Around 100 prompts are designed for each dimension, with content-specific construction intended to test the targeted ability efficiently and comprehensively.The prompt suite addresses the high cost of video generation while preserving diversity across evaluation dimensions and content categories.
  • Human Preference Annotation: The benchmark includes human preference annotation protocols that compare videos from the same prompt while instructing annotators to judge only the specified dimension.Annotators choose whether video A, video B, or neither is better, and the study applies training and quality-assurance checks.
  • Evaluation Results: VBench reports per-dimension scores for video models alongside Empirical Min and Max reference baselines.The table compares four video generation models across all 16 dimensions and uses higher scores to indicate relatively better performance.

4. Experiments

Experiments apply VBench to multiple text-to-video models, reference baselines, content categories, and text-to-image models. The evaluations test human alignment and compare frame-wise generation capabilities across modalities.

  • Model Evaluation: LaVie, ModelScope, VideoCrafter, and CogVideo are evaluated with VBench, alongside Empirical Max, Empirical Min, and WebVid-Avg baselines.The baselines approximate achievable score bounds or reflect WebVid-10M dataset quality for each dimension.
  • Human Alignment: VBench computes model win ratios from pairwise human comparisons and from per-dimension benchmark results, then correlates the two.A win ratio is the total score divided by the number of pairwise comparisons, with ties contributing 0.5 to both models.
  • Human Alignment: VBench evaluations are highly correlated with human preference annotations across individual evaluation dimensions.Figure 5 plots human-preference win ratios against VBench win ratios and reports Spearman correlation coefficients for each dimension.
  • Content Categories: The benchmark evaluates text-to-video models across eight content categories using category-specific prompt suites.Performance is calculated across evaluation dimensions for videos generated from prompts organized by content type.
  • Video–Image Comparison: The study compares frame-wise generation capabilities of text-to-video and text-to-image models using ten VBench dimensions.The comparison includes Stable Diffusion 1.4, Stable Diffusion 2.1, and Stable Diffusion XL alongside video generation models.

5. Insights and Discussions

VBench reveals trade-offs across temporal consistency, dynamic motion, content categories, and dataset composition. Category-specific evaluation exposes model strengths that aggregate scores can obscure.

  • Trade-off across Ability Dimensions: Models often trade temporal consistency against Dynamic Degree, with static scenes attaining stronger consistency while large motions remain difficult.LaVie performs well on Background Consistency and Temporal Flickering but has low Dynamic Degree, whereas VideoCrafter shows the opposite pattern.
  • Evaluation Across Content Categories: Figure 7 presents VBench results across eight content categories using category-specific prompt suites, with scores linearly normalized from 0 to 1.Comprehensive numerical results and normalization details are provided in the supplementary file.
  • Content-Specific Capabilities: CogVideo achieves strong Aesthetic Quality for Food but underperforms for Animal and Vehicles, revealing content-specific capability variation.Average scores can mask strong performance in individual categories.
  • Complex Content Categories: Spatially complex categories, including Animal, LifeStyle, Human, and Vehicles, show relatively poor Aesthetic Quality.The passage attributes this to challenges in harmonious color schemes, articulated structures, and appealing layouts amid complex elements.
  • Data Quality and Quantity: Food almost always receives the highest Aesthetic Quality, despite representing only 11% of WebVid-10M, suggesting data quality may matter more than quantity at million-scale.VBench dimensions may also help clean datasets according to specified quality dimensions.
  • Compositionality: T2I versus T2V: T2V models significantly underperform T2I models in Multiple Objects and Spatial Relationship, highlighting a compositionality gap.The comparison especially identifies SDXL as a stronger T2I reference.

6. Conclusion

The paper presents VBench as a multi-dimensional, human-aligned benchmark for evaluating video generation models. It positions the benchmark as a source of insights for future advances while identifying expansion and safety gaps.

  • Conclusion: VBench is proposed as a comprehensive benchmark suite with multi-dimensional, human-aligned, and insight-rich properties.The conclusion frames it as a contribution to the video generation and evaluation community.
  • Limitations and Future Work: The authors plan to add more models and extend VBench to additional video generation tasks, including image-to-video.These plans define the benchmark’s current scope and future expansion.
  • Potential Negative Societal Impacts: VBench currently omits safety and equality dimensions, so users are urged to exercise caution with open-sourced video generation models.The conclusion identifies ethical considerations as important for future benchmark iterations.

Supplementary Material

The supplementary material expands the benchmark’s methodological, prompting, annotation, experimental, visualization, and societal-impact documentation. It also points readers to a demo illustrating VBench dimensions and video examples.

  • Supplementary Sections: Section G details the Evaluation Dimension Suite and Evaluation Method Suite, while Section H elaborates on Prompt Suite details.These sections provide supplementary methodological and prompting information.
  • Annotations and Experiments: Section I explains Human Preference Annotations, and Section J provides further implementation details on experiments and visualizations.The supplementary file organizes annotation and experiment documentation into separate sections.
  • Impact and Limitations: Section K discusses potential societal impacts, Section L discusses limitations, and Section M provides additional material.The supplementary file includes dedicated sections for impact, limitations, and further information.
  • Demo Video: A demo video illustrates VBench and shows video examples for each evaluation dimension.The demo accompanies the supplementary file.

G.1. Video Quality

VBench evaluates video quality through disentangled measures of temporal consistency, motion, dynamics, aesthetics, and image quality. Its temporal flickering evaluation uses static scenes, while human rankings remain nearly unchanged across dynamic degrees.

  • Temporal Quality: Subject consistency measures whether a video's subject maintains the same appearance across frames using DINO feature similarity.The score averages cosine similarities between frame features and preceding or first-frame features.
  • Temporal Quality: Background consistency evaluates whether the scene remains stable across frames using CLIP image features.The metric parallels subject consistency but focuses on the background scene.
  • Temporal Quality: Temporal flickering is measured on filtered static videos to isolate local high-frequency temporal inconsistencies from motion and other artifacts.Frame-to-frame mean absolute error is normalized so higher scores indicate less flickering and better perceptual quality.
  • Temporal Quality: Around 99% correlation across dynamic, semi-dynamic, and static benchmarks indicates that human temporal-flickering rankings are largely independent of motion strength.The three benchmarks draw videos from the Subject Consistency, Background Consistency, and Temporal Flickering prompt suites.
  • Motion and Dynamics: Motion smoothness assesses whether generated movement is physically smooth using motion priors from video frame interpolation models.Dynamic degree distinguishes videos with obvious camera or object motion from nearly unchanged videos.

J.1. Video Generation Models in Evaluation

The evaluation compares video-generation models using VBench across categories, dimensions, and reference baselines, while also illustrating human alignment and VLM-tuning applications. Supplementary visualizations include model examples, category-specific results, and video–image comparisons.

  • Video Generation Models: The evaluation adopts four video-generation models and samples videos under each model's stated inference settings.The supplied examples include CogVideo and LaVie, with model-specific frame, resolution, FPS, and sampling configurations.
  • VLM Tuning: Human preference annotations are used to fine-tune a VLM for evaluating video-generation capabilities in specific dimensions.The tuned VLM selects relevant metrics, describes videos, and predicts scores within those metrics.
  • Category Evaluation: VBench evaluates models across eight content categories using category-specific prompt suites and reports performance across evaluation dimensions.Category-level charts show results for different models within the same content category.
  • Human Alignment: VBench evaluations across all dimensions closely match human perceptions, as shown by corresponding VBench and human win ratios.Table A5 reports these win ratios for each dimension and model.
  • Reference Baselines: The study uses Empirical Max, Empirical Min, and WebVid-Avg baselines to contextualize attainable, minimum, and average reference scores.Empirical references are approximated from WebVid-10M videos for dimensions where ideal scores are difficult to achieve.
  • Visualization: Radar charts normalize scores for relative visualization by mapping selected maximum and minimum model scores to fixed axis values.The stated mappings use 0.8 for maxima and 0.3 for minima in the main comparison charts.

L. Limitations and Future Work

VBench initially focuses on text-to-video evaluation and is limited by the currently small number of open-sourced T2V models. Future work will broaden model participation and support additional controlled video-generation tasks.

  • Open-Sourced Models: The limited number of open-sourced T2V models currently restricts the breadth of available evaluations and annotated generation results.The authors plan to open-source VBench and encourage more T2V models to participate.
  • Task Scope: VBench is initially built for text-to-video, while future extensions will accommodate video-driven, image-driven, personalized, and other multimodal-controlled generation.Existing Video Quality dimensions are readily applicable, while Video-Condition Consistency dimensions are planned for extension.

M. Additional Experimental Results

Additional experiments provide numerical and distributional views of VBench evaluations across video and image models, human alignment, and WebVid-10M content categories. These results support comparisons across categories and modalities.

  • Video–Image Comparison: Table A8 compares four video-generation models with three image-generation models across VBench dimensions.Overall Consistency uses CLIP instead of ViCLIP to enable image-model evaluation.
  • Human Alignment: Table A5 reports VBench and human win ratios for each dimension and model to assess human alignment.The caption states that evaluations across all dimensions closely match human perceptions.
  • WebVid-10M Categories: Figures A27 and A28 show the category distribution and aesthetic scores of the eight WebVid-10M content categories.These visualizations support observations about category composition and category-specific aesthetics.
Loading 2311.17982v1…