Source-linked AI summary

VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

Longteng Jiang, DanDan Zheng, Qianqian Qiao, Heng Huang, Huaye Wang, Yihang Bo, Bao Peng, Jingdong Chen, Jun Zhou, Xin Jin

arXiv:2608.25452v1cs.CVcs.AI

TL;DR

Video-generation metrics and benchmarks provide limited fine-grained, human-aligned coverage of aesthetics and offer little supervision for directly improving generators. VGA-BenchV2 expands VGA-Bench with large-scale annotations, hybrid evaluators, and reinforcement-learning optimization, while experiments report human-aligned evaluation across diverse models and a closed-loop improvement framework.

  • Problem

    Existing metrics and benchmarks provide limited fine-grained aesthetic insight, constrained human-labeled supervision, and mainly passive evaluation rather than actionable optimization signals.

  • Method

    VGA-BenchV2 preserves 52 evaluation sub-dimensions and adds human annotations, VAQA-Net, Qwen-based VTag-Net and VGQA-Net, and reinforcement-learning-based generator optimization.

  • Results

    Experiments demonstrate precise, human-aligned evaluation across diverse generation models and facilitate multi-dimensional optimization.

  • Takeaways & Limitations

    VGA-BenchV2 forms a closed-loop framework linking benchmark design, human annotation, automated evaluation, and model improvement.

  • Takeaways & Limitations

    Annotations follow an explicit-trigger principle, so annotators evaluate only dimensions specified by each prompt.

Abstract

from arXiv · show

We introduce VGA-BenchV2, an extended human-aligned benchmark and optimization framework for jointly evaluating and improving video generation quality and aesthetic value. Built upon VGA-Bench, VGA-BenchV2 preserves the original fine-grained taxonomy with two primary dimensions-Aesthetic and Generation-and 52 sub-dimensions. Guided by this taxonomy, we curate 1,016 diverse prompts and collect over 60,000 videos generated by 12 mainstream video generation models. More importantly, VGA-BenchV2 substantially expands human-labeled supervision by adding 36,000 task-level annotations, including 16,200 for aesthetic quality, 13,200 for aesthetic tagging, and 6,600 for generation quality, corresponding to 13.46x, 11.15x, and 1.55x scale-ups over VGA-Bench, respectively. Leveraging this enlarged annotation corpus, we develop a hybrid evaluator architecture consisting of VAQA-Net for continuous aesthetic scoring and two Qwen-based Large Vision-Language Model evaluators, VTag-Net and VGQA-Net, for aesthetic tagging and generation quality assessment. Extensive experiments demonstrate strong alignment with human judgments across diverse generation models. Beyond evaluation, VGA-BenchV2 further introduces an evaluation-to-optimization pipeline, where the learned aesthetic evaluator serves as a reward model for reinforcement learning-based generator fine-tuning. This closes the loop from benchmark construction and human supervision to automated evaluation and model optimization, enabling video generators to improve not only in realism but also in aesthetic quality and human preference alignment. Resources are available at https://huggingface.co/datasets/BestiVictoryLab/VGA-Bench.

1 Introduction

VGA-BenchV2 addresses the limited human alignment and coarse aesthetic coverage of existing video evaluation by expanding supervision and unifying fine-grained assessment with optimization. It combines specialized and vision-language evaluators within a framework spanning benchmark construction, cross-model analysis, and reinforcement-learning-based improvement.

  • Existing metrics capture technical properties but provide limited insight into composition, lighting, color harmony, cinematic expression, and style controllability.
  • VGA-BenchV2 extends VGA-Bench while preserving its two primary dimensions and 52 sub-dimensions for jointly evaluating aesthetic and generation quality.
  • 36,000 newly collected task-level annotations expand supervision across aesthetic quality, aesthetic tagging, and generation quality.The additions include 16,200 aesthetic-quality, 13,200 aesthetic-tagging, and 6,600 generation-quality annotations.
  • The hybrid evaluator combines VAQA-Net for continuous aesthetic scoring with Qwen-based VTag-Net and VGQA-Net for tagging and generation-quality assessment.
  • The framework uses the learned aesthetic evaluator as a reward model for reinforcement-learning-based generator fine-tuning toward higher aesthetic quality and human preference alignment.
  • VGA-BenchV2 unifies 52 dimensions, 1,016 prompts, over 60,000 generated videos, human annotations, automated evaluation, and reinforcement-learning optimization.This supports systematic analysis, fair cross-model comparison, and closed-loop improvement.

2 Related Work

Related benchmarks broaden video-generation evaluation, but VGA-BenchV2 positions itself as a closed-loop extension that connects fine-grained assessment, human supervision, automated evaluation, and model optimization.

  • Earlier metrics mainly assess distributional similarity, temporal consistency, prompt alignment, or low-level artifacts rather than fine-grained aesthetic factors.
  • V-Bench and related benchmarks systematize evaluation across dimensions including temporal coherence, compositional binding, and narrative consistency.
  • VGA-BenchV2 extends VGA-Bench by strengthening human supervision, adding hybrid evaluation, and connecting evaluation with reinforcement-learning-based generator optimization.
  • The resulting framework links benchmark design, human annotation, automated evaluation, and model improvement beyond passive benchmarking.

3 VGA-BenchV2 Construction

VGA-BenchV2 preserves VGA-Bench’s 52-dimension taxonomy while expanding prompts, human supervision, evaluator training, and evaluation-to-optimization capabilities.

  • Inherited Evaluation Taxonomy: The taxonomy contains 52 sub-dimensions organized into Aesthetic and Generation categories, preserving continuity with VGA-Bench.Aesthetic covers continuous quality scores and discrete tags, while Generation covers semantic faithfulness, plausibility, and stability.
  • Prompt Suite and Video Pool: The benchmark uses explicit-attribute prompts so annotations and evaluations focus only on dimensions grounded in each prompt.This design reduces ambiguity and supports consistent comparisons across versions and downstream components.
  • Prompt Suite and Video Pool: The prompt suite contains 1,016 prompts spanning aesthetic quality, aesthetic tagging, and generation quality, with each dimension covered by at least 50 prompts.The suite also supports single-dimension, multidimensional, and lightweight evaluation subsets.
  • Annotation Protocol and Quality Control: VGA-BenchV2 adds 36,000 task-level annotations across aesthetic quality, aesthetic tagging, and generation quality.Annotations use expert exemplars, trained annotators, batch audits, averaging for scores, and majority voting for tags.
  • Hybrid Automated Evaluation Networks: The hybrid evaluator combines VAQA-Net for continuous aesthetic regression with VTag-Net and VGQA-Net for tagging and generation-quality assessment.The design combines specialized aesthetic scoring with LVLM-based semantic reasoning.
  • Evaluation-to-Optimization Interface: VAQA-Net’s aesthetic score can serve as a scalar reward in reinforcement learning, connecting automated evaluation with generator fine-tuning.The reward is intended to guide higher aesthetic quality while maintaining training stability.

4 Experiments and Results

VGA-BenchV2’s evaluators align closely with human judgments across aesthetic, tagging, and generation-quality tasks, while also supporting reward-driven optimization of video generators.

  • Automated Evaluator Validation: 87.6% SROCC: VAQA-Net’s overall aesthetic scores show strong ranking consistency with human experts.VTag-Net also reaches 93.6% on Light Color and 83.7% on Number of Light Sources.
  • Automated Evaluator Validation: 71.3% average accuracy: VGQA-Net aligns consistently with human judgments across 31 fine-grained generation-quality dimensions.It achieves 89.2% on Scene Realism and 85.7% on Abnormal Lighting Detection.
  • VGA-BenchV2 Evaluation: All generated videos are held out from evaluator training, preventing data leakage during cross-model assessment.The evaluation uses VAQA-Net, VTag-Net, and VGQA-Net across diverse aesthetic and quality dimensions.
  • VGA-BenchV2 Evaluation: Models are ranked using average aesthetic score, aesthetic-tag alignment accuracy, and mean generation-quality score across sub-dimensions.Higher values indicate better aesthetic quality, controllability, or generation fidelity and consistency, respectively.
  • Aesthetic Optimization via Reinforcement Learning: 0.49 to 0.52: fine-tuning Wan2.1 increases the average aesthetic score on a held-out test set.The optimization uses VAQA-Net’s Overall Score as the reward with Flow-GRPO, LoRA, and ODE-to-SDE conversion; qualitative results show more appealing composition.

5 Conclusion

VGA-BenchV2 combines an expanded benchmark, human supervision, hybrid evaluators, and reinforcement-learning optimization into a unified framework for video quality and aesthetic assessment.

  • Conclusion: The RL aesthetic optimization pipeline uses VAQA-Net as a reward model to fine-tune video generation models with Flow-GRPO.This connects aesthetic evaluation with generator optimization.
  • Conclusion: Table 5 compares text-to-video models using aesthetic score, tag classification accuracy, and generation level metrics.These metrics cover the benchmark’s aesthetic, controllability, and generation-quality evaluation categories.
  • Conclusion: VGA-BenchV2 preserves 52 sub-dimensions, 1,016 prompts, and over 60,000 videos while adding expanded supervision, hybrid evaluation, and RL-based optimization.The framework is intended to support precise, human-aligned evaluation and multi-dimensional optimization.

Ethical Statement

The benchmark excludes offensive content, and its human annotation procedures received ethical review and Institutional Review Board approval.

  • Ethical Statement: All benchmark prompts and generated videos were screened to exclude pornographic, violent, or otherwise offensive content.The screening applies across the benchmark’s included materials.
  • Ethical Statement: Human annotation procedures comply with ethical guidelines and were reviewed and approved by the relevant Institutional Review Board.The statement covers the benchmark’s human data-collection process.
Loading 2608.25452v1…