Source-linked AI summary
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
Songlin Yang, Haobin Zhong, Ruilin Zhang, Xiaotong Zhao, Shuai Li, Kai Zheng, Xuyi Yang, Zhe Wang, Zhenchen Tang, Yang Li, Bohai Gu, Zhengwei Peng, Yidan Huang, Mengzhou Luo, Yihang Bo, Dalu Feng, Yujia Zhang, Juntao Ma, Ruiqi Wang, Lvmin Zhang, Yuwei Guo, Frank Guan, Maneesh Agrawala, Hongbo Fu, Alan Zhao, Anyi Rao
TL;DR
Professional cinematic video generation lacks evaluation that captures cinematic quality beyond prompt correctness. EvalVerse provides a pipeline-aware, expert-calibrated evaluator and reports strong alignment with human experts across advanced cinematic dimensions.
Problem
Existing video benchmarks emphasize prompt correctness while lacking scalable, domain-specific assessment of professional cinematic quality and aesthetics.
Method
EvalVerse combines a filmmaking-workflow taxonomy with expert-calibrated VLM fine-tuning that produces step-by-step Chain-of-Thought reasoning before scoring.
Results
EvalVerse predictions show strong, robust alignment with expert annotations across cinematic dimensions, with abstract and temporally entangled dimensions achieving the highest agreement.
Takeaways & Limitations
EvalVerse provides trustworthy diagnostic signals and expert-aligned reward vectors for reward modeling and agentic video-generation workflows.
Takeaways & Limitations
Current VLMs process discrete keyframes rather than continuous streams, limiting temporal perception and evaluation of long-form narratives and diverse avant-garde styles.
Abstract
from arXiv · showhide
The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community transitions towards Reinforcement Learning (RL) and agentic workflows. However, reliable evaluation has emerged as a critical bottleneck. Existing benchmarks predominantly evaluate ''whether it is right'' (basic prompt-following) while fundamentally neglecting ''whether it is good'' (cinematic quality, acting, and aesthetics). Furthermore, current automated metrics lack the domain-specific rigor required to provide trustworthy signals, creating a severe credibility gap between human aesthetic perception and machine scoring. To bridge this gap, we introduce EvalVerse, a comprehensive, pipeline-aware, and expert-calibrated evaluation framework. We treat video generation assessment not merely as an engineering task, but as a core scientific problem: the systematic digitization of subjective cinematic expertise. First, we organize domain knowledge into an evaluation taxonomy aligned with the professional filmmaking workflow (pre-production, production, and post-production). Second, we distill human expert judgments into a curated dataset with large-scale human annotations. Third, we inject this knowledge into Vision-Language Models (VLMs) through an expert-calibrated fine-tuning strategy, enabling the VLM to perform explicit Chain-of-Thought reasoning. Compared to previous works, EvalVerse not only retains compatibility with foundational ''rightness'' metrics, but also significantly expands the criteria to ''goodness'' and broaden the task coverage to complex multi-shot sequencing and audio-visual integration. Consequently, by providing granular diagnostic signals, EvalVerse transcends a static leaderboard and establishes a fundamental infrastructure for future work, such as reward models and evaluator agent.
1 Introduction
EvalVerse addresses evaluation’s dual gap: benchmarks emphasize prompt-following “rightness” while neglecting professional cinematic “goodness,” and automated metrics lack expert-calibrated credibility. It introduces a pipeline-aware taxonomy and expert-calibrated VLM reasoning to provide broader, more trustworthy diagnostic evaluation.
- Evaluation gaps: Existing benchmarks focus on whether generated videos are “right,” assessing prompt following and basic visual-element presence rather than professional cinematic quality.They neglect nuanced aesthetic, physical, and cinematic qualities required for professional production.
- Evaluation gaps: The shift from “rightness” to “goodness” creates a methodological and credibility bottleneck because cinematic quality depends on domain-specific expertise and subjective judgment.Previous automated metrics fail to capture these expert-dependent perceptual nuances.
- Core contributions: EvalVerse introduces a pipeline-aware cinematic taxonomy that audits final generated videos through the professional filmmaking stages of pre-production, production, and post-production.The taxonomy provides a structured diagnostic lens for defining and measuring cinematic “goodness.”
- Core contributions: EvalVerse calibrates human expert judgments with current VLM perceptual and analytical boundaries to produce expert-aligned Chain-of-Thought evaluation.The process involves filmmakers, artists, algorithm scientists, and engineers in a human-in-the-loop calibration process.
- Coverage and impact: EvalVerse expands coverage beyond silent, single-shot benchmarks to include multi-shot sequencing and audio-visual integration while retaining “rightness” and “goodness” evaluation.The framework is presented as a source of trustworthy diagnostic signals for reward modeling and agentic evaluator workflows.
2 Related Work
Generative video models have progressed from early 3D U-Nets toward scalable architectures and controllable, professional-grade production. Evaluation has likewise evolved from holistic consistency metrics toward decomposed benchmarks and specialized cinematography-focused assessment.
- Model Evolution: Generative video foundation models progressed from early 3D U-Nets to scalable DiT and Flow Matching architectures.The passage also describes a shift from stochastic, silent generation toward highly controllable, professional-grade production.
- General Benchmarks: Early benchmarks relied on holistic FVD and CLIP-Score metrics that often missed temporal dynamics and semantic precision.VBench shifted evaluation toward decomposing video quality into multiple hierarchical dimensions.
- Cinematography and Aesthetics: Specialized benchmarks now assess professional cinematography, including camera-control precision, lighting, and cinematographic understanding.Stable Cinemetrics introduced a structured taxonomy, while CineTechBench focused more narrowly on specific cinematographic tasks.
3 Taxonomy
EvalVerse introduces a hierarchical, pipeline-aware taxonomy that bridges AI video synthesis and professional filmmaking standards. It uses the traditional filmmaking workflow as a diagnostic lens rather than assuming multi-step generation or treating videos as flat visual-attribute collections.
- Taxonomy: EvalVerse organizes evaluation through a hierarchical, pipeline-aware taxonomy aligned with professional filmmaking standards.The taxonomy is designed to bridge AI video synthesis and professional filmmaking.
- Taxonomy: The taxonomy uses the traditional filmmaking workflow as a diagnostic lens for end-to-end video generation.This approach does not assume that modern foundation models follow a multi-step generation process.
- Taxonomy: EvalVerse mirrors the professional cinematic workflow with comprehensive evaluation dimensions spanning the full production pipeline.The framework presents its taxonomy as pipeline-aware and workflow-oriented.
3.1 Pre-Production
Pre-Production evaluates visual development and asset-design logic before dynamic synthesis, emphasizing identifiable, logically consistent characters and environments. It audits conceptual integrity against the intended worldview, narrative setting, physical laws, spatial logic, and artistic style.
- Pre-Production: Pre-Production evaluates foundational visual development and asset-design logic before dynamic synthesis, ensuring generated assets are identifiable and logically consistent.This stage focuses on foundational assets rather than dynamic synthesis.
- Pre-Production: Conceptual integrity is audited across characters and environments to ensure alignment with the intended worldview and narrative settings.This dimension is positioned as a cornerstone of directing and art design.
- Pre-Production: Character evaluation includes Identifiability, requiring recognizable visual anchors that distinguish subjects without unintended identity morphing.Examples of visual anchors include facial structures, body types, and silhouettes.
- Pre-Production: Scene evaluation audits environmental plausibility through physical laws, spatial logic, and genre distinctiveness in the visual language.It addresses gravity, collisions, support, perspective, scale, object relations, and clear artistic styles.
3.2 Production
Production evaluates the execution of the virtual shoot across subject performance, camera language, visual rendering, and emotional atmosphere. Its taxonomy examines consistency, physical and psychological expressiveness, visual storytelling, technical fidelity, material and lighting realism, and emotional continuity.
- Subject Performance: Subject performance is assessed through character consistency, kinetic action, and nuanced expression, including stable identity and attributes, physically logical movement, emotional alignment, and continuous facial transitions.Consistency covers face identity, hair, clothing, and accessories; action evaluates tension and emotion synergy; expression evaluates accuracy, facial tension, diversity, and continuity.
- Virtual Camera: Camera evaluation audits composition, optical validity, and movement pacing to ensure framing, lens behavior, focus, exposure, and trajectories serve visual storytelling.Criteria include shot-size rationality, subject prominence, spatial layering, depth of field, focal length, focus, exposure, and movement rationality.
- Visual Aesthetics: Visual aesthetics assess render fidelity, color grading, surface realism, and illumination logic, requiring detailed, temporally stable imagery with coherent palettes, materials, shaders, and lighting.The criteria penalize noise, compression, aliasing, ghosting, plastic textures, generative artifacts, inconsistent materials, and unexplained lighting behavior.
- Emotional Atmosphere: Emotional atmosphere is evaluated as a continuous experience, from an identifiable tonal baseline and visual-emotional synergy to smooth transitions and layered intensity over time.The emotional arc should build or recede rhythmically through setup, development, and climax without causeless jumps, flatness, or excessive explosiveness.
3.3 Post-Production
Post-production evaluation focuses on natively generated multi-shot sequences and synthesized audio, assessing shot assembly, temporal coherence, and audio-visual integration. Its taxonomy covers sequential logic, editing rhythm, vocal integration, and immersive soundscape quality.
- 3.3 Post-Production: The scope evaluates natively generated multi-shot sequences and synthesized audio because complex traditional editing and dubbing interventions remain challenging to assess.The final stage centers on shot assembly and multimodal integration while explicitly excluding broader post-processing interventions.
- Logic: Sequential logic assesses scene, narrative, and spatial continuity across cuts, including stable environments, causal action sequencing, subject states, positioning, and the 180-degree rule.These criteria target narrative flow, spatial coherence, and avoidance of generation errors across multiple shots.
- Rhythm: Editing rhythm evaluates shot-duration rationality and rhythmic layering by matching information consumption, emotional tension, audio rhythm, cutting tempo, and narrative-arc progression.The framework considers both dynamic and static shots, from setup through climax.
- Vocal: Vocal integration evaluates acoustic quality, lip-sync, and narrative sound design for voice fidelity, phoneme-mouth alignment, emotional matching, and off-screen storytelling cues.It also checks character-consistent voice attributes, technical purity, spatial reverb, volume-to-mouth alignment, and audience attention guidance.
- Soundscape: Soundscape evaluation audits ambient sound fidelity and musical score alignment for realistic Foley, spatial depth, environmental matching, and synchronization with visual cuts and emotional beats.These criteria target an immersive sonic environment coordinated with the image.
4 Dataset Curation: Test Pair Construction
EvalVerse constructs “Real-to-Gen” test pairs from professional cinematic videos through structured annotation, proportional sampling, and task-specific multimodal prompt construction. The process produces verified ground-truth metadata and balanced coverage across nine cinematic dimensions.
- Data Engine: The data engine transforms raw cinematic videos into “Real-to-Gen” test pairs through structured annotation, strategic sampling, and test pair construction.This pipeline is designed to capture professional filmmaking complexities in the benchmark.
- Annotation: A multimodal perception suite extracts structured metadata spanning the evaluation taxonomy, including camera parameters, character attributes, and environments.The database contains professional films and animations, and the labels undergo rigorous manual verification.
- Annotation: Manually verified labels provide robust ground-truth metadata for downstream sampling and prompt generation.Verification follows industrial-grade processing and supports the reliability of later pipeline stages.
- Sampling: Proportional sampling across nine core cinematic dimensions maintains balanced, industry-representative coverage rather than relying on stochastic selection.The sampling operates on the annotated database to keep the benchmark comprehensive.
- Construction: Task-specific multimodal test pairs combine structured metadata and raw video captions into professional-grade prompts, while reference-based tasks use extracted keyframes.Gemini 3.1 Pro synthesizes prompts using cinematic terminology, and Nano Banana Pro generates high-fidelity references from keyframes.
5 Benchmark: Expert Evaluation Results
Expert evaluation uses strict, multidisciplinary side-by-side ranking across predefined cinematic dimensions to compare leading closed-source, open-source, multi-shot, and audio-visual video models. Seedance 2.0 achieves the strongest overall performance, while other models exhibit distinct strengths, moderate capabilities, or specialized weaknesses across generation settings.
- Human Evaluation Protocol: A multidisciplinary team evaluates prompt, ground-truth, and model videos through strict side-by-side discriminative rankings across predefined dimensions.The protocol is designed to capture both cinematic aesthetics and algorithmic fidelity.
- Video Generation Model Selection: The benchmark covers closed-source, open-source, multi-shot, and audio-visual models, including Seedance 2.0, Kling-v3-Omni, Happy Horse 1.0, Hunyuan 1.5, LTX2, and Wan2.2.The evaluated set also includes Vidu-Q2-Pro, Hailuo 2.3, HoloCine, UniVideo, and MultiShotMaster.
- Overall: Seedance 2.0 achieves the best comprehensive performance, followed by Kling-v3-Omni and Happy Horse 1.0, while Wan2.2, Hunyuan 1.5, and LTX2 show moderate overall capability.HoloCine, UniVideo, and MultiShotMaster exhibit more uneven or specialized performance profiles.
- Text-to-Video: Seedance 2.0 remains strongest in Text-to-Video, with leading results in soundscape fidelity, identity preservation, attribute consistency, visual quality, and camera control.Kling-v3-Omni also maintains a leading position across fine-grained criteria.
- Text-to-Video: HoloCine is strong in scene plausibility and multi-shot consistency, whereas MultiShotMaster performs relatively well in shot-duration rationality but trails on broader visual, acting, aesthetic, and sound dimensions.HoloCine’s lower acting and rendering-related scores constrain its overall quality.
- Reference-to-Video: In Reference-to-Video, Seedance 2.0, Kling-v3-Omni, and Happy Horse 1.0 form the leading group, with Seedance 2.0 scoring highest overall.Seedance 2.0 is particularly strong in vocal acoustic quality, chromatic harmony, attribute consistency, visual concept preservation, and camera-related criteria.
6 Machine Evaluation Suite
The Machine Evaluation Suite approximates expert cinematic judgment with a multidimensional score vector produced by a pipeline that combines deterministic perception operators and a fine-tuned VLM. The VLM uses expert-guided multi-questioning, step-by-step Chain-of-Thought reasoning, self-reflection, and context-aware metric gating, learned through pairwise and pointwise expert supervision.
- Evaluation Objective: The framework computes a multidimensional cinematic score vector S to approximate expert judgment from video, audio, prompts, and references.The evaluation pipeline is described in Sec. 6.1 and powered by a fine-tuned VLM in Sec. 6.2.
- Evaluation Pipeline: Inference first extracts deterministic evidence with specialized operators, then applies expert-guided, step-by-step reasoning through the fine-tuned VLM.The operators provide contextual perception priors because VLMs struggle with fine-grained temporal tracking and low-level perception.
- Evaluation Pipeline: The perception operators include DINO and InsightFace for identity tracking, YOLO for semantic anchoring, SyncNet for audio-visual synchronization, and Whisper for speech emotion recognition.These operators form the suite Φ used to extract objective professional evidence Eprof.
- VLM Reasoning: Given multimodal context and expert-designed questions, the fine-tuned VLM generates detailed Chain-of-Thought reasoning rather than blindly outputting scores.Self-reflection prompts the model to re-examine its judgments for hallucinations, while context-aware gating bypasses metrics unsupported by the narrative context.
- Expert-Calibrated Fine-Tuning: Two-stage fine-tuning first learns relative cinematic preferences from pairwise comparisons, then maps them to absolute metrics while generating expert rationales and scores.The second stage uses pointwise supervision with ground-truth expert Chain-of-Thought rationales and absolute expert scores, trained autoregressively with cross-entropy.
7 Human-Machine Calibration
EvalVerse uses progressive human-machine calibration to adapt evaluation prompts, weights, and model parameters to cinematic expertise. Its CoT and SFT strategies yield strong alignment on perceptually grounded criteria while addressing abstract, temporal, and cross-modal dimensions.
- Calibration mechanism: EvalVerse proposes a progressive three-tier calibration mechanism spanning prompt-level rationale replacement, fusion-level weight optimization, and an additional calibration tier.Prompt calibration replaces overly abstract dimensions and multi-questions that exceed VLM perception or reasoning capabilities.
- Alignment validation: Human-machine alignment is assessed through pairwise win-ratio comparisons, correlation coefficients, and trend-consistency visualization.Pairwise win ratio is used as the unified comparison signal across evaluations.
- Alignment results: EvalVerse predictions show striking absolute proximity to expert annotations, while correlation coefficients remain within a tight band across sub-dimensions.The reported analysis uses both Spearman Rank Correlation Coefficient and Pearson Linear Correlation Coefficient.
- Alignment results: Prompt-level CoT provides strong alignment for pixel-grounded dimensions, including Visual Concept Design, Cinematography, Acting, Aesthetics, and Affectivity.The passage characterizes CoT-based digitization as a reliable backbone for most cinematic criteria.
- Limitations and complementarity: Frozen-VLM prompt-level CoT reaches a perceptual ceiling on subjective, temporally entangled, and cross-modal aspects such as multi-shot rhythm.Abstract concepts such as rhythmic layering cannot be robustly decomposed into zero-shot observable tokens through prompt elaboration alone.
- Limitations and complementarity: Task-specific SFT injects human scoring distributions into VLM parameters, complementing CoT by bridging the perception-reasoning gap for complex dimensions.CoT supplies transparent reasoning across the pipeline, while SFT provides last-mile alignment for abstract cinematic expertise.
8 Conclusion
EvalVerse reframes video-generation evaluation from basic prompt-following to auditing professional filmmaking, using pipeline structure and human–machine calibration to encode nuanced human preferences. Future work targets temporal perception, long-form narrative reasoning, and evaluation of diverse avant-garde styles.
- Conclusion: EvalVerse shifts assessment from whether a video is right to whether it is good within a professional filmmaking workflow.The framework structurally mirrors the real-world pipeline and audits cinematic quality beyond basic prompt-following.
- Conclusion: Human–machine calibration provides a principled mechanism for injecting nuanced human preferences into algorithmic scoring and digitizing subjective expertise.The conclusion presents calibration as the basis for translating human judgments into computational evaluation.
- Limitations and Future Work: Future work must address VLM keyframe-based temporal limits, 10+ minute narrative reasoning, and assessment of boundless avant-garde artistic styles.Current VLMs process discrete keyframes rather than continuous streams, while macro-narratives require advanced long-context reasoning.