Source-linked AI summary
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjing Chen, Chongxiao Wang, Yanhao Wu, Chenglong Huang, Xiaoqian Zhu, Jie Tian, Hua Li, Jingjing Fan, Mingshuang Tang, Zhong Li, Hengxia Qiang, Weibin Chen, Jinyang Zhen, Bing Zhao, Lin Qu, Jing Li, Hu Wei
TL;DR
Existing video benchmarks do not adequately evaluate professional Cinematic Language or cross-shot film structure. FilmBench addresses this with film-grounded prompts, an academy-aligned taxonomy, and an expert-grade evaluator, whose rankings match expert rankings while exposing persistent cinematic-generation gaps.
Problem
Existing benchmarks lack professional Cinematic Language criteria and end-to-end evaluation of cross-shot narrative structure against verified cinematic references.
Method
FilmBench uses prompts reverse-engineered from award-winning films, an academy-aligned three-level taxonomy, and the FilmOps expert-grade automatic evaluator for T2V and R2V.
Results
Models show no saturation across T2V and R2V: dynamic aesthetics are lowest-scoring, and multi-shot prompts are uniformly harder, with top scores of 88.93 for T2V and 86.66 for R2V.
Takeaways & Limitations
FilmBench provides a cinematic benchmark that distinguishes models on professional Cinematic Language and supports research toward film-grade generative systems.
Takeaways & Limitations
FilmBench’s evaluation dimensions and model rankings may not reflect performance on non-cinematic video-generation tasks.
Abstract
from arXiv · showhide
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.
1 Introduction
FilmBench addresses the gap between generic video benchmarks and professional cinematic creation with a director-grounded T2V/R2V benchmark built from award-winning film clips. It combines an academy-aligned Cinematic Language taxonomy, reverse-engineered multi-shot prompts, and an expert-grade evaluator whose rankings closely match film-industry experts.
- Motivation: FilmBench evaluates cinematic T2V and R2V generation against professional standards, addressing benchmarks’ reliance on web- or LLM-sourced prompts and generic scene distributions.The benchmark was built jointly with Beijing Film Academy directors and faculty and the Hujing Digital Media & Entertainment Group film studio.
- Benchmark design: 20 film categories curated by expert directors anchor FilmBench’s prompts and fine-grained evaluation hierarchy in professional cinematic genres and Cinematic Language.The taxonomy contains 3 L1 axes, 12 L2 components, and 35+3 R2V-only L3 sub-metrics.
- Benchmark design: Directors reverse-engineer structured shot-list prompts from award-winning film clips, while R2V additionally evaluates scene-space, character-appearance, and prop-reference fidelity.R2V has 13 L2 components and 35+3 L3 sub-metrics, with only the sub-metric matching each prompt’s reference type scored.
- Evaluation: Spearman ρ=0.95 for T2V and 0.96 for R2V shows that FilmBench’s in-house evaluation agent reproduces expert film-industry model rankings.The benchmark open-sources the evaluator’s core Cinematic Language operator suite, FilmOps.
- Findings: Top scores of 88.93 for T2V and 86.66 for R2V indicate no evaluated model approaches the cinematic-quality ceiling.FilmBench benchmarks 9 T2V models and 7 R2V models, including Seedance 2.0, HappyHorse, Kling, Veo, Grok Imagine Video, Vidu, and Hailuo.
2 Related Work
Video-generation evaluation progressed from generic quality scores to capability-specific tests and increasingly film-like assessment, but each stage left gaps in professional cinematic grounding. FilmBench addresses the accumulated shortcomings of generic taxonomies, clip-level scope, and fragmented evaluation of film aspects.
- Evaluation Evolution: Evaluation evolved through three stages—generic quality scoring, capability-specific stress tests, and increasingly film-like assessment—each closing one gap while exposing the next.This progression motivates FilmBench.
- Generic Evaluation: Early benchmarks used distribution metrics and generic multidimensional criteria that did not localize failures or reflect professional film-production demands.Prompts were largely drawn from web users, and evaluation axes remained generic rather than production-grounded.
- Capability-Specific Evaluation: Capability-specific suites exposed brittle attribute, motion, spatial-relation, temporal, and motion-binding failures that generic quality scores masked.Examples include T2V-CompBench, FETV, ChronoMagic-Bench, and VMBench.
- Capability-Specific Evaluation: MSVBench still evaluates shots with generic quality criteria rather than director-intended shot lists, leaving cross-shot narrative structure and director intent unjudged against verified cinematic references.This constitutes the second identified gap.
- Toward Film: Film-oriented benchmarks separately address reference consistency, audio, aesthetics, and Cinematic Language, leaving coverage fragmented across individual aspects.Reference/image-conditioned suites raise reference fidelity as a concern, while the broader positioning identifies generic taxonomy, clip-level scope, and fragmented film-aspect coverage as the accumulated gaps.
3 FilmBench: Construction and Dimensions
FilmBench is constructed by reverse-engineering director-selected award-winning cinematic clips into structured, predominantly multi-shot prompts for T2V and R2V. Its evaluation uses a professionally co-designed, three-level Cinematic Language taxonomy spanning 3 axes, 12 components, and 35 sub-metrics, with R2V adding reference-specific visual-following measures.
- Cinematic Language taxonomy: The taxonomy organizes cinematic competence hierarchically into 3 L1 axes, 12 L2 components, and 35 L3 sub-metrics.This multi-level structure supports conclusions at coarse, medium, and fine granularity and is co-designed with film-school and studio faculty.
- Reverse-engineering pipeline: Professional directors select award-winning, mostly multi-shot clips across 20 cinematic genres to stress Cinematic Language and visual generation.The selection involves the Beijing Film Academy and Hujing Digital Media & Entertainment Group film studio.
- Tasks and scale: 1,056 of 1,169 prompts script multiple shots, including 402 of 515 T2V prompts and all 654 R2V prompts.The prompts follow real shot lists, unlike single-clip benchmarks.
- Task-specific extension: R2V adds Visual Following with 3 L3 sub-metrics: scene-space, character-appearance, and prop-reference fidelity.Only the fidelity sub-metric matching the prompt’s scene, prop, or character reference type is scored.
4 Evaluation Method
FilmBench uses an expert-grade evaluator grounded in FilmOps, a practitioner-vetted Cinematic Language operator suite trained across genres and paired with sample-level score aggregation. Fine-grained scoring also reveals diagnostic model trade-offs that aggregate scores can obscure.
- Motivation: Generic multimodal judges and existing domain models misjudge professional cinematic categories and fail to transfer reliably across genres.The identified categories include shot scale, composition, and camera movement.
- FilmOps design: FilmOps defines six practitioner-vetted dimensions spanning shot scale, composition, viewing angle, tone & color, character layout, and camera movement.Five dimensions operate at frame level and camera movement operates at shot level; together they cover 55 sub-categories.
- FilmOps design: FilmOps is trained on 5,000+ real film/TV works spanning live-action, 3D animation, 2D animation, stylized/VFX content, and varied narrative types.Each operator uses ∼40K–60K annotated samples with shot-based testing and practitioner-led quality control.
- FilmOps release: FilmOps open-sources its classification standard, model weights, and inference scripts for reusable, modular self-evaluation pipelines.Users can reuse individual operators or embed selected dimensions into their own evaluation systems.
- Score aggregation: 1–5 sub-metric scores map linearly to 0–100, aggregate directly from L3 to L1 with equal L3 weighting, and average across videos at model level.Each video is first aggregated across dimensions, preserving sample variance for significance analysis.
- Qualitative analysis: 86.11 vs. 55.97, Seedance 2.0 outperforms Grok Imagine Video on a multi-shot T2V sci-fi mech battle, with gaps varying by cinematic dimension.The example demonstrates that fine-grained per-dimension scores expose trade-offs hidden by aggregate scores.
5 Experiments
FilmBench finds that leading video models remain far from saturation and differ most in Cinematic Language, especially camera craft, while fine-grained evaluation exposes complementary strengths hidden by aggregate rankings. Performance is also substantially harder on multi-shot prompts, particularly for weaker models, although overall automatic rankings closely match expert judgments.
- Evaluator validation: Overall automatic rankings achieve ρ ≥0.95 on both tasks, although agreement is lower for audio quality at 0.68, audio coherence at 0.68 and editing at 0.60.These lower-agreement components involve subjective temporal behavior such as sound fidelity and cut rhythm.
- Overall results: Seedance 2.0 leads the overall FilmBench score on T2V at 88.93 and R2V at 86.66, but no model approaches saturation.On T2V, HappyHorse 1.1 scores 87.42 and Hailuo 2.3 scores 68.94; on R2V, HappyHorse 1.1 scores 85.51.
- Fine-grained differentiation: 99.4 Instruction Following variance exceeds Aesthetic Quality at 12.6 and Temporal Continuity at 5.6, making it the largest axis-level differentiator.At the component level, Cinematic Language has variance 187.9, ahead of Audio at 148.4 and Scene at 71.3.
- Fine-grained differentiation: Camera movement, focus, shot scale, viewing angle and fore/mid/background are the five highest-variance sub-metrics, with camera movement spanning 357.6.Seedance 2.0 scores 86.5 on camera movement, 85.6 on focus and 79.8 on shot scale, leading the cited head-model comparisons.
- Model specialization: No single system dominates every axis of film craft: Seedance 2.0 leads camera craft, HappyHorse leads scene consistency and character performance, and Seedance 2.0 wins only 18/35 T2V or 20/38 R2V sub-metric championships.HappyHorse 1.1 reaches 96.7 on Scene and 97.6 on Character & performance, while Seedance 2.0 reaches 96.3 on Character & performance.
- Multi-shot evaluation: 7.9 points is the average multi-shot score drop, from 89.2 to 81.3, with larger losses for weaker models.Seedance 2.0 drops 2.3 points and HappyHorse 1.1 drops 2.7, versus Grok Imagine Video at 10.2, Veo 3.1 at 10.3 and Hailuo 2.3 at 22.8.
6 Conclusion
FilmBench evaluates T2V and R2V generation through professional Cinematic Language, combining film-grounded prompts, an academy-aligned taxonomy, and expert-grade evaluation. Its conclusions target cinematic content creation and research toward higher-quality video generation, not non-cinematic production scenarios.
- Conclusion: FilmBench benchmarks T2V and R2V generation using professional Cinematic Language and prompts reverse-engineered from award-winning real films.The benchmark was built with directors and Beijing Film Academy faculty and includes a three-level taxonomy with 3 L1 axes, 12 L2 components, and 35 L3 sub-metrics.
- Limitations: FilmBench’s evaluation dimensions and model rankings are not necessarily representative of non-cinematic tasks such as vertical short-form video or casual user-generated content.The benchmark is designed for professional film and cinematic content creation capabilities.
- Broader Impact: FilmBench does not itself generate or deploy content and is intended to support research toward higher-quality video generation systems.The authors do not anticipate negative societal impacts from the work.
A Qualitative Evaluation Examples
Figures 16–17 show how FilmBench’s fine-grained R2V evaluation exposes dimension-level trade-offs hidden by aggregate scores. Character-reference fidelity is generally strong, while prop-reference fidelity reveals a shared bottleneck caused by effects altering the reference prop.
- Dimension-Level Trade-offs: Figures 16–17 show that R2V score gaps concentrate on Cinematic Language and character performance within Instruction Following, while temporal and aesthetic scores cluster.The examples demonstrate trade-offs that aggregate scores alone would conceal.
- Reference Fidelity: Reference-character appearance fidelity scores remain relatively high across models because reference identity is generally preserved.This is illustrated by the reference-character case in Figure 16.
- Reference Fidelity: 0: both models score 0 on prop-reference fidelity when a golden energy aura bleeds onto the reference prop and alters its color.The reference-prop case in Figure 17 exposes this as a field-wide bottleneck.
B L3 Dimension Definitions
The benchmark defines 35 L3 sub-metrics for T2V and adds three reference-grounded sub-metrics for R2V, for 38 total, all scored on a 1–5 anchor-rubric scale.
- L3 Dimension Definitions: 35 L3 sub-metrics are defined for T2V and grouped by their L1 axis and L2 component.Tables 3 and 4 list the complete T2V metric set.
- L3 Dimension Definitions: 3 further L3 sub-metrics bring the R2V total to 38 under a Visual-Following component of Instruction Following.The additions cover scene-space, character-appearance, and prop-reference fidelity against corresponding reference types.
- L3 Dimension Definitions: 1–5 anchor-rubric scoring is used for every sub-metric.Each metric is rated against an anchor rubric.
C Per-Task Variance Breakdown
This section reports per-task cross-model variance across every L3 sub-metric, with a merged two-task view provided in the main text.
- Per-Task Variance: Figures 18 and 19 show per-task cross-model variance for each L3 sub-metric.The breakdown is organized by task and sub-metric.
- Per-Task Variance: The variance analysis covers every L3 sub-metric.It measures cross-model variation separately for each task.
- Merged View: Figure 6 presents the merged view combining both tasks in the main text.The main-text figure complements the per-task breakdowns in Figures 18 and 19.
D Degradation Analysis
FilmBench’s degradation analysis examines performance losses in single- to multi-shot and dialogue-to-action transitions. It also illustrates cases where R2V gaps arise from Cinematic Language and character performance despite high reference fidelity.
- Degradation analysis: Tables 5 and 6 analyze T2V degradation from single- to multi-shot and dialogue-to-action transitions across taxonomy levels.Both tables report mean scores macro-averaged over nine models and define ∆ as the corresponding performance drop.
- R2V examples: 85.43 vs. 63.57: Seedance 2.0 outscored Grok Imagine Video in an R2V reference-character example.Character-appearance fidelity was high, while the score gap was driven by Cinematic Language and character performance.
- R2V examples: 82.75 vs. 79.36: Seedance 2.0 outscored Kling 3.0 Omni in an R2V reference-prop example.Both models scored 0 on prop-reference fidelity because a golden energy aura altered the prop’s color; the overall gap was driven by Cinematic Language.
E Per-axis Scores · F Full L3 / L2 Heatmaps · G FilmOps Taxonomy Details
The appendix details per-axis and reference-type scores, market robustness, and full L3/L2 heatmaps for T2V and R2V. It also specifies FilmOps’ practitioner-validated taxonomy of six dimensions and 55 fixed sub-categories.
- E Per-axis Scores: The appendix evaluates machine scores over the full prompt sets using sample-level stand aggregation: T2V N = 515 and R2V N = 654.Model abbreviations include Seedance 2.0, HappyHorse 1.1/1.0, Kling 3.0 Omni, Kling 3.0, Grok Imagine Video, Veo 3.1, Vidu Q3-Pro/Q2-Pro, and Hailuo 2.3.
- E Per-axis Scores: Tables 7–9 report per-axis T2V and R2V scores plus R2V visual-following fidelity by scene, character, and prop reference type.The R2V reference-type subsets are disjoint: scene N = 232, character N = 213, and prop N = 209.
- E Per-axis Scores: Seedance 2.0 ranks first on both Chinese and international T2V source-title subsets.The market-restricted subsets contain 1,431 Chinese and 1,764 international records, excluding prompts without a market label.
- F Full L3 / L2 Heatmaps: Complete per-model heatmaps corroborate the aggregated rankings while exposing low-variance temporal/audio rows and discriminative camera/framing, dynamic-aesthetics, and R2V prop-fidelity rows.Figures 20–23 provide complete score matrices at L3 and L2 granularities, with rows grouped by axis or component and columns ordered by overall rank.
- G FilmOps Taxonomy Details: FilmOps maps each frame or shot into structured cinematic labels across six core dimensions and 55 fixed sub-categories.Character layout is an open-ended natural-language field and is excluded from the 55-category count; definitions align with classical production references and were validated by practitioners.
- G FilmOps Taxonomy Details: R2V heatmaps add three Visual Following fidelity rows at L3 and one Visual Following component at L2.The R2V maps contain 38 sub-metrics across seven models at L3 and 13 components across seven models at L2.
- G FilmOps Taxonomy Details: The FilmOps inventories define eight shot-scale classes, 12 composition classes, and seven viewing-angle classes.The listed shot scales range from extreme close-up to extreme long shot, while composition includes center, rule of thirds, symmetry, framing, leading lines, and depth of field.
H FilmOps Operator Evaluation
FilmOps uses trained operators evaluated against four strong zero-shot general MLLM baselines, with classification measured by macro-F1 and character layout by precision. The trained operators outperform every baseline across all six dimensions, especially on camera movement, composition, and tone & color.
- FilmOps Operator Evaluation: FilmOps compares each operator with GPT-4.1, Qwen3-VL-235B, Gemini 3.5 Flash, and Gemini 3.1 Pro.The comparison reports each operator’s backbone and accuracy.
- FilmOps Operator Evaluation: The trained operators outperform every baseline across all six dimensions.The strongest advantages occur on temporally and professionally demanding dimensions.
- FilmOps Operator Evaluation: The largest performance advantages occur for camera movement, composition, and tone & color.These results corroborate the design discussion in Section 4.1.
- FilmOps Operator Evaluation: Classification operators use macro-F1, while the natural-language character-layout operator uses precision.Unsupported settings are marked with “–”; “F” and “P” denote Flash and Pro.