Source-linked AI summary
DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
Haoyuan Shi, Mingtao Chen, Shuo Jiang, Ziyan Chen, Xuyi Sheng, Yiming Liu, Ying Zhang, Miao Wang, Jianxiang Lu, Fanyang Lu, Songyuanyi Lu, Xiele Wu, Zhichao Hu, Yuhong Liu, Richeng Xuan
TL;DR
Existing benchmarks do not evaluate the complete short-drama production chain on artefacts produced by that chain, leaving stage adherence and cross-stage coherence insufficiently assessed. DramaChain Bench addresses this with shared cross-stage dimensions, a commercially calibrated pipeline, localised human annotation, and an agentic judge. Human annotations show upstream defects accumulate downstream, while automated scoring reproduces model rankings at a mean PLCC of 0.918.
Problem
Existing benchmarks largely evaluate isolated video-generation stages with authored inputs and lack localised, stage-linked evidence about defects across the production chain.
Method
DramaChain Bench evaluates pipeline-produced items across six stages using five shared axes, professional localised annotations, and an agentic judge validated against human references.
Results
Upstream defects accumulate along the chain, and DramaChain Agentic Judge reproduces model rankings at a mean PLCC of 0.918.
Takeaways & Limitations
Final short-drama quality is not governed by video generation alone, and new models can be evaluated without additional human annotation.
Takeaways & Limitations
The 63-dimension taxonomy does not provide exhaustive coverage; the paper explicitly documents negative coverage.
Abstract
from arXiv · showhide
Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pipeline outputs. This leaves two critical questions unanswerable: whether each stage adheres to the original script intent (rather than only its immediate input prompt), and whether disparate shots remain coherent after assembly into multi-episode releases. We present DramaChain Bench, the first short-drama benchmark that evaluates every stage of the complete production chain. It is built upon three in-house systems sharing one dimension system, DramaChain Dimensions: five evaluation axes instantiated at every stage, resolving into 63 leaf dimensions. DramaChain Agent is calibrated against commercial short-drama platforms in both workflow and finished short-drama quality, enabling stage-wise fair comparison across models. DramaChain Labeling System has each of the 5,785 items scored independently by three professional annotators, with all defects spatio-temporally localised and selected from a predefined defect list. This process produces 17,488 valid scores and 255,925 traceable attribution records. The human annotations confirm that upstream defects cascade across the pipeline, demonstrating that final episode quality is not governed by video generation alone. DramaChain Agentic Judge then scores every leaf dimension automatically, gathering evidence over multiple agentic rounds before judging against a per-item checklist; it reproduces the model ranking at a mean PLCC of 0.918, enough to admit new models at no annotation cost.
1 Introduction
DramaChain Bench addresses gaps in stage-isolated evaluation by scoring pipeline-produced artefacts across the complete short-drama chain. Its calibrated generation, structured annotation, and automated judging systems support fair stage-wise comparison and traceable evaluation.
- Motivation: Existing benchmarks test stages in isolation using authored inputs, leaving defect causes and downstream propagation unmeasured.They also lack localisation of which dimension failed, where it failed, and why.
- Benchmark: DramaChain Bench evaluates every stage on items produced by the tested production chain.Its five axes are instantiated at six granularities, yielding 63 leaf dimensions over 5,785 items.
- Benchmark: DramaChain Agent generates 20 dramas and 60 episodes, forks only at the tested stage, and gives competing models identical upstream artefacts and prompts.Its workflow and output quality are benchmarked against commercial short-drama platforms.
- Evaluation: DramaChain Agentic Judge reproduces human model rankings at a mean PLCC of 0.918, enabling new models to enter without additional human annotation.The framework also supports extensible evaluation through stage-level branching and consistent automatic scoring.
2 Related Work
Prior work studies narrative generation, staged production systems, and video-quality benchmarks, but existing corpora and evaluations do not capture complete model-produced hand-offs with localised causal attribution. DramaChain Bench is positioned to fill that chain-level gap.
- Generation systems: Multi-shot narrative systems either generate scenes jointly or use memory mechanisms to maintain cross-shot information.Examples include joint scene denoising, causal shot-boundary generation, and memory-based refinement.
- Generation systems: Staged generation systems decompose production into roles that pass artefacts between stages, sometimes adding multi-agent search.Their differences mainly concern stage boundaries and whether search is applied.
- Datasets and benchmarks: Existing corpora are assembled from footage or references not produced by the pipeline under test and lack complete generation-chain records.They therefore do not provide one pipeline’s script, storyboard, keyframes, shot videos, and finished drama as linked artefacts.
- Datasets and benchmarks: Prior benchmarks score capability dimensions, film language, consistency, and other finished artefacts, but generally evaluate outputs rather than production hand-offs.VideoWeaver reads traces and intermediate files for workflow completion, not rubric-level quality of each intermediate artefact.
- Datasets and benchmarks: Table 1 compares related benchmarks using disclosed or estimated dimension and item counts, with explicit markers for unavailable bases and embedded evaluation suites.The comparison also applies the paper’s definition of stages scored.
3 DramaChain Bench
DramaChain Bench combines a shared cross-stage rubric, a commercially calibrated production agent, exhaustive localised annotation, and an automated judge. These components make stage comparisons and defect tracing possible across pipeline-produced short-drama items.
- DramaChain Dimensions: Shared axes let the benchmark compare an upstream instruction with downstream outputs instead of charging every observed defect to the final generation stage.For example, storyboard executability and video plausibility can be evaluated under the same generation-plausibility axis.
- DramaChain Dimensions: DramaChain Dimensions applies five shared evaluation axes across six stages, resolving into 63 leaf dimensions.The axes cover input fidelity, internal consistency, generation plausibility, visual quality, and cinematic expressiveness.
- DramaChain Agent: DramaChain Agent produces pipeline-generated items, including 20 dramas and 60 episodes spanning channel, period, genre, visual style, and language.The system also creates intermediate production artefacts such as reference sheets, shot descriptions, and a camera tree.
- DramaChain Agent: The agent matches commercial production through a fully automatic six-stage chain and four video-input paradigms, while expert review found comparable output-quality bands.The evaluated paradigms include first/last frames, grid panels, and multireferences.
- DramaChain Agent: Stage-level branching gives every competing model identical upstream artefacts and prompts, so new-model evaluation requires only one stage of generation.Each shot slot is generated once per participating model at every granularity.
- DramaChain Labeling System: Three professional annotators score every item on applicable dimensions, while non-perfect scores require spatial or temporal localisation and fixed attribution tags.The system yields 17,488 scored entries and 255,925 verifiable attributions, with quality-control review and rule-based filtering.
- Automated evaluation: A single holistic VLM verdict is structurally limited because local frame defects and inter-segment discontinuities can escape impression-based assessment.The benchmark therefore audits whether specialist metrics align with human judgements before using them in hybrid evaluation.
- Automated evaluation: Most audited metrics correlate weakly with human annotations, with |SRCC| no greater than 0.2; fourteen photography metrics show negative correlation.The reported explanation is that photography and AIGC metrics apply priors misaligned with short-drama stylisation.
4 Experiment
Experiments show that performance declines and defects accumulate across the production chain, while model strengths vary substantially by evaluation dimension. Automated judging agrees strongly with human model rankings, but item-level reliability is modality-dependent and some criteria remain difficult to automate.
- Main results: 3.30: The leading end-to-end chain barely clears the 3.0 usable line, leaving negligible safety margin.The chain uses gpt-5.5-xhigh for storyboarding, gpt-image-2 for keyframes, and seedance-2.0 for video.
- Main results: 3.78 to 2.93: Scores decay from storyboard design through episode video before rebounding modestly to 3.03 for the finished drama.The three video granularities remain within 0.1 of the usable line, and episode video falls below it.
- Main results: 1.22, 0.70, 0.64, 0.53, and 0.42: Model gaps are widest for input fidelity and narrowest for cinematic expressiveness, following reference explicitness.Small gaps on weakly anchored metrics may reflect benchmark limitations rather than comparable model performance.
- Main results: 0.02 and 0.01 composite gaps conceal larger per-dimension trade-offs, so composite scores support tier grouping but can mislead model selection.For single keyframes and single-shot video, models with nearly identical composites differ substantially across fidelity, plausibility, and expressiveness.
- Agreement validation: 0.918 mean PLCC: DramaChain Agentic Judge reproduces model-level agreement with human evaluation across six stages, but video-stage item agreement reaches only 0.79–0.85× the human baseline.Storyboard design reaches 2.35× baseline consistency, and the validated scorer admits additional models without extra annotation.
5 Conclusion
DramaChain Bench evaluates short-drama production as a multi-stage chain, using in-chain artefacts to support stage-specific comparison and defect attribution. Its reusable infrastructure is intended for diagnosis across inter-stage boundaries and within episodes, rather than leaderboard ranking.
- Conclusion: DramaChain Bench evaluates each production stage using outputs produced within the chain itself.This design enables defects to be attributed to their originating stage rather than merely to where they are observed.
- Conclusion: Three systems provide calibrated generation, professional item-wise annotation, and validated automated scoring for the benchmark.DramaChain Agent generates and branches items at the tested stage; the Labeling System localises defects; Agentic Judge ranks models against the human reference.
- Conclusion: New models can be plugged into individual stages without additional human annotation, while human-dependent components remain explicitly exposed.The pipeline, dimension system and scoring configurations are reusable for extensible evaluation.
- Conclusion: The benchmark encourages evaluation across inter-stage boundaries and throughout episode shots instead of evaluating isolated clips.Its intended use is diagnosis of actual short-drama failure points.
6 Limitations
The benchmark’s 63 scored dimensions should not be treated as exhaustive coverage. Its limitations are documented explicitly through a table describing what the benchmark does not cover and why.
- Limitations: 63 scored dimensions do not imply exhaustive coverage of short-drama evaluation.The paper explicitly counters this assumption in its coverage analysis.
- Limitations: Table 10 records the benchmark’s uncovered areas and the reasons for those boundaries.The table is presented as an explicit negative-coverage statement.
Ethics and Responsible Use
The benchmark uses synthetic source material and professional third-party annotators, while screening generated content against commercial release policies. It does not certify public-release fitness, perform a broader harm audit, or predict commercial performance.
- Ethics and Responsible Use: All benchmark artefacts are synthetic and use no real-person likeness, voice reference, or reused commercial drama assets.The corpus consists of project-original scripts and model-generated character sheets, keyframes and clips.
- Ethics and Responsible Use: Generated episodes were screened against commercial release policies, with unevaluable episodes marked missing rather than imputed.The corpus retains common genre tropes including conflict, coercion and revenge.
- Ethics and Responsible Use: No broader harm audit was performed, and the benchmark does not certify artefacts as fit for public release.These are explicit scope boundaries for responsible use.
- Ethics and Responsible Use: Annotations were produced by professional contracted staff from third-party vendors rather than volunteers or anonymous crowd-workers.Recruitment was filtered by the background required for each stage, and median per-item working time was 23 minutes.
- Ethics and Responsible Use: The benchmark supports stage-wise model comparison and root-cause diagnosis but cannot predict commercial performance or evaluate safety and content policy.These two use cases are explicitly out of scope.
A The Dimension System and Its Implementation
DramaChain Dimensions uses a shared declarative taxonomy and rubric across six production stages, with routed evidence and model-callable tools implementing each leaf dimension. The implementation combines text, image, episode, and video-specific checks, while reporting cost and availability constraints for automated evaluation.
- Dimension System: DramaChain Dimensions defines 66 leaf dimensions, of which 63 are scored after retiring three with near-zero human baselines.Each scored dimension has a five-level rubric and an attribution-tag vocabulary shared by humans and the automated scorer.
- Implementation: Every leaf dimension is declaratively specified by a rubric, closed tag vocabulary and method line that determines its routed evidence producers.No per-dimension code is required because the router mounts producers from the method description.
- Implementation: The implementation distinguishes routed evidence from the resolved model-callable tool set when reading dimension tables.A dash indicates no selected producer or no closer examination on request, respectively.
- Video Dimensions: Video dimensions combine frame sampling, temporal and audio analysis, OCR, optical flow, action continuity, perceptual quality and camera or performance checks.Video-specific routing begins with video_probe and extract_frame, then adds tools according to the dimension’s needs.
- Text and Storyboard Dimensions: Text and storyboard dimensions use LLM, VLM and alignment checks for beats, dialogue, executability, additions, narrative flow and shooting rhythm.The listed implementations include dialogue edit distance and embedding similarity, spatial-physical checks, and sequence analysis.
- Keyframe Image Dimensions: Image dimensions use reference comparison, segmentation, embeddings, zoom inspection and rule-based checks for identity, appearance, layout, style and physical plausibility.The taxonomy includes both single-panel and cross-panel consistency dimensions, with some dimensions explicitly retired.
- Cost and Reliability: Automated evaluation has a median 14.6-minute wall-clock cost per item and reaches 17 of 21 video leaf dimensions through metric routing.Median tool time is 198.7 seconds per dimension, with 20 sub-agent calls and about 19 measurement calls per job; ocr_track availability is 71.1%.
A.5 Validating candidate metrics
Candidate metrics are retained as low-level evidence only after validation against human consensus, because most fail the scoring-anchor threshold across metric families. The judging model receives measurements directionally, while the aesthetic scorer is discarded.
- Every candidate metric was tested against three-annotator consensus across applicable granularities and human dimensions.Each scalar field was evaluated separately with SRCC and p, both pooled and within generation paradigms.
- Most candidate metrics failed the |SRCC| ≥0.3 scoring-anchor threshold over at least 100 items.Failures spanned structure, subject, identity, consistency, quality, aesthetics and AIGC-specific discriminators.
- Genre and art direction make frame-level visual quality non-uniform, causing the same feature to be judged differently across items.Deliberately saturated xianxia imagery illustrates why a visual feature can signal commercial style in one case but an AI artefact in another.
- Validated small models provide directional low-level evidence rather than scores, while the aesthetic scorer is omitted from judging prompts.Physics, reference-similarity and voiceprint fields retain their measurements; the aesthetic scorer is dropped before prompting.
- Model coverage is incomplete and non-random for several video models, whereas storyboard coverage is nearly complete.Failures resulted from provider-side content review and generation errors; affected models are scored only on the slices they ran.
- Storyboard +0.28, single keyframe +0.41, episode keyframes +0.57, single-shot video −0.35, episode video +0.14, short drama +0.42 are automated scorer signed offsets by stage.
D.1 By visual style: live action scores higher
Automated scores vary with visual style, dialogue language and input paradigm rather than reflecting a uniform difficulty shift. Input paradigm has the largest average effect, while style changes which defects are detected.
- By visual style: Live action scores higher on the automated composite at five of six granularities, by 0.08 to 0.22.Single keyframes are the exception, where animation is higher by 0.08.
- By visual style: At single-shot video, live action is 0.32 higher on C while animation is 0.25 higher on Q, with F identical.The larger dimensional extremes show that style changes defect type rather than making items uniformly easier.
- By dialogue language: The automated composite is higher for English dialogue at every granularity, never by more than 0.12.Nearly all movement sits on language-bearing dimensions; cross-shot character and cross-episode consistency do not move in the same direction.
- By input paradigm: Input paradigm is the largest average effect and reverses direction along the chain: first/last frame leads grid panel by 0.37 on single keyframes but trails by 0.42 on episode keyframes.Grid panel leads all three video granularities, though narrowly at single-shot video and short drama.
- By input paradigm: First/last-frame inputs lead keyframe F at 3.57 against 2.73, while grid panels lead episode-keyframe consistency at 3.53 against 3.11.Each paradigm is strongest where its supplied artefact constrains the criterion most directly.
E Agreement Slices of Automated Evaluation
Automated agreement depends strongly on stage, axis and modality, with image-side agreement exceeding the human baseline but video-side agreement remaining below it. Across dimensions, automated and human judgments tend to succeed and fail together.
- By stage: Per-item agreement is 2.35× on storyboard design, 1.19× on single keyframes and 1.48× on episode keyframes, then 0.79×, 0.85× and 0.85× across video granularities.The 1× crossing occurs between image and video stages.
- By axis: F, C, P, E and Q have pooled agreement values of 0.423, 0.235, 0.146, 0.132 and 0.121 respectively.Only F exceeds the human baseline net, by +0.107 and in 9 of 12 dimensions; Q wins none and is 0.086 below.
- By modality: E is +0.175 above the human baseline on storyboard text but 0.190 below it on video, while F declines from +0.317 on text to +0.012 on time and audio.An axis validated in one modality does not necessarily transfer to another.
- By human consensus: Across 63 leaf dimensions, automated and human agreement correlates at PLCC 0.418 and SRCC 0.416.The two systems are accurate on the same dimensions rather than providing complementary strengths.
- By human consensus: The human baseline is the appropriate filter for retiring dimensions because irreproducible human judgments provide no stable target for automated alignment.
F.1 Equal scores, different failure modes
Composite scores conceal distinct, spatially localised failure modes and miss cross-shot consistency failures that are visible only across an episode. Attribution and controlled storyboard branching expose these differences.
- Equal scores, different failure modes: 1.33 on I-B2 and 1.67 on I-C1 can represent different defects despite similar low scores.The examples localise failures to limbs and faces, while the corresponding alternative dimensions score 3.0 and 3.33.
- Equal scores, different failure modes: Four items scoring 1.3–1.7 on one dimension fail in disjoint locations, including faces, limbs, unwanted people and hands.In each pair, models trade which dimension they fail, showing why forced attribution is needed.
- Equal scores, different failure modes: The model losing identity holds action at 3.0, while the model losing action holds identity at 3.33.
- Cross-shot consistency: The same episode and character sheets yield a 2.66-point MI-D1 gap even though every failing panel is individually defensible.The failure is relational: a shot becomes wrong relative to another shot rather than in isolation.
- Storyboard control: Changing storyboard shot count from three to six restores dialogue fidelity, while increasing from six to twelve improves only executability.At three shots, all annotators scored executability and shooting rhythm at 1.00 with zero disagreement.
F.4 Each input paradigm carries its own failure mode
The first/last-frame and grid paradigms exhibit distinct, paradigm-specific failure modes, while upstream scene-setting defects can persist through video generation into the finished drama.
- F.4 Each input paradigm carries its own failure mode: 1/10 versus 9/15 failing dimensions distinguishes first/last-frame from grid inputs, respectively.The two paradigms fail differently rather than according to visual style.
- F.4 Each input paradigm carries its own failure mode: Grid panels can omit people because one image must fill several cells, whereas first/last-frame shots can change person identity between frames.These failures arise from how each paradigm represents or produces its visual inputs.
- F.4 Each input paradigm carries its own failure mode: A grid-based pixverse-c1 shot scored 1.67 on keyframe adherence, with all three annotators marking localized deviations from the given keyframe.The defects were localized to 0.29–11.04 s, 2.98–4.87 s, and 6.56–8.15 s.
- F.4 Each input paradigm carries its own failure mode: Because the prompt specifies only camera moves and dialogue, grid-shot appearance and blocking depend solely on the split image and adhere more weakly than first/last-frame pairs.Splitting one image into cells weakens the visual anchor for the shot.
- F.4 Each input paradigm carries its own failure mode: Upstream set-dressing drift is not repaired by the video stage; it is re-expressed and survives into the short drama.The benchmark’s one-stage replacement design isolates this cascading behavior while keeping other stages unchanged.
F.6 hy3 reproduces the script and does not dramatise it
The storyboard comparison separates script reproduction from dramatization, while later results show that model behavior can trade single-shot quality for cross-shot coherence and that perceptual scoring may overestimate quality.
- F.6 hy3 reproduces the script and does not dramatise it: 3.43 overall places hy3 behind doubao-seed-2.1-pro at 3.60, with hy3 stronger on script reproduction but weaker on television-oriented expressiveness.Hy3 reaches 4.44 on dialogue fidelity, while emotional expression is 3.64 and audiovisual style treatment is 3.91.
- F.6 hy3 reproduces the script and does not dramatise it: On one identical episode script, hy3 reproduced all 20 dialogue turns verbatim, whereas doubao-seed-2.1-pro reproduced 3 of them verbatim.Hy3 used 10 shots with limited additions; doubao-seed-2.1-pro used 7 shots and wrote them frame by frame.
- F.6 hy3 reproduces the script and does not dramatise it: The storyboard difference reflects a fidelity/expressiveness trade-off rather than a defect in either system.Across 533 items, both reproduce-the-script dimensions favor hy3, while none of the four expressiveness dimensions does.
- F.6 hy3 reproduces the script and does not dramatise it: Seedance-2.5 loses on single-shot video but gains at episode and short-drama granularity, including +1.143 in cross-shot style consistency.Audio-visual sync changes by −0.658, and the reported mechanism is more literal prompt following.
- F.6 hy3 reproduces the script and does not dramatise it: A reference-checked item receives correct scores on every referenced dimension but is overestimated by 3.67 points on the purely perceptual dimension.This item-level failure matches the broader metric-validation pattern.