Source-linked AI summary

CineForge: Self-Improving Agents for Long-Horizon Video Generation

Junxiang Liu, Lin Wang, Haiyu Shi, Hongxu Ma, Xiaoyu Yang, Chunjie Chen, Xiaoxiao Xu, Kaiqiao Zhan, Boao Wang, Shuizhou Shi, Tianyun Zhu, Jie Li, Jiangtong Li

arXiv:2608.29621v1cs.CVcs.AI

TL;DR

Long-horizon story video generation lacks persistent, stage-targeted improvement across productions despite requiring coordination across interdependent scenes. CineForge couples structured trajectory-based production with replay-gated policy evolution and evaluates it with CineScope, reporting higher story-level quality, benchmark gains, and reduced review workload.

  • Problem

    Long-horizon story production requires preserving narrative, character, spatial, directorial, and pacing state across interdependent scenes, while existing adaptation remains limited to instance repair or reusable-skill evolution.

  • Method

    CineForge-Produce records typed production states and canonical trajectories, while CineForge-Evolve applies CPPE to convert recurring cases into bounded patches validated by replay and paired evaluation.

  • Results

    The evolved policy reaches 4.380 CineScope-Metric versus 4.024 in Round 0, outperforms three long-video baselines with consistent ScriptAgent gains, and reduces review LLM calls on new stories.

  • Takeaways & Limitations

    Production trajectories provide actionable experience for cumulative policy improvement across long-form storytelling tasks.

  • Takeaways & Limitations

    The structural certificate does not establish causal attribution or rendered-video improvement, which is evaluated empirically with frozen outcome evaluators.

Abstract

from arXiv · show

Long-horizon story-driven video generation requires a production agent to coordinate narrative decomposition, state tracking, shot design, prompt construction, rendering, and revision across interdependent scenes. Existing adaptive video systems primarily refine requests or reusable skills, leaving recurring production failures disconnected from persistent, stage-targeted improvements across stories. We introduce CineForge, a self-evolving video-production agent framework that couples CineForge-Produce for video generation with CineForge-Evolve for cross-story policy evolution. CineForge-Produce organizes each source story into typed narrative, character, spatial, and cinematic states, uses them to coordinate asset and clip generation, and records the process as a canonical production trajectory. CineForge-Evolve applies Case-to-Pattern-to-Policy Evolution (CPPE) to review trajectory evidence, consolidate recurrent findings into bounded stage-local patches, and deploy validated updates through structural replay and confidence-controlled paired evaluation. To measure complete story realization, we introduce CineScope, which combines a 100-script CineScope-Data suite with a human-aligned, multiscale CineScope-Metric spanning causal state, directorial orchestration, pacing and resource allocation, and character arc. Across CineScope-Data and two public benchmarks, the evolved CineForge policy improves CineScope-Metric from 4.024 to 4.380, outperforms three long-video baselines with consistent gains under ScriptAgent, and reduces review LLM calls by 37.0% on new stories. These results establish production trajectories as actionable experience for video agents that improve cumulatively across long-form storytelling tasks.

1 INTRODUCTION

CineForge addresses long-horizon story production by evolving persistent production policies from trajectory evidence across stories. It combines structured production, CPPE-based policy updates, and CineScope evaluation, with reported gains in quality and review efficiency.

  • Long-horizon story videos require preserving event coverage, character and spatial state, directorial intent, and pacing across interdependent scenes and shots.
  • Existing adaptive systems repair individual generations or evolve reusable skills, whereas CineForge targets persistent trajectory-level policy evolution across stories.
  • CineForge-Produce structures stories into narrative, character, spatial, and cinematic states, while CineForge-Evolve converts production trajectories into versioned policy improvements.
  • CineScope combines 100 long-form scripts with a human-aligned multiscale metric covering causal state, directorial orchestration, pacing and resource allocation, and character arc.
  • 4.380 CineScope-Metric, up from 4.024 in Round 0, accompanied consistent gains over three long-video baselines and under ScriptAgent.
  • CineForge’s contributions include CPPE, CineScope, persistent policy improvement across stories, transfer to new stories, and reduced production-review workload.

2 PRELIMINARIES

The framework represents each production with a typed trajectory and controls policy evolution through bounded, stage-local patches. Replay and paired evaluation separate structural admission from final-video quality measurement.

  • CineForge-Produce executes a persistent policy on each story and returns both the generated video and a canonical trajectory containing stage inputs, outputs, validators, provenance, and dependencies.
  • The persistent policy contains stage contracts, prompt-rendering rules, production thresholds, review-routing rules, and bounded style profiles, with accepted patches creating the next version.
  • Findings record observed symptoms and evidence, while normalized issues add the earliest supported causal stage, failure family, repair specification, confidence, lineage, and status.
  • Recurring patterns require sufficiently confident, policy-editable issues across distinct story-level opportunities and compile into typed, stage-local patches with rollback metadata.
  • Every candidate patch undergoes deterministic structural replay, with patches affecting stochastic generation additionally evaluated through paired stochastic comparison.
  • Final-video quality is measured only after the policy is fixed, using frozen outcome evaluators such as CineScope-Metric and ScriptAgent.

3 CINEFORGE FRAMEWORK

CineForge couples observable story production with cross-story CPPE evolution: trajectories generate evidence, recurring patterns yield bounded patches, and replay gates control policy versioning. Its formal boundary distinguishes structural admission from empirical rendered-video quality.

  • 3.1 OVERVIEW: CineForge-Produce generates story videos and canonical trajectories, while CineForge-Evolve aggregates trajectory evidence, compiles bounded patches, and validates the next policy version.
  • 3.2 PRODUCTION: Typed production states organize narrative atoms across episode, scene, and shot units while tracking character, spatial, cinematic, and prompt fields for coverage and continuity checks.
  • 3.2 PRODUCTION: The production path coordinates assets, provider-specific prompts, clip generation, visual-anchor propagation, composition, deterministic validation, and bounded Review-Rewrite repair.
  • 3.3 CPPE: CPPE forms production cases from pipeline, local-video, and global-video findings, retaining linked evidence, provenance, repair actions, and outcomes.
  • 3.3 CPPE: Cross-story cases are grouped by causal stage, failure family, and context signature; recurring patterns then compile into typed patches targeting contracts, fields, thresholds, review routes, rendering rules, or style profiles.
  • 3.4 EVALUATION BOUNDARY: Replay evaluates frozen candidates against protected guard metrics, while fresh admission stories support structural certification and stochastic evaluation is added for patches invoking new generation calls.
  • 3.4 EVALUATION BOUNDARY: The certificate requires fresh admission stories after candidate freezing and does not establish causal attribution or rendered-video improvement; final-video quality is evaluated empirically.

4 CINESCOPE BENCHMARK

CineScope evaluates complete story realization rather than isolated clip quality through a 100-script suite and a frozen, human-aligned multiscale metric. Its criteria cover causal development, directorial organization, pacing, and character progression across dependent scenes.

  • CineScope combines CineScope-Data, a 100-script suite designed to expose long-range dependencies, with CineScope-Metric, a frozen output-only measure of complete-episode story realization.
  • The 100 scripts span diverse genres, narrative structures, character configurations, temporal spans, and scene-transition patterns, exercising continuity, causal development, pacing, and character progression.
  • CineScope-Metric evaluates a source narrative, task package, and rendered episode using complete-episode, 150-second, and 30-second evidence.
  • Eligibility excludes temporal windows lacking assessable source-text evidence rather than scoring them as failures, and reported scores renormalize over eligible scopes.
  • The metric’s twenty criteria are grouped into Causal State, Directorial Orchestration, Pacing and Resource Allocation, and Character Arc.
  • CineScope-Metric is calibrated on development data, frozen before held-out evaluation, and excluded from CPPE formation, patch admission, checkpoint selection, and stopping.

5 EXPERIMENTS

The experiments validate CineScope-Metric against expert judgments, compare CineForge with video-generation and optimization baselines, and test whether CPPE improves quality and review efficiency across evolution rounds.

  • CineScope-Metric human alignment: CineScope-Metric reaches 0.710 Spearman correlation and 81.0% pairwise agreement with held-out expert judgments.The metric is evaluated on system-unseen videos and remains frozen during subsequent system evaluation.
  • Video generation agents: CineForge-Produce achieves the highest CineScope-Metric weighted score, 4.024, and leads all four Global dimensions against three long-video baselines.The comparison uses the system-test split under a common generation and evaluation protocol.
  • Self-evolution strategies: CineForge (Full) reaches 4.380, compared with 4.024 for the common CineForge-Produce baseline, exceeding GenMAC at 4.118 and VideoWeaver at 4.183.All optimization methods share the same production baseline, backbone, generation budget, and final test protocol.
  • Self-evolution scaling: Across Rounds 0–3, unseen CineScope-Metric rises from 4.0243 to 4.3797, while unseen ScriptAgent rises from 3.9130 to 4.1009.The system-test split is excluded from evolution, and evaluation rollouts cannot update persistent state or policy selection.
  • Self-evolution scaling: Unseen Review LLM calls fall from 46.68 to 35.63 per episode by Round 1, while review issues fall from 103.79 to 71.87.Both efficiency measures continue declining across later rounds under the same frozen reviewer configuration.

6 CONCLUSION

CineForge frames long-horizon story-to-video generation as cross-production self-evolution, combining trajectory-aware production, persistent policy updates, and episode-level evaluation. The framework complements existing planning and consistency methods by targeting recurring cross-stage failures across stories.

  • Conclusion: CineForge couples trajectory-aware production with CPPE-based policy evolution, validating bounded patches before they change future behavior.CineForge-Produce records canonical production trajectories, while CineForge-Evolve localizes failures and admits updates through replay and confidence-controlled evaluation.
  • Conclusion: CineScope evaluates complete episodes across causal state, directorial orchestration, pacing and resource allocation, and character arcs.Its data suite and metric are designed for story-level dependencies rather than clip-level perceptual quality alone.
  • Conclusion: Existing story-visualization and long-video systems establish planning, memory, consistency, and temporal context, but do not define persistent updates for recurring cross-stage production failures.CineForge addresses this gap through production trajectories and cross-story policy evolution.

A.2 LONG-VIDEO AND NARRATIVE EVALUATION

The paper positions CineForge against long-video evaluation and adaptive-agent methods by emphasizing persistent, trajectory-level policy evolution rather than only current-request repair or reusable-skill updates.

  • Narrative video benchmarks increasingly evaluate event completion, sequential tasks, story visualization, narrative richness, and event-level alignment.
  • CineScope-Metric complements specialized perceptual and narrative benchmarks with four episode-level Global dimensions and twenty human-aligned criteria.It does not replace specialized perceptual, audio-visual, entity-consistency, or cinematic taxonomies.
  • Existing adaptive video systems operate at instance-level repair or reusable-skill evolution, whereas CineForge targets persistent production-policy evolution across stories.The distinction is framed as current-generation refinement versus cross-story adaptation.
  • CineForge converts recurring failures into typed, bounded patches over prompt rules, thresholds, genre profiles, or review policies, with risk levels and regression gates.This separates in-run repair from persistent rule evolution.
  • Unlike generic prompt or workflow optimizers, CineForge updates a multi-stage production policy using explicit intermediate-state failures and auditable replay-tested edits.The policy can affect decomposition, shot planning, spatial consistency, safety rewriting, and prompt rendering.
  • CineForge claims no new video backbone, diffusion architecture, generic LLM optimizer, or open-ended self-modifying agent.Its stated novelty is the combination of explicit states, cross-production case memory, typed patches, and replay gates.

B.1.1 DATA CONSTRUCTION AND SPLIT ALLOCATION

CineScope-Data is a 100-script long-form suite designed to test cross-scene dependencies, pacing, character development, and state continuity through separated evaluation roles.

  • CineScope-Data contains 100 long-form scripts spanning diverse genres, narrative organizations, character configurations, temporal spans, and scene-transition patterns.Each item includes a complete script or script-like narrative and a stable identifier for deriving task packages.
  • The 100 scripts are allocated across evolution collection, human alignment, replay admission, and system testing using disjoint role-specific splits.The allocation prevents collection, admission, and final testing from sharing the same designated items.
  • The evolution-collection split contains 60 CineScope-Data scripts for case formation and candidate compilation.
  • The human-alignment split contains 30 scripts, with 10 from each source, and supports metric calibration plus held-out alignment evaluation.The sources are CineScope-Data, AnimeShooter, and ViStoryBench.
  • The replay-admission split contains 45 scripts used only for fresh candidate admission, while the disjoint 45-script system-test split supplies system comparisons and unseen measurements.
  • The corpus draws on publicly accessible or explicitly authorized long-form novels, stories, and screenplays, with item-level provenance and licensing details planned for release.The paper withholds individual platform names pending final corpus audit.

B.1.2 UNIFIED VIDEO-PRODUCTION AND AUTOMATIC-EVALUATION PROTOCOL

The evaluation protocol standardizes production backbones, generation randomness, automatic scoring, human alignment, and evaluator wrappers so comparisons target production frameworks rather than underlying models.

  • All compared production systems use GLM 5.1 for text-side calls and Seedance 2.0 for video generation.
  • CineScope-Metric uses a frozen Qwen3.6-Plus evaluator and gives systems the same source material, generation budget, output specification, and post-processing policy.This configuration is intended to isolate production frameworks or optimization mechanisms.
  • Each story-system pair produces two independently randomized videos, each evaluated twice per automatic evaluator, yielding 2 × 2 = 4 call-level scores.The four scores are equally averaged into the story-system score for that evaluator.
  • Human evaluation uses ten trained university evaluators and the same 20 definitions and 0-5 anchors as CineScope-Metric.Each generated video receives at least five independent ratings, balanced by source, generator, and replicate.
  • Development ratings calibrate temporal-scale weights, aggregation weights, confidence and abstention thresholds, and evaluator prompts, retaining preregistered defaults when candidates do not improve alignment.
  • VBench-Long and ScriptAgent use frozen input-only wrappers that preserve their original metric definitions while exposing partitioned long-video inputs.The wrappers add no CineScope-Metric dimensions, production traces, or optimization feedback.

B.1.5 SUPPLEMENTARY GENERATOR-EVALUATION RESULTS

The supplementary protocol separates story-level and perceptual evaluation from production-policy admission, while enforcing frozen splits, evidence-based scoring, bounded updates, and replay safeguards.

  • VBench-Long reports complementary perceptual properties, while ScriptAgent provides an additional story-level critic view on the same generator comparison and test set.Neither evaluator is combined with CineScope-Metric or used as optimization feedback.
  • CineScope-Metric receives source materials and final rendered output but not production trajectories, review findings, evidence memory, candidate rules, or replay decisions.This keeps the metric independent of the optimization process.
  • The metric uses a 0-5 scale with eligibility gates, evidence timestamps, confidence, and abstention; unsupported evidence is rerun rather than scored without video grounding.
  • Its rubric contains four Global dimensions, each aggregating five scored subdimensions, including causal state, directorial orchestration, pacing and resource allocation, and character-related criteria.
  • Multiscale scoring combines aligned 150-second and 30-second windows with complete-episode evidence, while renormalizing frozen temporal weights over eligible evidence.Long-range criteria rely primarily on complete-episode evidence unless a relevant milestone appears in an aligned window.
  • CineForge-Evolve restricts automatic updates to bounded, stage-local targets, routing higher-risk or system-level changes to stricter replay, human approval, or exclusion.
  • Low-risk structural deployment rejects any candidate with a negative protected deterministic delta, while stochastic-impact candidates require the corresponding paired evaluation pathway.
  • Replay admission freezes candidate classes before fresh stories are revealed, prevents replay artifacts from entering memory or testing, and selects at most one patch per round.

B.6 STOCHASTIC, RECURRENCE, AND TRANSFER GUARANTEES

The guarantees bound stochastic admission, recurrence-trigger reliability, and cause-level transfer under separately stated conditions. Together, they distinguish admission evidence from rendered-video quality claims and require empirical confidence support for transfer claims.

  • With probability at least 1 −δV, accepted stochastic-impact patches satisfy the stated expected CPPE admission-score improvement guarantee.The bound applies simultaneously across adaptive rounds when the per-round failure probabilities satisfy the proposition’s conditions.
  • The stochastic admission proposition concerns CPPE admission only and makes no CineScope-Metric claim.The theory-to-evidence map separates structural, stochastic, recurrence, and transfer claims under their respective conditions.
  • Positive held-out transfer occurs if and only if total removed error mass exceeds total newly introduced error mass.The decomposition separates errors removed after an update from errors newly introduced by it.
  • An empirical transfer claim additionally requires the lower endpoint of the story-bootstrap 95% confidence interval for ∆transfer to exceed zero.The special-case expression applies only under the stated no-new-target-error, no-removed-non-target-error, and unchanged-unknown-mass conditions.

B.7 DETAILED METHOD NOTES

CineForge exposes a typed, auditable production graph and converts trajectory evidence into stage-local policy candidates. Replay gates, deterministic safeguards, and stochastic admission controls govern which patches can update the persistent policy.

  • CineForge-Produce: CineForge-Produce decomposes stories into narrative atoms, coverage maps, typed states, and auditable stage records before rendering.Deterministic checks enforce coverage and preserve identifiers, dependencies, and validation records, while bounded creative infill supplies semantic fields.
  • Trajectory records: The canonical trajectory stores versioned plans, prompts, assets, outputs, validator findings, repairs, provenance, and dependency links for later diagnosis and replay.Paired replay fixes task inputs and uses the changed field’s downstream dependency slice.
  • CineForge-Evolve: CPPE traces video symptoms backward through the production chain and assigns the earliest evidence-supported stage as the root-cause stage.This separates where a defect becomes visible from where the process introduced it; unsupported backend-only symptoms remain attributed to the backend/provider.
  • Policy evolution: Recurring diagnosed failures become finite, typed, stage-local patch candidates with targets, edit types, risk levels, dependency slices, and rollback metadata.One-off findings remain auditable but do not support persistent evolution.
  • Replay-gated deployment: Every candidate faces deterministic replay gates, while stochastic-impact patches additionally undergo repeated paired admission before at most one accepted patch updates the policy.Rejected candidates leave the policy unchanged and are recorded in the audit.

B.9 FULL PROOFS

The proofs formalize dependency confinement, replay-based structural guarantees, stochastic admission bounds, recurrence reliability, and transfer decomposition. Their conclusions remain conditional on fixed protocols, independence or calibration assumptions, and separation from final-video quality evaluation.

  • Replay assumptions: Exact replay applies only when affected downstream transformations are deterministic; stochastic model calls are handled by a separate admission proposition.The cache boundary fixes upstream artifacts, task inputs, and validator versions while editable rules remain in the policy.
  • Scope boundary: Theorem 1 is stated for a fixed deployment distribution, so materially changing that distribution requires rebuilding the replay pool and certificate.
  • Dependency closure: Dependency closure reaches a least fixed point over changed trace fields and downstream transformations, proving confinement of deterministic changes to the audited dependency slice.The construction stabilizes because the finite trace-field set grows monotonically.
  • Dependency closure: The confinement result does not identify the cause of an observed failure.It establishes only that changed deterministic fields lie within the executed dependency closure of the audited patch support.
  • Structural certificate: Theorem 1’s cumulative structural-welfare guarantee assumes candidate sets are fixed before fresh replay-admission stories are revealed.The proof uses conditional bounded replay variables, candidate-wise union bounds, and a union bound across adaptive rounds.
  • Structural certificate: The structural certificate does not establish rendered-video improvement because replay welfare is not linked to final-video utility.Rendered outcomes are evaluated empirically after the policy version is fixed.
  • Transfer: Transfer is positive exactly when removed error mass exceeds added error mass under matched pre/post evaluation.The proof decomposes the total difference across the frozen operational cause set.
  • Recurrence: Recurrence reliability bounds false triggers and missed systematic defects under conditionally independent story-level opportunities, but correlated streams require empirical calibration outside the lemma.The family-wise false-trigger bound scales with the number of monitored pattern families.
Loading 2608.29621v1…