Source-linked AI summary
AniMaster: From Story Texts to Animated Videos via Cinematic Script Generation and Interactive Authoring
Ruiqi Yu, Dekun Qian, Jiale Xu, Sizhe Cheng, Yize Li, Xiangyang Wu, Zhiguang Zhou, Wei Chen, Yong Wang
TL;DR
Everyday creators struggle to produce polished long-form animated videos from brief stories because translating narrative intent into cinematic scripts and visual parameters requires specialized expertise. AniMaster addresses this with a narratology- and film-studies-based Story–Script–Video framework, staged translation, and interactive editing. Evaluations report effective and usable authoring support, while the method remains limited for nonlinear and cross-scene narrative planning.
Problem
Everyday creators lack expertise for translating brief story texts into cinematic scripts, videos, and visual parameters such as composition, camera controls, and shot sequencing.
Method
AniMaster uses Events, Beats, and Shots in a three-layer framework, translating story events into editable beat sequences and beats into executable visual parameters.
Results
The evaluation found AniMaster effective and usable, with intermediate representations supporting professional knowledge scaffolding and shifting attention toward higher-level shot and narrative decisions.
Takeaways & Limitations
Visible intermediate representations can support both novice learning and experienced creators’ organization across guided-to-autonomous authoring.
Takeaways & Limitations
Translation is conservative for experimental expression, with limited support for nonlinear narratives and cross-scene dependencies beyond short approximately 170-word stories.
Abstract
from arXiv · showhide
Recent advances in Video Generation Models (VGMs) have demonstrated strong capabilities in producing short video clips. However, it is still challenging for everyday creators to leverage these models to produce polished long-form animated videos from brief story texts. Informed by a formative study with both novice creators and film experts, we identify two major challenges of interactive video authoring: (1) the lack of expertise in translating free-form story texts to professional cinematic scripts and finally high-quality animated videos, and (2) the absence of effective ways to convey video design intents to key variables of visual storytelling, such as shot composition, camera controls and shot sequencing. Drawing on narratology and film studies, we propose a three-layer design framework that defines the key design dimensions across three layers (i.e., story texts, cinematic scripts, and animated videos) as well as the translation between them. Built on this framework, we present AniMaster, a VGM-powered authoring tool to enable everyday creators to easily produce smooth animated videos from free-form story texts. AniMaster automatically expands brief story texts to detailed cinematic scripts, and further translates cinematic scripts into polished videos by following professional visual storytelling principles. It also allows users to interactively edit the scripts and refine the generated videos via text instructions and intuitive interactions. We extensively evaluated AniMaster through an in-depth user study with 16 participants, two case studies, and expert interviews with 2 film professionals. The results demonstrate the effectiveness and usability of AniMaster in helping everyday creators create polished animated videos from free-form story texts.
1 Introduction
Existing VGMs can generate short clips, but everyday creators still struggle to turn brief stories into polished long-form animated videos. AniMaster addresses this gap by making cinematic translation stages and decisions explicit, editable, and actionable.
- Motivation: Everyday creators lack the filmmaking expertise needed to translate free-form story texts into professional scripts and polished animated videos.The challenge persists despite recent advances in video generation models.
- Motivation: Prior AI video-authoring systems do not reflect professional filmmaking workflows or incorporate foundational narratological and cinematic guidelines.This limits their support for structured story-to-video creation.
- Design Challenges: The formative study identified translation from story texts to scripts and videos, and mapping design intents to shot composition, camera controls, and sequencing, as major challenges.The study included novice creators and film experts.
- AniMaster: AniMaster uses a three-layer framework and staged translation to expand story events into editable beats, then translate beats into executable shots.The tool also exposes intermediate decisions through a canvas-based interface for inspection and revision.
- Evaluation: Evaluation with 16 participants, two case studies, and two film professionals demonstrated AniMaster’s effectiveness and usability for creating polished animated videos.The evaluation also suggests value in making professional narrative and cinematic knowledge explicit and computationally actionable.
2 Related Work
Related work spans narrative structuring, shot-level authoring, and generative cinematic production, but existing systems still leave cinematic planning and sequence-level coherence insufficiently controlled.
- Research Landscape: Prior work organizes the story-to-video pipeline into narrative and scriptwriting, previsualization and authoring interfaces, and generative cinematic production.These groups address structure, rapid shot-level iteration, and translating creative intent into producible video.
- Cross-Modal Authoring: Cross-modal systems map textual units to visual units, but existing approaches lack systematic encoding of narratological and cinematic knowledge into an intermediate representation.This limits direct text-to-video bridging for structured cinematic authoring.
- Generative Cinematic Production: Recent work improves visual consistency, camera control, text-driven previsualization, and composition optimization across storyboard and spatial-temporal planning.Film idioms have also been formalized to map narrative functions to camera expressions.
- Generative Cinematic Production: Video-generation research advances generation quality and temporal coherence, while end-to-end cinematic pipelines chain script planning, shot design, and keyframe generation.Creators still cannot reliably constrain shot scheduling and montage logic with directorial thinking.
- Open Gap: Narrative coherence and pacing at the sequence level remain difficult to guarantee in current generative cinematic production.The limitation concerns long-form sequence organization rather than only individual clip quality.
3 Design Framework
The framework derives a three-step Story–Script–Video pipeline from formative studies and film scholarship, formalizing intermediate representations and translations that support coherent authoring.
- Formative Study: A formative study combined user observation of 12 non-experts with interviews involving 6 filmmaking experts to identify translation difficulties.The analysis coded observed behaviors and thematically analyzed expert interviews.
- Design Challenges: Novice creators could generate locally reasonable images or clips but struggled with ordering events, decomposing moments, and connecting adjacent shots coherently.Their outputs often lacked clear progression and coherence.
- Professional Pipeline: Experts described a professional workflow that identifies story events, decides how to tell them, and then converts those decisions into concrete shots.Each stage requires a different mode of thinking.
- Design Challenges: The framework targets two challenges: translating free-form stories into cinematic scripts and videos, and mapping narrative intent to visual storytelling parameters.The latter includes shot composition, camera controls, and shot sequencing.
- Three-Layer Framework: The framework defines Story Space, Script Space, and Video Space as layers for story content, audience-oriented orchestration, and executable visual parameters.Figure 2 represents these layers with Events, Beats, and Shots connected by translation logic.
3.3 Space Dimensions
The design framework assigns distinct operational units and dimensions to Story, Script, and Video Space, then links them through rules for narrative unfolding and visual translation.
- Dimension Derivation: The framework screens dimensions for Extractability, Operability, and Renderability so they can be extracted, edited, and executed by VGMs.These criteria connect theoretical concepts with practical authoring and generation requirements.
- Story Space: Story Space uses Events as its basic unit, representing what happens through actors, locations, events, and implicitly ordered time.The model draws on Bal’s fabula and treats Event as the most self-contained unit.
- Script Space: Script Space uses Beats as the smallest structural units carrying narrative functions, with each Event typically unfolding into an ordered sequence of Beats.Beat Types identify functions such as establishing environments or depicting actions, while four dimensions describe each Beat.
- Video Space: Video Space uses Shots as executable visual units organized across cinematography, mise-en-scène, editing, and sound.The nine parameters include shot size, camera angle, camera movement, composition, lighting, color tone, transition, duration, and audio.
- Translation: Story-to-script translation unfolds Events into ordered Beats using contextual grounding, causal progression, and narrative pacing.Seven Beat Types operationalize these conditions through environment, presentation, action, dialogue, reveal, reaction, and pause.
- Translation: Script-to-video translation turns Beats into executable parameters through Shot Continuity and Narrative Function Mapping.Continuity governs relations between shots, while Beat Type determines shot patterns, shot counts, and default visual parameters.
4 AniMaster
AniMaster helps everyday creators progressively build cinematic animations from story texts.
- AniMaster is a system for progressively building cinematic animations from story texts.
- The system is built on a three-layer design framework.
- AniMaster supports the transformation from story texts toward cinematic animations.
4.1 Interface Overview
AniMaster combines a cross-layer Script Reader with a Main Canvas that supports navigation from narrative structure to shot-level production.
- Script Reader: The Script Reader aligns Events in Story, Beats, and Shots across three columns.This cross-layer view lets creators trace how a story moment expands across layers.
- Main Canvas: The Main Canvas integrates navigation, breakdown, preview, asset management, and shot-composition workspaces.Its components include the Canvas Map, Breakdown view, Video Preview, Asset Library, and Beat Canvas.
- Semantic Zooming: Semantic zooming moves from global narrative structure through Beat-level planning to shot-level production.This progression matches AniMaster’s three design layers.
4.2 Feature Walkthrough
AniMaster guides creators from story-level organization to editable Beats and shots, then supports asset binding, parameter control, generation, and hierarchical refinement.
- From Story to Event: AniMaster identifies characters, locations, and events, presenting a global spatial overview of the story.Locations are arranged left-to-right according to the story’s spatial trajectory.
- From Event to Beat: An Event expands into a Beat sequence, such as Environment, Presentation, Dialogue, and Reaction.The Script Reader links these Beats to the original story text.
- From Beat to Shot: A Beat opens a local workspace where generated Shot nodes, frames, videos, and assets form an editable production pipeline.Creators can start with recommended Gen-mode parameters or switch to DIY mode.
- Crafting the Shots: Creators bind reusable character and location assets, whose reference images support visual consistency across shots.The Asset Library supplies variants that can be dragged into the Beat Canvas.
- Crafting the Shots: The Inspector edits nine visual parameters, while the Composition Editor adjusts element positions within a frame.Creators can spatially reposition characters and scene elements directly.
- Frame and Video Generation: Generation proceeds from a static Frame to a Video clip, with continuation chaining successive clips from the last keyframe.A Frame connected to a Video supplies its visual reference for generation.
- Cinematic Cues: Cinematic Cues annotate shot-to-shot relationships and let creators express directorial intent for parameter inference.Examples include shot-size continuity, axis-rule compliance, spatial consistency, and “push to close-up.”
- Editing and Refining: Creators can insert suggested Beats, revise narrative intent, and selectively regenerate affected downstream results while preserving explicit overrides.Editing is supported at Event, Beat, and Shot levels.
4.3 Two-Stage Translation Pipeline
AniMaster uses a two-stage pipeline that expands story Events into editable Beats, realizes Beats as editable Shots, and generates videos through intermediate Frames.
- S2S Translation: S2S expands an Event into an editable Beat sequence for revision before shot authoring.The sequence appears in the Breakdown and aligns with the Script Reader.
- S2V Translation: S2V realizes each Beat as editable Shot structures with initial visual parameters in the Beat Canvas.These structures are instantiated as Frame and Video nodes.
- Generation: Generation proceeds from a static Frame to a Video clip rather than directly from raw story text.Intermediate narrative and visual decisions remain available for interactive editing.
4.4 Implementation
AniMaster combines configurable language, image, video, and frame-interpolation components in a Vue/TypeScript and FastAPI/Python implementation.
- The front end uses Vue 3 and TypeScript, while the back end uses FastAPI and Python.
5 Evaluation
AniMaster was evaluated through a within-subjects user study, case studies, and expert interviews. Results indicate improved creativity support and authoring experience, professional alignment, and support for distinct creator workflows, while revealing limits for experimental and long-range coherence.
- Evaluation Design: Three complementary studies addressed creativity support, creator interaction patterns, and alignment with professional filmmaking practice.The evaluation included a within-subjects comparison, two exploratory case studies, and interviews with two film professionals.
- User Study: 16 participants used AniMaster and a Baseline condition to create complete animated videos in a within-subjects design.The Baseline used the same generative model but lacked AniMaster’s structured representations, semantic zooming, and rule-guided interactions.
- User Study: 7.16 vs. 5.70 overall CSI favored AniMaster significantly (p = .012, r = .71), while NASA-TLX scores did not differ significantly (p = .597).All 9 bipolar preference items favored AniMaster, and all 16 participants preferred it.
- Case Studies: Participants used AniMaster as scaffolding or an assistant, progressively moving from system suggestions toward deliberate shot-level judgments and local refinement.The two cases produced complete works in about 4 hours despite differing creative strategies.
- Expert Interviews: Both experts found that generated shots served distinct narrative functions, including causal progression, atmosphere, and pacing through pauses.Their assessments supported the professional validity of the system’s shot-level narrative logic.
- Expert Interviews: Experts recognized correspondence between AniMaster’s three-layer workflow and professional practice, but found its translation rules conservative for experimental expression.The authoring logic was considered most effective for plot-driven, continuity-based narratives, with long-range coherence and cross-shot consistency remaining open issues.
6 Discussion
AniMaster’s structured intermediate representations support professional knowledge scaffolding, redirected creative effort, and narrative-structured exploration. However, its translation rules and generation pipeline remain limited in experimental storytelling, cross-scene planning, and long-range visual consistency.
- Professional Knowledge Scaffolding: Beats, Shots, and visual parameters make professional filmmaking decisions explicit and editable for creators with different experience levels.Participants described the AI as a collaborator, while case studies positioned the structure as either a teacher or assistant.
- Cognitive Load Redistribution: AniMaster did not reduce total cognitive effort, but shifted attention from low-level prompt iteration toward shot logic and narrative pacing.NASA-TLX scores did not differ significantly (p= .597), while observed authoring effort moved toward higher-level creative decisions.
- Narrative-Structured Free-Form Authoring: Semantic zoom levels support spatial, exploratory authoring while narrative structure maintains global coherence across the hierarchy.The Main Canvas provides Overview, Breakdown, and Beat Canvas levels for navigating and editing narrative elements.
- Theoretical Scope of Translation: Translation rules grounded in classical Hollywood continuity provide a reliable baseline but conservatively support experimental and nonlinear narratives.Support for flashbacks and parallel montage remains limited, and Translation currently uses a uniform static prompt that does not adapt to creators.
- Cross-Scene Narrative Planning: AniMaster expands shots event-by-event, leaving cross-scene techniques and scalability to longer narratives with complex dependencies insufficiently validated.The controlled experiment used short stories of approximately 170 words, while foreshadowing, callbacks, and cross-Event narrative planning remain future work.
- Generation-Layer Constraints: Even coherent Beat sequences and shot plans may produce outputs with cross-shot inconsistency and weak scene-to-scene continuity.The system is therefore better understood as a structured authoring and previsualization scaffold than as a fully reliable end-to-end production pipeline.
7 Conclusion
AniMaster combines a three-layer Story–Script–Video framework with rule-based translation, semantic-zoom canvases, and editable intermediate representations. Evaluation with 16 participants, two case studies, and expert interviews found improved creativity support and authoring experience.
- 7 Conclusion: AniMaster organizes Story, Script, and Video dimensions through translation rules and editable Events, Beats, and Shots.Its canvas interface lets creators review, modify, and override system decisions at every layer.
- 7 Conclusion: A user study (N= 16), two case studies, and expert interviews demonstrate significantly improved creativity support and authoring experience.The structured representations support both guided authoring for novices and more autonomous organization for experienced creators.