Source-linked AI summary

A Text-Native Interface for Generative Video Authoring

Xingyu Bruce Liu, Mira Dontcheva, Dingzeyu Li

arXiv:2603.09072v1cs.HCcs.AI

TL;DR

Generative video authoring remains difficult because existing tools fragment scripts, prompts, visuals, audio, and timelines, while generated clips often lack narrative coherence and precise control. Doki addresses this gap with a text-native document that serves as a shared, editable, and executable representation for end-to-end video creation. A week-long study found faster idea-to-content workflows, improved coherence and narrative comprehension, while also exposing limits in precise visual control and asynchronous audio editing.

  • Problem

    Existing video tools fragment authoring across specialized interfaces, while generative video systems focus on isolated clips and provide limited support for coherent longer narratives.

  • Method

    Doki uses a text-native document as a shared human–AI representation in which users define assets, structure scenes, generate shots, refine edits, and add audio.

  • Results

    A week-long diary study with 10 participants produced 46 videos and found faster idea-to-content workflows, improved coherence through parameterization, and clearer narrative comprehension.

  • Takeaways & Limitations

    Doki supports accessible visual-story authoring for novices and complements experts’ rapid ideation and storyboarding without replacing high-fidelity production environments.

  • Takeaways & Limitations

    Doki provides limited precise visual control, often requiring repeated regeneration, and makes audio spanning multiple paragraphs or beginning asynchronously difficult.

Abstract

from arXiv · show

Everyone can write their stories in freeform text format -- it's something we all learn in school. Yet storytelling via video requires one to learn specialized and complicated tools. In this paper, we introduce Doki, a text-native interface for generative video authoring, aligning video creation with the natural process of text writing. In Doki, writing text is the primary interaction: within a single document, users define assets, structure scenes, create shots, refine edits, and add audio. We articulate the design principles of this text-first approach and demonstrate Doki's capabilities through a series of examples. To evaluate its real-world use, we conducted a week-long deployment study with participants of varying expertise in video authoring. This work contributes a fundamental shift in generative video interfaces, demonstrating a powerful and accessible new way to craft visual stories.

1 Introduction

Doki addresses the difficulty of authoring generative video by making text the primary interaction and consolidating creation in one structured document. Its design emphasizes flexible human–AI collaboration, consistency through parameterization, and simple end-to-end authoring.

  • Motivation: Traditional video tools support fine-grained editing but rely on complex interfaces, while generative video tools typically focus on isolated clips rather than complete narratives.Text-based editors reduce friction mainly for speech-track editing and retain multiple panes and timelines for visual authoring.
  • Doki: Doki makes writing the primary interaction for defining assets, structuring scenes, generating shots, refining edits, and adding audio within one document.The document functions both as a narrative and as an executable production script.
  • Design principles: Text serves as a shared human–AI medium that supports freeform ideation, generation, revision, and rapid in-place refinement.Humans can edit and review text while AI generates or suggests changes through the same representation.
  • Design principles: Doki consolidates scripts, prompts, visuals, audio, and timelines in one structured text representation to reduce the need to manage multiple views.The system supports many end-to-end creative tasks without aiming to replace professional toolchains completely.
  • Design principles: Parameterized definitions help preserve characters, styles, and other elements consistently across narratives and support longer, more cohesive stories.The representation includes propagation and context handling for cross-shot consistency.
  • Human–AI collaboration: The study found that creators retained a strong sense of authorship while delegating production tasks to AI, often viewing themselves as directors.The document remained readable, editable, transparent, and revisable by humans while executable by AI.

2 Related Work

Related video-authoring systems explore multiple panes, transcript editing, or limited additive generation, but they generally do not provide a single representation for complete multishot stories. Doki instead treats video creation as a dynamic document that is both narratively readable and executable as a production script.

  • Generative video models: Generative video models primarily generate individual short clips, leaving users to assemble longer narratives manually in separate editing environments.Doki uses these models as supporting engines while adding a higher-level interface for authoring complete video stories.
  • Interface paradigms: Many AI video-authoring systems distribute work across canvases, scripts, storyboards, and timelines, creating split-attention costs when users reconcile multiple views.This “bento box” paradigm contrasts with Doki’s text-native canonical representation.
  • Interface paradigms: Doki consolidates authoring steps into a single representation rather than distributing them across multiple panes.The document is the primary interface for coordinating the authoring process.
  • Additive workflows: Existing additive systems can synthesize assets from text but remain limited to specific shot types, whereas Doki supports additive authoring of complete multishot stories.Doki treats text as the substrate for generating shots and assets rather than only selecting or trimming existing footage.
  • Transcript-based editing: Transcript-centered tools are especially effective for dialogue-driven media, but they primarily edit transcripts or assemble existing footage rather than authoring generative visual narratives.Doki extends text-native authoring beyond speech-track manipulation.
  • Dynamic documents: Doki extends dynamic-document precedents into video by making one document readable as a narrative and executable as a script.This supports gradual enrichment from flexible authoring toward structured video production.

3 How People Create Generative Videos Today

Existing generative-video workflows vary from asset-first to script-first and iterative exploration, yet they share fragmentation, prompt-management burdens, and consistency problems. Doki responds by embedding prompts and structure in a text-centered representation that supports gradual authoring and revision.

  • Current workflows: Creators use asset-first, script-first, and iterative-exploratory workflows, but each requires substantial coordination across generation and editing tools.Asset preparation, repeated prompt specification, or repeated regeneration can all become bottlenecks during production.
  • Recurring challenges: Many creators use three to five tools for scripts, visuals, clips, audio, and timeline assembly, causing context switching and weakening project overview.Fragmentation raises the barrier to entry and slows iteration.
  • Recurring challenges: Prompt engineering can displace narrative development because a 30-shot video may require 30 separate, verbose prompts managed manually.The resulting workflow leaves less bandwidth for higher-level creative choices.
  • Recurring challenges: Maintaining visual and narrative consistency is both a modeling and authoring problem as projects grow.Reference images alone do not provide a structural backbone that propagates assets, styles, and relationships across an extended narrative.
  • Design principles: Doki embeds prompts directly into narrative text so users can develop stories rather than manage disconnected prompt lists.Text is familiar to humans, native to generative models, and flexible for in-place refinement.
  • Design principles: Doki’s example workflows combine slash-command setup and preview generation with sidebar drafting, review, and inline refinement.These workflows support both structured starts and progressively refined drafts.

4 Doki

Doki makes writing the primary medium for generative video authoring, combining narrative structure, visual generation, editing, and AI assistance in one document. Its hierarchical text representation and reusable definitions connect prose with shots while supporting continuity and revision.

  • 4.1 Example Workflows: Doki supports both deliberate manual authoring and AI-directed workflows, including document generation, inline edits, and multi-step revisions.Its sidebar and inline agents apply edits directly to the document and visually highlight those changes.
  • 4.2 A Structured Text Representation for Generative Videos: The representation maps documents to videos, paragraphs to sequences, and sentences to shots, aligning writing practices with video-production structure.A shot marker inserts an inline preview element; following sentences describe that shot until the next shot or paragraph break.
  • 4.3 Doki System: Parameterized definitions and automatic reference resolution preserve consistency across characters, styles, sequences, and shots.Definitions can propagate globally or within heading-scoped regions, reducing repeated references while allowing controlled deviations.
  • 4.3 Doki System: Users create shots and reusable definitions through a lightweight slash menu, with inline previews that support direct review and editing.Shots begin from a slash command and raw text prompt; expanded previews provide trimming, deletion, and regeneration controls.
  • 4.4 Doki’s Shot Generation Pipeline: The shot-generation pipeline transforms user text into structured and rewritten prompts before downstream image and video generation.

5 Diary Study

The authors evaluated Doki through a week-long mixed-methods diary study involving participants with varied video-authoring and generative-AI experience. The study examined perceived benefits and limitations, user workflows, and how prior experience shaped use.

  • Study Design: The mixed-methods diary study addressed perceptions of Doki, workflows, and the influence of prior video-editing and generative-AI experience.The authors chose an extended diary format because short lab sessions would mainly reveal initial impressions and would not capture evolving freeform workflows.
  • Participants: The study recruited 10 participants with varied professional roles, video-making frequencies, and experience using video and generative-AI tools.Participants included filmmakers, animators, designers, engineers, and creators, with video creation ranging from daily to a few times per year.
  • Procedure: Participants used Doki independently for five days after a 60-minute onboarding, aiming to produce 2–3 complete videos while following natural workflows.Daily surveys captured satisfaction, strategies, ease and difficulty, feature use, and optional project submissions.
  • Data Collection: The researchers combined surveys, submitted videos and documents, interaction logs, interviews, and System Usability Scale responses.Logs covered editing, text entry, definitions, previews, generation activity, exports, AI-assistant interactions, and generation costs.

6 Results

In a week-long study, participants used Doki extensively across video creation tasks and generally found it usable, learnable, fast, and supportive of narrative comprehension. Parameterized text representations helped organize stories and maintain consistency, while participants also reported limitations in precise visual control.

  • Usage: Participants spent 91.7 minutes per session editing with Doki and generated 45.5 images and 20.3 videos per session on average.Sessions ranged from 40 to 217 minutes, indicating engagement beyond the required one hour per day.
  • Usage: Export was used in 96% of sessions, while shot generation and the video player were each used in 90%.Defining hashtags appeared in 82% of sessions, mentions in 78%, and audio in 68%.
  • Created videos: Participants submitted 46 videos averaging 67.14 seconds, with most videos lasting 30–90 seconds.The videos covered storytelling, instructional, advertising, music, and experimental categories.
  • Ease of use: Doki received an average SUS score of 81.2, corresponding to the Excellent category and the 90–95th percentile.Nine of ten participants described Doki as very easy to use, and participants reported positive satisfaction throughout the study.
  • Fast to go from idea to content: Eight participants said Doki accelerated the transition from rough ideas to deliverable content, especially for fast, lightweight projects.One participant reported creating about one minute or longer of video in 15 minutes, while another produced five 30-second videos in roughly ten minutes each.
  • Document representation and parametrization: Participants said Doki’s document representation improved their understanding of narrative flow, while parametrized definitions supported reusable elements and cross-story coherence.Participants also found text made AI changes easier to understand and helped reduce repeated keywords and randomness in generated outputs.

7 Discussion and Future Work

Doki centralizes video authoring in a document that humans and AI can both read, edit, and execute, simplifying collaboration while exposing limits in narrative quality and temporal control.

  • Narrative quality: Lowering the barrier to creation does not guarantee compelling storytelling, leaving novices less certain how to structure and improve narrative arcs.The authors propose optional narrative scaffolds and lightweight diagnostics as future support.
  • Future work: Short video-model durations are not presented as the main obstacle because typical film shots average roughly four seconds; consistency, control, and context remain larger gaps.Current text-to-video systems are described as capping out at 5–10 seconds, while average contemporary film shots are roughly 4 seconds.
  • Human–AI common ground: Doki uses a document as an intermediate representation that is readable and editable by humans and interpretable and executable by AI.This shared format keeps AI operations legible, revisable, and open to human intervention.
  • Temporal expressivity: A linear document struggles with concurrent timing, cross-shot transitions, and duration-sensitive rhythm, especially for overlapping audio and cinematic edits.Inline audio notations address basic needs, but finer temporal control remains awkward; proposed extensions include offsets, tempo markers, and transition definitions.
  • Single representation: The single text-native substrate simplifies authoring and collaboration but trades away some task-specific efficiency available in specialized views.Doki supports immediate authoring and stronger version control while making precise pacing, spatial layout, and temporal control harder.

8 Conclusion

The paper presents Doki as a text-native interface that reframes generative video authoring around text as the substrate for narrative, structure, and production. A week-long diary study reports faster workflows, improved coherence through parameterization, and clearer narrative comprehension compared with traditional tools.

  • Conclusion: A week-long diary study found faster idea-to-content workflows, improved coherence via parameterization, and clearer narrative comprehension compared with traditional tools.These are the paper’s reported findings about Doki’s authoring experience.
  • Conclusion: Doki positions text as the primary substrate for narrative, structure, and production rather than merely as input to generative systems.The conclusion frames this as a paradigm for rethinking creative-tool interfaces.

9 Appendix

The appendix documents Doki’s data structures, agent-editing mechanisms, shot-generation logic, model configuration, rewriting modules, participant information, and usage analytics.

  • Data structure: Doki represents shots with identifiers, structured prompts, statuses, definitions, mentions, and hashtags in a JSON-based document structure.Examples include character and location definitions and a shared visual style.
  • Agent memory: Doki’s agents access structured document context containing definitions, headings, paragraphs, shots, and metadata, with the sidebar agent also maintaining conversational memory.This context supports document-aware interactions.
  • Editing API: Agent edits use typed JSON objects specifying an operation identifier, target, replacement content, and user-visible description.The API supports updating, inserting, deleting, and sequencing edits across document nodes.
  • Editing API: Natural-language requests can become structured insertions, such as adding an establishing shot after an introduction with specified content.The example encodes both the insertion location and new shot text.
  • Editing API: Creating a definition from selected text produces coordinated edits that first create the definition and then replace the original text with a reference.The sequence links reusable definitions to subsequent narrative text.
  • Shot generation: Shot-generation logic selects image or video generation based on shot state, definition status, image readiness, and whether generation is already underway.Video generation excludes definition shots and requires an image that is ready, outdated, or saved; image generation requires an idle or ready status.
  • Generation modules: Doki uses configured text, image, and video models while supporting additional model options, and its rewriters expand structured prompts into detailed image or cinematic video descriptions.The image rewriter emphasizes visual descriptiveness and consistency; the video rewriter emphasizes motion, temporal flow, camera, composition, and explicitly mentioned audio.
  • Study materials: The appendix includes participant demographics and usage analytics covering session cost, regenerations, and per-minute activity rates for images, videos, and regenerations.Cost data were available for 35 of 50 sessions, and images had the highest median activity rate.
Loading 2603.09072v1…