Source-linked AI summary

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang, Qingyan Bai, Ka Leong Cheng, Yue Yu, Yixuan Li, Yihao Meng, Zichen Liu, Yanhong Zeng, Yujun Shen, Qifeng Chen

arXiv:2607.02517v1cs.CV

TL;DR

Dynamic object memory remains underexplored in video world models, which struggle to preserve moving entities when they leave view. WorldDirector decouples semantic motion orchestration from video synthesis to enable controllable trajectories and persistent appearance consistency, and evaluations show dynamic consistency across extended sequences.

  • Problem

    Video world models still lack persistent dynamic object memory, leaving dynamic entities’ physical movements and existence underexplored when they move out of camera view.

  • Method

    WorldDirector uses an LLM to orchestrate semantic 3D object trajectories and camera movements, then conditions causal video generation on these plans and appearance information.

  • Results

    Extensive evaluations show WorldDirector synthesizes controllable dynamic scenarios while preserving object permanence and appearance consistency across prolonged out-of-view intervals.

  • Takeaways & Limitations

    WorldDirector establishes a highly controllable paradigm for interactive video world models with persistent dynamic memory and open-ended event design.

  • Takeaways & Limitations

    Synthetic game data creates a domain gap that can restrict visual fidelity, producing unnatural locomotion or blurry faces.

Abstract

from arXiv · show

We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. Unlike existing world models that entangle physical dynamics with pixel rendering and rely on continuous visual observation to sustain motion, our framework explicitly decouples semantic motion orchestration from visual generation. By leveraging an LLM to coordinate 3D trajectories with camera movements and subsequently employing these orchestrated trajectories as control signals for video generation, our approach ensures strict physical logic and appearance stability, successfully preserving the exact visual identities of dynamic entities even when they re-enter the scene after prolonged periods out of view. Experimental results demonstrate that our method supports the synthesis of complex and extended events with unprecedented controllability and persistent dynamic object memory. Project Page: https://worlddirector.github.io/

1 Introduction

WorldDirector frames persistent world simulation around independent, camera-unconstrained object motion and strict appearance consistency for entities returning from out of view. It decouples semantic motion planning from video synthesis to enable controllable dynamic scenarios with persistent memory across extended sequences.

  • Interactive video world simulation requires memory that preserves static scenes and continuous dynamic-object movement both in view and out of view.
  • Robust dynamic memory requires entities to follow independent trajectories governed by continuous physical logic regardless of camera visibility.
  • WorldDirector decouples dynamic-object motion planning from video synthesis and conditions generation on semantic-level planning results.
  • WorldDirector synthesizes controllable dynamic scenarios while maintaining object permanence and appearance consistency across prolonged out-of-view intervals.

2 Related Works

Prior work spans generative video synthesis, interactive world models, controllable motion guidance, and temporal memory, while highlighting failures in sustaining unseen objects and dynamics. WorldDirector is positioned against these limitations through explicit decoupling and persistent object memory.

  • Generative video and world models: Diffusion- and transformer-based video synthesis has expanded toward interactive video world models for simulating environments.Examples named include Sora, Genie, Oasis, and DIAMOND, alongside game-like simulators and long-sequence interactive generators.
  • Generative video and world models: Implicitly memorizing object states, actions, and appearances overloads generative models, causing off-screen entities to freeze or vanish.The cited limitation arises when active entities exit the camera’s field of view.
  • Controllable synthesis: Controllable synthesis has progressed from image control to video, bounding-box and camera control, motion tracking, and dense or point-trajectory guidance.Dense and point trajectories are described as flexible interfaces for fine-grained, entity-level motion control.
  • Memory and object permanence: Long-video generators use sliding windows but struggle with extended occlusion, while prior consistency methods assume static scenes.Static-scene consistency has been pursued through FOV retrieval and 3D representations.
  • Memory and object permanence: Object permanence requires objects to persist and evolve while unobserved, remaining a core challenge in physical reasoning and complex active scenes.The passage frames this challenge as extending beyond immediate temporal context.

3 Method

WorldDirector builds a controllable world simulator by combining curated dynamic-object data with location, appearance, textual, and contextual-memory conditions. Its architecture and training procedure preserve object identity and scene consistency during causal video generation and free-viewpoint exploration.

  • Model design: The model extends LingBot-World-Base with spatial and appearance feature channels, encoding both through a 3D VAE before concatenating them with noisy latent sequences.Instance-specific color masks provide geometric trajectory priors, while appearance features anchor dynamic-object identity.
  • Data curation: The data curation pipeline creates training tuples containing dynamic-object bounding boxes, appearance references, behavior-focused captions, and contextual signals.These components target spatial grounding, visual conditioning, behavioral modeling, and causal generation.
  • Data curation: A game-based platform generates 15-second videos with scripted object disappearances and reappearances, while SAM3 extracts re-identifiable 2D bounding-box trajectories.The tracking procedure is designed to follow objects through temporary field-of-view exits.
  • Data curation: Each training sequence receives location and appearance conditioning videos to encode object trajectories, positional priors, and appearance consistency across re-entry events.Location conditions use identity-specific color-coded boxes, while appearance conditions provide visual references for dynamic objects.
  • Model design: WorldDirector conditions denoising on location masks, sparse appearance features, multi-granularity behavioral prompts, and retrieved contextual memory frames.The contextual frames are selected through dual-stream retrieval and paired with aligned location and appearance conditions.
  • Training and inference: Training applies flow matching with MSE only to the target segment, leaving historical context clean so inference can generate new content anchored by prior frames.Inference proceeds through world planning via an LLM after this memory-preserving training setup.

4 Experiments

Experiments evaluate WorldDirector on unseen scenes using reconstruction, consistency, and dynamic-subject metrics, alongside quantitative comparisons, ablations, and promptable-event demonstrations. Results attribute reconstruction performance to location conditioning, show the importance of explicit appearance conditioning, and extend the framework to novel objects introduced through text-defined trajectories.

  • Evaluation Protocol: The evaluation uses 100 test videos with novel scenes and subjects unseen during training, measuring PSNR, SSIM, LPIPS, VBench consistency, and Dynamic Subject Consistency.The protocol combines pixel-wise reconstruction fidelity, frame-level subject and background coherence, and dynamic-subject consistency based on detected bounding boxes.
  • Quantitative Results: WorldDirector achieves state-of-the-art performance across all three reconstruction metrics, attributed to location conditioning that captures object positions and camera poses.The reported explanation links continuous position and camera-pose conditioning to generation that more accurately aligns with ground truth.
  • Quantitative Results: WorldDirector maintains the expected scene and the consistency of a man reappearing after a long period of disappearance.The passage contrasts this behavior with baseline VBench results, which are described as favoring Yume, HY-World, and Infinite-World despite generated-video analysis.
  • Ablation Studies: Without explicit Appearance Condition, the model fails to exploit contextual color-coded masks, causing severe identity loss for re-entering dynamic objects.The ablation examines whether visual consistency can be maintained implicitly from the Location Condition and reports failure in a complex movement case.
  • Promptable World Events: The LLM can introduce novel objects by defining identities, entrance timings, and 3D trajectories, with synthesized appearances added to the Appearance Condition pool for later consistency.This promptable-world-events setup is not restricted to entities present in the initial frame.

5 Conclusion

WorldDirector enables free exploration and flexible event design in video world models by decoupling semantic orchestration from latent synthesis while preserving rigorous dynamic memory. Its reliance on synthetic game data creates a domain gap that can reduce visual fidelity, motivating future use of real-world datasets.

  • Conclusion: WorldDirector supports free exploration and flexible event design while preserving rigorous dynamic memory.The framework decouples semantic orchestration from latent synthesis.
  • Conclusion: LLMs plan complex 3D trajectories and open-world events, which are visually realized through causal chunk-based context routing with spatial and appearance conditioning.
  • Conclusion: Synthetic game data introduces a domain gap that can restrict visual fidelity through unnatural locomotion or blurry faces.Future work will incorporate real-world datasets to bridge this gap and enhance visual realism.

A Training and Compute Details.

WorldDirector is trained on a distributed 64-GPU cluster with memory- and throughput-oriented system optimizations. The training configuration uses 832 × 480 videos at 16 fps, a 10-frame context, AdamW with a 1 × 10−5 learning rate, BF16 mixed precision, and a global batch size of 64.

  • Training and Compute Details: Training uses 8 compute nodes, each equipped with 8 NVIDIA A100 (80GB) GPUs, for a total of 64 GPUs.Fully Sharded Data Parallel (FSDP) and activation checkpointing improve memory efficiency and training throughput.
  • Training and Compute Details: The model processes training videos at 832 × 480 pixels and 16 fps with context length N = 10 frames.These settings define the video resolution, frame rate, and temporal context used during training.
  • Training and Compute Details: Optimization uses AdamW with a constant learning rate of 1 × 10−5, BF16 mixed precision, and a global batch size of 64.BF16 mixed precision is used to accelerate computation while maintaining numerical stability during denoising.

B Details of Static and Dynamic Context Retrieval. · C Details of Inference System

The retrieval mechanism selects N memory frames through parallel static-viewpoint and dynamic-entity scoring, then interleaves and temporally orders them into the final context. The inference system comprises World Planning via LLM and Causal Chunk-Based Generation.

  • B Details of Static and Dynamic Context Retrieval.: Algorithm 1 constructs the final context M from candidate frames F using training frames V, camera poses C, and dynamic-entity 2D bounding boxes B.Its output is the static and dynamic context, with context length N controlling the number of memory frames selected.
  • B Details of Static and Dynamic Context Retrieval.: Static retrieval ranks candidate frames by their maximum Field-of-View overlap with training-frame camera poses.Each candidate’s score is maxv∈V FoV_Overlap(Cc, Cv), and candidates are sorted in descending score order.
  • B Details of Static and Dynamic Context Retrieval.: Dynamic retrieval greedily selects frames for the least-covered dynamic entity, maximizing its visible 2D bounding-box area.Coverage counters start at zero and selection continues until the requested context length N is reached.
  • B Details of Static and Dynamic Context Retrieval.: Selected-frame coverage updates add normalized bounding-box areas for every entity present in the chosen frame.The normalizer is the maximum bounding-box area Amax in that frame, producing visibility-weighted coverage updates.
  • B Details of Static and Dynamic Context Retrieval.: The final memory set alternately draws from camera-based and box-based retrieval lists while discarding duplicates.Frames are appended from Pcam and Pbox until |M| = N, then returned in temporal order.
  • B Details of Static and Dynamic Context Retrieval.: Interleaving enforces a minimum temporal stride of four frames to guarantee uniform context-frame distribution.The context frames are returned in chronological order after interleaving terminates.
  • C Details of Inference System: The inference system has two main components: World Planning via LLM and Causal Chunk-Based Generation.The section introduces these as the system’s two principal inference operations.

C.1 World Planning via LLM · C.2 Causal Chunk-Based Generation

WorldDirector uses an LLM to convert scene structure and narrative prompts into 3D object and camera trajectories, including motion for novel objects, then projects them into spatial controls. It generates videos causally in chunks, using appearance, location, camera, and historical context conditions to maintain continuous long-horizon streams.

  • C.1 World Planning via LLM: WorldDirector uses Gemini to plan 3D trajectories from selected objects’ estimated 3D boxes, orientations, camera parameters, and customized narrative prompts.SAM and DepthAnything v2 provide rough 3D bounding boxes and initial orientations for entities and the camera.
  • C.1 World Planning via LLM: The LLM generates physically plausible kinematics for completely novel entities specified in prompts, enabling Promptable World Events beyond initially selected objects.This extends trajectory planning to objects absent from the initial frame.
  • C.1 World Planning via LLM: The planner produces complete 3D object trajectories and camera poses, then projects them into 2D bounding-box sequences used as the deterministic Spatial Location Condition (B).The 2D projections are applied in the subsequent generative stage.
  • C.2 Causal Chunk-Based Generation: The synthesis stage uses the projected location conditions and camera trajectories in a causal autoregressive process that partitions sequences into N chunks of size K.A global video buffer V is initialized with the starting reference frame I0.
  • C.2 Causal Chunk-Based Generation: For the first chunk, appearance is extracted from I0 and historical memory is empty; later chunks construct appearance and retrieve context from the previously generated buffer V.The conditioning strategy changes according to the temporal state of generation.
  • C.2 Causal Chunk-Based Generation: Each chunk is synthesized from location, caption, appearance, memory, camera, and initial-frame conditions supplied to WorldDirector.The last frame of the current buffer serves as the next chunk’s conditional initial frame.
  • C.2 Causal Chunk-Based Generation: Generated chunks are appended without their overlapping first frame, progressively unrolling a continuous long-horizon video stream without inherent length limitations.Excluding the first frame prevents redundancy at chunk boundaries.

D Ablation on Dynamic Context and Appearance Condition Drop Mechanism

The ablation study evaluates retrieved dynamic context and the Temporal Drop Mechanism, showing that dynamic-context retrieval is necessary for preserving the identities of re-entering entities. Using only Appearance Condition produces semantically similar but non-identical entities.

  • The study evaluates both retrieved dynamic context and the Temporal Drop Mechanism for re-entering dynamic entities.
  • Ablating the dynamic-context stream confirms that retrieving dynamic objects within contextual memory is necessary for identity preservation.
  • Using only Appearance Condition degrades identity preservation, generating semantically similar but non-identical entities.

E Flexible Viewpoint Control

WorldDirector enables flexible viewpoint control by incorporating spatial location conditions into 3D trajectory planning. It supports both third-person tracking and independent first-person exploration, including dynamic switches between these perspectives.

  • Viewpoint Control: Spatial location conditions enable flexible viewpoint control and seamless transitions between first- and third-person exploration.The framework incorporates these conditions during 3D trajectory planning.
  • Third-Person Exploration: Anchoring a target entity’s 2D bounding box near the camera-field center yields a third-person perspective.This configuration supports tracking shots such as a running dog with a 360◦ panoramic sweep.
  • First-Person Exploration: Decoupling the camera trajectory from dynamic objects enables independent first-person navigation and backward movement.The viewpoint can switch from third-person tracking to independent first-person backward movement.

F More Qualitative Comparisons · G Impact Statement

Additional qualitative comparisons support the main findings: several baselines generate less subject motion, while LingBot-World produces dynamic results but lacks fine-grained interactive control. The work targets applications in virtual reality, gaming, filmmaking, and interactive design, while acknowledging unaddressed risks such as deceptive video generation.

  • F More Qualitative Comparisons: Yume, HY-World, and Infinite-World tend to generate significantly less subject motion than the proposed results.
  • F More Qualitative Comparisons: LingBot-World produces highly dynamic results that align well with textual prompts.
  • F More Qualitative Comparisons: LingBot-World lacks fine-grained interactive control precision, making strict user-intent matching difficult.
  • F More Qualitative Comparisons: Figure S3 provides qualitative comparisons with baselines, including a HyDRA setup using the initial 10s of the results as reference video.
  • G Impact Statement: The paper focuses on controllable video world simulation with persistent dynamic memory.
  • G Impact Statement: The work aims to enhance virtual reality, gaming, filmmaking, and interactive design, while not directly addressing potential negative societal impacts.Examples include malicious or unintended uses such as generating deceptive or fake video content.

H Responsible Release and Safeguards

WorldDirector proposes a staged, documented release limited to academic research and evaluation. The paper recommends established safeguards for downstream use because its long-horizon, persistent-entity generation lowers the barrier to creating complex, logical events and raises misuse concerns.

  • Release Plan: The authors plan a staged, documented release of inference code, LLM prompt templates, and pretrained checkpoints strictly for academic research and evaluation.The release materials are intended for research and evaluation rather than unrestricted deployment.
  • Downstream Safeguards: Downstream applications should combine WorldDirector with prompt safety filters, video watermarking, content provenance metadata, and deployment-time monitoring.These established safety mechanisms are outside the scope of the foundational paper but are strongly recommended for application use.
  • Risk Considerations: Persistent entities and continuous long-horizon event generation lower the barrier to creating complex, logical content, requiring mitigation of video-generation misuse risks.The paper identifies this capability as an important consideration for responsible deployment.
Loading 2607.02517v1…