Source-linked AI summary

CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models

Hao He, Ceyuan Yang, Shanchuan Lin, Yinghao Xu, Meng Wei, Liangke Gui, Qi Zhao, Gordon Wetzstein, Lu Jiang, Hongsheng Li

arXiv:2503.10592v1cs.CV

TL;DR

Existing camera-conditioned video models struggle to preserve dynamic content and support broad, sequential viewpoint exploration. CameraCtrl II addresses this with dynamic camera-annotated data, lightweight initial-layer camera injection, dynamics-preserving joint training, and clip-wise autoregressive generation. The framework generates camera-controlled dynamic videos with coherent sequential clips and broader scene exploration, while occasional camera–geometry conflicts and distillation-related quality degradation remain limitations.

  • Problem

    Existing camera-controlled video models suffer diminished dynamics, short-clip constraints, and limited viewpoint exploration.

  • Method

    CameraCtrl II combines dynamic camera-annotated data, lightweight initial-layer camera injection, joint training, and clip-wise autoregressive generation.

  • Results

    CameraCtrl II generates camera-controlled dynamic videos while maintaining high quality and temporal consistency across sequential clips.

  • Takeaways & Limitations

    The framework enables users to explore dynamic scenes through precise camera control and sequentially specified trajectories.

  • Takeaways & Limitations

    The method occasionally produces physically implausible camera paths when camera movement conflicts with scene geometry, and distillation can degrade conditional-generation quality.

Abstract

from arXiv · show

This paper introduces CameraCtrl II, a framework that enables large-scale dynamic scene exploration through a camera-controlled video diffusion model. Previous camera-conditioned video generative models suffer from diminished video dynamics and limited range of viewpoints when generating videos with large camera movement. We take an approach that progressively expands the generation of dynamic scenes -- first enhancing dynamic content within individual video clip, then extending this capability to create seamless explorations across broad viewpoint ranges. Specifically, we construct a dataset featuring a large degree of dynamics with camera parameter annotations for training while designing a lightweight camera injection module and training scheme to preserve dynamics of the pretrained models. Building on these improved single-clip techniques, we enable extended scene exploration by allowing users to iteratively specify camera trajectories for generating coherent video sequences. Experiments across diverse scenarios demonstrate that CameraCtrl Ii enables camera-controlled dynamic scene synthesis with substantially wider spatial exploration than previous approaches.

1. Introduction

CameraCtrl II targets dynamic scene exploration by improving camera-controlled video generation and extending it across sequential clips. It combines dynamic camera-annotated data, lightweight camera conditioning, joint training, and autoregressive video extension.

  • Motivation: Existing camera-controlled video models often lose dynamic content and generate only short clips, limiting scene types and explorable spatial range.CameraCtrl generates 25-frame clips, while AC3D generates 49-frame clips; existing methods cannot continue a scene from prior content and new camera trajectories.
  • Method: The authors construct a dynamic dataset by extracting camera trajectory annotations from real videos with Structure-from-Motion.The annotations are extracted using VGGSfM, addressing the reliance on predominantly static annotated video datasets.
  • Method: Camera parameters are injected only at the diffusion model’s initial layer, while joint training on labeled and unlabeled videos preserves dynamic and diverse generation.The training strategy also maintains general video generation without camera input and enables camera classifier-free guidance during inference.
  • Method: A clip-wise autoregressive scheme generates new clips from clean frames of previous clips and newly specified camera trajectories.Training optimizes newly generated frames, enabling sequential generation conditioned on prior scene content.
  • Contributions: The stated contributions are dynamic camera-annotated data curation, lightweight camera control with dynamics-preserving training, and extended clip-wise exploration.Together, these contributions address dynamic content generation and broader scene-range exploration.

2. Related Work

Video diffusion research has advanced general-purpose generation through stronger architectures, datasets, benchmarks, and training methods. CameraCtrl II instead focuses these models on dynamic scene exploration over broad camera-controlled viewpoints.

  • Video Diffusion Models: Video diffusion models have progressed through improved architectures, large-scale datasets, comprehensive benchmarks, and training techniques.Transformer-based models are highlighted for improving temporal consistency and generation quality at scale.
  • Video Diffusion Models: Earlier work transformed text-to-image models into video generators by adding temporal modeling layers, while newer systems increasingly use transformer architectures.These efforts primarily target general-purpose video generation.
  • Camera-Controlled Video Diffusion Models: Camera-controlled video diffusion methods provide camera-pose control, whereas CameraCtrl II applies video diffusion to dynamic scene exploration across a large range.The paper positions scene exploration as distinct from general-purpose video pretraining.

3. CameraCtrl II

CameraCtrl II combines dynamic-video data curation, lightweight camera conditioning, and sequential clip generation to support camera-controlled exploration of dynamic scenes.

  • Camera Representation: The model represents camera geometry with per-pixel Plücker embeddings aligned to encoded visual-token dimensions.The embedding is computed from camera extrinsics and intrinsics for each pixel and frame.
  • Dataset Curation: CameraCtrl II uses a dynamic video dataset with camera trajectory annotations, addressing the static-scene bias of prior camera-annotated datasets.The pipeline detects dynamic foregrounds, estimates optical flow, calibrates scene scales, and balances trajectory types.
  • Camera Control: CameraCtrl II injects camera features at the model’s initial layer and trains with camera-labeled and unlabeled videos to preserve dynamic generation.The camera patchify module produces features matching visual features, while joint training keeps base parameters trainable.
  • Sequential Video Generation: Sequential generation conditions each new clip on clean tokens from a previous clip, concatenates them with noised current tokens, and computes loss only on newly generated tokens.This clip-wise autoregressive scheme uses new camera trajectories during inference for extended scene exploration.
  • Efficiency: Reducing neural function evaluations from 96 to 16 decreases 4-second video sampling time from 13.83 seconds to 2.61 seconds while maintaining visual quality.The reported setting uses 12fps videos and 4 H800 GPUs; camera-control accuracy does not significantly degrade.
  • Efficiency: One-step APT distillation provides further acceleration but degrades conditional-generation quality under the reported setup.The authors attribute possible improvement to using more computational resources and larger batch sizes.

4. Experiments

Experiments compare CameraCtrl II with prior methods across quantitative, qualitative, ablation, and visualization studies. Results show improved camera control, dynamics, consistency, sequential exploration design, and applicability across diverse scenes.

  • Comparisons with other methods: The evaluation compares CameraCtrl II with MotionCtrl and CameraCtrl for I2V generation and with AC3D for camera-controlled T2V generation.The study evaluates both single-clip and extended-scene capabilities using quantitative metrics, qualitative comparisons, ablations, and visualizations.
  • Comparisons with other methods: CameraCtrl II follows camera trajectories more completely while preserving stronger video dynamics than CameraCtrl and AC3D in qualitative comparisons.CameraCtrl misses upward movement, whereas AC3D misses forward motion at the trajectory’s end.
  • Comparisons with other methods: CameraCtrl II significantly outperforms previous methods across all evaluated metrics in both I2V and T2V settings.In I2V, it improves FVD, Motion strength, TransErr, RotErr, and Geometric consistency relative to MotionCtrl and CameraCtrl.
  • Ablation Study: Removing scale calibration increases TransErr from 0.1830 to 0.2121 and RotErr from 1.74 to 2.14, while reducing geometric consistency.The ablation supports normalizing scene scales to improve geometric relationships, camera control, and scene reconstruction.
  • Ablation Study: Balancing camera-trajectory distributions is essential for robust camera control and geometric consistency across diverse movement patterns.Removing distribution balancing causes notable degradation in both measures.
  • Ablation Study: A simple patchify camera encoder performs better across most metrics than a more complex encoder, while multilayer injection significantly reduces Motion strength.The results support injecting camera features only at the initial DiT layer to guide overall generation without restricting local dynamic content.
  • Ablation Study: Joint training with unlabeled videos preserves diverse dynamics and improves camera control, while global reference frames improve cross-clip camera and appearance consistency.Noised conditioning creates a training–inference mismatch that degrades FVD and Appearance consistency.
  • Visualization Results: Across varied scenarios, CameraCtrl II controls diverse camera movements, preserves appropriate dynamics, and produces videos suitable for high-quality 3D point-cloud reconstruction.The visualizations include game-like, historical, indoor, fantasy, and anime-style scenes; FLARE estimates detailed point clouds from generated frames.

5. Discussion

CameraCtrl II combines dynamic-video camera control with sequential clip extension for broader scene exploration. The framework preserves dynamic generation while maintaining quality and temporal consistency across clips, but can produce physically implausible paths when camera movement conflicts with scene geometry.

  • REALCAM provides dynamic videos with camera pose annotations for training CameraCtrl II.
  • A lightweight injector integrates camera conditions at DiT’s initial layers, while joint training preserves pretrained dynamic-scene generation.
  • Clip-level extension generates new video clips conditioned on previously generated content and new camera trajectories.
  • Experiments show camera-controlled dynamic videos retain high quality and temporal consistency across sequential video clips.
  • The model sometimes generates physically implausible camera paths that intersect scene structures when camera movement conflicts with scene geometry.
Loading 2503.10592v1…