Source-linked AI summary

MotionCtrl: A Unified and Flexible Motion Controller for Video Generation

Zhouxia Wang, Ziyang Yuan, Xintao Wang, Tianshui Chen, Menghan Xia, Ping Luo, Ying Shan

arXiv:2312.03641v2cs.CVcs.AIcs.LGcs.MM

TL;DR

Existing video-generation methods often focus on one motion type or fail to distinguish camera and object motion, limiting fine-grained control. MotionCtrl addresses this gap with separate camera and object motion modules and reports superior qualitative and quantitative control across both motion types, while simultaneous complex control remains difficult.

  • Problem

    Existing methods often focus on one motion type or do not clearly distinguish camera motion from object motion, limiting fine-grained and diverse control.

  • Method

    MotionCtrl integrates a temporal camera-pose module and a spatial object-trajectory module into a pretrained video-generation model, using tailored datasets and multi-step training.

  • Results

    MotionCtrl demonstrates superior camera and object motion control over previous methods in qualitative and quantitative evaluations.

  • Takeaways & Limitations

    Camera poses and object trajectories enable independent or combined motion control while minimally affecting generated objects’ visual appearance.

  • Takeaways & Limitations

    Simultaneously controlling complex camera motion and complex object trajectories has a relatively low success rate and requires careful trajectory design.

Abstract

from arXiv · show

Motions in a video primarily consist of camera motion, induced by camera movement, and object motion, resulting from object movement. Accurate control of both camera and object motion is essential for video generation. However, existing works either mainly focus on one type of motion or do not clearly distinguish between the two, limiting their control capabilities and diversity. Therefore, this paper presents MotionCtrl, a unified and flexible motion controller for video generation designed to effectively and independently control camera and object motion. The architecture and training strategy of MotionCtrl are carefully devised, taking into account the inherent properties of camera motion, object motion, and imperfect training data. Compared to previous methods, MotionCtrl offers three main advantages: 1) It effectively and independently controls camera motion and object motion, enabling more fine-grained motion control and facilitating flexible and diverse combinations of both types of motion. 2) Its motion conditions are determined by camera poses and trajectories, which are appearance-free and minimally impact the appearance or shape of objects in generated videos. 3) It is a relatively generalizable model that can adapt to a wide array of camera poses and trajectories once trained. Extensive qualitative and quantitative experiments have been conducted to demonstrate the superiority of MotionCtrl over existing methods. Project Page: https://wzhouxiff.github.io/projects/MotionCtrl/

1 INTRODUCTION

MotionCtrl addresses the limited distinction between camera-induced and object motion by introducing a unified controller with separate modules, datasets, and training strategies. It aims to provide fine-grained, flexible control while preserving video appearance and supporting diverse motion combinations.

  • Video generation requires temporally consistent motion, including global camera motion and local object motion.
  • Previous methods often focus on one motion type or use shared conditions, limiting fine-grained and diverse motion control.
  • MotionCtrl uses separate Camera Motion Control and Object Motion Control Modules to independently or jointly control camera and object motion.
  • The adapter-like modules can be trained separately, reducing dependence on a dataset containing captions, camera poses, and object trajectories together.
  • MotionCtrl uses camera poses and trajectories as motion conditions, supporting varied motion combinations while minimally affecting object appearance.

2 RELATED WORKS

Prior video-generation methods use diffusion models and template or generalized motion guidance, but often require motion-specific training or insufficiently disentangle camera and object motion. MotionCtrl instead combines camera poses and object trajectories for finer, more flexible control.

  • Recent video-generation research increasingly uses latent diffusion models for computationally efficient, high-fidelity video synthesis.
  • Template-based motion-control methods can be effective for specific motions but typically require a new model for different templates.
  • Generalized methods provide broader motion control but may fail to finely disentangle camera and object motion.
  • MotionCtrl uses camera poses, object trajectories, or their combination to control generated-video motion with finer and more flexible guidance.

3 METHODOLOGY

MotionCtrl extends a latent video diffusion model with separate temporal camera and spatial object controls, trained using complementary datasets and sparse trajectory supervision. Its design matches camera motion’s global character and object motion’s local character while addressing incomplete annotations.

  • 3.1 Preliminary: LVDM generates text-guided videos by denoising noisy latent features with a U-Net and noise-prediction objective.
  • 3.2 MotionCtrl: MotionCtrl adds CMCM and OMCM to LVDM, aligning temporal camera control with temporal transformers and spatial object control with convolutional layers.
  • 3.2 MotionCtrl: CMCM receives camera poses represented by 3×3 rotation and 3×1 translation matrices, concatenating them with temporal features before the second self-attention module.
  • 3.2 MotionCtrl: OMCM represents object trajectories as framewise spatial positions and injects their multi-scale convolutional features into LVDM’s convolutional layers.
  • 3.2.3 Training Strategy and Data Construction: MotionCtrl trains its modules with separate annotation requirements, using Realestate10K for captions and camera poses and WebVid trajectories synthesized with ParticleSfM.
  • 3.2.3 Training Strategy and Data Construction: The OMCM is trained with sparse trajectories sampled from dense trajectories and refined with Gaussian filtering to make user-provided control more practical.

4 EXPERIMENTS

Experiments show that MotionCtrl provides strong camera and object motion control, supports combined control, and preserves video quality and text similarity. Ablations also validate its module placement and multi-step training strategy.

  • Camera Motion Control: MotionCtrl produces adjustable-speed camera motion for basic poses and more natural complex-pose videos than VideoComposer.VideoComposer’s dense motion vectors can capture unintended shapes, whereas MotionCtrl uses camera poses to produce motion close to the reference.
  • MotionCtrl outperforms competing approaches in both camera and object motion control while preserving text similarity and video quality.
  • Object Motion Control: MotionCtrl follows specified object trajectories more precisely than VideoComposer across frames.The expected object locations are marked by green points against the given red trajectory.
  • Integrated Motion Control: MotionCtrl independently controls camera and object motion and can combine both controls within one video.Adding zoom-out camera poses to a trajectory-controlled rose animates both the rose and background.
  • Camera Motion Module Ablation: CMCM integrated with LVDM temporal transformers achieves a CamMC score of 0.0289, while spatial and time-embedding integrations remain close to the original LVDM.The temporal placement matches camera motion’s global transformations over time.
  • Training Strategy: Training OMCM first with dense trajectories and then fine-tuning with sparse trajectories improves object motion control over dense-only or sparse-only training.The sequence CMCM before OMCM also avoids later CMCM training disrupting object-motion adaptation.

5 LIMITATIONS

MotionCtrl remains limited when complex camera and object trajectories must be controlled simultaneously in the same video. Such cases require careful trajectory design and currently have a relatively low success rate.

  • Simultaneously controlling complex camera and object trajectories requires careful trajectory design and has a relatively low success rate.

6 CONCLUSION

The paper introduces MotionCtrl as a unified controller for independent or combined camera and object motion. Its tailored modules, multi-step training, and augmented datasets are supported by qualitative and quantitative experiments showing superiority in both control tasks.

  • MotionCtrl uses tailored camera and object motion modules with multi-step training and augmented datasets to control motion independently or jointly.
  • Comprehensive qualitative and quantitative evaluations show MotionCtrl’s superiority in both camera and object motion control.

A DETAILS OF TRAINING DATA CONSTRUCTION

MotionCtrl constructs separate augmented datasets for camera and object motion control by adding captions to RealEstate10K and trajectories to WebVid.

  • Augmented-RealEstate10K adds Blip2-generated captions to videos already annotated with camera poses for CMCM training.Captions are formed from frames sampled at the first, quarter, half, three-quarter, and final positions.
  • Augmented-WebVid adds object movement trajectories synthesized with ParticleSfM to captioned WebVid videos for OMCM training.

B DETAILS OF EVALUATION DATASETS

The evaluation suite independently measures camera and object motion control across diverse poses, trajectories, prompts, and dataset sources.

  • Camera Motion Control Evaluation Dataset: The camera evaluation dataset contains 407 samples spanning basic and relatively complex camera-pose sequences.Complex poses include camera turning or self-rotation within a single sequence.
  • Camera Motion Control Evaluation Dataset: Basic camera evaluation covers pan, zoom, and rotation sequences paired with multiple prompts.
  • Camera Motion Control Evaluation Dataset: Additional camera samples use complex poses from RealEstate10K, WebVid, and HD-VILA with prompts from the evaluation protocol.
  • Object Motion Control Evaluation Dataset: The object evaluation dataset contains 283 samples built from 74 diverse trajectories and 77 prompts, pairing trajectories and prompts in varied combinations.
  • The constructed datasets primarily assess camera and object control quantitatively, while MotionCtrl supports poses and trajectories beyond those included.

C.1 More Quantitative Results

Additional quantitative evaluations compare MotionCtrl with VideoComposer on complex camera poses and user preferences across motion-control assessment aspects.

  • Complex Camera Motion Control: MotionCtrl outperforms VideoComposer on complex camera poses from RealEstate10K, WebVid, and HD-VILA.Table 4 reports statistics separately for each source dataset.
  • The evaluation includes camera-pose and object-trajectory datasets designed to quantify control effectiveness.
  • User Study: Over 90 percent of participants preferred MotionCtrl results across all assessment aspects.
  • User Study: VideoComposer’s motion-vector conditioning could produce unnatural object shapes, contributing to stronger user preference for MotionCtrl’s relatively natural results.

C.2 More Qualitative Results

Qualitative results show MotionCtrl following complex camera poses and object trajectories while preserving prompt alignment, supporting combined and flexible motion control.

  • Comparisons with VideoComposer: MotionCtrl follows camera poses and object trajectories more effectively than VideoComposer while producing higher-quality videos.
  • MotionCtrl does not require extra fine-tuning for different camera poses or trajectories.
  • Basic Camera Motion: A single MotionCtrl model integrates eight basic camera motions, unlike AnimateDiff’s separate LoRA model for each camera motion.
  • Complex Camera Motion: For complex camera-pose sequences involving turning or self-rotation, generated motion follows the poses and content aligns with text prompts.
  • Object Motion: With identical object trajectories and different prompts, MotionCtrl generates different objects exhibiting identical object motion.
  • Combined Motion Control: Combining camera and object controls with the same trajectory but different camera poses changes the horse’s performance.

D MORE RESULTS OF MOTIONCTRL DEPLOYED ON AIMATEDIFF [Guo et al. 2023]

MotionCtrl is deployed on AnimateDiff with different LoRA models to control camera and object motion, including basic and relatively complex motion settings.

  • MotionCtrl is combined with AnimateDiff and various LoRA models to control motion in generated videos.The paper reports deployments using a fine-tuned AnimateDiff together with different LoRA models.
  • The deployment includes results for basic camera motion control and relatively complex camera and object motion control.Basic camera-control results are shown in Figures 19 and 20, while more complex results are reported in the manuscript.

E MORE DISCUSSIONS ABOUT THE RELATED WORKS

Compared with related methods, MotionCtrl uses one model to control camera and object motion independently through camera poses and trajectories. The supplied results also illustrate its flexibility across motion types, models, and simultaneous controls.

  • Related-work comparison: MotionCtrl distinguishes camera motion from object motion, unlike related unified approaches guided by motion vectors or trajectories.MotionDirector and VideoComposer use unified models but do not separate the two motion types.
  • Related-work comparison: Unlike several template-video methods, MotionCtrl uses a unified model rather than training distinct models for different template videos.AnimateDiff, Tune-a-video, LAMP, and MotionDirector extract motion from template videos and require different models for different templates.
  • Motion conditions: Camera poses and trajectories guide camera and object motion, respectively, providing more fine-grained control over video generation.The two condition types support independent and flexible control of a wide range of camera and object motions.
  • Camera-motion results: The same MotionCtrl model controls eight basic camera poses on LVDM without extra fine-tuning for each pose.The poses are pan up, pan down, pan left, pan right, zoom in, zoom out, anticlockwise rotation, and clockwise rotation.
  • Object-motion results: Given the same trajectories, MotionCtrl can generate different objects while maintaining their motion and can control multiple object motions simultaneously.The figure describes trajectory-guided results on LVDM with green starting points and multiple trajectories in one video.
  • AnimateDiff deployment: On AnimateDiff, MotionCtrl controls camera motion and motion speed using basic camera-pose guidance.The reported AnimateDiff results are presented in Figures 19 and 20.
Loading 2312.03641v2…