Source-linked AI summary

Video-P2P: Video Editing with Cross-attention Control

Shaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin, Jiaya Jia

arXiv:2303.04761v1cs.CV

TL;DR

Video editing needs attention-controlled methods that preserve video structure, but large-scale pre-trained video-generation models are not publicly available. Video-P2P adapts an image diffusion model through Text-to-Set tuning, shared embedding optimization, and decoupled guidance, enabling detailed real-world edits while preserving original poses and scenes. The framework supports multiple text-driven editing applications and is reported to outperform previous approaches.

  • Problem

    Real-world video editing requires attention control, but no large-scale pre-trained video-generation models are publicly available.

  • Method

    Video-P2P tunes a Text-to-Set model, optimizes a shared unconditional embedding for inversion, and applies decoupled guidance with source and target attention maps.

  • Results

    Video-P2P supports word swap, prompt refinement, and attention re-weighting while preserving original poses and scenes, and significantly outperforms previous approaches.

  • Takeaways & Limitations

    Adapting an image diffusion model enables local and global real-world video editing with cross-attention control.

  • Takeaways & Limitations

    Video-P2P cannot perform video motion editing because it lacks temporal priors.

Abstract

from arXiv · show

This paper presents Video-P2P, a novel framework for real-world video editing with cross-attention control. While attention control has proven effective for image editing with pre-trained image generation models, there are currently no large-scale video generation models publicly available. Video-P2P addresses this limitation by adapting an image generation diffusion model to complete various video editing tasks. Specifically, we propose to first tune a Text-to-Set (T2S) model to complete an approximate inversion and then optimize a shared unconditional embedding to achieve accurate video inversion with a small memory cost. For attention control, we introduce a novel decoupled-guidance strategy, which uses different guidance strategies for the source and target prompts. The optimized unconditional embedding for the source prompt improves reconstruction ability, while an initialized unconditional embedding for the target prompt enhances editability. Incorporating the attention maps of these two branches enables detailed editing. These technical designs enable various text-driven editing applications, including word swap, prompt refinement, and attention re-weighting. Video-P2P works well on real-world videos for generating new characters while optimally preserving their original poses and scenes. It significantly outperforms previous approaches.

1. Introduction

Video-P2P adapts an image diffusion model into a video-editing pipeline that combines structured inversion with attention control. This design targets semantically consistent local and global edits while improving reconstruction through shared unconditional embedding optimization.

  • Video-P2P addresses the challenge of editing local video objects without influencing the environment, while supporting local and global edits.
  • Adapting a pre-trained image diffusion model into a Text-to-Set model preserves semantic consistency across video frames.Frame-attention replaces self-attention, enabling the same edited object type across frames.
  • Tuning the Text-to-Set model provides an approximate video inversion after model inflation degrades generation quality.The tuned model is sufficient for approximate inversion, although inversion errors accumulate.
  • Optimizing one shared unconditional embedding across frames improves inversion quality by aligning denoising and diffusion latent features.Experiments identify the shared embedding as an efficient and effective choice for video inversion.
  • Compared with frame-wise Image-P2P editing, Video-P2P maintains the same robotic penguin type across frames.
  • Video-P2P introduces cross-attention control with decoupled guidance to improve video-editing performance.The framework is presented as the first video-editing framework with attention control and is evaluated through ablations and comparisons.

2. Related Work

Related work spans text-to-video generation, single-video adaptation, image editing, and generative video editing. Existing methods provide useful capabilities but remain limited by artifacts, availability, temporal consistency, or the magnitude of semantic changes they support.

  • 2.1. Text Driven Generation: Text-to-video methods generate reasonable short videos but still contain artifacts and generally do not support real-world video editing.Most described approaches are also not publicly available.
  • 2.2. Single Video Generation: Single-video models adapt generative models to individual videos, but Tune-A-Video has limited temporal consistency despite allowing semantic changes.
  • 2.3. Image Editing: Image-editing research includes diffusion and attention-control methods for preserving unrelated image regions during text-driven edits.Prompt-to-Prompt and Plug-and-Play use attention control, while Null-Text Inversion improves real-image editing.
  • 2.4. Video Editing: Generative video-editing methods modify video content through texture editing, image-to-video or video-to-video editing, and joint image-video training.Text2Live struggles with significant semantic changes, while Dreamix can also change motion.

3. Method

Video-P2P adapts an image diffusion model into a video-editing pipeline by combining video inversion with cross-attention control. Its design uses a shared unconditional embedding and separate source/edited branches to balance reconstruction and editability.

  • Decoupled-guidance Attention Control: Cross-attention control uses separate reconstruct and editable branches, incorporating their attention maps to create the target video.The reconstruct branch uses an optimized shared embedding, while the editable branch uses an initialized embedding.
  • Video Inversion: A tuned Text-to-Set model uses frame-attention to generate semantically consistent image sets and provide an approximate video inversion.Fine-tuning recovers generation quality degraded by model inflation and enables successful approximate inversion.
  • Video Inversion: The method optimizes one unconditional embedding shared across all frames to align denoising and diffusion latent features while reducing memory use.Sharing the embedding also avoids destabilizing semantic consistency during attention control.
  • Editing Applications: The framework applies attention control methods from Image-P2P to video editing, including prompt refinement, attention re-weighting, and word swapping.Algorithm 1 takes source and target prompts and outputs source and edited videos from DDIM-inverted latent features.

14 Return (z0, z∗

The method converts an attention map into a binary mask for subsequent editing operations. Pixels above a threshold receive value 1.

  • 14 Return (z0, z∗: The attention map is converted into a binary mask for the editing process.A value is assigned based on whether the attention exceeds a threshold.
  • 14 Return (z0, z∗: The binary mask assigns value 1 where the attention-map value is larger than the threshold.Values below the threshold do not receive value 1.
  • 14 Return (z0, z∗: The thresholded attention map provides a binary selection of regions used by the editing pipeline.The passage specifies the thresholding rule but not the subsequent operation on selected regions.

4. Experiments

Video-P2P supports diverse text-driven video edits while preserving temporal consistency and limiting changes to unrelated regions. Ablations show that model initialization, shared unconditional embeddings, and decoupled guidance improve reconstruction or editing quality.

  • Applications: Video-P2P supports prompt refinement, attention re-weighting, and word swapping while maintaining semantic consistency across frames and temporal coherence.Examples include changing object properties, adjusting an object's generated extent, and replacing entities while preserving surrounding content.
  • Applications: Video-P2P replaces entities while preserving unrelated regions, including the motorbike, surrounding grass, and background details.Examples include replacing a motorbike rider with Spider-Man and a dog with a cat while preserving gestures and surroundings.
  • Applications: Video-P2P performs both local and global edits, including weather changes, flooding, and style transfer to a watercolor painting.Local edits minimize changes to grass and sky, while global edits alter broader scene properties.
  • Comparison: Video-P2P produces more temporally consistent, structure-preserved results than TAV+DDIM and preserves background details and original motion in the reported comparisons.Compared with TAV+DDIM and Dreamix, the method preserves complex background structure and avoids the motion slowdown reported for Dreamix.
  • Quantitative results: Video-P2P performs well across CLIP Score, Masked PSNR, LPIPS, and OSV, while ranking first on average in the user study.It achieves higher Masked PSNR and lower LPIPS than TAV+DDIM, and lower OSV than the other compared methods.
  • Ablation Study: Fine-tuning the inflated T2S model on the input video recovers generation quality, while shared unconditional embedding optimization improves inversion with lower parameter usage than per-frame embeddings.Per-frame embeddings increase PSNR by only 0.2, use n times more parameters, and yield a lower post-attention-control Masked PSNR of 20.51.

5. Conclusion

Video-P2P enables local and global video editing with cross-attention control by adapting image diffusion models for video inversion and attention control.

  • Video-P2P optimizes a shared unconditional embedding on a tuned T2S model for video inversion.It also uses different unconditional embeddings for source and target prompts and integrates attention maps from both branches.
  • The framework supports word swapping, prompt refinement, and attention re-weighting for video editing.
  • Video-P2P enables both local and global video editing while preserving semantic consistency across frames.
  • Future work will address more complex editing tasks, including injecting extra objects.
Loading 2303.04761v1…