Source-linked AI summary

Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Weichen Fan, Chenyang Si, Junhao Song, Zhenyu Yang, Yinan He, Long Zhuo, Ziqi Huang, Ziyue Dong, Jingwen He, Dongwei Pan, Yi Wang, Yuming Jiang, Yaohui Wang, Peng Gao, Xinyuan Chen, Hengjie Li, Dahua Lin, Yu Qiao, Ziwei Liu

arXiv:2501.08453v1cs.CVcs.LG

TL;DR

Video diffusion models face challenges in maintaining spatial and temporal coherence while scaling training to long sequences and large distributed systems. Vchitect-2.0 combines a multimodal diffusion block, memory-efficient parallel training, and a curated dataset to address these challenges. Its evaluations report stronger video-generation quality and scalability, while substantial hardware demands remain a limitation.

  • Problem

    Long video generation requires spatial fidelity, temporal consistency, scalable computation, and high-quality annotated data.

  • Method

    Vchitect-2.0 combines a multimodal diffusion block, hybrid parallelism with memory optimization, and a curated dataset evaluated through annotation and aesthetics scoring.

  • Results

    The full model improves Total Score, Overall Consistency, Aesthetic Quality, and Imaging Quality over a vanilla attention baseline in VBench evaluation.

  • Takeaways & Limitations

    Vchitect-2.0 provides a scalable framework for high-fidelity, text-aligned video generation and long-sequence training on distributed systems.

  • Takeaways & Limitations

    Broader applicability remains constrained by token length, data diversity, and computational efficiency, with substantial hardware requirements for training and deployment.

Abstract

from arXiv · show

We present Vchitect-2.0, a parallel transformer architecture designed to scale up video diffusion models for large-scale text-to-video generation. The overall Vchitect-2.0 system has several key designs. (1) By introducing a novel Multimodal Diffusion Block, our approach achieves consistent alignment between text descriptions and generated video frames, while maintaining temporal coherence across sequences. (2) To overcome memory and computational bottlenecks, we propose a Memory-efficient Training framework that incorporates hybrid parallelism and other memory reduction techniques, enabling efficient training of long video sequences on distributed systems. (3) Additionally, our enhanced data processing pipeline ensures the creation of Vchitect T2V DataVerse, a high-quality million-scale training dataset through rigorous annotation and aesthetic evaluation. Extensive benchmarking demonstrates that Vchitect-2.0 outperforms existing methods in video quality, training efficiency, and scalability, serving as a suitable base for high-fidelity video generation.

I. INTRODUCTION

Vchitect-2.0 addresses text-to-video challenges in temporal consistency, scalability, computational efficiency, and annotated-data availability through a parallel transformer architecture, hybrid training framework, and curated dataset. Its experiments and ablations report improvements in video quality, training scalability, and efficiency.

  • Video generation requires both high spatial fidelity and seamless temporal consistency, while long sequences increase computational and memory demands.
  • Vchitect T2V DataVerse uses detailed annotation and aesthetic evaluation to preserve text-video alignment across diverse and complex tasks.
  • Vchitect-2.0 uses a multimodal diffusion block to align text prompts with frame-wise features while maintaining temporal consistency.
  • Its hybrid parallelism combines data and sequence parallelism with recomputation and offloading to train long-duration, high-resolution videos on distributed systems.
  • Benchmark evaluations and qualitative analyses report improvements in temporal coherence, spatial fidelity, computational efficiency, smoother transitions, and reduced motion artifacts.
  • Ablation studies identify the multimodal diffusion block and parallelism strategies as important contributors to overall performance.

II. RELATED WORKS

Related work adapts image diffusion models for video through temporal modules and develops sequence-parallel methods for long-context transformers. These approaches trade off attention-head limits, cross-node communication, and scalability.

  • Make-A-Video, Imagen Video, and MagicVideo extend text-to-image diffusion with spatio-temporal attention, directed temporal attention, or 3D convolutions.
  • Sequence parallelism addresses long-context memory demands through head parallelism or context parallelism with distributed attention communication.
  • Head parallelism is limited by attention-head count, whereas context parallelism incurs cross-node communication overhead.
  • USP and LoongTrain improve sequence-parallel scalability through hybrid attention strategies and flexible context- and head-parallel configurations.

C. Memory-Efficient Training

Memory-efficient video training combines distributed memory strategies with single-device optimization techniques. Diffusion and latent-diffusion preliminaries frame the computational demands that these methods address.

  • ZeRO distributes parameters, gradients, and optimizer states across devices, while tensor and pipeline parallelism require more extensive model refactoring.
  • Mixed precision, selective gradient checkpointing, and activation offloading reduce memory overhead by lowering precision, recomputing activations, or moving them to host memory.
  • DDPMs learn data distributions through forward Gaussian-noise diffusion followed by reverse denoising from a noisy sample.
  • A time-conditional UNet estimates reverse-process parameters and is trained with a denoising objective.
  • LDMs compress inputs into learned latent representations before diffusion, reducing parameter count and memory consumption while maintaining performance.

B. Parallel Transformer Architecture

Vchitect-2.0 extends a multimodal diffusion transformer to jointly process text and video features through full-sequence, spatial, and temporal attention. These branches are combined to support text-frame alignment and spatial-temporal consistency.

  • Vchitect-2.0 replaces separate self-attention and cross-attention layers with unified multimodal diffusion blocks augmented by learnable context latents.
  • Text and visual embeddings are interleaved in a checkerboard sequence for full-sequence cross-attention anchored by the first-frame text embedding.
  • Spatial and temporal attention reshape the [B, F, L, H, W] input differently to model within-frame structure and across-frame dynamics.
  • The architecture combines full-sequence, temporal, and spatial attention over text and video features to maintain spatial and temporal consistency.
  • The outputs of the spatial, temporal, and full-sequence branches are aggregated by element-wise addition into the final feature representation.

IV. VCHITECT T2V DATAVERSE

The Vchitect T2V DataVerse combines public and internally sourced videos with staged segmentation, filtering, annotation, and quality assessment. Its pipeline targets coherent, dynamic, well-described, and aesthetically suitable clips for text-to-video training.

  • Vchitect-2.0 combines public datasets with 1 million internally sourced videos and applies filtering and re-annotation for longer, higher-quality training data.
  • The pipeline segments long videos into shots, stitches related segments into coherent events, and removes clips dominated by static frames.
  • Aesthetic Evaluation uses a predictor trained on 441k samples to assign frame scores from 1 to 10 for subsequent filtering.
  • Dynamic Estimation uses RAFT-based motion analysis to assess movement patterns and identify clips with excessive shake or minimal scene changes.
  • Video Captioning uses LLaVA-Next-Video and fine-tuned VideoChat with 16 input frames to generate detailed descriptions for processed clips.
  • Text Localization and watermark classification identify unwanted text and watermarked content for dataset quality control.

B. Training Data Statistics

Vchitect-2.0’s training data combines five public and in-house datasets, with the in-house collection emphasizing longer, higher-resolution videos and richer annotations. Its aesthetic scores and captions improve substantially over earlier Vchitect data.

  • Nearly half of the in-house videos score above 6 aesthetically, compared with 16.89% in Vchitect.
  • Vchitect-XL increased video training duration while maintaining category diversity and recaptioned existing data with an average caption length of 100 tokens.
  • Vchitect-XL annotations capture camera movement and describe content before and after picture changes more dynamically than the original annotations.
  • The training-data table covers five datasets across domain, video count, duration, caption length, and resolution, including in-house videos reaching 4K.
  • The in-house data spans diverse categories, while its aesthetic distribution is compared directly with public datasets.

V. SEQUENCE PARALLELIZED VIDEO TRAINING

Long video sequences create severe activation-memory and communication challenges for distributed training. Vchitect-2.0 addresses them by limiting sequence-parallel groups within nodes, using data parallelism across nodes, and adding memory-saving techniques.

  • Long, high-quality videos produce excessive forward-activation memory because their 3D visual representations contain many vision tokens.
  • Inter-node sequence parallelism is difficult because cross-node bandwidth is much lower than intra-node connectivity, making communication slow and unstable.
  • The model’s concatenated text-visual attention and separately parameterized operators make simplistic sharding vulnerable to load imbalance.
  • The training framework combines parallelism and memory optimization to support longer video sequences on distributed systems.
  • Activation reduction is central because parameter, gradient, optimizer, and EMA-state memory is relatively small compared with activation memory.
  • Vchitect-2.0 limits sequence-parallel groups within a node and applies data parallelism across nodes to achieve training throughput without complex refactoring.

B. Memory Efficient Video Training

Vchitect-2.0 combines hybrid parallelism with activation offloading, recomputation, and distributed memory management to train long videos efficiently. Specialized sequence and VAE designs further improve scalability and reduce overhead.

  • Activation offloading and recomputation are combined because CPU-GPU bandwidth limits offloading overlap, while selective offloading can replace recomputation in some layers.
  • Spatial sequence slicing, head parallelism, and separate sharding of text and visual tokens are specialized for the model’s 3D multimodal sequences.
  • Frame-wise data-parallel VAE inference parallelizes processing of the much larger original video inputs and helps avoid GPU out-of-memory failures.
  • The training system combines intra-node sequence parallelism, inter-node data parallelism, offloading, recomputation, and FSDP to reduce memory footprint.
  • LiteGen supports flexible combinations of parallelism and training techniques with simple configurations and minimal code refactoring.

A. Implementation Details

The implementation combines staged training, distributed workflow components, and VBench-based evaluation. Results show stronger consistency and visual-quality metrics than CogVideoX-2B, while post-processing raises performance above Kling.

  • Training progresses from WebVid-10M pretraining through Panda70M, longer-sequence and higher-resolution stages, and refinement on InternVid-18M-aes, Vimeo, and internal data.
  • The workflow uses Framewise VAE processing, all-to-all communication, parallel Transformer Blocks, replicated unpatchify operations, and final all-gather reconstruction.
  • VBench evaluation reports Total Score together with Dynamic Degree, Overall Consistency, Aesthetic Quality, Imaging Quality, Human Action, and Spatial Relationship.
  • 1.35% higher Overall Consistency, 0.65% higher Aesthetic Quality, and 3.92% higher Imaging Quality distinguish Vchitect-2.0-2B from CogVideoX-2B.
  • Vchitect-2.0-2B [E] exceeds the commercial model Kling by 0.39% in Total Score after VEhancer post-processing.

C. Human Evaluation

Human evaluation compares models on video-text alignment, temporal quality, and frame-wise quality, while efficiency experiments examine memory strategies and scalable parallel training. The reported results favor the proposed workflow’s throughput and long-sequence handling.

  • Human Evaluation: Human evaluators compare paired videos on video-text alignment, temporal quality, and frame-wise quality across five models.
  • Memory-Efficient Training: A combination of activation offloading and recomputation reduces the overhead of simple recomputation by about 3% for long input sequences.
  • SP Workflow: Up to 20% performance gain comes from separately sharding multimodal sequences, followed by another 30% throughput improvement from frame-wise DP-VAE and patchifying.
  • Model Architecture: Full-sequence cross-attention improves Total Score, Overall Consistency, Aesthetic Quality, and Imaging Quality over the vanilla temporal-and-spatial-attention baseline.

VII. DISCUSSION

The discussion identifies remaining limits in long-sequence coherence, dataset coverage, and computational accessibility. Despite these constraints, the paper concludes that Vchitect-2.0 advances scalable, high-fidelity text-to-video generation.

  • Performance may decline on longer video sequences as errors accumulate and temporal inconsistencies emerge.
  • Dependence on Vchitect T2V DataVerse may limit generalization because its diversity and quality may not cover the full spectrum of real-world scenarios.
  • Despite memory optimization, training and deployment remain computationally demanding and require substantial hardware resources.
  • Broader applicability requires addressing token length, data diversity, and computational efficiency limitations.
  • The proposed architecture combines multimodal diffusion, sequence-parallel training, memory-efficient strategies, and curated data to improve video quality, efficiency, and scalability.
Loading 2501.08453v1…