Source-linked AI summary
SkyReels-V3 Technique Report
Debang Li, Zhengcong Fei, Tuanhui Li, Yikun Dou, Zheng Chen, Jiangping Yang, Mingyuan Fan, Jingtao Xu, Jiahua Wang, Baoxuan Gu, Mingshan Chang, Wenjing Cai, Yuqiang Xie, Binjie Mao, Youqiang Zhang, Nuo Pang, Hao Zhang, Yuzhe Jin, Zhiheng Xu, Dixuan Lin, Guibin Chen, Yahui Zhou
TL;DR
SkyReels-V3 addresses the challenge of multimodal, coherent video generation for world-model development. It uses a unified multimodal in-context learning framework with diffusion Transformers to support reference-image synthesis, video extension, and audio-guided talking avatars, with evaluations reporting strong performance across diverse tasks.
Problem
Video generation must model complex real-world dynamics and multimodal context to support world models and practical artificial-intelligence deployment.
Method
SkyReels-V3 unifies visual, video, audio, and textual conditioning with diffusion Transformers, hybrid image-video training, and spatiotemporal modeling across three generation paradigms.
Results
Extensive evaluations demonstrate strong subject consistency, motion fidelity, instruction following, and generalization across reference synthesis, video extension, and talking-avatar generation.
Takeaways & Limitations
SkyReels-V3 provides a unified open-source model family for coherent narrative-level video creation across diverse tasks and aspect ratios.
Abstract
from arXiv · showhide
Video generation serves as a cornerstone for building world models, where multimodal contextual inference stands as the defining test of capability. In this end, we present SkyReels-V3, a conditional video generation model, built upon a unified multimodal in-context learning framework with diffusion Transformers. SkyReels-V3 model supports three core generative paradigms within a single architecture: reference images-to-video synthesis, video-to-video extension and audio-guided video generation. (i) reference images-to-video model is designed to produce high-fidelity videos with strong subject identity preservation, temporal coherence, and narrative consistency. To enhance reference adherence and compositional stability, we design a comprehensive data processing pipeline that leverages cross frame pairing, image editing, and semantic rewriting, effectively mitigating copy paste artifacts. During training, an image video hybrid strategy combined with multi-resolution joint optimization is employed to improve generalization and robustness across diverse scenarios. (ii) video extension model integrates spatio-temporal consistency modeling with large-scale video understanding, enabling both seamless single-shot continuation and intelligent multi-shot switching with professional cinematographic patterns. (iii) Talking avatar model supports minute-level audio-conditioned video generation by training first-and-last frame insertion patterns and reconstructing key-frame inference paradigms. On the basis of ensuring visual quality, synchronization of audio and videos has been optimized. Extensive evaluations demonstrate that SkyReels-V3 achieves state-of-the-art or near state-of-the-art performance on key metrics including visual quality, instruction following, and specific aspect metrics, approaching leading closed-source systems. Github: https://github.com/SkyworkAI/SkyReels-V3.
1 Introduction
SkyReels-V3 is a unified multimodal video-generation framework supporting reference-image synthesis, video extension, and audio-guided generation. Its design targets coherent, high-fidelity video creation across subjects, contexts, and practical scenarios.
- Video generation frameworks encode geometric, semantic, and physical knowledge through visual-sequence synthesis for multimodal world modeling.
- SkyReels-V3 unifies visual reference, video, audio, and textual inputs within one in-context learning framework.
- The framework supports reference images-to-video, video-to-video extension, and audio-guided talking-avatar generation.
- Diffusion Transformers, multimodal alignment, and spatiotemporal consistency modeling support instruction following, motion generation, identity preservation, and audio-visual synchronization.
- Extensive evaluations report industry-leading or higher performance across key metrics and applicability to professional production, avatars, commerce, and entertainment.
2 Methods and Evaluation
SkyReels-V3 unifies reference image-to-video synthesis, video extension, and audio-guided generation, with evaluation spanning consistency, instruction following, and visual quality. Its methods combine multi-reference conditioning, hybrid multi-resolution training, shot-aware extension modeling, and audio–visual alignment.
- Reference Images to Video: SkyReels-V3 supports multi-reference image-to-video synthesis conditioned on up to four images and textual prompts.The model preserves identity attributes, spatial composition, and narrative continuity while following semantic instructions.
- Reference Images to Video: A dedicated data pipeline constructs reference image-to-video pairs from high-quality, dynamic clips using cross-frame pairing.The passage identifies reference-pair quality as central and describes filtering and frame selection from continuous sequences.
- Reference Images to Video: Image–video hybrid training and multi-resolution joint optimization improve generalization across static and dynamic cues, spatial scales, and aspect ratios.The training scheme jointly uses image and video datasets and supports varied output configurations.
- Video Extension: The video extension model supports both single-shot continuation and shot switching while preserving motion dynamics, scene structure, visual style, and narrative coherence.A shot-switching detector classifies single-shot, cut-in, cut-out, multi-angle, shot/reverse-shot, and cut-away transitions for training-data construction.
- Video Extension: Extension results cover cinematic creation, short-form series, game cutscenes, and long-form enhancement, producing high-definition outputs with sharp details and natural motion.Figures 5–10 present single-shot and multiple shot-switching results.
- Audio-Guided Generation: The talking avatar model generates audio-conditioned video from a portrait and audio, jointly modeling audio, visual inputs, and textual cues for expressions, head movements, camera dynamics, and lip synchronization.The model supports long-form generation and multi-character interactions, with region masking used to model speech–facial-motion correspondence.
3 Conclusion
SkyReels-V3 unifies reference-based synthesis, video extension, and audio-driven talking-avatar generation within one multimodal in-context learning framework. Across diverse tasks and aspect ratios, it achieves strong subject consistency, high-fidelity motion, robust instruction following, and competitive benchmark performance.
- SkyReels-V3 integrates reference-based video synthesis, video extension, and audio-driven talking-avatar generation in one in-context learning framework.
- Jointly modeling visual, temporal, and auditory signals advances generation from short frame-level synthesis toward coherent narrative-level content creation.
- Multimodal conditioning, hybrid image–video training, hierarchical spatiotemporal modeling, and efficient token fusion support subject consistency, high-fidelity motion, and instruction following.
- Extensive evaluations show competitive performance across multiple benchmarks, rivaling leading closed-source models in various domains.
Contributors and Acknowledgments
The report lists the contributors to SkyReels-V3. The listed contributors are Debang Li, Zhengcong Fei, Tuanhui Li, Yikun Dou, Zheng Chen, Jiangping Yang, Mingyuan Fan, Jingtao Xu, Jiahua Wang, Baoxuan Gu, Mingshan Chang, Wenjing Cai, Yuqiang Xie, Binjie Mao, Youqiang Zhang, Nuo Pang, Hao Zhang, Yuzhe Jin, Zhiheng Xu, Dixuan Lin, Guibin Chen, and Yahui Zhou.
- The contributor list begins with Debang Li, Zhengcong Fei, Tuanhui Li, Yikun Dou, Zheng Chen, and Jiangping Yang.
- The contributor list continues with Mingyuan Fan, Jingtao Xu, Jiahua Wang, Baoxuan Gu, Mingshan Chang, and Wenjing Cai.
- Additional listed contributors are Yuqiang Xie, Binjie Mao, Youqiang Zhang, Nuo Pang, Hao Zhang, Yuzhe Jin, Zhiheng Xu, Dixuan Lin, Guibin Chen, and Yahui Zhou.