Source-linked AI summary
Seedance 2.0: Advancing Video Generation for World Complexity
Team Seedance, De Chen, Liyang Chen, Xin Chen, Ying Chen, Zhuo Chen, Zhuowei Chen, Feng Cheng, Tianheng Cheng, Yufeng Cheng, Mojie Chi, Xuyan Chi, Jian Cong, Qinpeng Cui, Fei Ding, Qide Dong, Yujiao Du, Haojie Duanmu, Junliang Fan, Jiarui Fang, Jing Fang, Zetao Fang, Chengjian Feng, Yu Gao, Diandian Gu, Dong Guo, Hanzhong Guo, Qiushan Guo, Boyang Hao, Hongxiang Hao, Haoxun He, Jiaao He, Qian He, Tuyen Hoang, Heng Hu, Ruoqing Hu, Yuxiang Hu, Jiancheng Huang, Weilin Huang, Zhaoyang Huang, Zhongyi Huang, Jishuo Jin, Ming Jing, Ashley Kim, Shanshan Lao, Yichong Leng, Bingchuan Li, Gen Li, Haifeng Li, Huixia Li, Jiashi Li, Ming Li, Xiaojie Li, Xingxing Li, Yameng Li, Yiying Li, Yu Li, Yueyan Li, Chao Liang, Han Liang, Jianzhong Liang, Ying Liang, Wang Liao, J. H. Lien, Shanchuan Lin, Xi Lin, Feng Ling, Yue Ling, Fangfang Liu, Jiawei Liu, Jihao Liu, Jingtuo Liu, Shu Liu, Sichao Liu, Wei Liu, Xue Liu, Zuxi Liu, Ruijie Lu, Lecheng Lyu, Jingting Ma, Tianxiang Ma, Xiaonan Nie, Jingzhe Ning, Junjie Pan, Xitong Pan, Ronggui Peng, Xueqiong Qu, Yuxi Ren, Yuchen Shen, Guang Shi, Lei Shi, Yinglong Song, Fan Sun, Li Sun, Renfei Sun, Wenjing Tang, Boyang Tao, Zirui Tao, Dongliang Wang, Feng Wang, Hulin Wang, Ke Wang, Qingyi Wang, Rui Wang, Shuai Wang, Shulei Wang, Weichen Wang, Xuanda Wang, Yanhui Wang, Yue Wang, Yuping Wang, Yuxuan Wang, Zijie Wang, Ziyu Wang, Guoqiang Wei, Meng Wei, Di Wu, Guohong Wu, Hanjie Wu, Huachao Wu, Jian Wu, Jie Wu, Ruolan Wu, Shaojin Wu, Xiaohu Wu, Xinglong Wu, Yonghui Wu, Ruiqi Xia, Xin Xia, Xuefeng Xiao, Shuang Xu, Bangbang Yang, Jiaqi Yang, Runkai Yang, Tao Yang, Yihang Yang, Zhixian Yang, Ziyan Yang, Fulong Ye, Bingqian Yi, Xing Yin, Yongbin You, Linxiao Yuan, Weihong Zeng, Xuejiao Zeng, Yan Zeng, Siyu Zhai, Zhonghua Zhai, Bowen Zhang, Chenlin Zhang, Heng Zhang, Jun Zhang, Manlin Zhang, Peiyuan Zhang, Shuo Zhang, Xiaohe Zhang, Xiaoying Zhang, Xinyan Zhang, Xinyi Zhang, Yichi Zhang, Zixiang Zhang, Haiyu Zhao, Huating Zhao, Liming Zhao, Yian Zhao, Guangcong Zheng, Jianbin Zheng, Xiaozheng Zheng, Zerong Zheng, Kuan Zhu, Feilong Zuo
TL;DR
Video generation research seeks more controllable synthesis for complex multimodal creative workflows. Seedance 2.0 addresses this with unified multimodal generation, reference, editing, and evaluation capabilities, and reports leading performance across evaluated tasks alongside remaining artifact and synchronization issues.
Problem
Seedance 2.0 is developed to address the shift from short video clips with limited controllability toward highly controllable synthesis with diverse control signals.
Method
The model uses a unified multimodal framework supporting text, image, video, and audio references alongside multimodal generation, editing, and evaluation capabilities.
Results
Seedance 2.0 ranks first across all evaluated video and audio dimensions on T2V, I2V, and R2V tasks, with leading performance in multimodal capabilities and human-preference evaluation.
Takeaways & Limitations
The model provides a broad creative-production system combining multimodal controllability, reference-based workflows, video editing, continuation, and audiovisual generation.
Takeaways & Limitations
Remaining issues include minor deformation artifacts, edge-case motion plausibility, visual noise, audio distortion or noise, and multi-speaker lip-sync errors.
Abstract
from arXiv · showhide
Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro, Seedance 2.0 adopts a unified, highly efficient, and large-scale architecture for multi-modal audio-video joint generation. This allows it to support four input modalities: text, image, audio, and video, by integrating one of the most comprehensive suites of multi-modal content reference and editing capabilities available in the industry to date. It delivers substantial, well-rounded improvements across all key sub-dimensions of video and audio generation. In both expert evaluations and public user tests, the model has demonstrated performance on par with the leading levels in the field. Seedance 2.0 supports direct generation of audio-video content with durations ranging from 4 to 15 seconds, with native output resolutions of 480p and 720p. For multi-modal inputs as reference, its current open platform supports up to 3 video clips, 9 images, and 3 audio clips. In addition, we provide Seedance 2.0 Fast version, an accelerated variant of Seedance 2.0 designed to boost generation speed for low-latency scenarios. Seedance 2.0 has delivered significant improvements to its foundational generation capabilities and multi-modal generation performance, bringing an enhanced creative experience for end users.
1 Introduction
Seedance 2.0 advances video generation from limited controllability toward highly controllable multimodal audio-video synthesis. It combines broad reference and editing capabilities with improvements in motion, audiovisual fidelity, and production usability.
- Core contribution: Seedance 2.0 supports diverse control signals and multimodal reference or editing workflows for complex creative production.Capabilities include subject control, motion manipulation, style transfer, special effects, creative generation, and video extension.
- Deployment and usability: Seedance 2.0 supports 4–15-second audio-video generation at 480p or 720p, references up to 3 videos, 9 images, and 3 audio clips, and offers a Fast variant.The Fast version is designed for low-latency scenarios.
- Generation of Real-world Complexity: The model improves human motion modeling through greater naturalness, temporal coherence, physical plausibility, and stability in complex interactions.It targets artifacts common in recent video generation models and supports multi-subject and complex-motion scenarios.
- Strong multimodal capability: Seedance 2.0 interprets text, image, video, and audio references while following complex instructions and preserving subject identity across detailed scripts.It also supports storyboard references, shot sequencing, visual presentation planning, and new video editing capabilities.
- High-Fidelity Audio-Video Generation: The model generates synchronized, immersive binaural audio with multiple tracks, including background audio, sound effects, and character narration.The supplied passage describes precise temporal alignment between audio content and visual rhythm.
- Applications and responsibility: The model is positioned for applications including advertising, cinematic effects, game animation, and commentary videos, while safety assessment remains part of its development lifecycle.The passages describe cross-scene adaptability and structured efforts to evaluate and mitigate potential risks.
2.1 Overview
Seedance 2.0 is evaluated across T2V, I2V, and R2V tasks using multimodal, visual, audio, and instruction-following dimensions. The overview reports leading performance across these tasks while identifying several remaining artifact and synchronization issues.
- Evaluation scope: The evaluation benchmark covers audio-video generation, reference-based generation, and video editing across core multimodal and production-oriented dimensions.These include reference generation, complex instruction following, motion stability, cinematographic expression, and audio-related performance.
- Video results: Video results emphasize improved motion stability, instruction following, visual aesthetics, complex movement, and micro-expression rendering.The model is described as mitigating common structural inaccuracies and visual artifacts.
- Audio results: Audio results emphasize layered expressiveness, audiovisual synchronization, and improved following of dialect, opera, singing, rap, and instrumental prompts.The reported improvements include alignment among lip movements, dialogue, sound effects, background audio, and visual rhythm.
- Reference-based generation: Reference-based generation covers multimodal references, editing, and continuation, with improved understanding and response accuracy for reference content.The overview describes this as competitive leading comprehensive performance.
- Overall results: Seedance 2.0 ranks first across all evaluated video and audio dimensions on T2V, I2V, and R2V tasks.Figure 1 summarizes comprehensive leading performance over competing models across every evaluated dimension.
- Limitations: Remaining issues include minor deformation artifacts, edge-case motion plausibility, high-frequency visual noise, audio distortion or noise, and multi-speaker lip-sync errors.These are explicitly identified as areas for improvement.
2.2 Evaluation Framework
The paper upgrades its evaluation framework to SeedVideoBench 2.0 and complements it with human-preference evaluation on Arena. The framework broadens multimodal and narrative assessment while distinguishing objective metrics from subjective expert review.
- SeedVideoBench 2.0: SeedVideoBench 2.0 adds multimodal generation, narrative quality, multilingual coverage, and refined audio-expressiveness assessment for production-oriented evaluation.Expert evaluators from advertising and game production provide subjective ratings focused on narrative and aesthetic quality.
- Framework design: The framework formally evaluates multimodal task following and generation consistency alongside baseline prompt-following and motion-quality measures.Consistency covers reference alignment and preservation of non-edited regions, supported by specialized datasets.
- Evaluation procedures: Objective metrics use automated pipelines, subjective metrics use blind expert review, and a realism study compares generated outputs with real video clips.Realism-study results were fed back into aesthetic tuning.
- Multimodal task coverage: Multimodal task following spans reference, editing, extension, and combination tasks representing real workflows.Examples include subject, motion, visual-effects, style, scene, audio editing, timeline extension, and paired reference-plus-editing tasks.
- Narrative assessment: Narrative quality assesses coherent storytelling, character performance, visual-effects impact, and fit between aesthetic choices and content.Its dimensions include cinematographic language, plot design, and stylistic aesthetics.
- Arena evaluation: Arena evaluates anonymous side-by-side model outputs through user votes, producing an Elo-style leaderboard that captures human preferences at scale.This provides a complementary human-preference perspective to automated benchmarks.
- Arena results: Dreamina Seedance 2.0 720p ranks #1 on both T2V and I2V Arena leaderboards, with Elo scores of 1450 (±15) and 1449 (±11).It leads the second-place models by 79 points on T2V and 29 points on I2V.
2.3 Text-to-Video Evaluation on SeedVideoBench 2.0
Seedance 2.0 leads text-to-video evaluation across six dimensions, with particularly strong motion, audio, synchronization, prompt-following, and aesthetics results. It also achieves broad category-level advantages over competing models and Seedance 1.5.
- Overall Results: Seedance 2.0 ranks first on all six T2V dimensions and improves over Seedance 1.5 by an average of 0.86 points.The largest gain is on motion quality (+1.36), while motion quality and audio-visual sync both reach 3.75.
- Overall Results: Seedance 2.0 is the only model above 3.4 on every T2V dimension, while competitors generally remain below 2.9 on audio dimensions.Seedance 2.0 exceeds 3.5 on all three audio dimensions.
- Motion Quality: Seedance 2.0 ranks first on 29 of 30 fine-grained motion categories, scoring 3.29–4.43.It leads on motion stability, editing rhythm, and multi-entity interaction, with fewer subject deformations and physically implausible motions than Seedance 1.5.
- Video Prompt Following: Seedance 2.0 ranks first on 27 of 30 video prompt-following categories, scoring 2.71–4.29.Its largest gains over Seedance 1.5 are in creative text, short text, text overlay, physical phenomena, and natural phenomena.
- Video Aesthetics: Seedance 2.0 ranks first or tied for first on 28 of 30 aesthetics categories, scoring 2.79–4.14.Its strongest categories include visual style, long script, framing/composition, cinematic visual effects, editing rhythm, natural phenomena, and multi-entity feature match.
- Audio and Synchronization: Seedance 2.0 ranks first on 16 of 17 audio-visual synchronization categories, with strong lip synchronization and action-audio alignment.It ranks first on English, singing/rap, dual-channel audio, and non-verbal voice; audio prompt following is the dimension where competitors score lowest.
2.4 Image-to-Video Evaluation on SeedVideoBench 2.0
Seedance 2.0 leads image-to-video evaluation across visual and audio dimensions, with especially strong performance on complex instructions, motion, interaction, and voice generation. Its main weakness is advanced camera movement, which remains difficult for all models.
- Overall Results: Seedance 2.0 ranks first on all six I2V dimensions, scoring 3.31–3.70, while no competitor exceeds 3.18.It leads on both video and audio, with the largest separation occurring in audio dimensions.
- Complex Instructions: Seedance 2.0 leads on compound multi-instruction handling, scoring MQ 4.00 and VPF 3.75, over 1 point above Kling 3.0 on MQ.Its MQ/VPF scores improved from Seedance 1.5 Pro’s 2.13/2.25.
- Complex Camera: Advanced camera movement is the hardest camera sub-category: Seedance 2.0 and Kling 3.0 tie on MQ at 2.71, and no model exceeds 3.14 on any metric.Camera flexibility remains an area for improvement across all models.
- Complex Motion: Seedance 2.0’s strongest motion results include sports at MQ 3.73 and VPF 3.93, and micro-expression and emotion at VPF 4.00.Combat visual-effects MQ reaches 3.63, compared with 2.25 for Kling 3.0 and Seedance 1.5 Pro.
- Complex Interaction: Same-type interaction reaches MQ 3.64, IP 3.82, and VPF 3.91, while cross-type interaction reaches VPF 4.00.Group motion remains difficult, with Seedance 2.0 scoring 3.00/3.00/2.88.
- Detailed Audio Evaluation: Seedance 2.0 leads all four composite voice sub-categories, including singing/rap APF 4.10 and off-screen voice scores of 3.75/3.75/3.88.Its off-screen narration synchronization remains stronger than competitors’ results.
2.5 Reference-to-Video Evaluation on SeedVideoBench 2.0
Seedance 2.0 leads reference-to-video evaluation across the broadest set of multimodal tasks and most quality dimensions. However, video extension is a notable weakness, and continuation still has specific consistency issues.
- Quantitative Results: Seedance 2.0 leads all evaluated models on R2V across multimodal task following, prompt following, editing consistency, reference alignment, and motion quality.It scores 2.50 and 2.52 on the two 1–3 dimensions, and 3.54, 3.03, and 3.24 on the 1–5 dimensions.
- Multimodal Task Support: Seedance 2.0 supports 20 of 22 input modalities, including seven visual-effects, creative-reference, continuation, and extension tasks unavailable from competitors.The two unsupported tasks are also unsupported by every other evaluated model.
- Reference Alignment: Seedance 2.0 scores 2.64 on motion-reference alignment, while all competitors fall below 2.0; Kling 3 Omni instead leads first-frame preservation at 4.31 versus 2.71.The comparison reflects a trade-off between dynamic subsequent motion and first-frame fidelity.
- Video Editing: In video editing, Kling O1 slightly leads task following at 2.29 versus Seedance 2.0’s 2.20, but Seedance 2.0 leads reference alignment at 3.79 and editing consistency at 3.75.Seedance 2.0 handles long-text and multi-edit instructions more completely.
- Video Continuation: Video continuation is supported only by Seedance 2.0, which scores 2.88 on task following and 3.18 on reference alignment.Remaining issues include color consistency, multi-subject omission, and subject duplication.
- Video Extension: For extension, Veo 3.1 scores 2.78 on task following versus Seedance 2.0’s 1.93, making extension Seedance 2.0’s weakest R2V task.Seedance 2.0 supports arbitrary uploaded videos, whereas Veo 3.1 only extends videos it generated itself.
2.6 Visualization Results
Figure 4 visualizes Seedance 2.0’s text-to-video and image-to-video generation results. The accompanying description emphasizes temporally precise, physically coherent motion in complex interactive scenes.
- Visualization Results: The described generations maintain motion quality by adhering to real-world physical laws and avoiding physical anomalies common in earlier AI video models.The examples involve temporally precise and complex interactive scenes rendered with high fidelity.
- Visualization Results: Figure 4 presents visualizations of text-to-video and image-to-video generation.The figure covers both T2V and I2V outputs.
3 Contributions and Acknowledgments
The contribution and acknowledgment section states that all Seedance-2.0 authors are listed alphabetically by last name.
- Contributions and Acknowledgments: All Seedance-2.0 authors are listed in alphabetical order by their last names.
De Chen
This passage lists contributors whose names include multiple individuals with the surnames Chen and Cheng.
- The passage lists Liyang Chen, Xin Chen, Ying Chen, Zhuo Chen, and Zhuowei Chen.
- It also lists Feng Cheng, Tianheng Cheng, Yufeng Cheng, Mojie Chi, and Xuyan Chi.
- Additional listed contributors include Jian Cong, Qinpeng Cui, Fei Ding, Qide Dong, and Yujiao Du.
Yueyan Li
These passages list contributors including Liang, Liu, Wang, Zhang, Zhao, Zheng, and Zhu surnames, along with many additional named individuals.
- The passages list contributors including Chao Liang, Han Liang, Jianzhong Liang, Ying Liang, and Wang Liao.
- They include Ashley Kim, Shanshan Lao, Yichong Leng, Bingchuan Li, Gen Li, and numerous additional contributors.
- The list includes Dongliang Wang, Feng Wang, Hulin Wang, Ke Wang, and Qingyi Wang.
- Additional contributors include Bowen Zhang, Chenlin Zhang, Heng Zhang, Jun Zhang, and Manlin Zhang.
- The passages also list Xiaoying Zhang, Xinyan Zhang, Xinyi Zhang, Yichi Zhang, and Zixiang Zhang.
- The final passage lists Yian Zhao, Guangcong Zheng, Jianbin Zheng, Xiaozheng Zheng, Zerong Zheng, Kuan Zhu, Feilong Zhu, and Zuo.