Source-linked AI summary
Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model
Team Seedance, Heyi Chen, Siyan Chen, Xin Chen, Yanfei Chen, Ying Chen, Zhuo Chen, Feng Cheng, Tianheng Cheng, Xinqi Cheng, Xuyan Chi, Jian Cong, Jing Cui, Qinpeng Cui, Qide Dong, Junliang Fan, Jing Fang, Zetao Fang, Chengjian Feng, Han Feng, Mingyuan Gao, Yu Gao, Dong Guo, Qiushan Guo, Boyang Hao, Qingkai Hao, Bibo He, Qian He, Tuyen Hoang, Ruoqing Hu, Xi Hu, Weilin Huang, Zhaoyang Huang, Zhongyi Huang, Donglei Ji, Siqi Jiang, Wei Jiang, Yunpu Jiang, Zhuo Jiang, Ashley Kim, Jianan Kong, Zhichao Lai, Shanshan Lao, Yichong Leng, Ai Li, Feiya Li, Gen Li, Huixia Li, JiaShi Li, Liang Li, Ming Li, Shanshan Li, Tao Li, Xian Li, Xiaojie Li, Xiaoyang Li, Xingxing Li, Yameng Li, Yifu Li, Yiying Li, Chao Liang, Han Liang, Jianzhong Liang, Ying Liang, Zhiqiang Liang, Wang Liao, Yalin Liao, Heng Lin, Kengyu Lin, Shanchuan Lin, Xi Lin, Zhijie Lin, Feng Ling, Fangfang Liu, Gaohong Liu, Jiawei Liu, Jie Liu, Jihao Liu, Shouda Liu, Shu Liu, Sichao Liu, Songwei Liu, Xin Liu, Xue Liu, Yibo Liu, Zikun Liu, Zuxi Liu, Junlin Lyu, Lecheng Lyu, Qian Lyu, Han Mu, Xiaonan Nie, Jingzhe Ning, Xitong Pan, Yanghua Peng, Lianke Qin, Xueqiong Qu, Yuxi Ren, Kai Shen, Guang Shi, Lei Shi, Yan Song, Yinglong Song, Fan Sun, Li Sun, Renfei Sun, Yan Sun, Zeyu Sun, Wenjing Tang, Yaxue Tang, Zirui Tao, Feng Wang, Furui Wang, Jinran Wang, Junkai Wang, Ke Wang, Kexin Wang, Qingyi Wang, Rui Wang, Sen Wang, Shuai Wang, Tingru Wang, Weichen Wang, Xin Wang, Yanhui Wang, Yue Wang, Yuping Wang, Yuxuan Wang, Ziyu Wang, Guoqiang Wei, Wanru Wei, Di Wu, Guohong Wu, Hanjie Wu, Jian Wu, Jie Wu, Ruolan Wu, Xinglong Wu, Yonghui Wu, Ruiqi Xia, Liang Xiang, Fei Xiao, XueFeng Xiao, Pan Xie, Shuangyi Xie, Shuang Xu, Jinlan Xue, Shen Yan, Bangbang Yang, Ceyuan Yang, Jiaqi Yang, Runkai Yang, Tao Yang, Yang Yang, Yihang Yang, ZhiXian Yang, Ziyan Yang, Songting Yao, Yifan Yao, Zilyu Ye, Bowen Yu, Jian Yu, Chujie Yuan, Linxiao Yuan, Sichun Zeng, Weihong Zeng, Xuejiao Zeng, Yan Zeng, Chuntao Zhang, Heng Zhang, Jingjie Zhang, Kuo Zhang, Liang Zhang, Liying Zhang, Manlin Zhang, Ting Zhang, Weida Zhang, Xiaohe Zhang, Xinyan Zhang, Yan Zhang, Yuan Zhang, Zixiang Zhang, Fengxuan Zhao, Huating Zhao, Yang Zhao, Hao Zheng, Jianbin Zheng, Xiaozheng Zheng, Yangyang Zheng, Yijie Zheng, Jiexin Zhou, Jiahui Zhu, Kuan Zhu, Shenhan Zhu, Wenjia Zhu, Benhui Zou, Feilong Zuo
TL;DR
Seedance 1.5 pro addresses the need for unified audio-video generation that produces more complete, production-ready works rather than fragmented visuals. It combines a unified multimodal architecture, curated audio-visual data, post-training, and inference acceleration, reporting precise synchronization, multilingual support, cinematic control, and narrative coherence, while retaining evolving mastery of specific opera vocal styles.
Problem
As video generation moves toward multimodal integration, unified audio-video generation requires coherent synchronization and emotional alignment for production-ready outputs.
Method
The model combines a unified MMDiT-based multimodal architecture with a multi-stage audio-visual data framework, SFT, audio-video RLHF, and accelerated inference.
Results
The model reports precise audio-visual synchronization and multilingual dialect support, cinematic camera control, enhanced narrative coherence, and end-to-end inference acceleration exceeding 10×.
Takeaways & Limitations
Seedance 1.5 pro demonstrates potential for professional-grade content creation and practical multi-shot video generation workflows.
Takeaways & Limitations
Mastery of specific vocal styles across different opera sub-genres is still evolving.
Abstract
from arXiv · showhide
Recent strides in video generation have paved the way for unified audio-visual generation. In this work, we present Seedance 1.5 pro, a foundational model engineered specifically for native, joint audio-video generation. Leveraging a dual-branch Diffusion Transformer architecture, the model integrates a cross-modal joint module with a specialized multi-stage data pipeline, achieving exceptional audio-visual synchronization and superior generation quality. To ensure practical utility, we implement meticulous post-training optimizations, including Supervised Fine-Tuning (SFT) on high-quality datasets and Reinforcement Learning from Human Feedback (RLHF) with multi-dimensional reward models. Furthermore, we introduce an acceleration framework that boosts inference speed by over 10X. Seedance 1.5 pro distinguishes itself through precise multilingual and dialect lip-syncing, dynamic cinematic camera control, and enhanced narrative coherence, positioning it as a robust engine for professional-grade content creation. Seedance 1.5 pro is now accessible on Volcano Engine at https://console.volcengine.com/ark/region:ark+cn-beijing/experience/vision?type=GenVideo.
1 Introduction
Seedance 1.5 pro is a native joint audio-video foundation model built around multimodal data, unified generation, post-training optimization, and accelerated inference. It targets synchronized multilingual audio-visual generation, cinematic control, narrative coherence, and practical multi-shot content creation.
- Core model and architecture: Seedance 1.5 pro is a foundational model with native support for joint video-audio generation across text-to-audio-video and image-guided audio-video tasks.Its scope also includes unimodal text-to-video and image-to-video generation through multi-task pre-training.
- Data and optimization: The audio-visual data framework combines multi-stage curation, advanced captioning, curriculum-based scheduling, and scalable infrastructure, prioritizing coherence and motion expressiveness.The captioning system supplies professional-grade descriptions for both video and audio modalities.
- Core model and architecture: The unified MMDiT-based architecture enables cross-modal interaction intended to preserve temporal synchronization and semantic consistency between visual and auditory streams.The model is trained on large-scale mixed-modality datasets for generalization across T2VA, I2VA, T2V, and I2V.
- Data and optimization: The training pipeline applies supervised fine-tuning and audio-video-specific RLHF with multi-dimensional rewards, while infrastructure optimizations improve RLHF training speed by nearly 3×.The reward model targets motion quality, visual aesthetics, and audio fidelity in T2V and I2V tasks.
- Inference efficiency: Inference distillation, quantization, and parallelism deliver end-to-end acceleration exceeding 10× while preserving model performance.The optimization reduces the Number of Function Evaluations required during generation.
- Capabilities and applications: The model supports precise lip-syncing, intonation, and performance-rhythm alignment across languages and regional dialects, alongside dynamic camera scheduling and stronger narrative coordination.Reported camera capabilities include continuous long takes, dolly zooms, cinematic transitions, and professional color grading; applications include Chinese film production, short dramas, and traditional performing arts.
2 Evaluation
Seedance 1.5 pro is evaluated as a native audio–video model across video and audio dimensions, with benchmarks covering diverse tasks and application scenarios. The reported results emphasize improvements in video generation, Chinese-language audio, synchronization, expressive performance, camera control, and narrative coherence.
- Evaluation framework: SeedVideoBench 1.5 expands evaluation beyond video generation to include audio dimensions and application scenarios such as advertising, social media, and short-form narratives.The framework includes taxonomies for subjects, motion, interactions, camera movements, human voices, non-speech audio, and audio quality.
- Video evaluation: Seedance 1.5 pro improves over Seedance 1.0 Pro, leading T2V instruction following and remaining competitive in visual aesthetics, motion dynamics, and I2V.The comparison uses absolute evaluation scores for Text-to-Video and Image-to-Video tasks.
- Audio evaluation: Seedance 1.5 pro outperforms Veo 3.1 in Chinese dialogue, dialect, and monologue generation while largely avoiding syllable dropping and mispronunciation.The reported advantage includes accurate responses and high articulatory clarity in Chinese-language contexts.
- Audio–visual synchronization: Seedance 1.5 pro surpasses Veo 3.1 and Kling 2.6 in lip–audio synchronization by aligning speaking characters, mouth motion, and sound effects with visual cues.The model mitigates mouth-motion redundancy and omission that can produce temporal misalignment.
- Application scenarios: Seedance 1.5 pro maintains consistent lip synchronization, vocal tonality, and performance rhythm across continuous shots, including Sichuanese, Taiwan Mandarin, Cantonese, and Shanghainese.The paper reports natural prosody and speech patterns in these dialect-rich settings.
- Application scenarios: Seedance 1.5 pro controls orbital, arc, and tracking shots while preserving style consistency and introducing genre-aligned subjects and actions for narrative continuity.The paper also reports coherent facial micro-expressions and expressive performance in realistic close-ups.
- Application scenarios: Seedance 1.5 pro captures distinctive operatic speech cadence and Eastern theatrical atmosphere, although mastery of specific vocal styles across opera sub-genres is still evolving.Reported performance details include orchid hand gestures and stylized eye expressions in comedic roles.
- Application scenarios: The reported capabilities support audio–visual integration and narrative expressiveness in Chinese film production, short-form drama, and opera-inspired storytelling.The paper frames these capabilities as enhancing creative controllability across these application domains.
A Contributions and Acknowledgments
The acknowledgments section states that all Seedance authors are listed alphabetically by last name.
- All authors of Seedance are listed in alphabetical order by their last names.
Authors
The paper lists a large author group spanning the surnames represented in the author list.
- The author list includes contributors with surnames ranging from Chen, Cheng, and Cui through Huang, Jiang, Wang, Yang, Zhang, and Zhu.
- The listed contributors include multiple researchers sharing surnames such as Li, Liu, Wang, Wu, Yang, and Zhang.