Source-linked AI summary
Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis
Yikang Ding, Jiwen Liu, Wenyuan Zhang, Zekun Wang, Wentao Hu, Liyuan Cui, Mingming Lao, Yingchao Shao, Hui Liu, Xiaohan Li, Ming Chen, Xiaoqiang Liu, Yu-Shen Liu, Pengfei Wan
TL;DR
Existing avatar systems often condition on local acoustic or visual cues without adequately grounding communicative intent, limiting semantic coordination and expressiveness. Kling-Avatar addresses this with an MLLM Director and a blueprint-guided, parallel cascaded generation pipeline. On a 375-sample benchmark, it reports stronger performance across synchronization, expressiveness, instruction following, identity preservation, and generalization, while supporting long-duration output up to 1080p and 48 fps.
Problem
Existing avatar-generation methods condition primarily on local acoustic or visual cues rather than modeling the communicative purpose of multimodal instructions.
Method
Kling-Avatar uses an MLLM Director to create a semantic blueprint, then generates parallel first–last-frame-conditioned sub-clips for long-duration portrait animation.
Results
Kling-Avatar reports superior lip synchronization, expressiveness, instruction controllability, identity preservation, and cross-domain generalization on a 375-sample benchmark.
Takeaways & Limitations
The framework provides semantically grounded, high-fidelity avatar synthesis with coherent long-duration generation and diverse multimodal instruction control.
Takeaways & Limitations
The evaluation primarily uses subjective GSB judgments, and the authors plan to add objective metrics in future work.
Abstract
from arXiv · showhide
Recent advances in audio-driven avatar video generation have significantly enhanced audio-visual realism. However, existing methods treat instruction conditioning merely as low-level tracking driven by acoustic or visual cues, without modeling the communicative purpose conveyed by the instructions. This limitation compromises their narrative coherence and character expressiveness. To bridge this gap, we introduce Kling-Avatar, a novel cascaded framework that unifies multimodal instruction understanding with photorealistic portrait generation. Our approach adopts a two-stage pipeline. In the first stage, we design a multimodal large language model (MLLM) director that produces a blueprint video conditioned on diverse instruction signals, thereby governing high-level semantics such as character motion and emotions. In the second stage, guided by blueprint keyframes, we generate multiple sub-clips in parallel using a first-last frame strategy. This global-to-local framework preserves fine-grained details while faithfully encoding the high-level intent behind multimodal instructions. Our parallel architecture also enables fast and stable generation of long-duration videos, making it suitable for real-world applications such as digital human livestreaming and vlogging. To comprehensively evaluate our method, we construct a benchmark of 375 curated samples covering diverse instructions and challenging scenarios. Extensive experiments demonstrate that Kling-Avatar is capable of generating vivid, fluent, long-duration videos at up to 1080p and 48 fps, achieving superior performance in lip synchronization accuracy, emotion and dynamic expressiveness, instruction controllability, identity preservation, and cross-domain generalization. These results establish Kling-Avatar as a new benchmark for semantically grounded, high-fidelity audio-driven avatar synthesis.
1 Introduction
Kling-Avatar addresses the limited semantic grounding of avatar instruction conditioning with an MLLM-guided cascaded framework. It combines global storyline planning, local refinement, curated evaluation data, and strong reported performance across synchronization, expressiveness, instruction following, and generalization.
- Motivation: Existing methods largely track acoustic or visual cues without modeling communicative intent, limiting semantic coordination and expressive control.The motivation identifies realism, fine-grained controllability, and reliable synchronization as core challenges.
- Framework: Kling-Avatar unifies multimodal instruction understanding with long-duration portrait animation through an MLLM Director and cascaded generation.The Director grounds instructions into semantic plans, while the cascade supports coherent and expressive long-duration videos.
- Evaluation: The authors construct a 375-sample benchmark with diverse instructions and challenging scenarios for comprehensive evaluation.The benchmark includes varied image categories, languages, speech rates, emotions, and dynamics.
- Contribution: The MLLM Director converts multimodal instructions into unified global plans that capture scene layout, camera positioning, character motion, emotions, and atmosphere.This shifts portrait generation from low-level cue tracking toward semantic and intent understanding.
- Contribution: The two-stage cascade first establishes high-level semantic guidance and then refines local dynamics, enabling coherent and expressive long-duration video generation.Parallel sub-clip generation follows the blueprint produced from the global plan.
- Results: Kling-Avatar reports state-of-the-art coherent and vivid portrait animation with precise lip synchronization, rich facial expressions, accurate multimodal instruction response, and strong generalization.The reported strengths span synchronization, expressiveness, controllability, and cross-scenario performance.
2 Method
Kling-Avatar grounds image, audio, and text inputs in a shared semantic plan, then uses blueprint-guided first–last-frame generation to synthesize detailed long videos. Its method combines parallel cascaded generation with quality-filtered data, lip-alignment training, identity-preserving inference, and benchmark evaluation against existing systems.
- Method overview: Kling-Avatar converts image, audio, and text inputs into fluent portrait animations with precise lip synchronization, instruction following, and long-term extrapolation.The method section frames these capabilities as the system’s generation objectives.
- 2.1 Grounding Multimodal Instructions with MLLMs: Per-modality conditioning can create semantic conflicts because existing methods rely on local cues and shallow multimodal fusion.The paper gives angry speech with unconstrained text as an example where emotion may be weakened.
- 2.1 Grounding Multimodal Instructions with MLLMs: The MLLM Director combines audio captions, image captions, and user prompts into a coherent storyline that provides global control signals.Qwen2.5-Omni supplies audio transcription and emotion, while Qwen2.5-VL supplies image descriptions.
- 2.2 Cascaded Generation Framework: The cascade generates a semantic blueprint first, extracts expressive identity-consistent keyframes, and uses them as first–last-frame conditions for local sub-clip refinement.The blueprint supplies global structure while sub-clips refine dynamics and visual details.
- 2.2 Cascaded Generation Framework: Independent parallel clip generation enables arbitrarily long videos with nearly the runtime of a single clip as anchor count increases.The framework is designed for long-duration applications including digital human podcasting, public speaking, and online education.
- 2.3 Data Construction: The training data pipeline uses expert-model filtering and manual curation to retain high-quality portrait videos, emphasizing data quality over indiscriminate scale.Filtering targets lip clarity and temporal continuity among other quality dimensions.
- 2.3 Data Construction: The benchmark contains 375 image–audio-prompt pairs spanning diverse human and non-human references, resolutions, languages, speeches, and songs.It is designed as a demanding testbed for multimodal instruction control and vivid, coherent portrait generation.
- 2.4 Training and Inference Strategy: Sliding-window audio injection and mouth-region weighting strengthen alignment between speech and lip movements during training.The method restricts attention toward temporally aligned audio tokens and emphasizes the mouth region in diffusion denoising.
3 Experiments
Experiments evaluate Kling-Avatar against competitive baselines using human-preference GSB judgments and visualizations of lip synchronization, multimodal control, diverse scenarios, and long-duration synthesis. The method consistently outperforms OmniHuman-1, improves over HeyGen on lip synchronization and visual quality, and preserves coherent identity and dynamics in long videos.
- Experimental Settings: The GSB protocol measures overall preference plus Lip Synchronization, Visual Quality, Control Response, and Identity Consistency.Three participants independently compare each method-baseline pair, with majority vote determining the final Good/Same/Bad label.
- Comparison with Baselines: Kling-Avatar consistently outperforms OmniHuman-1 across all reported GSB dimensions and benchmark settings.The evaluation covers the overall benchmark and English speeches, Chinese speeches, and bilingual singing, with smaller Japanese and Korean subsets included only overall.
- Comparison with Baselines: Kling-Avatar improves Lip Synchronization and Visual Quality over HeyGen while supporting arbitrary resolutions up to 1080p at 48 fps.HeyGen relies on repeated five-second action loops and fixed reference-image crops, whereas Kling-Avatar produces precise syllable-aligned lip movements across scenarios.
- Results on Diverse Scenarios: Multimodal instruction conditioning controls emotions, camera movements, lip synchronization, and motion dynamics across diverse human, cartoon, anime, and non-human scenarios.The MLLM Director integrates multimodal instruction intents into high-level plans that guide vivid and fine-grained generation.
- Long-Duration Video Synthesis: Long-duration generations maintain stable identity, coherent visual quality, and rich character dynamics over time.Every 10 seconds, sampled frames preserve identity while showing lighting changes, head movements, and hand gestures.
4 Related Work
Video generation has progressed toward high-fidelity controllable synthesis, while audio-driven avatar methods improve facial animation but remain constrained in broader motion generation. Existing approaches use diffusion-based generation and explicit facial or head representations, leaving limitations in natural upper-body animation.
- Video Generation: Video Diffusion Transformers enable high-fidelity video generation conditioned on multimodal signals including images, speech, and prompts.Prior work targets facial expression, lip synchronization, coordinated body motion, and data scaling.
- Audio-Driven Digital Human Synthesis: Explicit facial landmarks and 3D head models drive realistic facial expressions and lip movements but are typically limited to facial animation.These approaches cannot produce natural upper-body motion according to the supplied related-work passage.
5 Conclusion
Kling-Avatar unifies multimodal instruction understanding with long-duration lifelike portrait-video generation through a cascaded two-stage framework. On a 375-sample benchmark, it produces fluent videos up to 1080p and 48 fps with precise lip synchronization, strong controllability, and robust open-scenario generalization.
- Conclusion: Kling-Avatar combines an MLLM-generated semantic blueprint with parallel keyframe-guided sub-clip synthesis to preserve global intent and local detail.The benchmark contains 375 samples spanning diverse instructions and challenging scenarios.
- Conclusion: The framework delivers vivid, fluent long-duration videos up to 1080p and 48 fps with precise lip synchronization, strong controllability, and robust generalization.Human preference-based comparisons further confirm superior performance.