Source-linked AI summary
Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans
Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen, Xin Wang, Ye Shi, Jingya Wang
TL;DR
The paper addresses the gap between offline co-speech gesture generation and interactive digital humans that must generate synchronized motion from streaming speech under strict latency constraints. It proposes a coupled streaming speech and causal gesture framework, supported by interactive data synthesis and user-feedback-driven self-evolution, and reports better latency-quality trade-offs, synchronization, and user preference than existing baselines.
Problem
Existing co-speech gesture methods predominantly assume complete speech segments or future acoustic cues, while interactive digital humans require causal speech-synchronous generation under strict latency constraints.
Method
The framework couples streaming response speech with a causal multimodal autoregressive gesture generator, topic- and emotion-aware offline data synthesis, and a user-feedback-driven self-evolution loop.
Results
The framework achieves better latency-quality trade-offs, stronger speech-motion synchronization, and higher user preference than competitive existing baselines.
Takeaways & Limitations
The proposed design supports low-latency, speech-synchronous body-motion generation without future speech access and supports continual adaptation to user preferences.
Takeaways & Limitations
Strict online generation limits long-range future context, leaving gestures less globally structured for rapid topic shifts, dynamic prosody, long-form emotional evolution, or anticipatory gestures.
Abstract
from arXiv · showhide
Existing co-speech gesture generation methods are predominantly studied in offline settings, where gestures are synthesized from complete speech segments. However, interactive digital humans in real-world scenarios are required to generate speech-synchronous gestures online, using only currently available response audio under strict latency constraints. As a result, prior methods are unsuitable for real-time interaction, as they either rely on future speech information or incur substantial inference delay. In this paper, we formulate online co-speech gesture generation for interactive digital humans and propose a real-time interactive framework that couples a streaming speech response module with an online gesture generation module. Specifically, the gesture generator is designed as a causal multimodal autoregressive model that predicts body motion from streaming response speech and motion history, enabling low-latency and speech-aligned gesture synthesis without access to future speech. To support this setting, we further propose an offline data synthesis pipeline tailored to virtual companion scenarios, which leverages topic- and emotion-aware subject corpora to construct diverse human-agent dialogues and then generates co-speech gestures conditioned on the agent responses. Moreover, to bridge the gap between offline data construction and online deployment, we establish a self-evolving training loop by incorporating user feedback collected during online interaction into the data generation process, enabling continual adaptation to user preferences. Extensive experiments demonstrate that our framework achieves superior better latency-quality trade-off, stronger speech-motion synchronization, and higher user preference than competitive existing baselines. Project Page: https://super-star-2026.github.io/
1 Introduction
The paper targets online co-speech gesture generation for interactive digital humans, where gestures must remain speech-synchronous under strict latency constraints and incomplete future context. It combines streaming response speech, causal gesture prediction, interactive data synthesis, and user-feedback-driven self-evolution.
- Most existing co-speech gesture methods assume complete speech segments are available before generating motion, limiting their use in real-time interaction.
- Future speech dependence and long-segment processing create inference delays that undermine immediacy in live interaction.
- Interactive virtual companion scenarios require diverse topic- and emotion-conditioned responses, richer posture variation, and stronger affective expression than existing datasets provide.
- Offline interactive data synthesis and online user feedback form a self-evolving loop for continual adaptation to user preferences.
- The framework couples a streaming speech response module with a causal autoregressive gesture generator conditioned on available speech and motion history.
2 Related Work
Prior work models speech-motion correspondence across modalities and generation paradigms, but is mainly designed for offline synthesis rather than causal online interaction. The paper addresses this gap while also targeting the richer affective and posture requirements of virtual companion scenarios.
- Earlier methods generate upper-body or pose-level gestures from speech, text, or both using recurrent, variational, and adversarial models.
- Despite strong performance, most existing methods require complete speech segments and are less compatible with strict real-time causal generation.
- Multimodal interactive systems address perception, reasoning, response generation, reaction synthesis, and social modeling, but often do not design synchronization for strict online co-speech generation.
- Virtual companion applications additionally require richer posture variation and stronger affective expression, motivating interactive data synthesis and professional-actor motion capture.
3 Method
Super Star unifies streaming speech response with causal online gesture generation and an offline, virtual-companion-oriented data pipeline. The system predicts synchronized body motion from available speech and motion history while using user feedback to continually adapt training data.
- Framework Overview: The unified framework couples a streaming speech response module with an online gesture generator for low-latency interactive digital humans.The response module produces immediately playable speech, while the gesture module predicts motion incrementally from streaming speech and motion history.
- Online Co-Speech Gesture Generation: Compositional tokenization divides motion into lower body, upper body, hands, and head-related facial expressions, each represented through independently trained VQ-VAEs.Cross-modal autoregressive generation then models dependencies among body parts to preserve holistic coordination during online prediction.
- Online Co-Speech Gesture Generation: The online generator predicts current body motion autoregressively under a strict causal constraint, without accessing future speech.Causal masking is applied during training and inference so motion at time t attends only to speech available through time t.
- Offline Interactive Data Synthesis: The offline synthesis pipeline samples topic and emotion segments to prompt an LLM to generate diverse human-agent dialogues and corresponding co-speech motions.Preferred online interaction samples are later incorporated as in-context examples to guide subsequent synthesis toward user preferences.
- Offline Interactive Data Synthesis: User feedback and preferred online interactions are fed back into offline data construction, enabling continual adaptation and narrowing the gap between training and deployment.The self-evolving loop expands the training set with preference-aware interaction data and newly synthesized offline data.
4 Experiments
Experiments evaluate the framework under a strict causal online protocol using BEATv2 and the interaction-oriented JIYI dataset, with adapted offline methods as baselines. Ablations, user studies, and qualitative comparisons assess causal modeling, self-evolution, latency, synchronization, and preference.
- Experimental protocol: The evaluation masks future audio, requiring causal body-motion generation from only current and past response speech.This protocol is designed to reflect real-time interactive deployment.
- Datasets: Experiments use public BEATv2 and JIYI, which contains about 6 hours of synchronized speech, transcripts, and co-speech body motions.JIYI serves as both a cold-start training dataset and an interactive evaluation benchmark.
- Baselines: Because no standard online benchmark exists, representative offline methods are re-implemented or adapted as online baselines.The comparison follows the strict online protocol described for the proposed method.
- Ablation study: Removing the causal audio attention mask verifies that explicit cross-modal interaction is essential for continuously aligning motion with incomplete future context.The model must align motion prediction with currently available response speech during online generation.
- Ablation study: One and two rounds of self-evolution both improve the model over the base version, validating the user-feedback-driven training loop.Each round expands training data with preference-aware online interaction data and newly synthesized offline data.
- User and qualitative studies: The method receives the highest overall online user preference and produces more natural, expressive, smooth, coordinated, and temporally aligned gestures qualitatively.Compared with baselines, it better preserves gesture continuity during streaming generation.
5 Conclusion
The paper formulates causal online co-speech gesture generation for interactive digital humans and presents a unified framework for low-latency, speech-synchronous motion without future speech. Experiments on BEATv2 and JIYI report improved latency-quality trade-offs, synchronization, and user preference over competitive baselines.
- Conclusion: The framework couples streaming speech responses with a causal multimodal autoregressive gesture generator using motion history and no future speech.It also includes offline interactive data synthesis and a user-feedback-driven self-evolution loop for virtual companion scenarios.
- Conclusion: Experiments on BEATv2 and JIYI demonstrate a better latency-quality trade-off, stronger speech-motion synchronization, and higher user preference than competitive baselines.
A Details of Our JIYI Dataset
JIYI is collected with an OptiTrack motion-capture setup using cameras and markers arranged around professional actors. The dataset records synchronized user-agent dialogue speech and co-speech gestures for interactive scenarios.
- Motion capture: The OptiTrack setup uses 16 high-resolution cameras positioned around the actor from bird-view and side-view perspectives.The cameras operate at 120 FPS during capture.
- Motion capture: Professional actors perform user-agent dialogues with synchronized speech and co-speech gestures for JIYI collection.
- Motion capture: Fifty markers are placed on one person and documented from front and rear views during motion capture.Actors wear motion-capture suits with markers.
A.2 Dataset statistics
JIYI contains approximately six hours of motion data across diverse topics and emotions, with 1,570 sequences split into training, validation, and test sets.
- Dataset statistics: JIYI contains about 6 hours of motion data captured at 120 FPS and downsampled to 30 FPS for training.
- Dataset statistics: The dataset includes 1,570 sequences split into 1,256 training, 157 validation, and 157 test sequences.
- Dataset statistics: Topics span weather, travel, relationships, sports, health, games, food, professions, and daily phrases, while emotions include happiness, anger, fear, sadness, calmness, and rationality.
B Experimental Details
The implementation uses streaming Qwen-Omni-3 audio, compositional motion tokenization, and separate quantization settings for online and offline models.
- Qwen-Omni-3 produces the first audio token after approximately 234 ms, meeting the system’s real-time requirements.
- Motion representation: Motion is represented compositionally across face, hands, upper body, and lower body using separate 6D-rotation components.The four parts contain 1 face joint plus 100 expression parameters, 30 hand joints, 13 upper-body joints, and 9 lower-body joints.
- Motion tokenization: The tokenizer reconstructs motion with reconstruction, temporal smoothness, mesh, and commitment losses.
- Offline quantization: Offline models use residual VQ-VAE without causal masking to represent complex motion more finely when future speech is available.Residual quantization adds multiple layers to capture motion at different granularities, especially for finger motions.
- Offline quantization: The offline configuration uses 4 residual quantization layers and codebook size 512.
B.3 Implementation details
Implementation details describe a 12-layer online transformer, a self-evolving dataset, and lightweight preference filtering for continual adaptation.
- Model configuration: The online gesture generator uses 12 transformer decoder layers on JIYI, while the offline generator uses residual quantization with four layers and codebook size 512.
- Self-evolution: The self-evolution pipeline expands an initial dataset with preferred online interaction samples and newly synthesized offline interaction data.Its purpose is to progressively adapt the online gesture model to real deployment scenarios.
- Self-evolution: After round one, the training set contains 2,434 samples: 50% original, 47% synthesized offline, and 3% preferred online data.
- Self-evolution: After round two, the training set contains 3,852 samples: 33% original, 62% synthesized offline, and 5% preferred online data.The increased evolved-data proportion is intended to adapt to the target interaction domain while retaining original data as a quality anchor.
- Preference feedback: User satisfaction or dissatisfaction selects preferred online samples, which provide direct supervision and in-context examples for later dialogue synthesis.Synthesized data is filtered using Beat Consistency, joint-limit, foot-skating, and abnormal-motion or velocity checks.
B.6 Extra data-source ablation
The ablation separates data-source effects, while the evaluation protocol measures gesture quality, synchronization, responsiveness, and user preference in interactive settings.
- Data-source ablation: Offline synthesized data improves coverage and diversity, whereas preferred online feedback provides deployment-specific alignment.
- Data-source ablation: Sparse preferred samples alone yield only modest gains, confirming the need for synthetic amplification.
- Data-source ablation: The non-causal offline gesture generator achieves FGD/BC/Div. of 2.512/7.587/11.713 on the held-out test split.This result verifies the quality of its pseudo-labels.
- Metrics: FGD measures distributional closeness to ground-truth gestures, while L1 Diversity measures variation among generated gesture clips.
- Metrics: Beat Consistency measures rhythmic alignment between generated gesture beats and audio beats.Audio onset and upper-body velocity minima define the respective beat sets.
- User study: The user study compares paired videos from the same response speech using randomized left-right order and records preference ratios across naturalness, synchronization, responsiveness, and overall preference.
D Limitation and Future Work
The paper identifies limits in long-horizon online planning, synthesized-data fidelity, preference adaptation, and the scope of embodied channels addressed.
- Limitations: Strict causal generation limits access to long-range future context, leaving rapid shifts, dynamic prosody, and long emotional utterances less globally structured than offline outputs.
- Limitations: Synthesized dialogues and motions may retain distribution gaps from real human interaction despite improving data diversity and scenario coverage.
- Limitations: The framework does not explicitly learn a dedicated user preference model, reward model, or fine-grained adaptation policy.Consequently, it may not fully capture subtle or personalized long-term user preferences.
- Future work: The current scope centers on speech-conditioned body motion rather than facial micro-expressions, gaze, head attention, turn-taking, and broader environmental grounding.
- Future work: Future work targets online planning, richer preference-aware adaptation, and more comprehensive multimodal virtual-companion interaction.