Source-linked AI summary
Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
Vida Adeli, Soroush Mehraban, Jacob Rommann, Harrison Sanborn, Cole Clifford, Babak Taati
TL;DR
Co-speech gesture generation must align with speech while respecting posture and surrounding objects, gaps left by scene-agnostic prior models. Puppeteer addresses this with causal latent autoregressive diffusion conditioned on speech, history, posture, and object geometry; experiments report improved realism, synchronization, and diversity, supported by SCENEGES and new metrics.
Problem
Existing gesture models emphasize audio–gesture alignment but do not explicitly model posture constraints or surrounding objects, limiting object-grounded co-speech gesture generation.
Method
Puppeteer decomposes gestures into primitives, encodes them as causally ordered latent tokens, and performs conditional diffusion using speech, motion history, posture, and object geometry.
Results
Puppeteer generates more realistic, synchronized, and diverse gestures than prior methods while enabling object-grounded gesture synthesis.
Takeaways & Limitations
SCENEGES and Puppeteer provide a framework and dataset for modeling the coupling between communicative gestures, posture, and surrounding objects.
Takeaways & Limitations
The method focuses on stationary conversational gestures and removes root trajectory, while minor physical penetrations may still occur in a small fraction of frames.
Abstract
from arXiv · showhide
Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.
1. Introduction
PUPPETEER addresses the need for co-speech gestures that align with speech while remaining temporally coherent, posture-aware, and physically consistent with surrounding objects. It combines causal latent diffusion, temporal control, posture-aware training, SCENEGES, and task-specific evaluation metrics.
- Natural embodied communication requires gestures that are speech-synchronized, temporally coherent, and physically consistent with nearby objects.
- Existing speech-driven models are scene-agnostic and struggle to maintain posture consistency, object awareness, long-horizon coherence, and precise speech alignment.Prior token-based methods can introduce quantization errors, while motion-space diffusion requires denoising in a high-dimensional space.
- PUPPETEER models posture-aware, object-grounded co-speech gestures through conditional diffusion in a causal latent primitive space.The framework uses temporally ordered latent tokens and conditions generation on speech, motion history, posture, and scene information.
- The framework introduces temporal cross-attention masking for precise audio and word-level text alignment, posture-aware training, SCENEGES, and new evaluation metrics.
2. Related Work
Related work spans speech-driven gesture synthesis, physically grounded scene-aware motion generation, and causal or streaming motion generation. Existing approaches generally address these dimensions separately or target different motion settings.
- Speech-driven Gesture Generation: Speech-driven gesture models integrate audio, text, emotion, speaker identity, rhythmic structure, semantic content, and conversational context in different ways.
- Speech-driven Gesture Generation: Diffusion-based gesture methods explore stochastic denoising frameworks for gesture generation and controllability.
- Motion Generation: Scene-aware motion generation models physically consistent human–scene interactions, but primarily target locomotion or task-oriented motions rather than communicative gestures.
- Streaming motion: Recent work also investigates causal and streaming motion generation.
3. PUPPETEER Architecture
PUPPETEER generates long-horizon gestures by encoding motion primitives into causally ordered continuous latents and applying autoregressive conditional diffusion. Posture, speech, temporal masks, and object fusion provide control over physical and communicative consistency.
- 3.1. Stage1: Causal Latent Gesture Primitive Space: The causal continuous latent space avoids high-dimensional motion-space diffusion and discrete VQ tokenization artifacts while preserving temporal control.
- 3.1. Stage1: Causal Latent Gesture Primitive Space: A CausalVAE compresses gesture primitives into compact continuous latent tokens with explicit temporal ordering and causal receptive fields.Each token depends only on past and current frames within its receptive field, supporting autoregressive modeling.
- 3.3. Stage 3: Gated Object Fusion: Figure 2 presents a three-stage pipeline: causal latent encoding, masked autoregressive diffusion, and gated object fusion using BPS scene representations.
- 3.2. Stage 2: Autoregressive Gesture Diffusion: Autoregressive latent diffusion conditions each gesture primitive on motion history, audio and text features, and a fixed initial posture reference.Long sequences are represented as ordered latent primitives, with the preceding motion window and speech features supplied for each primitive.
- 3.2. Stage 2: Autoregressive Gesture Diffusion: Posture conditioning is injected into future latent tokens through gated cross-attention, while speech conditioning uses a temporally controlled cross-attention mask.
- 3.2. Stage 2: Autoregressive Gesture Diffusion: Audio-window attention and text-span masking improve gesture–speech synchronization, with their combination producing the best overall ablation result.The audio window localizes conditioning around aligned positions, while text-span masking provides word-level control.
4. SCENEGES Dataset
SCENEGES is a synthetic dataset for scene-aware co-speech gesture generation, pairing 3D gesture motions with diverse chair and table assets. Its construction includes scene layouts, scripted dialogue and gestures, synthesized videos, and refined 3D motion.
- SCENEGES pairs diverse chair and table assets with 3D SMPL-X gesture motions for scene-aware co-speech gesture generation.
- The dataset pipeline curates object assets, builds simple scenes, generates scenario-driven dialogue and gesture scripts, synthesizes videos, and recovers physically plausible 3D motion.
3. SCENARIO + SCRIPT
SCENEGES scenarios combine varied conversational settings, pose labels, object layouts, and scene-grounded gesture suggestions before motion synthesis.
- Scenario design: Conversation scenarios specify setting, roles, vibe, and one of four pose labels.
- Scenario design: Dialogues are divided into 8-second chunks annotated with emotion and tone tags.
- Scene construction: The scene assets include chairs, tables, countertops, and diverse geometries and topologies curated from HSSD.
- Gesture scripting: Gesture suggestions are grounded in scene context, including resting hands on tables, leaning on armrests, and avoiding obstacles.
- Initialization: Rendered frontal starting frames vary human pose and appearance, including male and female characters.
- Scene construction: The object collection contains 106 chairs and 47 tables.
4. VIDEO SYNTHESIS
The video-synthesis pipeline combines scene layouts, generated dialogue and gesture scripts, and rendered motion recovery to construct SCENEGES object-grounded sequences.
- Dataset output: SCENEGES contains 26 interaction scenarios instantiated across 153 object assets, averaging 8 seconds per sequence for 28 minutes of motion.
- Video synthesis pipeline: The pipeline arranges objects into scenes, generates scenarios and gesture scripts with Gemini, synthesizes videos with Veo 3, and recovers refined SMPL-X motion.
- Motion refinement: The resulting motions are refined to produce collision-free, object-grounded co-speech gestures.
- Validation: SCENEGES is validated through a user study and kinematic comparison with Embody3D, matching real gesture statistics and being perceptually difficult to distinguish from real motion.
5. Experiments
Experiments evaluate Puppeteer across speech alignment, realism, diversity, posture, and object awareness, including quantitative comparisons, ablations, and qualitative analyses.
- Quantitative evaluation: Puppeteer achieves the best FGD, BC, and Div on the BEAT2 Speaker2 test set while remaining competitive in ∆BC.
- Object awareness: LoM produces 2.8× higher MeanPen, 2.0× higher MaxPen, and 8.4× higher LL1 than the object-aware model in the structured-scene proxy comparison.
- Object-awareness ablation: Adding the object module reduces unrealistic body–object interactions, while Lcollision substantially reduces MeanPen and MaxPen.
- CausalVAE ablation: CausalVAE outperforms standard VAE in reconstruction and generation, including lower FGD, higher BC, lower ∆BC, and higher Div.
- Conditioning ablation: Combining temporal audio-window conditioning with text span masking gives the best overall masking result, increasing diversity and reducing the synchronization gap.
- Qualitative evaluation: Qualitatively, Puppeteer better matches speech timing and emphasis than LOM and GestureLSM, while object awareness adapts gestures to surrounding furniture.
6. Conclusion
The paper presents Puppeteer as an object-grounded co-speech gesture framework with causal temporal control, supported by SCENEGES and new evaluation metrics.
- Puppeteer explicitly models the coupling between communicative gestures, posture, and surrounding objects.
- Its causal latent space enables autoregressive generation with explicit temporal control over speech–gesture alignment.
- Experiments show improved realism, synchronization, and diversity over prior methods.
Supplementary Material
The supplementary material details the motion, posture, and scene representations, training constraints, dataset construction, and perceptual validation used by Puppeteer.
- Representation: Each gesture frame combines global and joint rotations, joint positions, and frame-to-frame displacements into a 666-dimensional representation.Gesture primitives group these frames into fixed-length sequences.
- Motion Constraints: Standing-mode generation fixes the lower body through inverse kinematics while preserving full-body gesture generation and posture awareness.This prevents unintended trajectory shifts while allowing natural upper-body motion.
- Scene Representation: BPS encodes person–object spatial relationships using 1024 distances within a forward- and laterally-biased ellipsoidal support region.The support is shifted toward the front and upper body because hand movements are more likely to cause collisions there.
- CausalVAE Training: Training enforces reconstruction, latent regularization, velocity, and SMPL-X geometric consistency between reconstructed and represented motion.Velocity consistency compares predicted displacements and relative rotations with values recomputed from reconstructed joints and rotations.
- Dataset: SCENEGES contains 26 interaction scenarios across 153 chair and table assets, totaling 28 minutes of object-grounded co-speech motion.The scenes vary conversational contexts, object topologies, and human–furniture contact configurations.
- Perceptual Evaluation: SCENEGES motion was difficult to distinguish from real captured motion and received a naturalness score of 4.28 versus 4.09 for Embody3D.In a 23-participant study, balanced accuracy was 51.41%, near random chance; motion statistics also closely matched real conversational motion distributions.
G. More Experiments and Ablations
Ablations identify balanced latent capacity, moderate guidance, intensity weighting, and additional sampling as important for generation quality, synchronization, and diversity. Further experiments show posture and object conditioning support robust, adaptable gestures, while the method remains scoped to stationary conversational motion and retains occasional physical penetrations.
- G.3. Effect of CFG Scale: A CFG scale of 2.5 provides the best trade-off between realism, diversity, and speech–gesture synchronization.Higher guidance strengths overly constrain generation, increasing FGD and ∆BC while reducing diversity.
- G.4. Effect of Sampling Steps: Increasing sampling steps from 1 to 10 substantially reduces FGD and increases diversity by allowing further refinement of latent motion.Even one sampling step produces competitive results, but very few steps limit motion refinement and speech–gesture alignment.
- G.5. Effect of coefficient λI: λI = 3 yields the highest BC and diversity together with the lowest ∆BC for the single-speaker setting and the best synchronization on the full dataset.Intensity-based reweighting emphasizes high-intensity motion segments; larger values provide no consistent improvement and slightly degrade some metrics.
- Object and Posture Generalization: PUPPETEER preserves spatial precision on unseen furniture, achieving MeanPen of 3.280 × 10−4 and LL1 error of 0.86.The model also generates diverse, speech-aligned gestures across standing and sitting scenarios, and posture conditioning prevents drift toward a neutral standing pose.
- Temporal Editing: Causal latent modeling supports gesture in-betweening and completion by producing smooth, speech-matched transitions between surrounding ground-truth frames.The regenerated segment can replace less expressive motion with a more appropriate gesture.