Source-linked AI summary
IAM: Identity-Aware Human Motion and Shape Joint Generation
Wenqi Jia, Zekun Li, Abhay Mittal, Chengcheng Tang, Chuan Guo, Lezi Wang, James Matthew Rehg, Lingling Tao, Size An
TL;DR
Existing motion generators often ignore how body morphology shapes motion dynamics, producing identity-neutral movement. IAM jointly generates motion and body shape from multimodal identity cues, improving motion realism, identity consistency, and motion quality across motion-capture and in-the-wild evaluations.
Problem
Existing motion generators commonly assume motion is independent of identity and morphology, overlooking how body proportions and mass distribution shape movement dynamics.
Method
IAM jointly models motion and body shape, conditioning generation on multimodal textual and visual identity cues.
Results
Experiments report improved motion realism, identity consistency, and motion quality, including state-of-the-art FID and strong zero-shot generalization to unseen identities.
Takeaways & Limitations
Identity-aware synthesis supports more realistic and controllable human animation by preserving body-shape precision and motion fidelity across diverse action sequences.
Takeaways & Limitations
Shape reconstruction is sensitive to loose clothing and occlusions, while absolute error increases for extreme body types outside the training distribution.
Abstract
from arXiv · showhide
Recent advances in text-driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. However, most existing approaches assume identity-neutral motion and generate movements using a canonical body representation, ignoring the strong influence of body morphology on motion dynamics. In practice, attributes such as body proportions, mass distribution, and age significantly affect how actions are performed, and neglecting this coupling often leads to physically inconsistent motions. We propose an identity-aware motion generation framework that explicitly models the relationship between body morphology and motion dynamics. Instead of relying on explicit geometric measurements, identity is represented using multimodal signals, including natural language descriptions and visual cues. We further introduce a joint motion-shape generation paradigm that simultaneously synthesizes motion sequences and body shape parameters, allowing identity cues to directly modulate motion dynamics. Extensive experiments on motion capture datasets and large-scale in-the-wild videos demonstrate improved motion realism and motion-identity consistency while maintaining high motion quality. Project page: https://vjwq.github.io/IAM
1 Introduction
Existing text-driven motion models largely treat motion as identity-independent and use canonical body representations, overlooking how morphology shapes movement. IAM addresses this by jointly generating motion and body shape from multimodal identity cues, improving physical consistency and reported quality.
- Introduction: Body attributes such as height, limb proportions, and mass distribution directly modulate movement, but existing approaches commonly assume universal, identity-independent motion.Individuals following the same jogging instruction can exhibit different stride lengths and joint trajectories.
- Introduction: Canonical skeletons and post-hoc retargeting or rescaling treat identity as a superficial visual attribute rather than a structural prior governing motion dynamics.This decoupled design either ignores identity during synthesis or introduces it after generation.
- Introduction: IAM represents identity through multimodal natural-language and visual signals instead of rigid geometric measurements.Text conveys semantic attributes such as “a tall, athletic male,” while images provide fine-grained cues about proportions and mass distribution.
- Introduction: IAM jointly models motion sequences and body parameters so identity cues directly inform generation and synthesized trajectories remain consistent with the character’s physical frame.The framework learns the joint distribution of motion and body morphology rather than applying identity as a post-hoc constraint.
- Introduction: IAM is model-agnostic across discrete and continuous backbones and significantly improves FID and β Dist. on diffusion architectures, achieving state-of-the-art results.The reported gains target motion quality through FID and identity consistency through β Dist.
2 Related Works
Text-to-motion methods generate human motion from natural-language descriptions using diverse generative paradigms, but most rely on canonical body representations that ignore identity and morphology. Recent approaches incorporate body shape or identity, yet remain limited by numerical conditioning, motion-retargeting focus, or the lack of joint text-based motion-and-shape synthesis.
- Text-to-Motion Synthesis: Text-to-motion synthesis generates human motion sequences from natural-language descriptions, using high-dimensional pose representations and diffusion- or transformer-based generative paradigms.Common pose formats include 263-dimensional and 272-dimensional parameterizations, while diffusion methods use iterative denoising for motion quality and diversity.
- Identity and Morphology: Most text-to-motion methods assume a fixed canonical body representation, ignoring how body proportions, age, and gender influence motion dynamics.This limits the realism and identity consistency of synthesized motions.
- Identity- and Shape-Aware Generation: Recent methods incorporate identity or body shape, but Shape My Moves depends on accurate numerical conditioning and HUMOS focuses on motion retargeting without joint text-based motion-and-shape synthesis.These limitations affect flexibility, robustness in identity transfer, and the ability to synthesize motion and shape together from text.
3 Method
IAM formulates identity-consistent motion synthesis as multimodally conditioned joint generation of motion and body shape. It integrates textual and visual identity cues with diffusion- or VQ-based generative paradigms to couple morphology and motion dynamics.
- Task Formulation: The task generates a motion sequence M from a motion prompt Tm and identity condition Ci = {Ti, Ii}, combining semantic identity text with an optional visual prior.Each motion frame xt ∈ R272 represents pose and motion features.
- Multimodal Identity Representation: Identity is represented through physique-focused text descriptions and image-based structural priors capturing fine-grained cues such as limb proportions and torso-to-leg ratios.HumanML3D uses SMPL-derived references, while IdentityMotion uses multimodal annotations and representative video keyframes.
- Multimodal Conditioning: A frozen text encoder and projected image encoder produce textual and visual embeddings, which are concatenated as distinct tokens into a unified conditional representation C.The representation is C = [Etxt; Eimg] ∈ R(L+1)×d, with joint dropout supporting classifier-free guidance.
- Joint Motion–Shape Generation: The diffusion paradigm estimates p(M, β|Ci) by concatenating 10-dimensional shape parameters with each 272-dimensional motion representation, forming a 282-dimensional joint state.A unified denoising objective couples temporal dynamics with static morphology and promotes identity consistency across the generated sequence.
- Joint Motion–Shape Generation: The VQ paradigm adds a shape-regression head to a generative Transformer, jointly predicting discrete motion tokens and continuous shape parameters while keeping the RVQ frozen.Its multi-task loss combines token cross-entropy with shape MSE, using γ = 0.1 to balance the objectives.
4 Experiments
The experiments evaluate motion quality, identity–shape consistency, and generalization to unseen identities using quantitative and qualitative analyses. Results show that diffusion-based joint generation improves motion quality, shape accuracy, zero-shot identity generalization, and controllable identity–motion synthesis.
- Evaluation Goals: Experiments assess motion generation quality, identity–shape consistency, and generalization to unseen identities through quantitative and qualitative evaluations.Evaluation covers motion quality, text alignment, body-shape reconstruction accuracy, and unseen-identity generalization.
- Datasets: HumanML3D contains 14,616 motions and 44,970 text descriptions, augmented with SMPL shape parameters covering 449 identities across genders and body types.The augmented dataset includes 263 males, 186 females, 116 slim, 269 average, and 64 heavyset identities.
- Metrics: The evaluation reports FID, R-Precision, MM-D, Diversity, body-measurement errors, and SMPL/SMPL-X β reconstruction distance.These metrics measure motion quality, text alignment, motion variation, and body-shape accuracy in parameter and geometry spaces.
- HumanML3D Results: The dual-conditioned diffusion model achieves an FID of 7.371 and the lowest β Dist. of 0.647 on HumanML3D.It outperforms the VQ-based baseline in both motion quality and body-shape accuracy.
- Zero-shot Generalization: On IdentityMotion, the dual-conditioned model achieves the lowest FID score of 23.174 and a β Dist. of 1.279 on unseen identities.The results indicate that multimodal identity descriptors map to the underlying body-shape space rather than merely memorizing training identities.
- Qualitative and Controllable Generation: Qualitative evaluations show that diffusion-based generation preserves target body shape while maintaining motion-prompt consistency and adapting body proportions to independently specified identities.On unseen identities, Shape My Moves frequently fails to follow motion prompts; randomly composed identity–motion prompts demonstrate controllable generation.
5 Limitations
The framework remains sensitive to visual ambiguity during shape reconstruction and performs less accurately on extreme body types outside the training distribution. Future work should improve identity encoding robustness.
- Visual ambiguity: Shape reconstruction is sensitive to loose clothing and occlusions in reference images, which can introduce noise into predicted parameters.These visual conditions affect the reliability of the reconstructed body shape.
- Out-of-distribution bodies: Zero-shot evaluations show high motion consistency but increased absolute error for extreme body types outside the training distribution.Examples include exceptional height or mass.
- Future directions: Future work could explore more robust identity encoders to address these limitations.The passage identifies robust identity encoders as one direction for improvement.
6 Conclusion
IAM presents a diffusion-based framework for identity-aware human motion generation, using multimodal identity descriptors to improve body-shape precision and motion fidelity. Experiments show state-of-the-art FID, strong zero-shot generalization, and controllable identity-aware synthesis.
- Framework: IAM is a diffusion-based framework for identity-aware human motion generation evaluated on HumanML3D and IdentityMotion.The model is conditioned on multimodal identity descriptors.
- Results: The model significantly outperforms baselines in preserving body shape precision and motion fidelity.These improvements were demonstrated through extensive experiments.
- Results: IAM achieves state-of-the-art FID on HumanML3D and strong zero-shot generalization to unseen identities.
- Identity-aware synthesis: Qualitative comparisons show that IAM disentangles identity from motion and enables fine-grained control over body proportions across diverse action sequences.The results underscore the importance of explicitly modeling body morphology in motion generation.
A Video Demonstration
The supplementary video visualizes the paper’s results, including animated versions of all main-paper figures, to facilitate evaluation of execution and identity-aware motion dynamics.
- Video contents: The supplementary video includes animated results for all figures in the main paper.It is recommended for evaluating the method’s execution and identity-aware motion dynamics.
- User study interface: The user study presents anonymous video pairs, an input prompt, and a frontal mesh reference for judging motion, body shape, and overall motion–shape realism.Participants select which video better matches each criterion.
B Human Perception Study
A perception study compared the proposed method with Shape My Move using side-by-side judgments on motion and identity-related criteria. The proposed method was significantly preferred across all criteria, including identity-motion synergy and physical plausibility.
- Study Protocol: The study collected 25 valid responses, with each participant evaluating 10 trials from 30 randomly sampled HumanML3D test prompt-video pairs.The model was trained exclusively on HumanML3D for a fair baseline comparison.
- Study Protocol: Participants performed side-by-side comparisons using three criteria, including Motion Plausibility and Realism.Realism assessed the physical synergy between motion and shape, and the protocol included a “Cannot judge” option to minimize bias.
- Results: The proposed method was significantly preferred across all criteria, with p < 0.05 in every comparison.The results indicate superior identity-motion synergy and more physically plausible coupling between body builds and action dynamics than the baseline.
C IdentityMotion Annotation Prompt
The annotation pipeline uses Gemini 2.5 Pro to convert human-motion videos into structured motion descriptions and generation prompts, then uses Llama 3.2 to anonymize identity-related descriptors.
- Annotation process: Gemini 2.5 Pro analyzes each human-motion video and produces concise prompts for text-to-motion generation.The annotation task begins with detailed, accurate, and unambiguous motion analysis.
- Identity anonymization: Llama 3.2 neutralizes identity-related descriptors from the initial Gemini-generated annotations.This anonymization step follows the initial annotation process.
- Annotation process: The structured analysis records body-part involvement, action sequence, temporal progression, body description, and classification fields.Classification includes age, gender, action type, and scene.
- Clarity rules: Descriptions must specify movement details and express directional references from the performer’s perspective.Stationary people should be described explicitly as holding a posture rather than performing movement.
- Output format: The output must be a valid JSON object containing five distinct text-only motion prompts in the motion_prompt field.No text should appear outside the JSON object, and prompts should not include labels such as “Base Prompt:” or “Styled Prompt:”.
Llama Neutralization Prompt
The Llama neutralization prompt rewrites motion descriptions to remove identity, role, scene, and body-build information while preserving motion details. It enforces gender-neutral language and returns only the rewritten descriptions in the input list format.
- The task removes identity-related, role-related, scene-related, and body-build information while keeping all motion details intact.
- Gendered pronouns must be replaced with neutral forms, and identity or gender terms must become neutral human references with grammatical consistency.Examples include replacing “he” or “she” with “they,” and “a man” with “a person.”
- The output must preserve the input list format using a neutralized_prompt list and must not include commentary or code.