Source-linked AI summary
The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
Jing Lin, Ruisi Wang, Junzhe Lu, Ziqi Huang, Guorui Song, Ailing Zeng, Xian Liu, Chen Wei, Wanqi Yin, Qingping Sun, Zhongang Cai, Lei Yang, Ziwei Liu
TL;DR
Text-to-motion models struggle to generalize beyond standard benchmarks, motivating transfer of video-generation knowledge into motion generation. The paper introduces a coordinated dataset, gated flow-based model, distilled variant, and benchmark, and reports state-of-the-art performance across action accuracy and generalization while noting important scope limitations.
Problem
Existing text-to-motion models have limited generalization to diverse and long-tail instructions despite stronger generalization in adjacent video-generation models.
Method
The framework combines the ViMoGen-228K dataset, a gated flow-based diffusion transformer unifying MoCap and ViGen priors, ViMoGen-light, and the MBench evaluation suite.
Results
The framework achieves state-of-the-art performance in action accuracy and generalization, while ViMoGen-light reduces computational overhead by eliminating video-generation dependencies.
Takeaways & Limitations
The paper shows that coordinated data, modeling, and evaluation innovations can improve generalizable text-driven human motion generation.
Takeaways & Limitations
The method supports only single-person motion generation, and complex high-dynamic motions may depend heavily on potentially distorted video priors.
Abstract
from arXiv · showhide
Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing text-to-motion models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generative fields, most notably video generation (ViGen), have demonstrated remarkable generalization in modeling human behaviors, highlighting transferable insights that MoGen can leverage. Motivated by this observation, we present a comprehensive framework that systematically transfers knowledge from ViGen to MoGen across three key pillars: data, modeling, and evaluation. First, we introduce ViMoGen-228K, a large-scale dataset comprising 228,000 high-quality motion samples that integrates high-fidelity optical MoCap data with semantically annotated motions from web videos and synthesized samples generated by state-of-the-art ViGen models. The dataset includes both text-motion pairs and text-video-motion triplets, substantially expanding semantic diversity. Second, we propose ViMoGen, a flow-matching-based diffusion transformer that unifies priors from MoCap data and ViGen models through gated multimodal conditioning. To enhance efficiency, we further develop ViMoGen-light, a distilled variant that eliminates video generation dependencies while preserving strong generalization. Finally, we present MBench, a hierarchical benchmark designed for fine-grained evaluation across motion quality, prompt fidelity, and generalization ability. Extensive experiments show that our framework significantly outperforms existing approaches in both automatic and human evaluations. The code, data, and benchmark will be made publicly available. Homepage: https://motrixlab.github.io/2026_iclr_vimogen.
1 INTRODUCTION
ViMoGen addresses limited semantic generalization in text-to-motion generation by coordinating a diverse dataset, gated model, and fine-grained benchmark. Its framework transfers complementary knowledge from motion capture and video generation to improve generalization.
- ViMoGen-228K combines 171.5K text–motion pairs and 56.6K text–video–motion triplets from complementary sources to expand semantic coverage.The dataset integrates optical MoCap, in-the-wild video, and synthetic video-derived motions.
- ViMoGen uses gated fusion and dual branches to transfer motion priors from MoCap, in-the-wild videos, and synthetic video data.The Text-to-Motion branch uses MoCap priors, while the Motion-to-Motion branch uses ViGen-derived tokens.
- ViMoGen and ViMoGen-light substantially improve generalization over prior approaches on challenging prompts including martial arts, dynamic sports, and multi-step behaviors.ViMoGen-light bypasses expensive video-generation inference while retaining strong generalization.
- MBench evaluates motion quality, motion-condition consistency, and generalization with a curated open-world vocabulary.It addresses the limited granularity and simple prompt coverage of existing evaluation protocols.
2 METHOD
ViMoGen is a flow-based diffusion transformer that combines precise motion priors with the broader semantics of video generation. Its gated branches adapt conditioning during generation, while ViMoGen-light distills the video prior to remove inference overhead.
- ViMoGen unifies high-quality MoCap knowledge and broad ViGen semantic knowledge in a flow-based diffusion transformer for generalizable text-driven motion generation.The approach targets the tension between MoCap fidelity and video-model semantic diversity.
- 2.1 PRELIMINARIES: Flow matching trains the model to predict the velocity from noisy motion toward clean motion under textual conditioning.The interpolation is xt = (1 − t)ϵ + tx0, with velocity vt = x0 − ϵ.
- 2.2 UNIFYING VIDEO AND MOTION GENERATION MODEL PRIOR: Gated diffusion blocks fuse text and video motion tokens with noisy motion inputs through mutually exclusive Text-to-Motion and Motion-to-Motion branches.The branches share approximately 66% of DiT parameters and differ at cross-attention.
- 2.2 UNIFYING VIDEO AND MOTION GENERATION MODEL PRIOR: Adaptive branch selection uses semantic alignment to activate Motion-to-Motion for video-motion refinement or Text-to-Motion otherwise.Training regulates branch usage by dataset characteristics and simulates video tokens with compound noise.
- 2.3 DISTILLING MOTION PRIOR FROM VIDEO GENERATION MODEL: ViMoGen-light distills the video prior into a lightweight model that exclusively uses Text-to-Motion and eliminates video-generation dependencies.This addresses the computational overhead of video generation during ViMoGen inference.
3 VIMOGEN-228K DATASET
ViMoGen-228K balances motion fidelity and semantic diversity by combining filtered optical MoCap, in-the-wild videos, and strategically generated synthetic videos. Its construction uses standardization, quality filtering, and complementary annotation and synthesis pipelines.
- ViMoGen-228K balances 172K high-fidelity text-motion pairs with 56K diverse text-video-motion triplets from complementary data sources.The dataset aggregates optical MoCap, filtered web videos, and synthetic videos to combine quality with semantic diversity.
- The dataset unifies 30 optical MoCap datasets, filters web videos for motion fidelity, and generates synthetic videos to expand semantic coverage.These strategies address the limited scale and semantic diversity of optical MoCap and the quality compromises of web-video datasets.
- Table 1 compares ViMoGen-228K with existing datasets and identifies unified, aggressively filtered, and strategically generated data sources.Its annotations distinguish the 29-dataset union, filtering from 10M clips, and synthetic semantic-coverage expansion.
- The optical MoCap pipeline standardizes data to SMPL-X and 20 fps, segments sequences into five-second clips, and retains 172K clips after quality filtering.Filtering removes T-poses, short sequences, jitter, and low-dynamics motions; 17 of 30 candidate datasets are selected.
- In-the-wild video data broadens semantic coverage, while synthetic video generation targets controllable human motions that are difficult to find in real-world videos.The source video pool is quality-filtered, and synthetic prompts are designed for easier vision-based motion capture.
4 MBENCH
MBench is a hierarchical benchmark that evaluates motion generation across nine dimensions spanning generalization, prompt consistency, and motion quality. It uses diverse prompts and fine-grained metrics to assess capabilities beyond memorization.
- 4.1 EVALUATION DIMENSIONS: MBench decomposes evaluation into nine dimensions across Motion Generalization, Motion-Condition Consistency, and Motion Quality.The benchmark is designed for granular, multifaceted assessment and includes human preference annotations to validate alignment with perception.
- 4.1 EVALUATION DIMENSIONS: MBench uses balanced data and substantially different prompt designs from HumanML3D to evaluate motion quality, prompt-following, and generalization systematically.The benchmark overview contrasts its prompt distribution and designs with HumanML3D.
- 4.1 EVALUATION DIMENSIONS: Motion Generalizability evaluates rare or unseen actions to measure whether models produce diverse, contextually coherent movements beyond memorized patterns.The dimension targets actions absent from commonly used MoGen and ViGen datasets.
- 4.1 EVALUATION DIMENSIONS: Motion-Condition Consistency measures semantic fidelity between generated motion and text by selecting the correct label from one ground-truth and nine distractors.The evaluation uses a VLM-based description and classification pipeline to reduce reliance on biased action-recognition metrics.
- 4.1 EVALUATION DIMENSIONS: Motion Quality covers temporal consistency and frame-wise properties, including jitter, foot contact, and dynamic degree.Temporal quality assesses cross-frame stability and physical plausibility, while dynamic degree measures overall motion intensity through joint velocities.
5 EXPERIMENTS
ViMoGen outperforms prior methods on semantic consistency and generalization, while ViMoGen-light transfers much of this benefit without video generation at inference. Ablations link these gains to adaptive branch selection, diverse data, and descriptive prompts, with a trade-off in some motion-quality metrics.
- 5.1 COMPARISON WITH SOTA METHODS: ViMoGen significantly outperforms all baselines on Motion Condition Consistency and Generalizability, while ViMoGen-light matches the strongest baseline in Generalization Score without video generation at inference.The full model’s semantic advantage is attributed to richer textual guidance, while the efficient variant preserves transferred generalization knowledge.
- 5.1 COMPARISON WITH SOTA METHODS: The generalization gains introduce a motion-quality trade-off, with diverse video-sourced actions producing higher Jitter Degree but lower Dynamic Degree.The authors relate this pattern to complex actions with less global movement than locomotion-heavy datasets.
- 5.1 COMPARISON WITH SOTA METHODS: For out-of-domain prompts such as “body surfing,” ViMoGen produces plausible motions, whereas T2M-GPT generates implausible or generic motions.The qualitative comparison attributes ViMoGen’s behavior to semantic knowledge from video-generation priors.
- 5.2 ABLATION STUDY: Adaptive branch selection outperforms single-branch strategies in generalization and accuracy.Video-derived motion is semantically relevant but noisy, MoCap-only T2M has strong quality but limited generalization, and the adaptive design combines their complementary strengths.
- 5.2 ABLATION STUDY: Adding diverse data sources progressively improves action accuracy and generalization, with 14k synthetic video clips raising Generalization Score from 0.50 to 0.55.The data ablation starts from a ViMoGen-light model trained solely on HumanML3D.
- 5.2 ABLATION STUDY: Training on descriptive video-style text and testing on concise motion-style descriptions yields the best performance.The authors interpret rich descriptions as data augmentation that improves robustness and alignment with pretrained text encoders.
6 CONCLUSION AND DISCUSSION
The paper addresses generalizable 3D human motion generation through coordinated advances in data, modeling, and evaluation. ViMoGen achieves state-of-the-art action accuracy and generalization, while ViMoGen-light reduces computational overhead and MBench enables fine-grained assessment.
- 6 CONCLUSION AND DISCUSSION: The framework combines a 228K-clip dataset, a gated diffusion model, and MBench to address generalizable 3D human motion generation.ViMoGen-228K combines high quality with broad semantic coverage, while MBench evaluates generalization, motion-condition consistency, and motion quality.
- 6 CONCLUSION AND DISCUSSION: ViMoGen achieves state-of-the-art action accuracy and generalization by unifying video-generation priors with motion-specific knowledge.The lightweight ViMoGen-light variant distills these generalization capabilities while reducing computational overhead.
- 6 CONCLUSION AND DISCUSSION: The method currently supports single-person motion generation but not multi-person interactions.This is identified as a primary limitation of the current architecture.
- 6 CONCLUSION AND DISCUSSION: For complex high-dynamic motions, distorted video priors can leave the M2M branch unable to fully correct the dynamics.The paper also reports that stronger generalization does not always produce the highest scores on specific quality metrics such as Dynamic Degree.
ETHICS STATEMENT
The study addresses ethical risks from human-subject motion data and possible misuse of motion generation. It limits data sources to publicly available research-friendly datasets and excludes personally identifiable information from internet video data.
- ETHICS STATEMENT: The study uses publicly available optical motion-capture datasets under research-friendly licenses to mitigate privacy and consent concerns.For internet video, the authors state that they use motion information only and no personally identifiable information.
- ETHICS STATEMENT: The authors acknowledge that motion generation could be misused for deceptive content creation or surveillance.They state that the model is intended solely for scientific research and beneficial applications.
REPRODUCIBILITY STATEMENT
The paper commits to reproducibility through public release plans and detailed documentation, while situating its approach against motion-generation data and model limitations.
- Code, datasets, benchmark configurations, training protocols, hyperparameters, implementation choices, and evaluation metrics are planned for public release or documentation.These materials are intended to enable independent verification and fair comparison.
- Optical MoCap datasets provide precise, physically plausible motion but remain small because collection is labor-intensive and expensive.HumanML3D contains approximately 14,000 clips, and controlled collection limits semantic coverage.
- Video-based datasets address the MoCap scale bottleneck but can remain limited in scale or concentrated in specific scenarios.Motion-X is described as an early large-scale effort focused primarily on sports and gaming.
- Existing motion-generation models are dominated by diffusion and autoregressive paradigms, while video-prior transfer methods face quality, prior-utilization, and efficiency challenges.The cited related work motivates systematic transfer from video generation to motion generation.
C.1 HUMAN PREFERENCE ANALYSIS
MBench combines human preference annotation with automatic per-dimension evaluation to test metric alignment and motion-model generalization across a curated 450-prompt suite.
- DATA PREPARATION: Five models produce 10 pairwise combinations per prompt, yielding N × 10 comparisons across N prompts.Win ratios assign 1 to the preferred model, 0 to the other, and 0.5 to each model in a tie.
- HUMAN PREFERENCE ANALYSIS: Human evaluation uses pairwise comparisons for condition consistency and generalizability, while quality dimensions receive 3-point individual ratings with randomized repeated presentation.Each video appears five times, and contradictory pairwise judgments may be revised using single-video scores.
- VALIDATING HUMAN ALIGNMENT OF MBENCH: MBench automatic evaluation win ratios show strong positive correlations with human-preference win ratios across dimensions, although Foot Floating correlates less strongly.The lower Foot Floating correlation reflects limited human sensitivity to subtle artifacts, whereas the physics-based metric distinguishes them precisely.
- PROMPT SUITE: The benchmark contains 450 prompts spanning temporal quality, frame-wise quality, motion-condition consistency, and motion generalizability.The suite includes 150 temporal-quality prompts and 100 prompts for each of the other three listed dimensions.
- PROMPT SUITE: Prompts target open-world semantic generalization through underrepresented everyday actions, linguistic nuances, semantic clustering, automated filtering, and expert curation.The construction pipeline uses video-derived descriptions, multiple embedding encoders, K-means clustering into 1000 clusters, and final expert selection.
- DATA PREPARATION: The dataset curation process excluded several datasets whose high-speed or stylistic motions and weak descriptions degraded MBench performance and text-motion consistency.Sixteen additional datasets with positive impact were selected for the final training set.
D.2.2 IN-THE-WILD VIDEO DATA
The in-the-wild video pipeline scales motion collection through large video sources, multi-stage quality filtering, synthetic augmentation, motion extraction, and structured text annotation.
- IN-THE-WILD VIDEO DATA: An initial pool of approximately 60 million video clips is reduced through quality filtering and selection of clips with at least 80% visible human skeletons.Filtering removes low-resolution, poorly lit, blurry, and heavily occluded clips before more granular selection.
- IN-THE-WILD VIDEO DATA: Synthetic video augmentation uses a long-tail action vocabulary and Wan2.1 to generate 81-frame clips at 16 fps under visual-MoCap-friendly conditions.Generated motions pass through the standard filtering pipeline and are interpolated from 16 fps to 20 fps.
- IN-THE-WILD VIDEO DATA: Human motion is extracted from in-the-wild and synthetic videos by tracking people with YOLOV8 and recovering SMPLX motion with CameraHMR.The pipeline also canonicalizes global orientation to face the y+ axis initially.
- IN-THE-WILD VIDEO DATA: Gemini 2.0 Flash generates structured annotations covering whole-sequence motion, style, trajectory, temporal actions, and fine-grained frame intervals.The system prompt specifies seven structured annotation components, including descriptions every two and four frames and key motion statuses.
- IN-THE-WILD VIDEO DATA: The annotation format represents motion with narrative descriptions, qualitative styles, environmental trajectory, chronological actions, and detailed body-posture changes.The example describes a controlled drag curl through progressive lowering, squatting, and stabilization.
E VIMOGEN IMPLEMENTATION DETAILS
ViMoGen is implemented with a Wan2.1-based motion architecture, adaptive multimodal training, standard HumanML3D evaluation, and noise simulation for visual-MoCap conditions.
- IMPLEMENTATION DETAILS: ViMoGen uses a 1.3B-parameter Wan2.1 foundation model and 276-dimensional SMPL-based motion vectors.Training uses 8 H800 GPUs, AdamW, a 0.0002 learning rate, batch size 128, and FSDP.
- IMPLEMENTATION DETAILS: Adaptive training double-weights synthetic data, restricts visual-MoCap supervision to local poses, and adjusts branch probabilities according to data quality.These choices are designed to use heterogeneous sources while mitigating visual-MoCap artifacts and quality variation.
- IMPLEMENTATION DETAILS: Visual-MoCap noise is simulated with Gaussian corruption, temporal jitter, and dropout, while global translation is masked to remove trajectory-estimation domain gaps.The specified corruption probabilities and scales emulate tracking imperfections during M2M training.
- EXPERIMENTAL SETTINGS: The proposed network replaces MLD’s diffusion denoiser while retaining its pretrained motion VAE and most baseline hyperparameters for fair comparison.A supplementary HumanML3D experiment isolates the architectural contribution from the large-scale ViMoGen-228K dataset.
- EXPERIMENTAL SETTINGS: The HumanML3D evaluation uses FID, Diversity, MultiModality, R-Precision, and Multimodal Distance to measure motion quality, diversity, and text-motion consistency.The test set contains 14,616 motion sequences paired with 44,970 text annotations.
F.2 RESULTS AND ANALYSIS
ViMoGen-light establishes state-of-the-art text-motion consistency while preserving strong motion quality on HumanML3D. Its gains also transfer as an effective component within existing frameworks.
- MLD + ViMoGen-light surpasses prior methods on every reported text-alignment metric, achieving the best R-Precision and lowest Multimodal Distance.The comparison uses 20 repetitions and reports means with 95% confidence intervals.
- FID 0.114 improves substantially over the MLD baseline at 0.473 while remaining competitive with top-performing methods such as T2M-GPT.The model also maintains a strong MultiModality score.
- The results indicate that ViMoGen-light can advance existing frameworks independently, rather than benefiting only from the large-scale OmniMotion dataset.
G ADDITIONAL QUALITATIVE RESULTS
Additional qualitative results examine adaptive branch selection, prompt-style effects, and comparisons on detailed MBench prompts. Across these analyses, the visualizations emphasize robust semantic alignment and generation quality.
- Figure 11 visualizes adaptive selection between M2M and T2M branches according to the quality of motion extracted from generated videos.
- Figure 10 presents qualitative HumanML3D comparisons for complex multi-step prompts, where ViMoGen-light is described as more plausible and better aligned than prior works.
- Figure 12 compares video-style and motion-style training and testing prompts for “swaggering into the room” and “chopping wood.”
- ViMoGen and ViMoGen-light consistently adhere more faithfully to detailed MBench descriptions than prior methods in additional qualitative comparisons.