Source-linked AI summary

Motus: A Unified Latent Action World Model

Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, Hongyan Zhao, Hanyu Liu, Zhizhong Su, Lei Ma, Hang Su, Jun Zhu

arXiv:2512.13030v2cs.CVcs.LGcs.RO

TL;DR

Embodied agents need unified modeling because existing systems isolate understanding, world modeling, and control and rely heavily on heterogeneous labeled data. Motus connects pretrained understanding, video-generation, and action experts with coordinated multimodal modeling and optical-flow latent actions. It reports improvements over state-of-the-art methods in both simulation and real-world scenarios, while identifying broader motion-prior learning from internet-scale videos as future work.

  • Problem

    Existing embodied methods isolate understanding, world modeling, and control, while heterogeneous embodiments make control signals difficult to reuse and limit integration of large-scale internet video.

  • Method

    Motus uses a Mixture-of-Transformers with understanding, video-generation, and action experts, a UniDiffuser-style scheduler, optical-flow latent actions, and staged training over a six-layer data pyramid.

  • Results

    +15~45% in simulation and +11~48% in real-world scenarios compared with existing state-of-the-art embodied models.

  • Takeaways & Limitations

    Unifying multimodal generative capabilities and shared motion priors is reported to benefit downstream robotic tasks and policy-learning generalization.

  • Takeaways & Limitations

    Future work includes learning latent actions from internet-scale general videos for embodied intelligence.

Abstract

from arXiv · show

While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative capabilities and hinders learning from large-scale, heterogeneous data. In this paper, we propose Motus, a unified latent action world model that leverages existing general pretrained models and rich, sharable motion information. Motus introduces a Mixture-of-Transformer (MoT) architecture to integrate three experts (i.e., understanding, video generation, and action) and adopts a UniDiffuser-style scheduler to enable flexible switching between different modeling modes (i.e., world models, vision-language-action models, inverse dynamics models, video generation models, and video-action joint prediction models). Motus further leverages the optical flow to learn latent actions and adopts a recipe with three-phase training pipeline and six-layer data pyramid, thereby extracting pixel-level "delta action" and enabling large-scale action pretraining. Experiments show that Motus achieves superior performance against state-of-the-art methods in both simulation (a +15% improvement over X-VLA and a +45% improvement over Pi0.5) and real-world scenarios(improved by +11~48%), demonstrating unified modeling of all functionalities and priors significantly benefits downstream robotic tasks.

1. Introduction

Motus unifies understanding, video generation, and action modeling through a shared architecture and coordinated multimodal generation. Its optical-flow latent actions and staged data recipe support scalable pretraining, with reported gains over state-of-the-art methods in simulation and real-world scenarios.

  • Motivation: Existing embodied methods isolate vision-language-action, world-model, and generative capabilities, while F1 excludes world and video-generation models.This fragmentation leaves the five-way unification of embodied modeling incomplete.
  • Architecture: Motus integrates understanding, video-generation, and action experts through Mixture-of-Transformers and shared multi-head self-attention.Tri-model Joint Attention preserves specialized functionalities while enabling cross-modal knowledge fusion.
  • Architecture: A UniDiffuser-style scheduler coordinates modality-specific timesteps and noise scales, enabling marginal, conditional, and joint distributions across inference modes.The modes include world modeling, vision-language-action, inverse dynamics, video generation, and video-action joint prediction.
  • Training recipe: Optical-flow latent actions encode pixel-level “delta action” representations that bridge visual dynamics and control signals.Low-dimensional latents receive supervision from limited task-related and task-agnostic action labels.
  • Training recipe: Motus uses three training phases and a six-layer data pyramid to align behaviors across embodiments and share interaction knowledge with target robots.The phases are video pretraining, latent-action pretraining, and embodiment-specific action finetuning.
  • Results: +15% over X-VLA and +45% over π0.5 in simulation, with +11~48% improvements in real-world scenarios.The reported results indicate that general and domain-specific priors can be fused for policy-learning generalization.

2. Related Works

Related work spans unified multimodal models, embodied policies, and latent-action representations. Motus builds on these directions to connect multimodal experts and use visual dynamics for scalable action pretraining.

  • Unified multimodal models: Unified multimodal models share generation and understanding components, but embodied foundation models have largely developed as separate paradigms.Bagel is cited as using Mixture-of-Transformers to share attention layers between understanding and generation experts.
  • Latent actions: Existing latent-action methods often couple inverse dynamics and forward dynamics models to reconstruct future frames, but RGB supervision can include task-irrelevant appearance information.Low-dimensional autoencoder latents are commonly used to reduce redundant information.
  • Latent actions: Motus uses optical-flow-based motion representations to align cross-embodiment behaviors and learn latent actions for large-scale pretraining.The approach connects motion information with transferable robotic learning.

3. Problem Formulation and Challenges

The paper formulates language-conditioned robotic manipulation across observations, language, proprioception, and action, then identifies unification and heterogeneous-data utilization as central challenges.

  • Embodied policies: Language-conditioned policies predict future action chunks from visual observations, proprioception, and language instructions.The policy models pθ(at+1:t+k | ot, pt, ℓ) or pθ(at+1:t+k | ot, ℓ).
  • Embodied policies: The five embodied modeling types are VLA, world model, inverse dynamics, video generation, and video-action joint prediction.They correspond to action prediction, future-observation prediction, action inference from observation sequences, video prediction, and joint video-action prediction.
  • Challenge 1: Unifying Multimodal Generative Capabilities: Current systems struggle to jointly model the five distributions because they lack the complete combination of visual-understanding and physical-interaction priors.The stated gap concerns unified modeling of vision, language, and action within one framework.
  • Challenge 2: Utilization of Heterogeneous Data: Heterogeneous embodiments differ in action dimensions, ranges, semantics, morphology, actuation, and sensing, making control signals difficult to reuse directly.Existing approaches remain primarily dependent on labeled robotic trajectories and cannot readily integrate large-scale internet video.

4. Methodology

Motus unifies video, action, and vision-language experts through shared attention and timestep scheduling, while latent actions connect optical-flow motion to executable control. Its training recipe combines staged learning with heterogeneous embodied data organized across six levels.

  • Model Architecture: Motus gives each expert a Transformer while sharing multi-head self-attention for cross-modal feature fusion without task interference.The shared design is termed Tri-model Joint Attention.
  • Model Architecture: A UniDiffuser-like scheduler assigns different timesteps and noise scales to videos and actions, supporting switching among five embodied modeling modes.The modes include VLA, world model, IDM, VGM, and joint video-action prediction.
  • Action-Dense Video-Sparse Prediction: Motus downsamples video frames relative to actions to balance token counts and reduce redundant video prediction.One example sets the video frame rate to one-sixth of the action frame rate.
  • Latent Actions: Optical-flow latent actions encode pixel-level motion into a lower-dimensional control representation for action pretraining on heterogeneous videos and robot trajectories.A DC-AE reconstructs optical flow, and a lightweight encoder projects four 512-dimensional tokens into a 14-dimensional vector.
  • Latent Actions: The latent-action objective combines flow reconstruction, latent-to-real-action alignment, and KL regularization, with task-agnostic data providing real action supervision.The training mix uses 90% unlabeled reconstruction data and 10% labeled trajectories.
  • Model Training and Data: Motus trains through video pretraining, latent-action pretraining, and target-robot fine-tuning across six data levels spanning web, human, simulation, and robotic sources.The data pyramid organizes sources by richness and policy relevance, with quantity decreasing and quality increasing toward higher levels.

5. Experiments

Motus is evaluated in simulated and real-world robotic manipulation settings against established baselines, including randomized multi-task simulation and two dual-arm platforms. The experiments also examine partial success, training-stage ablations, and task-level performance.

  • Simulation Evaluation: Motus is compared with π0.5 and X-VLA in RoboTwin 2.0 simulation, including clean and randomized settings.The evaluation covers single-task and multi-task manipulation tasks.
  • Simulation Evaluation: The randomized benchmark tests generalization across varied backgrounds, clutter, table heights, lighting, and instructions under distribution shift.All models receive only 40k fine-tuning steps from pretrained checkpoints.
  • Real-World Experiments: Motus is evaluated on two real-world dual-arm platforms, AC-One and Agilex-Aloha-2, across spatial, deformable-object, fluid-control, visual-understanding, and long-horizon tasks.Examples include folding towels, brewing coffee, and grinding coffee beans.
  • Real-World Experiments: Motus significantly outperforms π0.5 across all reported real-world tasks on both robotic arms using partial success rate.Partial success decomposes long-horizon tasks into subtasks and assigns scores for achieved subgoals.
  • Ablation Studies: Ablations compare the original Stage 2-pretrained Motus with models having no pretraining or only Stage 1 pretraining.Simulator results are summarized in Figure 6, while real-world comparisons include a from-scratch counterpart.

6. Conclusion and Limitations

Motus unifies major embodied-model capabilities through a latent-action generative framework that combines pretrained experts, multimodal scheduling, and optical-flow motion representations. Across simulation and real-world environments, it consistently outperforms existing state-of-the-art embodied models, while future work targets broader architectures and motion priors.

  • Conclusion: Motus integrates vision-language understanding, video generation, inverse dynamics, world modeling, and video-action joint prediction in one framework.The model connects pretrained experts through Mixture-of-Transformers and uses latent actions as pixel-level motion representations.
  • Conclusion: +15~45% improvement in simulation and +11~48% in real-world scenarios are reported over existing state-of-the-art embodied models.The paper presents these results as evidence for unified multimodal capabilities and shared motion priors.
  • Limitations and Future Work: Future work will explore more advanced unified architectures, universal motion priors, and latent actions learned from internet-scale general videos.These directions extend Motus toward broader embodied-intelligence pretraining.

7. Training and Inference of the Unified Model

Motus uses a unified training and inference framework that supports five modeling modes, combining multimodal generation with action prediction. Experiments report strong video, future-prediction, inverse-dynamics, VLA, and joint video-action capabilities.

  • Inference Modes: Motus supports five inference modes by switching between observation and action timesteps in a shared generative framework.The modes include world model, IDM, VLA, VGM, and video-action joint prediction.
  • Experimental Results: Motus’s VGM mode produces high-quality visualizations across Agilex-Aloha-2 and AC-One embodiments.
  • Experimental Results: Motus’s world model mode generates high-quality future videos across two real-world robotic embodiments.The evaluation uses real-world robot data and is reported with Table 6 metrics.
  • Experimental Results: Motus achieves lower action MSE than specifically trained ResNet-18 and DINOv2 IDM baselines on RoboTwin 2.0 randomized data.The baselines predict 16-action chunks from current observations.
  • Experimental Results: Motus demonstrates competitive VLA performance and can generate videos and precise actions simultaneously in video-action joint prediction mode.

8. More Experiments Results

Additional evaluations cover RoboTwin 2.0 simulation, LIBERO-Long, and VLABench, extending assessment across long-horizon and language-conditioned manipulation tasks. Motus reaches state-of-the-art performance on LIBERO-Long and is evaluated across multiple VLABench tracks.

  • RoboTwin 2.0: RoboTwin 2.0 simulation evaluates Motus and baselines across 50 tasks in both clean and randomized scenes.
  • LIBERO-Long: 97.6 average success on LIBERO-Long matches X-VLA’s best reported performance and reaches state-of-the-art results.LIBERO-Long contains 10 language-conditioned, long-horizon manipulation tasks.
  • VLABench: VLABench evaluates a single multi-task-finetuned Motus model by success rate across three tasks and two tracks: In Distribution and Cross Category.The benchmark covers manipulation, vision understanding, semantic comprehension, common sense, and reasoning.
  • Real-World Evaluation: Real-world task results provide detailed subtask breakdowns for AC-One and Agilex-Aloha-2 experiments.Figure 8 visualizes Motus execution for the tasks listed in Table 3.

9. Implementation Details

Implementation details document Motus’s architecture, training data, and three-stage training configuration, alongside visualizations of simulation and real-world behavior. The materials cover both robotic platforms and multiple operating modes.

  • Architecture: Table 11 specifies Motus architecture hyperparameters and key configuration settings.
  • Real-World Execution: Figure 8 visualizes Motus executing real-world tasks with two robots.The figure covers nine tasks.
  • Data: Table 12 details the pretraining and finetuning datasets used by Motus.
  • Training: Table 13 reports training configurations across Motus’s three stages.
  • World Model: Figures 10 and 11 visualize Motus’s world model mode on Agilex-Aloha-2 and AC-One datasets.
  • Inference Visualizations: Figures 12 and 9 provide visualizations of video-action joint prediction and VGM modes during real-world inference and on AC-One.
Loading 2512.13030v2…