Source-linked AI summary

Riemann-1.0: An Embodied World Action Model for Physical AI

Haofeng Sun, Jiangbo Pei, Fei Kang, Zexiang Liu, Yaokun Li, Boyi Jiang, Hua Xue, Cindy Zhou, Wei Li, Yichen Wei, Mengyin An, Fanliang Zhao, Biao Jiang, Zile Wang, Yang Liu, Yangguang Li

arXiv:2608.27033v1cs.RO

TL;DR

Embodied World Action Models need to scale heterogeneous experience while jointly representing observations, states, actions, and world evolution. Riemann-1.0 combines a fully causal autoregressive model with progressive embodied pretraining, achieving state-of-the-art results across simulation and real-world manipulation benchmarks.

  • Problem

    Existing approaches do not fully leverage heterogeneous embodied experience or jointly support executable robot policy learning and causally consistent action-conditioned world simulation.

  • Method

    Riemann-1.0 jointly models visual observations, robot states, and embodiment-specific actions causally, while progressive pretraining transfers supervision from human videos through demonstrations to robot control.

  • Results

    Riemann-1.0 achieves state-of-the-art performance across simulation and real-world tasks, including 62.6% on RoboCasa365 and 85.0% SR with 94.43% PSR in long-horizon real-world manipulation.

  • Takeaways & Limitations

    The results support combining Fully Causal World Action Modeling with Progressive Embodied Pretraining to transfer large-scale embodied experience into generalizable long-horizon robot manipulation.

Abstract

from arXiv · show

We introduce Riemann-1.0, a fully causal autoregressive World Action Model for embodied intelligence. Riemann-1.0 jointly models multi-view visual observations, robot states, and embodiment-specific actions within a unified causal autoregressive sequence, representing robot actions and world evolution as causal state transitions. Unlike existing WAMs based on joint generation, video-first prediction, or decoupled modeling paradigms, Riemann-1.0 unifies online robot policy execution and action-conditioned world simulation within a single model, enabling it to function as both an executable robot policy and a multi-embodiment visual world simulator. To scale embodied experience across heterogeneous data sources, we further develop a progressive embodied pretraining framework that unifies learning from egocentric human videos, handheld-gripper demonstrations, and heterogeneous robot trajectories under a shared World Action Modeling objective. Built upon 200K+ hours of interaction data, Riemann-1.0 progressively transfers large-scale embodied experience into executable robot manipulation capabilities. Riemann-1.0 achieves state-of-the-art performance across both simulation benchmarks and real-world manipulation tasks. It achieves success rates of 94.3% on RoboTwin2.0, 99.0% on LIBERO, and 62.6% on the long-horizon compositional benchmark RoboCasa-365, outperforming the previous best method by 8.4% On long-horizon real-world manipulation tasks, Riemann-1.0 achieves a Success Rate (SR) of 85.0% and a Progress Success Rate (PSR) of 94.4%, exceeding the strongest open-source baseline by 15% in SR. These results demonstrate that unified World Action Modeling together with progressive embodied pretraining effectively transforms large-scale embodied experience into generalizable robot manipulation capabilities.

1. Introduction

Riemann-1.0 addresses the challenge of scaling heterogeneous embodied experience and unifies robot policy execution with action-conditioned visual world simulation in one fully causal model. Its progressive pretraining transfers supervision from human videos through demonstrations to executable robot control, yielding strong simulation and real-world results.

  • Motivation: Existing WAMs and pretraining strategies struggle to jointly leverage heterogeneous embodied data while supporting both robot policy learning and causally consistent world simulation.The data sources differ in modalities, action spaces, and supervision, while prior WAM paradigms model only selected aspects of interaction.
  • Method: Riemann-1.0 jointly models multi-view observations, robot states, and embodiment-specific actions as causal state transitions in one autoregressive sequence.This formulation aligns the model with real-world interaction and supports both executable policies and action-conditioned visual simulation.
  • Method: Progressive Embodied Pretraining transfers knowledge from weakly supervised human interaction data to executable control through a three-stage curriculum.The framework unifies egocentric videos, handheld-gripper demonstrations, and robot trajectories under a shared World Action Modeling objective.
  • Results: 94.3% success on RoboTwin2.0, 99.0% on LIBERO, and 62.6% on RoboCasa-365 establish strong simulation performance.The RoboCasa-365 result exceeds the best method by 8.4 percentage points.
  • Results: 85.0% Success Rate (SR) and 94.4% Progress Success Rate (PSR) are achieved on long-horizon real-world manipulation tasks.The real-world SR exceeds the strongest open-source baseline by 15 percentage points.
  • Data and scaling: 200K+ hours of embodied experience underpin the framework, spanning heterogeneous sources and thousands of interaction skills.The corpus is designed to transfer scalable interaction knowledge, cross-embodiment alignment, and embodiment-specific executable supervision.

2. Data Infrastructure

The unified embodied data infrastructure converts heterogeneous human, handheld-gripper, and robot experience into a common action-level representation. Processing, alignment, filtering, and semantic-aware balancing create a scalable supervision source for World Action Modeling.

  • Challenge: Heterogeneous embodied sources differ in observations, action representations, temporal granularity, and supervision fidelity, complicating unified large-scale pretraining.The corpus combines egocentric human videos, handheld-gripper demonstrations, and heterogeneous robot trajectories.
  • Unified representation: The data infrastructure converts all sources into a common action-level video–state–action representation with language, embodiment, observations, states, actions, and semantic metadata.State and action streams are temporally aligned and normalized under embodiment-specific canonical definitions, with validity masks for variable-dimensional actions.
  • Corpus: 230K+ hours of embodied experience span human videos, handheld-gripper and wearable demonstrations, and heterogeneous robot trajectories across diverse environments and embodiments.The corpus covers household, office, and industrial settings and thousands of interaction skills.
  • Source-specific processing: Human videos receive hierarchical task and action annotations, while handheld-gripper and robot trajectories undergo temporal alignment, action normalization, and trajectory quality filtering.Gripper-state transitions refine action boundaries, and additional filters remove camera shake, stationary end effectors, and abnormal controls.
  • Processing pipeline: The processing pipeline applies visual preprocessing, semantic annotation, quality filtering, 3D hand reconstruction, geometric filtering, and semantic-aware data balancing.Handheld-gripper and robot data additionally undergo action calibration and robot-specific filtering.
  • Balancing: Semantic-aware sampling preserves the scale of human datasets while increasing exposure to long-tail skills and low-resource robot embodiments.Balancing operates across scenes, tasks, skills, objects, and embodiments rather than raw data volume.

3. Riemann-1.0 Design

Riemann-1.0 models actions, robot states, and visual latents in a fully causal autoregressive sequence, supporting both robot policy execution and action-conditioned visual simulation. Its shared backbone combines flow-matching action and visual predictions with embodiment-specific interfaces and causal online inference.

  • Causal action-video modeling: A WAM jointly models robot actions and their visual consequences to represent coupled control and observation dynamics over time.
  • Causal action-video modeling: Riemann-1.0 factorizes interaction causally by predicting actions from prior visual, state, and action histories, then generating the corresponding visual latent conditioned on the current action.This ordering follows real robot interaction, where actions precede observed visual consequences.
  • Causal action-video modeling: The model replaces predicted visual latents with real post-action observations during deployment, while recursively feeding predicted latents back during visual simulation.
  • Multi-embodiment interfaces: Riemann-1.0 shares a transformer backbone across robots while using embodiment-specific projections, prediction heads, and padded canonical action-state interfaces.
  • Training and inference: The training sequence serializes visual, state, and action tokens, pairing each visual latent with a temporally aligned chunk of low-level actions under structured causal masking.
  • Training and inference: Separate flow-matching heads predict visual-latent and action velocities, with the objective balancing their losses through λ.

4. Training Recipe

Riemann-1.0 uses a three-stage curriculum that moves from video-only pseudo-action supervision to mixed real trajectories and finally high-quality robot demonstrations. This progression transfers broad visual dynamics into executable, embodiment-specific control while retaining a shared action-video objective.

  • Three-stage curriculum: The three-stage curriculum progresses from LAM-derived pseudo actions, through mixed multi-embodiment trajectories, to high-quality robot-only demonstrations.
  • Training scale and supervision: The framework is built on more than 200K+ hours of embodied experience and unifies heterogeneous sources under one action-video objective.
  • LAM-Action Bootstrap: Stage I trains on unlabeled human videos using a frozen Latent Action Model as a pseudo-action annotator for manipulation dynamics.The stage initializes visual dynamics rather than learning a deployable robot policy.
  • LAM-Action Bootstrap: The frozen LAM converts adjacent-frame transitions into deterministic pseudo-action codes that are grouped into action chunks aligned with the WAM interface.
  • Trajectory-Grounded Alignment: Stage II replaces pseudo actions with UMI demonstrations, robot trajectories, and human videos annotated with 3D hand poses and keypoints.These sources provide more reliable action supervision across viewpoints, objects, tasks, and embodiments.
  • Robot-Policy Enhancement: Stage III uses only high-quality robot demonstrations and increases the action weight to λ=0.9 to sharpen executable action predictions.

5.1. Real-World Experiments

Real-world experiments evaluate Riemann-1.0 on four challenging manipulation tasks and held-out instructions. The model achieves strong final-task completion and intermediate progress, while transferring performance to compositional and out-of-domain settings.

  • Post-Training Data and Setup: The real-world suite covers ordered cube stacking, clothes folding, desk organization, and kitchen organization.These tasks span rigid and deformable objects, precise insertion, cluttered rearrangement, and long-horizon sequencing.
  • Real-World Evaluation: Success Rate measures final goal completion, whereas Progress Success Rate measures completion of task-specific intermediate milestones.PSR captures partial progress when long-horizon tasks do not reach their final state.
  • Real-World Evaluation: 85.00% average SR and 94.43% average PSR give Riemann-1.0 the best average real-world performance among compared models.It maintains at least 80.0% SR and exceeds 91.0% PSR on all four tasks.
  • Real-World Evaluation: 85.0% SR and 91.6% PSR on ordered cube stacking improve over G0.5 by 5.0 and 2.1 points, respectively.On kitchen organization, the model reaches 90.0% SR and improves PSR over LingBot-VLA by 18.2 points.
  • Compositional Generalization and Out-of-Domain Evaluation: 65.0% average SR across compositional tasks and 75.0% across four held-out tasks show performance beyond the post-training task set.The stricter compositional setting reaches 50.0% SR, while OOD evaluation reaches 85.0% average SR, including 70.0% on towel-to-basin placement.

5.2. Simulation Evaluation

Simulation evaluation measures task success on RoboCasa365, RoboTwin 2.0, and LIBERO using a shared temporal interface. Riemann-1.0 achieves leading performance across the reported benchmarks and RoboCasa365 categories.

  • Simulation Evaluation: The evaluation covers RoboCasa365, RoboTwin 2.0, and LIBERO, reporting task success rate under each benchmark’s standard protocol.The same WAM temporal interface maps one visual latent to 16 low-level action steps.
  • RoboCasa365: 62.6% average success rate on RoboCasa365 improves from the strongest baseline’s 54.2% and leads across Atomic-Seen, Composite-Seen, and Composite-Unseen.The gains on Composite-Seen and Composite-Unseen are 11.7 and 10.7 percentage points.
  • LIBERO: 99.0% overall average success rate on LIBERO comprises 99.6% Spatial, 100.0% Object, 97.6% Goal, and 98.6% Long.The benchmark contains four official suites with 10 tasks each.
  • Multi-Embodiment Evaluation: Multi-embodiment rollouts span humanoid, dual-arm, single-arm, dexterous-hand, and simulated platforms while preserving action-consistent scene and object motion.The examples vary in viewpoint, embodiment geometry, gripper appearance, and task type.

5.3. Action-Conditioned Visual Rollout as a Multi-Embodiment Simulator

Riemann-1.0 operates as an action-conditioned visual simulator by predicting future visual consequences from current context and a candidate action trajectory. Qualitative rollouts cover diverse embodiments and viewpoints.

  • Action-Conditioned Visual Rollout: The simulator receives visual observation, task prompt, robot state, and a candidate future action trajectory, then predicts visual consequences.The trajectory uses the same embodiment-specific action interface as policy training and is embedded as action tokens.
  • Action-Conditioned Visual Rollout: Future visual latents are rolled out from encoded visual context and decoded back into RGB video.Action chunks may come from the policy head, a sampled candidate plan, or a recorded trajectory.
  • Multi-Embodiment Simulation: Qualitative rollouts span heterogeneous embodiments and camera layouts while preserving main scene structure and action-consistent object and robot motion.Examples include humanoid, dual-arm, single-arm, dexterous-hand, RoboTwin, and RoboCasa-GR1 settings.

6. Related Work

Related work includes vision-language-action policies, World Action Models, action-conditioned video systems, and approaches that increasingly combine policy execution with predictive modeling and staged training.

  • Vision-Language-Action Policies: Vision-language-action policies scale robot manipulation through broad demonstrations, semantic pretraining, continuous action prediction, and cross-embodiment learning.Representative systems include RT-1, RT-2, OpenVLA, π0, GR00T-N1, and LingBot-VLA.
  • World Action Models: World Action Models jointly model robot actions and future observations for planning, evaluation, or policy learning.Existing formulations differ in joint generation, video-action pretraining, action-centered inference, and separated visual supervision and action decoding.
  • Policies and Action-Conditioned Simulators: Recent robot-learning systems increasingly use staged training recipes that combine semantic pretraining with robot imitation.This trend complements architectural efforts to connect policy and world-modeling capabilities.
  • Policies and Action-Conditioned Simulators: Robot policies prioritize fast action prediction, whereas model-based simulators prioritize future observations or state transitions.Video-based robot learning increasingly connects these roles through planning, future representation learning, and policy supervision.

7. Conclusion

Riemann-1.0 is presented as a fully causal World Action Model that addresses heterogeneous embodied learning and unified causal modeling. It achieves state-of-the-art results across simulation and real-world manipulation evaluations.

  • Riemann-1.0 is a fully causal autoregressive World Action Model for embodied intelligence.
  • The model addresses unifying heterogeneous embodied experience and jointly modeling actions, states, and world evolution in a causal form aligned with real interaction.Its pretraining uses more than 200K+ hours spanning egocentric human videos, handheld-gripper demonstrations, and robot trajectories.
  • 62.6% on RoboCasa365, 94.3% on RoboTwin 2.0, and 99.0% on LIBERO demonstrate state-of-the-art simulation performance.
  • 85.0% SR and 94.43% PSR were achieved on long-horizon real-world manipulation tasks.
  • The results support a scalable path for transforming large-scale embodied experience into generalizable robot manipulation capabilities.
Loading 2608.27033v1…