Source-linked AI summary

$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation

Pengfei Zhou, Shengcong Chen, Di Chen, Jiaxu Wang, Rongjun Jin, Bingwen Zhu, Yike Pan, Songen Gu, Kuanning Wang, Shufeng Nan, Xingyu Qiu, Chenhao Qiu, Pu Yang, Yunuo Cai, Jianxiong Gao, Yifan Li, Yanwei Fu, Xiangyu Yue, Zhi Chen, Jianlan Luo

arXiv:2606.01027v1cs.RO

TL;DR

Robotic manipulation needs executable actions that anticipate uncertain physical consequences before execution. τ0-WM unifies action generation, video prediction, and action evaluation, achieving the best average success rate among evaluated baselines across four manipulation tasks.

  • Problem

    Robotic manipulation lacks predictive models that connect executable actions with future visual and task outcomes before physical execution.

  • Method

    τ0-WM shares a video diffusion backbone for action generation, future prediction, and action-conditioned evaluation, with deployment-time proposal, simulation, and revision.

  • Results

    τ0-WM achieves the best average success rate among evaluated baselines across four manipulation tasks and performs best on most tasks.

  • Takeaways & Limitations

    The results support using jointly learned video prediction and executable action generation for deployment-time imagining, scoring, and refining of manipulation actions.

Abstract

from arXiv · show

Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present $τ_0$-World Model ($τ_0$-WM), a unified video-action world model that integrates policy learning, video prediction, and action evaluation within a single future-predictive framework. Built on a shared video diffusion backbone, $τ_0$-WM provides two complementary interfaces. First, a video action model jointly predicts future visual latents and continuous action chunks from multi-view observations, language instructions, and robot state. Second, an action-conditioned video simulator rolls out candidate action chunks into multi-view futures and predicts dense task-progress scores. The model is trained on approximately $27{,}300$ hours of real-robot teleoperation, UMI-style interaction, egocentric human videos, and rollout or failure trajectories using modality-specific supervision masks. At inference time, $τ_0$-WM uses test-time computation to sample action candidates, rank them with re-denoising consistency, and invoke simulator-based rectification for low-quality candidates. On challenging long-horizon and fine-grained robotic manipulation tasks, $τ_0$-WM shows superior performance over other relevant baselines.

I. INTRODUCTION · II. RELATED WORK

τ0-WM frames robotic manipulation as predictive action under uncertain physical consequences, unifying video prediction, executable action generation, and action-conditioned evaluation in one shared world model. It combines heterogeneous data and test-time proposal–evaluation–revision to improve manipulation across challenging tasks and embodiments.

  • I. INTRODUCTION: Robotic manipulation requires anticipating how actions alter scenes through contact, object motion, and multi-step interactions.This motivates connecting manipulation policies with predictive control and world models.
  • I. INTRODUCTION: Egocentric videos provide broad visual evidence about object motion, contacts, and long-horizon task organization but lack deployable robot control actions.These data cover diverse objects, scenes, and behaviors without specifying actions in the robot’s control space.
  • I. INTRODUCTION: Robot demonstrations ground observations in continuous actions but cover narrower objects, environments, and tasks than broad video data.The paper argues that robot-only training is grounded but narrow, whereas video-only training is predictive but action-ungrounded.
  • I. INTRODUCTION: τ0-WM places future observations, robot actions, and task progress within one predictive model using modality-specific supervision.Video-only data train visual dynamics, robot trajectories train executable actions, and progress or failure trajectories train action-conditioned evaluation.
  • I. INTRODUCTION: τ0-WM integrates action generation, video prediction, and action-conditioned future evaluation around a shared video diffusion backbone.Its complementary interfaces are a Video Action Model and an Action-Conditioned Video Simulator.
  • I. INTRODUCTION: Approximately 27,300 hours of heterogeneous data train τ0-WM across real-robot teleoperation, UMI-style demonstrations, egocentric videos, and rollout or failure trajectories.The sources provide different degrees of supervision and action fidelity, including deployment-aligned continuous actions from real-robot demonstrations.
  • I. INTRODUCTION: At inference, τ0-WM samples action chunks, ranks them with re-denoising consistency, and uses simulator-based rectification for unreliable candidates.This proposal–evaluation–revision procedure uses predicted futures directly to improve actions before execution.
  • II. RELATED WORK: τ0-WM unifies robotic video action models and action-conditioned video simulators within a single video-action world-modeling framework.The related-work discussion organizes prior research around these two areas and their intersections.

A. Robotic Video Action Models · B. Action-Conditioned Video Simulators for Robotics · III. DATA SOURCES FOR PREDICTIVE ROBOT LEARNING

τ0-WM unifies video-action policy learning, action-conditioned simulation, and evaluation through shared predictive representations. It trains on heterogeneous robot, UMI-style, and egocentric data with modality-specific supervision to support executable, future-aware manipulation.

  • A. Robotic Video Action Models: Prior video action models typically use joint denoising to generate future visual latents and action chunks together, yielding dynamics-aware representations for manipulation.These methods commonly build on pretrained video-generation diffusion models.
  • A. Robotic Video Action Models: τ0-WM jointly predicts multi-view future visual latents and executable action chunks while sharing predictive representations with its action-conditioned simulator.This extends future prediction beyond auxiliary policy learning to support action evaluation at test time.
  • B. Action-Conditioned Video Simulators for Robotics: τ0-WM’s Action-Conditioned Video Simulator shares the VAM’s action interface and backbone, trains on the same heterogeneous mixture, and predicts multi-view rollouts plus task-progress scores.The simulator is integrated rather than treated as a separate module.
  • B. Action-Conditioned Video Simulators for Robotics: At test time, τ0-WM samples candidate actions, ranks them by re-denoising consistency, and uses ACVS to evaluate and rectify low-quality candidates.This adds simulator-based decision support beyond feed-forward action prediction.
  • III. DATA SOURCES FOR PREDICTIVE ROBOT LEARNING: 27.3K hours of training data combine 17.8K hours of real-robot teleoperation, 6.5K hours of filtered UMI-style demonstrations, and 3.0K hours of additional interaction data.The real-robot data spans AGIBOT-G01, ARX manipulators, and dual-arm Franka systems.
  • III. DATA SOURCES FOR PREDICTIVE ROBOT LEARNING: Real-robot teleoperation supplies deployment-aligned action supervision across household, retail, and industrial settings, with actions aligned to robot kinematics, control, sensing, and deployment conditions.Data comes from AGIBOT-G01, ARX, and dual-arm Franka platforms using head-view and wrist-mounted cameras.
  • III. DATA SOURCES FOR PREDICTIVE ROBOT LEARNING: UMI-style demonstrations provide scalable, diverse visual interaction data and action-like device-motion signals, while egocentric videos broaden behavioral coverage but supervise only visual prediction.Egocentric videos lack robot-compatible action labels and differ in embodiment and viewpoint.
  • III. DATA SOURCES FOR PREDICTIVE ROBOT LEARNING: Modality-specific supervision masks specify observed inputs, predicted targets, and active losses, enabling heterogeneous sources to train one end-to-end video-action objective.The unified representation respects differences in supervision reliability and availability.

IV. VIDEO ACTION MODEL … C. Joint Flow-Matching Objective

The Video Action Model (VAM) is τ0-WM’s policy-facing interface, jointly predicting future visual latents and executable continuous action chunks from observations, language, and robot state. Its shared predictive representation couples video dynamics and action generation, while masked joint flow matching unifies heterogeneous supervision.

  • A. Model Interface and Problem Formulation: VAM predicts future visual latents and executable action chunks from multi-view observations, language instructions, and robot state.Future visual prediction learns transferable interaction dynamics, while action prediction grounds representations in executable robot control.
  • A. Model Interface and Problem Formulation: Future visual prediction transfers interaction dynamics from heterogeneous data, including videos without action annotations, while action prediction supports robot control.The two objectives complement one another by combining broad dynamics learning with executable supervision.
  • B. Architecture: VAM couples a video branch for future visual prediction with an action branch for executable action generation through a shared predictive representation.Feature-level cross-attention lets future visual dynamics directly support action generation.
  • B. Architecture: VAM uses Wan2.2-TI2V-5B with a Wan VAE and a 5B-parameter Wan video DiT backbone to process synchronized multi-view latent canvases.The current observation latent remains clean as visual context, while future slots are noised and denoised.
  • B. Architecture: Action tokens model temporal dependencies and cross-attend to instruction-aware, dynamics-relevant video features conditioned on clean visual context and language.This feature-level coupling supports action generation while preserving the video backbone as the shared predictive foundation.
  • C. Joint Flow-Matching Objective: VAM applies flow matching to both future video latents and continuous action chunks using separate noise levels, noised inputs, and velocity targets.The objective jointly trains video and action vector-field heads, with intermediate video features consumed by the action branch.
  • C. Joint Flow-Matching Objective: Supervision masks unify robot trajectories and egocentric human videos: robots provide visual and action supervision, whereas human videos provide only visual dynamics.All experiments set λz = λa = 1, allowing heterogeneous samples with missing modalities to participate in one training process.

D. Inference and Deployment … B. Architecture

τ_0-WM deploys a Video Action Model (VAM) that predicts executable action chunks from observations, instructions, and robot state, alongside an Action-Conditioned Video Simulator (ACVS) that evaluates candidate actions through imagined futures and dense rewards. ACVS reuses the shared video backbone to condition future latent rollouts on proposed actions without generating actions itself.

  • D. Inference and Deployment: VAM takes multi-view observations, a language instruction, and robot state to predict executable action chunks for receding-horizon execution.Future latents may be decoded into frames for explicit visual rollouts or retained as latent representations for action generation.
  • A. Simulator Interface and Problem Formulation: ACVS evaluates candidate action chunks by predicting action-conditioned future visual rollouts and dense reward trajectories instead of physically executing every candidate.It provides an action-conditioned proxy for deployment-time evaluation.
  • A. Simulator Interface and Problem Formulation: Given memory observations, a language instruction, and a candidate action chunk, ACVS predicts future video latents together with dense reward scores.The imagined future latent rollout is denoted ˆz, and the predicted reward trajectory is denoted ˆr.
  • A. Simulator Interface and Problem Formulation: The architecture couples VAM’s future-latent and action prediction through a shared video backbone and an Action DiT branch with cross-attention.VAM is the policy interface, while ACVS is the evaluation interface.
  • A. Simulator Interface and Problem Formulation: ACVS is an evaluation interface rather than a policy: it treats the candidate action chunk as a clean condition and estimates the future it induces.Different candidate actions can therefore produce different imagined futures under the same observation and instruction.
  • B. Architecture: ACVS reuses the Wan VAE and video transformer backbone, removes the Action DiT policy branch, and denoises noisy future latent slots from encoded observation context.Memory and current observations form clean latent context, while future slots are initialized with noise.
  • B. Architecture: For each future latent slot, temporally aligned actions are grouped into action blocks and projected through lightweight MLPs for diffusion-time and AdaLN conditioning.The resulting conditions are broadcast across spatial tokens and camera views, while observation slots remain unconditioned.

C. Reward and Progress Scoring · D. Training Objective

ACVS scores candidate action chunks with dense subtask-level progress rewards, including failures and recovery trajectories to distinguish meaningful progress from plausible but unsuccessful motion. Its training jointly supervises future video latents and reward trajectories through flow matching, while test-time computation samples, ranks, and optionally evaluates candidates with ACVS.

  • C. Reward and Progress Scoring: ACVS predicts dense reward trajectories for candidate action chunks using subtask-level progress labels and Monte Carlo propagation for frame-level supervision.This replaces a single terminal success label with dense reward supervision across each subtask segment.
  • C. Reward and Progress Scoring: Failed subtask segments receive negative trajectory rewards, teaching ACVS to recognize unsuccessful contact, incorrect object motion, and task regression.The reward construction distinguishes meaningful task progress from visually plausible motion that does not advance the task.
  • C. Reward and Progress Scoring: Failure-heavy and recovery trajectories improve simulator fidelity by exposing ACVS to off-distribution actions, failed interactions, and recovery behaviors.These trajectories are valuable for simulator learning even when they are suboptimal as direct policy supervision.
  • D. Training Objective: ACVS uses the same flow-matching formulation as VAM to jointly supervise future video latents and dense reward trajectories.The formulation conditions future latent rollouts and target reward trajectories on candidate actions and clean visual context.
  • D. Training Objective: At test time, the procedure samples N candidate actions, ranks them by RCS, returns the top candidate when its score reaches γ, and otherwise evaluates candidates with ACVS using rollout value J(i).The algorithm specifies candidate sampling, RCS-based selection, thresholding, ACVS evaluation, and rollout-value computation.
  • D. Training Objective: Training constructs noised video and reward inputs with corresponding velocity targets and optimizes their prediction from action-conditioned context.The notation includes future latent rollouts z_t+1:t+H_v, reward trajectories r_t:t+H_a−1, and noise levels u_z and u_r.
  • D. Training Objective: The model uses separate video and reward velocity predictors, with action-conditioned video features consumed by the reward head and λ_z = λ_r = 1.The equal loss weights are used in all experiments.

VI. TEST-TIME COMPUTATION · A. Re-denoising Consistency Score

τ0-WM uses coarse-to-fine test-time computation to select reliable actions from multimodal candidates, applying re-denoising consistency first and invoking simulator-based rectification only when candidates appear unreliable. The Re-denoising Consistency Score filters candidates by consistency with the learned conditional action manifold at negligible overhead.

  • VI. TEST-TIME COMPUTATION: Multimodal conditional action distributions yield multiple feasible sequences that differ in precision, robustness, and likelihood of success, making deployment-time selection important.
  • VI. TEST-TIME COMPUTATION: τ0-WM samples multiple action candidates from VAM and applies a lightweight self-consistency filter before using ACVS for expensive rollout evaluation and rectification.This coarse-to-fine strategy preserves real-time performance in most situations while recovering from difficult states.
  • VI. TEST-TIME COMPUTATION: Algorithm 2 rectifies low-quality actions by converting the selected future latent into conditioning, re-querying VAM, and generating a refined action chunk.
  • A. Re-denoising Consistency Score: Given context C_t = (o_t, p, s_t), VAM samples N candidate action chunks for test-time evaluation.
  • A. Re-denoising Consistency Score: For each candidate, the method samples K flow timesteps, re-noises the action using training’s flow-matching process, and computes average re-denoising error E^(i) with VAM’s action vector field.
  • A. Re-denoising Consistency Score: The Re-denoising Consistency Score (RCS) is defined from the candidate’s average re-denoising error.
  • A. Re-denoising Consistency Score: RCS favors candidates consistent with the learned conditional action manifold while adding negligible computational overhead relative to rollout-based evaluation.

B. Low-quality Action Rectification

Low-quality Action Rectification addresses states where even the most self-consistent sampled action may be poor. It evaluates candidate futures with ACVS, selects the highest-value rollout, and conditions a second policy query on that future to generate a refined action chunk.

  • B. Low-quality Action Rectification: LAR is introduced because RCS may select a self-consistent candidate even when all sampled actions are poor in challenging states.This motivates rectification after candidate selection.
  • B. Low-quality Action Rectification: When the selected candidate meets the reliability condition, ACVS evaluates all candidate actions using imagined rollouts and dense reward trajectories.ACVS is invoked based on a reliability threshold γ.
  • B. Low-quality Action Rectification: Each candidate receives a rollout value based on the maximum task progress achieved by its imagined rollout, denoted J(i).The rollout value aggregates the progress represented by the imagined trajectory.
  • B. Low-quality Action Rectification: The candidate with the highest rollout value is selected as the most promising future.Selection is based on predicted future task progress rather than direct execution of the original candidate.
  • B. Low-quality Action Rectification: Rather than executing the selected action directly, VAM receives the selected rollout latent as an additional future condition and generates a refined action chunk.This second policy query explicitly guides the action toward the selected high-value future.

VII. EXPERIMENTAL EVALUATION · A. Main Results

τ0-WM is evaluated on challenging, long-horizon, fine-grained real-robot manipulation tasks across three robot embodiments using task success rate and task-accomplishment progress. It achieves the highest average success rate and performs best on most tasks, while its explicit future modeling supports additional corrective actions beyond binary completion.

  • VII. EXPERIMENTAL EVALUATION: The evaluation tests whether τ0-WM enables strong policies, benefits from heterogeneous pre-training, and improves closed-loop execution through deployment-time computation.Experiments use robot, UMI-style, and egocentric interaction data and examine deployment-time action selection.
  • VII. EXPERIMENTAL EVALUATION: Experiments span AGIBOT-G01, ARX manipulators, and a dual-arm Franka system on language-conditioned, multi-view packing and assembly tasks.τ0-WM is compared with π0.5 and Fast-WAM, while deployment-time reasoning is compared with standard execution, CFG, and ACG.
  • A. Main Results: Four precision-sensitive tasks excluded from pre-training require long-horizon reasoning, multi-stage interaction, and precise geometric alignment.School Bag involves sequential zipper manipulation and object placement, while Faucet requires accurate hose alignment and secure attachment.
  • A. Main Results: The tasks cover multiple embodiments: Badminton uses the ARX manipulator, Faucet uses the dual-arm Franka platform, and the remaining tasks use AGIBOT-G01.This setup evaluates embodiment diversity across the main task suite.
  • A. Main Results: τ0-WM achieves the highest average success rate and performs best on most tasks, remaining consistently strong across all four tasks.π0.5 is competitive on Toolbox but degrades on longer-horizon and fine-grained tasks; Faucet remains challenging for every method.
  • A. Main Results: Task-accomplishment progress reveals qualitative differences that binary success misses, including incomplete tool insertion by baseline policies on Toolbox.The figure evaluates both task success rate and stepwise task-accomplishment progress.
  • A. Main Results: τ0-WM often performs corrective pushing or pressing after tool insertion, targeting the quality of the final scene configuration rather than an intermediate completion state.The paper hypothesizes that explicit future visual-outcome modeling produces this behavior.

B. Ablation Studies · VIII. CONCLUSION AND FUTURE WORK

The ablations show that heterogeneous pretraining data and test-time computation substantially improve manipulation performance, while τ0-WM unifies action generation, future prediction, and deployment-time reasoning. The conclusion identifies tactile sensing, stronger reasoning, and longer-horizon prediction as key directions for more capable robot foundation models.

  • B. Ablation Studies: Pretraining with UMI and egocentric data improves τ0-WM performance in both zero-shot execution and supervised fine-tuning settings.The comparison contrasts robot-teleoperation-only training with the complete pretraining corpus.
  • B. Ablation Studies: 0.14 to 0.55 average success rate is the largest gain from adding UMI and egocentric data in zero-shot manipulation.The benefit remains visible after fine-tuning, particularly under cluttered conditions, indicating improved robustness.
  • B. Ablation Studies: Test-time computation combines RCS candidate selection with LAR, which uses ACVS rollouts to rectify unreliable actions.Unless otherwise specified, the strategy uses four action proposals per decision step.
  • B. Ablation Studies: 0.43 to 0.50 average success rate results from adding RCS, while enabling LAR further raises it to 0.60 under the single-attempt setting.The results indicate that candidate selection and future-conditioned rectification address distinct execution failures.
  • B. Ablation Studies: RCS+LAR consistently achieves the best performance across tasks compared with CFG and ACG generation-time guidance.The method evaluates candidate actions and their induced futures before execution, with larger improvement on Pen→Box.
  • VIII. CONCLUSION AND FUTURE WORK: τ0-WM unifies action generation, future prediction, and deployment-time reasoning within a single predictive framework.Its training uses heterogeneous interaction data, including robot teleoperation, UMI-style demonstrations, egocentric videos, and rollout trajectories.
  • VIII. CONCLUSION AND FUTURE WORK: Future work includes tactile sensing for contact-rich manipulation, more reliable deployment-time reasoning, and predictive modeling over longer horizons.Suggested reasoning improvements include better uncertainty estimation, longer-horizon evaluation, and more effective search strategies.
  • VIII. CONCLUSION AND FUTURE WORK: Predictive robot learning is presented as a promising path toward more capable and reliable robot foundation models, with τ0-WM as a practical step.Experiments demonstrate strong policy performance on long-horizon, fine-grained manipulation tasks.

APPENDIX … 1) Cross-Attention KV Cache:

The appendix details τ0-WM’s two-stage training and deployment setup, including heterogeneous pre-training data, standardized optimization and execution settings, and inference-efficiency optimizations. These optimizations include cached cross-attention key-value tensors and text representations that reduce redundant computation and latency.

  • A. Training and Deploymeny Details: τ0-WM uses two training stages: pre-training on 27.3K hours of heterogeneous interaction data followed by downstream-task post-training.The pre-training data include real-robot teleoperation, UMI-style interactions, egocentric human videos, and rollout or failure trajectories.
  • 1) Training Configuration:: Both training stages use AdamW with a 5 × 10−5 learning rate, while global batch sizes are 12,288 and 384 for pre-training and post-training.The same optimization hyperparameters are used across embodiments and tasks unless otherwise specified.
  • 2) Deployment Details:: Real-robot experiments use language-conditioned multi-view observations and receding-horizon closed-loop execution with fixed-length action chunks of 30.These settings apply unless otherwise stated.
  • 2) Deployment Details:: 220 ms is the standard end-to-end action-generation latency on a single RTX 5090 GPU, reduced to 180 ms by caching reusable text representations without changing outputs.The reduced-latency configuration preserves model outputs.
  • B. Inference Acceleration: Inference acceleration uses implementation-level optimizations to improve deployment efficiency.The appendix introduces these optimizations before describing cross-attention key-value caching.
  • 1) Cross-Attention KV Cache:: During action-only inference, video features condition the action branch through cross-attention, whose key and value tensors are computed once because visual context remains unchanged.The cached tensors are reused across all sampling steps, avoiding repeated computation.
  • 1) Cross-Attention KV Cache:: The cached key and value tensors are reused throughout inference, eliminating redundant projection operations.The passage defines x(l) as the video feature at transformer layer l before describing the cached tensors.

2) Fused QKV Projection: · 3) Simplified Rotary Position Embedding: · 4) Torch Compile Optimization:

The deployment stack combines fused QKV projection, precomputed rotary embeddings, and per-block torch.compile to reduce attention and positional-encoding overhead. Together, these optimizations reduce deployment latency from approximately 180 ms to 140 ms, while compiler transformations may introduce small numerical differences.

  • 2) Fused QKV Projection:: Query, key, and value projections are fused into a single matrix multiplication.This replaces three independent projection layers.
  • 2) Fused QKV Projection:: Fused QKV projection reduces kernel launch overhead.The reduction comes from combining the three projections into one matrix multiplication.
  • 2) Fused QKV Projection:: Fused QKV projection improves memory throughput compared with three independent projection layers.
  • 3) Simplified Rotary Position Embedding:: One-dimensional rotary position embeddings are precomputed for the temporal sequence of action tokens.The precomputed embeddings are reused throughout deployment.
  • 3) Simplified Rotary Position Embedding:: Reusing one-dimensional rotary embeddings avoids repeated frequency construction.This reduces positional encoding overhead during deployment.
  • 4) Torch Compile Optimization:: Per-block torch.compile further reduces latency when combined with the preceding optimizations.
  • 4) Torch Compile Optimization:: 180 ms to 140 ms: deployment latency is reduced by the combined optimizations.The reported latency change is approximate.
  • 4) Torch Compile Optimization:: Compiler-level graph transformations and kernel fusion may introduce small numerical differences compared with eager execution.For diffusion-based models, these differences can occasionally propagate through the sampling process.
Loading 2606.01027v1…