Source-linked AI summary
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
Tianxing Chen, Yue Chen, Zixuan Li, Junyuan Tang, Kailun Su, Haoran Lu, Weijie Wan, Baijun Chen, Songling Liu, Haowen Yan, Honghao Su, Zhiyang Dou, Kaixuan Wang, Dandan Zhang, Yunze Liu, Yan Qin, Qiwei Liang, Qiwei Wu, Zijian Lin, Wenwei Lin, Yuran Wang, Minghua He, Tianshu Wu, Ruihai Wu, Jingquan Zhou, Kai-Chong Lei, Haibao Yu, Yuanfeng Ji, Weiyang Jin, Guanyu Lin, Xiaofan Li, Qi Xiong, Renjing Xu, Zhongyu Li, Wenhao Chai, Enze Xie, Ziwei Wang, Yao Mu, Hao Dong, Wojciech Matusik, Mingyu Ding, Wenbo Ding, Ping Luo, Masayoshi Tomizuka
TL;DR
Existing robot-manipulation benchmarks provide limited capability coverage and usually separate scalable simulation from costly, hard-to-reproduce real-world evaluation. RoboDojo unifies 42 simulation and 18 real-world tasks to evaluate generalist policies across complementary capabilities and deployment conditions, revealing that current policies remain insufficiently comprehensive and robust for challenging manipulation.
Problem
Existing benchmarks provide limited capability coverage and often separate scalable simulation from costly, difficult-to-reproduce real-world evaluation.
Method
RoboDojo unifies 42 simulation and 18 real-world tasks across five capability dimensions with scalable simulation and standardized physical-world evaluation.
Results
Evaluated policies show insufficient capability coverage across key dimensions and poor performance on complex real-world manipulation tasks.
Takeaways & Limitations
RoboDojo provides a reproducible platform for diagnosing generalist manipulation progress across diverse simulation capabilities and challenging physical deployment conditions.
Takeaways & Limitations
Real-world task outcomes can exhibit substantial variance when physical trials are limited, especially for contact-rich or multi-stage execution.
Abstract
from arXiv · showhide
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited capability coverage, and are often conducted only in simulation or only in the real world. Simulation enables scalable feedback but misses physical deployment challenges, while real-world evaluation is costly, time-consuming, and difficult to reproduce. We introduce RoboDojo, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies. RoboDojo includes 42 simulation tasks and 18 real-world tasks covering diverse and complementary manipulation capabilities. The simulation benchmark evaluates five dimensions: generalization, memory, precision, long-horizon execution, and open-vocabulary instruction following, while the real-world benchmark exposes policies to challenging physical-world deployment conditions. RoboDojo supports scalable evaluation through heterogeneous parallel simulation in Isaac Sim and provides RoboDojo-RealEval, a reproducible real-world evaluation system with remote cloud access, standardized hardware, scene reset, evaluation protocol, and deployment interface. Together with XPolicyLab, policies can be integrated once and evaluated across simulation and real-world settings with minimal adaptation. We integrate 30 policies into XPolicyLab and evaluate them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance. The website is available at http://robodojo-benchmark.com/.
5 Dimensions Real-World
RoboDojo evaluates generalist robot manipulation across five simulation dimensions and reproducible real-world testing. Its unified framework combines broad task coverage, parallel simulation, RoboDojo-RealEval, XPolicyLab, and leaderboard governance.
- The simulation evaluation covers generalization, memory, long-horizon execution, precision, and open-vocabulary instruction following.
- The real-world evaluation standardizes layout, lighting, robot setup, and hardware to support reproducibility.
- RoboDojo unifies simulation evaluation with reproducible real-world testing for generalist robot manipulation.The framework is described as comprehensive, fast, and high-quality while addressing real-world challenges.
- 42 simulation tasks and 18 real-world tasks provide broad benchmark coverage alongside heterogeneous parallel simulation and RoboDojo-RealEval.XPolicyLab supports policy integration across settings, and the leaderboard is continuously updated.
- AI MMLab Club maintains the leaderboard, with global academic partners co-governing and operating it without commercial funding or sponsorship.
1. Introduction
RoboDojo addresses fragmented and incomplete evaluation of generalist robot manipulation policies with a unified, reproducible sim-and-real benchmark. It combines broad capability coverage, standardized physical testing, shared policy infrastructure, and systematic evaluation of existing policies.
- Motivation and limitations: Simulation benchmarks are scalable for rapid feedback but miss contact-rich dynamics, actuation errors, perception noise, and other deployment factors, while real-world evaluation is costly and time-consuming.The introduction identifies the separation between simulation and real-world evaluation as a key limitation.
- RoboDojo benchmark: RoboDojo includes 42 simulation tasks and 18 real-world tasks for efficient, comprehensive, and reproducible evaluation of generalist robot manipulation policies.Its simulation benchmark covers Generalization, Memory, Long-Horizon, Precision, and Open, with heterogeneous parallel evaluation in Isaac Sim.
- Real-world evaluation: RoboDojo-RealEval standardizes hardware, workspace layout, lighting, scene reset, evaluation protocol, and deployment interface for remote, reproducible physical testing.The system complements simulation by exposing policies to challenging physical-world conditions.
- RoboDojo benchmark: The benchmark evaluates simulation capabilities and challenging physical deployment conditions across three robot embodiments.Real-world tasks include contact-rich interactions, perception noise, actuation errors, and environmental variations.
- Unified infrastructure: XPolicyLab integrates 30 robot manipulation models into a shared framework, enabling cross-benchmark development, deployment, and evaluation with minimal policy-side adaptation.The unified system uses a shared policy interface and evaluation pipeline across simulation and real-world testing.
- Evaluation and analysis: Evaluations on RoboDojo establish a public leaderboard and systematically analyze current policy limitations across simulation and real-world settings.RoboDojo is intended to support diagnosis, policy improvement, and analysis of transfer to physical deployment conditions.
2. Related Work
Robot manipulation research has progressed from classical imitation learning to generalist policies, while existing benchmarks evaluate specialized capabilities or settings. RoboDojo addresses the resulting need for structured evaluation across complementary capabilities and deployment conditions.
- Policy Evolution: Robot manipulation policies have evolved from classical imitation learning toward large-scale generalist robot policies and vision-language-action foundation models.Imitation learning methods achieve strong visuomotor manipulation by learning from expert demonstrations.
- Evaluation Motivation: Existing policies remain vulnerable beyond controlled task distributions because of limited temporal reasoning, spatial precision, generalization, and robustness to physical-world variation.These limitations motivate evaluation protocols that diagnose complementary manipulation capabilities rather than reporting aggregate success alone.
- Simulation Benchmarks: Simulation benchmarks cover language-conditioned long-horizon manipulation, lifelong learning, bimanual generalization, memory-dependent manipulation, garment manipulation, and sim-to-real correlation.Examples include CALVIN, LIBERO, RoboTwin, RMBench, GarmentLab, and SimplerEnv.
- Benchmark Limitations: Existing benchmarks are often specialized by task family, capability axis, or evaluation environment, limiting comprehensive assessment of manipulation policies.The cited benchmark landscape spans both simulation and real-world evaluation, but no single benchmark in these passages covers all described dimensions.
- Simulation and Real-World Evaluation: Simulation-only benchmarks enable efficient, controlled stress testing but miss physical deployment uncertainties, whereas real-world benchmarks provide direct evidence yet remain difficult to reproduce.Reproducibility challenges include hardware setups and scene reset procedures.
3. RoboDojo Benchmark
RoboDojo is a unified benchmark with 42 simulation tasks and 18 real-world tasks for evaluating generalist robot manipulation policies across five capability dimensions and diverse physical deployment conditions. Its multi-embodiment real-world benchmark tests reliability across different robot platforms, kinematics, workspaces, camera placements, and data distributions.
- Benchmark Overview: RoboDojo comprises 42 simulation tasks and 18 real-world tasks organized to evaluate generalist robot manipulation policies across simulation and physical deployment.The simulation benchmark targets capability-oriented diagnosis, while the real-world benchmark evaluates challenging deployment conditions.
- Simulation Benchmark: The 42 simulation tasks span Generalization, Memory, Long-Horizon, Precision, and Open, enabling rapid policy feedback beyond narrow skill variations.The tasks are designed on the ARX X5 bimanual platform with arm bases separated by 0.6 m and are supported by heterogeneous parallel simulation.
- Simulation Benchmark: The simulation dimensions evaluate robustness to unseen scenes, partial observability, dependent action sequences, fine-grained control, and language-conditioned transfer.They include 12 Generalization tasks, 6 Memory tasks, 8 Long-Horizon tasks, 8 Precision tasks, and 8 Open tasks.
- Evaluation Protocol: Policies are evaluated on all 42 simulation tasks for 50 episodes each, totaling 2,100 episodes, with success rate and average score reported separately.Success rate measures binary task completion, while average score captures partial task progress.
- Real-World Benchmark: The 18-task real-world benchmark covers ARX X5, Piper, and Piper X to test reliability across robot embodiments and deployment conditions difficult to fully capture in simulation.These conditions include contact-rich interactions, perception noise, actuation errors, workspace constraints, and temporal variations; RoboDojo-RealEval provides open-source hardware and evaluation specifications for reproducible deployment.
4. Technical Implementation
RoboDojo’s technical implementation combines configurable, physically grounded simulation with heterogeneous parallel execution, reproducible real-world evaluation, and XPolicyLab for unified policy integration. Together, these components support diverse task construction, scalable evaluation, data collection, and consistent deployment across simulation and physical settings.
- Platform Overview: RoboDojo comprises a simulation platform, a real-world evaluation platform, and XPolicyLab for unified policy development and deployment.The simulation platform supports diverse task construction, digital-twin asset generation, heterogeneous parallel evaluation, and simulation data collection.
- Simulation Platform: The Isaac Sim and Isaac Lab simulation platform uses configuration-driven task construction, physically grounded digital assets, and heterogeneous parallel execution.Tasks are instantiated from modular YAML specifications defining assets, layouts, initialization distributions, randomization ranges, and success conditions.
- Simulation Platform: Heterogeneous parallel simulation independently samples scene configurations while sharing a vectorized interface, varying objects, distractors, articulations, and task layouts across environments.This preserves scene-level diversity while improving evaluation speed compared with homogeneous cloned environments.
- Data Collection: RoboDojo collects data through automated trajectory synthesis and VR teleoperation, using shared affordance annotations for skill grounding and task validation.Automated demonstrations compose reusable skills including grasp, place, handover, insert, open, close, stack, and push_up with the cuRobo v2 planner.
- Real-World Evaluation: RoboDojo-RealEval standardizes hardware and software to improve reproducibility and fairness in real-world evaluation despite sensitivity to lighting, placement, camera pose, surfaces, and scene resets.The platform fixes relative robot-arm and camera poses and controls other evaluation conditions.
- XPolicyLab: XPolicyLab standardizes policy data, preprocessing, training, action, and runtime workflows, integrating 30 models for consistent training, deployment, and comparison across RoboDojo environments.A standardized observation-action interface connects policy servers to simulation and RoboDojo-RealEval, enabling the same implementation to move from simulation to remote physical evaluation.
5. RoboDojo Leaderboard
RoboDojo establishes standardized simulation and real-world leaderboards for comparing robot manipulation policies across broad capabilities and physical deployment conditions. The evaluation integrates 30 simulation policies and 10 real-world policies using repeated trials, expert teleoperation references, and standardized protocols.
- Leaderboard scope: RoboDojo provides complementary simulation and real-world leaderboards for standardized policy comparisons across broad capabilities and physical deployment conditions.The simulation leaderboard supports repeated trials across five capability dimensions, while the real-world leaderboard evaluates standardized deployment across multiple robot embodiments.
- Simulation leaderboard: 30 representative robot manipulation policies are integrated, trained, and evaluated on the RoboDojo Simulation Benchmark Leaderboard.Three expert teleoperators with more than 1,000 hours of simulation teleoperation experience provide a human-level reference under matching success criteria and execution horizons.
- Simulation leaderboard: 150 trials per task are used for most simulation policies across three seeds, with results reported as mean success rate, mean score, and corresponding standard deviations.Generalization trials are split evenly between standard and randomized settings, with 50 trials per seed per task.
- Real-world leaderboard: 10 representative robot manipulation policies are evaluated on the RoboDojo Real-World Benchmark Leaderboard.Expert teleoperators with more than 1,500 hours of real-world robot teleoperation experience provide a human-level reference using the same leader-follower setup and task constraints.
- Real-world leaderboard: 18 tasks across three embodiments and 10 trials per task produce 180 real-world evaluation trials per policy under standardized scene reset, evaluation, and deployment procedures.Each policy uses one random seed for each robot embodiment, and both task success rate and score are reported.
6. Experiments
RoboDojo experiments show that current policies make partial progress on structured manipulation but remain weak in precision, memory-conditioned execution, open-semantic tasks, and reliable real-world deployment. Simulation and real-world results expose complementary limitations, including incomplete task execution, unstable control, and safety-critical behaviors.
- Simulation benchmark: Leading simulation policies remain clustered in a low-score regime, revealing a substantial gap toward balanced generalist manipulation.
- Simulation benchmark: 14.92% success rate and 25.74 score make Hy-Embodied-0.5-VLA the strongest Long-Horizon policy, ahead of π0.5 and Spatial Forcing.π0.5 achieves a 14.67% success rate and 23.54 score, while Spatial Forcing reaches 14.58% and 23.26.
- Simulation benchmark: 12.00% success rate makes X-VLA the best Precision policy, showing that global task execution does not automatically produce accurate local control.Precision failures often reflect open-loop execution, weak state-conditioned correction, action jitter, and missing verification of stage preconditions.
- Simulation benchmark: 12.11% success rate and 13.37 score make Hy-Embodied-0.5-VLA the strongest Memory policy, while explicit and implicit memory mechanisms remain unreliable.EventVLA reaches a 4.78% success rate and 4.92 score; X-WAM achieves a 4.67% success rate and 6.32 score.
- Simulation benchmark: 1.67% success rate and 1.98 score are the best Open results, indicating that open-semantic manipulation and semantic-to-action grounding remain largely unsolved.Policies struggle to align open-ended instructions, visual affordances, executable actions, relevant objects, functional relations, and physical constraints.
- Real-world evaluation: 12.8% overall success rate and 22.9 score make π0.5 the strongest policy across 18 real-world tasks, far below human teleoperation’s 100.0% success rate and 100.0 score.Higher scores than success rates indicate partial progress without completion; deployment also reveals alignment, contact, transition, jitter, oscillation, instability, and safety failures.
7. Future Extensions
RoboDojo is designed as an extensible benchmarking platform that will expand beyond its current task collection toward broader manipulation scenarios, embodiments, and evaluation settings. Planned directions include dexterous hand, humanoid whole-body, tactile, and mobile manipulation, with multiple robot embodiments supported for each.
- Platform expansion: RoboDojo will continuously expand toward broader manipulation scenarios, embodiments, and evaluation settings rather than remain a fixed task collection.The benchmark is explicitly designed as an extensible platform for future releases.
- Future benchmark directions: Planned larger-scale benchmarks will cover dexterous hand manipulation, humanoid whole-body manipulation, tactile manipulation, and mobile manipulation.These four directions are identified as future extensions in the paper and Figure 9.
- Embodiment support: Each future direction will support multiple robot embodiments to improve accessibility and enable broader evaluation.The passage specifies multi-embodiment support across the planned extensions.
8. Conclusion · Appendix · A.1. Contributors
RoboDojo provides a unified 42-task simulation and 18-task real-world benchmark, and its evaluation shows current robot policies remain unreliable across diverse manipulation demands and physical deployment. The appendix acknowledges contributors to the RoboDojo-RealEval infrastructure and policy integration and evaluation.
- 8. Conclusion: 42 simulation tasks and 18 real-world tasks cover diverse, long-horizon, precision-demanding, memory-dependent, and open-ended manipulation scenarios.The benchmark evaluates generalist robot manipulation policies across complementary simulation and real-world settings.
- 8. Conclusion: Heterogeneous parallel execution enables efficient large-scale simulation evaluation with rapid feedback and fine-grained capability diagnosis.This supports scalable assessment of policy capabilities in simulation.
- 8. Conclusion: Current robot policies remain far from reliable general-purpose manipulation according to the reported experiments.The conclusion identifies substantial limitations in present policy reliability.
- 8. Conclusion: RoboDojo exposes insufficient coverage in generalization, long-horizon execution, precise manipulation, memory, and open-ended task understanding, while real-world policies perform poorly on complex tasks.These findings indicate that current models are not yet comprehensive across diverse requirements or robust to challenging physical deployment conditions.
- A.1. Contributors: RoboDojo-RealEval infrastructure contributors supported development of the real-world evaluation infrastructure.The acknowledged contributors are Tian Nian, Zijian Cai, Kehe Ye, Yukun Liao, Shaolong Zhu, Qiangyu Chen, Jiahao Zhang, and Zichun Chen.
- A.1. Contributors: Policy contributors provided policy models and supported their integration and evaluation in RoboDojo.The acknowledgment lists Jun Guo, Zongzheng Zhang, Hongzhe Bi, Jisong Cai, Xiaofeng Wang, Zheng Zhu, Yuhang Tang, Weijie Ke, Mingleyang Li, Ganlin Yang, Shuai Yang, Hengtao Li, Wenxuan Song, Zhangzheng Tu, Kaidong Zhang, Yu Sun, Shuhe Huang, Junliang Guo, Tong Zhang, Yixing Chen, Pengxiang Ding, Rongxu Cui, and Hengkai Tan.
A.2. Evaluation Integrity and Leaderboard Publication Details
RoboDojo requires official remote evaluation through its standardized sim-and-real pipeline, with scores computed by the evaluation system. Simulation uses repeated seeded runs, while hidden verification and artifact release support leaderboard integrity and reproducibility.
- Official remote evaluation: Official scores are computed through RoboDojo’s online evaluation system using released layouts, standardized scene resets, the evaluation protocol, and deployment interface.Policies connect through a deployable package or remote policy server.
- Repeated simulation evaluation: Simulation policies are evaluated under three random seeds, with RoboDojo reporting the mean and standard deviation across runs.Participants may submit three training-seed checkpoints or one checkpoint evaluated under three evaluation seeds.
- Hidden verification: Hidden verification uses randomized layouts to detect overfitting, hand-tuning, or gaming, invalidating submissions whose hidden and public success rates differ significantly.Invalid submissions are excluded from the official verified leaderboard.
- Open-source artifact for verified publication: Publishing on the official verified leaderboard requires releasing the full evaluation artifact through XPolicyLab.The artifact includes training and deployment code, the evaluated checkpoint, configuration files, and model loading, inference, deployment, and evaluation instructions; this applies only at publication, not private evaluation.
B. Comparison with Existing Benchmarks
RoboDojo is distinguished from existing robot manipulation benchmarks by its unified sim-and-real design, broad capability coverage, task diversity, reproducibility, and remote evaluation support. It connects rapid simulation-based diagnosis with standardized real-world evaluation rather than specializing in only one setting.
- RoboDojo unifies simulation and real-world evaluation, connecting rapid simulation-based diagnosis with standardized real-world evaluation.
- Compared with prior benchmarks, RoboDojo emphasizes broader capability coverage and task diversity.
- RoboDojo additionally supports real-world reproducibility and remote evaluation.
C. Real-World Evaluation Stability · D. Simulation Training and Evaluation Details · D.1. Simulation Training Data Details
The merged sections specify reproducible real-world stability evaluation and a simulation-data pipeline spanning large-scale bimanual demonstrations, multimodal sensing, task-dependent collection, and controlled visual diversification.
- C. Real-World Evaluation Stability: Three independent RoboDojo-RealEval runs evaluate π0.5, GalaxeaVLA, and InternVLA-A1 for stability using success-rate and score standard deviations across R1–R3.Success-rate deviations are reported in percentage points, while the Overall row measures standard deviation of average performance across runs.
- C. Real-World Evaluation Stability: RoboDojo is compared with other benchmarks using Skill, Num. Policies, and HP, with N/R indicating unreported values and HP denoting heterogeneous parallelization.Skill counts operation primitives, whereas Num. Policies counts integrated or evaluated baseline policies.
- D. Simulation Training and Evaluation Details: The simulation training set contains 35 task directories, 3,500 trajectories, and 1,859,602 frames, totaling 20.66 hours of bimanual manipulation data recorded at 25 Hz.Each trajectory includes synchronized RGB-D observations from one head-mounted and two wrist cameras at 640 × 480 resolution.
- D.1. Simulation Training Data Details: Depth images are normalized and stored as integer millimeter values, and a third-view RGB video is also recorded for each demonstration.The passage specifies the storage format and additional third-view recording as part of the simulation training data.
- D.1. Simulation Training Data Details: Demonstrations use automated trajectory synthesis or VR-based teleoperation, with 100 trajectories per task for Generalization, Memory, Long-Horizon, and Precision.The Open dimension is evaluation-only and has no task-specific training demonstrations because it tests recombination and transfer to unseen specifications.
- D.1. Simulation Training Data Details: Generalization training uses normal settings plus one auxiliary DLC task with 100 domain-randomized trajectories covering backgrounds, lighting, and clutter layouts.The auxiliary data broadens visual exposure without leaking task-specific solutions from evaluation tasks.
D.2. Simulation Evaluation Details · E. Real-World Training and Evaluation Details
The paper sets simulation horizons from task-specific demonstration lengths and reports simulation training-data statistics, while also documenting repeated-run real-world evaluation stability. Simulation training excludes the Open dimension to assess skill recombination and transfer to unseen task specifications.
- D.2. Simulation Evaluation Details: Simulation horizons use 1.2× the 90th-percentile task-specific trajectory length.The 90th percentile is computed from demonstration trajectories.
- D.2. Simulation Evaluation Details: Short automated-trajectory tasks instead use a 1.5× horizon multiplier.The larger multiplier reduces premature termination from small trajectory-length variations.
- D.2. Simulation Evaluation Details: The horizon setting balances sufficient execution time with consistent evaluation horizons.This rationale applies to the task-specific horizon construction.
- E. Real-World Training and Evaluation Details: Table 9 reports full per-task real-world evaluation stability across repeated runs.The supplied passage identifies the table’s scope but provides no per-task values.
- D.2. Simulation Evaluation Details: Simulation training contains 1,859,602 frames from 3,500 trajectories.The data correspond to 20.66 hours of bimanual manipulation recorded at 25 Hz.
- D.2. Simulation Evaluation Details: The Open dimension is excluded from training and reserved for evaluating skill recombination and transfer to unseen task specifications.DLC is an auxiliary domain-randomized task used to broaden visual exposure during training.
E.1. Real-World Training Data Details … J.3. Piper
RoboDojo combines standardized real-world demonstrations and evaluation with a configurable, heterogeneous simulation platform, unified policy interfaces, and documented task suites. Its infrastructure supports reproducible data collection, deployment, scalable evaluation, and diverse manipulation scenarios across simulation and real-world settings.
- E.1. Real-World Training Data Details: 1,800 trajectories and 1,611,841 frames comprise the real-world dataset across ARX X5, Piper, and Piper X embodiments.Each task contributes 100 demonstrations from four operators; Table 11 reports 17.91 hours of bimanual data at 25 Hz, with 600 trajectories per embodiment across six tasks.
- F. Simulation Platform Details: The simulation platform builds on MagicSim and uses configuration-first task definitions with deterministic seeded scene generation and unified support for rigid, articulated, and deformable assets.Configurations specify assets, initialization, cameras, lighting, textures, articulation states, success conditions, and evaluation seeds, separating task specification from simulator execution.
- F.3. Digital Asset Processing and Validation: RoboDojo validates a metadata-rich digital asset library through simulation rollouts to support task construction, automated demonstrations, and physically plausible interactions.Assets include semantic descriptions, placement and success annotations, manipulation affordances, collision geometry, joint properties, material parameters, and realistic textures.
- F.4. Heterogeneous Parallelism Implementation: Heterogeneous parallelism preserves vectorized execution while allowing environments to vary in objects, geometries, clutter, seeds, robot initialization, observations, and success conditions.Multiple GPU processes shard episode seeds and aggregate results, avoiding cloned scene templates during large-scale evaluation and data collection.
- F.5. Simulation Demonstration Collection: Simulation demonstrations come from automated skill synthesis or VR teleoperation, using shared affordance annotations and cuRobo v2 for physically feasible motion generation.The automated library includes push_up, place, handover, grasp, insert, open, close, and stack; VR control maps 6D controller-pose deltas to end-effector targets, with approximately 3 ms planning latency for a single ARX X5 arm.
- G. RoboDojo-RealEval Platform Details: RoboDojo-RealEval standardizes hardware, scene replay, safety controls, cloud video collection, and scoring through fixed workspaces, controlled lighting, cameras, interchangeable bimanual robots, and touchscreen operation.The platform includes ARX X5, Piper, and Piper X embodiments, while reference layouts reduce reset variation and emergency stops return the robot to a safe reset state.
- H. XPolicyLab Design Details: XPolicyLab unifies data formats, policy interfaces, deployment packages, communication, and evaluation loops so heterogeneous policies can run across single real-world and batched simulation settings.It supports diffusion, VLA, world-model, and classical imitation-learning policies through standardized methods and WebSocket/MessagePack communication, while task documentation specifies instructions, demonstrations, and train/test usage.
K. Policy Training Details
The section specifies each evaluated policy’s initialization checkpoint, training data or configuration, batch size, and optimization schedule for simulation and real-world evaluation. Training setups vary substantially across policies, including distinct simulation/real-world step counts and specialized memory or multistage optimization settings.
- Hy-Embodied-0.5-VLA: Hy-Embodied-0.5-VLA uses its foundation checkpoint, batch size 128, 200K simulation steps, and 6 historical images sampled every 20 steps for memory encoding.The simulation setup follows the official RoboTwin 2.0 evaluation setting.
- Spatial Forcing: Spatial Forcing uses the pi05_base foundation checkpoint with batch size 256 and 60K simulation steps.The checkpoint is hosted at gs://openpi-assets/checkpoints/pi05_base.
- Simulation and real-world schedules: π0.5, X-VLA, and Xiaomi-Robotics-0 use batch size 256, with simulation schedules of 60K, 100K, and 100K steps and real-world schedules of 30K, 50K, and 50K steps, respectively.π0.5 uses pi05_base; X-VLA uses 2toInf/X-VLA-Pt; Xiaomi-Robotics-0 uses Xiaomi-Robotics-0-Pretrain.
- Multistage and simulation schedules: X-WAM uses batch size 32 for 40K simulation steps, while GigaWorld-Policy uses batch size 32 with 50K video optimization and 250K action optimization steps.GigaWorld-Policy totals 300K optimization steps across 3 epochs.
- Additional policy schedules: StarVLA-α, GalaxeaVLA, LingBot-VLA, EventVLA, Fast-WAM, and AHA-WAM use simulation schedules of 100K steps, 4 epochs, 15K steps, 150K steps, 20K steps, and 25K steps, respectively.Their batch sizes are 128, 32, 256, 128, 256, and 768, respectively.
- Additional policy schedules: The remaining policies use varied simulation configurations, including ACT trained from scratch for 6K steps and GO-1 trained for 4 epochs corresponding to 77,484 optimization steps.Other listed schedules range from 30K to 300K steps, with some real-world schedules specified separately.