Source-linked AI summary
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Open X-Embodiment Collaboration, Abby O'Neill, Abdul Rehman, Abhinav Gupta, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew Wang, Andrey Kolobov, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Beomjoon Kim, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu, Charlotte Le, Chelsea Finn, Chen Wang, Chenfeng Xu, Cheng Chi, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Driess, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Foster, Fangchen Liu, Federico Ceola, Fei Xia, Feiyu Zhao, Felipe Vieira Frujeri, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Ge Yan, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guangwen Yang, Guanzhi Wang, Hao Su, Hao-Shu Fang, Haochen Shi, Henghui Bao, Heni Ben Amor, Henrik I Christensen, Hiroki Furuta, Homanga Bharadhwaj, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim, Jaimyn Drake, Jan Peters, Jan Schneider, Jasmine Hsu, Jay Vakil, Jeannette Bohg, Jeffrey Bingham, Jeffrey Wu, Jensen Gao, Jiaheng Hu, Jiajun Wu, Jialin Wu, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jingyun Yang, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kaiyuan Wang, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Ken Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Lin, Kevin Zhang, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Lawrence Yunliang Chen, Lerrel Pinto, Li Fei-Fei, Liam Tan, Linxi "Jim" Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma, Moo Jin Kim, Muhammad Zubair Irshad, Naoaki Kanazawa, Nicklas Hansen, Nicolas Heess, Nikhil J Joshi, Niko Suenderhauf, Ning Liu, Norman Di Palo, Nur Muhammad Mahi Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R Sanketi, Patrick "Tree" Miller, Patrick Yin, Paul Wohlhart, Peng Xu, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Quan Vuong, Rafael Rafailov, Ran Tian, Ria Doshi, Roberto Martín-Martín, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Ruohan Zhang, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante, Sean Kirmani, Sergey Levine, Shan Lin, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham Sonawani, Shubham Tulsiani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vikash Kumar, Vincent Vanhoucke, Vitor Guizilini, Wei Zhan, Wenxuan Zhou, Wolfram Burgard, Xi Chen, Xiangyu Chen, Xiaolong Wang, Xinghao Zhu, Xinyang Geng, Xiyuan Liu, Xu Liangwei, Xuanlin Li, Yansong Pang, Yao Lu, Yecheng Jason Ma, Yejin Kim, Yevgen Chebotar, Yifan Zhou, Yifeng Zhu, Yilin Wu, Ying Xu, Yixuan Wang, Yonatan Bisk, Yongqiang Dou, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yue Cao, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zhuo Xu, Zichen Jeff Cui, Zichen Zhang, Zipeng Fu, Zipeng Lin
TL;DR
Robotic learning commonly trains separate models for each task, robot, and environment, motivating generalist policies trained across embodiments. The paper assembles standardized multi-robot data and trains RT-X models, finding positive transfer and improved generalization across robots, while leaving substantially different modalities and new-robot generalization for future work.
Problem
Robotic learning conventionally trains separate models for each application, robot, and environment, leaving open whether generalist X-robot policies can adapt efficiently across these settings.
Method
The paper pools diverse robotic data, organizes it in an X-embodiment repository, and trains RT-1 and RT-2 Transformer-based policies across 9 robotic manipulators.
Results
RT-X exhibits positive transfer: RT-1-X improves over robot-specific policies, while RT-2-X achieves approximately 3× generalization improvement over a model trained only on the evaluation embodiment.
Takeaways & Limitations
The results demonstrate that multi-robot co-training can improve performance and enable additional capabilities on individual robots by leveraging data from other platforms.
Takeaways & Limitations
The experiments do not cover robots with very different sensing or actuation modalities, do not study generalization to new robots, and lack a criterion for predicting positive transfer.
Abstract
from arXiv · showhide
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train generalist X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. More details can be found on the project website https://robotics-transformer-x.github.io.
I. INTRODUCTION
The paper asks whether diverse multi-robot data can support generalizable robotic policies, rather than separate models for each task, robot, and environment. It evaluates positive transfer and provides shared data, models, and tools for X-embodiment robotic learning.
- Large, diverse pretraining datasets often produce capable general-purpose models that outperform narrowly targeted systems.
- X-embodiment training combines data from multiple robotic platforms to cover more variations in robots and environments.
- The paper evaluates whether multi-robot policies achieve positive transfer and organizes large robotic datasets for future X-embodiment research.
- RT-1 and RT-2 are trained on 9 robotic manipulators, producing RT-X models that improve generalization and capabilities over evaluation-domain policies.
- The Open X-Embodiment Repository contains data from 22 robotic embodiments and 21 institutions, plus open-source tools for research.
II. RELATED WORK
The repository extends prior robot-learning datasets and transfer work by consolidating multi-embodiment data with pretrained RT-X checkpoints and supporting tools. Its dataset spans many robots, labs, trajectories, skills, and objects in a standardized resource.
- Prior transfer methods often address embodiment gaps through shared actions, representation learning, embodiment adaptation, or separated robot and environment representations.
- Unlike most prior open-source datasets, the repository focuses on data spanning multiple robot embodiments rather than robots of one type.
- The Open X-Embodiment Repository provides large-scale data and pretrained model checkpoints for X-embodied robot-learning research.
- The Open X-Embodiment Dataset contains 1M+ robot trajectories from 22 embodiments, pooled from 60 datasets and converted to a consistent RLDS format.
- The project includes RT-X checkpoints for inference and finetuning and is intended as a foundation for community-driven X-embodiment research.
B. Dataset Analysis
The dataset spans 60 datasets across 22 embodiments, with uneven contributions across robots. Its metadata and language annotations reveal diversity in scenes, trajectories, skills, and household objects.
- Franka has the largest diversity of visually distinct scenes because many Franka datasets contribute to the repository.
- Most skills belong to the pick-place family, while the long tail includes wiping and assembling.
- The data covers household objects ranging from appliances to food items and utensils.
IV. RT-X DESIGN
The RT-X design uses high-capacity Transformer-based robotic policies to make productive use of large, heterogeneous X-embodiment datasets. The paper builds its experiments around RT-1 and RT-2 and adapts them to shared cross-robot learning.
- The experiments build on RT-1 and RT-2, two Transformer-based robotic policies adapted to the X-embodiment setting.
- The models require sufficient capacity to productively use large and heterogeneous multi-robot datasets.
- The design is intended to evaluate how X-embodiment training improves policies on individual robots.
A. Data format consolidation
The datasets use a coarsely aligned interface that standardizes visual observations and end-effector actions across robots.
- Each model receives recent images and language instructions, then predicts a 7-dimensional end-effector action vector.The action controls x, y, z, roll, pitch, yaw, and gripper opening or their rates.
- A canonical camera view is selected, resized to a common resolution, and paired with converted 7-DoF end-effector actions.
B. Policy architectures
The paper evaluates RT-1, an efficient Transformer controller, and RT-2, a large vision-language-action model that outputs tokenized robot actions.
- Both architectures take visual input and a natural-language task instruction and output a tokenized action.
- RT-1 is a 35M-parameter Transformer that combines image histories and language through FiLM before decoding tokenized actions.It processes 15 images with an ImageNet-pretrained EfficientNet and a USE language embedding.
- RT-2 is a family of vision-language-action models trained on Internet-scale vision-language data and robotic control data.The paper uses the RT-2-PaLI-X variant with ViT and UL2 components, pretrained primarily on WebLI.
C. Training and inference details
Training uses a robotics mixture spanning multiple manipulators, while inference runs at robot-specific rates; experiments compare transfer and capacity effects.
- The robotics data mixture combines data from 9 manipulators and multiple named robotic-learning datasets.
- Both models use categorical cross-entropy over discrete action buckets for RT-1 or language tokens for RT-2.
- RT-1-X achieves a 50% higher mean success rate than both the Original Method and RT-1 with the same network architecture as RT-1.The caption attributes the increase to co-training on the robotics data mixture.
- RT-1-X underfits large-scale Bridge and RT-1 datasets, whereas the larger RT-2-X obtains strong performance in both scenarios.RT-2-X is co-fine-tuned with an approximately one-to-one split of original VLM data and robotics data.
- RT-1 runs locally and RT-2 is queried through a cloud service at 3–10 Hz, depending on robot requirements.
V. EXPERIMENTAL RESULTS
The experiments test whether X-embodiment training enables positive transfer, improves unseen-task generalization, and benefits from particular model or dataset designs.
- The study evaluates positive transfer, unseen-task generalization, and design effects across 6 different robots.The evaluation comprises 3600 total trials.
- The design analysis examines model size, model architecture, and dataset composition as influences on performance and generalization.
- The experiments use the same robotics data training mixture across the evaluations reported in this section.
A. In-distribution performance across different embodiments
Across embodiment domains, X-embodiment training improves performance and emergent skills, especially when data are limited or the model has sufficient capacity.
- RT-1-X outperforms robot-specific Original Method models on 4 of 5 small-scale datasets.The results show a large average improvement in domains with limited data.
- RT-1-X does not outperform the embodiment-specific RT-1 baseline in large-data domains, indicating underfitting for that model class.
- The larger RT-2-X model outperforms both the Original Method and RT-1 in data-rich domains.This suggests that multi-embodiment training benefits data-rich settings when paired with sufficient model capacity.
- RT-2 and RT-2-X perform roughly on par for unseen objects, backgrounds, and environments.RT-2 already generalizes well along these dimensions because of its VLM backbone.
- RT-2-X outperforms RT-2 by ∼3× on emergent skills involving objects and skills absent from the RT-2 dataset but present in Bridge data.
- Removing the Bridge dataset significantly reduces performance on Google Robot hold-out tasks.The ablation suggests that WidowX data contributes to the additional skills performed by RT-2-X on the Google Robot.
C. Design decisions
Ablations identify model history, web pretraining, and capacity as important design choices for RT-2-X generalization and emergent skills, while fine-tuning strategies perform similarly in this diverse-data setting.
- A short history of images significantly improves generalization performance.
- Web-based pre-training is critical for high performance in the large RT-2-X models.
- The 55B model has a significantly higher Emergent Skills success rate than the 5B model.The comparison demonstrates that higher model capacity enables a higher degree of transfer.
- Co-fine-tuning and fine-tuning have similar performance in both Emergent Skills and Generalization Evaluation.The authors attribute this result to the greater diversity of robotics data used in RT-2-X.
VI. DISCUSSION, FUTURE WORK, AND OPEN PROBLEMS
The paper presents a consolidated multi-robot dataset and demonstrates positive transfer between robots using Transformer-based policies. It also identifies important open problems concerning modality diversity, new-robot generalization, and predicting when transfer occurs.
- 22 robotic embodiments from 21 institutions contribute data demonstrating 527 skills across 160266 tasks.
- Transformer-based policies trained on the consolidated data exhibit significant positive transfer between the different robots.
- Open problems: The experiments do not include robots with very different sensing and actuation modalities.
- Open problems: The study does not evaluate generalization to new robots or provide a decision criterion for when positive transfer occurs.
- Future work: The work presents X-robot learning as feasible and practical while providing tools for future research.