Robotics

Papers filed under cs.RO on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,941 to 3,000 of 3,248

  1. Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

    Jinghui Lu, Jiayi Guan, Zhijian Huang +47

    cs.CVcs.CLcs.ROarXiv:2604.18486v32026
  2. Rethinking Demonstration Unlearning in Imitation Learning for Robotics

    Jiazhuo Li, Yu Zhang, Yiming Fei +4

    cs.ROcs.LGarXiv:2608.20784v12026
  3. Cortex 2.0: Grounding World Models in Real-World Industrial Deployment

    Adriana Aida, Walid Amer, Katarina Bankovic +25

    cs.ROcs.AIarXiv:2604.20246v12026
  4. UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling

    Boyu Chen, Yi Chen, Lu Qiu +3

    cs.ROcs.AIarXiv:2604.19734v12026
  5. dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model

    Yaxuan Li, Zhongyi Zhou, Yefei Chen +2

    cs.ROarXiv:2604.22152v12026
  6. OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism

    Xiangyu Li, Huaizhi Tang, Xin Ding +3

    cs.ROcs.AIarXiv:2603.14371v22026
  7. SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control

    Ruihua Han, Rui Gao, Zhe Liu +6

    cs.ROcs.AIarXiv:2608.21175v12026
  8. Neural-Primitive: An Efficient End-to-end Local Planner with Primitive-based Imitation Learning for Autonomous Flight

    Zhitao Liu, Guangtong Xu, Zihan Wang +3

    cs.ROcs.AIarXiv:2608.20948v12026
  9. Anatomy-Informed Neural Networks: Encoding Anatomic Priors in Loss and Architecture, with an SE(3) Formulation of Guidewire-Induced Aortoiliac Deformation

    David P. Stonko

    cs.AIcs.CVcs.ROarXiv:2608.21332v12026
  10. Unified Branch-and-Bound Search for the Steiner Traveling Salesman Problem on Graphs of Convex Sets

    Jingtao Tang, Hang Ma

    cs.AIcs.ROarXiv:2608.21319v12026
  11. Unsupervised Learning for Physical Interaction through Video Prediction

    Chelsea Finn, Ian Goodfellow, Sergey Levine

    cs.LGcs.AIcs.CVarXiv:1605.07157v42016
  12. Graph-Operator World Models for Morphology-Parameter Generalization in Continuous Control

    Xu Yang, Yiqin Yang, Qianchuan Zhao

    cs.AIcs.ROarXiv:2608.20936v12026
  13. ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

    Siyuan Ma, Yutian Zhang, Boshi Zhang +4

    cs.AIcs.ROarXiv:2608.20735v12026
  14. Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting

    Benjamin Wilson, William Qi, Tanmay Agarwal +10

    cs.CVcs.AIcs.LGarXiv:2301.00493v12023
  15. A Generalist Agent

    Scott Reed, Konrad Zolna, Emilio Parisotto +17

    cs.AIcs.CLcs.LGarXiv:2205.06175v32022
  16. TrackingNet: A Large-Scale Dataset and Benchmark for Object Tracking in the Wild

    Matthias Müller, Adel Bibi, Silvio Giancola +2

    cs.CVcs.ROarXiv:1803.10794v12018
  17. Learning robust perceptive locomotion for quadrupedal robots in the wild

    Takahiro Miki, Joonho Lee, Jemin Hwangbo +3

    cs.ROarXiv:2201.08117v12022
  18. Open X-Embodiment: Robotic Learning Datasets and RT-X Models

    Open X-Embodiment Collaboration, Abby O'Neill, Abdul Rehman +291

    cs.ROarXiv:2310.08864v92023
  19. DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries

    Yue Wang, Vitor Guizilini, Tianyuan Zhang +3

    cs.CVcs.AIcs.LGarXiv:2110.06922v12021
  20. Robots that can adapt like animals

    Antoine Cully, Jeff Clune, Danesh Tarapore +1

    cs.ROcs.AIcs.LGarXiv:1407.3501v42014
  21. ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

    Mohit Shridhar, Jesse Thomason, Daniel Gordon +5

    cs.CVcs.AIcs.CLarXiv:1912.01734v22019
  22. Learning to Identify Out-of-Distribution Objects for 3D LiDAR Anomaly Segmentation

    Simone Mosco, Daniel Fusaro, Alberto Pretto

    cs.CVcs.ROarXiv:2604.23604v12026
  23. DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion

    Chen Wang, Danfei Xu, Yuke Zhu +4

    cs.CVcs.ROarXiv:1901.04780v12019
  24. Structural-RNN: Deep Learning on Spatio-Temporal Graphs

    Ashesh Jain, Amir R. Zamir, Silvio Savarese +1

    cs.CVcs.LGcs.NEarXiv:1511.05298v32015
  25. Supersizing Self-supervision: Learning to Grasp from 50K Tries and 700 Robot Hours

    Lerrel Pinto, Abhinav Gupta

    cs.LGcs.CVcs.ROarXiv:1509.06825v12015
  26. Recovering Hidden Reward in Diffusion-Based Policies

    Yanbiao Ji, Qiuchang Li, Yuting Hu +7

    cs.ROarXiv:2605.00623v22026
  27. Learning to Track at 100 FPS with Deep Regression Networks

    David Held, Sebastian Thrun, Silvio Savarese

    cs.CVcs.AIcs.LGarXiv:1604.01802v22016
  28. Diversity is All You Need: Learning Skills without a Reward Function

    Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz +1

    cs.AIcs.ROarXiv:1802.06070v62018
  29. RigidFormer: Learning Rigid Dynamics using Transformers

    Zhiyang Dou, Minghao Guo, Haixu Wu +3

    cs.CVcs.AIcs.GRarXiv:2605.09196v12026
  30. Taskonomy: Disentangling Task Transfer Learning

    Amir Zamir, Alexander Sax, William Shen +3

    cs.CVcs.AIcs.LGarXiv:1804.08328v12018
  31. Zero-Shot Sim-to-Real Robot Learning: A Dexterous Manipulation Study on Reactive Catching

    Kejia Ren, Gaotian Wang, Andrew S. Morgan +1

    cs.ROarXiv:2605.09789v12026
  32. CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models

    Wenxuan Song, Han Zhao, Fuhao Li +7

    cs.CVcs.ROarXiv:2605.10903v12026
  33. Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics

    Jeffrey Mahler, Jacky Liang, Sherdil Niyaz +5

    cs.ROarXiv:1703.09312v32017
  34. Inner Monologue: Embodied Reasoning through Planning with Language Models

    Wenlong Huang, Fei Xia, Ted Xiao +14

    cs.ROcs.AIcs.CLarXiv:2207.05608v12022
  35. Solving Rubik's Cube with a Robot Hand

    OpenAI, Ilge Akkaya, Marcin Andrychowicz +16

    cs.LGcs.AIcs.CVarXiv:1910.07113v12019
  36. A Topology-Aware Spatiotemporal Handover Framework for Continuous Multi-UAV Tracking

    Jianlin Ye, Christos Kyrkou, Panayiotis Kolios

    cs.ROcs.AIarXiv:2605.15779v12026
  37. Planning-oriented Autonomous Driving

    Yihan Hu, Jiazhi Yang, Li Chen +13

    cs.CVcs.ROarXiv:2212.10156v22022
  38. DexHoldem: Playing Texas Hold'em with Dexterous Embodied System

    Feng Chen, Tianzhe Chu, Li Sun +6

    cs.ROcs.AIarXiv:2605.18727v12026
  39. Learning Rich Features from RGB-D Images for Object Detection and Segmentation

    Saurabh Gupta, Ross Girshick, Pablo Arbeláez +1

    cs.CVcs.ROarXiv:1407.5736v12014
  40. Minimalist Visual Inertial Odometry

    Francesco Pasti, Jeremy Klotz, Nicola Bellotto +1

    cs.ROcs.CVcs.LGarXiv:2605.19990v12026
  41. Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models

    Kurtland Chua, Roberto Calandra, Rowan McAllister +1

    cs.LGcs.AIcs.ROarXiv:1805.12114v22018
  42. Deep Reinforcement Learning for Robotic Manipulation with Asynchronous Off-Policy Updates

    Shixiang Gu, Ethan Holly, Timothy Lillicrap +1

    cs.ROcs.AIcs.LGarXiv:1610.00633v22016
  43. Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments

    Jie Jia, Yaofeng Su, Zeyu Bao +4

    cs.ROarXiv:2605.22189v12026
  44. Can Predicted Dynamics Exist in the Physical World?

    Barak Or

    cs.ROcs.AIarXiv:2606.00089v22026
  45. Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems

    Barak Or

    cs.ROcs.AIarXiv:2606.00090v12026
  46. Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view Generation

    Aviral Chharia, Fernando De la Torre

    cs.CVcs.GRcs.ROarXiv:2605.25220v12026
  47. Argoverse: 3D Tracking and Forecasting with Rich Maps

    Ming-Fang Chang, John Lambert, Patsorn Sangkloy +8

    cs.CVcs.ROarXiv:1911.02620v12019
  48. Octo: An Open-Source Generalist Robot Policy

    Octo Model Team, Dibya Ghosh, Homer Walke +16

    cs.ROcs.LGarXiv:2405.12213v22024
  49. Robot Operating System 2: Design, Architecture, and Uses In The Wild

    Steve Macenski, Tully Foote, Brian Gerkey +2

    cs.ROarXiv:2211.07752v12022
  50. Learning Quadrupedal Locomotion over Challenging Terrain

    Joonho Lee, Jemin Hwangbo, Lorenz Wellhausen +2

    cs.ROcs.LGeess.SYarXiv:2010.11251v12020
  51. QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation

    Dmitry Kalashnikov, Alex Irpan, Peter Pastor +8

    cs.LGcs.AIcs.CVarXiv:1806.10293v32018
  52. A Survey of Deep Learning Techniques for Autonomous Driving

    Sorin Grigorescu, Bogdan Trasnea, Tiberiu Cocias +1

    cs.LGcs.ROarXiv:1910.07738v22019
  53. Generation and Comprehension of Unambiguous Object Descriptions

    Junhua Mao, Jonathan Huang, Alexander Toshev +3

    cs.CVcs.CLcs.LGarXiv:1511.02283v32015
  54. Frequency-Guided Action Diffusion via Sub-Frequency Manifold Traversal

    Junlin Wang

    cs.ROcs.LGarXiv:2605.27919v12026
  55. Objaverse: A Universe of Annotated 3D Objects

    Matt Deitke, Dustin Schwenk, Jordi Salvador +7

    cs.CVcs.AIcs.GRarXiv:2212.08051v12022
  56. Code as Policies: Language Model Programs for Embodied Control

    Jacky Liang, Wenlong Huang, Fei Xia +5

    cs.ROarXiv:2209.07753v42022
  57. Sim-to-Real Transfer of Robotic Control with Dynamics Randomization

    Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba +1

    cs.ROeess.SYarXiv:1710.06537v32017
  58. Zero-1-to-3: Zero-shot One Image to 3D Object

    Ruoshi Liu, Rundi Wu, Basile Van Hoorick +3

    cs.CVcs.GRcs.ROarXiv:2303.11328v12023
  59. Benchmarking Deep Reinforcement Learning for Continuous Control

    Yan Duan, Xi Chen, Rein Houthooft +2

    cs.LGcs.AIcs.ROarXiv:1604.06778v32016
  60. Prior Availability in Industrial Visual Sim-to-Real: A Review of CAD-Guided and CAD-Unavailable Regimes

    Chenxi Tao, Seung-Kyum Choi

    cs.CVcs.AIcs.ROarXiv:2605.30581v22026