Robotics

Papers filed under cs.RO on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,261 to 1,320 of 3,243

  1. Behind the Curtain: Learning Occluded Shapes for 3D Object Detection

    Qiangeng Xu, Yiqi Zhong, Ulrich Neumann

    cs.CVcs.AIcs.LGarXiv:2112.02205v12021
  2. AdaWorld: Learning Adaptable World Models with Latent Actions

    Shenyuan Gao, Siyuan Zhou, Yilun Du +2

    cs.AIcs.CVcs.LGarXiv:2503.18938v42025
  3. MEM: Multi-Scale Embodied Memory for Vision Language Action Models

    Marcel Torne, Karl Pertsch, Homer Walke +14

    cs.ROcs.LGarXiv:2603.03596v22026
  4. Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

    Hao Luo, Yicheng Feng, Wanpeng Zhang +7

    cs.CVcs.LGcs.ROarXiv:2507.15597v12025
  5. Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

    Zhixuan Liang, Yizhuo Li, Tianshuo Yang +9

    cs.CVcs.LGcs.ROarXiv:2508.20072v42025
  6. Interactive Post-Training for Vision-Language-Action Models

    Shuhan Tan, Kairan Dou, Yue Zhao +1

    cs.LGcs.AIcs.CVarXiv:2505.17016v12025
  7. Real-Time Shape Control of Multi-Segment Soft Robotic Arms Using Koopman Operators with Global and Local Observables

    Jiahe Wang, Eron Ristich, Sultan Haidar Ali +5

    cs.ROeess.SYarXiv:2609.03175v12026
  8. RoboBrain 2.0 Technical Report

    BAAI RoboBrain Team, Mingyu Cao, Huajie Tan +50

    cs.ROarXiv:2507.02029v52025
  9. A Large Scale Event-based Detection Dataset for Automotive

    Pierre de Tournemire, Davide Nitti, Etienne Perot +2

    cs.CVcs.LGcs.ROarXiv:2001.08499v32020
  10. DriveTransformer: Unified Transformer for Scalable End-to-End Autonomous Driving

    Xiaosong Jia, Junqi You, Zhiyuan Zhang +1

    cs.LGcs.CVcs.ROarXiv:2503.07656v22025
  11. Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

    Shaoyuan Xie, Lingdong Kong, Yuhao Dong +5

    cs.CVcs.ROarXiv:2501.04003v12025
  12. HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation

    Yi Li, Yuquan Deng, Jesse Zhang +9

    cs.ROcs.AIcs.CVarXiv:2502.05485v42025
  13. AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control

    Jialong Li, Xuxin Cheng, Tianshu Huang +3

    cs.ROcs.AIcs.LGarXiv:2505.03738v12025
  14. Hardware-Accelerated Instance Segmentation for Resource-Constrained Space Robotics with Criticality Analysis

    Siddhant Shete, Hilmi Dogu Kücüker, Udo Frese +1

    cs.ROcs.ARcs.CVarXiv:2609.02219v12026
  15. Random vector functional link network: recent developments, applications, and future directions

    A. K. Malik, Ruobin Gao, M. A. Ganaie +2

    cs.NEcs.LGcs.ROarXiv:2203.11316v22022
  16. Robotic Ultrasound Imaging: State-of-the-Art and Future Perspectives

    Zhongliang Jiang, Septimiu E. Salcudean, Nassir Navab

    cs.ROarXiv:2307.05545v22023
  17. HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit

    Qingwei Ben, Feiyu Jia, Jia Zeng +3

    cs.ROcs.AIcs.HCarXiv:2502.13013v22025
  18. Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation

    Param Thakkar, Parsika Paresh Shah, Manisha Sushant Gote

    cs.ROcs.AIarXiv:2609.02046v12026
  19. Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

    Gemini Robotics Team, Abbas Abdolmaleki, Saminda Abeyruwan +169

    cs.ROarXiv:2510.03342v32025
  20. Aggressive Quadrotor Flight through Narrow Gaps with Onboard Sensing and Computing using Active Vision

    Davide Falanga, Elias Mueggler, Matthias Faessler +1

    cs.ROarXiv:1612.00291v62016
  21. More than a Million Ways to Be Pushed: A High-Fidelity Experimental Dataset of Planar Pushing

    Kuan-Ting Yu, Maria Bauza, Nima Fazeli +1

    cs.ROarXiv:1604.04038v22016
  22. Zeva: In-Context Causal Learning for Generalizable Embodied Manipulation

    Fu Chen, Xin Ding, Bingjia Huang +8

    cs.ROarXiv:2608.30880v12026
  23. HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

    Jiaming Liu, Hao Chen, Pengju An +12

    cs.CVcs.ROarXiv:2503.10631v32025
  24. JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation

    Shuang Zeng, Dekang Qi, Xinyuan Chang +7

    cs.CVcs.ROarXiv:2509.22548v22025
  25. A Review on Cooperative Adaptive Cruise Control (CACC) Systems: Architectures, Controls, and Applications

    Ziran Wang, Guoyuan Wu, Matthew Barth

    eess.SYcs.ITcs.ROarXiv:1809.02867v12018
  26. AutoTAMP: Autoregressive Task and Motion Planning with LLMs as Translators and Checkers

    Yongchao Chen, Jacob Arkin, Charles Dawson +3

    cs.ROcs.CLcs.HCarXiv:2306.06531v32023
  27. Flow Matching Policy Gradients

    David McAllister, Songwei Ge, Brent Yi +5

    cs.LGcs.ROarXiv:2507.21053v22025
  28. Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models

    Xingyu Ding, Yuzhong Zhao, Chunhai Zhao +3

    cs.ROarXiv:2608.30643v12026
  29. TriSAR: Task Coordination and Collision Avoidance for Aerial Robot Teams in Disaster Response

    Aditya Anil Kapile, Pedro Machado, Isibor Kennedy Ihianle

    cs.ROarXiv:2609.01731v12026
  30. Self-Supervised Policy Adaptation during Deployment

    Nicklas Hansen, Rishabh Jangir, Yu Sun +5

    cs.LGcs.CVcs.ROarXiv:2007.04309v32020
  31. Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

    Anthony Liang, Yigit Korkmaz, Jiahui Zhang +14

    cs.ROcs.AIcs.LGarXiv:2603.02115v22026
  32. Conditional Affordance Learning for Driving in Urban Environments

    Axel Sauer, Nikolay Savinov, Andreas Geiger

    cs.ROcs.LGeess.SYarXiv:1806.06498v32018
  33. Training Agents Inside of Scalable World Models

    Danijar Hafner, Wilson Yan, Timothy Lillicrap

    cs.AIcs.LGcs.ROarXiv:2509.24527v12025
  34. BiTraP: Bi-directional Pedestrian Trajectory Prediction with Multi-modal Goal Estimation

    Yu Yao, Ella Atkins, Matthew Johnson-Roberson +2

    cs.CVcs.ROarXiv:2007.14558v22020
  35. Learning Dexterous Manipulation for a Soft Robotic Hand from Human Demonstration

    Abhishek Gupta, Clemens Eppner, Sergey Levine +1

    cs.LGcs.ROarXiv:1603.06348v32016
  36. EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data

    Ruijie Zheng, Dantong Niu, Yuqi Xie +12

    cs.ROarXiv:2602.16710v12026
  37. DNC-IMM: Early Lane-Change Intention Recognition via Neural Calibration Based on Driving Context Information

    Woong-Chan Byun, Seung-Hyun Kong

    cs.ROcs.AIarXiv:2609.01120v12026
  38. NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

    Chia-Yu Hung, Qi Sun, Pengfei Hong +5

    cs.ROcs.AIcs.CVarXiv:2504.19854v12025
  39. Programming and execution of skill-based human-robot-crane collaborative tasks

    Taneli Lohi, Markku Suomalainen, Roope Mellanen +1

    cs.ROarXiv:2609.03392v12026
  40. GeoWAM: Visual Geometry World Action Models for Autonomous Driving

    Yiren Lu, Xin Ye, Jiaming Liu +9

    cs.CVcs.ROarXiv:2608.23486v12026
  41. Learning to infer and manipulate through distributed whole-arm interaction in a soft robot

    Chuhan Zhang, Ebrahim Shahabi, Kseniia Khomenko +2

    cs.ROarXiv:2608.30773v12026
  42. Active Surface-Driven Reconfigurable Gripper: Robust Grasping and Sequential Manipulation of Thin Objects

    Ziyi Zheng, Keqi Zhu, Hao Wu +2

    cs.ROarXiv:2608.26883v12026
  43. Decoupling Policy Extraction for Offline Reinforcement Learning

    Xuyao Lin, Yixiang Shan, Jinru Duan +7

    cs.LGcs.ROarXiv:2608.20909v12026
  44. Demonstration-Guided Humanoid Stand-Up on an Emulated Deformable Surface

    Aniruddh Kushwah, Vyankatesh Ashtekar, Ashish Dutta

    cs.ROarXiv:2608.20852v12026
  45. Pre-training Visual Dexterity in Simulation

    Sarthak Kamat, Adam Rashid, Satvik Sharma +4

    cs.ROcs.AIcs.CVarXiv:2608.15917v12026
  46. Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade

    Malo de Pastor

    cs.LGcs.ROarXiv:2608.14650v12026
  47. DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects

    Tianshan Zhang, Yijia Duan, Yanjun Li +2

    cs.ROcs.CVarXiv:2606.15133v12026
  48. Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

    Yi Wang, Xinchen Li, Pengwei Xie +13

    cs.ROarXiv:2605.00416v22026
  49. Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications

    Kento Kawaharazuka, Jihoon Oh, Jun Yamada +2

    cs.ROcs.AIcs.CVarXiv:2510.07077v12025
  50. Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models

    Zichen Jeff Cui, Omar Rayyan, Haritheja Etukuru +16

    cs.ROcs.LGarXiv:2602.09017v12026
  51. Reinforcement Learning with Action Chunking

    Qiyang Li, Zhiyuan Zhou, Sergey Levine

    cs.LGcs.AIcs.ROarXiv:2507.07969v42025
  52. GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving

    Lloyd Russell, Anthony Hu, Lorenzo Bertoni +4

    cs.CVcs.AIcs.ROarXiv:2503.20523v12025
  53. OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction

    Lujie Yang, Xiaoyu Huang, Zhen Wu +6

    cs.ROcs.AIcs.LGarXiv:2509.26633v32025
  54. Learning a visuomotor controller for real world robotic grasping using simulated depth images

    Ulrich Viereck, Andreas ten Pas, Kate Saenko +1

    cs.ROcs.AIarXiv:1706.04652v32017
  55. Learning to Act from Actionless Videos through Dense Correspondences

    Po-Chen Ko, Jiayuan Mao, Yilun Du +2

    cs.ROcs.CVcs.LGarXiv:2310.08576v12023
  56. Being-H0.7: A Latent World-Action Model from Egocentric Videos

    Hao Luo, Wanpeng Zhang, Yicheng Feng +6

    cs.ROcs.CVcs.LGarXiv:2605.00078v12026
  57. Unified Vision-Language-Action Model

    Yuqi Wang, Xinghang Li, Wenxuan Wang +5

    cs.CVcs.ROarXiv:2506.19850v12025
  58. Navigating to Objects in the Real World

    Theophile Gervet, Soumith Chintala, Dhruv Batra +2

    cs.ROcs.CVcs.LGarXiv:2212.00922v12022
  59. Open-vocabulary Queryable Scene Representations for Real World Planning

    Boyuan Chen, Fei Xia, Brian Ichter +5

    cs.ROcs.AIcs.CVarXiv:2209.09874v22022
  60. A Comparative Study of Nonlinear MPC and Differential-Flatness-Based Control for Quadrotor Agile Flight

    Sihao Sun, Angel Romero, Philipp Foehn +2

    cs.ROarXiv:2109.01365v62021