Robotics

Papers filed under cs.RO on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,821 to 2,880 of 3,237

  1. FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment

    Han Zhao, Jingbo Wang, Wenxuan Song +5

    cs.ROarXiv:2602.17259v12026
  2. Learning to Navigate in Complex Environments

    Piotr Mirowski, Razvan Pascanu, Fabio Viola +9

    cs.AIcs.CVcs.LGarXiv:1611.03673v32016
  3. VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory

    Shaoan Wang, Yuanfei Luo, Xingyu Chen +6

    cs.ROcs.CVarXiv:2601.08665v12026
  4. R3M: A Universal Visual Representation for Robot Manipulation

    Suraj Nair, Aravind Rajeswaran, Vikash Kumar +2

    cs.ROcs.AIcs.CVarXiv:2203.12601v32022
  5. RMA: Rapid Motor Adaptation for Legged Robots

    Ashish Kumar, Zipeng Fu, Deepak Pathak +1

    cs.LGcs.AIcs.CVarXiv:2107.04034v12021
  6. MWM: Mobile World Models for Action-Conditioned Consistent Prediction

    Han Yan, Zishang Xiang, Zeyu Zhang +1

    cs.CVcs.ROarXiv:2603.07799v12026
  7. VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models

    Zixuan Wang, Yuxin Chen, Yuqi Liu +6

    cs.ROarXiv:2603.22003v32026
  8. TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers

    Bin Yu, Shijie Lian, Xiaopeng Lin +8

    cs.ROcs.CVarXiv:2601.14133v22026
  9. DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

    Alexander Khazatsky, Karl Pertsch, Suraj Nair +98

    cs.ROarXiv:2403.12945v22024
  10. SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

    Boyuan Chen, Zhuo Xu, Sean Kirmani +6

    cs.CVcs.CLcs.LGarXiv:2401.12168v12024
  11. CLIPort: What and Where Pathways for Robotic Manipulation

    Mohit Shridhar, Lucas Manuelli, Dieter Fox

    cs.ROcs.CLcs.CVarXiv:2109.12098v12021
  12. MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation

    Yejin Kim, Wilbert Pumacay, Omar Rayyan +23

    cs.ROcs.AIcs.CVarXiv:2602.11337v22026
  13. MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation

    Abhay Deshpande, Maya Guru, Rose Hendrix +23

    cs.ROarXiv:2603.16861v22026
  14. Illuminating search spaces by mapping elites

    Jean-Baptiste Mouret, Jeff Clune

    cs.AIcs.NEcs.ROarXiv:1504.04909v12015
  15. Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

    Xianjin Wu, Dingkang Liang, Tianrui Feng +5

    cs.CVcs.ROarXiv:2603.19235v32026
  16. What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

    Ajay Mandlekar, Danfei Xu, Josiah Wong +7

    cs.ROcs.AIcs.LGarXiv:2108.03298v22021
  17. Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning

    Jiaheng Hu, Jay Shim, Chen Tang +4

    cs.LGcs.ROarXiv:2603.11653v32026
  18. Sim-to-Real: Learning Agile Locomotion For Quadruped Robots

    Jie Tan, Tingnan Zhang, Erwin Coumans +5

    cs.ROcs.AIarXiv:1804.10332v22018
  19. FASTER: Rethinking Real-Time Flow VLAs

    Yuxiang Lu, Zhe Liu, Xianzhe Fan +5

    cs.ROcs.CVarXiv:2603.19199v32026
  20. RLBench: The Robot Learning Benchmark & Learning Environment

    Stephen James, Zicong Ma, David Rovick Arrojo +1

    cs.ROcs.AIcs.CVarXiv:1909.12271v12019
  21. TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments

    Zhiyu Huang, Yun Zhang, Johnson Liu +3

    cs.ROarXiv:2602.02459v22026
  22. VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs

    Haoran Yuan, Weigang Yi, Zhenyu Zhang +9

    cs.ROcs.AIcs.CVarXiv:2603.23481v12026
  23. DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving

    Yang Zhou, Hao Shao, Letian Wang +3

    cs.CVcs.AIcs.ROarXiv:2601.01528v22026
  24. Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning

    Lukas Brunke, Melissa Greeff, Adam W. Hall +4

    cs.ROcs.LGeess.SYarXiv:2108.06266v22021
  25. IMU-Free Body-Frame State Estimation with Sparse Scene Flow for Quadcopters

    Daniel Grønhaug, Sofie Markeset, Mathias Kolberg

    cs.ROcs.CVarXiv:2608.20891v12026
  26. Gibson Env: Real-World Perception for Embodied Agents

    Fei Xia, Amir Zamir, Zhi-Yang He +3

    cs.AIcs.CVcs.GRarXiv:1808.10654v12018
  27. MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation

    Yang Liu, Pengxiang Ding, Tengyue Jiang +10

    cs.ROarXiv:2603.25406v32026
  28. Logic-VLA: A Temporal Logic Conditioned Vision-Language-Action Model

    Celina Shiyu Wang, Yiqi Zhao, Junjie Ye +2

    cs.ROcs.LOeess.SYarXiv:2608.20556v12026
  29. Humanoid Musical Robots as Experimental Interfaces for Music-Evoked Emotion

    Vincent K. M. Cheung, Jia-Yeu Lin

    cs.ROcs.HCcs.MMarXiv:2608.20433v12026
  30. The Coastline as a Structural Constraint: Harnessing Scene Geometry for Autonomous Surface Vessel Localization

    Derek R. Benham, Joshua G. Mangelson

    cs.ROcs.CVarXiv:2608.21276v12026
  31. Learning-Based Measurement-Robust Control Barrier Functions for Obstacle Avoidance under State Estimation Error

    Nicholas Rober, Yixuan Jia, Jonathan P. How

    eess.SYcs.ROarXiv:2608.20467v12026
  32. VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for Manipulation

    Congsheng Xu, Qiaochu Yang, Fangyuan Shi +7

    cs.ROcs.CVarXiv:2608.21290v12026
  33. Relation-Shape Convolutional Neural Network for Point Cloud Analysis

    Yongcheng Liu, Bin Fan, Shiming Xiang +1

    cs.CVcs.AIcs.CGarXiv:1904.07601v32019
  34. VLANeXt: Recipes for Building Strong VLA Models

    Xiao-Ming Wu, Bin Fan, Kang Liao +6

    cs.CVcs.AIcs.ROarXiv:2602.18532v22026
  35. An Algorithmic Perspective on Imitation Learning

    Takayuki Osa, Joni Pajarinen, Gerhard Neumann +3

    cs.ROcs.LGarXiv:1811.06711v12018
  36. Roadside-Cooperative Autonomous Driving: From Data Platform to Vision-Language End-to-End Reasoning

    Yitao Xu, Tong Wu, Yiyan Wu +7

    cs.ROarXiv:2608.21032v12026
  37. Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation

    Mutian Xu, Tianbao Zhang, Tianqi Liu +3

    cs.ROcs.CVarXiv:2603.16669v12026
  38. Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning

    Yalcin Tur, Jalal Naghiyev, Haoquan Fang +4

    cs.ROarXiv:2602.07845v12026
  39. Learning Native Continuation for Action Chunking Flow Policies

    Yufeng Liu, Hang Yu, Juntu Zhao +9

    cs.ROcs.AIarXiv:2602.12978v22026
  40. On Evaluation of Embodied Navigation Agents

    Peter Anderson, Angel Chang, Devendra Singh Chaplot +8

    cs.AIcs.CVcs.LGarXiv:1807.06757v12018
  41. Masked Depth Modeling for Spatial Perception

    Bin Tan, Changjiang Sun, Xiage Qin +8

    cs.CVcs.ROarXiv:2601.17895v12026
  42. Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning

    Nikita Rudin, David Hoeller, Philipp Reist +1

    cs.ROcs.LGarXiv:2109.11978v32021
  43. Natural Sit-to-Stand Motion Synthesis For Humanoids via Guided Assistance Curricula and Staged Rewards

    Meet Pal Singh, Vyankatesh Ashtekar, Ashish Dutta

    cs.ROarXiv:2608.20823v12026
  44. Joint 2D-3D-Semantic Data for Indoor Scene Understanding

    Iro Armeni, Sasha Sax, Amir R. Zamir +1

    cs.CVcs.ROarXiv:1702.01105v22017
  45. SimToolReal: An Object-Centric Policy for Zero-Shot Dexterous Tool Manipulation

    Kushal Kedia, Tyler Ga Wei Lum, Jeannette Bohg +1

    cs.ROcs.AIarXiv:2602.16863v22026
  46. NeSAM: Neuro-Symbolic Kinodynamics with Soil Adaptation for Off-Road Mobility

    Chenhui Pan, Tong Xu, Francesco Cancelliere +1

    cs.ROarXiv:2608.21330v12026
  47. DS-SLAM: A Semantic Visual SLAM towards Dynamic Environments

    Chao Yu, Zuxin Liu, Xinjun Liu +4

    cs.ROarXiv:1809.08379v22018
  48. Scalable Distributed Simulation-Based Testing for Automated Driving Systems

    Christian Geller, Benedikt Haas, Lutz Eckstein

    cs.ROcs.DBcs.SEarXiv:2608.20904v12026
  49. Koala Gripper: Co-designing Robotic Grippers and Data-Capture Devices for Scaling Dexterous Manipulation Learning

    Amar Hajj-Ahmad, Zubin Kremer Guha, Tim Fofonoff +8

    cs.ROarXiv:2608.20546v12026
  50. EndoLIFT: Language-Disambiguated Latent-Conditioned Rectified Flow for Bidirectional Endoscopic Control

    Chi Kit Ng, Yidong Zhang, Lui Siu Hing +9

    cs.ROarXiv:2608.20478v12026
  51. LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries

    Shijie Lian, Bin Yu, Xiaopeng Lin +6

    cs.AIcs.CLcs.CVarXiv:2601.15197v72026
  52. Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning

    Chi-Pin Huang, Yunze Man, Zhiding Yu +4

    cs.CVcs.AIcs.LGarXiv:2601.09708v22026
  53. Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization

    Chelsea Finn, Sergey Levine, Pieter Abbeel

    cs.LGcs.AIcs.ROarXiv:1603.00448v32016
  54. VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

    Wenlong Huang, Chen Wang, Ruohan Zhang +3

    cs.ROcs.AIcs.CLarXiv:2307.05973v22023
  55. Teaching is a Process: The TOSS Framework for Modeling Human Teaching Decisions in Human-Interactive Robot Learning

    Bernhard Hilpert, Kim Baraka, Joost Broekens

    cs.ROcs.HCarXiv:2608.21083v12026
  56. Fast Coordinated Bimanual Motion Planning With Hard Constraints

    Borna Paro, Luka Petrović, Ivan Marković

    cs.ROarXiv:2608.20946v12026
  57. Pneumatic Units for Logic-based Sequential Excitation (PULSE) in Wearable Haptic Devices

    Jessica Healey, Anoush Sepehri, Michael T. Tolley +1

    cs.HCcs.ROarXiv:2608.20626v12026
  58. SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes

    Nicholas Pfaff, Thomas Cohn, Sergey Zakharov +2

    cs.ROcs.AIcs.CVarXiv:2602.09153v22026
  59. ABot-N0: Technical Report on the VLA Foundation Model for Versatile Embodied Navigation

    Zedong Chu, Shichao Xie, Xiaolong Wu +41

    cs.ROcs.AIcs.CVarXiv:2602.11598v12026
  60. Continuous Deep Q-Learning with Model-based Acceleration

    Shixiang Gu, Timothy Lillicrap, Ilya Sutskever +1

    cs.LGcs.AIcs.ROarXiv:1603.00748v12016