Robotics
Papers filed under cs.RO on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,821 to 2,880 of 3,237
FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment
Han Zhao, Jingbo Wang, Wenxuan Song +5
cs.ROarXiv:2602.17259v12026Learning to Navigate in Complex Environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola +9
cs.AIcs.CVcs.LGarXiv:1611.03673v32016VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory
Shaoan Wang, Yuanfei Luo, Xingyu Chen +6
cs.ROcs.CVarXiv:2601.08665v12026R3M: A Universal Visual Representation for Robot Manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar +2
cs.ROcs.AIcs.CVarXiv:2203.12601v32022RMA: Rapid Motor Adaptation for Legged Robots
Ashish Kumar, Zipeng Fu, Deepak Pathak +1
cs.LGcs.AIcs.CVarXiv:2107.04034v12021MWM: Mobile World Models for Action-Conditioned Consistent Prediction
Han Yan, Zishang Xiang, Zeyu Zhang +1
cs.CVcs.ROarXiv:2603.07799v12026VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
Zixuan Wang, Yuxin Chen, Yuqi Liu +6
cs.ROarXiv:2603.22003v32026TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers
Bin Yu, Shijie Lian, Xiaopeng Lin +8
cs.ROcs.CVarXiv:2601.14133v22026DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
Alexander Khazatsky, Karl Pertsch, Suraj Nair +98
cs.ROarXiv:2403.12945v22024SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani +6
cs.CVcs.CLcs.LGarXiv:2401.12168v12024CLIPort: What and Where Pathways for Robotic Manipulation
Mohit Shridhar, Lucas Manuelli, Dieter Fox
cs.ROcs.CLcs.CVarXiv:2109.12098v12021MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
Yejin Kim, Wilbert Pumacay, Omar Rayyan +23
cs.ROcs.AIcs.CVarXiv:2602.11337v22026MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
Abhay Deshpande, Maya Guru, Rose Hendrix +23
cs.ROarXiv:2603.16861v22026Illuminating search spaces by mapping elites
Jean-Baptiste Mouret, Jeff Clune
cs.AIcs.NEcs.ROarXiv:1504.04909v12015Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
Xianjin Wu, Dingkang Liang, Tianrui Feng +5
cs.CVcs.ROarXiv:2603.19235v32026What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
Ajay Mandlekar, Danfei Xu, Josiah Wong +7
cs.ROcs.AIcs.LGarXiv:2108.03298v22021Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning
Jiaheng Hu, Jay Shim, Chen Tang +4
cs.LGcs.ROarXiv:2603.11653v32026Sim-to-Real: Learning Agile Locomotion For Quadruped Robots
Jie Tan, Tingnan Zhang, Erwin Coumans +5
cs.ROcs.AIarXiv:1804.10332v22018FASTER: Rethinking Real-Time Flow VLAs
Yuxiang Lu, Zhe Liu, Xianzhe Fan +5
cs.ROcs.CVarXiv:2603.19199v32026RLBench: The Robot Learning Benchmark & Learning Environment
Stephen James, Zicong Ma, David Rovick Arrojo +1
cs.ROcs.AIcs.CVarXiv:1909.12271v12019TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments
Zhiyu Huang, Yun Zhang, Johnson Liu +3
cs.ROarXiv:2602.02459v22026VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
Haoran Yuan, Weigang Yi, Zhenyu Zhang +9
cs.ROcs.AIcs.CVarXiv:2603.23481v12026DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving
Yang Zhou, Hao Shao, Letian Wang +3
cs.CVcs.AIcs.ROarXiv:2601.01528v22026Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning
Lukas Brunke, Melissa Greeff, Adam W. Hall +4
cs.ROcs.LGeess.SYarXiv:2108.06266v22021IMU-Free Body-Frame State Estimation with Sparse Scene Flow for Quadcopters
Daniel Grønhaug, Sofie Markeset, Mathias Kolberg
cs.ROcs.CVarXiv:2608.20891v12026Gibson Env: Real-World Perception for Embodied Agents
Fei Xia, Amir Zamir, Zhi-Yang He +3
cs.AIcs.CVcs.GRarXiv:1808.10654v12018MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
Yang Liu, Pengxiang Ding, Tengyue Jiang +10
cs.ROarXiv:2603.25406v32026Logic-VLA: A Temporal Logic Conditioned Vision-Language-Action Model
Celina Shiyu Wang, Yiqi Zhao, Junjie Ye +2
cs.ROcs.LOeess.SYarXiv:2608.20556v12026Humanoid Musical Robots as Experimental Interfaces for Music-Evoked Emotion
Vincent K. M. Cheung, Jia-Yeu Lin
cs.ROcs.HCcs.MMarXiv:2608.20433v12026The Coastline as a Structural Constraint: Harnessing Scene Geometry for Autonomous Surface Vessel Localization
Derek R. Benham, Joshua G. Mangelson
cs.ROcs.CVarXiv:2608.21276v12026Learning-Based Measurement-Robust Control Barrier Functions for Obstacle Avoidance under State Estimation Error
Nicholas Rober, Yixuan Jia, Jonathan P. How
eess.SYcs.ROarXiv:2608.20467v12026VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for Manipulation
Congsheng Xu, Qiaochu Yang, Fangyuan Shi +7
cs.ROcs.CVarXiv:2608.21290v12026Relation-Shape Convolutional Neural Network for Point Cloud Analysis
Yongcheng Liu, Bin Fan, Shiming Xiang +1
cs.CVcs.AIcs.CGarXiv:1904.07601v32019VLANeXt: Recipes for Building Strong VLA Models
Xiao-Ming Wu, Bin Fan, Kang Liao +6
cs.CVcs.AIcs.ROarXiv:2602.18532v22026An Algorithmic Perspective on Imitation Learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann +3
cs.ROcs.LGarXiv:1811.06711v12018Roadside-Cooperative Autonomous Driving: From Data Platform to Vision-Language End-to-End Reasoning
Yitao Xu, Tong Wu, Yiyan Wu +7
cs.ROarXiv:2608.21032v12026Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation
Mutian Xu, Tianbao Zhang, Tianqi Liu +3
cs.ROcs.CVarXiv:2603.16669v12026Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning
Yalcin Tur, Jalal Naghiyev, Haoquan Fang +4
cs.ROarXiv:2602.07845v12026Learning Native Continuation for Action Chunking Flow Policies
Yufeng Liu, Hang Yu, Juntu Zhao +9
cs.ROcs.AIarXiv:2602.12978v22026On Evaluation of Embodied Navigation Agents
Peter Anderson, Angel Chang, Devendra Singh Chaplot +8
cs.AIcs.CVcs.LGarXiv:1807.06757v12018Masked Depth Modeling for Spatial Perception
Bin Tan, Changjiang Sun, Xiage Qin +8
cs.CVcs.ROarXiv:2601.17895v12026Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning
Nikita Rudin, David Hoeller, Philipp Reist +1
cs.ROcs.LGarXiv:2109.11978v32021Natural Sit-to-Stand Motion Synthesis For Humanoids via Guided Assistance Curricula and Staged Rewards
Meet Pal Singh, Vyankatesh Ashtekar, Ashish Dutta
cs.ROarXiv:2608.20823v12026Joint 2D-3D-Semantic Data for Indoor Scene Understanding
Iro Armeni, Sasha Sax, Amir R. Zamir +1
cs.CVcs.ROarXiv:1702.01105v22017SimToolReal: An Object-Centric Policy for Zero-Shot Dexterous Tool Manipulation
Kushal Kedia, Tyler Ga Wei Lum, Jeannette Bohg +1
cs.ROcs.AIarXiv:2602.16863v22026NeSAM: Neuro-Symbolic Kinodynamics with Soil Adaptation for Off-Road Mobility
Chenhui Pan, Tong Xu, Francesco Cancelliere +1
cs.ROarXiv:2608.21330v12026DS-SLAM: A Semantic Visual SLAM towards Dynamic Environments
Chao Yu, Zuxin Liu, Xinjun Liu +4
cs.ROarXiv:1809.08379v22018Scalable Distributed Simulation-Based Testing for Automated Driving Systems
Christian Geller, Benedikt Haas, Lutz Eckstein
cs.ROcs.DBcs.SEarXiv:2608.20904v12026Koala Gripper: Co-designing Robotic Grippers and Data-Capture Devices for Scaling Dexterous Manipulation Learning
Amar Hajj-Ahmad, Zubin Kremer Guha, Tim Fofonoff +8
cs.ROarXiv:2608.20546v12026EndoLIFT: Language-Disambiguated Latent-Conditioned Rectified Flow for Bidirectional Endoscopic Control
Chi Kit Ng, Yidong Zhang, Lui Siu Hing +9
cs.ROarXiv:2608.20478v12026LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries
Shijie Lian, Bin Yu, Xiaopeng Lin +6
cs.AIcs.CLcs.CVarXiv:2601.15197v72026Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
Chi-Pin Huang, Yunze Man, Zhiding Yu +4
cs.CVcs.AIcs.LGarXiv:2601.09708v22026Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization
Chelsea Finn, Sergey Levine, Pieter Abbeel
cs.LGcs.AIcs.ROarXiv:1603.00448v32016VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
Wenlong Huang, Chen Wang, Ruohan Zhang +3
cs.ROcs.AIcs.CLarXiv:2307.05973v22023Teaching is a Process: The TOSS Framework for Modeling Human Teaching Decisions in Human-Interactive Robot Learning
Bernhard Hilpert, Kim Baraka, Joost Broekens
cs.ROcs.HCarXiv:2608.21083v12026Fast Coordinated Bimanual Motion Planning With Hard Constraints
Borna Paro, Luka Petrović, Ivan Marković
cs.ROarXiv:2608.20946v12026Pneumatic Units for Logic-based Sequential Excitation (PULSE) in Wearable Haptic Devices
Jessica Healey, Anoush Sepehri, Michael T. Tolley +1
cs.HCcs.ROarXiv:2608.20626v12026SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes
Nicholas Pfaff, Thomas Cohn, Sergey Zakharov +2
cs.ROcs.AIcs.CVarXiv:2602.09153v22026ABot-N0: Technical Report on the VLA Foundation Model for Versatile Embodied Navigation
Zedong Chu, Shichao Xie, Xiaolong Wu +41
cs.ROcs.AIcs.CVarXiv:2602.11598v12026Continuous Deep Q-Learning with Model-based Acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever +1
cs.LGcs.AIcs.ROarXiv:1603.00748v12016