Robotics

Papers filed under cs.RO on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,061 to 3,120 of 3,241

  1. DA-WAM: Decision-Aligned Future Latents for Driving World Models

    Ruiguo Zhong, Benshan Ma, Xiaolong Chen +5

    cs.ROcs.AIarXiv:2608.19085v12026
  2. High-Dimensional Continuous Control Using Generalized Advantage Estimation

    John Schulman, Philipp Moritz, Sergey Levine +2

    cs.LGcs.ROeess.SYarXiv:1506.02438v62015
  3. World Pilot: Steering Vision-Language-Action Models with World-Action Priors

    Zefu Lin, Rongxu Cui, Junjia Xu +4

    cs.ROarXiv:2606.12403v12026
  4. GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

    Ziyang Cheng, Tianshu Tang, Jinxin Lan +17

    cs.ROcs.AIcs.LGarXiv:2608.18234v12026
  5. ShapeNet: An Information-Rich 3D Model Repository

    Angel X. Chang, Thomas Funkhouser, Leonidas Guibas +10

    cs.GRcs.AIcs.CGarXiv:1512.03012v12015
  6. CARLA: An Open Urban Driving Simulator

    Alexey Dosovitskiy, German Ros, Felipe Codevilla +2

    cs.LGcs.AIcs.CVarXiv:1711.03938v12017
  7. Zetta $ζ$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

    Xin Ding, Liang Mi, Mingzhe Huang +12

    cs.ROarXiv:2608.16590v12026
  8. iMaC: Translating Actions into Motion and Contact Images for Embodied World Models

    Zhenyu Wu, Xiuwei Xu, Yukun Zhou +8

    cs.ROcs.CVarXiv:2606.09813v12026
  9. WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation

    Arnav Kumar Jain, Yilin Wu, Jesse Farebrother +2

    cs.ROarXiv:2606.13672v22026
  10. VLCP: Vision Language Control Policy Closed-Loop Code Replanning for Robot Manipulation

    Dhia Naouali, Minghan Wu, Claudia Wong +2

    cs.ROcs.LGarXiv:2608.16978v12026
  11. Repetition as Reinforcement: Enhancing Sample Efficiency via Instant Episode Repetition in Reinforcement Learning

    Hoda Yamani, Yuning Xing, Koen van Rijnsoever +2

    cs.LGcs.ROarXiv:2608.17347v12026
  12. Communication Reduction via Semantic-Based Encoding in DMPC Using LSTMs

    Torben Schiz, Pedro H. J. Nardelli, Henrik Ebel

    eess.SYcs.DCcs.LGarXiv:2608.17592v12026
  13. Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups

    Zeyun Deng, Yuzhe Lu, Yawei Wang +6

    cs.ROcs.LGarXiv:2608.17423v12026
  14. UAV Video Deblurring via Motion-Aware Diffusion: A Path to Robust Target Detection

    Zhiqiang Hu, Shouren Huang, Masatoshi Ishikawa

    cs.CVcs.AIcs.ROarXiv:2608.15259v12026
  15. Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning

    Zihang Wang, Yishan Wang

    cs.ROcs.AIarXiv:2608.15088v12026
  16. Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

    Ayoub Kirouane, Christos Petrocheilos

    cs.CLcs.LGcs.ROarXiv:2608.17744v12026
  17. Training with synthetic data for drone detection in thermal imagery

    Tanel Liiv, Sander Soodla, Nzamba Bignoumba +2

    cs.CVcs.AIcs.ETarXiv:2608.17799v12026
  18. EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video

    Hyunjin Kim, Ri-Zhao Qiu, Guangqi Jiang +1

    cs.CVcs.AIcs.ROarXiv:2606.16202v12026
  19. Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes

    Tongyan Fang, Siyuan Huang, Naiyu Fang +6

    cs.ROcs.LGarXiv:2606.17043v12026
  20. Adaptive Volumetric Mechanical Property Fields Invariant to Resolution

    Rishit Dagli, Donglai Xiang, Vismay Modi +4

    cs.CVcs.LGcs.ROarXiv:2606.18231v12026
  21. Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City

    Adrian Cespedes, Marcelo Chincha, Dunant Cusipuma +3

    cs.CVcs.AIcs.ROarXiv:2606.20980v12026
  22. Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

    Nikita Kachaev, Andrey Moskalenko, Matvey Skripkin +10

    cs.LGcs.ROarXiv:2606.19297v12026
  23. Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement

    Kinam Kim, Namiko Saito, Heecheol Kim +3

    cs.ROarXiv:2606.18953v12026
  24. GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning

    Haoyu Wang, Guoqing Ma, Zeyu Zhang +3

    cs.CVcs.ROarXiv:2606.17480v12026
  25. EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

    Yifan Zhong, Zhang Chen, Tianrui Guan +13

    cs.ROarXiv:2607.09701v12026
  26. PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning

    Youngjoon Jeong, Jihwan Yu, Minsoo Jo +2

    cs.ROcs.AIcs.LGarXiv:2606.21139v12026
  27. A Theoretical Framework for Parallel Lifelong MAPF Using Group Decentralized Planning

    Alex DeWeese, Jiaoyang Li, Guannan Qu

    cs.MAcs.AIcs.ROarXiv:2608.17928v12026
  28. Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision

    Amir Arsalan Nematollahi, Shayan Ahmadi, Mehdi Tale Masouleh +1

    cs.ROcs.AIcs.LGarXiv:2608.17628v12026
  29. Dijkstra as an Oracle for Online Stochastic Shortest Path Navigation with Provable Guarantees

    Mansur M. Arief, Ali Akarma, Ahmad Alfan Alfian Irfan

    cs.ROcs.AImath.OCarXiv:2608.17703v12026
  30. ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control

    Yanghao Zhou, Jingyu Ma, Yibo Peng +3

    cs.ROarXiv:2604.27711v12026
  31. ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving

    Huimin Wang, Yue Wang, Bihao Cui +7

    cs.ROarXiv:2605.04647v22026
  32. MistyPilot: Enabling Social-Robot Control through Multi-Agent LLM Skill Orchestration

    Xiao Wang, Lu Dong, Ifeoma Nwogu +2

    cs.ROcs.AIarXiv:2608.15549v12026
  33. Geometric Action Model for Robot Policy Learning

    Jisang Han, Seonghu Jeon, Jaewoo Jung +7

    cs.ROcs.CVcs.LGarXiv:2606.17046v22026
  34. Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models

    Yanyan Zhang, Chaoda Song, Vikash Singh +6

    cs.ROcs.AIcs.CVarXiv:2605.11459v22026
  35. Guava: An Effective and Universal Harness for Embodied Manipulation

    Haowen Liu, Xirui Li, Shaoxiong Yao +5

    cs.ROcs.AIarXiv:2606.18363v12026
  36. NPU Offloading of a Frozen Visual Encoder for Robot Policy Training

    Hyojun Yun, Seungjae Won, Hyungpil Moon

    cs.ROcs.ARcs.LGarXiv:2608.15002v12026
  37. Teach and Grow: An Agent-Centered Architecture for General Robot Learning

    Chang Nie, Zhe Liu, Hesheng Wang

    cs.ROcs.AIcs.CVarXiv:2608.17209v12026
  38. ORPA: Online Residual Policy Adaptation for Robot Manipulation Control with Human Feedback

    Muhammad A. Muttaqien, Tomohiro Motoda, Ryo Hanai +1

    cs.ROcs.AIarXiv:2608.17323v12026
  39. MuJoCo-Drones-Gym: A GPU-Accelerated Multi-Drone Simulator for Control and Reinforcement Learning

    Manan Tayal

    cs.ROarXiv:2606.08039v12026
  40. APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies

    Kechun Xu, Zhenjie Zhu, Anzhe Chen +2

    cs.ROarXiv:2606.12366v12026
  41. LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

    Baochang Ren, Xinjie Liu, Xi Chen +15

    cs.CLcs.AIcs.LGarXiv:2606.13578v22026
  42. Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time

    Jeongeun Park, Juhan Park, Taekyung Kim +3

    cs.ROcs.AIarXiv:2606.15631v12026
  43. Human Universal Grasping

    Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu +5

    cs.ROcs.AIcs.CVarXiv:2606.17054v12026
  44. ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

    Hao Li, Ganlong Zhao, Yufei Liu +8

    cs.ROarXiv:2606.17200v12026
  45. ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue

    Daoxuan Zhang, Ping Chen, Jianyi Zhou +1

    cs.ROarXiv:2605.01371v12026
  46. StableVLA: Towards Robust Vision-Language-Action Models without Extra Data

    Yiyang Fu, Chubin Zhang, Shukai Gong +7

    cs.CVcs.ROarXiv:2605.18287v12026
  47. Learning Visual Feature-Based World Models via Residual Latent Action

    Xinyu Zhang, Zhengtong Xu, Yutian Tao +3

    cs.CVcs.AIcs.LGarXiv:2605.07079v12026
  48. PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects

    Ziang Cao, Yinghao Liu, Haitian Li +5

    cs.CVcs.ROarXiv:2605.21572v12026
  49. SkillComposer: Learning Reusable Skills for Natural-Language Robot Programming

    John Woods, Hasti Seifi

    cs.ROcs.CLcs.LGarXiv:2608.14944v12026
  50. HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents

    Shen Liu, Zhenguo Xu, Shaopu Wang +2

    cs.AIcs.ROarXiv:2608.16447v12026
  51. GaussMemory: Task-Driven 3D Gaussian Scene Memory for Long-Horizon Robotic Manipulation

    Zhiqiang Hu, Shouren Huang, Masatoshi Ishikawa

    cs.ROcs.AIarXiv:2608.14986v12026
  52. Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring

    Seongheon Park, Wendi Li, Changdae Oh +4

    cs.ROcs.AIarXiv:2605.30834v12026
  53. Learning High-Frequency Continuous Action Chunks in Latent Space

    Kunyun Wang, Yuhang Zheng, Yupeng Zheng +2

    cs.ROarXiv:2605.24931v12026
  54. EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints

    Ao Zhou, Bo Dai, Le Yu +5

    cs.AIcs.ROarXiv:2608.15502v12026
  55. Flash-WAM: Modality-Aware Distillation for World Action Models

    Arman Akbari, Ci Zhang, Arash Akbari +6

    cs.LGcs.CVcs.ROarXiv:2606.05254v12026
  56. Playful Agentic Robot Learning

    Junyi Zhang, Jiaxin Ge, Hanjun Yoo +17

    cs.ROcs.AIarXiv:2606.19419v12026
  57. Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation

    Yijie Xu, Haopeng Jin, Run Zhou +8

    cs.ROcs.AIarXiv:2608.15680v12026
  58. FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge

    Rajat Bhattacharjya, Yoomee Jung, Minwoo Kim +4

    cs.DCcs.AIcs.CVarXiv:2608.15410v12026
  59. Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning

    Yuxing Long, Lei Kang, Ziyan Yu +8

    cs.ROcs.AIcs.CLarXiv:2608.15863v12026
  60. Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

    Hongyan Feng, Sunlai Chen, Xuanyu Liu +9

    cs.ROarXiv:2608.17512v12026