Robotics

Papers filed under cs.RO on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

661 to 720 of 3,243

  1. Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

    Moritz Reuss, Ömer Erdinç Yağmurlu, Fabian Wenzel +1

    cs.ROarXiv:2407.05996v12024
  2. ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation

    Weisheng Dai, Kai Lan, Jianyi Zhou +5

    cs.ROarXiv:2602.00557v12026
  3. Flow as the Cross-Domain Manipulation Interface

    Mengda Xu, Zhenjia Xu, Yinghao Xu +4

    cs.ROcs.AIarXiv:2407.15208v22024
  4. On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning

    Changyu Liu, Yiyang Liu, Taowen Wang +7

    cs.ROarXiv:2601.06748v32026
  5. Trajectory Forecasts in Unknown Environments Conditioned on Grid-Based Plans

    Nachiket Deo, Mohan M. Trivedi

    cs.CVcs.ROarXiv:2001.00735v22020
  6. MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision

    Zijie Zhu, Weiren Cai, Yizhou Wang +4

    cs.CVcs.ROarXiv:2609.04958v12026
  7. DynaRetarget: Dynamically-Feasible Retargeting using Sampling-Based Trajectory Optimization

    Victor Dhedin, Ilyass Taouil, Shafeef Omar +4

    cs.ROarXiv:2602.06827v32026
  8. Deep Whole-body Parkour

    Ziwen Zhuang, Shaoting Zhu, Mengjie Zhao +1

    cs.ROcs.AIarXiv:2601.07701v12026
  9. Walk the PLANC: Physics-Guided RL for Agile Humanoid Locomotion on Constrained Footholds

    Min Dai, William D. Compton, Junheng Li +2

    cs.ROarXiv:2601.06286v12026
  10. Scaling World Model for Hierarchical Manipulation Policies

    Qian Long, Yueze Wang, Jiaxi Song +13

    cs.ROarXiv:2602.10983v22026
  11. Coupled Control and Wireless World Models for Resilient Remote Robotic Control

    H. P. Madushanka, Sumudu Samarakoon, Mehdi Bennis

    cs.ROcs.LGarXiv:2609.04851v12026
  12. Low Level Control of a Quadrotor with Deep Model-Based Reinforcement Learning

    Nathan O. Lambert, Daniel S. Drew, Joseph Yaconelli +3

    cs.ROcs.LGarXiv:1901.03737v22019
  13. Performance, Precision, and Payloads: Adaptive Nonlinear MPC for Quadrotors

    Drew Hanover, Philipp Foehn, Sihao Sun +2

    cs.ROarXiv:2109.04210v22021
  14. A Survey on Global LiDAR Localization: Challenges, Advances and Open Problems

    Huan Yin, Xuecheng Xu, Sha Lu +5

    cs.ROarXiv:2302.07433v52023
  15. Make Tracking Easy: Neural Motion Retargeting for Humanoid Whole-body Control

    Qingrui Zhao, Kaiyue Yang, Xiyu Wang +6

    cs.ROarXiv:2603.22201v32026
  16. Torque Saturation in Bipedal Robotic Walking through Control Lyapunov Function Based Quadratic Programs

    Kevin Galloway, Koushil Sreenath, Aaron D. Ames +1

    eess.SYcs.ROmath.OCarXiv:1302.7314v12013
  17. DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI

    En Yu, Haoran Lv, Jianjian Sun +46

    cs.ROarXiv:2602.14974v12026
  18. Abstracting Robot Manipulation Skills via Mixture-of-Experts Diffusion Policies

    Ce Hao, Xuanran Zhai, Yaohua Liu +1

    cs.ROarXiv:2601.21251v12026
  19. Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals

    Nate Gillman, Yinghua Zhou, Zitian Tang +6

    cs.CVcs.AIcs.ROarXiv:2601.05848v22026
  20. Emergent Dexterity via Diverse Resets and Large-Scale Reinforcement Learning

    Patrick Yin, Tyler Westenbroek, Zhengyu Zhang +9

    cs.ROarXiv:2603.15789v32026
  21. Orbeez-SLAM: A Real-time Monocular Visual SLAM with ORB Features and NeRF-realized Mapping

    Chi-Ming Chung, Yang-Che Tseng, Ya-Ching Hsu +6

    cs.ROcs.CVarXiv:2209.13274v22022
  22. Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation

    Huanyu Li, Kun Lei, Sheng Zang +5

    cs.ROcs.AIcs.LGarXiv:2601.07821v12026
  23. One Hand to Rule Them All: Canonical Representations for Unified Dexterous Manipulation

    Zhenyu Wei, Yunchao Yao, Mingyu Ding

    cs.ROarXiv:2602.16712v22026
  24. QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization

    Yuhao Xu, Yantai Yang, Zhenyang Fan +4

    cs.CVcs.ROarXiv:2602.03782v12026
  25. Cross-Hand Latent Representation for Vision-Language-Action Models

    Guangqi Jiang, Yutong Liang, Jianglong Ye +6

    cs.ROarXiv:2603.10158v12026
  26. FAVLA: A Force-Adaptive Fast-Slow VLA model for Contact-Rich Robotic Manipulation

    Yao Li, Peiyuan Tang, Wuyang Zhang +7

    cs.ROarXiv:2602.23648v12026
  27. Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

    Wonje Jeung, Sangyeon Yoon, Hyesoo Hong +6

    cs.ROcs.CLarXiv:2609.05401v12026
  28. Simultaneous Contact, Gait and Motion Planning for Robust Multi-Legged Locomotion via Mixed-Integer Convex Optimization

    Bernardo Aceituno-Cabezas, Carlos Mastalli, Hongkai Dai +7

    cs.ROcs.AIeess.SYarXiv:1904.04595v12019
  29. ProgressVLA: Progress-Guided Diffusion Policy for Vision-Language Robotic Manipulation

    Hongyu Yan, Qiwei Li, Jiaolong Yang +1

    cs.ROcs.AIarXiv:2603.27670v12026
  30. Force Policy: Learning Hybrid Force-Position Control Policy under Interaction Frame for Contact-Rich Manipulation

    Hongjie Fang, Shirun Tang, Mingyu Mei +9

    cs.ROarXiv:2602.22088v22026
  31. Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild

    Hao Luo, Ye Wang, Wanpeng Zhang +5

    cs.ROarXiv:2602.21736v12026
  32. Dynamic Occupancy Grid Prediction for Urban Autonomous Driving: A Deep Learning Approach with Fully Automatic Labeling

    Stefan Hoermann, Martin Bach, Klaus Dietmayer

    cs.ROarXiv:1705.08781v22017
  33. CoSTAR: Instructing Collaborative Robots with Behavior Trees and Vision

    Chris Paxton, Andrew Hundt, Felix Jonathan +2

    cs.ROarXiv:1611.06145v12016
  34. Combining Optimal Control and Learning for Visual Navigation in Novel Environments

    Somil Bansal, Varun Tolani, Saurabh Gupta +2

    cs.ROcs.AIcs.CVarXiv:1903.02531v22019
  35. TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance

    Zhemeng Zhang, Jiahua Ma, Xincheng Yang +13

    cs.ROarXiv:2601.20239v62026
  36. Video Generation Models in Robotics -- Applications, Research Challenges, Future Directions

    Zhiting Mei, Tenny Yin, Ola Shorinwa +9

    eess.SYcs.ROarXiv:2601.07823v12026
  37. Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

    Homanga Bharadhwaj, Debidatta Dwibedi, Abhinav Gupta +7

    cs.ROcs.CVcs.LGarXiv:2409.16283v12024
  38. ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation

    Xialin He, Sirui Xu, Xinyao Li +4

    cs.ROcs.CVarXiv:2603.03279v12026
  39. FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions

    Peng Li, Zihan Zhuang, Yangfan Gao +16

    cs.ROcs.CLcs.CVarXiv:2601.12799v12026
  40. Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization

    Ye Wang, Sipeng Zheng, Hao Luo +9

    cs.ROarXiv:2602.09722v12026
  41. Safe Large-Scale Robust Nonlinear MPC in Milliseconds via Reachability-Constrained System Level Synthesis on the GPU

    Jeffrey Fang, Glen Chou

    cs.ROcs.AIeess.SYarXiv:2604.07644v12026
  42. ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation

    Zhenyang Liu, Yongchong Gu, Yikai Wang +2

    cs.ROarXiv:2601.08325v12026
  43. Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills

    Yevgen Chebotar, Karol Hausman, Yao Lu +8

    cs.ROcs.LGarXiv:2104.07749v32021
  44. CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding

    Chenyang Ma, Kai Lu, Guangyu Yang +6

    cs.ROarXiv:2601.02295v22026
  45. Robo3D: Towards Robust and Reliable 3D Perception against Corruptions

    Lingdong Kong, Youquan Liu, Xin Li +6

    cs.CVcs.ROarXiv:2303.17597v42023
  46. RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

    Zhenxuan Fan, Bo Zhang, Yutong Lin +9

    cs.ROcs.AIcs.CVarXiv:2609.05324v12026
  47. HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing

    Konstantin Gubernatorov, Mikhail Sannikov, Ilya Mikhalchuk +7

    cs.ROarXiv:2603.15257v22026
  48. Interactively Picking Real-World Objects with Unconstrained Spoken Language Instructions

    Jun Hatori, Yuta Kikuchi, Sosuke Kobayashi +5

    cs.ROcs.CLarXiv:1710.06280v22017
  49. What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

    Vivek Chavan, Pengtao Xie, Yahuan Shi +3

    cs.ROcs.AIcs.CVarXiv:2609.05376v12026
  50. S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight

    Haodong Yan, Zhide Zhong, Jiaguan Zhu +10

    cs.CVcs.ROarXiv:2603.16195v22026
  51. FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation

    Ruiteng Zhao, Wenshuo Wang, Yicheng Ma +4

    cs.ROcs.CVarXiv:2602.02142v22026
  52. Learning to drive from a world on rails

    Dian Chen, Vladlen Koltun, Philipp Krähenbühl

    cs.ROcs.CVcs.LGarXiv:2105.00636v32021
  53. STORM: An Integrated Framework for Fast Joint-Space Model-Predictive Control for Reactive Manipulation

    Mohak Bhardwaj, Balakumar Sundaralingam, Arsalan Mousavian +4

    cs.ROarXiv:2104.13542v22021
  54. A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning

    Chongwen Dong, Mithun Paul Saint-Germain, Pinjari Asif +1

    cs.ROcs.AIarXiv:2609.05133v12026
  55. LIBERO-X: Robustness Litmus for Vision-Language-Action Models

    Guodong Wang, Chenkai Zhang, Qingjie Liu +4

    cs.CVcs.AIcs.ROarXiv:2602.06556v12026
  56. Not All Features Are Created Equal: A Mechanistic Study of Vision-Language-Action Models

    Bryce Grant, Xijia Zhao, Peng Wang

    cs.ROarXiv:2603.19233v12026
  57. Language-Conditioned World Modeling for Visual Navigation

    Yifei Dong, Fengyi Wu, Yilong Dai +10

    cs.CVcs.AIcs.ROarXiv:2603.26741v12026
  58. One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation

    Arka Pal, Rajesh Kumar, Hannes Eriksson +4

    cs.CVcs.AIcs.LGarXiv:2609.04921v12026
  59. Relational Graph Learning for Crowd Navigation

    Changan Chen, Sha Hu, Payam Nikdel +2

    cs.ROcs.AIcs.LGarXiv:1909.13165v32019
  60. TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation

    Yuzhe Huang, Pei Lin, Wanlin Li +5

    cs.ROarXiv:2601.20321v22026