Robotics

Papers filed under cs.RO on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,701 to 2,760 of 3,243

  1. Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping

    Konstantinos Bousmalis, Alex Irpan, Paul Wohlhart +9

    cs.LGcs.AIcs.CVarXiv:1709.07857v22017
  2. RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning

    Seungku Kim, Suhyeok Jang, Byungjun Yoon +3

    cs.ROcs.AIcs.CVarXiv:2602.18742v12026
  3. Learning Smooth Time-Varying Linear Policies with an Action Jacobian Penalty

    Zhaoming Xie, Kevin Karol, Jessica Hodgins

    cs.ROcs.GRarXiv:2602.18312v12026
  4. Enhancing Sim2Real Transfer for Torque-Controlled Robots through Real2Sim Dynamics Estimation and Reinforcement Learning

    Davide Bargellini, Alex Pasquali, Andrea Govoni +2

    cs.ROeess.SYarXiv:2608.22629v12026
  5. Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation

    Zipeng Fu, Tony Z. Zhao, Chelsea Finn

    cs.ROcs.AIcs.CVarXiv:2401.02117v12024
  6. DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model

    Zhenhua Xu, Yujia Zhang, Enze Xie +5

    cs.CVcs.ROarXiv:2310.01412v52023
  7. Mollified Value Learning

    Hrishikesh Viswanath, Juanwu Lu, S. Talha Bukhari +4

    cs.LGcs.ROarXiv:2602.23280v22026
  8. Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation

    Tyler Han, Bat Nemekhbold, Siyang Shen +6

    cs.ROarXiv:2602.24121v22026
  9. The Event-Camera Dataset and Simulator: Event-based Data for Pose Estimation, Visual Odometry, and SLAM

    Elias Mueggler, Henri Rebecq, Guillermo Gallego +2

    cs.ROcs.CVarXiv:1610.08336v42016
  10. TNT: Target-driveN Trajectory Prediction

    Hang Zhao, Jiyang Gao, Tian Lan +9

    cs.CVcs.ROarXiv:2008.08294v22020
  11. EmboAlign: Aligning Video Generation with Compositional Constraints for Zero-Shot Manipulation

    Gehao Zhang, Zhenyang Ni, Payal Mohapatra +3

    cs.ROarXiv:2603.05757v12026
  12. GCS-Bridging: Restoring Connectivity of Disconnected Convex Sets for Graph-of-Convex-Sets Motion Planning

    Xiaokai Zhou, Baoshi Cao, Yang Liu +4

    cs.ROarXiv:2608.22326v12026
  13. TartanAir: A Dataset to Push the Limits of Visual SLAM

    Wenshan Wang, Delong Zhu, Xiangwei Wang +6

    cs.ROarXiv:2003.14338v22020
  14. EMPIRE: Explicit Manipulation Planning as a Learnable Intermediate Representation for Egocentric Hand-Motion Forecasting

    Wen Wang, Ruibing Hou, Hong Chang +2

    cs.ROcs.AIarXiv:2608.22449v12026
  15. InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation

    Mengao Zhao, Ziang Li, Chaodong Huang +15

    cs.ROarXiv:2608.22990v12026
  16. Free-Energy-Gated Plasticity for Real-Time Online Motor Learning in Physical Human-Robot Interaction

    Hiroki Sawada, Jun Tani

    cs.ROcs.HCarXiv:2608.23000v22026
  17. Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks

    Henggang Cui, Vladan Radosavljevic, Fang-Chieh Chou +5

    cs.ROcs.CVcs.LGarXiv:1809.10732v22018
  18. Macro-Action Topological Navigation under Noisy Localization using Reinforcement Learning

    Simon Hakenes, Tobias Glasmachers

    cs.LGcs.ROarXiv:2608.23055v12026
  19. One-Shot Imitation Learning

    Yan Duan, Marcin Andrychowicz, Bradly C. Stadie +5

    cs.AIcs.LGcs.NEarXiv:1703.07326v32017
  20. AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent Forecasting

    Ye Yuan, Xinshuo Weng, Yanglan Ou +1

    cs.AIcs.CVcs.LGarXiv:2103.14023v32021
  21. LLM-Based Selection of Incongruent Verbal and Nonverbal Behavior for Virtual Humans

    Parisa Ghanad Torshizi, Stacy Marsella

    cs.AIcs.HCcs.ROarXiv:2608.22731v12026
  22. Autonomous Vehicles that Interact with Pedestrians: A Survey of Theory and Practice

    Amir Rasouli, John K. Tsotsos

    cs.ROcs.CVcs.HCarXiv:1805.11773v12018
  23. Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

    Sangoh Lee, Sangwoo Mo, Wook-Shin Han

    cs.ROcs.AIcs.CVarXiv:2608.23478v12026
  24. SIMART: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLM

    Chuanrui Zhang, Minghan Qin, Yuang Wang +3

    cs.CVcs.GRcs.ROarXiv:2603.23386v12026
  25. Deep Object Pose Estimation for Semantic Robotic Grasping of Household Objects

    Jonathan Tremblay, Thang To, Balakumar Sundaralingam +3

    cs.ROarXiv:1809.10790v12018
  26. $π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs

    Siting Wang, Xiaofeng Wang, Zheng Zhu +7

    cs.ROcs.CVarXiv:2603.02083v22026
  27. FlyPose: Towards Robust Human Pose Estimation From Aerial Views

    Hassaan Farooq, Marvin Brenner, Peter Stütz

    cs.CVcs.ROarXiv:2601.05747v22026
  28. Enhancing Underwater Imagery using Generative Adversarial Networks

    Cameron Fabbri, Md Jahidul Islam, Junaed Sattar

    cs.CVcs.ROarXiv:1801.04011v12018
  29. LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action

    Dhruv Shah, Blazej Osinski, Brian Ichter +1

    cs.ROcs.AIcs.CLarXiv:2207.04429v22022
  30. Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

    Alexandre Chapin, Bruno Machado, Emmanuel Dellandréa +1

    cs.ROarXiv:2601.21416v22026
  31. Multi-Modal Fusion Transformer for End-to-End Autonomous Driving

    Aditya Prakash, Kashyap Chitta, Andreas Geiger

    cs.CVcs.AIcs.LGarXiv:2104.09224v12021
  32. Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs

    Xinyuan Liu, Eren Sadikoglu, Riana Chatterjee +1

    cs.ROcs.AIcs.MAarXiv:2608.22657v12026
  33. SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models

    Hyeonbeom Choi, Daechul Ahn, Youhan Lee +3

    cs.ROcs.AIcs.LGarXiv:2602.04208v22026
  34. Cognitive Mapping and Planning for Visual Navigation

    Saurabh Gupta, Varun Tolani, James Davidson +3

    cs.CVcs.AIcs.LGarXiv:1702.03920v32017
  35. VAD: Vectorized Scene Representation for Efficient Autonomous Driving

    Bo Jiang, Shaoyu Chen, Qing Xu +7

    cs.ROcs.CVarXiv:2303.12077v32023
  36. RACO: Reliability-Aware Coarse-Goal Optimization for Inspection-Oriented UAV Vision-Language Navigation

    Sen Wang, Yiming Sun, Jiaxuan He +1

    cs.ROcs.AIcs.CVarXiv:2608.22678v12026
  37. UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models

    Lars Osterberg, Maggie Wang, Mac Schwager

    cs.ROcs.CVarXiv:2608.22869v12026
  38. LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models

    Chan Hee Song, Jiaman Wu, Clayton Washington +3

    cs.AIcs.CLcs.CVarXiv:2212.04088v32022
  39. MultiNet: Real-time Joint Semantic Reasoning for Autonomous Driving

    Marvin Teichmann, Michael Weber, Marius Zoellner +2

    cs.CVcs.ROarXiv:1612.07695v22016
  40. Large-Scale Study of Curiosity-Driven Learning

    Yuri Burda, Harri Edwards, Deepak Pathak +3

    cs.LGcs.AIcs.CVarXiv:1808.04355v12018
  41. Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving

    Jiangxin Sun, Feng Xue, Teng Long +4

    cs.CVcs.AIcs.ROarXiv:2602.23259v12026
  42. NaviDriveVLM: Decoupling High-Level Reasoning and Motion Planning for Autonomous Driving

    Ximeng Tao, Pardis Taghavi, Dimitar Filev +2

    cs.ROcs.LGarXiv:2603.07901v12026
  43. What is the effect of running-specific prostheses on long jumps? Optimization-based prediction and analysis using biomechanical models

    Anna Lena Emonds, Johannes Funken, Wolfgang Potthast +1

    cs.ROarXiv:2608.22507v12026
  44. ESPIRE: A Diagnostic Benchmark for Embodied Spatial Reasoning of Vision-Language Models

    Yanpeng Zhao, Wentao Ding, Hongtao Li +2

    cs.CVcs.LGcs.ROarXiv:2603.13033v12026
  45. HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions

    Yukang Cao, Haozhe Xie, Fangzhou Hong +4

    cs.CVcs.ROarXiv:2603.15612v12026
  46. Triplet2Track: A Hierarchical System with Object-Centric Representations for Reliable Long-Horizon Manipulation

    Jianxiang Liu, Gaojing Zhang, Chuan Wen +4

    cs.ROcs.AIarXiv:2608.22800v12026
  47. Habitat 2.0: Training Home Assistants to Rearrange their Habitat

    Andrew Szot, Alex Clegg, Eric Undersander +18

    cs.LGcs.ROarXiv:2106.14405v22021
  48. Motion Attribution for Video Generation

    Xindi Wu, Despoina Paschalidou, Jun Gao +5

    cs.CVcs.AIcs.LGarXiv:2601.08828v22026
  49. CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence

    Tianle Zeng, Yanci Wen, Hong Zhang

    cs.ROcs.AIcs.CVarXiv:2603.28032v22026
  50. Meta-Ctrl: Guaranteed Plan Generation by Decoupling Syntactic and Semantic Constraints

    Gwen Yidou-Weng, Edward Sun, Tianyi Ma +5

    cs.ROcs.AIarXiv:2608.22149v12026
  51. Socially Aware Motion Planning with Deep Reinforcement Learning

    Yu Fan Chen, Michael Everett, Miao Liu +1

    cs.ROcs.AIcs.HCarXiv:1703.08862v22017
  52. Google Scanned Objects: A High-Quality Dataset of 3D Scanned Household Items

    Laura Downs, Anthony Francis, Nate Koenig +5

    cs.ROcs.GRarXiv:2204.11918v12022
  53. DeepIM: Deep Iterative Matching for 6D Pose Estimation

    Yi Li, Gu Wang, Xiangyang Ji +2

    cs.CVcs.ROarXiv:1804.00175v42018
  54. Lidar for Autonomous Driving: The principles, challenges, and trends for automotive lidar and perception systems

    You Li, Javier Ibanez-Guzman

    cs.ROarXiv:2004.08467v12020
  55. DeFM: Learning Foundation Representations from Depth for Robotics

    Manthan Patel, Jonas Frey, Mayank Mittal +5

    cs.ROcs.CVarXiv:2601.18923v12026
  56. VikPath: A Vision Kansformer Framework for Effective Obstacle Avoidance in Self-Supervised Pathfinding

    Junyao Wang, Yulin Xu, Mohammad Abdullah Al Faruque

    cs.ROarXiv:2608.22675v12026
  57. SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-body Manipulation

    Mu Huang, Hui Wang, Kerui Ren +5

    cs.ROcs.AIcs.CVarXiv:2602.02402v22026
  58. TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving

    Kashyap Chitta, Aditya Prakash, Bernhard Jaeger +3

    cs.CVcs.AIcs.LGarXiv:2205.15997v12022
  59. Learning to Drive in a Day

    Alex Kendall, Jeffrey Hawke, David Janz +6

    cs.LGcs.AIcs.ROarXiv:1807.00412v22018
  60. Guided Riemannian Optimization (GuRO): Bridging Model Predictive Control and Decision Transformers

    Hossein Abdi, Satya Prakash Dash, Mingfei Sun

    cs.ROarXiv:2608.23204v12026