Source-linked AI summary

Augmented Reality and Robotics: A Survey and Taxonomy for AR-enhanced Human-Robot Interaction and Robotic Interfaces

Ryo Suzuki, Adnan Karim, Tian Xia, Hooman Hedayati, Nicolai Marquardt

arXiv:2203.03254v1cs.ROcs.CVcs.HC

TL;DR

Existing AR-HRI and robotic-interface research has grown across individual explorations without systematic analysis of its design space. This paper surveys 460 papers to build an eight-dimensional taxonomy, finding recurring patterns in proximity, visual widgets, and interaction modalities while identifying deployment and technology challenges.

  • Problem

    AR-HRI research spans many individual explorations, but its design strategies and research questions have rarely been analyzed systematically.

  • Method

    The paper synthesizes a corpus of 460 papers into a taxonomy covering eight design-space dimensions for AR and robotics research.

  • Results

    The corpus most commonly features co-located users without contact with robots, labels and annotations as widgets, and pointers/controllers or spatial gestures as interaction modalities.

  • Takeaways & Limitations

    The taxonomy provides common ground for positioning AR-HRI work and identifies design space categories and future research opportunities.

  • Takeaways & Limitations

    The review corpus is not exhaustive, and many AR-HRI systems have primarily been studied in controlled laboratory conditions rather than real-world settings.

Abstract

from arXiv · show

This paper contributes to a taxonomy of augmented reality and robotics based on a survey of 460 research papers. Augmented and mixed reality (AR/MR) have emerged as a new way to enhance human-robot interaction (HRI) and robotic interfaces (e.g., actuated and shape-changing interfaces). Recently, an increasing number of studies in HCI, HRI, and robotics have demonstrated how AR enables better interactions between people and robots. However, often research remains focused on individual explorations and key design strategies, and research questions are rarely analyzed systematically. In this paper, we synthesize and categorize this research field in the following dimensions: 1) approaches to augmenting reality; 2) characteristics of robots; 3) purposes and benefits; 4) classification of presented information; 5) design components and strategies for visual augmentation; 6) interaction techniques and modalities; 7) application domains; and 8) evaluation strategies. We formulate key challenges and opportunities to guide and inform future research in AR and robotics.

1 INTRODUCTION

The paper establishes a broad scope for AR-enhanced interaction with robots and robotic interfaces, then develops a taxonomy from a representative corpus of 460 papers. It positions the taxonomy as a common ground for organizing research, locating related work, and identifying future opportunities.

  • Motivation: Traditional robot interaction relies on internal physical or visual feedback, but robot form factors limit expressive physical feedback beyond those capabilities.Examples include movement, gestures, gaze, physical transformation, lights, and small displays.
  • Contribution: The taxonomy organizes AR and robotics research across eight dimensions spanning augmentation approaches, robot characteristics, purposes, information, visual design, interaction, applications, and evaluation.It covers both AR-enhanced HRI and robotic user interfaces.
  • Scope: The review includes robotic systems and actuated interfaces designed to interact with people, including adaptive environments, swarm interfaces, and shape-changing interfaces.The scope includes traditional robots, self-driving cars, and actuated user interfaces, while excluding externally actuated passive objects.
  • Corpus: The literature review selected 460 papers by combining database search, screening, and 64 additional papers identified through expert knowledge.The initial search produced 925 papers after deduplication; 396 remained after screening before the additional papers were merged.
  • Limitations: The corpus is representative rather than exhaustive because inclusion boundaries are not clear-cut, so the authors make their coding and dataset open-source for expansion.This limitation applies to the systematic compilation used to build the taxonomy.
  • Research aid: The paper uses systematic coding and provides color-coded figures, citation lookup, and appendix tables to help researchers navigate the design space.The appendix compiles citations and paper counts for design-space categories and subcategories.

3 APPROACHES TO AUGMENTING REALITY IN ROBOTICS

The taxonomy classifies reality augmentation by where AR hardware overrides the optical path and whether robots or their surroundings are augmented. It distinguishes on-body, on-environment, and on-robot implementations, each with different sharing, mobility, and device requirements.

  • Classification: The classification uses hardware placement and target location to organize augmentation approaches into on-body, on-environment, and on-robot categories.The placement dimension concerns where the optical path is overridden with digital information.
  • Augment Robots: Augmenting robots overlays or anchors additional information on the robots themselves.The approach includes on-body devices, environment-embedded devices, and robots that augment their own appearance.
  • Augment Robots: On-body augmentation uses HMDs or mobile AR interfaces to add appearances or expressive content to robots.Examples include overlaying a remote user on a telepresence robot and adding an animated face to a Roomba.
  • Augment Robots: On-environment augmentation uses projectors or see-through displays embedded around robots to overlay information on robotic interfaces and shape-changing systems.Examples include projection mapping on drones, shape-shifting walls, and handheld shape-changing interfaces.
  • Augment Robots: On-robot augmentation lets robots modify their own appearance using attached projectors or displays without external AR devices.Furhat demonstrates a back-projected animated face, while robot bodies can serve as projection screens.
  • Augment Surroundings: Augmenting surroundings targets surrounding mid-air 3D space, physical objects, or physical environments such as walls, floors, and ceilings.This is the alternative top-level approach to augmenting robots themselves.
  • Augment Surroundings: On-body surroundings augmentation uses HMDs, mobile AR devices, or handheld projectors, enabling expressive 3D rendering and spatial scene understanding.Drone Augmented Human Vision changes a wall’s appearance for drone control, while RoMA overlays interactive 3D models on a robotic printer.
  • Augment Surroundings: On-environment augmentation shares AR experiences easily with co-located users through projection mapping or surface displays, but installed equipment fixes the location and can limit outdoor mobility.Examples communicate robot intentions or show additional information in the robot’s surroundings.

4 CHARACTERISTICS OF AUGMENTED ROBOTS

The taxonomy characterizes augmented robots by form factor, people-to-robot relationship, scale, and interaction proximity. These dimensions cover diverse robotic systems and interaction settings.

  • Form Factor: Augmented robots span robotic arms, drones, mobile robots, humanoids, vehicles, actuated objects, multiple form factors, and fabrication machines.
  • Relationship: People-to-robot relationships include one person with one robot, one person with multiple robots, multiple people with one robot, and multiple people with multiple robots.
  • Scale: Robot scale ranges from handheld and tabletop systems to body-scale robots, vehicles, and building-construction robots.
  • Proximity: Interaction proximity ranges from co-located to remote, affecting whether robots are touchable, visible, or out of sight.

5 PURPOSES AND BENEFITS OF VISUAL AUGMENTATION

The survey organizes visual augmentation by its purposes and by the information presented. AR supports robot programming, control, safety, intent communication, expressiveness, and four broad information categories.

  • Purposes and Benefits: Visual augmentation purposes divide into programming and control, and understanding, interpretation, and communication.
  • Purposes and Benefits: AR facilitates robot programming by simulating programmed behaviors, including trajectories that show users how robots will behave.
  • Purposes and Benefits: AR supports real-time control, navigation, and teleoperation through interactive visual feedback in remote or co-located settings.
  • Purposes and Benefits: Visual augmentation can improve safety awareness, communicate robot intent spatially, and increase robot expressiveness through virtual arms, facial expressions, remote users, or interactive content.
  • Presented Information: Presented information comprises robot internal information, external environmental information, plan and activity, and supplemental content.
  • Presented Information: Internal and external information can cover robot status, software and hardware condition, sensor data, camera feeds, objects, obstacles, localization, and reconstructed environments.
  • Presented Information: Plan and activity information represents future motion or behavior, simulated programmed behavior, goals, targets, and task progress, while supplemental content supports expressive interaction and telepresence.

7 DESIGN COMPONENTS AND STRATEGIES FOR VISUAL AUGMENTATION

The taxonomy groups visual augmentation practices into UIs and widgets, spatial references and visualizations, and embedded visual effects. These strategies range from controls and labels to data overlays and expressive graphical content.

  • UIs and Widgets: AR interfaces use UIs and widgets to help users see, understand, and interact with robot-related information.
  • UIs and Widgets: Menus present selectable options and can support robot control or communication through menu and gestural interaction.
  • UIs and Widgets: Information panels present precise textual or visual robot and environmental status, including altitude, measured length, and task-program network graphs.
  • UIs and Widgets: Labels, annotations, controls, and handles provide object information or user mechanisms for interacting with virtual and robotic elements.
  • Spatial References and Visualizations: Spatial references overlay data on physical referents using points, paths, areas, boundaries, and more complex color or force maps.
  • Embedded Visual Effects: Embedded visual effects place graphical content directly in the real world without necessarily encoding data, including anthropomorphic effects, virtual replicas, and texture mapping.
  • Embedded Visual Effects: Anthropomorphic effects add robot bodies, faces, human avatars, or character animations, while replicas visualize simulated behaviors, environments, or hidden spaces.
  • Embedded Visual Effects: Texture mapping overlays interactive terrain, landscapes, game elements, colored textures, or surface effects onto shape-changing interfaces and surrounding backgrounds.

8 INTERACTIONS

The survey classifies AR-robotics interaction by interactivity level and modality. Systems range from visual output without input to direct manipulation, using touch, controllers, gestures, gaze, voice, proximity, and tangible interaction.

  • Level of Interactivity: Interaction level ranges from no interaction, through implicit and indirect manipulation, to explicit direct physical manipulation.
  • Level of Interactivity: No-interaction systems provide visual output independently of user action, whereas implicit interaction responds to user position or proximity.
  • Level of Interactivity: Indirect manipulation uses remote input such as pointing, selecting, drawing, or body motion to determine robot actions.
  • Level of Interactivity: Direct physical manipulation includes touch, embodied interaction, deformation, grasping, robot manipulation, and physical demonstration.
  • Interaction Modalities: Interaction modalities include tangible interaction, touch, pointers and controllers, spatial gestures, gaze, voice, and proximity.
  • Interaction Modalities: Touch supports precise robot control and programming, while controllers provide tactile feedback that reduces manipulation effort.
  • Interaction Modalities: Spatial gestures manipulate waypoints, menus, and swarm robots; gaze can accompany gestures or point to locations in 3D space.
  • Interaction Modalities: Voice executes robot-operation commands, especially in co-located settings, while proximity communicates implicitly through user behavior and position.

9 APPLICATION DOMAINS

AR and robotics research spans domestic, industrial, entertainment, educational, social, creative, medical, telepresence, mobility, and other application domains. Industry is the largest category, while AR supports workload reduction and robot programming in manufacturing and maintenance contexts.

  • Application domains: The taxonomy also identifies social interaction, design and creative tasks, medical and health, telepresence, mobility, and transportation as application clusters.Figure 10 and the appendix provide detailed use cases and references for each category.
  • Application domains: Industry is the largest application category, with 166 papers spanning manufacturing, maintenance, safety, inspection, automation, and teleoperation.Examples include joint assembly, robot maintenance, remote repair, monitoring, calibration, debugging, and factory automation.
  • Application domains: Domestic and everyday use comprises 35 papers covering home automation, item delivery, photography, wearables, assistance, companionship, and navigation.
  • Application domains: Entertainment includes 32 papers on games, storytelling, immersive media, music, and interactive experiences.
  • Application domains: Education and training includes 22 papers involving remote instruction, military and robotic training, tangible learning, programming education, and posture analysis.

10 EVALUATION STRATEGIES

The survey organizes evaluation strategies for AR and robotics into demonstration-based, technical, and user evaluations. These categories cover proof-of-concept presentation, internal system measurements, and studies of user interaction and performance.

  • Evaluation strategies: Evaluation through demonstration assesses how a system may work by showing applications, proof-of-concept systems, workshops, focus groups, case studies, or conceptual ideas.
  • Evaluation strategies: Technical evaluation measures internal system performance using latency, tracking accuracy, success rate, or comparisons with other systems.
  • Evaluation strategies: User evaluation measures system effectiveness through user studies, including NASA TLX, quantitative and qualitative lab studies, interviews, and questionnaires.
  • Evaluation strategies: Some studies combine user evaluation with demonstration or technical evaluation, and observational studies gather feedback through researcher observation.

11 DISCUSSION AND FINDINGS

The taxonomy highlights recurring design and interaction patterns across AR-HRI research. Co-located interaction without contact, labels and annotations, and pointers, controllers, and spatial gestures are especially common.

  • Summary of findings: Figure 11 summarizes paper counts for characteristics across the taxonomy’s dimensions, supporting comparisons of recurring strategies and gaps.
  • Common patterns: 251 papers use co-located users and robots with distance, making co-located interaction without direct contact the preferred proximity category.The survey associates this pattern with interacting with robots through AR without directly touching them, including manipulation programming.
  • Common patterns: 241 papers use labels and annotations, the most common UI and widget choice for displaying information about robots, environments, and objects.
  • Common patterns: Pointers and controllers appear in 136 papers, while spatial gestures appear in 116 papers, making them the most common interaction modalities.
  • Common patterns: Touch and tangibles appear in 66 and 68 papers respectively, indicating continued use of these modalities in AR-HRI systems.

12 FUTURE OPPORTUNITIES

The paper identifies practical and deployment challenges that must be addressed for AR-HRI to become more widespread. It emphasizes reliable technology, user-centered design, and evaluation beyond controlled laboratory settings.

  • Making AR-HRI Practical and Ubiquitous: AR-HRI faces challenges in realistic superimposition, occlusion, tracking, display and tracking technology, reliability, visual clutter, and critical-information obstruction.Failures in safety-critical settings may expose users to device malfunctions, misalignment, or inappropriate content overlap.
  • Making AR-HRI Practical and Ubiquitous: Improved display and tracking technologies could broaden practical applications requiring precise alignment, including robotic-assisted surgery and medical applications.
  • Deployment and Evaluation In-the-Wild: Most prior AR-HRI research was conducted in controlled laboratories, leaving direct applicability to real-world situations questionable.
  • Deployment and Evaluation In-the-Wild: Outdoor contexts such as search-and-rescue and construction may require different technical requirements from indoor scenarios, including visibility and tracking coverage.
  • Deployment and Evaluation In-the-Wild: User-centered design should use repeated cycles of interviews, prototyping, and evaluation while accounting for HMD limitations such as resolution, comfort, battery life, weight, field of view, and latency.

Designing and Exploring New AR-HRI

AR-HRI is presented as a space for re-imagining robot appearance, enabling immersive authoring, real-time data visualization, and explainable, explorable robot behavior.

  • Re-imagining Robot Design: AR-HRI can re-imagine robot designs beyond physical constraints, including fictional appearances, animated behavior, virtual skins, and dynamic visual transformations.Examples include humanoid or animated nonhumanoid robots, robots that disappear or change color, and expressive telepresence drones.
  • Immersive Authoring and Prototyping: Immersive authoring and prototyping tools could lower the software and hardware barriers that make functional AR-HRI systems time-consuming to develop.Direct manipulation within AR is proposed to reduce movement between computer screens and the real world and broaden design exploration.
  • Real-time Embedded Data Visualization for AR-HRI: Real-time embedded data visualizations could aggregate internal, external, and goal-related information to support operators’ complex decision-making.The paper highlights embedding data directly into the real world and combining visualizations with world-in-miniature views for large-area tasks such as drone search-and-rescue navigation.
  • Explainable and Explorable Robotics: AR-based explainable and explorable robotics could make sensing, decision-making, and actions visible while allowing users to explore how robot behavior changes.Users might inspect recognized obstacles, path choices, or updated paths after physically manipulating obstacles.

Novel Interaction Design Enabled by AR-HRI

AR-HRI broadens interaction design through expressive multimodal inputs and tighter coupling between virtual overlays and programmable physical actuation.

  • Natural Input Interactions with AR-HRI Devices: HMD-based AR-HRI supports casual, location-independent inputs including gesture, gaze, head, voice, and proximity-based interaction.The paper gives mid-air drawing and other expressive gestures as examples for controlling drone swarms across several application contexts.
  • Natural Input Interactions with AR-HRI Devices: Combining voice, gaze, gesture, and AR feedback can clarify ambiguous references and help users express or register robot commands.Gaze and gesture can disambiguate phrases such as “this” and “there,” while visual feedback can clarify user intent.
  • Further Blending the Virtual and Physical Worlds: AR and physically reconfigurable robots could couple virtual pixels with physical atoms across robotic furniture, wearable robots, haptic devices, shape-changing displays, and actuated interfaces.The paper envisions programmable augmentation and actuation that makes the physical world more dynamic and reconfigurable.
  • Future Research Directions: The survey identifies eight future directions spanning technical challenges, in-the-wild evaluation, robot redesign, authoring, visualization, explainability, interaction techniques, and programmable augmentation.These directions frame AR-HRI as both a visual-augmentation field and a broader interaction-design space.

Appendix Table: Full Citation List

The appendix organizes cited work across the paper’s major sections and application domains, including domestic use, mobility, search and rescue, and evaluation strategies.

  • Application Citations: Domestic and everyday-use examples include home automation, item movement and delivery, cartoon overlays while sweeping, and multipurpose tables.The appendix also lists photography drones, mid-air advertising, haptic interaction, fog screens, and head-worn projectors for sharable AR scenes.
  • Application Citations: Mobility and search-and-rescue citations include projected vehicle guidance, pedestrian interaction, passenger displays, augmented wheelchairs, navigation maps, and collaborative ground search.The list also references target detection and notification within search-and-rescue work.
Loading 2203.03254v1…