Source-linked AI summary

Teleoperation of Humanoid Robots: A Survey

Kourosh Darvish, Luigi Penco, Joao Ramos, Rafael Cisneros, Jerry Pratt, Eiichi Yoshida, Serena Ivaldi, Daniele Pucci

arXiv:2301.04317v1cs.RO

TL;DR

Humanoid teleoperation is needed because autonomous robots remain insufficiently socially and physically competent for many complex or hazardous environments, while humanoid robots can perform human-oriented tasks. This survey organizes the field’s systems, methods, communication issues, evaluation practices, and applications, concluding that robust teleoperation remains constrained by dynamics, latency, and evaluation challenges.

  • Problem

    Autonomous solutions remain far from socially and physically competent robot behaviors, while humanoid teleoperation faces complexity in unstructured environments with limited communication.

  • Method

    The paper conducts an extensive survey of humanoid teleoperation systems, architectures, devices, modeling, control, communication, evaluation, and applications.

  • Results

    The survey synthesizes technological and methodological advances and identifies unresolved challenges involving dynamic environments, communication delays, robot stability, and system evaluation.

  • Takeaways & Limitations

    Reliable humanoid teleoperation can support hazardous-environment and space tasks where human operation is unsafe or autonomous robots are not yet sufficiently reliable.

Abstract

from arXiv · show

Teleoperation of humanoid robots enables the integration of the cognitive skills and domain expertise of humans with the physical capabilities of humanoid robots. The operational versatility of humanoid robots makes them the ideal platform for a wide range of applications when teleoperating in a remote environment. However, the complexity of humanoid robots imposes challenges for teleoperation, particularly in unstructured dynamic environments with limited communication. Many advancements have been achieved in the last decades in this area, but a comprehensive overview is still missing. This survey paper gives an extensive overview of humanoid robot teleoperation, presenting the general architecture of a teleoperation system and analyzing the different components. We also discuss different aspects of the topic, including technological and methodological advances, as well as potential applications. A web-based version of the paper can be found at https://humanoid-teleoperation.github.io/.

I. INTRODUCTION

Humanoid teleoperation combines human expertise with humanoid physical capabilities for hazardous, complex, and socially interactive tasks. The survey addresses the field’s multidisciplinary challenges through a structured review of systems, control, communication, evaluation, and applications.

  • Motivation: Teleoperated humanoids can replace humans in hazardous environments while performing mobility, manipulation, inspection, maintenance, and human interaction tasks.Relevant settings include construction sites, chemical plants, contaminated areas, space, and telenursing.
  • Challenges: Humanoid teleoperation must coordinate dynamic locomotion, dexterous manipulation, safety, social interaction, haptic feedback, and communication constraints.Humanoid models are highly redundant, nonlinear, hybrid, and underactuated, making teleoperation particularly demanding.
  • Research gap: An up-to-date survey was missing because the prior broad survey was a decade old and humanoid teleoperation remained an active, far-from-solved research area.Earlier work addressed humanoid robotics, dynamics, control, motion generation, interfaces, metrics, and interactive robots.
  • Survey scope: The survey covers teleoperation systems and devices, modeling and control, assisted teleoperation, non-ideal communication, evaluation, applications, and associated challenges.Its architecture includes human sensing, retargeting, communication, robot control, and feedback.
  • Terminology: Teleoperation describes remote human control of a robot, while telexistence describes virtually existing at a remote location through an avatar; the survey uses the terms interchangeably.The survey considers a real environment and a surrogate humanoid robot as the avatar.

B. Human Sensory Measurement Devices

Humanoid teleoperation uses multiple sensing technologies to capture human motion, interaction forces, and physiological state. The survey describes wearable, optical, exoskeletal, locomotion, EMG, and EEG-based approaches, along with their practical constraints.

  • Measurement requirements: Natural interfaces are needed because keyboard, mouse, and joystick commands may not provide enough control authority for complex humanoid teleoperation.The requirement becomes stronger as the user seeks more detailed control over the robot.
  • Human kinematics and dynamics measurements: IMU-based wearables measure human motion through segregated sensors or integrated sensor networks distributed across the body.They offer high accuracy and frequency without occlusion, but magnetic disturbances and sensor displacement can reduce accuracy.
  • Human kinematics and dynamics measurements: Optical sensors capture human motion using depth sensors, optical motion capture, RGB cameras, or stereo cameras to generate a body skeleton.The passage distinguishes active sensing from passive camera-based sensing.
  • Human kinematics and dynamics measurements: Exoskeletons estimate human link poses and velocities by combining an exoskeleton model and encoder data with forward kinematics.They are often used to track user motion in bilateral teleoperation.
  • Human kinematics and dynamics measurements: Treadmills support gait measurement on even terrain, whereas cockpit-like setups address locomotion retargeting on uneven terrain.Force-torque sensors can be integrated into ground plates or shoes to estimate interaction forces.
  • Human physiological measurements: EMG measures muscle activity and can estimate human effort, muscle forces, or stiffness while anticipating motion before force generation.Low signal-to-noise ratio is identified as the main barrier to desirable surface-EMG performance.
  • Human physiological measurements: EEG sensors identify user mental state by measuring brainwave-related electrical fluctuations through scalp electrodes.They are widely used in non-invasive brain-machine interfaces.

C. Feedback Interfaces: Robot to Human

Feedback interfaces convey the robot’s state and remote environment through visual, auditory, haptic, and balance-related signals. Their design must account for situational awareness, forceful interaction, latency, motion sickness, and operator responsiveness.

  • Visual feedback: Visual feedback helps operators localize themselves, people, and objects in the remote environment using GUIs and robot camera information.DRC teams used GUIs to remotely pilot robots through competition tasks.
  • Visual feedback: VR camera feedback can be effective but may cause motion sickness when unstabilized images, bandwidth limits, and delays affect locomotion.Human reaction time is approximately 250-300 ms for visual input and around 100-150 ms for haptic information.
  • Haptic feedback: Haptic feedback is required alongside visual feedback when power manipulation or human interaction makes robot dynamics and contact forces crucial.It exploits the operator’s motor skills to augment robot performance.
  • Haptic feedback: Force, tactile, vibrotactile, and air-pressure interfaces communicate forces, touch, texture, temperature, directions, or user alerts.Kinesthetic force feedback can be delivered through exoskeleton-like or cable-driven interfaces.
  • Haptic feedback: TELESAR V combines haptic feedback types to provide complete cutaneous sensations through force, vibration, and temperature stimuli.These stimuli are treated as sufficient combinations for reproducing cutaneous sensations without directly touching the object.
  • Balance feedback: Balance feedback transfers information about disturbances affecting robot dynamics or stability rather than directly mapping disturbance forces to the operator.Examples include a vibrotactile belt and a torso-force feedback interface.
  • Auditory feedback: Auditory feedback supports remote communication, situational awareness, sound localization, and collision detection.It is delivered through headphones or one or more speakers, with sound localization using multiple microphones.

D. Graphical User Interfaces (GUIs)

GUIs support both supervision and robot commands by combining task controls with visualized robot, plan, and perception information. They also accommodate different autonomy levels, from manual corrections to high-level task specification.

  • D. Graphical User Interfaces (GUIs): GUIs let operators supervise execution and manually correct tasks using panels that display robot states, motion plans, and perception data.DRC interfaces included 2D and 3D visualization, current and goal states, and hardware-driver status.
  • D. Graphical User Interfaces (GUIs): Operators can guide perception by annotating image regions, point-cloud positions, and target poses for robot arms, configurations, or manipulation objects.These annotations help fit object models to point clouds and define goals at different autonomy levels.
  • D. Graphical User Interfaces (GUIs): GUIs encode frequent high-level robot tasks as state machines or task sequences, using tools such as RViz and Choreonoid with custom plugins.Many of these functions can also be integrated with VR and AR devices.
  • D. Graphical User Interfaces (GUIs): Teleoperation retargeting maps perceived human actions from A into robot actions in A′ while minimizing intent-action differences under scenario constraints.The mapping reflects the robot’s autonomy and the human’s authority in the teleoperation scenario.
  • D. Graphical User Interfaces (GUIs): Online retargeting and control are difficult because humanoids have nonlinear, hybrid, underactuated, high-degree-of-freedom dynamics and uncertain models and environments.The survey identifies model-based optimal control architectures as the foremost technique used to address these challenges.

A. Modeling

Humanoid modeling ranges from complete multi-body dynamics to simplified models that support intuitive, real-time planning and control. The survey describes rigid-body configuration, velocity, inertial, actuation, gravity, and contact terms, then introduces the LIPM and DCM formulations.

  • A. Modeling: A humanoid is modeled as n + 1 rigid links connected by n one-degree-of-freedom joints, with configuration q combining floating-base pose and joint angles.The floating base is usually associated with the pelvis, and link poses are obtained through forward kinematics.
  • A. Modeling: The complete model uses q for configuration and ν for base and joint velocities, while forward kinematics computes each link frame pose.The velocity vector has dimension n + 6 because it includes floating-base linear and angular velocity.
  • A. Modeling: The dynamics equation represents inertia, Coriolis and centrifugal effects, gravity, actuator torques, and contact wrenches in the humanoid model.M(q) is the inertia matrix; C(q,ν), g(q), B, τ, and contact-wrench terms encode the remaining components.
  • A. Modeling: Simplified models are used to derive intuitive control heuristics and enable real-time lower-order motion planning.The inverted pendulum approximation assumes a support foot connected to the center of mass by a variable-length link.
  • A. Modeling: Under constant CoM height, the Linear Inverted Pendulum Model describes CoM motion relative to the LIPM base, often identified with the robot’s ZMP or CoP.Its dynamics separate stable and unstable modes, with the unstable mode described by the Capture Point or DCM.

B. Retargeting & Planning

Retargeting and planning transform teleoperation inputs into robot link and locomotion references while accounting for task requirements, obstacles, contacts, and robot constraints. Approaches span search-based and reactive planning, plus lower-body, upper-body, and whole-body retargeting.

  • B. Retargeting & Planning: The retargeting block morphs human commands or measurements into robot link references and locomotion references such as footstep locations, timings, and support-leg assignments.Retargeting policies vary with task and system requirements and may be selected online as a classification problem.
  • B. Retargeting & Planning: Teleoperation inputs range from high-level base or end-effector goal regions to low-level CoM, base, footstep, end-effector, or whole-body motion references.At low level the user handles obstacle avoidance, whereas high-level planning and retargeting address feasible paths.
  • B. Retargeting & Planning: High-level planning must find feasible obstacle-free or step-over paths while planning humanoid footsteps or arm motions toward GUI-specified goals.Search-based methods include A*, D* Lite, RRT variations, and dynamic programming.
  • B. Retargeting & Planning: Exhaustive search is inefficient for real-time execution, so heuristics prune the search space and produce greedy searches.Method selection considers completeness, global optimality, and real-time replanning in dynamic environments.
  • B. Retargeting & Planning: Reactive retargeting and path planning can be formulated as optimization or dynamical-system problems, including MPC with foot poses as continuous QP decisions.Adding end-effector rotation or obstacle avoidance increases the optimization problem’s complexity.
  • B. Retargeting & Planning: Retargeting approaches are grouped into lower-body footstep generation, upper-body retargeting, and whole-body retargeting.Lower-body methods use teleoperated CoM and base references to generate foot locations and timings.
  • B. Retargeting & Planning: Upper-body retargeting maps human limb motion in task or configuration space and solves inverse kinematics under robot constraints.Whole-body retargeting maps complete human motion, may include contact constraints, and provides outputs to the stabilizer for feasibility and stability.
  • B. Retargeting & Planning: The stabilizer adapts retargeted references to improve the stability and balance of the robot’s centroidal dynamics.It receives reference trajectories from retargeting and planning because those trajectories may destabilize the robot.

1) ZMP approach:

Stability control corrects retargeted references before whole-body control, using criteria such as ZMP, DCM, and contact-wrench regulation. The ZMP approach tracks a desired support-polygon trajectory but has stated disturbance and terrain limitations.

  • 1) ZMP approach:: The ZMP approach keeps the robot’s ZMP inside the support polygon and computes a desired ZMP from kinematic footstep trajectories.During single support it remains near the middle of the supporting foot; during double support it moves smoothly between support legs.
  • 1) ZMP approach:: Under high disturbances, the ZMP controller does not adapt footstep locations online to prevent falling and extends poorly to non-flat, rotating, or sliding support feet.These limitations constrain the approach beyond nominal support conditions.
  • 1) ZMP approach:: DCM control stabilizes the unstable DCM dynamics while also regulating ZMP, extending the capture-point concept to three-dimensional uneven terrain.The CoM dynamics converge toward the DCM value, motivating stabilization of the DCM state.
  • 1) ZMP approach:: A contact-wrench controller chooses desired wrenches at each contact point to enhance robot stability under contact constraints.The approach is related to momentum-based control and has been used as a stability-augmentation criterion.
  • 1) ZMP approach:: Whole-body control receives stabilizer-corrected human references and outputs robot joint states or torques through QP, LQR, or MPC formulations.Its outputs may include joint angles, velocities, accelerations, and/or joint torques.
  • 1) ZMP approach:: Inverse kinematics and inverse dynamics solve constrained robot-control problems for task poses, velocities, torques, and feasible contacts.Redundant inverse kinematics commonly uses constrained QP formulations, while inverse dynamics minimizes motion-task and contact-wrench errors.

4) Momentum-based control:

Momentum-based control formulates humanoid whole-body behavior through desired momentum-rate tracking, configuration acceleration, and ground-reaction-force computation. Teleoperation remains limited by unresolved integration challenges, simplified models, computational demands, and real-world environmental variability.

  • Momentum-based control: The controller computes configuration-space acceleration and ground reaction forces so the robot follows a desired whole-body momentum rate.Newton–Euler dynamics relate centroidal-momentum change to the sum of external wrenches, enabling a quadratic-program formulation.
  • Low-level control: Low-level joint torque tracking can produce different behavior in simulation and on the real robot.The joint controller converts higher-level position, velocity, or torque commands into motor current or voltage.
  • State estimation: Humanoid state estimation combines joint encoders, dynamics models, force/torque sensors, and model-based methods such as Kalman filters.Base estimation may assume a grounded, nonslipping contact link, but modeling errors propagate through integrated odometry.
  • Challenges and future directions: Whole-body retargeting, planning, stability, and control are not yet effectively integrated for agile teleoperation in unstructured real-world environments.Dividing retargeting and planning can be ineffective, while simplified models impose assumptions such as fixed CoM height and at least one nonslipping contact.
  • Challenges and future directions: MPC can combine whole-body retargeting and planning as constrained optimization, but its nonlinear, non-convex problems are computationally demanding and vulnerable to local minima.Real-world deployment also requires predicting future human motion and differences between human and robot terrain.
  • Challenges and future directions: Data-driven approaches that account for robot dynamics and stability, with explicit safety enforcement, are presented as promising for varied teleoperation settings.The survey cites neural-network, reinforcement-learning, and cycle-consistency approaches for related retargeting and mobile-robot problems.

A. Shared Control

Shared-control and bilateral strategies distribute authority between the human and humanoid robot, using autonomy, feedback, and whole-body regulation to improve task execution, safety, and telepresence. The robot may assist with balance, intent-based control, safeguarding, or dynamic feedback while preserving operator involvement.

  • Shared Control: Full manual control can cause clunky motions, failures, and repeated attempts because humanoid operators must manage hands, feet, and balance simultaneously.Shared control modifies operator input by sharing control authority with robot autonomy to improve performance or safety.
  • Shared Control: Intent confidence can blend operator input u and robot input r, increasing preference for r as the robot approaches a predicted goal.The confidence function uses distance d to the goal and threshold D, with confidence reaching zero beyond D.
  • Shared Control: With full task autonomy, the operator supervises the robot, intervening when unexpected or uncovered situations arise.This supervisory strategy was used in the DARPA Robotics Challenge to guide robots through failures when necessary.
  • Shared Control: A safeguarding system can modify or override operator commands in hazardous situations while monitoring collision, rollover, and system health.In benign situations, the operator retains full vehicle-motion control.
  • Whole-body bilateral teleoperation: Bilateral teleoperation sends kinodynamic references to the robot and returns kinesthetic feedback about motion reproduction and external disturbances.Upper-body bilateral control can be paired with a separate lower-body balancing controller and admittance control.
  • Whole-body bilateral teleoperation: Whole-body bilateral teleoperation maps human kinematic and dynamic references to the robot while feeding back whole-body dynamics, despite the robot’s unstable dynamics.Feedback strategies may represent CoM state, tipping proximity, or external forces through waist forces and vibrotactile cues.
  • Whole-body bilateral teleoperation: Assisted teleoperation balances two priorities: following human task commands and maintaining bipedal stability.Some approaches regulate balance autonomously, while others require the operator to detect destabilization through feedback and adapt motion.

D. Impedance Control

Impedance-based strategies regulate robot stiffness and damping to support dexterous interaction, while communication delays and bandwidth limitations complicate bilateral teleoperation. Passivity-based and shared-control methods address stability, but variable network conditions remain challenging.

  • D. Impedance Control: Autonomous impedance control adjusts joint stiffness and damping according to manipulation loading conditions.It uses a virtual mass-spring-damper system, whereas admittance control generates references from the corresponding dynamics.
  • D. Impedance Control: Tele-impedance sends an impedance profile to the robot through a brain-machine interface using non-intrusive position and EMG measurements.The approach links operator arm measurements to remote compliance behavior.
  • Communication challenges: Communication channels introduce transmission delays and information distortion that affect teleoperation stability and performance.These effects are especially consequential when human and robot exchange force-reflected information.
  • Communication challenges: Supervisory commands, stored subroutines, and predictive displays avoid closing the force loop and remain useful for complex humanoid robots under delay.The DARPA Robotics Challenge used either a continuous 9600 bps bidirectional link or an intermittent 300 Mbps unidirectional link, boosting robot autonomy.
  • Communication challenges: Force reflection is required for a strong sense of telepresence, but time delay can destabilize the system; shared control can help manage it.Round-trip delays around 100 ms can induce instability, unless motion bandwidth is reduced or advanced methods are used.
  • Communication challenges: Passivity and network-theoretic methods can stabilize teleoperation with constant delays, but packet-switched networks add variable delays and discrete-time exchange difficulties.These limitations motivate analysis beyond the constant-delay assumption used for simple communication channels.

B. Distortion

Network-based models represent teleoperation communication through interconnected ports relating mechanical effort and flow. Passivity and scattering theory provide tools for analyzing stability and energy behavior in nonlinear teleoperation systems.

  • B. Distortion: An n-port models a network component through the relationship between effort f, such as force or voltage, and flow v, such as velocity or current.The communication channel can therefore be represented as interconnected n-ports using a mechanical–electrical analogy.
  • B. Distortion: For linear time-invariant 2-port networks, flows and efforts can be related in the frequency domain by a hybrid matrix.Control can modify the matrix characteristics to address communication-channel difficulties.
  • B. Distortion: For nonlinear humanoid bilateral teleoperation, frequency-domain analysis is replaced by equivalent analysis using Hilbert-space efforts and flows in a Hilbert network.This extends network-based reasoning beyond linear-system assumptions.
  • B. Distortion: Passivity analyzes nonlinear-system stability by requiring that a system may dissipate energy but cannot increase its total energy.With force inputs and velocity outputs, power is represented as P(t) = f^T(t)v(t).
  • B. Distortion: A lossless n-port is one in which no power is dissipated.Otherwise, non-negative dissipation accounts for power that is not stored.
  • B. Distortion: Scattering theory relates network effort and flow through a scattering operator, while combined passive systems add their individual stored energies and dissipations.The formalism imitates physical systems obeying energy conservation.

E. Stability of Time-delayed Humanoid Robot Teleoperation

Time delay affects humanoid teleoperation stability, balance, velocity, and performance. The survey reviews passivity, stability, delay-dependent balance, and compensating strategies, while noting unresolved challenges and trade-offs.

  • Passivity and stability: Passivity does not guarantee good performance, and stability and transparency remain conflicting objectives.Energy tanks can enforce passivity and stability in whole-body control, but may not preserve teleoperation quality.
  • Delay effects and compensation: Delay is relevant to humanoid robot balance and velocity, with larger delays requiring higher autonomy.For mid-range delays, move-and-wait can mitigate delay but reduces robot velocity and task efficiency; its use is highly limited for bilateral manipulation.
  • Delay effects and compensation: Predictive approaches are envisioned to compensate for large delays by predicting human motion, robot motion, interaction forces, and visual feedback.The survey identifies prediction on both human and robot sides as a future approach rather than an established solution.
  • Delay effects and compensation: Lower center-of-mass height reduces the tolerated delay for maintaining humanoid balance.A lower center of mass implies faster linear inverted pendulum dynamics and therefore a lower allowable delay; tolerance also depends on human performance, retargeting, control, and actuators.
  • Evaluation considerations: System deployment requires human-centered design and evaluation metrics that identify problems, limitations, and design needs.The survey discusses usability, situation awareness, and workload as evaluation considerations for teleoperation systems.

3) Workload:

Workload and related human-centered metrics are important for designing and evaluating humanoid teleoperation systems. The survey connects these measures to autonomy selection, coordinated system behavior, and deployment readiness, while highlighting gaps in systematic evaluation and non-expert usability.

  • 3) Workload:: Workload varies across operators and tasks, is multidimensional, and includes mental, physical, information, perceptual, and communication loads.This variability makes workload difficult to measure reliably.
  • 3) Workload:: Physiological signals, oxygen consumption, heart rate, and task performance provide approaches for assessing physical or mental workload.Examples include cardiac, eye, respiratory, speech, and brain activity measurements.
  • Human-centered autonomy: Intermediate human-centered autonomy can enhance situation awareness, reduce workload, and improve performance against failures.High automation may degrade situation awareness, whereas low autonomy in complex tasks intensifies workload.
  • Evaluation challenges: Systematic evaluation remains limited, with fewer behavioral and physiological measurements used in humanoid teleoperation studies.ANA Avatar XPRIZE evaluated task performance and subjective measures across locomotion, manipulation, and social interaction.
  • Evaluation challenges: Latency can arise in perception, communication, control, or actuation, so poor performance requires examining interconnections among components.Unsynchronized arms, torso, locomotion, hands, and social cues can further degrade performance and user experience.
  • Evaluation challenges: Non-expert usability remains a challenge because state-of-the-art systems often rely on one expert operator’s training and expertise.The survey calls for evaluation metrics that support deployment meeting users’ needs.

VII. APPLICATIONS AND PERSPECTIVES

The survey identifies promising humanoid teleoperation applications in social interaction, disaster response, healthcare, space, and domestic environments. These settings expose practical requirements involving safety, mobility, communication, autonomy, reliability, and human factors.

  • Social interaction: Humanoid teleoperation may convey social cues and support presence and manipulation in remote social-interaction scenarios.These needs became especially salient when in-person meetings and visits were disrupted during the COVID-19 pandemic.
  • Hazardous environments: Disaster response requires teleoperated robots for surveillance, search and rescue, and manipulation in hazardous, dynamic, and unstructured environments.Bipedal locomotion offers mobility, agility, and a wide range of motion, while response time and technology readiness are critical.
  • Deployment constraints: Humanoid robots remain expensive, and hardware or control software often lacks adequate handling of control failures, balance loss, or incorrect interactions.These constraints have limited study and deployment to a restricted robotics community.
  • Healthcare: Teleoperated humanoids could improve healthcare-worker safety while requiring supervision and avoidance of patient discomfort or interference with other workers.Clinical deployment must balance task performance with patient and workplace considerations.
  • Space applications: Space teleoperation supports servicing, maintenance, experiments, exploration, and construction where human access is costly or unsafe and autonomy is not yet sufficiently reliable.Communication latency, bandwidth, unknown target dynamics, and human factors create additional challenges; greater distances require trade-offs among autonomy, risk, and efficiency.
  • Domestic environments: Domestic applications include care, housekeeping, shelf restocking, teleducation, hotel guidance, and teletourism in environments built for human use.Their intrinsic unstructuredness and human-oriented design motivate the use of humanoid platforms.
Loading 2301.04317v1…