Source-linked AI summary

SEED-UMI: Sharing the Exoskeleton between human and robot for onE-to-one Dexterous demonstration

Tengbo Yu, Jiahao Wu, Daohan Li, Bingxu Chen, Hao Liu, Xiaojian Ma, Hangxin Liu

arXiv:2609.11753v1cs.RO

TL;DR

Dexterous imitation learning is limited by the difficulty of collecting contact-rich demonstrations that transfer faithfully across differing human and robot embodiments. SEED-UMI shares a kinematically matched exoskeleton between them, using paired robot replay for retargeting and raw exoskeleton-centric wrist observations for policy learning. Across five tasks, it collects data 3.0× faster than exoskeleton teleoperation and reaches 70.0% mean success.

  • Problem

    Dexterous imitation learning lacks reliable contact-rich demonstration transfer because human and robot embodiments differ in kinematics, sensing, and contact behavior.

  • Method

    SEED-UMI shares a kinematically isomorphic exoskeleton between human and robot, using paired replay to refine retargeting and raw exoskeleton-centric wrist observations for policies.

  • Results

    Across five dexterous tasks, data collection is 3.0× faster than exoskeleton teleoperation and paired fine-tuning achieves 70.0% mean success.

  • Takeaways & Limitations

    Shared physical measurement turns retargeting and visuomotor policy learning into paired cross-embodiment supervision for more scalable dexterous robot learning.

  • Takeaways & Limitations

    The system remains specialized to one target hand and still requires engineering for different link layouts, joint limits, actuation ranges, calibration, and paired regression targets.

Abstract

from arXiv · show

Imitation learning for dexterous hands is bottlenecked by the difficulty of collecting contact-rich demonstrations that transfer faithfully to the robot. Prior wearable-exoskeleton systems record only on the human side and retarget via open-loop mappings calibrated in free space, which degrade under contact. We present SEED-UMI, a framework in which both the human and the robot wear the same exoskeleton: joint encoders become a physically shared measurement, and wrist cameras mounted to the exoskeleton observe the same outer mechanism during both human data collection and robot policy rollouts. This turns retargeting into paired cross-embodiment supervision and lets policies train directly on raw exoskeleton-centric wrist images, without segmentation or inpainting. On five contact-rich tasks, SEED-UMI achieves 3.0x greater data collection efficiency than exoskeleton-based teleoperation and a 70.0% average rollout success rate.

1 Introduction

SEED-UMI addresses dexterous imitation learning’s demonstration-transfer bottleneck by sharing a mechanically matched exoskeleton between human and robot. Paired robot replay and exoskeleton-centric observations support contact-aware retargeting and direct policy learning, yielding faster collection and strong task performance.

  • Motivation: Dexterous imitation remains difficult because high-dimensional hands interact through contacts while human and robot embodiments differ in kinematics, sensing, and contact behavior.Existing pipelines therefore separate human demonstrations from robot execution through increasingly complex retargeting models.
  • Contribution: SEED-UMI uses one kinematically isomorphic exoskeleton for human demonstrations and robot execution, producing shared encoder measurements across 20 degrees of freedom.The exoskeleton’s link lengths, joint axes, and motion ranges are jointly optimized for both embodiments.
  • Method: Paired human–robot encoder discrepancies supervise contact-aware retargeting through motor babbling and paired trajectory fine-tuning.This replaces a one-time free-space mapping with supervision from robot replay under contact.
  • Method: Policies train directly on raw exoskeleton-centric wrist observations, avoiding post-hoc segmentation and generative image translation.The shared wrist-camera arrangement preserves aligned observations across human collection and robot rollout.
  • Results: 3.0× faster collection and 70.0% mean success across five tasks show SEED-UMI’s central empirical outcome.The success rate approaches the 71.7% teleoperation baseline, with advantages on precise and dynamic tasks.
  • Contribution: SEED-UMI reframes retargeting and visuomotor learning as paired cross-embodiment supervision through a shared physical measurement interface.The authors present this interface redesign as a direction for more scalable and efficient dexterous robot learning.

2 Related Work

Prior wearable systems improve human hand measurement but retain a human-only, open-loop transfer process and often require visual translation. SEED-UMI instead mounts the same exoskeleton on the robot, enabling paired supervision from real robot replay without simulation or reinforcement learning.

  • Wearable exoskeletons: Glove-based systems lack one-to-one joint correspondence, while wearable exoskeletons constrain the hand into a robot-matched kinematic structure.DexUMI instruments the exoskeleton with joint encoders and adapts it to multiple robot hands.
  • Wearable exoskeletons: SEED-UMI converts prior one-shot free-space regression into supervised learning with paired cross-embodiment data by mounting the same exoskeleton on the robot.Earlier systems generally mounted the exoskeleton only on the human side.
  • Visual alignment: Inpainting-based approaches and DexUMI use generative models to bridge embodiment appearance gaps, whereas SEED-UMI reduces that gap through shared exoskeleton-mounted cameras.The shared camera mounting aligns the observed outer mechanism across embodiments.
  • Encoder-to-motor mapping: SEED-UMI obtains real-world retargeting supervision by replaying demonstrations on the robot through the shared exoskeleton, requiring neither simulation nor reinforcement learning.This positions the method between free-space regression and simulation-trained low-level control.

3 Hardware: A Shared Exoskeleton

SEED-UMI’s hardware makes encoder readings a shared cross-embodiment measurement by matching contact-facing geometry and the encoder frame on human and robot sides. The design is deliberately optimized for one fully actuated target hand rather than cross-platform transfer.

  • Shared hardware: Matching exoskeleton geometry and encoder frames lets discrepancies under identical commands supervise the residual command-to-joint mapping.Cuffs and mounts are adapted to each embodiment while the contact-facing structure remains matched.
  • Design scope: The exoskeleton is co-designed for a single target hand, allowing identical contact geometry and 1:1 joint correspondence instead of cross-platform transferability.Link lengths, joint axes, and range limits can therefore be fixed to the robot’s kinematic tree.

E.1 Isomorphic kinematic skeleton.

The skeleton mirrors the target hand’s active degrees of freedom through matched links, axes, and motion limits. It measures 20 independently controlled joints across five fingers.

  • Isomorphic kinematic skeleton: 20 independently measured revolute joints mirror the target hand’s five-finger kinematic structure.Link lengths, joint axes, and hard range-of-motion stops match the robot, so reachable operator configurations match robot-reachable configurations.

E.3 Identical contact geometry on both sides.

SEED-UMI uses matched exoskeleton geometry and sensing frames across human and robot embodiments. This shared physical interface makes encoder discrepancies informative about contact-dependent effects.

  • E.3 Identical contact geometry on both sides.: Figure 4 depicts collision avoidance during operation and the mechanism of the exoskeleton linkage.The caption does not report a quantitative comparison.
  • E.3 Identical contact geometry on both sides.: Shared link lengths, joint axes, dorsal linkage, and fingertip pads give both embodiments the same external contact geometry.Cuffs and mounts adapt the skeleton without changing joint frames.
  • E.3 Identical contact geometry on both sides.: Encoder readings are shared cross-embodiment measurements, while discrepancies under identical commands supervise the residual command-to-joint mapping.The design uses these discrepancies to capture load-dependent effects such as contact force, friction, and hysteresis.
  • E.3 Identical contact geometry on both sides.: All sensing is mounted on the exoskeleton frame to preserve a common physical frame across human-side and robot-side recordings.The sensor layout targets both downstream action-observation capture and reduced distribution shift.
  • E.3 Identical contact geometry on both sides.: Contactless magnetic rotary encoders at every actuated joint produce a 20-dimensional signal, whose mapping to robot commands is nonlinear.The nonlinearity arises from four-bar linkage routing and manufacturing tolerances.

S.2 Wrist-mounted cameras.

SEED-UMI rigidly mounts wrist sensing to the shared exoskeleton, aligning observations between human collection and robot execution. This supports policy learning from raw wrist images without segmentation or inpainting.

  • S.2 Wrist-mounted cameras.: A dorsal RealSense T265 and ventral 150° fisheye camera are rigidly mounted to the exoskeleton, fixing their intrinsics and encoder-frame extrinsics across embodiments.The sensors provide wrist pose and workspace views.
  • S.2 Wrist-mounted cameras.: Figure 5 evaluates five contact-rich tasks spanning tool stabilization, finger coordination, release timing, table wiping, and thumb-involved actuation.The caption identifies the task coverage but not comparative outcome values.
  • S.2 Wrist-mounted cameras.: Because collection and rollout use the same exoskeleton-camera assembly, policies receive aligned raw observations with only photometric and crop augmentations.The alignment preserves contact pixels and removes the need for hand segmentation or inpainting.
  • S.2 Wrist-mounted cameras.: Robot execution uses the shared encoder signal but requires an encoder-to-command mapping, initialized through motor babbling and refined under real contact.Aligned wrist images are reserved for policy learning rather than proprioceptive mapping.
  • S.2 Wrist-mounted cameras.: The robot exploration protocol covers single-joint sweeps, coupled motions at varied speeds, and selected self-contact poses.These recordings provide the initial mapping data.

M.2 Paired replay and fine-tuning.

Paired replay compares human encoder trajectories with robot-side responses under contact, then fine-tunes the command mapping using those discrepancies. Repeated replay and refinement produces a more contact-robust mapping without manual retuning.

  • M.2 Paired replay and fine-tuning.: Paired replay replays human object-grasping demonstrations on the robot and records robot-side encoder trajectories alongside the human traces.These paired trajectories form the mapping dataset for refinement.
  • M.2 Paired replay and fine-tuning.: A forward model maps robot commands and encoder history to robot-side encoder readings, enabling fine-tuning of the inverse command mapping.The paired dataset supplies the human-side target trajectory.
  • M.2 Paired replay and fine-tuning.: Figure 6 compares human-side and robot-side encoder traces during the AirPods task.The supplied caption identifies the traces and colors but does not state an outcome.
  • M.2 Paired replay and fine-tuning.: The refinement loss matches predicted robot-side encoders to human traces, regularizes commands toward paired replay, and penalizes temporal jumps.The procedure can repeat replay and refinement once to improve contact robustness.
  • M.2 Paired replay and fine-tuning.: Replay need not perfectly reproduce the human trajectory because the human–robot encoder discrepancy itself supplies the correction signal.This makes the discrepancy operational rather than merely diagnostic.

M.3 Policy training and rollout.

SEED-UMI trains dexterous policies on raw wrist observations and encoder vectors from human-collected real data. Direct physical and visual feedback supports precise contact-rich demonstrations compared with separated teleoperation.

  • M.3 Policy training and rollout.: Policies use raw wrist images and 20-dimensional encoder vectors to predict action chunks containing 6-DoF wrist actions and 20-DoF hand commands.Training uses ACT, Diffusion Policy, and π0.5 on human-collected real data without simulation, reinforcement learning, segmentation, or inpainting.
  • M.3 Policy training and rollout.: During AirPods insertion, direct visual and physical feedback lets the operator regulate delicate contacts, whereas teleoperation separates the operator from object contact.The separated interface can over-drive motion and apply excessive force to fragile objects.

5 Evaluation

SEED-UMI is evaluated on five real-world dexterous tasks using autonomous rollouts across teleoperation and two shared-exoskeleton conditions. Paired fine-tuning improves contact-heavy performance and enables substantially faster demonstration collection while retaining similar overall success to teleoperation.

  • Experiment Setup: The evaluation compares teleoperation, SEED-UMI without paired fine-tuning, and SEED-UMI with paired fine-tuning across tool use, fine coordination, dynamic release, and sustained contact.The five tasks are Screw Driving, AirPods Case Insertion, Ball Basket Throwing, Table Cleaning, and Air Freshener Spray.
  • Key Findings: 70.0% average success for P+ nearly matches teleoperation at 71.7% and exceeds P− at 57.3% across five tasks and policy backbones.Each task–backbone–condition combination is evaluated over 20 autonomous robot rollouts.
  • Key Findings: P+ outperforms teleoperation on AirPods Case Insertion at 73.3% versus 65.0% and Ball Basket Throwing at 68.3% versus 60.0%.These tasks involve fine orientation corrections or fast grasp-to-release timing that are difficult through delayed real-time robot control.
  • Key Findings: Paired fine-tuning raises average success by 12.7 percentage points over the babbling-only mapping, with gains of 16.7 points on Screw Driving and 15.0 points on both Table Cleaning and Air Freshener Spray.The comparison isolates the contribution of paired-replay fine-tuning under contact load.
  • Key Findings: 52 successful demonstrations versus 18 in 30 minutes yields nearly 3.0× higher collection throughput for SEED-UMI than teleoperation.Because overall success is nearly matched, the success-normalized useful-data gain remains approximately 2.9×.

6 Conclusions

SEED-UMI uses a shared exoskeleton to connect human demonstrations with robot execution through paired encoder supervision and aligned raw wrist observations. Figure 7 frames the approach as an efficiency–quality trade-off.

  • Figure 7 presents SEED-UMI through an efficiency–quality trade-off.
  • SEED-UMI uses one shared exoskeleton as the interface between human demonstrations and robot execution.The shared mechanism transfers contact-rich human motion into robot actions.
  • Paired encoder supervision transfers natural human motion while preserving aligned raw wrist observations.

7 Limitations

SEED-UMI is not yet fully automated across dexterous hands, and practical deployment remains shaped by engineering and operator-wearability requirements.

  • SEED-UMI still requires engineering choices to adapt across dexterous hands.These choices concern exoskeleton geometry, calibration motions, and paired regression targets.
  • Different link layouts, joint limits, and actuation ranges constrain cross-hand adaptation.
  • Wearability, donning time, and task convenience affect long-horizon operation because operators wear the exoskeleton during collection.
Loading 2609.11753v1…