Source-linked AI summary
Fiber Optic Sensing Glove for High Performance Dexterous Manipulation Capture
J. D. Peiffer, Taylor Niehues, Li Guan, Ziyi Kou, Ergys Ristani
TL;DR
Dexterous manipulation makes accurate hand tracking difficult because vision suffers occlusion and existing gloves may not recover full pose robustly. The paper introduces a multi-core fiber-optic glove with shape registration and curve-constrained inverse kinematics, achieving sub-5 mm fingertip error after one-time tunnel alignment. The system supports dexterous data capture and bimanual virtual teleoperation.
Problem
Accurate full-hand tracking during dexterous manipulation remains difficult because vision is vulnerable to occlusion and existing FBG gloves do not recover out-of-plane motion.
Method
The glove reconstructs full 3D shapes from multi-core FBG fibers, registers them to a common frame, and fits a hand model using shape-constrained inverse kinematics.
Results
4.5-4.9 mm F.MKPE was achieved after tunnel alignment across range-of-motion and manipulation phases, compared with 6.5 to 7.2 mm before alignment.
Takeaways & Limitations
The glove provides a practical tool for capturing complex dexterous behavior and supports real-time bimanual virtual teleoperation.
Takeaways & Limitations
The best-accuracy results depend on one-time mocap-aided tunnel alignment, while model mismatch and fiber or fabric motion remain additional error sources.
Abstract
from arXiv · showhide
Capturing hand pose during dexterous manipulation remains difficult: vision-based methods degrade under occlusion and challenging lighting, while sensorized gloves, though occlusion-free, are prone to drift and magnetic interference and rarely match motion-capture accuracy. We introduce a fiber optic sensing glove for full hand pose tracking that targets these failure modes, using multi-core shape-sensing fibers that capture each fiber's full 3D shape rather than curvature alone. A novel pipeline registers each reconstructed fiber shape to a common hand reference frame, and a new inverse-kinematics solver reconstructs full hand pose at 60 Hz using curve constraints. Benchmarked on a 2-hour dataset of dexterous object manipulation tasks across 5 subjects, the glove achieves 7.2 mm mean fingertip position error against motion capture ground truth, reduced to 4.9 mm by a one-time factory calibration of the fiber routing hub that transfers across users and sessions. These capabilities enable high-fidelity data capture and bimanual virtual teleoperation - both essential to advancing the robotics field.
I. INTRODUCTION
The paper develops a fiber-optic glove for full-hand pose tracking during dexterous manipulation, targeting occlusion, drift, and magnetic-interference limitations. Multi-core fibers, shape registration, and curve-constrained inverse kinematics support high-accuracy capture and virtual teleoperation.
- Vision-based tracking is often ill-posed during manipulation because hands or objects occlude the camera view.
- Single-core FBG gloves generally estimate flexion at selected joints rather than full hand pose because they do not measure bending direction.
- Multi-core FBGs measure curvature magnitude and direction, enabling reconstruction of each fiber’s full 3D shape while largely canceling common-mode temperature effects.
- The proposed pipeline registers reconstructed fiber shapes to the hand and recovers full-hand pose with an inverse-kinematics solver constrained by those shapes.
- The glove is benchmarked on representative manipulation tasks and demonstrated for dexterity data capture and bimanual virtual teleoperation.
II. METHODS
The methods comprise glove hardware, a hand model, fiber-shape-based full-hand pose estimation, and user self-calibration.
- The paper presents hardware, a hand model, full-hand pose estimation from reconstructed fiber shapes, and self-calibration as its main method components.
A. Hardware
The glove combines passive multi-core optical fibers, dorsal fabric routing, a routing hub, and remote FBG interrogation to reconstruct fiber shape.
- Each finger uses an optical fiber fixed near the fingernail and routed through dorsal fabric loops, with a nitinol tube providing protection.
- Each fiber has three sensing cores containing 26 FBG sensors spaced at 1 cm intervals.
- The fibers are reconstructed in local frames and registered to a common routing-hub frame using known tunnel geometry.
- The passive sensors connect through 5 meter patch cables to a remotely placed interrogator tailored for five multi-core fibers.
- Three-core strain measurements provide local curvature and bending direction, which are integrated into fiber shape at 1 mm spatial resolution.
B. Hand Model
The hand model represents wrist and finger kinematics, while registration aligns measured fiber curves with routing-hub tunnels before pose fitting.
- B. Hand Model: The hand is modeled as a kinematic tree with wrist translation, wrist rotation, 22 joint angles, and 21 anatomical landmarks.
- 1) Fiber to Routing Hub Registration:: Fiber shapes are initially local to the first FBG sensor and are registered to pre-measured digit-specific hub tunnel centerlines.
- 1) Fiber to Routing Hub Registration:: The registration algorithm searches a sliding-window fiber segment whose curvature matches the corresponding routing-hub tunnel.
- 1) Fiber to Routing Hub Registration:: Candidate alignments are rigidly transformed using an SVD-based Procrustes solution and selected by the lowest post-alignment point residual.
- 1) Fiber to Routing Hub Registration:: Cubic B-spline upsampling reduces registration jitter, while the returned fiber shapes remain non-upsampled.
2) Inverse Kinematic Solution:
The solver reconstructs full-hand pose by fitting a scaled hand model to 3D fiber shapes using geometric constraints and weighted optimization. Temporal regularization, dorsum constraints, and joint-limit penalties support stable, anatomically plausible solutions.
- Inverse kinematic formulation: Full-hand pose is estimated by fitting a scaled Momentum Hand Model to reconstructed 3D optical-fiber shapes.The model uses subject-specific hand scale and per-digit landmark locators.
- Geometric constraints: Fingertip position and distal-phalanx orientation are constrained using the spline endpoint and its distal direction.The distal direction is computed from the endpoint and a point five samples earlier on the spline.
- Geometric constraints: Intermediate and proximal digit landmarks are constrained to follow the reconstructed fiber splines through nearest-point associations.Nearest-neighbor associations are recomputed at each objective evaluation.
- Geometric constraints: Dorsum constraints enforce orientation, position, and a planar relationship using the calibrated routing-hub reference.These constraints anchor the hand model to the measured dorsum geometry.
- Optimization: The objective combines distal-axis, spline, dorsum, temporal, and joint-limit terms, optimized per frame with Gauss–Newton updates.Temporal regularization is applied when a recent pose exists and the timestamp gap is small.
D. Routing Hub Self-Calibration
The routing hub is self-calibrated from a short sequence by first obtaining unconstrained initializations, then solving a sequence objective with dorsum constraints enabled.
- Routing hub self-calibration: Routing-hub position and orientation are calibrated using a short sequence of recorded samples.The calibration uses a Gauss–Newton sequence solver.
- Routing hub self-calibration: Each sample first receives a per-frame IK solve without dorsum constraints to initialize the hand parameters.The resulting parameters are stored in the sequence state before calibration optimization.
- Routing hub self-calibration: The sequence objective then adds dorsum position, orientation, and plane terms alongside the usual IK errors.The resulting problem is solved jointly across the recorded sequence.
III. EXPERIMENTS
The experiments evaluate the glove during varied hand poses and functional object-manipulation tasks across five participants and two sessions per participant. Re-donning variability is included through a full glove doff-and-don between sessions.
- III. EXPERIMENTS: Five participants performed varied hand poses and object interactions in the dexterous-manipulation dataset.Each participant completed two sessions.
- III. EXPERIMENTS: Each participant fully doffed and re-donned the glove between sessions to capture re-donning variability.This design exposes the system to session-to-session changes in glove placement.
- III. EXPERIMENTS: The range-of-motion phase covered pinch, palm-flat, and thumb-to-finger side-rub configurations.These poses were designed to span common finger–finger contact configurations.
- III. EXPERIMENTS: Manipulation tasks included buzz wire, box and blocks, cup stacking, in-hand rotation, the 9-Hole peg test, and object squeezing.Each session included both range-of-motion and manipulation phases.
- III. EXPERIMENTS: Ground truth was collected with 26 mocap markers and 20 calibrated OptiTrack cameras surrounding each participant.Nineteen markers supported hand pose estimation, while seven were placed on the routing hub.
B. Evaluation Methodology
Evaluation separates fiber reconstruction and registration accuracy from whole-hand pose reconstruction, using mocap-based measurements while excluding frames with incomplete finger-marker detection. A rigid-transform correction is then used to address consistent spatial bias in the fiber registration.
- B. Evaluation Methodology: Two errors are reported: fiber-to-marker error and hand-landmark Euclidean error against mocap-based reconstruction.The first evaluates the fiber pipeline, while the second evaluates hand pose reconstruction.
- B. Evaluation Methodology: Frames with any undetected mocap marker on a finger are excluded because marker occlusion can corrupt the reference pose.Glove and mocap frames are aligned using a per-frame rigid transform from routing-hub markers.
- B. Evaluation Methodology: Fiber-to-marker error averages each finger marker’s Euclidean distance to its closest point on the reconstructed 3D fiber.It isolates combined fiber-reconstruction and registration error before model fitting.
- B. Evaluation Methodology: Mean Keypoint Position Error averages Euclidean distances over 21 model landmarks, while Fingertip MKPE uses fingertip landmarks only.Both compare the glove-estimated and mocap-derived Momentum hand models.
- B. Evaluation Methodology: A single rigid transform is estimated during range of motion and applied to routing-hub tunnel templates to correct consistent spatial bias.The corrected templates are used to rerun shape registration.
D. Quantitative Results
The glove achieves low hand-pose and fiber-to-marker errors across manipulation tasks, with one-time tunnel alignment improving reconstruction accuracy and preserving real-time operation. It also maintains visually consistent poses during marker occlusion.
- 3.3–3.4 mm average fiber-to-marker error indicates accurate shape reconstruction and registration across range-of-motion and manipulation phases.Tunnel alignment reduced manipulation error from 3.4 to 2.1 mm.
- 4.5–4.9 mm fingertip MKPE after tunnel alignment improved on 6.5–7.2 mm before alignment across both phases.Alignment decreased error in every subject and session; the thumb remained highest at 7.2 mm after alignment.
- 60 Hz tracking fits within the 16.7 ms frame budget, with 32.5 ms device latency, 5.3 ms upsampled registration, and 0.12 ms inverse kinematics.
- 34% of manipulation-phase mocap samples had at least one marker occluded, while the glove produced visually consistent hand poses during self and object occlusion.The reported examples include grasping a coffee cup handle and other manipulation tasks.
- The glove tracks fingertip landmarks closely in the 9 Hole Peg Test and in-hand object rotation, including during mocap marker occlusion.
- One-time routing tunnel alignment corrects fiber-rotation bias and improves pinky and ring-finger alignment with ground truth during a pinch.
F. Application Demonstration
A 6-DoF tracker on the routing hub places both hands in a shared world frame, combining real-time finger pose estimates for bimanual virtual-reality interaction. The demonstration shows virtual jar opening and closing.
- A 6-DoF routing-hub tracker supplies root transforms that place both hands in a shared world coordinate frame for bimanual virtual-reality interaction.
- The participant opened and closed a virtual jar, demonstrating real-time virtual teleoperation.
IV. DISCUSSION
The glove reconstructs full hand pose from multi-core fiber shapes, achieving sub-5 mm fingertip error after tunnel alignment across dexterous manipulation tasks. Remaining limitations include calibration dependence, modeling assumptions, and imperfect comparability with prior gloves.
- Applications and significance: 34% of motion-capture frames had partial occlusion during object manipulation, underscoring the glove’s value for capturing contact-rich finger adjustments hidden from cameras.The system was also demonstrated in real-time bimanual virtual teleoperation.
- Accuracy and calibration: F.MKPE fell from 6.5–7.2 mm before tunnel alignment to 4.5–4.9 mm afterward across range-of-motion and manipulation phases.Tunnel alignment reduced error in every subject and session.
- Error sources: The dominant residual error source was fiber-to-hub shape registration, while hand-model mismatch, fiber motion, and approximate IK solutions also contribute discrepancies.The model assumes medial fiber routing and does not capture all fabric or sensor shifts relative to anatomy.
- Comparison with prior work: The method is the first reported FBG glove to reconstruct full hand pose with a hand model, whereas prior FBG gloves primarily estimated flexion at limited joints.Single-core sensing cannot resolve bending direction and therefore cannot disentangle several out-of-plane motions.
- Comparison with prior work: Direct comparison with prior gloves is limited because their sensing targets, evaluation protocols, and reported errors differ.For example, a comparable fingertip error was measured using a simpler tracing task rather than dexterous manipulation.
- Accuracy and calibration: A one-time mocap-aided tunnel alignment transferred across users and sessions, differing from matched calibration by only 0.1 mm on average.This suggests factory calibration can reduce the need for motion capture during later captures.