Source-linked AI summary

Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation

Han Xue, Jieji Ren, Wendi Chen, Gu Zhang, Yuan Fang, Guoying Gu, Huazhe Xu, Cewu Lu

arXiv:2503.02881v3cs.ROcs.AIcs.LG

TL;DR

The paper addresses the difficulty of making robots responsive and force-aware in contact-rich tasks, where visual imitation learning often executes action chunks without instant tactile feedback. It introduces TactAR for AR-based tactile teleoperation and RDP, a slow-fast policy combining latent action-chunk planning with fast tactile refinement. Across three tasks, RDP improves performance over visual imitation-learning baselines and works across multiple tactile and force sensors.

  • Problem

    Visual imitation-learning policies model complex behaviors with action chunks but cannot respond instantly to tactile feedback during chunk execution, while teleoperation systems often lack fine-grained tactile or force feedback.

  • Method

    The paper introduces TactAR for real-time AR tactile feedback and RDP, which combines a slow latent diffusion policy with a fast tactile-feedback network for closed-loop control.

  • Results

    RDP improves overall scores by more than 35% over Diffusion Policy baselines across three challenging contact-rich tasks and applies across GelSight Mini, MCTac, and joint torque sensors.

  • Takeaways & Limitations

    The slow-fast hierarchy addresses the trade-off between complex trajectory modeling and rapid tactile reaction in contact-rich manipulation.

  • Takeaways & Limitations

    TactAR is designed for two-finger grippers, while RDP is limited to single-task scenarios and its fast policy does not yet incorporate high-frequency visual inputs.

Abstract

from arXiv · show

Humans can accomplish complex contact-rich tasks using vision and touch, with highly reactive capabilities such as fast response to external changes and adaptive control of contact forces; however, this remains challenging for robots. Existing visual imitation learning (IL) approaches rely on action chunking to model complex behaviors, which lacks the ability to respond instantly to real-time tactile feedback during the chunk execution. Furthermore, most teleoperation systems struggle to provide fine-grained tactile / force feedback, which limits the range of tasks that can be performed. To address these challenges, we introduce TactAR, a low-cost teleoperation system that provides real-time tactile feedback through Augmented Reality (AR), along with Reactive Diffusion Policy (RDP), a novel slow-fast visual-tactile imitation learning algorithm for learning contact-rich manipulation skills. RDP employs a two-level hierarchy: (1) a slow latent diffusion policy for predicting high-level action chunks in latent space at low frequency, (2) a fast asymmetric tokenizer for closed-loop tactile feedback control at high frequency. This design enables both complex trajectory modeling and quick reactive behavior within a unified framework. Through extensive evaluation across three challenging contact-rich tasks, RDP significantly improves performance compared to state-of-the-art visual IL baselines. Furthermore, experiments show that RDP is applicable across different tactile / force sensors. Code and videos are available on https://reactive-diffusion-policy.github.io.

I. INTRODUCTION

The paper targets contact-rich manipulation by combining real-time AR tactile feedback with a slow-fast visual-tactile imitation-learning policy. This design preserves action-chunk modeling while adding high-frequency tactile responsiveness.

  • Motivation: Visual imitation-learning methods use action chunks to model temporally consistent, non-Markovian behavior but execute chunks open loop.This prevents instant responses to environmental changes during chunk execution.
  • Motivation: Existing tactile imitation-learning methods commonly add tactile observations while retaining conventional action-chunk prediction.Their action-level control therefore remains limited in response speed during chunk execution.
  • Contributions: TactAR provides fine-grained tactile and force feedback in real time through AR using a unified 3D deformation-field representation.The representation is applicable to multiple tactile and force sensors and is rendered attached to the robot end-effector.
  • Contributions: RDP combines a slow latent diffusion policy for low-frequency high-level action chunks with a fast network that refines them using high-frequency tactile feedback.The slow stage operates at 1–2 Hz and the fast stage at 20–30 Hz.
  • Results: Experiments across three contact-rich tasks show more than 35% improvement over state-of-the-art imitation-learning baselines, with applicability across multiple sensors.The evaluation covers precision, adaptive force control, disturbance response, and bimanual coordination.

B. Robot Data Collection System

TactAR addresses limitations of visually driven teleoperation by combining low-cost AR hardware with real-time tactile sensing. Its design supports precise contact-rich data collection across sensors and robot embodiments.

  • Background: Most existing teleoperation systems rely primarily on visual feedback, making precise contact-rich tasks difficult.Alternative haptic systems may require isomorphic hardware, costly force-feedback devices, or provide only coarse vibration feedback.
  • System Overview: TactAR combines a consumer VR controller with tactile sensing to provide AR-based feedback while preserving accuracy for precise contact-rich tasks.The system is described as flexible and easy to deploy across tactile sensors and robot platforms.
  • System Overview: The system renders 3D deformation or force fields in AR and attaches them to the robot end-effector after calibration.This provides users with tactile, force, and torque feedback during teleoperation.
  • Versatility and Accessibility: TactAR supports deployment across different tactile or force sensors and robot arms, including single-arm and bimanual configurations.Its feedback representation requires the robot TCP pose and 3D deformation field rather than hardware-specific configuration parameters.
  • Versatility and Accessibility: The teleoperation setup uses a consumer-level Meta Quest3 headset and does not require specialized or isomorphic haptic hardware.The paper reports a headset cost of $500 for the system description.

A. 3D Deformation Field Extraction

TactAR extracts or directly represents tactile and force information as a 3D field, aligns it with the robot through AR calibration, and supports diverse sensing and control pipelines.

  • 3D Deformation Field Extraction: Tactile images are converted into normalized marker positions, then optical flow is computed between an initial and current frame.The marker deformation provides contact information without directly deriving force values through complex calibration.
  • 3D Deformation Field Extraction: The resulting 3D deformation field uses displacement components and a z-offset, while force sensors directly provide 3D force components.These fields are rendered in AR for feedback.
  • 3D Deformation Field Extraction: Marker deformation fields in GelSight Mini vary across contact modes and encode contact-related motion patterns.The figure presents examples of these fields during different contact modes.
  • Real-time Tactile / Force Feedback Rendering in AR: AR calibration aligns the virtual coordinate system with the predefined robot TCP position and the world-coordinate origin.Quest3 pose tracking and synchronized sensor, robot, and camera streams support this alignment and rendering pipeline.
  • Versatility and Accessibility: TactAR is designed for cross-sensor and cross-embodiment deployment, supporting optical tactile sensors, joint torque sensors, different robot arms, and bimanual control.The paper reports experiments using GelSight Mini, MCTac, and built-in joint torque sensors.
  • Versatility and Accessibility: The pipeline is positioned as accessible through low-cost hardware, including a Meta Quest3 headset and lower-cost optical tactile sensors.The paper reports approximately $50 for customized MCTac fabrication and about $3,000 for an ATI mini45 force/torque sensor.
  • Versatility and Accessibility: RDP contrasts open-loop action chunking and temporal ensembling with a slow-fast pipeline that supports high-frequency closed-loop adjustments.Its fast tactile-feedback component is intended to preserve reactive behavior during chunk execution.

A. Tactile / Force Representation

The paper represents tactile and force feedback for a slow-fast policy using sensor-specific inputs, while an asymmetric tokenizer compresses action chunks and reconstructs actions conditioned on high-frequency tactile signals.

  • Tactile / Force Representation: Optical tactile sensors use PCA features from marker deformation fields, improving robustness to tracking noise, texture changes, lighting changes, and gel replacements.The representation is generated from the marker deformation field F.
  • Tactile / Force Representation: Force feedback is represented by concatenating a 6-D wrench containing force and torque into the observation vector.
  • Slow-Fast Policy: RDP samples latent action chunks at low frequency, then uses high-frequency tactile representations to predict each next action autoregressively within the chunk.This separates complex sequence modeling from reactive closed-loop control.
  • Fast Policy: The asymmetric tokenizer encodes an action chunk into a downsampled latent sequence and decodes actions using the latent chunk plus tactile representations.Its 1D-CNN encoder preserves sequence structure, while the GRU decoder predicts precise actions from tactile input.
  • Fast Policy: The fast policy takes less than 1 ms for inference and theoretically supports inputs above 300 Hz.The small KL coefficient smooths the tokenizer latent space rather than making it a generative model.
  • Slow Policy: The slow Latent Diffusion Policy operates on latent action chunks, reducing computational cost while retaining temporally consistent complex or non-Markovian behavior.The policy is trained by denoising latent action chunks conditioned on observations.

3) Implementing Suggestions for Slow-Fast Policy:

The section identifies design choices and evaluation questions for slow-fast policy learning, emphasizing trajectory representation, latency alignment, tactile feedback, sensor compatibility, perturbation response, and teleoperation data quality.

  • Relative trajectory: Relative end-effector trajectories use transformations from the last observation frame instead of consecutive-frame deltas to reduce compounding errors.
  • Latency matching: Latency matching discards initial predicted action steps so policy inference and execution align, producing smoother transitions between action chunks.The procedure also prevents out-of-distribution tactile signals from inducing abnormal fast-policy actions.
  • Research questions: The experiments ask whether tactile signals improve contact-rich policy performance and whether RDP enables complex behavior, immediate perturbation response, and compatibility with different sensors.
  • Research questions: The study evaluates why small chunk sizes or temporal ensembling may not substitute for the slow-fast design's closed-loop control frequency.
  • Research questions: The study also examines how TactAR's tactile / force feedback affects teleoperation data quality and how data quality influences policy performance.

A. Setup

The experiments use dual robotic arms, multiple cameras, and three tactile or force sensing modalities to evaluate RDP on peeling, wiping, and bimanual lifting under varied conditions.

  • Setup: The platform uses two Flexiv Rizon 4 arms, Flexiv Grav grippers, and wrist cameras, with an additional fixed external camera for bimanual trials.
  • Sensors: Experiments use GelSight Mini, improved MCTac, and built-in joint torque sensors, covering optical tactile and estimated force / torque feedback.The joint-torque-derived signals are streamed at 120 Hz and downsampled to 24 FPS, with comparatively larger noise.
  • Baselines: Baselines include visual-only Diffusion Policy, tactile-image and tactile-embedding variants, plus RDP using tactile embeddings or force feedback.
  • Tasks: The three evaluation tasks are Peeling, Wiping, and Bimanual Lifting, each requiring contact-rich capabilities such as precision, adaptive or precise force control, rapid response, or coordination.
  • Evaluation protocol: Peeling and Wiping test no perturbation, pre-contact perturbation, and post-contact perturbation, while Bimanual Lifting tests soft and hard paper cups.Each test-time variation uses 10 trials.

4) Evaluation Protocols:

Across three contact-rich tasks, the evaluation examines precision, force control, disturbance response, and bimanual coordination. RDP combines low-frequency trajectory modeling with high-frequency tactile feedback and outperforms visual imitation baselines while supporting varied sensors.

  • RDP uses a slow latent diffusion policy for low-frequency action chunks and a fast policy for high-frequency tactile or force updates.The slow policy predicts at 1–2 Hz, while the fast policy updates at 24 FPS because of GelSight’s frame-rate limitation.
  • Adding tactile observations alone does not improve complex-task performance, although tactile embeddings are more robust than raw tactile images to gel texture changes.
  • The evaluation covers Peeling, Wiping, and Bimanual Lifting, testing precision, adaptive force control, disturbance response, and bimanual coordination.
  • RDP improves overall scores by more than 35% over Diffusion Policy baselines across all three tasks.The tasks require precision, adaptive force control, or bimanual coordination, capabilities associated in the paper with high-frequency closed-loop feedback.
  • RDP transfers across tactile and force sensors, achieving 0.9 with GelSight Mini and 0.88 with MCTac on Peeling and performing well with mixed sensors in Bimanual Lifting.The paper also reports that simple force input achieves the best results across all three tasks despite noise during rapid movements.
  • On Peeling under contact perturbation, RDP with GelSight scores 0.8 versus 0.15 for Diffusion Policy with tactile embedding.RDP changes the predicted trajectory immediately through closed-loop fast-policy control, whereas baselines continue executing open-loop chunks.
  • Reducing chunk size or using temporal ensembling introduces trade-offs: chunk reduction lowers grasp success from 100% to 20%, while ensemble factors τ = 0.2 and τ = 0.8 produce 30% grasp rate and over-smoothing, respectively.
  • Relative trajectory prediction performs better than absolute action prediction, while TactAR data produces more stable contact forces and higher-quality data improves policy performance.Using traditional VR teleoperation data causes the Peeling score to decrease by more than 30%.

VI. LIMITATIONS AND FUTURE WORKS

The paper identifies limitations in teleoperation efficiency, hardware scope, input frequency, and task scope. Future work targets lower latency, dexterous hands, high-frequency vision, and integration with vision-language-action models.

  • TactAR feedback is less intuitive or efficient than direct human-hand operation, and the system is designed for two-finger grippers.Future work could reduce sensor and system latency and extend the system to dexterous hands with tactile sensors.
  • RDP’s fast policy responds to high-frequency tactile or force signals but cannot yet process high-frequency image inputs swiftly.The authors propose adding high-frequency visual inputs to broaden task applicability.
  • RDP is currently limited to single-task scenarios.Future work could integrate it with a vision-language-action model using an asymmetric tokenizer for reactive tactile control.

VII. CONCLUSION

The paper presents TactAR for real-time AR-based tactile and force feedback and RDP for slow-fast visual-tactile imitation learning. Across three contact-rich tasks, RDP outperforms visual imitation baselines and generalizes across tactile and force sensors.

  • TactAR is a low-cost AR teleoperation system, while RDP is a slow-fast imitation learning algorithm for contact-rich manipulation.
  • RDP separates complex trajectory planning from reactive tactile feedback using slow and fast hierarchical networks.
  • Across three challenging tasks, RDP outperforms state-of-the-art visual imitation learning baselines in task completion and tactile reactivity.
  • Cross-sensor experiments indicate that the approach generalizes across different tactile and force sensors.

APPENDIX

TactAR integrates an improved camera-based tactile sensor with AR visualization of 3D deformation fields, while supporting low-latency feedback for tactile and force sensors.

  • Tactile Sensor: The improved MCTac sensor redesigns the gripper interface and camera mount, while using white-side illumination to emphasize shear force.It is based on the open-source MCTac design and is adapted for the Flexiv Grav Gripper.
  • AR Visualization: The AR system receives robot TCP poses, 3D deformation fields, and undeformed marker locations to construct deformation arrows.Arrow start and end points are first computed in the robot TCP coordinate system.
  • AR Visualization: The deformation arrows are transformed into AR world coordinates using the robot TCP pose and a visibility scale factor.The transformed arrows are rendered in Quest3, with adaptive coloring based on arrow length.
  • AR Visualization: Tangential force and torsional torque are readily recognized from the flow field, whereas normal force can be difficult to observe intuitively.Large normal forces may produce an outward diffusion pattern that dominates the marker pattern.
  • Latency: Marker tracking and Quest3 rendering each take about 10ms, while network latency is about 1ms-6ms.Force sensors can provide more intuitive AR feedback because their internal latency is below 1ms, compared with 10ms-60ms for optical tactile sensors.

C. Details of the Tactile Representation

The tactile representation converts optical sensor marker deformation fields into compact PCA features that preserve interpretable force and torque information while improving robustness to noise and sensor-surface variation.

  • Deformation Field: Each optical tactile frame produces a deformation field matrix F with shape (n, 2) from n marker dots.The deformation field is less affected by lighting, texture, and moderate gel damage or replacement.
  • PCA Embedding: PCA reduces concatenated deformation fields from m frames and 2n features into a matrix with d principal components.The reduced representation is computed as F_reduced = T_proj F_concat, where T_proj is the PCA projection matrix.
  • PCA Embedding: For a new frame, the tactile PCA feature is obtained by projecting its deformation field with T_proj.The resulting feature is f_reduced_t′ = T_proj F_t′.
  • PCA Embedding: Setting d = 15 provides reconstruction without losing much detail, while leading components correspond to interpretable force and torque directions.The first components represent tangential force, torsional torque, and normal force.
  • Datasets: The PCA tactile embedding uses 30 minutes of Gelsight Mini data and 40 minutes of MCTac data collected through random interactions with 20 objects.The three evaluation tasks use 60, 80, and 50 demonstrations for Peeling, Wiping, and Bimanual Lifting, respectively.

E. Details of the Evaluation Protocols

Evaluation uses blinded human-involved testing, matched action-observation settings for baselines and RDP, and ablations that remove shear information or the asymmetric tokenizer.

  • Testing Protocol: Testing uses a single-blind protocol in which evaluators do not know which randomly selected model is being assessed.This protocol is used because human interaction is involved during testing.
  • Inference Protocol: Diffusion Policy baselines and LDP use two observations, with baselines predicting open-loop 12FPS action sequences and approximately 0.67-second chunks.The slow policy predicts at 1-2Hz, while the fast policy receives tactile or force observations and outputs actions at 24FPS.
  • Ablations: Removing shear force reduces the Peeling-task policy performance from 0.95 to 0.48.The ablation retains only the normal-force dimension, matching the 3D-ViTac force representation.
  • Ablations: Replacing the asymmetric tokenizer with concatenated action and tactile or force inputs reduces average policy performance from 0.95 to 0.58.This ablation tests the contribution of the asymmetric tokenizer.
  • Teleoperation Study: The Peeling teleoperation study involves 10 users, each completing 10 trials under traditional VR teleoperation and TactAR.Users have different levels of VR teleoperation or imitation-learning experience, and cucumber pose varies across trials.
Loading 2503.02881v3…