Source-linked AI summary
AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception
Ruoxuan Feng, Yuxuan Zhou, Siyu Mei, Dongzhan Zhou, Pengwei Wang, Shaowei Cui, Bin Fang, Guocai Yao, Di Hu
TL;DR
Existing tactile resources largely emphasize static object properties and lack systematic coverage of temporal dynamics and force-aware interactions. The paper introduces ToucHD and AnyTouch 2 to organize and learn hierarchical dynamic tactile perception, achieving consistently strong performance across static, dynamic, and manipulation tasks. A limitation is that paired touch–force collection remains restricted to simplified indenter-based interactions.
Problem
Existing tactile datasets and models provide limited coverage of fine-grained temporal dynamics and force-related physical principles during interactions.
Method
The paper introduces a five-tier tactile dynamic pyramid, the ToucHD dataset, and AnyTouch 2 with objectives for deformation, action, and force-aware perception.
Results
AnyTouch 2 delivers consistently strong performance across static and dynamic tactile perception benchmarks and real-world manipulation tasks spanning all pyramid tiers.
Takeaways & Limitations
The combined data and representation framework supports general tactile perception across object-level understanding, dynamic interaction modeling, and physical reasoning.
Takeaways & Limitations
Paired tactile–force data collection is restricted to simplified interactions with specially designed indenters, excluding many everyday objects.
Abstract
from arXiv · showhide
Real-world contact-rich manipulation demands robots to perceive temporal tactile feedback, capture subtle surface deformations, and reason about object properties as well as force dynamics. Although optical tactile sensors are uniquely capable of providing such rich information, existing tactile datasets and models remain limited. These resources primarily focus on object-level attributes (e.g., material) while largely overlooking fine-grained tactile temporal dynamics during physical interactions. We consider that advancing dynamic tactile perception requires a systematic hierarchy of dynamic perception capabilities to guide both data collection and model design. To address the lack of tactile data with rich dynamic information, we present ToucHD, a large-scale hierarchical tactile dataset spanning tactile atomic actions, real-world manipulations, and touch-force paired data. Beyond scale, ToucHD establishes a comprehensive tactile dynamic data ecosystem that explicitly supports hierarchical perception capabilities from the data perspective. Building on it, we propose AnyTouch 2, a general tactile representation learning framework for diverse optical tactile sensors that unifies object-level understanding with fine-grained, force-aware dynamic perception. The framework captures both pixel-level and action-specific deformations across frames, while explicitly modeling physical force dynamics, thereby learning multi-level dynamic perception capabilities from the model perspective. We evaluate our model on benchmarks that covers static object properties and dynamic physical attributes, as well as real-world manipulation tasks spanning multiple tiers of dynamic perception capabilities-from basic object-level understanding to force-aware dexterous manipulation. Experimental results demonstrate consistent and strong performance across sensors and tasks.
1 Introduction
Dynamic tactile perception remains underdeveloped because existing datasets and models emphasize static object properties rather than temporal interactions and force dynamics. The paper addresses this gap with ToucHD and AnyTouch 2, evaluated across static, dynamic, and manipulation tasks.
- Motivation: Existing tactile datasets and models largely focus on static object properties, overlooking temporal touch dynamics and force-related principles.Many datasets rely on press-only actions, with limited random sliding or rotation and only preliminary touch–force grounding.
- Contributions: The tactile dynamic pyramid organizes data into five tiers according to data rarity and the complexity of supported perception capabilities.Higher tiers cover specific actions, manipulation, and force-grounded perception, but remain substantially scarcer than lower-tier data.
- Contributions: ToucHD provides 2,426,174 contact samples spanning atomic actions, real-world manipulation, and touch–force paired data.It combines simulated data, manipulation recordings, and force-paired interactions to enrich higher-tier tactile data.
- Contributions: AnyTouch 2 unifies sensor-invariant object understanding with fine-grained deformation, action-specific dynamics, and force-related physical perception.Its components include frame-difference reconstruction, action matching, and temporal force prediction from touch–force pairs.
- Results: AnyTouch 2 delivers consistently strong performance across static and dynamic tactile benchmarks and real-world manipulation tasks spanning the pyramid.The evaluation covers static object properties, dynamic physical prediction, and manipulation capabilities across multiple tiers.
2 Related Works
Prior tactile datasets and models primarily use pressing or simple random actions and adapt visual representation learning, providing limited coverage of dynamic interaction and physical reasoning.
- Datasets: Early tactile datasets collected through pressing mainly support static semantic properties such as material and hardness.Their limited dynamic variation constrains the tactile features learned from these datasets.
- Datasets: Some datasets add random actions on object surfaces, but these interactions remain basic.These extensions provide an initial understanding of tactile dynamics without establishing broad dynamic coverage.
- Models: Recent tactile models leverage visual self-supervised learning and vision-language alignment for feature learning and semantic understanding.Optical tactile sensors motivate these approaches because their data are image-based.
- Models: Dynamic tactile models process continuous inputs to model temporal variation, but often lack designs tailored to tactile-specific characteristics.This leaves a gap between generic visual temporal modeling and the requirements of contact-rich manipulation.
3 Tactile Hierarchical Dynamic Dataset
ToucHD is designed as a hierarchical dynamic tactile dataset that fills higher-tier data gaps with action-specific, manipulation, and force-paired interactions.
- Dataset Overview: ToucHD contains 2,426,174 contact samples and targets the highest three tiers of the tactile dynamic pyramid.The dataset is explicitly designed to enrich higher-tier dynamic tactile data.
- Dataset Subsets: Simulated Atomic Action Data collects multi-sensor frames from four atomic actions performed on 1,043 objects.The actions are left/right sliding and clockwise/counterclockwise rotation, with additional rotated sliding samples.
- Dataset Subsets: Real-World Manipulation Data contains 584,842 contact frames from 46 manipulation tasks collected with two sets of tactile sensors.The modified FastUMI setup simultaneously records interaction videos and tactile data.
- Dataset Subsets: Touch-Force Paired Data provides 722,436 touch–force pairs collected from five sensors and 71 robotic-arm-mounted indenters.Indenters slide in four directions while a wrist-mounted force sensor records 3D contact-force sequences.
- Dataset Coverage: Together, the three subsets cover action-specific, real-world manipulation, and force-paired interactions across objects, sensors, and dynamics.Combined with existing lower-tier datasets, ToucHD forms a complete ecosystem for hierarchical dynamic tactile perception.
4 Method
AnyTouch 2 builds a general tactile representation by combining sensor-invariant object understanding with pixel-, action-, and force-level dynamic perception. Its multi-level objectives and curriculum align learning with the tactile dynamic pyramid, supporting both subtle deformation modeling and physical reasoning.
- Framework overview: AnyTouch 2 unifies sensor-invariant object-level understanding with pixel-, semantic-, action-, and force-level dynamic perception.The framework is organized around the hierarchical tiers of the tactile dynamic pyramid.
- Pixel-Level Dynamic Details: Masked video reconstruction and frame-difference reconstruction capture global deformation patterns and subtle temporal variations across optical tactile sensors.The model normalizes tactile frames by subtracting the background, masks spatio-temporal tokens, and reconstructs both frames and frame differences.
- Semantic-Level Tactile Features: Multi-modal alignment, cross-sensor matching, and action matching connect tactile signals with object, material, interaction, and atomic-action semantics.Action matching groups pressing, leaving, sliding, and rotating sequences while cross-sensor matching aligns contacts with the same object across sensors.
- Dynamic Physical Properties: Force prediction uses touch–force pairs to estimate 3D contact forces and temporal force variations from tactile videos.Joint force and delta-force decoders are trained with an L1 loss, linking dynamic deformations to physical magnitudes.
- Training Recipe: Curriculum task scheduling introduces higher-level semantic and physical objectives after low-level patterns, with task-specific weights increasing over training iterations.The strategy is intended to reduce task interference while first establishing robust low-level tactile representations.
5 Experiments
The experiments evaluate tactile perception across static properties, dynamic physical attributes, and real-world manipulation tasks spanning all tiers of the tactile dynamic pyramid. Results show that dynamic, force-aware modeling and hierarchical training data improve performance across sensors and task levels.
- Evaluation scope: The evaluation covers object properties, force prediction, pose estimation, slip detection, and four real-world manipulation tasks spanning the tactile dynamic pyramid.The offline benchmarks use GelSight, DIGIT, and GelSight Mini, while the manipulation evaluation spans DIGIT and GelSight Mini.
- Offline benchmarks: AnyTouch 2 matches AnyTouch 1 on static Object Bench while consistently outperforming prior approaches on fine-grained dynamics and force-sensitive reasoning.The comparison supports a unified evaluation of object-level understanding, action-aware dynamics, and force-grounded perception.
- Offline benchmarks: Consecutive-frame models outperform single-frame baselines on dynamic benchmarks, whose temporal ordering cannot be captured without dynamic inputs.Single-frame baselines can perform worse than CLIP on Force Prediction and Slip Detection because they lack temporal position embeddings.
- Online real-world manipulation: Static single-frame models perform significantly worse than dynamic models in real-world manipulation, especially on higher-tier tasks.Models trained only on lower-tier dynamic data perform poorly on Tier 1 and Tier 2 tasks outside their training coverage.
- Online real-world manipulation: Adding ToucHD to MAE (S) improves every evaluated manipulation task over the original model, except accurate Tier 1 force perception.The result indicates that hierarchical dynamic data supports capabilities across Tier 2 and lower-tier tasks, while Tier 1 force perception remains difficult.
- Sensor comparison: Sensor strengths differ by task: GelSight Mini excels on Tier 5 detail capture, whereas DIGIT’s 30 Hz acquisition supports higher-tier manipulation.GelSight Mini operates at 18 Hz and provides sharper deformation imaging; DIGIT supplies denser dynamic information through its higher acquisition frequency.
- Ablation study: Removing action matching, force prediction, or frame-difference reconstruction reduces performance on their associated dynamic tasks.Frame-difference reconstruction affects all dynamic perception tasks, while removing multimodal alignment harms static Object Bench performance.
6 Conclusion
The paper presents a hierarchical framework for dynamic tactile perception, combining the ToucHD data ecosystem with AnyTouch 2 and simulated multi-sensor contact data. Its tiered design links increasingly challenging data collection to richer annotations and dynamic capabilities.
- Conclusion: The tactile dynamic pyramid organizes data collection and model design around five tiers of hierarchical tactile perception capabilities.The tiers range from press-only data to force-labeled interactions, with increasing collection difficulty and data rarity.
- Tactile dynamic pyramid: Tier 1 is the only tier containing paired force labels, while lower tiers provide progressively different action structures and annotations.Tier 3 includes detailed action labels without paired force data, and Tier 2 covers real object manipulation without paired force data.
- Tactile dynamic pyramid: Higher-tier data is harder and rarer to collect but provides richer annotations or more realistic manipulation scenarios for stronger dynamic perception.The hierarchy is defined by collection effort, action type, and label difficulty.
- Simulated data: The IMPM simulation platform combines elastomer–object contact simulation with rendering to generate dynamic tactile data across sensors.It processes more than 1,000 objects spanning over 10 material types and five environments, then renders simulated tactile images from deformed elastomer meshes.
- Related data generation: Implicit neural tactile fields can generate many static frames but cannot directly render tactile images during dynamic contact, placing them at Tier 5.Their main limitation is the absence of dynamic-contact rendering rather than a lack of material diversity.
B.2 Manipulation Data
The manipulation data collection adapts UMI for multiple optical tactile sensors and emphasizes diverse tasks that elicit fine-grained dynamic tactile variation. It also includes synchronized multimodal streams, touch–force pairs, and atomic action labels, while trading dual-UMI collection for UMI–hand collaboration.
- Sensor and platform design: Three commercial optical tactile sensors—GelSight Mini, DIGIT, and DM-Tac W—are integrated into the adapted UMI setup.The Mini is used with and without markers.
- Multimodal data: Synchronized external-camera and tactile streams illustrate dynamic interactions across tasks including capping a pen, placing a test tube, detaching Velcro, and closing a box lid.The examples combine GelSight Mini with DIGIT or DM-Tac W.
- Task design: The dataset covers 46 manipulation tasks involving pushing, pulling, squeezing, rotating, sliding, and aligning, with both sensor groups completing identical tasks.This design enables direct comparison across sensors under matched task conditions.
- Collection trade-off: The collection uses UMI plus left-hand collaboration instead of two UMIs because the modified dual-sensor setup became too bulky for many tasks.The authors identify this as a trade-off that may bias the visual modality.
- Touch–force data: Touch–force data are collected from five optical tactile sensors: GelSight Mini, DIGIT, DuraGel, DM-Tac W, and GelStereo BioTip.The collection targets force as a physical property relevant to dexterous manipulation.
- Additional annotations and sensors: The data include atomic action labels for some samples, and the sensors span planar optical designs as well as GelStereo BioTip’s 3D-marker representation.These additions support action matching and potential future integration of non-planar tactile sensors.
C Training Dataset Statistics
Training uses nine tactile datasets spanning different dynamic-pyramid tiers, modalities, sensors, and scales. To reduce redundant static frames and improve efficiency, the pipeline selects informative changes and samples intervals.
- Dataset composition: Nine tactile datasets are used for training, differing in dynamic-pyramid tier, paired modalities, sensors, and data scale.Most datasets occupy lower dynamic tiers and contain many contact-static frames.
- Redundancy reduction: Frame selection uses the variance of the Laplacian relative to the preceding frame to retain more informative dynamic contact events.A threshold is applied to identify frames with greater changes.
- Redundancy reduction: Interval sampling is additionally applied to YCB-Slide, Touch-Slide, and ToucHD to further reduce data redundancy and improve training efficiency.
D Benchmark and Baseline Details
Evaluation covers static object properties and dynamic physical understanding across several optical tactile sensors, while baselines range from single-frame to multi-frame representation learners. Training uses OpenCLIP-based encoders with a ViT tactile decoder and specified optimization settings.
- Benchmarks: Touch and Go and Cloth evaluate material recognition and clothing texture classification, while Sparsh and ToucHD Bench evaluate dynamic physical understanding.Touch and Go and Cloth each contain 20 categories; the benchmarks cover GelSight, DIGIT, and GelSight Mini.
- Dynamic force benchmark: ToucHD Bench selects 10 probes from 71 touch–force paired probes, using seven indenters for training and three unseen indenters for testing.The split targets force-related physical-property perception.
- Representation-learning baselines: Baselines include UniTouch and T3 with single-frame inputs, plus MAE, VJEPA, and AnyTouch 1 with consecutive-frame inputs.Single-frame methods receive consecutive frames unfolded along the batch dimension for comparison.
- Implementation: The encoders build on OpenCLIP-Base, while the tactile decoder is a six-layer ViT with eight attention heads and hidden dimension 512.Optimization uses AdamW with learning rate 3 × 10^-4 and batch size 64.
- Training procedure: Samples lacking labels for a training objective are excluded from that objective’s loss computation.This handles incomplete supervision across the multiple training objectives.
F Real-world task setup
Four real-world manipulation tasks span the tactile dynamic pyramid, from object-property recognition to force-sensitive dexterous control. The experiments use different sensor–robot embodiments and action-chunked control for real-time execution.
- Tactile Grasping: Tactile Grasping requires distinguishing balls by material and texture while monitoring deformation feedback during placement.The task includes balls with different tactile properties, including one smooth surface.
- Whiteboard Wiping: Whiteboard Wiping requires structured directional contact and sufficient force because the robot has only one opportunity to erase the letters completely.Insufficient force leaves the letters uncleared with no correction opportunity.
- USB Insertion: USB Insertion tests multidirectional deformation perception under tight socket tolerances, with collisions potentially changing the connector pose.The task includes both extraction and reinsertion.
- Chip Moving: Chip Moving demands sensitivity to minute force variations while transferring a fragile chip under partial visual occlusion.The robot must rely primarily on tactile feedback during manipulation.
- Embodiments and control: Experiments use Piper with GelSight Mini and DIGIT for three tasks, and xArm 6 with GelSight Mini for Chip Moving; actions are chunked with horizon 8 at 3 Hz.Only the first two predicted actions are executed at each inference step.
H Force prediction Evaluation
The force-probe comparisons show that multi-frame and force-aware training improve tactile force prediction, with AnyTouch 2 achieving precise predictions across all three directions.
- Force-probe comparison: MAE(Sparsh) and VJEPA(Sparsh) outperform single-frame baselines but retain substantial bias in tangential X- and Y-direction forces.Both use multi-frame inputs, yet their errors indicate insufficient perception of sliding dynamics.
- Force-probe comparison: AnyTouch 1 predicts the Z-axis normal force relatively accurately but performs poorly on tangential force prediction.This contrasts its static-feature emphasis with AnyTouch 2’s dynamic force modeling.
- Evaluation setup: AnyTouch 2 is evaluated alongside baselines using 3D force probes on DIGIT and GelSight Mini subsets of ToucHD Bench.The visual comparison covers CLIP, T3, MAE, VJEPA, AnyTouch 1, and AnyTouch 2.
I Ablation study
The ablations show that pyramid-guided data, force labels, temporal context, and task-specific modules contribute to dynamic tactile perception, while some design choices trade off static and dynamic performance.
- Module ablations: Removing multi-modal alignment lowers material-understanding performance but improves most dynamic physical perception datasets.Coarse text labels can pull samples with different contact forces closer together, which is undesirable for force-sensitive tasks.
- Dataset ablations: ToucHD removal causes consistent performance drops across real-world manipulation tasks, especially higher-tier tasks corresponding to ToucHD data.The reported impact is strongest for Tier 1, 2, and 3 tasks.
- Dataset ablations: Tier 1 data and shear-force labels clearly contribute to force prediction and real-world manipulation tasks.ToucHD provides both normal and shear force labels.
- Input-frame study: Four-frame inputs outperform two-frame inputs across all benchmarks, while nearly doubling token sequence length.Denser dynamic tactile information improves perception at increased computational cost.
- Hyper-parameter study: Within each feasible hyper-parameter range, AnyTouch 2 consistently outperforms baselines despite minor fluctuations and peaks.The study varies one hyper-parameter at a time from an anchor configuration.
- Training behavior: Training losses decrease smoothly, and rare zero Force Loss or Matching Loss iterations do not affect training.The zero-loss batches arise from sampler constraints and are extremely rare, especially across multiple GPUs.
- Sensor robustness: Replacing the gel pad causes only a minor performance drop in tested GelSight Mini manipulation tasks.This supports robustness to gel-pad changes at test time.
L limitations
The authors identify four limitations involving unused sensor data, constrained tactile–force collection, underused multi-sensor data, and limited modality coverage.
- Unexplored sensors: Two marker-based optical sensors in ToucHD, including spherical GelStereo BioTip, remain unexplored and unused in this study.Using these data could further increase sensor diversity.
- Force data collection: Paired tactile–force data require a specially designed indenter moving across the sensor, restricting interactions and excluding many everyday objects.The authors propose capturing tactile–force data during natural object manipulations.
- Multi-sensor data: Multi-sensor paired manipulation data are aligned with visual inputs without specialized architectures exploiting cross-sensor synergies.The visual modality could also predict future tactile signals, an underexplored direction.
- Modality coverage: The framework supports general optical tactile sensors but does not yet integrate common array-based tactile modalities.Future work would need to handle heterogeneous tactile data formats.