Source-linked AI summary

SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation

Bowen Jing, Mingxin Wang, Ruiyang Hao, Chenchen Ge, Hanwen Shen, Junjie He, Yang Cui, Yiming Hou, Weitao Zhou, Jiawei Wang, Minglei Li, Dandan Zhang, Ding Zhao, Houde Liu, Xiaofan Li, Si Liu, Ping Luo, Haibao Yu

arXiv:2608.18701v1cs.RO

TL;DR

Most manipulation benchmarks measure task completion without recording the physical consequences of contact. SoftVTBench pairs visuo-tactile demonstrations with deformation-aware evaluation, finding higher OOD task success in all six comparisons and higher OOD DSR in five.

  • Problem

    Most benchmarks measure task completion but lack policy-visible contact sensing paired with independent records of physical interaction consequences.

  • Method

    SoftVTBench provides 4,000 demonstrations across 40 tasks and over 50 assets, pairing synchronized visuo-tactile observations with evaluator-only FEM states and a deformation-aware success metric.

  • Results

    Higher OOD TSR occurred in all six comparisons and higher OOD DSR in five of six, despite mixed in-distribution effects.

  • Takeaways & Limitations

    Task success alone can conceal excessive deformation, and tactile observations do not guarantee effective multimodal fusion.

  • Takeaways & Limitations

    The benchmark evaluates simulated visuo-tactile learning and does not validate transfer of its tactile signals to physical sensors.

Abstract

from arXiv · show

Physical interaction quality is central to deformable-object manipulation, yet most benchmarks evaluate task success alone. A policy may complete the task while allowing slip or causing excessive compression. A primary bottleneck is the absence of visuo-tactile datasets that pair policy-visible contact observations with independent physical ground truth over complete tasks. We introduce SoftVTBench, a visuo-tactile dataset for physical-interaction-aware deformable-object manipulation. It contains 4,000 expert demonstrations and more than 50 assets, including volumetric deformable objects and visually matched rigid twins. At 20 Hz, each episode synchronizes multi-view RGB, dual-finger tactile RGB and marker motion, proprioception, language, and binary and continuous gripper actions, alongside evaluator-only finite-element (FEM) states. Building upon this dataset, we establish a closed-loop benchmark that uses fixed object-specific calibration to define the Deformation-aware Success Rate (DSR), which counts a rollout as successful only when it completes the task and keeps peak normalized deformation within tolerance. Across Diffusion Policy, $π_{0.5}$, and FastWAM, all 12 in-distribution configurations contain successful rollouts that violate the deformation tolerance, accounting for 0.7--24% of each configuration's successes. Under distribution shift, visuo-tactile variants achieve higher task success in all six policy--suite comparisons and higher DSR in five, whereas their in-distribution benefits are mixed. These results show that making touch available does not by itself ensure effective multimodal fusion. SoftVTBench therefore provides a common visuo-tactile resource for studying not only whether a policy succeeds, but how it physically interacts with deformable objects and when touch improves that interaction.

1 Introduction

SoftVTBench addresses the gap between task completion and physical-interaction quality in deformable-object manipulation by pairing policy-visible visuo-tactile trajectories with evaluator-only physical ground truth. Its closed-loop benchmark uses deformation-aware success to reveal when successful rollouts violate calibrated deformation tolerances and to study when touch improves performance.

  • Motivation: Task success alone cannot distinguish loose-grasp slip from task completion achieved through excessive compression of deformable objects.This limitation motivates evaluating both outcome and physical interaction, especially for deformation-sensitive objects.
  • Dataset: The dataset separates policy-visible observations from evaluator-only FEM, object, and contact states, enabling contact-aware learning and auditable physical-interaction evaluation.The evaluator-only stream prevents tactile observations from serving as their own evaluation target.
  • Dataset: SoftVTBench contains 4,000 expert demonstrations across 40 pick-and-place tasks and more than 50 assets, including volumetric deformable objects and matched rigid twins.Each trajectory synchronizes multi-view RGB, dual-finger tactile observations, proprioception, language, actions, and evaluator-only physical states.
  • Benchmark: The benchmark defines DSR by combining task completion with FEM-grounded deformation compliance under policy-independent, object-specific calibration.Matched rigid twins and in-distribution/out-of-distribution protocols support controlled physical comparisons.
  • Results: Across all 12 in-distribution deformable-object configurations, successful rollouts exceeded the calibrated deformation tolerance, while touch was most consistently associated with stronger results under distribution shift.The findings show that tactile availability does not guarantee effective multimodal fusion, and that task success can conceal excessive deformation.

2 Related Work

Prior benchmarks cover complete-task manipulation, deformable-object simulation, and visuo-tactile policy learning, but they generally do not evaluate gentle handling through deformation-aware closed-loop tasks. SoftVTBench addresses this gap by combining complete manipulation, volumetric soft bodies, policy-visible touch, and evaluator-only deformation scoring.

  • Robotic Manipulation Benchmarks: General-purpose suites standardize complete-task evaluation across multi-task, language-conditioned, and generalization settings, but primarily measure task success.Representative benchmarks include Meta-World, RLBench, LIBERO, CALVIN, RoboCasa, RoboTwin, THE COLOSSEUM, and SIMPLER.
  • Deformable-Object Manipulation: Deformable-object benchmarks model shape change across cloth, rope, garments, and soft bodies, including differentiable simulation and broader manipulation suites.Examples include SoftGym, DEDO, GarmentLab, PlasticineLab, DaXBench, ManiSkill2, and MoDeSuite.
  • Deformable-Object Manipulation: SoftVTBench targets gentle handling: completing a separate task while keeping deformation within a calibrated tolerance, unlike settings where changing shape is the objective.Closest prior studies measure grasp-induced deformation or stress, but their evaluation is limited to isolated grasps rather than complete closed-loop tasks.
  • Visuo-Tactile Sensing and Policy Learning: Tactile sensing exposes contact geometry, shear, slip, and compression that vision alone may not capture, yet prior visuo-tactile work typically evaluates benefits through task success.Relevant simulators and policy-learning systems include Taxim, FOTS, TacEx, TacSL, DiffTactile, ManiFeel, TacO, and VTDexManip.
  • Visuo-Tactile Sensing and Policy Learning: Tabero evaluates interaction quality alongside task success using closed-loop force feedback, but on rigid objects; no compared benchmark jointly provides the four SoftVTBench capabilities.The four capabilities are complete manipulation tasks, volumetric deformable objects, policy-visible touch, and evaluator-only deformation scoring.

3 The SoftVTBench Dataset

SoftVTBench is a 4,000-demonstration visuo-tactile dataset for contact-rich deformable-object manipulation, combining policy-visible observations with evaluator-only physical states. Its matched assets, diagnostic suites, controlled shifts, and per-object deformation calibration support evaluating both task execution and physical interaction.

  • Tasks, Objects, and Scenes: SoftVTBench contains 4,000 expert demonstrations across 40 pick-and-place tasks in four diagnostic suites, together with more than 50 assets.The suites are Object-Soft, Spatial-Soft, Object-Rigid, and Spatial-Rigid, with 10 tasks and 1,000 demonstrations per suite.
  • Dataset Splits: In-distribution evaluation uses 500 held-out initial states per suite, while out-of-distribution evaluation applies nine single-factor conditions to the two deformable suites.The shifted factors are lighting, mass, and Young’s modulus, with one factor changed at a time.
  • Visuo-Tactile Observations and Actions: Each episode synchronizes multi-view RGB, dual-finger tactile RGB and marker motion, proprioception, language, binary and continuous gripper actions, and evaluator-only physical states at 20 Hz.Evaluator-only states include FEM nodal positions, object poses, contact events, and drop events, and are never exposed to policies.
  • Tasks, Objects, and Scenes: The four suites form diagnostic controls over object type and variation type, while Object suites test asset adaptation and Spatial suites test language-guided target grounding.Object suites vary the manipulated asset under fixed layouts; Spatial suites place visually identical instances in one scene while varying the language-selected target.
  • Deformation from Physical State: The benchmark measures normalized peak rigid-motion-removed FEM displacement and uses per-object calibration to define a deformation tolerance independent of evaluated policy behavior.The tolerance is the 90th percentile of peak displacements over stable grasps; calibration precedes policy training and supplies the benchmark criterion.

4 The SoftVTBench Benchmark

SoftVTBench evaluates closed-loop manipulation by requiring both instructed task completion and compliance with calibrated deformation tolerances. It separates policy-visible observations from evaluator-only physical states and reports paired in-distribution and out-of-distribution results.

  • Benchmark Protocol: SoftVTBench executes policies closed loop on fixed evaluation sets with a shared physics and observation interface, defining policy inputs, evaluator-only states, DSR, and paired ID/OOD protocols.All methods are evaluated under the benchmark’s standardized protocol.
  • Benchmark Protocol: Policies receive RGB, proprioception, language when applicable, and dual-finger tactile images with marker fields for visuo-tactile variants.The simulator runs at a 20 Hz control rate and executes actions from the current predicted action chunk.
  • Benchmark Protocol: Evaluator-only FEM positions, object poses, contact events, and drop events remain hidden, requiring policies to infer slip, contact, and compliance from visible inputs.Identical episodes, initial states, and seeds are used for all methods.
  • Metrics: DSR is the fraction of episodes that both satisfy the task-success predicate and remain within calibrated deformation tolerance, while TSR reports task success alone.The task-success predicate requires the instructed object to rest inside the designated target region and is purely kinematic.
  • Evaluation Regimes: In-distribution evaluation uses 500 held-out episodes per suite and configuration, while out-of-distribution evaluation applies nine held-out conditions across illumination, mass, and stiffness.OOD conditions shift one parameter at a time and reuse each episode’s task, initial state, and seed from in-distribution evaluation.

5 Experiments

SoftVTBench shows that completion-only evaluation can accept physically unsafe rollouts and misrank policies, while deformation-aware scoring reveals policy- and task-dependent interaction quality. Tactile sensing is more consistently associated with robustness under distribution shift than with peak in-distribution performance, and rigid twins enable controlled attribution of deformability effects.

  • 5.1 Deformation-aware evaluation: 24% of Diffusion Policy VT-C successes violate deformation tolerance on Object-Soft, while the TSR–DSR gap is non-zero in all twelve configurations.On Object-Soft and Spatial-Soft, the gaps reach 9.6 and 8.0 percentage points, respectively; these correspond to exact counts of 48 and 40 episodes.
  • 5.2 Policy and task effects: 10–24% of Diffusion Policy and 8–20% of π0.5 successful rollouts violate the safety zone, compared with 0.7–6.5% for FastWAM.FastWAM keeps TSR and DSR within 0.4 percentage points on both spatial configurations, equivalent to two episodes out of 500.
  • 5.2 Policy and task effects: 18.4 percentage points is π0.5’s object-variation loss from rigid to soft, versus 2.6 points for Diffusion Policy and 2.0 for FastWAM.Matched rigid twins control for task difficulty and help separate deformability effects from target grounding or spatial generalization.
  • 5.3 Ablations: 41.4% TSR for VT-C on Object-Soft matches continuous control alone, showing that comparing VO-B against VT-C would overattribute gains to touch.Continuous control alone raises TSR by 11.4 points, while tactile input alone raises it by 10.8 points.
  • 5.3 Ablations: 10.4 points is the Object-Soft DSR advantage of VO-C over VT-B despite only a 0.6-point TSR difference, with safety-zone violations at 7.7% versus 31.7%.Continuous control raises DSR in all four within-modality binary-to-continuous comparisons across the two suites.
  • 5.4 Distribution shift: VT-C exceeds VO-C in TSR across all six policy–suite comparisons under distribution shift and in DSR in five, but touch can widen the TSR–DSR gap.For Diffusion Policy, the gap grows to 6.2 points on Object-Soft and 7.4 on Spatial-Soft; FastWAM stays within one percentage point in every pooled comparison.

6 Conclusion

SoftVTBench combines a synchronized visuo-tactile dataset with a closed-loop benchmark for physical-interaction-aware deformable-object manipulation. Experiments across three policy families show that successful rollouts can exceed calibrated deformation tolerance, underscoring the benchmark’s diagnostic value.

  • Dataset and benchmark: SoftVTBench provides 4,000 expert demonstrations across 40 pick-and-place tasks for physical-interaction-aware deformable-object manipulation.The dataset pairs multi-view RGB, dual-finger tactile RGB and marker motion, proprioception, language, actions, and evaluator-only FEM states.
  • Dataset and benchmark: Matched rigid twins, controlled variations, synchronized recording, and auditable curation support both policy learning and diagnostic evaluation.
  • Benchmark findings: 0.7–24% of each configuration’s successes exceed calibrated deformation tolerance despite achieving task success, across all twelve ID deformable-object configurations.These results come from experiments across three policy families.
  • Benchmark findings: Experiments across three policy families demonstrate the benchmark’s necessity and diagnostic value.

1 Overview

The supplement presents implementation details, calibration results, task-suite specifications, training settings, out-of-distribution results, additional analyses, and tactile-simulation scope. Table 7 summarizes these contents in document order.

  • The supplement provides implementation details, per-object calibration results, and task-suite specifications.
  • It reports baseline training settings, per-condition out-of-distribution results, and additional analyses.
  • It also defines the scope of tactile simulation, with Table 7 summarizing the supplementary sections in document order.

2 Simulation and Implementation Details

SoftVTBench is implemented in Isaac Sim and Isaac Lab using GPU-accelerated PhysX 5. Physics runs at 60 Hz with synchronized control and logging at 20 Hz across observation, action, and evaluator-only physical-state streams.

  • Simulation platform: SoftVTBench uses Isaac Sim 4.5.0, Isaac Lab 0.41.3, and the GPU-accelerated PhysX 5 pipeline.The implementation covers the benchmark’s simulation environment and physics backend.
  • Simulation platform: 60 Hz physics with control decimation of 3 produces a 20 Hz control and logging rate.This rate is used for benchmark interaction and data recording.
  • Synchronized sensing: At 20 Hz, visual, tactile, proprioceptive, action, and evaluator-only physical-state streams are synchronized.The configuration includes robot, control, sensing, simulation, and hardware components.

3 Assets and Per-Object Interaction Safety Zones · 4 Task Suite Details

SoftVTBench combines ten calibrated volumetric deformable assets with matched rigid twins and evaluates them through four matched task suites. The suites total 4,000 demonstrations and include object, spatial, and held-out physical-parameter variation under shared protocols.

  • 3 Assets and Per-Object Interaction Safety Zones: The dataset includes ten volumetric deformable assets: six bakery-style meshes and four procedurally generated geometric primitives simulated as PhysX GPU FEM soft bodies.Each asset has authored density, friction, elasticity, and damping parameters that are withheld from the policy.
  • 3 Assets and Per-Object Interaction Safety Zones: Rigid twins share each deformable counterpart’s mesh, texture, and mass, but their stiffness makes deformation negligible under achievable grasps.They populate Object-Rigid and Spatial-Rigid suites using the same layouts, instructions, and recording stack.
  • 3 Assets and Per-Object Interaction Safety Zones: Each deformable asset has a calibrated gripper envelope, measured aperture span, and object-specific deformation tolerance τo expressed relative to its reference bounding-box diagonal.The aperture span is computed from measured loose and tight jaw apertures rather than a universal command-to-width conversion.
  • 3 Assets and Per-Object Interaction Safety Zones: All safety thresholds use a single globally fixed 90th-percentile calibration value, applied uniformly across assets before policy training and not tuned per object or method.Changing the percentile tightens or loosens every threshold uniformly and rescales each object’s Rmax by a per-object constant.
  • 4 Task Suite Details: The four suites form a matched 2 × 2 design crossing deformable versus rigid objects with object-identity versus spatial-layout variation.Each suite contains 10 tasks and 100 expert demonstrations per task, yielding 1,000 demonstrations per suite and 4,000 overall under shared robot, sensing, rollout, and evaluation protocols.
  • 4 Task Suite Details: Object-Soft varies the manipulated object across fixed layouts, whereas Spatial-Soft uses two visually identical instances and language-based spatial referring expressions to identify the target.Success requires the instructed instance, rather than merely any instance, to reach the target; rigid suites replicate both structures with twins.
  • 4 Task Suite Details: Held-out values lie strictly outside training ranges, and each out-of-distribution condition shifts one variation factor while keeping all others nominal.The nine conditions cover three held-out levels each for dome-light intensity, mass, and Young’s modulus.

5 Baseline Training Details

The baselines comprise six continuous-control vision-only and visuo-tactile configurations across Diffusion Policy, π0.5, and FastWAM, plus two binary-control π0.5 ablations. Visuo-tactile variants extend their corresponding vision-only observation interfaces with tactile and marker-motion inputs while preserving task-specific policy and action frameworks.

  • Six primary continuous-control baselines pair vision-only and visuo-tactile variants of Diffusion Policy, π0.5, and FastWAM.
  • VO-B and VT-B are binary-control ablations of π0.5 that differ from continuous counterparts only in gripper-action encoding.All other training and inference settings remain unchanged.
  • Diffusion Policy: Diffusion Policy’s visuo-tactile variant adds left- and right-finger tactile RGB streams and a 792D marker-motion input alongside the 7D robot state.ResNet-18 encoders process tactile streams, whose representations fuse with visual features.
  • π0.5: π0.5’s visuo-tactile variant retains vision, robot state, and language inputs while adding both-finger tactile RGB and marker-motion observations.Eight past tactile frames per finger are tiled into a single 4×4 mosaic processed by the shared visual encoder.
  • π0.5: Both π0.5 variants predict 50-step chunks of 7D actions, execute 10 actions before replanning, and are LoRA-fine-tuned for 7k steps.Training uses eight NVIDIA A100-80GB GPUs and a global batch size of 256 unless otherwise specified.
  • FastWAM: FastWAM adds a third tactile DiT expert and an anchor-only tri-branch attention mask that prevents action prediction from using future tactile tokens.RGB, tactile, and action branches are jointly optimized with flow-matching loss weights (1.0, 0.2, 1.0).

6 Per-Condition Out-of-Distribution Results

Tables 14–16 report out-of-distribution results separately for each held-out condition, policy family, input modality, and deformable suite. They use TSR for task completion and DSR for task completion while maintaining Rmax ≤1, with DSR nested inside TSR.

  • Per-condition breakdown: Tables 14–16 decompose aggregate out-of-distribution results across individual held-out conditions, policy families, input modalities, and deformable suites.These conditions correspond to those listed in Table 11.
  • Metrics: TSR is the fraction of episodes that complete the task.
  • Metrics: DSR is the fraction of episodes that complete the task and keep Rmax ≤1 throughout.
  • Metrics: Because DSR is nested inside TSR, DSR cannot exceed TSR.

7 Additional Analyses · 8 Tactile Simulation Pipeline and Scope

Additional analyses show that process-level evaluation distinguishes successful rollouts by whether they respect calibrated deformation limits, while tactile simulation combines contact-based rendering components with explicitly limited sim-to-real claims. The benchmark therefore exposes physical interaction differences without claiming validation of its simulated tactile signals on real sensors.

  • 7 Additional Analyses: All displayed rollouts achieve task success, but two exceed the calibrated object-specific deformation limit during interaction.Figure 5 contrasts successful trajectories that satisfy the deformation tolerance with successful trajectories whose normalized deformation exceeds Rt = 1.
  • 7 Additional Analyses: Peak interaction is selected at arg max_t R_t, with synchronized visual and tactile observations drawn from the same rollout.Tactile marker fields expose contact evolution alongside third-person, wrist, and tactile observations.
  • 7 Additional Analyses: At nominal stiffness, both Diffusion Policy variants achieve TSR 42, yet differ by 9 points in DSR and 0.06 in median Rmax.The independent movement of stiffness and deformation axes shows that completion can separate less than deformation-aware evaluation.
  • 8 Tactile Simulation Pipeline and Scope: TacEx generates tactile observations from Isaac Sim contact states, while Taxim renders optical tactile images and FOTS produces marker-motion fields.Each finger provides a 320×240 tactile RGB image and an 11×9 marker representation.
  • 8 Tactile Simulation Pipeline and Scope: The pipeline combines gel–object contact geometry with example-based optical tactile calibration and marker-motion rendering at each control step.The component sequence is TacEx contact simulation, Taxim tactile-image rendering, and FOTS marker-field generation.
  • 8 Tactile Simulation Pipeline and Scope: Published component validations do not establish sim-to-real validity for SoftVTBench’s specific assets, materials, contacts, or rendering parameters.The work performs no real-sensor comparison and makes no claim that the simulated tactile signals transfer to physical GelSight sensors.
Loading 2608.18701v1…