Source-linked AI summary

Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models

Hongyu Li, Wanjia Fu, Xiaoyan Cong, Zekun Li, Binghao Huang, Hanxiao Jiang, Xintong He, Yiqing Liang, Rao Fu, Tao Lu, Srinath Sridhar, Kevin A. Smith, George Konidaris, Yunzhu Li

arXiv:2607.05390v1cs.ROcs.CV

TL;DR

Deformable-object world modeling is difficult because of high-dimensional dynamics and occluded contact-induced deformations. Deform360 introduces a large-scale visuotactile benchmark with markerless 3D tracking and compares 2D video with 3D particle models, revealing a trade-off between visual scalability and structural priors.

  • Problem

    Accurately predicting deformable-object behavior remains challenging because deformable bodies have theoretically infinite degrees of freedom and contact deformations are often occluded.

  • Method

    Deform360 combines a large multi-view visuotactile dataset with markerless 3D tracking and systematic evaluation of 2D video and 3D particle world models.

  • Results

    3D particle models perform better in low-data future prediction, while 2D video models provide stronger visual reconstruction and zero-shot visual generalizability.

  • Takeaways & Limitations

    The benchmark identifies a trade-off between structural priors that support dynamics generalization and scalability that supports visual synthesis and generalizability.

  • Takeaways & Limitations

    Heavy self-occlusion, highly plastic materials, and visible contact slip can reduce tracking quality or violate the particle optimization assumptions.

Abstract

from arXiv · show

Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling deformable objects presents a particularly difficult case due to their high-dimensional state spaces and complex material properties. While current world models approach this through two distinct paradigms: learning the dynamics over the 2D pixel space or more explicit 3D geometric space. A systematic understanding of their relative strengths and limitations remains elusive due to the lack of diverse, large-scale real-world data. To address this, we present Deform360, a large-scale visuotactile dataset featuring 198 daily-life objects, 1,980 interaction sequences, and over 215 hours of observations from 41 surround-view cameras and bimanual tactile grippers to capture both global motion and contact-induced local deformations. Leveraging a novel markerless visuotactile 3D tracking pipeline to extract dense geometry and motion, we systematically evaluate current state-of-the-art world models, comparing 2D video models against 3D particle models. Finally, we provide a preliminary demonstration indicating the real-world applicability of our dataset by performing robot planning tasks on deformable objects. Our analysis reveals key insights into the trade-offs between structural priors and scalability, providing a solid benchmark for future research in generalizable deformable object-centric world modeling. Project website: https://deform360.lhy.xyz

1 Introduction

Deform360 addresses the difficulty of modeling deformable dynamics by providing a large-scale, diverse, multi-view visuotactile dataset with global and contact-induced deformation observations. It supports contact prediction, comparative 2D-versus-3D world-model evaluation, and real-world robot planning.

  • Motivation: Deformable dynamics remain difficult because bodies have theoretically infinite degrees of freedom and contact-induced local deformations are often occluded.Tactile sensing is therefore important for observing occluded interactions and providing ground-truth signals.
  • Motivation: Current approaches predict deformable dynamics in 2D pixel space or 3D geometric space, but assessing their relative strengths requires diverse, large-scale, multi-view real-world datasets.Such datasets are needed to verify generalizability and extract ground-truth 3D states, while existing datasets often lack object diversity or rely on synthetic data.
  • Dataset and Contributions: Deform360 contains 198 daily-use objects and 1,980 interaction sequences captured with 41 surround-view cameras and bimanual tactile-equipped UMI grippers.The dataset delivers over 215.7 hours and 23.3 million frames, spanning diverse materials and enabling 360° observation of global dynamics and fine-grained contact-induced deformations.
  • Evaluation Tasks: The dataset supports three evaluation tasks: contact prediction, systematic benchmarking of state-of-the-art 2D video models against 3D particle models, and real-world robot planning.These tasks use synchronized tactile data to model visual-physical contact coupling and assess generalization across multiple settings.
  • Perception Pipeline: A markerless multi-view visuotactile perception pipeline produces high-quality 3D reconstructions and tracking results for each interaction episode.The resulting annotations support modeling and evaluation across the dataset’s deformable-object interactions.

2 Related Work

Prior deformable-object datasets are often narrow in object categories, viewpoints, or sensing modalities, whereas Deform360 provides large-scale, multimodal real-world data for dynamics evaluation. Related world-modeling approaches span scalable action-conditioned 2D video models and structurally grounded 3D geometry models.

  • Dataset Benchmarks: Earlier datasets targeted thin-shell objects such as cloth or linear objects such as ropes and cables, often with limited object diversity and camera viewpoints.Robo360 expanded categories and provided 86 camera views but lacked h
  • Dataset Benchmarks: Deform360 provides 23.3 million frames and 215.7 hours of 41-view calibrated, synchronized video-tactile data for 198 diverse objects, with high-fidelity 2D/3D annotations.This scale and sensory richness substantially exceed those of existing benchmarks and support comprehensive dynamics evaluation.
  • Action-conditioned 2D Video Models: Action-conditioned video diffusion models, including, Vid2World, PAN, and Cosmos, predict 3D dynamics implicitly in latent space while offering notable scalability.These models build on video diffusion as a paradigm for world simulation [15] [16].
  • 3D World Models: 3D world models represent objects with explicit meshes or particles, using physics-based optimization or data-driven learning to incorporate stronger structural priors.Differentiable-simulator approaches such as DiffCloth, PhysTwin, and PhysGaussian embed physical priors like mass conservation and achieve longterm stability, but remain suscept

3 The Deform360 Dataset

Deform360 is a large-scale real-world dataset with 198 daily-life objects and 1,980 interaction sequences, captured through synchronized multi-view video, bimanual tactile sensing, and robot proprioception. Its broad material categories and dense annotations target generalizable modeling of global deformation and contact physics.

  • Capture setup: Each sequence synchronizes 41 RGB cameras with bimanual tactile signals at 30 Hz to capture occluded local deformations and contact-induced material responses.The cameras record at 720×1280 resolution and use fixed intrinsic and extrinsic calibration; tactile sensors are mounted on UMI grippers.
  • Object categories: The 198 objects are categorized into 28 1D deformables, 98 2D deformables, and 72 3D volumetric deformables spanning diverse materials and geometric complexities.The categories cover linear objects, thin-shell materials, and squeezable volumetric objects with significant shape change.
  • Interactions and annotations: Episodes include poking, squeezing, stretching, folding, and twisting, while recording gripper pose and openness alongside visual and tactile streams for action reasoning.The annotation pipeline reconstructs dynamic geometry, tracks objects markerlessly in 2D and 3D, and enforces temporal, multi-view, and physical consistency.
  • Dataset scale: 198 objects and 1,980 interaction sequences yield 74,850 raw videos, 23.3 million frames, and 215.7 hours of cumulative multi-view footage.Each object contributes 5 unimanual and 5 bimanual episodes; the setup uses 41 viewpoints at 720p and 30 FPS.

4 Annotation Pipeline

The annotation pipeline decouples per-frame 3D geometry recovery from temporally consistent particle tracking by lifting multi-view 2D trajectories into 3D and refining them with visuotactile physics. It produces warped point clouds that preserve particle identity under deformation and occlusion.

  • Calibration: Multi-view calibration, undistortion, and RANSAC aggregation of ArUco-tracked gripper poses establish metric-scale geometry and gripper motion.
  • Geometry Recovery: Object segmentation filters each camera stream before per-frame 3D Gaussian Splatting reconstruction, restricting recovered geometry to the interacting object.3DGS represents scenes with anisotropic Gaussians and rasterizes color and depth maps.
  • Motion Tracking: Up to M = 1,600 mask-filtered points are tracked across views with CoTracker3, then lifted into 3D and fused with RANSAC into a global velocity field.Tracking uses 15-frame clips with stride 5 to address invisible points caused by self-occlusion.
  • Physics-Informed Tracking: The physics-informed tracking objective combines shape alignment, local ARAP rigidity, Laplacian smoothness, and tactile constraints to make particle motion temporally coherent and physically plausible.Tactile-influenced particles are excluded from shape loss, while tactile feedback acts as a soft localized no-slip regularizer around activated taxels.
  • Output: The optimized velocities are integrated from initial particle positions into warped point clouds that preserve particle identity under large deformations and self-occlusions.The experiments use λlocal = 20.0, λlap = 0.1, and λtactile = 1.0.

5 Experiments

Deform360 is evaluated through perception fidelity, visual-to-tactile contact prediction, cross-paradigm world-model benchmarking, and zero-shot robot planning. The experiments show complementary strengths of physics-based and video models across data regimes, generalization settings, and deployment conditions.

  • 5.1 Perception evaluation: The perception pipeline achieves consistently high reconstruction fidelity across 198 objects and 1D, 2D, and 3D volumetric categories using PSNR, SSIM, and LPIPS on held-out views.Performance is reported by object category in Table 2.
  • 5.2 Visuotactile coupling: Visual observations predict binary tactile contact with 88.67% mean accuracy across 36 synchronization-filtered views, demonstrating the dataset’s synchronized visuotactile coupling.A transformer encoder maps visual streams and robot actions to contact predictions.
  • 5.3 World-model evaluation: PhysTwin outperforms PGND and ParticleFormer in per-episode reconstruction and prediction, highlighting the importance of explicit physical and structural priors in low-data regimes.Cosmos is excluded because extremely limited per-episode data cannot support stable post-training.
  • 5.3 World-model evaluation: Cosmos achieves superior multi-episode reconstruction, whereas ParticleFormer outperforms it in future prediction, revealing a trade-off between visual fidelity and dynamics forecasting.Cosmos better preserves textures and surface details through direct 2D modeling, while ParticleFormer benefits from limited interaction episodes for prediction.
  • 5.3 World-model evaluation: Cosmos generalizes better to novel object categories in the zero-shot setting, leveraging large-scale pretraining to outperform 3D counterparts on image-quality metrics.PhysTwin cannot natively generalize to novel categories because it fits physical parameters to specific object geometries.
  • 5.3 World-model evaluation: Cosmos often produces physically reasonable long-horizon dynamics but fails to follow robot commands strictly, indicating action misalignment as a key failure mode.The passage suggests more fine-tuning data or more sophisticated action representations could help resolve this issue.
  • 5.4 Real-world planning: PhysTwin representations enable zero-shot MPC planning on a different xArm setup and in another laboratory environment, demonstrating preliminary real-world applicability.Cosmos is not deployed because of appearance-shift sensitivity and the difficulty of defining geometric rewards directly on generated videos.

6 Limitations

Deform360 still faces tracking and modeling challenges in difficult deformable interactions, including prolonged self-occlusion, highly plastic materials, and visible contact slip.

  • 6 Limitations: Heavy self-occlusion can reduce tracking quality when large object regions remain invisible to most cameras for extended periods.This limitation affects difficult cases despite Deform360’s broad coverage of everyday deformable interactions.
  • 6 Limitations: Highly plastic materials may violate the local rigidity and smoothness assumptions used in particle optimization.
  • 6 Limitations: Visible slip at contact can cause the tactile no-slip regularizer to over-constrain nearby particles.

7 Conclusion · A Gripper Pose and Openness Tracking · B Preliminaries

Deform360 introduces a large-scale multi-view visuotactile dataset and markerless 3D tracking pipeline for modeling deformable-object dynamics, while benchmarking 3D particle models against 2D action-conditioned video models. The supporting methods track gripper pose and openness through multi-view ArUco observations, and frame deformable dynamics through 2D video and explicit 3D particle paradigms.

  • 7 Conclusion: Deform360 provides multi-view visuotactile data for 198 daily-life objects, extracting high-fidelity particle trajectories and dynamic geometry with a markerless 3D tracking pipeline.The dataset supports systematic benchmarking of 3D particle dynamics models against 2D action-conditioned video generation models.
  • A Gripper Pose and Openness Tracking: The gripper-tracking method attaches eight ArUco markers per gripper, covering the wrist/base and both fingers for 6D pose and openness estimation.Markers use the DICT_4X4_250 dictionary, with distinct marker IDs assigned to wrist/base and finger locations.
  • A Gripper Pose and Openness Tracking: Each visible marker’s camera-frame 6D pose is transformed into the world frame using calibrated camera extrinsics and its CAD-measured gripper-frame offset.The transformation composes world-to-camera, camera-to-marker, and marker-to-gripper transforms.
  • A Gripper Pose and Openness Tracking: Robust multi-view fusion removes occlusion- and noise-induced outliers with RANSAC, then averages inlier translations and rotations to estimate the wrist pose.Rotation averaging uses SVD to keep the result in SO(3).
  • A Gripper Pose and Openness Tracking: Gripper openness is computed as the Euclidean distance between the tracked left and right finger positions.This openness signal complements the fused wrist pose during interaction tracking.
  • B Preliminaries: Deformable object-centric world models follow two paradigms: 2D video models predict future image- or latent-space observations, whereas 3D particle models predict explicit geometric dynamics.The 3D approach represents objects using structured geometric entities, most commonly particles.

B.1 Action-conditioned Video Models

Action-conditioned video world models predict future frames from past observations and robot actions, typically using latent diffusion for computational efficiency. In this evaluation, Cosmos Predict 2.5 is post-trained with 7D robot actions injected through cross-attention.

  • Action-conditioned Video Models: These models predict future video frames from past observations and robot actions.The conditioning includes past frames F1:t and future action sequence A1:t+k.
  • Action-conditioned Video Models: Latent diffusion models encode RGB videos into compressed latent representations with a VAE and model dynamics using a Diffusion Transformer.The DiT learns to reverse Gaussian denoising in latent space, conditioned on past latents and future actions.
  • Action-conditioned Video Models: Internet-scale training enables these models to capture complex visual and physical priors.Latent-space operation is intended to improve computational efficiency.
  • Action-conditioned Video Models: The evaluation post-trains Cosmos Predict 2.5, a DiT-based video model, by injecting 7D robot actions into DiT blocks through cross-attention.The action sequence comprises 6D wrist pose and gripper openness, enabling simulation of future object states.

B.2 Particle-based Dynamics Models

Particle-based dynamics models represent deformable objects as particles with positions, velocities, and attributes, whose states evolve through learned or physics-based transition functions. The section covers graph-based learning, differentiable spring-mass simulation, and the evaluated implementations.

  • Particle-based representation: A deformable object is represented by N particles, each storing position, velocity, and attributes such as mass or material properties.The system state G_t evolves through a transition function T(G_t, A_t).
  • Learning-based Dynamics: Learning-based models approximate particle-state transitions with neural networks trained on large-scale trajectories, using graph message passing to model interactions.The evaluation uses PGND, which combines particles for shape tracking with spatial grids.
  • Physics-based Differentiable Simulation: Physics-based differentiable models embed simulators into learning loops, commonly using spring-mass systems whose forces include elasticity, damping, and external forces.Differentiable simulation enables physical parameters to be optimized by backpropagating prediction-error gradients to the initial state and material properties.
  • Benchmark implementations: The benchmark evaluates Cosmos Predict 2.5, PhysTwin, PGND [86], and a faithful reproduction of ParticleFormer because its official implementation was unavailable.The first three use official open-source implementations, while ParticleFormer follows its paper’s implementation details.

C Multi-Episode Generalization Visualization

Figure 9 visualizes multi-episode generalization on unseen glove and sack interactions by comparing ground-truth sequences with action-conditioned predictions from four world models. PhysTwin* is evaluated differently because its optimization-based simulator cannot zero-shot predict unseen episodes and instead observes 80% of transitions.

  • Multi-Episode Generalization Visualization: Figure 9 compares ground-truth sequences with action-conditioned predictions from Cosmos, ParticleFormer, PGND, and PhysTwin on unseen glove and sack interactions.The visualization covers glove-cloth and sack-cloth objects under the episode-generalization setting.
  • Multi-Episode Generalization Visualization: PhysTwin* cannot directly predict an entirely unseen episode zero-shot because it requires system identification from an observed trajectory.For PhysTwin*, the optimization process is allowed to observe 80% of the transitions.

D Dataset Object Taxonomy and Diversity … D.3 3D Volumetric Deformables

Deform360 organizes 198 daily-life deformable objects into 1D, 2D, and 3D volumetric classes spanning 17 everyday semantic categories. The taxonomy covers diverse material responses and interaction challenges, from knotting and folding to internal-stress-driven volumetric deformation.

  • D Dataset Object Taxonomy and Diversity: Deform360 categorizes 198 objects into 28 1D, 98 2D, and 72 3D volumetric deformables spanning 17 everyday semantic categories.The category distribution is naturally long-tailed.
  • D.1 1D Deformables: 1D deformables include ropes, cables, and wires with varied stiffness, thickness, and surface texture that can knot, entangle, and self-occlude during manipulation.They are divided into thick rope-like, thin rope-like, and stiff cable-like subcategories.
  • D.1 1D Deformables: The 1D taxonomy distinguishes thick rope-like objects with higher bending stiffness, thin rope-like objects with high flexibility, and stiff cable-like objects with elastic memory and torsional stiffness.Examples include cotton and climbing ropes, threads and ribbons, and USB cables, hoses, and chains.
  • D.2 2D Deformables: 2D deformables comprise fabrics, cloths, garments, and paper-like materials exhibiting folding, wrinkling, and multi-modal contact dynamics.Their properties range from highly compliant silk to relatively stiff paper and plastic, enabling tests across elasticity and plasticity.
  • D.2 2D Deformables: The 2D collection further includes cloths and fabrics, garments, bag-like and paper-like objects, and other thin-shell deformables with complex geometry or nonlinear responses.Examples include shirts, socks, hats, gloves, airbags, trash bags, and paper.
  • D.3 3D Volumetric Deformables: 3D volumetric deformables have non-negligible thickness and complex internal structures, including plush toys, stuffed animals, and foam-based objects.Their interactions can produce volume-preserving or volume-changing shape changes, requiring modeling of internal stress and long-range dependencies within the object volume.
  • D.3 3D Volumetric Deformables: The 3D taxonomy divides objects into stuffed animals and plush toys, plus foam and squeezables with varying density and recovery rates.Examples include teddy bears, animal toys, sponges, stress balls, hand sanitizers, and shoes.

E Dataset Visualizations

The dataset visualizations present diverse 1D, 2D, and 3D volumetric deformable objects alongside synchronized multi-view captures. The 41-camera setup illustrates how dense viewpoints capture complex interactions and reduce self-occlusion.

  • E Dataset Visualizations: Figures 10–16 visualize the collected 1D, 2D, and 3D volumetric deformable objects, demonstrating the dataset’s broad object diversity.The visualizations include separate examples for 1D, 2D, and 3D deformables.
  • E Dataset Visualizations: Figures 14–16 specifically visualize 3D deformables, complementing the single-view coverage of 1D and 2D objects.These figures continue the dataset’s object-level visualization across volumetric deformables.
Loading 2607.05390v1…