Source-linked AI summary

ArtVIP: Articulated Digital Assets of Visual Realism, Modular Interaction, and Physical Fidelity for Robot Learning

Zhao Jin, Zhengping Che, Tao Li, Zhen Zhao, Kun Wu, Yuheng Zhang, Yinuo Zhao, Zehui Liu, Qiang Zhang, Xiaozhu Ju, Jing Tian, Yousong Xue, Jian Tang

arXiv:2506.04941v3cs.RO

TL;DR

Robot learning needs higher-quality articulated assets because existing simulation datasets have limited visual realism and physical fidelity. ArtVIP addresses this gap with standardized digital twins, tuned dynamics, embedded modular behaviors, annotations, and scene assets, and its fidelity is supported by real-world and simulation evaluations.

  • Problem

    Existing articulated-object simulation datasets have limited visual realism and physical fidelity, motivating higher-quality assets for robot learning.

  • Method

    ArtVIP combines professionally modeled digital-twin articulated objects, fine-tuned physical parameters, embedded modular interactions, pixel-level annotations, and indoor scenes.

  • Results

    ArtVIP shows close agreement between simulated and real drawer trajectories and supports simulation-trained policies achieving zero-shot real-world success.

  • Takeaways & Limitations

    ArtVIP provides reusable USD assets with production guidelines and demonstrates applicability in imitation learning and reinforcement learning.

  • Takeaways & Limitations

    Scaling remains limited by the intensive human labor required for asset modeling.

Abstract

from arXiv · show

Robot learning increasingly relies on simulation to advance complex ability such as dexterous manipulations and precise interactions, necessitating high-quality digital assets to bridge the sim-to-real gap. However, existing open-source articulated-object datasets for simulation are limited by insufficient visual realism and low physical fidelity, which hinder their utility for training models mastering robotic tasks in real world. To address these challenges, we introduce ArtVIP, a comprehensive open-source dataset comprising high-quality digital-twin articulated objects, accompanied by indoor-scene assets. Crafted by professional 3D modelers adhering to unified standards, ArtVIP ensures visual realism through precise geometric meshes and high-resolution textures, while physical fidelity is achieved via fine-tuned dynamic parameters. Meanwhile, the dataset pioneers embedded modular interaction behaviors within assets and pixel-level affordance annotations. Feature-map visualization and optical motion capture are employed to quantitatively demonstrate ArtVIP's visual and physical fidelity, with its applicability validated across imitation learning and reinforcement learning experiments. Provided in USD format with detailed production guidelines, ArtVIP is fully open-source, benefiting the research community and advancing robot learning research. Our project is at https://x-humanoid-artvip.github.io/ .

1 INTRODUCTION

ArtVIP addresses the quality bottleneck in articulated-object simulation by combining realistic, physically faithful assets with modular interactions, annotations, and deployable scenes. It releases a large, standardized collection validated for robot-learning applications.

  • Motivation: Simulation offers low-cost data, virtual environments, rapid deployment, and standardized testing without hardware-damage or safety concerns.These properties motivate simulation as an efficient alternative for advancing robot learning.
  • Motivation: High-quality articulated-object assets are needed because existing datasets provide limited visual realism, physical fidelity, or simulator flexibility.The paper identifies asset quality, rather than quantity, as the central bottleneck and specifies visual realism, modular interaction, physical fidelity, and simulation friendliness as requirements.
  • ArtVIP: ArtVIP embeds interaction semantics within assets and supplies pixel-level affordance annotations, enabling modular reuse and scalable behavior modeling.These features are designed to support manipulation skills including rotating, clicking, pulling, and pressing.
  • Applications: The dataset includes configured indoor scenes and supports imitation-learning, reinforcement-learning, and 3D-construction experiments.Assets are distributed in USD and can be converted to legacy formats such as URDF or MJCF for established robotics workflows.

2 RELATED WORKS

Prior simulation datasets provide useful assets but remain fragmented across navigation, rigid-body modeling, articulated manipulation, and asset-generation settings. ArtVIP is positioned against limitations in articulation support, standardization, visual realism, and physical fidelity.

  • Simulation Platforms: Simulation platforms combine physics and rendering engines, but differing platform capabilities make broadly compatible, ready-to-use assets valuable.The paper contrasts robotics-oriented simulators such as MuJoCo and Webots with game engines that do not natively support ROS.
  • Datasets for Robot Simulation: Indoor-scene datasets support navigation but lack GUI-based editing, while ShapeNet, Objaverse, and related assets generally function only as rigid bodies.Rigid-body-only assets cannot support articulated manipulation tasks.
  • Articulated Objects: Existing articulated-object datasets and construction methods do not consistently combine visual realism, physical fidelity, simulator compatibility, and scalable production.The related work discusses both dataset limitations and generation methods that remain reliable mainly for simple joints.
  • Construction and Generation: Public-repository assets often have inconsistent modeling quality, disorganized part hierarchies, and non-standardized coordinate systems requiring manual preprocessing.Generative and reconstruction methods scale more easily but are not yet mature enough to ensure quality.

3 ARTVIP COLLECTION AND METHODOLOGY

ArtVIP constructs articulated assets through standardized hierarchical modeling, tuned collision and joint dynamics, and embedded reusable interaction behaviors. The methodology prioritizes fidelity while retaining simulation efficiency and direct usability.

  • 3.1 VISUAL REALISM: Professional modelers build each object hierarchically from assembly to module to mesh under unified modeling, texture, material, and coordinate guidelines.The top-down mechanical approach organizes rigid-body modules and their mesh parts before bottom-up assembly and joint integration.
  • 3.2 PHYSICAL FIDELITY: ArtVIP balances contact fidelity and computational efficiency by selecting convex hulls, convex decompositions, or fine-tuned collision meshes for each object.The collision representation is chosen according to geometry complexity and affordance requirements.
  • 3.2 PHYSICAL FIDELITY: The joint model extends basic stiffness-and-damping dynamics by making stiffness and target terms functions of joint position, with velocity dependence when needed.This parameterization represents variable dynamics in complex joints such as door closers and light switches.
  • 3.3 MODULAR INTERACTION: ArtVIP embeds customizable behaviors directly in assets, covering five canonical interaction primitives across 394 assets and more than 900 joints.The primitives include latching or magnetic closure, damping, cross-asset effects, within-asset effects, and hover or hold position.
  • 3.3 MODULAR INTERACTION: Binding behaviors at design time lets users import a USD file and obtain interaction affordances without writing additional code.The modular design is intended to reduce development overhead and accelerate algorithm iteration.

4 EVALUATION

ArtVIP is evaluated for visual realism and physical fidelity through geometric, reconstruction, feature-distribution, and joint-motion comparisons. The evaluations show richer geometric detail, better reconstruction quality, closer alignment with real-world visual features, and close agreement between simulated and real drawer displacement.

  • Evaluation Scope: The evaluation combines digital-twin comparisons with quantitative visual and physical tests to assess ArtVIP’s realism and interaction fidelity.The visual tests include triangle counts, reconstruction performance, and feature distributions, while physical testing tracks joint motion against real-world trajectories.
  • Visual Realism Evaluations: ArtVIP assets show richer geometric detail than comparison datasets, with dense triangular meshes preserving surface smoothness and minimizing faceting.The comparison covers object categories appearing in all three datasets.
  • Visual Realism Evaluations: ArtVIP feature distributions align more closely with real-world data than OmniGibson features in the CLIP-based comparison.The analysis samples 100 ArtVIP models and corresponding or semantically similar OmniGibson and real-world objects using matched viewpoints.
  • Visual Realism Evaluations: ArtVIP produces reconstructions with higher structural fidelity and finer detail preservation than OmniGibson under identical multi-view sampling.The evaluation uses VGGT on matched asset views.
  • Physical Fidelity and Interaction Evaluations: Simulated and real drawer displacement trajectories closely agree across applied forces, demonstrating physical fidelity of ArtVIP joints.The real-world trials apply horizontal pulling forces of 1 N, 1.5 N, 2 N, and 2.5 N, while simulation compares default and optimized joint parameters.

5 APPLICATIONS

ArtVIP is applied to imitation and reinforcement learning in matched real-world and simulated robotic environments. Simulation-trained policies transfer to real tasks, mixed real-sim data improves imitation performance, higher-quality assets strengthen microwave-task transfer, and simulated and real RL performance are strongly correlated.

  • Experimental Setup: The imitation-learning study uses four tasks, 100 successful real trajectories and 100 simulated trajectories per task, and ACT and DP baselines.The tasks are PullDrawer, OpenCabinet, SlideShelf, and CloseOven, with RGB observations from four viewpoints and proprioceptive states.
  • Imitation Learning in Real World Environments: Simulation-trained policies achieve zero-shot real-world success, while mixing real and simulated data improves performance across articulated-object manipulation tasks.For example, ACT reaches 39% on PullDrawer with simulation-only training, and DP on SlideShelf rises from 44% to 59% with mixed data.
  • Imitation Learning in Real World Environments: Equal-volume real-world training still outperforms simulation, with DP reaching 49% versus 10% on OpenCabinet, showing persistent sim-to-real challenges.The experiments compare real-only, sim-only, and mixed datasets for ACT and DP across four tasks.
  • Digital-Cousin Comparison: Higher-quality ArtVIP assets produce stronger zero-shot sim-to-real transfer and higher mixed-data success than PartNet-Mobility on the microwave door-pull task.The digital-cousin comparison uses five microwaves from each dataset and evaluates real-only, sim-only, and mixed settings.
  • Reinforcement Learning in High-Fidelity Simulators: Simulation-to-real reinforcement-learning performance has a Pearson correlation coefficient of 0.9886 across checkpoints, supporting ArtVIP’s physical and perceptual fidelity.A policy trained in Isaac Sim is deployed unchanged in the real world and evaluated across five checkpoints.

6 LIMITATION AND CONCLUSION

ArtVIP combines visually realistic articulated objects, physically tuned joints, modular interactions, indoor scenes, and annotations for robotic manipulation. The authors validate its quality across visual, physical, and learning-oriented evaluations while identifying labor-intensive asset modeling as a scaling limitation.

  • Conclusion: ArtVIP provides visual realism, accurate physical properties, and modular interaction capabilities for articulated-object robotic manipulation.The dataset includes 992 objects across 9 categories and 37 subcategories, plus six physically interactive indoor environments.
  • Limitations: Scaling remains limited by the intensive human labor required for asset modeling, motivating future generative synthesis methods.The stated future direction is to reduce manual effort while broadening object diversity.
  • Annotations: ArtVIP’s annotations describe functional object parts to support task-appropriate interaction behavior.The annotation design depends on consistent modeling standards that preserve functionally distinct components.
  • Visual modeling standards: High-quality meshes, textures, and materials support realistic appearance and reliable collision-based interaction in simulation.ArtVIP uses manifold meshes, high-resolution UV-aligned textures, and physically based rendering with RTX Renderer.
  • Physical fidelity: ArtVIP models complex joint behavior by extending basic drive dynamics with position- and velocity-dependent stiffness, targets, damping, and friction.The friction model covers static, maximum static, and dynamic friction, while latch and door-closer mechanisms use state- and position-dependent transitions.

F CHART COMPARISON WITH EXISTING DATASETS

The paper presents a detailed comparison between ArtVIP and existing articulated-object datasets. The comparison is documented in Table 6.

  • Chart comparison with existing datasets: Table 6 provides a detailed comparison of ArtVIP with existing articulated-object datasets.

G VISUAL REALISM COMPARISON

ArtVIP’s visual realism comparison examines mesh complexity and simulation performance against PartNet-Mobility and BEHAVIOR-1K. The evaluation reports how triangle count and active joints affect frame rate across isolated assets and complex scenes.

  • Visual realism comparison: ArtVIP uses more triangular faces than BEHAVIOR-1K and PartNet-Mobility for categories such as toilets and refrigerators, improving geometric detail at a frame-rate cost.Profiling was conducted to optimize simulation frame rate for each object.
  • Visual realism comparison: Table 7 reports category-wise averages for triangle count, active joints, and FPS across ArtVIP, PartNet-Mobility, and BEHAVIOR-1K.
  • Visual realism comparison: Approximately 100k triangles and up to 20 active joints have negligible FPS impact for a single object, while both factors reduce FPS in complex scenes.The kitchen scene contains 65 actuated joints, making scene complexity an important performance condition.

H PHYSICAL FIDELITY AND INTERACTION EVALUATIONS

The evaluation examines modular interaction and articulated manipulation through real-versus-simulation motion analyses and four imitation-learning tasks. ACT and Diffusion Policy are trained from multi-view RGB and proprioceptive inputs to produce robot control signals.

  • Physical fidelity evaluation: Microwave button-press experiments compare real-world and simulated door-opening trajectories to validate modular interaction behavior.Real motion is captured with optical tracking, while simulation uses a virtual marker.
  • Physical fidelity evaluation: Door-closer evaluation analyzes linear and angular velocities as the door moves from a threshold state to full closure.
  • Task design: Four manipulation tasks—PullDrawer, OpenCabinet, SlideShelf, and CloseOven—require rotation, angled pushing, and horizontal translation.
  • Imitation learning application: ACT and Diffusion Policy receive multi-view RGB images and proprioceptive states, then output robot control signals for end-to-end task execution.Implementation details are provided in Tables 8 and 9.
  • Imitation learning application: The full imitation-learning experiment results are presented in Table 10.

J REINFORCEMENT LEARNING APPLICATION

ArtVIP is evaluated on the long-horizon CloseTrashcan task using a two-stage visual reinforcement-learning pipeline with PPO training and visuomotor distillation. EAGLE substantially outperforms vision-based PPO in the reported comparison.

  • Task and Training Setup: The CloseTrashcan task requires an agent to approach the lid and close it smoothly within a time limit.
  • Training Details: Stage 1 trains a PPO teacher from privileged low-level states, while Stage 2 distills that teacher into a visuomotor student.
  • Reward Functions: The multi-objective reward combines gripper–lid proximity, directional alignment, closure progress, and action smoothness.The reported weights are λ1 = 0.5, λ2 = 0.125, λ3 = 10, and λ4 = −0.01.
  • Baseline Comparison: EAGLE achieves a 98% success rate after 500k training iterations, whereas the vision-based PPO baseline performs poorly on CloseTrashcan.The comparison attributes the baseline’s poor performance to high computational complexity and low data diversity.

K PEARSON CORRELATION COEFFICIENT DETAILS

The Pearson correlation analysis compares simulated and real-world success rates at corresponding checkpoints. The resulting high correlation supports ArtVIP as a reliable simulated training and evaluation pipeline for reinforcement learning.

  • r = 0.9886 indicates a strong linear relationship between simulated and real-world success rates across checkpoints.The coefficient is computed from corresponding simulation and real-world success rates in Tab. 3.

L COMPARISON TO GENERATIVE PIPELINES

ArtVIP digital-twin assets outperform reproduced generative outputs in visual and physical simulation readiness. The comparison identifies real-image reconstruction limitations and efficiency trade-offs affecting current generative pipelines.

  • Generative outputs exhibit self-collisions, mesh distortions, incorrect joint parameters, unrealistic materials, and missing interior details, unlike ArtVIP digital-twin assets.The reported failures lead the authors to conclude that current generative baselines are not simulation-ready for articulated assets.
  • On real-world cabinet and fridge inputs, generative reconstruction performance degrades markedly compared with synthetic inputs, revealing a substantial sim-to-real gap.The comparison uses reconstruction metrics reported in Tab. 13 for the same object categories.
  • Generated cabinet and fridge assets contain approximately 100× more triangles than ArtVIP counterparts, reducing simulation frame rate from approximately 90 fps to 70 fps.
  • Generative pipelines commonly rely on synthetic training data with calibrated cameras and controllable materials, limiting their robustness on real images.
  • Real-image reconstruction accuracy degrades when camera extrinsics are estimated rather than known.
  • Generalization is constrained by PartNet-Mobility category coverage and articulation priors, while low-texture planar structures may appear warped.
  • Reconstruction outputs may lack physically consistent materials and textures, while spherical-harmonic colors require conversion or baking for simulator rendering.
  • Current metrics do not capture mesh efficiency or collision performance relevant to simulation frame rate.
Loading 2506.04941v3…