Source-linked AI summary

DARP: A Calibrated Dual-Arm RGB-D-IR Dataset for Multi-View Robotic Perception

Manish Kansana, Mohammed Yusuf Mujawar, Sudip Mittal, Shahram Rahimi, Noorbakhsh Amiri Golilarz

arXiv:2608.31002v2cs.ROcs.CV

TL;DR

Single-view robotic perception is limited by partial visibility, while combining multiple viewpoints requires temporal association and a common spatial frame. DARP provides a calibrated dual-arm multimodal dataset and deterministic measured-surface fusion pipeline, achieving millimetre-scale agreement with held-out RGB-D observations.

  • Problem

    Single-view robotic perception is limited by self-occlusion and incomplete surface visibility, motivating combined observations from multiple viewpoints in a common spatial frame.

  • Method

    DARP combines synchronized RGB-D-IR and robot-state recordings from two eye-in-hand arms with autonomous object-centered acquisition and calibration for shared-frame fusion.

  • Results

    2.13 mm median point-to-mesh distance was achieved over more than 1.56 million unused 3D query points using deterministic measured-surface fusion.

  • Takeaways & Limitations

    The released calibration, robot-state information, and multimodal measurements support metric dual-arm fusion of complementary observed surface regions.

  • Takeaways & Limitations

    Meshes may remain incomplete on unobserved surfaces, reconstruction depends on calibration stability, and the release contains ten objects from one laboratory setup.

Abstract

from arXiv · show

Robotic perception from a single viewpoint is often limited by self-occlusion and incomplete surface visibility. This paper presents DARP(Dual-Arm Robotic Perception) https://doi.org/10.21227/rmv3-be47, a calibrated dual-arm RGB-D-IR dataset for object-centered robotic perception using two independently moving eye-in-hand manipulators positioned on opposite sides of a shared tabletop workspace. Each arm carries an Intel RealSense sensor that continuously records RGB, depth, and stereo infrared data while synchronized robot joint states are logged for pose recovery. Objects are placed without fixed poses or marked locations, and the acquisition procedure performs automatic localization, cross-arm confirmation, adaptive viewpoint generation, and continuous multimodal recording. DARP contains ten unique tabletop objects and preserves the original sensor recordings, robot-state logs, object-level metadata, and calibration information required to reconstruct camera trajectories in a shared metric frame. To evaluate the geometric consistency of the acquisition, we implement a deterministic multi-view fusion pipeline that converts calibrated RGB-D observations into complementary partial point clouds and measured surface meshes without using learned or generative completion methods. Evaluation on 224 held-out RGB-D keyframes comprising 1,563,466 three-dimensional query points yields a median point-to-mesh distance of 2.13~mm and an RMSE of 4.04~mm, with 96.56\% of points within 10~mm of the measured-surface mesh. DARP is intended as a reusable resource for multi-view reconstruction, collaborative robotic perception, multimodal fusion, active perception, and future learning-based reasoning over partial object observations.

I. INTRODUCTION

DARP addresses incomplete single-view robotic perception with a calibrated dual-arm dataset that preserves synchronized multimodal observations and robot motion. Its deterministic fusion evaluation demonstrates geometrically consistent complementary measurements without learned completion.

  • Single-view robotic perception is limited by self-occlusion and incomplete surface visibility, motivating temporally associated observations in a common spatial frame.These limitations matter for object understanding, grasp planning, and interaction.
  • DARP records synchronized RGB, metric depth, stereo infrared, and robot-state observations from two independently moving eye-in-hand manipulators.The arms observe a shared tabletop workspace from opposite sides.
  • Objects are placed without fixed poses or marked locations, while autonomous acquisition localizes objects, confirms them across arms, generates adaptive viewpoints, and records multimodal data continuously.This procedure produces complementary object-centered observations.
  • ChArUco calibration, synchronized joint states, joint-zero corrections, and forward kinematics recover camera poses in a shared metric world frame.The preserved information supports cross-arm observation fusion without requiring a calibration target during normal collection.
  • 2.13 mm median point-to-mesh distance was achieved over more than 1.56 million unused 3D query points using deterministic measured-surface fusion.The pipeline uses calibrated RGB-D observations and does not use trained or generative models.
  • DARP is designed as a task-flexible resource for geometric processing and future applications including segmentation, classification, collaborative perception, active perception, and partial-shape completion.The released data preserve multimodal measurements, robot poses, and calibration information.

II. RELATED WORK

Prior work provides RGB-D object, manipulation, reconstruction, and point-cloud-completion resources, but DARP targets a complementary dual-arm robotic acquisition setting. Its strictly geometric reconstruction is positioned separately from learning-based completion methods.

  • Washington RGB-D, BigBIRD, and YCB provide multi-view RGB-D observations for object-centric perception, while ShapeNet and ScanNet cover digital models and reconstructed indoor scenes.
  • GraspNet-1Billion connects RGB-D sensing with large-scale grasp annotations, whereas ROBI studies multi-view sensing of reflective objects in robotic bin-picking.Their acquisition goals differ from continuous complementary recording by two independently moving eye-in-hand manipulators.
  • Point-cloud completion methods such as PCN, TopNet, PoinTr, and SnowflakeNet map incomplete observations toward more complete shape representations.
  • Generative completion methods include diffusion-based approaches such as PCDreamer and SuperPC, while benchmark performance may not directly transfer to real sensor observations.

C. Collaborative Robotic Perception

Collaborative robotic perception combines multiple agents' observations to reduce individual viewpoint limitations. DARP studies this principle in a close-range, object-centered setting where robot kinematics place complementary eye-in-hand measurements in a shared frame.

  • Collaborative perception combines observations from multiple agents to reduce the limitations of an individual viewpoint.
  • DARP uses two fixed-base manipulators that independently move eye-in-hand cameras around the same tabletop object from complementary viewpoints.
  • Robot kinematics provide the pose of each sensor measurement, enabling embodied viewpoints to be combined at the object-geometry level.The setting also supports later feature-level collaborative perception.
  • Reliable fusion requires camera, wrist, robot-base, and shared-world relationships established through ChArUco-based calibration.
  • Unlike visual-tracking pipelines, DARP derives camera poses primarily from robot calibration, synchronized joint states, and forward kinematics.
  • DARP connects RGB-D perception, robotic manipulation, collaborative perception, and multi-view reconstruction through preserved multimodal observations, robot motion, and calibration.

A. Robotic Platform and Sensors

DARP uses two fixed-base Unitree Z1 Pro arms with eye-in-hand Intel RealSense sensors to capture complementary views in a shared workspace. Calibration and synchronized robot states recover each camera pose in a common frame.

  • Two 6-DoF Unitree Z1 Pro manipulators are mounted on opposite sides of a shared tabletop workspace.
  • Each arm carries a rigidly mounted Intel RealSense camera recording color, depth, and stereo infrared at 640 × 480 resolution and 30 Hz.
  • The manipulators are separated by approximately 1.3986 m and oriented nearly antiparallel, enabling complementary surface observations.
  • ChArUco calibration estimates camera-to-wrist and base-to-base transformations needed to align both sensing systems geometrically.
  • Camera poses are recovered from interpolated joint configurations, joint-zero corrections, forward kinematics, and retained rigid transforms.
  • Held-out hand-eye calibration errors are approximately 3.58 mm and 3.93 mm for the two arms.

C. Autonomous Data Collection

DARP autonomously surveys freely placed tabletop objects by localizing them, confirming their position from both arms, and generating feasible complementary viewpoints for synchronized recording.

  • Objects are placed without manually measured poses, marked locations, or a turntable, so the system estimates position before trajectory generation.
  • The acquisition sequence performs object localization, cross-arm confirmation, adaptive viewpoint planning, synchronized multimodal recording, and post-capture verification.
  • A structured pan-based search obtains an initial depth-based object estimate, which is independently checked from both manipulators.
  • Each arm receives an object-centered trajectory whose candidates are filtered for inverse-kinematics feasibility, workspace limits, clearance, joint motion, and interference.
  • Only one manipulator moves at a time while the other remains parked, and both cameras continuously record changing viewpoints.
  • The collection strategy produces complementary partial observations rather than identical views, providing diversity for later fusion and evaluation.

D. Dataset Organization

DARP preserves raw, synchronized multimodal recordings together with robot states, object metadata, and calibration-related information. The organization supports both direct geometric processing and alternative downstream pipelines.

  • The released dataset contains ten unique physical tabletop objects, with each survey preserved in acquisition form.
  • Each survey includes two native RealSense recordings, a synchronized robot-state log, and object-centered acquisition metadata.
  • The continuous recordings contain RGB, depth, left infrared, and right infrared streams plus profiles, intrinsics, and depth scale information.
  • Optional compressed keyframes package aligned RGB, raw depth, stereo infrared, object-support masks, intrinsics, depth scale, and world-from-camera poses.
  • Raw acquisition files remain primary so alternative frame selection, calibration, segmentation, fusion, or learning pipelines can be applied.

E. Geometric Fusion and Surface Reconstruction

DARP evaluates shared-frame geometric consistency by fusing calibrated RGB-D observations into complementary partial point clouds and measured surface meshes, while preserving unobserved regions.

  • Calibrated RGB-D frames are converted to metric point clouds and transformed into the shared world frame.
  • Object envelopes, depth consistency, support frequency, and statistical filtering remove background, weakly supported, and outlier measurements.
  • Multi-scale ball pivoting converts the fused object cloud into a measured surface mesh without predicting unobserved geometry.
  • Held-out keyframes are excluded from fusion and used to evaluate agreement between reconstructed surfaces and unused RGB-D observations.
  • Across retained recordings, DARP contains 61,012 robot joint-state rows, approximately 18.4 min of acquisition, and more than 63,000 depth frames.
  • The dataset supports analysis at the observation, single-arm sequence, complementary partial-cloud, and dual-arm fused-representation levels.

B. Cross-Arm Geometric Consistency

DARP transforms independently recorded arm observations into a shared metric world frame using calibration and synchronized robot states. The resulting partial clouds show complementary surface coverage while remaining reasonably aligned in co-observed regions.

  • Pose recovery: Calibrated base and hand-eye transforms plus synchronized robot configurations recover each selected camera pose in the shared world frame.This avoids requiring visual re-registration of every frame.
  • Calibration validation: Overlap between independently reconstructed arm-specific point clouds provides an additional check of calibration consistency.Corresponding measurements in co-observed regions should occupy similar locations after world-frame transformation.
  • Visualization: Orthographic XY, XZ, and YZ projections visualize arm110 and arm111 partial point clouds in the shared world frame for the Ketchup object.Red points denote arm110 and blue points denote arm111.
  • Observed consistency: The partial clouds occupy complementary surface regions, while co-observed portions remain reasonably aligned after calibration and pose recovery.This supports coherent combined surfaces without eliminating the distinct geometric contributions of either arm.

C. Held-Out Surface Evaluation

DARP evaluates geometric repeatability by comparing held-out RGB-D observations against a measured surface mesh generated from separate observations. The evaluation reports millimeter-scale agreement while explicitly distinguishing held-out observation agreement from absolute error against a complete independent model.

  • Evaluation protocol: 224 held-out RGB-D keyframes provide 1,563,466 three-dimensional query points for point-to-mesh evaluation.The held-out observations are excluded from fusion before independent back-projection and comparison.
  • Geometric accuracy: 2.13 mm median point-to-mesh distance, 4.04 mm RMSE, and 6.38 mm 90th-percentile distance quantify held-out geometric agreement.These metrics are computed against the measured surface mesh.
  • Distance coverage: 83.82% of held-out points lie within 5 mm and 96.56% lie within 10 mm of the reconstructed surface.The thresholds summarize the distribution of held-out point-to-mesh distances.
  • Interpretation boundary: The reported values measure held-out observation agreement rather than absolute error against a complete object model from an independent high-accuracy scanner.The reconstructed mesh contains only surfaces supported by recorded observations.
  • Qualitative reconstruction: The qualitative reconstruction combines complementary viewpoints, increasing visible-surface coverage while preserving measured rather than deliberately closed geometry.Multi-scale ball-pivoting generates the mesh only over repeatedly supported measured regions.

A. Supported Applications

DARP is designed as a reusable resource for studying calibrated dual-arm, multimodal, object-centered perception. Its preserved measurements, trajectories, and calibration support reconstruction, multimodal fusion, active perception, collaborative manipulation, and future shape-completion research within the stated scope.

  • Reusable resource: Preserved RGB, depth, stereo infrared, joint states, and calibration information support image-, point-cloud-, trajectory-, and fused-surface-level use.The same acquisition can therefore be represented across multiple data modalities and geometric forms.
  • Multi-view reconstruction: Calibrated observations from both arms can be transformed into a shared world frame for more complete measured surfaces and single-arm versus dual-arm studies.Arm-specific observations also enable quantification of complementary-viewpoint benefits.
  • Multimodal perception: RGB, metric depth, and stereo infrared can support segmentation, recognition, feature fusion, and sensor-ablation studies with viewpoint or pose information.Retained robot configuration and camera pose prevent each image from being treated only as an isolated observation.
  • Active perception: Recorded camera trajectories and partial observations provide examples for active perception and next-best-view planning as visibility changes during eye-in-hand motion.Future methods could use these data to predict viewpoints that reduce occlusion or add geometric coverage.
  • Shape completion: Shape completion is a downstream learning task because DARP's geometric pipeline reconstructs only surfaces supported by measured RGB-D observations.Appropriate complete-shape targets or self-supervised objectives are required for completion methods.
  • Collaborative manipulation: The shared metric representation supports collaborative manipulation studies in which one manipulator observes surfaces occluded from the other.Both observations remain related through the common world frame.
  • Dataset positioning: DARP differs primarily through two independently moving opposite-side eye-in-hand sensors, continuous multimodal acquisition, synchronized states, and shared metric calibration.The comparison emphasizes acquisition characteristics rather than identical research problems across datasets.

VI. DISCUSSION AND LIMITATIONS

DARP aligns complementary dual-arm observations in a shared metric frame, but its reconstruction and evaluation remain bounded by visibility, calibration, dataset diversity, and same-acquisition validation. The release is therefore a reusable resource with explicit scope limits rather than a complete-object benchmark.

  • Discussion: DARP aligns complementary observations from independently moving eye-in-hand cameras using calibration, synchronized joint states, and forward kinematics.These components support reconstruction in a shared metric frame.
  • Limitations: Reconstructed meshes contain only observed surface geometry, leaving undersides and strongly occluded regions incomplete without inferred closure.The deterministic pipeline does not artificially complete regions unseen by either camera.
  • Limitations: Cross-arm consistency depends on the robotic calibration state, so mounting, base, bracket, or joint-zero changes may require recalibration.
  • Evaluation Scope: 2.13 mm median point-to-mesh distance and 96.56% of query points within 10 mm were obtained from held-out observations of the same acquisition, not an independently scanned complete model.The current release contains ten unique tabletop objects from a single laboratory setup and does not span broader material, shape, or environmental diversity.
Loading 2608.31002v2…