Source-linked AI summary

Benchmarking in Manipulation Research: The YCB Object and Model Set and Benchmarking Protocols

Berk Calli, Aaron Walsman, Arjun Singh, Siddhartha Srinivasa, Pieter Abbeel, Aaron M. Dollar

arXiv:1502.03143v1cs.RO

TL;DR

Manipulation research lacks a common basis for quantitatively comparing approaches because researchers often choose their own objects and tasks. This paper introduces the YCB Object and Model set, standardized protocol guidelines, and example benchmarks to support reproducible evaluation across manipulation research, while noting that simplified meshes remain future work.

  • Problem

    Researchers commonly select their own objects and tasks, making experimental results difficult to interpret quantitatively against a shared basis.

  • Method

    The paper provides a widely distributed object and model set, a protocol template, and example benchmarks spanning mechanical design, planning, dexterity, and learning.

  • Results

    The paper delivers standardized objects, high-resolution scans and models, and six example protocols intended to support replicable research and performance comparison.

  • Takeaways & Limitations

    Detailed tasks and protocols involving common objects are presented as the basis for replicable research, performance comparison, and continued benchmark evolution.

  • Takeaways & Limitations

    Simplified meshes are not yet provided, although they are identified as future work for studying mesh approximation in motion-planning benchmarks.

Abstract

from arXiv · show

In this paper we present the Yale-CMU-Berkeley (YCB) Object and Model set, intended to be used to facilitate benchmarking in robotic manipulation, prosthetic design and rehabilitation research. The objects in the set are designed to cover a wide range of aspects of the manipulation problem; it includes objects of daily life with different shapes, sizes, textures, weight and rigidity, as well as some widely used manipulation tests. The associated database provides high-resolution RGBD scans, physical properties, and geometric models of the objects for easy incorporation into manipulation and planning software platforms. In addition to describing the objects and models in the set along with how they were chosen and derived, we provide a framework and a number of example task protocols, laying out how the set can be used to quantitatively evaluate a range of manipulation approaches including planning, learning, mechanical design, control, and many others. A comprehensive literature survey on existing benchmarks and object datasets is also presented and their scope and limitations are discussed. The set will be freely distributed to research groups worldwide at a series of tutorials at robotics conferences, and will be otherwise available at a reasonable purchase cost. It is our hope that the ready availability of this set along with the ground laid in terms of protocol templates will enable the community of manipulation researchers to more easily compare approaches as well as continually evolve benchmarking tests as the field matures.

I. INTRODUCTION

The paper addresses the difficulty of comparing manipulation research conducted with disparate objects and tasks by proposing a common physical object set, associated models, and standardized benchmarking protocols. The YCB set is designed for broad manipulation applications while remaining practical to disseminate and use in simulation and real experiments.

  • Researchers commonly choose their own objects and tasks, preventing experimental results from being analyzed against a common basis.
  • Prior object datasets often provide only models or images, while few make physical objects available; this limits manipulation benchmarking because important manipulation phenomena cannot be fully modeled.
  • The YCB set combines physical objects spanning varied shapes, sizes, weights, rigidities, textures, and manipulation applications with practical constraints on shipping, storage, cost, durability, and availability.
  • The accompanying database provides textured mesh models, high-quality images, and physical properties, with integration into MoveIt and the ROS manipulation stack for realistic simulation and planning.
  • The paper provides a protocol template and example benchmarks for evaluating mechanical design, manipulation planning, dexterity, and learning on real robotic systems.

B. Prosthetics and Rehabilitation

The paper reviews rehabilitation and prosthetics tests whose objects and setups are often commercial, difficult to standardize, or unavailable. It responds by including commonly used assessment objects and established manipulation-test equipment in the YCB set.

  • Rehabilitation evaluation includes widely published commercial tests as well as tests proposed in the literature with varying adoption.
  • Commercial assessments such as the Box and Blocks Test and 9-hole-peg test use protocol-specific setups and timed movements of simple objects.
  • Several noncommercial tests use objects that are difficult to obtain or define consistently, including outdated items and nonstandardized objects.
  • The YCB set includes objects common to rehabilitation assessments, daily-living activities, and widely used tests such as the 9-hole peg, box-and-blocks, and clothes-peg allocation tasks.

III. THE OBJECT AND DATA SET

The YCB Object and Model Set combines diverse, practically selected objects with associated manipulation tasks to support broad, accessible benchmarking research.

  • Section scope: The paper presents the object set together with its selection rationale, models, data, and example manipulation tasks.The proposed set is illustrated in Figures 1–7 and organized into the stated object categories.
  • Object selection: Objects were selected from frequently used daily-life, simulation, experimental, and assessment objects.The set includes common rehabilitation objects and activities such as mugs, pitchers, washers, bolts, kitchen items, pens, keys, and padlocks.
  • Practical constraints: Practical constraints included durability, approximately $350 cost, and portability within a large suitcase and a 22kg airline limit.The design favors standard consumer products over custom-fabricated objects and avoids fragile or perishable items.
  • Object selection: The set includes food, kitchen, tool, shape, and task items chosen to cover varied manipulation properties and applications.Selection considered shape, size, transparency, deformability, texture, manipulation difficulty, and use across grasping, daily activities, assembly, and rehabilitation tests.

B. Scans

The YCB dataset provides standardized visual scans, segmentation data, calibration, and textured meshes for integrating physical objects into simulation and planning.

  • Model generation: Poisson surface reconstruction generates watertight meshes and projected segmentation masks from the scans.Transparent or reflective regions may lack depth data, causing reconstruction failures; improved RGB-based models are identified as future work.
  • Scan data: Each object has 600 RGBD images, 600 high-resolution RGB images, segmentation masks, calibration information, and a texture-mapped 3D mesh.The scanning rig uses five RGBD sensors and five high-resolution RGB cameras with 120 turntable orientations.
  • Software integration: The meshes can be integrated directly into MoveIt and OpenRAVE, while URDF files can be generated automatically for ROS.The representations support collision, visualization, mass-property specification, signed-distance fields, and mesh-based collision checking.
  • Acquisition setup: The scanning rig is a computer-controlled turntable system used to collect the object data.The rig is shown in Figure 8, while Figure 9 illustrates point-cloud and textural overlays on YCB objects.
  • Limitations: Simplified meshes are not yet provided, although they could reduce collision-checking time in motion-planning benchmarks.The paper identifies mesh approximation and its impact on standardized planning problems as future work.

D. Functional Demonstration

The functional demonstration connects YCB objects and models to a real robot’s planning pipeline, while the protocol framework specifies reproducible benchmark design.

  • Functional demonstration: The HERB robot plans and executes a grasp of the YCB power drill using ROS and OpenRAVE.ROS handles communication among hardware and software, while OpenRAVE handles planning and collision checking.
  • Functional demonstration: CBiRRT, OMPL, and CHOMP plan and optimize motion trajectories inside the simulation environment.The environment supports chains of sequential actions and feedback from perception systems.
  • Functional demonstration: The prepared physical objects and meshes reduce setup time by removing the need to scan or model new objects.The resulting benchmark environments streamline experimental design for planning and manipulation algorithm development.
  • Protocol design: Direct comparison requires specifying what should be done with standardized objects, so the paper provides protocol and formatting guidelines.The framework is intended to support community-driven evolution through shared and updateable protocols.
  • Protocol design: Too few experimental constraints can undermine repeatability and performance assessment by introducing measurement discrepancies.The protocol guidelines address the challenge of balancing reliable specification with broad applicability.
  • Protocol guidelines: Protocols organize benchmark information into task, setup, robot or hardware or subject, procedure, and execution-constraint categories.Setup descriptions include target objects, initial poses, obstacles, and clutter; accurate initial poses reduce uncertainty from non-standard setups.

5) Execution Constraints:

The protocols define execution constraints and benchmark specifications so manipulation performance can be compared using task-specific procedures and scoring. Five example protocols cover pouring, gripper assessment, table setting, block pick-and-place, and peg insertion learning.

  • Execution Constraints: Protocol detail varies with the problem, so authors must tailor task descriptions and constraints to the intended application.The paper anticipates regular protocol improvement through research-community feedback.
  • Benchmark Guidelines: Benchmarks specify performance details and should use descriptive scoring with partial credit when possible instead of binary success/fail measures.Scoring should provide reasonable insight into system performance and intermediate task execution.
  • Example Protocols: Five example protocols assess pouring, gripper capabilities, table setting, block pick-and-place, and peg-insertion learning.Each protocol has a corresponding benchmark, and the paper reports experimental implementations.
  • Pitcher-Mug Protocol: The Pitcher-Mug protocol standardizes liquid transfer with pitcher and mug configurations and evaluates execution across ten scenarios.The task targets smooth, precise manipulation while avoiding spills.
  • Gripper Assessment Protocol: The Gripper Assessment protocol evaluates grasping objects with varied shapes and sizes, including round, flat, tool, and articulated objects.The benchmark uses standardized objects to compare gripper designs and reveal design strengths and weaknesses.

4) Block Pick and Place Protocol and Benchmark:

The block pick-and-place benchmark tests small-object dexterity through grasping and precise transfer, while related protocols extend evaluation to simulation, learning under perturbations, and baseline robotic manipulation. The examples connect standardized objects and models with hardware, planning, and control assessment.

  • Block Pick and Place: The block pick-and-place protocol evaluates grasping small objects, transferring them precisely, and the combined contribution of hardware and motion planning.Points reward task completion and manipulation precision.
  • Table Setting: The table-setting benchmark scores final object-pose accuracy and can run in simulation using the supplied object models and Gazebo URDF.The scenario uses a mug, fork, knife, spoon, bowl, and plate placed into predefined configurations.
  • Block Pick and Place: HERB failed the precise pick-and-place task because open-loop push grasping left the block pose relative to the gripper unknown after grasping.The grasp strategy robustly acquired blocks, but inaccurate post-grasp pose estimation prevented accurate placement.
  • Peg Insertion Learning Assessment: The Peg Insertion Learning Assessment benchmark compares learned controllers under random positional perturbations of the peg board.The paper applies it to a learned linear-Gaussian controller on a PR2 robot.
  • Box and Blocks Test: The Box and Blocks Test measures how many blocks a robot transfers between boxes within a fixed time and provides a robotic baseline using PR2.The baseline uses repeated random-location grasp attempts and reports ten two-minute experiments.
  • Benchmark Scope: The YCB set combines physical objects, high-resolution scans, and models to support reproducible manipulation benchmarks across real and simulated systems.The paper presents community-driven protocols as an evolving complement to the standardized object set.

APPENDIX A. PROTOCOL AND BENCHMARK TEMPLATE FOR MANIPULATION RESEARCH:

The appendix template structures manipulation protocols around setup, hardware or subject, prior information, execution, and benchmark specification. The Pitcher-Mug example instantiates these fields with standardized layouts, constraints, and a spill-sensitive score.

  • Protocol Identification: The template records reference information, version, authors, institution, and contact details for each protocol.These fields identify the protocol and its provenance.
  • Setup: The setup section describes the manipulation environment, object list, and initial object poses.These fields establish the physical or simulated starting conditions.
  • Robot/Hardware/Subject Description: The robot, hardware, or subject section specifies targeted agents, their initial state, and prior information available before execution.The template separates agent characteristics from the environmental setup.
  • Pitcher-Mug Protocol: The Pitcher-Mug setup uses the mug, pitcher, and white table cloth across ten printable initial configurations.Experiments avoid background clutter and may substitute rice for water.
  • Robot/Hardware Description: The protocol targets robots with onboard sensors and prohibits fixed environmental sensors.The robot vision sensor begins aligned with the marker on the printable layout.
  • Prior Information: The robot receives semantic facts about the pitcher and mug but not their object models.Known facts include the presence and relative size of handles and the pitcher opening’s location.
  • Procedure and Execution Constraints: The procedure repeats filling, transferring, resetting, and running the system for all ten scenarios using the same strategy.The execution sequence standardizes scenario-by-scenario evaluation.
  • Benchmark: The benchmark computes a transfer ratio from before-and-after mug and pitcher weights, then sums scenario points into an overall score.The scoring aims to reward spill-free pouring and requires scenario-level and overall scores plus system comments.

APPENDIX B.2 GRIPPER ASSESSMENT PROTOCOL AND BENCHMARK:

The Gripper Assessment benchmark tests grasping across object categories, spatial offsets, and grasp conditions. Its scoring records stability and successful handling, while the protocol standardizes the initial setup and allows category selection based on assessment scope.

  • Benchmark Definition: The protocol is presented as a gripper assessment benchmark within the broader YCB manipulation framework.Its formal protocol and benchmark documents define the assessment artifacts.
  • Task Description: The benchmark asks grippers to grasp objects with varied sizes and shapes one by one.The protocol is intended to assess gripper performance across representative object categories.
  • Object Categories: The object set includes round objects, flat objects, tools, articulated objects, and an additional planar object for applying position offsets.Users may select any combination of categories according to the performance assessment scope.
  • Setup and Initial Positions: The setup uses a printed grid and four standardized start points that vary planar placement or height relative to the table.Round objects use four initial positions, while flat objects use three positions per object.
  • Robot/Hardware Description: The protocol permits any robot or hardware and requires the same chosen grasping pose across specified position conditions.No prior information is provided to the robot.
  • Scoring: Scoring awards points for retaining objects and penalizes visible object motion during prescribed manipulation steps.Articulated-object success is based on whether any object part remains on the table after step five.
  • Submission: Submissions include the scoring table, gripper description, advantages, disadvantages, and reasons for failed grasps.The benchmark captures both performance and qualitative design information.

APPENDIX B.4 BLOCK PICK AND PLACE PROTOCOL AND BENCHMARK:

The appendix specifies block pick-and-place and peg-insertion protocols as standardized manipulation assessments, including setup, assumptions, procedures, and scoring guidance.

  • Block Pick and Place Protocol: The block pick-and-place protocol assesses robotic-manipulator dexterity by arranging blocks into a specified pattern.The setup uses a printed template marking the target final positions.
  • Block Pick and Place Protocol: The robot may begin in any configuration at least ten centimeters from the blocks and template.The template and block configuration may be specified beforehand or perceived with onboard sensors.
  • Block Pick and Place Protocol: The block procedure consists of placing the blocks as described and running the system, with no execution constraints.
  • Peg Insertion Learning Assessment Protocol: The peg-insertion learning protocol inserts a single peg into a hole using a learned policy and the YCB 9-peg hole test.
  • Peg Insertion Learning Assessment Protocol: Peg-board perturbations are evaluated at 5 mm and 10 mm, with four execution durations from 0.5 to 5 seconds.The procedure includes repeated trials at defined initial positions and durations.

APPENDIX C.1 PITCHER-MUG BENCHMARKING RESULTS

The benchmarking results illustrate quantitative evaluation of pitcher-mug manipulation and gripper performance across object categories, shapes, sizes, and grasp conditions.

  • Gripper Assessment: The Model T gripper achieved 122/447 total scored outcomes, including 43/187 round objects and 17/20 articulated objects.
  • Gripper Assessment: Model T performance was limited by object size, flexure-joint torsional forces, and finger gaps that prevented grasping some small objects.Interlacing fingers helped with articulated objects and some tools, but the gripper lacked sufficient stiffness for certain grasps.
  • Gripper Assessment: The Model T42 gripper achieved 379/447 total scored outcomes, including 84/96 flat objects and 91/114 tools.
  • Gripper Assessment: Model T42 handled a larger range of object sizes and shapes, using adaptive power grasps, precision grasps, and nails for tiny flat objects.Its main disadvantage was non-interlacing fingers, which reduced performance on some objects such as hammers and drillers.

APPENDIX C.3 BLOCK PICK AND PLACE BENCHMARKING RESULTS:

The block placement benchmark records performance under different execution times and board perturbations, with success decreasing as perturbation increases.

  • Block Pick and Place Benchmarking Results: Average benchmark scores were 4.67 blocks touching target areas, 0.17 entirely within target areas, 3.0 partially off-template blocks, and 1.83 for the final score.
  • Block Pick and Place Benchmarking Results: The HERB implementation used CBiRRT with TSR constraints and a specialized open-loop strategy for grasping small blocks with large hands.

APPENDIX C.4 PEG INSERTION LEARNING BENCHMARKING RESULTS:

The peg-insertion benchmark evaluates a learned controller across execution durations and randomized board perturbations using repeated trials and recorded pick outcomes.

  • Peg Insertion Learning Benchmarking Results: The benchmark varies execution duration among 0.5, 1, 3, and 5 seconds and evaluates no, 5 mm, and 10 mm perturbations.
  • Peg Insertion Learning Benchmarking Results: The learned system uses a linear-Gaussian controller whose state includes robot joint angles, angular velocities, and three end-effector points.
  • Peg Insertion Learning Benchmarking Results: The robot samples block-pick locations uniformly, retries when the gripper closes fully, and continues after detecting a grasp.
  • Peg Insertion Learning Benchmarking Results: The benchmark records misses, successive picks, successful trials, and transfer failures using explicitly defined trial-count measures.
Loading 1502.03143v1…