Source-linked AI summary
Peg-in-Bench: A Modular Benchmark for High-Precision Robotic Insertion
Yosel Delgado, José G. Buenaventura-Carreón, Floris Erich, Roman Mykhailyshyn, Tomohiro Motoda, Koshi Makihara, Yukiyasu Domae
TL;DR
High-precision insertion lacks physical benchmarks that systematically test transfer beyond fixed peg-in-hole configurations. Peg-in-Bench introduces a reconfigurable, fully 3D-printable component system and scenario generator, enabling reproducible variation in geometry, tolerance, spatial layout, and assembly structure while controlling other physical conditions. The benchmark therefore supports evaluation of spatial, geometric, and task-compositional generalization across unseen insertion scenarios.
Problem
Existing peg-in-hole evaluations commonly rely on fixed task configurations, limiting assessment of insertion-skill transfer across varied physical task configurations.
Method
Peg-in-Bench uses modular, reconfigurable, fully 3D-printable components and a scenario generation tool to create standardized insertion and assembly tasks.
Results
The benchmark supports systematic evaluation of spatial, geometric, and task-compositional generalization across reproducible unseen task configurations.
Takeaways & Limitations
A common set of physical components can generate diverse insertion and assembly scenarios while maintaining controlled physical conditions for standardized evaluation.
Abstract
from arXiv · showhide
High-precision insertion remains a fundamental challenge in robotic manipulation due to the strict alignment requirements and contact-rich interactions involved. Although peg-in-hole tasks are widely used for evaluation, existing bench- marks often rely on fixed task configurations, limiting their ability to assess robustness and generalization across different insertion scenarios. This paper introduces a reconfigurable peg-in-hole benchmark designed to evaluate task generalization in high-precision insertion. The benchmark consists of a set of fully 3D-printable modular components, including multiple peg geometries, tolerance levels, and configurable base structures that can be combined to generate a large variety of insertion and assembly tasks. By varying object layouts, orientations, and task structures while maintaining controlled physical conditions, the benchmark enables systematic evaluation of adaptation to unseen scenarios. To support reproducibility, we additionally provide a scenario generation tool capable of producing standardized task configurations and machine-readable task descriptions. The scenario generation tool and the STL files of the benchmark pieces are available through the project repository: https://github.com/aistairc/peg-in-bench.
I. INTRODUCTION
Existing peg-in-hole benchmarks support standardized evaluation but often use fixed configurations, limiting assessment of transfer across task variations. Peg-in-Bench addresses this gap with a modular, reconfigurable, fully 3D-printable benchmark for controlled high-precision insertion generalization.
- Peg-in-hole insertion evaluates perception, motion planning, and precision control under geometric constraints in contact-rich robotic manipulation.
- Existing benchmarks improve standardization and reproducibility but generally rely on fixed layouts, orientations, and assembly structures.
- Peg-in-Bench varies peg geometry, insertion tolerance, target position, target orientation, and assembly structure through reconfigurable components.
- The benchmark generates many reproducible physical scenarios for evaluating spatial, geometric, and task-compositional generalization.
- Its design targets both classical compliant control and modern learning-based manipulation approaches.
B. Generalization-Oriented Manipulation Benchmarks
Generalization-oriented benchmarks broaden evaluation across objects, environments, and procedures, but dedicated physical evaluation of high-precision insertion under configurable geometry and task reconfiguration remains limited.
- FMB, FurnitureBench, and Open X-Embodiment evaluate object, task, long-horizon, or cross-platform generalization.
- Recent vision-language-action systems demonstrate transfer across robotic tasks using large-scale pretraining.
- Traditional insertion benchmarks standardize precision and dexterity tasks, while newer resources add broader insertion operations or quality-oriented metrics.
- Existing insertion benchmarks generally use fixed configurations, limiting evaluation of transfer to unseen layouts, orientations, and assembly structures.
- Learning-based insertion methods report generalization across object geometries and environmental conditions but are often evaluated on custom setups with limited task collections.
- These limitations make cross-study comparison difficult and motivate standardized, configurable physical benchmarks.
E. Comparison with Existing Benchmarks
Peg-in-Bench complements existing benchmarks by replacing fixed task collections with reusable components that generate controlled, diverse scenarios for task-level generalization.
- Existing benchmarks emphasize standardized performance, learning and object generalization, or insertion-specific algorithm evaluation.
- Peg-in-Bench varies peg geometry, tolerance, target position, target orientation, and assembly structure using reusable components.
- The same physical pieces generate many reproducible scenarios for testing transfer across previously unseen task configurations.
- The framework targets task-level generalization rather than performance within a single predefined benchmark setting.
- Modularity supports isolated skill evaluation and composition into more complex multi-stage assembly scenarios.
A. Design Objectives and Generalization Dimensions
Peg-in-Bench evaluates spatial, geometric, and task-compositional generalization through modular changes to positions, orientations, geometries, tolerances, and assembly structures. The first version holds several environmental and contact factors approximately constant.
- The benchmark’s primary objective is evaluating task-level generalization in high-precision peg-in-hole insertion.
- Spatial generalization changes target positions and orientations using modular bases and rotatable hole pieces.
- Geometric generalization varies peg shapes and three tolerance levels to study differing alignment constraints, contact dynamics, and precision requirements.
- Task-compositional generalization constructs larger structures requiring sequential composition of multiple insertion actions.
- Illumination, camera placement, sensor noise, friction, material properties, and surface finish are approximately held constant in the first version.
B. Benchmark components
Peg-in-Bench uses modular, printable components spanning multiple peg geometries and tolerance levels for controlled high-precision insertion evaluation.
- The benchmark comprises pegs, shaped-hole pieces, and bases as its three main component families.
- Pegs and hole pieces are resin-printed at 0.05 mm resolution, while bases and non-precision components can use standard filament-based printers.
- Five peg geometries—circular, rectangular, hexagonal, triangular, and L-shaped—span symmetric, asymmetric, and keyed insertion challenges.
- The benchmark uses 0.1 mm, 1 mm, and 3 mm tolerances, yielding 15 shaped-hole pieces across the five peg shapes.
- Each shaped-hole piece is an octagonal prism with eight distinct orientations and an off-center hole, making each orientation a unique insertion configuration.
- Embedded magnets retain hole pieces during insertion while allowing rapid replacement and release under excessive lateral forces.
3) Bases:
The bases provide reconfigurable holders and interlocking structures that support repositioning hole pieces, multi-base assemblies, and insertion tasks beyond a single peg.
- Base pieces hold shaped-hole components and allow them to be replaced, rotated, and repositioned for new tasks.
- Each base is a cube with interlocking side joints and holes, enabling multiple bases to connect into larger structures.
- The cross-base model includes an octagonal space for shaped-hole pieces and a bottom space for embedding a magnet.
- Four base variants—corner, empty, line, and cross—provide different joint and hole arrangements for structural configurations.
- Base joints and holes use 0.25 mm tolerance, creating strong yet manipulable connections for multi-object insertion.
- The benchmark supports single-peg demonstrations and scenario variations across shapes, tolerances, workspace arrangements, and manipulation objectives.
A. Single-peg insertion and shape generalization
Single-peg tasks evaluate precision across peg geometries and increasingly restrictive clearances, while randomized orientations add geometric alignment and spatial-generalization demands.
- The benchmark requires inserting five peg shapes into corresponding holes at a fixed tolerance to compare geometry-dependent performance.
- Unlike comparable insertion tasks, the same benchmark task varies peg geometries and spatial configurations to evaluate geometric and spatial generalization.
- Progressively smaller tolerances test insertion performance under increasingly restrictive clearance conditions.
- Randomizing hole orientation independently across trials jointly tests geometric alignment accuracy and handling of restrictive contact conditions.
- A Peg-holder provides a consistent initial peg pose, while fixing pieces secure bases to the workspace during interaction.
IV. BENCHMARK TASKS AND EVALUATION SCENARIOS
Peg-in-Bench generates related task families by varying geometry, tolerance, position, orientation, and assembly structure, extending evaluation from single insertions to compositional assembly.
- IV. BENCHMARK TASKS AND EVALUATION SCENARIOS: The benchmark generates families of related tasks that systematically vary geometry, tolerance, spatial layout, and assembly structure.
- IV. BENCHMARK TASKS AND EVALUATION SCENARIOS: Tolerance changes can be combined with target-position and orientation variations to assess insertion accuracy and generalization across configurations.
- IV. BENCHMARK TASKS AND EVALUATION SCENARIOS: Combining insertion operations into larger assembly objectives supports evaluation of skill composition, long-horizon manipulation, and task sequencing.
- IV. BENCHMARK TASKS AND EVALUATION SCENARIOS: The scenario-generation tool can produce single-peg tasks with 3 mm tolerance holes and random positions and orientations.
- C. Modular assembly tasks: Base connections form linear or irregular structures, with each connection evaluating dexterous manipulation, spatial reasoning, and sequential planning.
- C. Modular assembly tasks: The modular structure generates new assembly scenarios from the same components, supporting task-compositional generalization.
V. SCENARIO GENERATOR TOOL
The scenario generation tool automates reproducible creation of benchmark configurations and outputs both visual and machine-readable task descriptions. These configurations support standardized evaluation and controlled novelty splits for studying manipulation generalization.
- The tool automatically creates benchmark configurations by randomizing hole-piece positions and orientations while respecting physical constraints.
- Users specify tolerance level, task type, and a random seed that uniquely identifies each generated scenario.
- Each generated scenario includes a visual representation and a machine-readable JSON description containing the complete task specification.
- Generated configuration files support replication across laboratories and controlled training, validation, and testing splits with varying novelty.
VI. CONCLUSIONS AND DISCUSSIONS
Peg-in-Bench provides a reconfigurable framework for generating diverse insertion scenarios from reusable components while preserving reproducible physical conditions. Its task-level focus enables controlled generalization studies, but constrains environmental robustness and may require alternative mounting solutions.
- Peg-in-Bench generates diverse insertion scenarios from a common set of reusable components rather than introducing a new peg-in-hole task.
- The same scenario can be replicated with benchmark pieces in a real-world environment for data collection.
- Controlled variation of geometry, tolerance, spatial layout, and assembly structure supports task-level generalization under reproducible physical conditions.
- Keeping material, friction, illumination, and sensing conditions approximately constant isolates task-configuration effects on insertion performance.
- This controlled design does not evaluate robustness to perception degradation, manufacturing variability, or changing contact properties.
- The current fixation system is designed around a Vention aluminum profile table and may require alternative mounting solutions elsewhere.