Source-linked AI summary
FolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects
Chenhuan Liu, Yi Xu, Feng Wu, Hanyang Wang, Wenxiao Kuai, Weihao Ding, Shan Wang, Yang Liu, Shuyong Gao, Wenqiang Zhang
TL;DR
FolDeX addresses limited real-robot evidence for reliable long-horizon deformable manipulation and efficient reuse of costly physical experience. It introduces a large garment-folding benchmark organized across four transfer axes, together with FoldChallenge for standardized external evaluation. In the reported baseline study, recovery data raises average success from 80.75% to 95.00%, while the benchmark’s scope remains constrained by evaluation cost and the absence of explicit tactile sensing.
Problem
Existing real-robot benchmarks provide limited coverage of long-horizon deformable manipulation, while physical-robot data are costly and difficult to reuse across settings.
Method
FolDeX organizes 2,000+ hours of real-robot data across recovery, task, scene, and embodiment axes, and pairs them with FoldChallenge’s standardized physical evaluation protocol.
Results
Adding recovery trajectories increases average success from 80.75% to 95.00% and average FoldScore from 75.59 to 82.53.
Takeaways & Limitations
FolDeX provides a unified testbed for studying heterogeneous real-robot data reuse and reliable long-horizon deformable manipulation.
Takeaways & Limitations
Evaluation is constrained by the time and hardware cost of long-horizon real-robot experiments, and the current benchmark lacks explicit tactile sensing.
Abstract
from arXiv · showhide
Embodied AI, including vision-language-action and world-action models, must operate reliably in the physical world. Yet methods that perform well in simulation can degrade substantially on real robots, especially in long-horizon deformable-object manipulation, where policies must track changing states and execute reliable multi-stage bimanual interactions. Existing real-robot benchmarks mainly focus on short-horizon rigid-object tasks and offer limited coverage of long-horizon deformable manipulation. We introduce FolDeX, a physical-world benchmark built entirely from real-robot data, with garment folding as its primary task. Since real-robot data collection is costly, FolDeX studies how heterogeneous physical experience can be reused efficiently. The benchmark is organized around four research axes: leveraging human intervention and recovery data collected during deployment; transferring data across tasks, including across garment categories and from rigid to deformable-object manipulation; reusing data across scenes with changes in lighting, background, and layout; and transferring data across robotic embodiments. FolDeX provides 2,000+ hours of real-robot data spanning 20+ tasks and 10+ embodiments. We also establish a fair real-robot evaluation platform for externally submitted policies, with standardized tasks, held-out physical objects, controlled initializations, and a unified execution protocol. The platform is publicly accessible at https://ai.midea.com/#/fold-challenge. We hope FolDeX will serve as a unified testbed for heterogeneous real-robot data reuse and reliable long-horizon deformable manipulation.
Introduction
FolDeX addresses the limited real-robot coverage of long-horizon deformable manipulation by benchmarking garment folding and systematic reuse of heterogeneous physical experience. It combines a large real-robot dataset with standardized, externally evaluated physical protocols across recovery, task, scene, and embodiment axes.
- Motivation: The benchmark targets the sim-to-real and rigid-to-deformable gaps because simulated or rigid-object performance may degrade on physical deformable manipulation.Real execution is affected by factors including friction, placement, calibration, latency, sensing noise, and reset variation.
- Motivation: Existing real-robot benchmarks mainly emphasize short-horizon rigid-object manipulation, leaving complete long-horizon deformable tasks insufficiently covered.Such tasks require tracking changing states, coordinating contact-rich bimanual actions, and preserving progress across multiple stages.
- FoldChallenge: FoldChallenge evaluates externally submitted policies using standardized tasks, held-out physical objects, controlled initializations, unified execution, component-wise metrics, and auditable rollout videos.The platform is designed to hold the target evaluation protocol fixed while varying training data or test distribution.
- Benchmark Design: FolDeX organizes data reuse around recovery, cross-task, cross-scene, and cross-embodiment axes with standardized partitions and controlled comparisons.These axes cover intervention data, garment and rigid-to-deformable transfer, scene changes, and transfer across platforms with different kinematics, sensors, and action spaces.
Related Work
Prior work spans generalist robot policies, simulation and physical-robot benchmarking, and deformable-object evaluation. These efforts motivate standardized assessment of long-horizon manipulation across tasks, objects, scenes, embodiments, and physical evaluation conditions.
- Generalist robot policies: Generalist robot-policy research includes vision-language-action models, open pretraining frameworks, heterogeneous co-training, and world-action or predictive policies.These approaches use visual observations, language, robot states, or future visual prediction as learning signals.
- Evaluation standardization: Standardized evaluation remains difficult because studies vary in robots, task definitions, initial-state distributions, reset procedures, and scoring rules.These differences make results obtained under separate evaluation setups difficult to compare.
- Real-robot benchmarking infrastructure: Simulation benchmarks support scalable evaluation of long-horizon control, transfer, and household activities, while physical platforms evaluate generalist policies on real-world tasks.The supplied examples include CALVIN, LIBERO, BEHAVIOR-1K, RoboChallenge, and Table30.
- Deformable-object and garment benchmarks: Deformable-object benchmarks cover simulated manipulation, garment tasks, multi-manipulator settings, and selected real-world evaluations across grasping, folding, and flinging.Examples include SoftGym, GarmentLab, RGBench, and the ICRA 2024 Cloth Competition.
- FolDeX: FolDeX organizes physical-world coverage around folding stages, garment categories, other deformable and rigid-object tasks, robotic embodiments, scenes, and standardized external evaluation.Its described coverage includes semantic stages, intermediate and failure states, and final-state quality under a physical evaluation protocol.
FolDeX
FolDeX is a real-robot benchmark for long-horizon deformable manipulation that emphasizes efficient reuse of heterogeneous experience across recovery, tasks, scenes, and embodiments. It combines garment-folding episodes, standardized evaluation, and metrics that capture success, final-state quality, and execution efficiency.
- FolDeX: FolDeX provides a physical-world benchmark for long-horizon deformable manipulation, centered on garment folding and including other deformable and a few rigid-object tasks.Its episodes involve continuous deformation, multi-stage transitions, bimanual coordination, partial observability, and accumulating errors.
- FolDeX: The benchmark organizes heterogeneous real-robot data around recovery, cross-task, cross-scene, and cross-embodiment transfer.These tracks study intervention and recovery data, rigid-to-deformable transfer, adaptation to changed scenes, and sharing knowledge across robot platforms.
- FolDeX: A complete folding episode runs from retrieving a randomly configured garment through flattening and multiple folds to final placement or stacking.The policy receives synchronized wrist and workspace RGB observations, language instructions, and proprioceptive state, then predicts future end-effector trajectories.
- FolDeX: FolDeX contains 2,000+ hours of experience spanning 20+ manipulation tasks and 10+ robotic embodiments, with a common schema for observations, proprioception, actions, instructions, and metadata.The platforms include single-arm, dual-arm, mobile, and humanoid systems with differing kinematics, workspaces, cameras, interfaces, and action representations.
- FolDeX: Table 1 distinguishes benchmarks with full dependent long-horizon deformable episodes from suites with limited folding coverage or selected long-horizon tasks.The comparison emphasizes whether a benchmark provides a complete physical evaluation protocol rather than merely including an isolated folding task.
- Evaluation Metrics: Evaluation reports task success, completion time, final-state neatness, and FoldScore under shared physical tasks, instances, initialization, duration, and execution protocols.Final-state neatness ranges from 0 for largely unfolded garments to 5 for compact, flat, stable, well-aligned results; FoldScore combines normalized quality, success, and time.
Experiments
FolDeX experiments compare multi-task and task-specific folding policies, recovery-data augmentation, and preliminary transfer across tasks, scenes, and embodiments. Recovery data substantially improves performance, while naive cross-task and cross-embodiment reuse remains unreliable.
- Multi-Task and Task-Specific Baselines: RTC task-specific policies achieve the highest average success rate at 82.50%, while RTC multi-task policies achieve the highest average FoldScore at 75.59.The differing leaders show that binary success and physical folding quality capture different aspects of performance.
- Multi-Task and Task-Specific Baselines: Within multi-task training, RTC improves average success rate from 72.50% to 80.75% and average FoldScore from 72.30 to 75.59.The improvement is particularly visible on Pants and Towel, which require longer-horizon state tracking and substantial deformable manipulation.
- Multi-Task and Task-Specific Baselines: RTC multi-task training achieves 76.67% success on Towel versus 50.00% for task-specific training, although transfer benefits are not uniform across garments.Task-specific training remains strong on Shirt, Skirt, and Pants.
- Recovery-Data Utilization: Recovery data raises average success rate from 80.75% to 95.00% and average FoldScore from 75.59 to 82.53.Pants and Towel each reach 100.00% success after recovery-data augmentation.
- Preliminary Observations on Other Transfer Axes: Naive heterogeneous task and embodiment reuse generally fails to preserve reliable complete-task performance because of action interference and catastrophic forgetting.Cross-scene changes in lighting, background, and moderate workspace layout are less disruptive in these preliminary evaluations.
Discussion and Limitations
The comparison is limited by the time and hardware cost of long-horizon real-robot evaluation, and the benchmark currently relies primarily on visual observations and proprioception.
- Discussion and Limitations: Long-horizon real-robot evaluation is constrained by time and hardware cost, while current benchmark sensing lacks explicit tactile feedback.Future versions are intended to incorporate tactile feedback for contact states and local deformations.
Conclusion
FolDeX introduces a real-robot benchmark for long-horizon deformable manipulation alongside FoldChallenge, a fair evaluation platform. It organizes research around reusing heterogeneous robot data across recovery, task, scene, and embodiment axes.
- Conclusion: FolDeX is presented as the first real-robot benchmark focused on long-horizon deformable manipulation, together with the FoldChallenge evaluation platform.The benchmark structures data reuse across recovery, task, scene, and embodiment dimensions.