Source-linked AI summary

LIBERO-X: Robustness Litmus for Vision-Language-Action Models

Guodong Wang, Chenkai Zhang, Qingjie Liu, Jinjin Zhang, Jiancheng Cai, Junjie Liu, Xinmin Liu

arXiv:2602.06556v1cs.CVcs.AIcs.RO

TL;DR

Existing VLA benchmarks provide limited evidence under realistic distribution shifts, motivating a benchmark that jointly improves evaluation protocols and training-data diversity. LIBERO-X introduces hierarchical multi-level evaluation and a high-diversity teleoperated dataset, and experiments show performance degrades as task and scene complexity increase. The benchmark is intended to provide a more faithful basis for assessing VLA generalization and robustness.

  • Problem

    Existing benchmarks can provide overly optimistic or misleading assessments because their perturbations, train–test scenarios, and task distributions insufficiently capture realistic multi-source shifts.

  • Method

    LIBERO-X combines a progressively challenging multi-level evaluation protocol with a high-diversity human-teleoperation dataset containing multiple fine-grained tasks per scene.

  • Results

    Performance degrades notably as task and scene complexity increase, revealing limitations in scene comprehension and instruction grounding across representative VLA models.

  • Takeaways & Limitations

    LIBERO-X provides a more faithful and rigorous framework for assessing VLA generalization and robustness under multi-dimensional distribution shifts.

Abstract

from arXiv · show

Reliable benchmarking is critical for advancing Vision-Language-Action (VLA) models, as it reveals their generalization, robustness, and alignment of perception with language-driven manipulation tasks. However, existing benchmarks often provide limited or misleading assessments due to insufficient evaluation protocols that inadequately capture real-world distribution shifts. This work systematically rethinks VLA benchmarking from both evaluation and data perspectives, introducing LIBERO-X, a benchmark featuring: 1) A hierarchical evaluation protocol with progressive difficulty levels targeting three core capabilities: spatial generalization, object recognition, and task instruction understanding. This design enables fine-grained analysis of performance degradation under increasing environmental and task complexity; 2) A high-diversity training dataset collected via human teleoperation, where each scene supports multiple fine-grained manipulation objectives to bridge the train-evaluation distribution gap. Experiments with representative VLA models reveal significant performance drops under cumulative perturbations, exposing persistent limitations in scene comprehension and instruction grounding. By integrating hierarchical evaluation with diverse training data, LIBERO-X offers a more reliable foundation for assessing and advancing VLA development.

I. INTRODUCTION

Existing VLA benchmarks can give overly optimistic robustness assessments because they use limited perturbations, narrow train–test gaps, and insufficiently diverse data. LIBERO-X addresses these issues with hierarchical multi-dimensional evaluation and diverse teleoperated training data, revealing degradation as complexity increases.

  • Existing evaluations can produce overly optimistic or misleading conclusions because protocols and data distributions inadequately reflect real-world shifts.
  • Prior extensions model perturbation dimensions independently and lack a systematic easy-to-hard progression for measuring degradation under compound shifts.
  • LIBERO-X evaluates spatial generalization, object recognition, and instruction understanding through a progressively challenging multi-level test suite.
  • Its training dataset uses multi-task scenes with diverse trajectories and variations in object attributes, spatial relations, and language instructions.
  • 2,520 demonstrations across 600 tasks and 100 scenes increase scene diversity and task granularity through human teleoperation.
  • Experiments reveal notable performance degradation as task and scene complexity increase, exposing limitations in scene comprehension and instruction grounding.

II. RELATED WORK

Prior robotic-manipulation benchmarks established reproducible evaluation but often test isolated shifts, use limited training diversity, or rely on off-the-shelf checkpoints. LIBERO-X combines diverse teleoperated training data with progressively harder, multi-dimensional evaluation.

  • Simulation benchmarks provide experimental convenience and reproducibility for assessing robotic manipulation generalization and robustness.
  • High nominal success rates can decline sharply under minor perturbations, suggesting reliance on memorized training data rather than generalized skills.
  • Existing evaluations often test isolated visual or positional shifts rather than coupled multi-factor shifts encountered in real environments.
  • Training datasets often lack variation in scenes, tasks, and trajectories, constraining the generalization capacity of resulting policies.
  • LIBERO-X integrates teleoperation-collected training data with evaluation that increases difficulty through spatial perturbations, object substitutions, and instruction rewrites.
  • The dataset contains 2,520 demonstration trajectories spanning 600 tasks and 100 scenes, with finer-grained task extensions and multiple manipulation objectives.

1) Training Trajectory Collection:

LIBERO-X constructs its training data from human-operated demonstrations in diverse simulated scenes containing varied object types, tasks, and attribute-conditioned interactions.

  • Trajectories are collected through VR teleoperation in MuJoCo, recording 6D end-effector pose, gripper status, Cartesian increments, and gripper commands.
  • The dataset comprises 100 scenarios and 600 distinct tasks spanning rigid and articulated objects, multi-object interactions, and attribute-conditioned manipulation.
  • The test framework combines progressively difficult levels with fine-grained multi-label task annotations for quantitative analysis across perturbations and environments.

1) Multi-level Evaluation:

LIBERO-X uses a two-dimensional evaluation design: progressively harder scene perturbations form the vertical axis, while hierarchical multi-label annotations provide fine-grained task analysis.

  • Each evaluation level superimposes additional perturbations across spatial reasoning, object recognition, and instruction understanding to assess cumulative robustness.
  • Levels 1 and 2 progress from minor positional changes to larger-scale spatial randomization.
  • Level 3 reconstructs scene topology by altering object relationships, swapping targets with distractors, and adding irrelevant objects.
  • Level 4 varies texture, color, and size while introducing unseen categories and visually similar confounders.
  • Level 5 reformulates instructions through synonym replacement, compression, word-order changes, voice conversion, and redundant descriptions while preserving semantics.
  • Multi-label categories cover interaction type, subtask count, spatial relation, and object attributes for detailed capability analysis.

IV. EXPERIMENTS

LIBERO-X evaluates five VLA models under progressively compounded spatial, object, and semantic changes, exposing substantial degradation where standard LIBERO results often saturate. The evaluation highlights weaknesses in spatial extrapolation, topology changes, novel-object grounding, and instruction variation.

  • Five representative VLA models were evaluated under the multi-level framework using their official configurations.
  • 39.4% average success at Level 1 and 65.2% for π0.5 contrast with often ∼90% performance reported on LIBERO.
  • 9.8% average success decline from Level 1 to Level 2 indicates difficulty with larger spatial shifts beyond training configurations.
  • 12.7% degradation at Level 3 is the largest decline, revealing limited structural invariance under topological rearrangements.
  • Diversified training preserves non-zero performance on out-of-distribution tasks, but unseen-object success remains below confounding-object success.
  • Instruction rephrasing causes a 3.26% Level 4-to-Level 5 drop, while voice conversion is the least disruptive variation.

C. Multi-label Evaluation

The multi-label evaluation reveals synchronized performance degradation across models as task complexity increases, with long-horizon tasks producing especially severe failures. Time-limit analysis shows that modest temporal slack can buffer execution imperfections, while additional slack eventually saturates.

  • Across levels, models retain a stable relative performance hierarchy while their absolute success rates decline with increasing complexity.
  • Tasks with three or more subtasks drive nearly all models’ success rates toward zero, exposing difficulty with error accumulation and long-horizon execution.
  • Time limits are evaluated relative to the average duration of human operation, using multipliers from 0.8 to 1.5.
  • Stringent 0.8× limits severely penalize performance, whereas extending the limit improves tolerance to inference and control errors.

APPENDIX

The appendix details LIBERO-X’s fine-grained, multi-label evaluation system and its supporting analyses. Labels cover interaction type, subtask count, spatial relations, and object attributes, enabling performance analysis across task dimensions.

  • The supplementary materials include detailed multi-label results, time-limit analysis, model details, scene design, visual-attribute cases, and semantic-reformulation cases.
  • The fine-grained system assigns labels across Interaction Type, Subtask Count, Spatial Relation Type, and Object Attribute Recognition.
  • Interaction Type: Interaction labels distinguish placement, state change, and composite actions, capturing basic manipulation and coordinated multi-step behavior.
  • Subtask Count: Subtask Count distinguishes single-action, two-action, and three-or-more-action sequences to assess increasingly long-horizon planning.
  • Spatial Relation Type: Spatial labels cover left–right relations, ordinal regions, and relative distance, testing directional, hierarchical, and proximity reasoning.
  • Object Attribute Recognition: Object Attribute Recognition evaluates grounding instructions to color, size, pattern, and category properties amid distractors and visual complexity.

B. Further Analysis of the Impact of Time Limits

Time-limit analysis finds that a modest temporal buffer improves performance across all evaluation levels, but gains diminish after the main buffer is added. Beyond 1.5× human reference time, performance plateaus because perception and reasoning remain limiting factors.

  • Across Levels 1–5, success rates increase with extended time limits, and 1.1× human reference time produces the largest gain.
  • The 1.1× threshold is adopted as the default because its temporal buffer compensates for execution noise and enables error recovery.
  • Between 1.1× and 1.5×, improvements continue but diminish as added time permits trial-and-error rather than robust long-horizon planning.
  • For π0.5 on Three or More Actions, success rises 12.4% from 1.0× to 1.1× but only 1.9% from 1.1× to 1.5×.
  • Beyond 1.5×, performance plateaus because perception and reasoning bottlenecks, rather than time, constrain complex-task performance.

C. Model Details

The study fine-tunes diverse VLA architectures on LIBERO-X and evaluates them with fine-grained tests across progressively difficult levels and time limits.

  • The evaluation compares diverse VLA architectures initialized from official pre-trained checkpoints and fine-tuned on LIBERO-X.
  • OpenVLA-OFT combines parallel decoding, action chunking, continuous actions, and FiLM-based language grounding.
  • π0 generates continuous actions with flow matching, while π0.5 adds broader semantic generalization through internet-scale data and hierarchical subtask prediction.
  • The benchmark reports fine-grained evaluation results and time-limit thresholds across Levels 1–5.
  • π0 and π0.5 are fine-tuned on LIBERO-X using resized 224×224 wrist and third-person images, with distinct adaptation strategies and action-chunk configurations.
  • X-VLA uses learnable soft prompts, multimodal Transformer fusion, and flow matching for cross-embodiment adaptation.
  • GR00T N1.5 combines a frozen VLM, an alignment adapter, and a Diffusion Transformer, with FLARE aligning predicted future states to ground truth.

D. Scene Design Details

LIBERO-X expands scene assets and evaluation predicates to make containment, posture, and object variation more precise and diverse.

  • The scene framework expands the original LIBERO system by optimizing core features and enlarging the 3D asset library.
  • ExactIn replaces imprecise In detection with a vertical-displacement threshold that confirms an object is fully inserted and resting at the container bottom.
  • Upright and side-on predicates classify book posture using angles between the object up-vector and surface normal.Upright placement requires an angle below approximately 37°, whereas side-on placement is within 17° of perpendicular.
  • These posture predicates support more precise recognition and reliable execution in placement tasks requiring posture differentiation.
  • Novel objects increase color, texture, and size diversity, while functional zones such as shelves and drawers provide additional scene and task variation.
  • Training and test datasets track Predicate, Object, and Container frequencies across scenes and tasks.

E. Visual Attribute Variation Cases Analysis

Level 4 tests visual attribute variation using confounding and unseen objects, revealing failures in target selection, instruction following, and multi-step execution.

  • Level 4 evaluates color, size, and texture variation for confounding objects and objects absent from training, using π0.5 test cases.
  • Confounding objects can trigger an initial grasp of a visually similar but incorrect target, disrupting the scene and target position.
  • The model sometimes grasps irrelevant objects despite clear instructions, reflecting limitations in instruction interpretation and target alignment.
  • Multi-step failures arise when trajectory switching disrupts transitions or environmental interference makes later targets difficult to grasp or identify.

F. Semantic-Equivalent Reformulation Cases

Level 5 evaluates semantic-equivalent task reformulations that preserve meaning while changing wording and sentence structure.

  • Synonym replacement substitutes instruction phrases using a mapping table while maintaining the original meaning.
  • Word compression removes redundant elements such as auxiliary verbs and articles while retaining core semantics.
  • Word order adjustment changes sentence structure without altering the intended task meaning.
  • The benchmark also visualizes predicate, object, and container relationships alongside confounding- and unseen-object test samples.
Loading 2602.06556v1…