Source-linked AI summary

DexHoldem: Playing Texas Hold'em with Dexterous Embodied System

Feng Chen, Tianzhe Chu, Li Sun, Pei Zhou, Zhuxiu Xu, Shenghua Gao, Yuexiang Zhai, Yanchao Yang, Yi Ma

arXiv:2605.18727v1cs.ROcs.AI

TL;DR

Embodied systems need evaluation that combines changing-scene perception, context-appropriate action, dexterous execution, and scene preservation, rather than isolated primitives. DexHoldem introduces a real-world ShadowHand benchmark with physical policy, agentic perception, and closed-loop evaluations. Its results show strong but imperfect primitive performance, a gap between field-wise and complete state recovery, and accumulating operational failures in deployment.

  • Problem

    Existing embodied-agent and dexterous-manipulation benchmarks separately underrepresent precise real-world dexterous execution or instruction-conditioned sequential state reasoning.

  • Method

    DexHoldem combines 1,470 demonstrations across 14 primitives with standardized physical policy evaluation, structured agentic perception assessment, and closed-loop embodied-agent case studies.

  • Results

    The benchmark reveals a cross-task gap: π0.5 reaches 61.2% task completion and 47.5% scene-preserving success, while perception reaches 34.3% strict problem-level accuracy versus 66.8% field-wise accuracy.

  • Takeaways & Limitations

    DexHoldem evaluates dexterous tabletop execution, agentic state recovery, and embodied decision routing within one shared physical setting.

  • Takeaways & Limitations

    The benchmark uses a fixed ShadowHand–UR setup and 1,470 demonstrations, so it does not establish cross-embodiment transfer or broad policy scaling behavior.

Abstract

from arXiv · show

Evaluating embodied systems on real dexterous hardware requires more than isolated primitive skills: an agent must perceive a changing tabletop scene, choose a context-appropriate action, execute it with a dexterous hand, and leave the scene usable for later decisions. We introduce DexHoldem, a real-world system-level benchmark built around Texas Hold'em dexterous manipulation with a ShadowHand. DexHoldem provides 1,470 teleoperated demonstrations across 14 Texas Hold'em manipulation primitives, a standardized physical policy benchmark, and an agentic perception benchmark that tests whether agents can recover the structured game state needed for embodied decision making. On primitive execution, $π_{0.5}$ obtains the highest task completion rate ($61.2\%$), while $π_{0.5}$ and $π_0$ tie on scene-preserving success rate ($47.5\%$). On agentic perception, Opus 4.7 obtains the best strict problem-level accuracy ($34.3\%$), while GPT 5.5 obtains the best average field-wise accuracy ($66.8\%$), exposing a gap between isolated visual sub-capabilities and complete routing-relevant state recovery. Finally, we instantiate the full embodied-agent loop in three case studies, where waiting, recovery dispatches, human-help requests, and repeated primitive execution reveal how perception and policy errors accumulate during closed-loop deployment. DexHoldem therefore evaluates dexterous tabletop execution, agentic perception, and embodied decision routing in a shared physical setting. Project page: https://dexholdem.github.io/Dexholdem/.

1 Introduction

DexHoldem addresses the incomplete evaluation of embodied systems by combining instruction grounding, sequential tabletop-state tracking, and fine-grained dexterous execution in a real Texas Hold’em setting. It contributes standardized benchmarks spanning physical policies, agentic perception, and closed-loop system behavior.

  • Existing embodied-agent benchmarks often use simulation, coarse actions, or gripper-centric manipulation, limiting evidence for precise physical multi-finger control.
  • Dexterous benchmarks commonly assess isolated low-level skills without instruction-conditioned visual grounding, sequential state awareness, or progress verification.
  • Texas Hold’em provides semantically structured card-and-chip targets whose tabletop state changes after each action, requiring grounding, action selection, dexterous execution, and recovery.
  • DexHoldem provides 1,470 real-world demonstrations across 14 atomic card and chip primitives using a ShadowHand and jointly evaluates instruction grounding with fine-grained dexterous control.
  • Its protocol standardizes task descriptions, initial-state randomization, and objective post-conditions including successful manipulation and scene preservation.
  • π0.5 reaches 61.2% task completion, while π0.5 and π0 tie at 47.5% scene-preserving success; perception reaches 34.3% strict full-state accuracy versus 66.8% best field-wise average.

2 Related Work

Related work spans general robot-manipulation benchmarks, dexterous datasets, simulation frameworks, and embodied agents. DexHoldem sits at their intersection by evaluating real-world dexterous manipulation together with embodied perception and decision behavior.

  • Dexterous manipulation research covers contact-rich behaviors such as grasping, in-hand reorientation, articulated-object operation, and bimanual coordination.
  • Robot benchmarks including RLBench, Meta-World, CALVIN, LIBERO, and related environments evaluate generalization, language-conditioned manipulation, transfer, household tasks, or coordination.
  • Simulation frameworks improve scalability and standardization, while dexterous datasets provide testbeds for multi-finger control, grasping, articulation, handover, and large-scale learning.
  • Embodied-agent work uses multimodal models for perception, reasoning, and high-level action selection across simulated or real environments, including vision-language-action policies.

3 DexHoldem System Design

DexHoldem couples an embodied agent that parses and routes tabletop state with a multi-task dexterous policy that executes 14 atomic primitives. Its evaluation separates policy execution, agentic perception, and closed-loop trajectory behavior while retaining physical scene constraints.

  • System architecture: The system couples structured game-state memory and activity selection with a multi-task policy conditioned on visual observations, proprioception, and the requested primitive.
  • Policy benchmark: The policy benchmark contains 14 language-instructed primitives and 1,470 teleoperated demonstrations, with 105 demonstrations per primitive.
  • Policy benchmark: Policies receive synchronized top-down, third-person, and wrist-camera views, proprioceptive state, and a task condition, then output short-horizon joint-position targets.
  • Policy benchmark: Physical rollouts distinguish scene-preserving success, disruptive completion, and task failure to measure both nominal completion and preservation of a usable tabletop.
  • Agentic perception: Agentic perception parses one current image into eight structured challenges and scores each problem by exact match over the challenges applicable to that state.
  • System-level evaluation: Closed-loop execution routes captured states through workflow gates, dispatching dexterous primitives only when physical motion is required and retrying recoverable failures.
  • System-level evaluation: Operational counters track captured states, agent and dexterous-policy dispatches, waiting, human-help requests, and recovery dispatches to expose accumulated trajectory complexity.

4 Experiments

Experiments evaluate physical primitive execution, RDT data scaling, agentic perception, and closed-loop system behavior. Results show that scene preservation remains harder than nominal completion, pretraining offers limited low-data efficiency, isolated perception does not reliably compose into complete state recovery, and errors accumulate during rollout.

  • Policy Model Results: 61.2% task completion rate is highest for π0.5, while π0.5 and π0 tie at 47.5% scene-preserving success rate.The evaluation covers 80 physical trials spanning all 14 primitives; disruptive completions count toward task completion but not scene-preserving success.
  • Policy Model Results: 30.0% scene-preserving success rate and 46.2% task completion rate place RDT in an intermediate aggregate performance tier.DP (DINO) is the strongest task-specific imitation baseline but trails the best pretrained policies by more than 20 percentage points in scene-preserving success.
  • Policy Model Results: 47.5% to 61.2% is π0.5’s increase from scene-preserving success rate to task completion rate when disruptive completions are included.The gap reflects rollouts that achieve the local objective while disturbing surrounding cards or chips enough to block continuation.
  • RDT Fine-Tuning Data Scaling Study: 11.3% is the largest validation-loss reduction from pretrained initialization at 100% data, while random and pretrained initializations follow similar data-scaling trends.The reduction is 1.2%, 9.0%, and 10.7% at 10%, 20%, and 50% data, respectively; the probe does not support a strong low-data-efficiency interpretation.
  • Agentic Perception: 34.3% strict Overall accuracy is best for Opus 4.7, whereas GPT 5.5 achieves the best 66.8% Avg field-wise accuracy.Overall requires every applicable field to be correct, while Avg is the unweighted mean across eight sub-capability columns.
  • Agentic Perception: 45.8% and 43.8% are the peak accuracies for current bet chips and opponent chip inventory, the two weakest average sub-capabilities.These fields require exact denomination-level chip dictionaries, and missed opponent bet-chip changes can route an embodied system to the wait branch.
  • System-Level Evaluation: About a third of the 23 states in trajectory (iii) are spent in the wait branch, alongside one recovery retry and no human-help request.The case study contains eight high-level decisions and terminates after the second card reveal; longer hands add repeated waiting, verification, recovery opportunities, and primitive dispatches.

5 Limitations

DexHoldem is intentionally limited to a controlled, fixed hardware and tabletop configuration, with a small dataset and substantial physical-evaluation costs. These boundaries constrain claims about transfer, broad dexterity, scaling, and simulation-based replacement of real evaluation.

  • Scope: The benchmark uses a fixed ShadowHand–UR platform, camera arrangement, table layout, cards, and chip denominations.It therefore does not establish cross-embodiment transfer, robustness to substantially different geometries, or arbitrary-object dexterity.
  • Data: The dataset contains 1,470 demonstrations, which define the task suite but are insufficient to study broad policy scaling behavior.The paper states that larger collections would be needed for that purpose.
  • Evaluation cost: Qualitative simulator reconstruction supports scene replay and geometry inspection but does not validate contact dynamics or replace physical evaluation.Faithful evaluation still requires hardware access, scene setup, and human effort.
  • Future direction: The benchmark’s real-contact policy signal remains costly to evaluate, motivating future work on reducing setup and human effort without losing physical validity.

B Benchmark Documentation

The benchmark documentation separates low-level primitive skills from agent-level perception, routing, verification, and recovery, while providing public access to the dataset, code, and assets.

  • Availability: DexHoldem’s demonstration dataset is hosted on Hugging Face, while project code, benchmark assets, and related repositories are maintained under its GitHub organization.Data-collection setup and dataset contents are summarized in Section B.2.
  • Task levels: Primitive-level tasks define callable dexterous skills for data collection, policy training, and physical rollouts.
  • Task levels: Agent-level tasks define perception, routing, verification, and recovery problems arising when primitives compose into Texas Hold’em interaction.
  • Evaluation interpretation: This two-level separation distinguishes low-level manipulation results from closed-loop embodied-agent behavior.

B.1 Embodied System Design Details

The embodied system combines image capture, structured-state perception, rule-based routing, primitive dispatch, verification, recovery, and possible human intervention. Its components share a standardized robot interface while differing in policy architectures and deployment constraints.

  • Closed-loop agent: The agent repeatedly captures an agent-view image, parses structured game state, routes decisions, and translates high-level actions into dexterous or non-robot operations.The runtime includes waiting, verification, completion, continuation, recovery, and human-help branches.
  • Primitive translation: Chip-betting actions dispatch one push or pull primitive per chip in descending denomination order, allowing failed primitives to be retried individually.The mapping uses 100 →50 →10 →5 denomination order.
  • Runtime support: The sandbox bundles a workflow document, perception guidelines, and deterministic helpers for capture, state management, routing, dispatch, and side effects.
  • Policy interface: Policies receive synchronized three-camera observations, proprioception, and task conditions, and output short-horizon targets in a shared 30-dimensional arm-and-hand space.The interface allocates 6 dimensions to the arm and 24 to the dexterous hand.
  • Policy implementations: The benchmark includes diffusion-policy, Transformer, CVAE, action-token, and pretrained π-series implementations adapted to the shared interface.RDT uses diffusion training and DPMSolver deployment, while RDT-small is trained from scratch without pretrained weights.
  • Deployment boundary: BeingH is excluded from the main real-robot comparison because unstable motion would require platform-dependent filtering outside the standardized protocol.

B.2 Dexterous Hand Policy Bench Details

The policy benchmark evaluates 14 teleoperated card and chip primitives on fixed ShadowHand hardware using multi-view sensing, standardized splits, physical rollouts, and a scene-preservation-aware scoring rubric. Results show distinct failure patterns across pickup, chip motion, and card placement or revealing.

  • Data and evaluation: Each primitive provides 100 training and 5 validation teleoperated trajectories, while evaluation uses 80 physical trials per policy.Pickup primitives receive 10 rollouts each and seed downstream placement and revealing trials; other primitives receive 5.
  • Hardware and sensing: The setup uses a ShadowHand mounted on a UR10e arm and three RGB-D cameras covering tabletop layout, global geometry, and hand–object contact.Streams are recorded at approximately 15 Hz with aligned depth and color.
  • Skill suite: The 14-primitive suite covers card pickup, face-down placement, card revealing, and denomination-specific chip pushing and pulling.Each trajectory has an instruction ID, natural-language task description, and physical success condition.
  • Scoring: The scoring rubric separates scene-preserving success from disruptive completion, task failure, and disruptive failure.SPSR counts only reusable-scene successes, whereas TCR also counts disruptive completions.
  • Failure patterns: Pickup is solved more reliably by the π-series, whereas chip-motion success is lower, especially for chip pulling.
  • Failure patterns: Put-down/show tasks often complete the requested card action more frequently than they preserve the full scene, creating a large SPSR–TCR gap.The benchmark therefore distinguishes object-level task completion from maintaining a usable tabletop state.

B.3 Agentic Perception Bench Details

The agentic perception benchmark evaluates structured state recovery from tabletop observations using deterministic, applicability-conditioned scoring across multiple perception fields. It mirrors deployment-time workflows while grading only validated structured outputs against held-out labels.

  • Benchmark design: 36 problems evaluate structured visual summaries of representative tabletop states encountered during DexHoldem deployment.Each problem uses a uniform prompt and fixed schema for the perception stage.
  • Benchmark design: Each field follows a distinct visual guideline specifying where to look and how to convert visual evidence into a structured value.Perceivers also receive deployment-time workflow scripts and routing guidelines, but do not execute scripts during evaluation.
  • Scoring: Nine deterministic scoring columns comprise one Overall column and eight sub-capability columns grouped by applicability.Universal columns cover loop stage, turn ownership, and blind assignment; chip-state and outcome columns apply to designated problem subsets.
  • Scoring: Outcome_judge problems require all eight sub-capability columns, whereas several other problem types require only the three universal columns.Overall is conditioned on each problem’s applicable fields and therefore is not an average of sub-column accuracies.
  • Evaluation procedure: Evaluation runs are scored only after artifact validation, then compared with ground-truth labels using deterministic column-level checks.The checks cover the eight perception challenges defined for the benchmark.

C Experiment Details

Figure 5 is a standalone diagnostic relating policy pretraining scale and policy size to physical task completion, complementing—but not replacing—the main trial-count comparison.

  • Diagnostic interpretation: Figure 5 places task-trained imitation baselines and adapted pretrained policies on common axes for pretraining scale, policy size, and physical task completion.Models without pretraining occupy the zero-pretraining position, while pretrained models use stated or estimated pretraining hours.
  • Diagnostic interpretation: The main quantitative comparison should be read from Table 1 trial counts and rates rather than inferred solely from the diagnostic plot.The figure summarizes relative scale and observed completion behavior.

C.2 RDT Fine-Tuning Curve Details

The RDT fine-tuning probe compares random and pretrained initialization across increasing DexHoldem data ratios using matched validation conditions. Pretrained initialization shows no distinct low-data regime at 10% but yields a modest lower-loss offset with more demonstrations.

  • Study design: The probe compares identical RDT architectures initialized randomly or from a pretrained RDT checkpoint across 10%, 20%, 50%, and 100% data ratios.These settings use 10, 20, 50, and 100 training trajectories per primitive, respectively.
  • Study design: Each data ratio retains the same five held-out validation trajectories per primitive, enabling comparison under a fixed validation split.The paired curves report train-time validation loss over completed paired seeds.
  • Curve interpretation: Lower validation loss indicates better held-out prediction of normalized action sequences under the supervised objective.The shaded region represents one standard deviation across completed paired seeds.
  • Observed pattern: Pretrained initialization does not create a distinct low-data regime at 10% data, but produces a modest lower-loss offset once more dexterous-hand demonstrations are available.The result is interpreted as an optimization or initialization benefit rather than a fundamentally distinct low-data regime.

C.3 System-Level Trajectory Panels

The trajectory panels visualize three closed-loop DexHoldem rollouts using Codex GPT 5.5 with π0, including routing states, waits, retries, human-help requests, and primitive expansions. The cases vary from interrupted short play to a long showdown and chip-collection sequence.

  • Panel encoding: The panels encode top-level agent primitives, wait reasons, continuation gates, visual card-reading steps, verification, completion, recovery, retry, and terminal states.They use the agent-to-policy primitive mapping documented in Table 4.
  • Trajectory (i): Trajectory (i) contains 22 states, two human-help escalations after unsettled scenes, and an interrupted final call before chip-push completion.Its post-flop sequence is raise (10), check, check, call.
  • Trajectory (ii): Trajectory (ii) contains 54 states and expands from escalating raises through showdown and winnings collection into six chip-pull atoms.The sequence includes raise (5), raise (105), raise (100), raise (100), all_in, two show_card primitives, and collect_winnings.
  • Trajectory (iii): Trajectory (iii) contains 23 states, follows raise (10), check, check, call, and finishes with two show_card primitives after one router-level retry.This is the case study analyzed in Section 4.5.

D Simulation Check

The simulation check reconstructs the DexHoldem tabletop and robot setup, then replays recorded real trajectories under camera views matched to the physical system. It provides qualitative evidence of reconstruction and replay consistency, not quantitative policy validation.

  • Interpretation: The check documents simulator–task mapping and replay consistency rather than serving as an additional policy-evaluation experiment.Its purpose is to verify correspondence between the real trajectories, robot embodiment, tabletop layout, object categories, and primitive definitions.
  • Reconstruction and replay: The check reconstructs the UR10e arm, ShadowHand, poker table, task-relevant cards and chips, and their source and target regions.Recorded arm and hand joint trajectories are replayed in simulation as open-loop commands.
  • Qualitative visualization: Figures 7 and 8 show the reconstructed environment across three matched camera views and a six-frame top-down replay sequence.The views are top-down, third-person, and wrist perspectives, while the sequence illustrates successive replay stages.
  • Interpretation: The check is qualitative and does not remove the real-sim gap, especially for thin cards and chips whose contact behavior depends on friction, compliance, pose errors, and disturbances.Policy results therefore remain based on physical robot rollouts.
Loading 2605.18727v1…