Source-linked AI summary

DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects

Tianshan Zhang, Yijia Duan, Yanjun Li, Zeyu Zhang, Hao Tang

arXiv:2606.15133v1cs.ROcs.CV

TL;DR

Articulated-object manipulation requires motion to emerge from sustained hand–handle contact, which geometric replay and fixed-dynamics policies do not adequately capture. DragMesh-2 introduces contact-driven hand control with PICA training and shows stronger robustness to contact-load variation across seven GAPartNet objects while maintaining high task success.

  • Problem

    Existing methods inadequately model sustained hand–object contact dynamics and may overfit fixed contact loads without force or tactile feedback.

  • Method

    DragMesh-2 uses hand-only contact-driven control, while PICA injects physically informed signals and dynamics variation into training without force sensing.

  • Results

    Across seven GAPartNet objects and varied damping conditions, DragMesh-2 achieves stronger robustness than compared methods while maintaining high task success.

  • Takeaways & Limitations

    The study supports contact-aware training as a way to improve dexterous hand–articulated-object interaction under changing contact loads without tactile or force feedback.

  • Takeaways & Limitations

    Success drops from 0.89 at ×1 damping to 0.56 at ×4, with heterogeneous per-object results and action saturation under strong contact loads.

Abstract

from arXiv · show

Dexterous interaction with articulated objects is important for household, assistive, and humanoid manipulation, where multi-finger hands can provide compliant contact patterns beyond parallel-jaw grasping. However, articulated-object manipulation differs from static-object manipulation: the target part cannot be directly actuated, and its motion must emerge through sustained physical hand--handle contact. This makes the transition from object-centric articulated generation to hand-driven dexterous hand--object interaction non-trivial, since geometric trajectory replay or open-loop execution does not model the contact dynamics required to move the articulated part. Moreover, policies trained only for task completion under fixed dynamics can overfit nominal contact loads, especially without tactile or force feedback, and may degrade when the contact load changes. To address these challenges, we present DragMesh-2, a contact-driven framework for dexterous interaction with articulated objects that extends articulated interaction from object-centric generation to hand-driven dexterous hand--object interaction, where articulated motion must arise through physical contact. We further propose PICA, a physically informed contact-aware training mechanism that injects physical signals into policy learning without tactile or force feedback, improving robustness and task success under changing contact loads. Finally, we conduct systematic evaluation across multiple damping conditions and articulated-object categories to study robustness under contact-load variation, and provide a pure-geometry dexterous interaction resource to support future loco-manipulation and humanoid hand--object interaction research. Across seven GAPartNet objects, DragMesh-2 achieves stronger robustness under contact-load variation than the compared methods while maintaining high task success across damping conditions.

1 Introduction

DragMesh-2 frames articulated-object manipulation as dexterous hand-driven interaction in which articulated motion must emerge from sustained physical hand–handle contact. It introduces PICA for contact-aware training without tactile or force feedback, alongside systematic damping-condition evaluation and a pure-geometry interaction dataset.

  • Motivation: Dexterous hands enable compliant multi-finger contact patterns for articulated-object interaction in household, assistive, and humanoid manipulation.This establishes the application importance and distinguishes dexterous hands from parallel-jaw grippers.
  • Challenge: Articulated parts cannot be directly controlled, so their motion must emerge through sustained hand–object contact that geometric replay and open-loop execution fail to model.Existing reinforcement-learning methods trained under fixed dynamics can also miss changing contact loads.
  • DragMesh-2: DragMesh-2 controls only the dexterous hand, requiring articulated motion to emerge through physical hand–handle contact without a target-joint action channel.The framework extends articulated interaction from object-centric motion generation to dexterous hand–object interaction.
  • PICA: PICA injects contact maintenance, detachment risk, action-boundary regularization, damping variation, and temporal contact response into policy learning.It improves robustness under contact-load changes without tactile or force feedback.
  • Evaluation and resource: The evaluation spans multiple damping conditions and articulated-object categories, measuring task success alongside action saturation and detachment diagnostics.The accompanying pure-geometry trajectories support grasp initialization, task-scale normalization, and tracking references for reproducible interaction.

2 Related Work

Related work spans articulated-object understanding and manipulation, dexterous hand-object control, and physics-grounded learning. Existing approaches use geometric modeling, interaction data, domain randomization, or constrained optimization to address contact and dynamics challenges.

  • Articulated Object Understanding and Manipulation: Articulated-object research covers part-level perception, pose estimation, shape representation, joint parameter prediction, and interactive simulation platforms.
  • Dexterous Hand-Object Manipulation: Classical dexterous manipulation models contact mechanics, grasp stability, and force closure, whereas reinforcement learning learns high-DoF policies directly from interaction.Model-based methods require accurate geometry and contact models, while reinforcement learning has performed strongly in rigid-object in-hand manipulation.
  • Physics-Grounded Manipulation Learning: Vision-based teleoperation and domain-randomized learning enable dexterous control from visual observations without tactile sensing, while handling variation through data diversity or randomization.
  • Physics-Grounded Manipulation Learning: Regularized action generation helps prevent saturated joint commands from breaking contact, while constrained reinforcement learning separates task objectives from control constraints.Reward penalties require manual weight tuning, whereas Constrained MDPs and Lagrangian optimization formalize the separation.

3 Method

DragMesh-2 formulates articulated-object manipulation as contact-driven control of a 51-DoF hand, with object motion emerging solely through hand–handle contact. Its PICA mechanism adds observable physical proxies to PPO through contact-aware rewards and temporal prediction, while geometry-guided trajectories support initialization, benchmarking, and resource release.

  • Task formulation: The 51-DoF hand controls the task, while the object joint has no action channel and moves only through hand–handle contact.The hand comprises 6 virtual wrist DoFs and 45 finger joints.
  • Task formulation: Task success and progress are normalized by each object’s reference-trajectory motion range, making drawers, sliders, and doors comparable without fixed displacement.The target joint is selected as the object DoF with the largest motion range.
  • Policy interface: Observations exclude visual, force, and tactile signals, while the 51-dimensional action updates clipped hand PD targets.Inputs include hand states, handle pose, palm–handle geometry, target-joint state, and task-scale progress features.
  • PICA: PICA augments PPO with contact-maintenance rewards and auxiliary prediction of object response, palm–handle distance, detachment risk, and tracking stress.The reward includes task progress, distance, action, time, detachment, success, boundary, and contact terms.
  • Reference trajectories: A geometry-guided procedure generates 277 phased contact trajectories across 7 GAPartNet categories for initialization, a tracking baseline, and a released pure-geometry resource.Trajectories contain approach, grasp, drag, and release phases and are generated without learning.

4 Experiments

Experiments evaluate DragMesh-2 on seven GAPartNet objects under varying damping and execution modes, comparing contact-driven policies with geometric and replay references. Results show that PICA provides the strongest learned-policy robustness, while physical signals, temporal encoding, and contact diagnostics explain performance under changing loads.

  • Evaluation Setup: The benchmark covers 7 GAPartNet objects from three categories, with five revolute doors and two prismatic drawers, initialized from expert grasp states.The target part opens only through hand–handle contact, using reference contact trajectories from the heuristically generated dataset.
  • Compared Methods: The comparison includes trajectory-tracking replay, a GT-part-pose parallel-jaw primitive, four learned baselines, and ablations removing PICA’s physical signals or GLA encoder.Evaluation also logs action-saturation ratio and detachment-failure rate alongside task success.
  • Main Comparison: 1.00 deterministic success at ×1 falls to 0.71 at ×2 and ×4 for trajectory replay, as two objects lose contact under higher damping.This demonstrates that open-loop replay can drive the part through contact nominally but is not robust out of distribution.
  • Main Comparison: 0.14 mean success is achieved by the GT-part-pose parallel-jaw primitive, while PICA leads learned policies in every damping and execution-mode column.PICA’s deterministic success declines from 0.89 at ×1 to 0.56 at ×4, compared with 0.27, 0.32, 0.30, and 0.09 for the listed learned baselines.
  • Ablation: 0.89 deterministic success at ×1 and 0.56 at ×4 result when PICA combines physical signals with the GLA encoder, exceeding either component alone by at least 0.13 at ×4.Using only GLA reaches 0.65 and 0.36, whereas using only physical signals reaches 0.75 and 0.43.
  • Robustness Diagnostics: 1.00 nominal success after 500 training epochs coincides with strong-damping success falling to 0.10, while clip099 approaches 1.0.Across 150 to 500 epochs, nominal success rises from 0.90 to 1.00, but ×4 success drops from 0.55 to 0.10, motivating contact-aware diagnostics and richer contact interfaces.

5 Limitations and Future Work

DragMesh-2 remains limited by action saturation and indirect contact-state inference, with success degrading under stronger contact loads. Future work includes richer force/tactile interfaces and extending contact-driven interaction toward whole-body humanoid control.

  • Limitations: Success drops from 0.89 at ×1 to 0.56 at ×4 under strong contact load, with heterogeneous per-object results and no universally dominant policy.The position-increment action interface tends toward action saturation as contact load increases.
  • Limitations: Without force or tactile feedback, contact state is inferred indirectly from kinematic error, which is insufficient for stable light pulling at high damping.The observation channel lacks direct force or tactile information.
  • Future Work: Future work should add wrist force or torque outputs and contact-force or tactile feedback so the policy can regulate grip force directly.This would replace pushing to the actuator boundary with direct grip-force regulation.
  • Future Work: Future work should couple upper-body contact interaction with whole-body control for humanoid loco-manipulation, using the dataset as a motion-scale prior.The current task isolates contact-driven pulling from an expert grasp state and controls a floating dexterous hand.

6 Conclusion

DragMesh-2 is a contact-driven framework for dexterous hand–articulated-object interaction in which the target part moves only through physical hand–handle contact. It shows that task success alone does not ensure stable contact behavior and that PICA improves robustness under contact-load shifts.

  • Framework: DragMesh-2 extends DragMesh 1 from object-centric articulated interaction to hand-driven physical interaction.The framework requires target-part motion to arise through physical hand–handle contact.
  • Robustness: Policies trained only for task progress can degrade sharply under contact-load shifts, showing that nominal task success does not guarantee stable contact behavior.This identifies contact stability under changing loads as distinct from task completion.
  • PICA: PICA improves robustness under contact-load shifts.It does so by adding physical information to policy learning, as stated in the conclusion passage.

A Additional Method Details … A.3.5 Policy Optimization

The appendix specifies DragMesh 2’s state-only control interface, contact-grounded evaluation protocol, and PICA implementation. It details physical regularization, randomized damping, temporal contact encoding, auxiliary supervision, and PPO optimization for robust hand-driven interaction.

  • A Additional Method Details: The appendix expands implementation details for observations, controls, rewards, temporal encoding, auxiliary supervision, optimization, diagnostics, and parameters.These details supplement the task definition, PICA signals, and evaluation metrics presented in Section 3.
  • A.1 Observation and Control Details: The policy observes geometric and joint state, outputs 51-dimensional incremental hand control, and moves the articulated part only through hand–handle contact.Actions are clipped to [−1, 1], scaled into local hand PD-target increments, and sent to a position PD controller.
  • A.2 Evaluation-Protocol Details: The trajectory-tracking baseline supplies reference hand poses as PD targets while never replaying object-joint states, testing whether contact induces physical opening.The object joint remains driven only through hand–handle contact.
  • A.3.1 Physical-Plausibility Reward: The physical-plausibility reward penalizes detachment and saturated actions while rewarding task progress, with ARAM and Reconfig providing optional contact-stabilization ablations.Detachment failure is gated on prior effective contact, and Reconfig encourages small-amplitude hand reconfiguration during stalls.
  • A.3 Reference Policy: Physical Signal Mechanism (PICA): PICA combines contact-regularized rewards, damping randomization, temporal contact history, and causal-window auxiliary supervision to constrain policies toward physically plausible behavior.The mechanism uses observable physical proxies rather than direct force or tactile measurements.
  • A.3.2 Damping Randomization: Damping randomization samples the target-joint damping scale uniformly from [1.0, 2.0] during training, exposing the policy to varying resistance without friction randomization.This reduces dependence on a single nominal dynamics setting.
  • A.3.3 Contact-History Temporal Encoder: A Gated Linear Attention encoder processes recent PD tracking errors and previous actions, helping infer contact impedance, sliding, impacts, and impending detachment.The temporal feature is fused with the current state for both actor and critic outputs.
  • A.3.4 Causal-Window Contact-Response Auxiliary Supervision: The policy uses causal-window auxiliary targets for object response, contact distance, detachment risk, and tracking stress, and optimizes them jointly with PPO, value, and boundary losses.The objective is L = LPPO + cv LV + cb Lbounds + waux Laux, with waux linearly warmed up.

B Reference-Trajectory Generation and Dataset Construction · C Additional Experimental Results · C.1 Qualitative Simulation and Hardware Visualizations

The paper constructs geometry-guided reference contact trajectories directly from GAPartNet geometry and mobility annotations, then provides auxiliary diagnostics and visualizations covering diverse articulated interactions. The released trajectories support DragMesh-2 evaluation across prismatic drawers and revolute doors without learned generation parameters.

  • B Reference-Trajectory Generation and Dataset Construction: The generator selects a target part from semantic annotations, prioritizing sliders or drawers, then doors, with the last annotated part as fallback.It chooses the handle annotation nearest the target part’s front face and extracts its oriented bounding-box geometry.
  • B Reference-Trajectory Generation and Dataset Construction: It constructs wrist orientation and finger configurations for open, pre-grasp, and force-closure poses, adapting the grasp to handle thickness.The wrist uses the handle’s long axis and outward normal, while its position compensates for the palm-center offset.
  • B Reference-Trajectory Generation and Dataset Construction: The trajectory proceeds through approach, grasp, drag, and release phases, combining wrist interpolation, finger closure, joint-aligned motion, and retraction.For revolute joints, the interaction center rotates about the joint axis and the wrist orientation synchronously tracks the handle.
  • B Reference-Trajectory Generation and Dataset Construction: The generator has no learned parameters and runs from GAPartNet geometry and mobility annotations, with phase durations and motion parameters adapted to object type.The generated trajectory is used by the DragMesh 2 environment, trajectory-tracking baseline, and evaluation motion-scale reference.
  • C Additional Experimental Results: The appendix reports auxiliary evaluation-side diagnostics across success, progress, clip099, and detach_proxy under damping conditions rather than comparing raw training rewards.Raw reward curves are omitted because reward scales differ after contact regularizers, auxiliary terms, and fine-tuning modules are introduced.
  • C Additional Experimental Results: Checkpoint diagnostics define a PICA base policy and apply ARAM and Reconfig contact-stabilization modules as 50-epoch fine-tunes, with SELECTED and OVERTRAINED reserved for rollout visualizations.The base policy is trained with the full physical-signal mechanism on the diagnostic object.
  • C.1 Qualitative Simulation and Hardware Visualizations: Additional visualizations show the contact-driven formulation handling both prismatic drawers and revolute doors through approach, grasp, and drag stages across object instances.Terminal-stage renders replay stored right-hand trajectory states using the simulator’s 51-DoF floating SMPL-X right-hand model and show category and joint-type diversity.

C.2 Rollout-Level Diagnostics … C.5 Limits of Damping-Distribution Expansion

The diagnostics distinguish stable contact-oriented behavior from overtrained failure under resistance, while extended fine-tuning and damping-range expansion fail to deliver reliable out-of-distribution robustness. These results indicate that sustained high-load interaction may require capabilities beyond the current action and observation interfaces.

  • C.2 Rollout-Level Diagnostics: SELECTED usually preserves pulling contact under damping, but its failures reflect insufficient sustained output rather than immediate detachment.On 7310, SELECTED succeeds in 5 of 6 rollouts; on 45936, ×4 failures involve long episodes and continued contact interaction.
  • C.2 Rollout-Level Diagnostics: OVERTRAINED fails every 7310 rollout, rapidly entering abnormal contact with mean episode length 6.0 steps and best progress 0.288.Its perturbations do not translate into target-joint progress, consistent with ineffective responses under resistance.
  • C.3 OOD Damping Robustness: At diagnostic object 45936 under ×4 damping, early checkpoints retain 0.50–0.55 success, whereas base-policy training to 300 or 500 epochs drops to roughly 0.10.The collapse indicates relatively rapid out-of-distribution degradation in the base policy.
  • C.4 Limits of Extended Fine-Tuning: Extended Both-module training preserves near-1.0 training success but degrades deterministic ×2 success from 0.85 to 0.25, 0.60, and 0.30 at epochs 250, 300, and 399.Stochastic ×2 also falls from 0.70 to the 0.10–0.20 range, while deterministic ×4 remains near 0.0.
  • C.4 Limits of Extended Fine-Tuning: ARAM delays but does not prevent OOD collapse, and the epoch-399 training-reward-best checkpoint remains degraded, making training reward alone insufficient for OOD checkpoint selection.The Both-module fine-tune reproduces the base-policy collapse over roughly four times the time scale.
  • C.5 Limits of Damping-Distribution Expansion: Broadening the damping interval from [1.0, 2.0] to [1.0, 4.0] raises deterministic ×4 success only from 0.00 to 0.05 while reducing deterministic ×2 from 0.85 to 0.50.Stochastic ×2 similarly decreases from 0.70 to 0.35, and stochastic ×4 regresses overall.
  • C.5 Limits of Damping-Distribution Expansion: An intermediate broadened-training checkpoint transiently reaches 0.25 deterministic ×4 success, but its ×2 success falls to roughly 0.55 and later epochs do not reproduce the effect.Other tested damping ranges produce sharper degradation at ×4 and ×2.
  • C.5 Limits of Damping-Distribution Expansion: Damping-range expansion alone provides limited benefit because stable light pulling under sustained high load may exceed the representational capacity of the 51-dimensional position-increment action and force-free observation channels.The stated limitation motivates changes to the task or policy interfaces for improved OOD ×4 robustness.

C.6 Ablation Analysis · C.7 Summary of Diagnostic Findings · C.8 Relationship Between Physical Diagnostics and Temporal Encoding

The diagnostic experiments show that robustness under strong damping depends on coupled physical mechanisms rather than training duration or damping-range expansion alone. Together, the findings motivate joint design of contact diagnostics, robustness evaluation, and temporal encoding.

  • C.6 Ablation Analysis: Additional training improves nominal and mid-damping success but does not monotonically improve strong-damping robustness.The OOD ×4 degradation concentrates between epochs 200 and 300, with the 300-epoch checkpoint nearly matching the 500-epoch checkpoint on ×4 success.
  • C.6 Ablation Analysis: Damping-range expansion alone improves deterministic ×4 only from 0.00 to 0.05 while degrading ×2 performance.The extended-damping fine-tuning variant therefore provides limited benefit within the incremental-position control and reward-shaping framework.
  • C.6 Ablation Analysis: ARAM improves deterministic nominal-damping success but provides no ×4 robustness gain, whereas Reconfig raises deterministic ×2 success to 0.95.Reconfig also preserves higher ×4 progress, indicating a role for contact reconfiguration under moderate resistance.
  • C.7 Summary of Diagnostic Findings: PICA’s contact regularizers, dynamics randomization, and temporal auxiliary supervision function as a coupled protocol rather than independent mechanisms.GLA supplies temporal capacity, while stable contact response requires explicit contact-maintenance and action-regularization signals.
  • C.7 Summary of Diagnostic Findings: Task reward and nominal success are unreliable checkpoint selectors for strong-load robustness across extended fine-tuning and damping-range expansion.The diagnostic summary motivates reporting clip099, detach_proxy, progress, and damping-conditioned performance alongside success.
  • C.8 Relationship Between Physical Diagnostics and Temporal Encoding: Physical diagnostics, robustness summaries, and temporal encoding should be designed jointly for contact-driven articulated-object manipulation.Temporal encoding alone can accelerate a nominal-dynamics shortcut, while physical shaping alone provides robustness without fine-grained historical contact modeling.

C.9 Limitations · C.10 Future Work · D Implementation Details and Inference Settings

The paper qualifies its results by emphasizing uncertainty across small episode aggregates and the absence of direct force or tactile sensing. It identifies force-aware supervision and whole-body extensions as future directions, while relegating reproducibility parameters to the appendix.

  • C.9 Limitations: Each reported cell aggregates 20 episodes, giving a success-estimate standard error of approximately 0.10.Small differences under strong damping, such as 0.15 versus 0.30 at ×4, should be interpreted as the same order rather than strictly ranked.
  • C.9 Limitations: The conclusions therefore rely on aggregate trends across objects, damping multipliers, and execution modes rather than any single cell.
  • C.9 Limitations: Without force or tactile signals, the policy infers contact state indirectly from kinematic error.Under strong damping, this limitation contributes to residual light-pulling failures in diagnostic results.
  • C.10 Future Work: Future work proposes force-aware control and richer contact-response supervision through expanded action and observation interfaces.Suggested auxiliary targets include contact normals and sliding velocity.
  • C.10 Future Work: Within strong-load regimes, future studies could separate light and heavy pulling into distinct contact modes.The paper also outlines a whole-body loco-manipulation extension toward contact-rich articulated interaction.
  • D Implementation Details and Inference Settings: The appendix records implementation parameters omitted from the main text to support reproducibility.These values are presented as implementation details rather than methodological contributions, preserving the main text’s focus on task definition, evaluation, and diagnostics.

D.1 Task and Control Parameters … D.5 Damping Randomization and Inference

The appendix specifies shared task, reward, auxiliary-loss, and PPO parameters, then defines damping randomization and inference protocols for evaluating robustness. These settings include fixed cross-object training choices, varied damping ranges, and 20-episode quantitative rollouts.

  • D.1 Task and Control Parameters: Task and control settings are shared across all policies, covering the success threshold, action scaling, episode length, temporal encoder, and network dimensions.These constants are documented in Table 9.
  • D.2 Reward and Termination Parameters: The reference policy uses specified reward weights, contact thresholds, and one-shot bonuses and penalties for reward and termination.These parameters are documented in Table 10.
  • D.3 Auxiliary-Supervision Parameters: The causal-window auxiliary loss uses defined channel weights and a warmup schedule for its overall contribution.These auxiliary-supervision parameters are documented in Table 11.
  • D.4 PPO Training Parameters: PPO optimization uses standard on-policy hyperparameters that are not tuned separately for each object.The PPO settings are listed in Table 12.
  • D.5 Damping Randomization and Inference: During training, target-joint damping is sampled uniformly from [1.0, 2.0] by default, with extended ranges [1.0, 4.0], [1.0, 2.5], and [2.0, 4.0].These extended intervals probe the boundary of training-distribution variation; friction randomization is not enabled.
  • D.5 Damping Randomization and Inference: At evaluation, the checkpoint remains fixed while rollouts use ×1, ×2, and ×4 damping.Deterministic execution uses the policy mean, whereas stochastic execution samples from the learned Gaussian policy.
  • D.5 Damping Randomization and Inference: Each quantitative evaluation cell uses 20 episodes by default, while single-rollout visualizations are reserved for behavioral diagnostics.Visualizations are not pooled into multi-episode statistics.
  • D.5 Damping Randomization and Inference: Checkpoint selection should combine out-of-distribution damping evaluation with additional criteria rather than relying on training reward alone.The supplied passage begins an additional criterion with “acti,” but the remainder is unavailable.
Loading 2606.15133v1…