Source-linked AI summary
A Physics-Consistent Benchmark for Contact-Rich Human-Robot Interaction in Assistive Care
Chengxiao He, Shanghai Yuan, Liuqun Fan, Shenzhen Zhu
TL;DR
Conventional task-level evaluation does not capture physically invalid contact in assistive human–robot interaction. The paper introduces a calibrated, physics-consistent bathing benchmark with responsive human mechanics, physics-aware scoring, and leak-free observations. Its evaluations show that task success can diverge from correct-region and safety-gated success, motivating physical screening.
Problem
Conventional task-level metrics do not capture deformation, passive motion, force, joint limits, and contact quality in contact-rich assistive-care interaction.
Method
The benchmark combines a deformable passively responding human, manikin-calibrated contact response, physics-aware scores, and a frozen vision-only / scorer-only protocol.
Results
Task completion diverges from physically valid contact: State Machine reaches 72.9% task success but 56.4% after correct-region and force-safety screening, while VoxPoser reaches 27.9%.
Takeaways & Limitations
Task completion alone does not imply physically valid contact, so the benchmark can screen contact-rich care policies before deployment.
Takeaways & Limitations
The study uses 20 trials per task with a medical-care manikin rather than human tissue and limited evaluated methods.
Abstract
from arXiv · showhide
Conventional task-level evaluation asks whether a robot policy completes a specified action, but can miss failures that emerge only during physical human contact. This limitation is critical in contact-rich assistive tasks, where meaningful evaluation requires a physically responsive human, interaction-quality assessment beyond task success, and a leak-free observer-scorer protocol. We introduce a physics-consistent benchmark for contact-rich human-robot interaction, instantiated in robot-assisted bathing. The benchmark combines a deformable, passively responding human, physics-aware scores alongside task-level success, and a frozen vision-only / scorer-only evaluation protocol. To establish physical validity, region-wise simulated responses are calibrated against force-indentation measurements from Franka impedance pushes on a medical-care manikin. Under a frozen T1-T7 protocol with 140 runs per method, an LLM-augmented state machine (State Machine) achieves 72.9% task success but drops to 56.4% after correct-region and force-safety screening; VoxPoser produces lighter and more stable contact but completes only 27.9% of trials; and zero-shot pi0.5 achieves 0.7% task success with no correct-region or safety-gated successes. These results show that task completion alone does not imply physically valid contact and motivate physics-aware screening before deployment of contact-rich assistive robot policies.
I. INTRODUCTION
Conventional task-level evaluation misses physically invalid contact in assistive-care interaction. The benchmark addresses this gap with a responsive human, physics-aware scoring, calibrated contact response, and leak-free evaluation.
- Physical human contact involves deformation, passive joint motion, force, and joint-limit constraints that conventional task metrics do not capture.
- Existing benchmarks often simplify humans as rigid or motion-replayed and emphasize task completion, trajectories, or collision measures.
- The benchmark combines a deformable, passively responding human, physics-aware scores, task-level success, and a frozen vision-only / scorer-only protocol.
- Region-wise simulated responses are calibrated against force–indentation measurements on a medical-care manikin to establish physical validity.
- 102/140 State Machine successes fall to 79/140 under safety screening, while π0.5 falls from 1/140 original success to 0/140 gated success.
- The benchmark evaluates T1–T7 Touch / Scrub / Push and diagnoses failures in force, joint-limit margin, contact stability, and region grounding.
II. RELATED WORK
Prior systems contribute care execution, simulation, safety criteria, or policy evaluation separately, but do not jointly provide calibrated, leak-free, physics-aware contact evaluation. This benchmark integrates those capabilities for contact-rich assistive interaction.
- Care systems demonstrate local task completion but not a reusable protocol for scoring physical contact under a frozen observation contract.
- Language-grounded and end-to-end policy systems evaluate manipulation performance, but strong general scores do not imply safe sustained contact with deformable human bodies.
- The paper treats soft-body engines as infrastructure and judges the environment by real-to-sim contact agreement rather than solver contests.
- Existing work provides separate ingredients rather than a unified combination of passive response, calibrated contact, leak-free observations, and physics-aware scoring.
III. BENCHMARK DESIGN
The benchmark defines physics consistency at the benchmark level through responsive simulated human mechanics, calibrated physical correspondence, and policy isolation from privileged simulator state.
- A physics-consistent benchmark requires measurable real-to-sim response correspondence, dynamically generated contact states, and no privileged-state exposure to policies.
- The simulated care recipient combines an articulated rigid skeleton with region-wise deformable soft-tissue shells and enabled passive joints.
- Soft-region contact deforms shells, transfers force through regional anchors, and can move passive joints rather than treating the mannequin as rigid.
- The benchmark comparison concerns benchmark-level capabilities relevant to the evaluation setting.
- The observation interface exposes RGB-D images, point clouds, semantic masks, and visual skeletal landmarks while withholding joint angles, anchor states, and contact metadata.
B. Benchmark Tasks and Evaluation Dimensions
The task suite targets sustained contact with deformable body regions and evaluates both task completion and physical interaction quality. Safety gates use conservative region-wise force thresholds.
- Seven micro-tasks evaluate continuous physical interaction across three soft-relevant body regions from randomized supine care-recipient states.
- The upper-arm push is omitted because its short moment arm provides limited displacement diversity.
- Scrub tasks mean sustained sliding cleaning contact, while rigid-landmark, language-ambiguity, and conventional arm-reposition tasks are excluded.
- Scorer-only ground truth provides task-level success together with physics-aware scores describing physical contact rather than clinically validated care outcomes.
- 35 N, 35 N, and 21 N are the conservative force gates for the forearm, upper arm, and abdomen, respectively.
C. Evaluation Protocol
The protocol freezes the scorer–logger contract across methods while keeping scorer-only physical streams hidden from policies.
- All evaluated methods use the planner-visible snapshot interface, while scorer-only streams exclude future contact labels.
IV. EXPERIMENTS
The experiments evaluate methods on the full T1–T7 suite under a frozen protocol, using scripted probes to validate the scoring apparatus before policy comparison.
- The benchmark treats passive-joint response as required rather than as an optional ablation.
- Scripted probes verify that push depth, joint-envelope violations, stop–go execution, and force fluctuation affect their intended scores.
- 140 runs per method cover the full T1–T7 suite, with 20 initializations across seven tasks.
- The State Machine uses an LLM-augmented finite-state controller with vision-derived, scripted Cartesian contact candidates.
- Zero-shot π0.5 is evaluated as a transfer stress test without additional teleoperation or finetuning.
- Reported metrics distinguish valid scoreable contact or trajectory data from C3 and use Wilson intervals for binary rates and medians for continuous metrics.
B. Physical Validity of the Benchmark
Physical validity is assessed by replaying measured manikin contact conditions in simulation and comparing force–indentation responses, while recognizing calibration limits.
- Real trials use a Franka Panda under Cartesian impedance control with a compliant, movable-joint medical-care manikin.
- Matched contact schedules, locations, and directions are replayed in simulation to calibrate Flex parameters against measured force–indentation curves.
- Logged contact force, wrist displacement, and local deformation illustrate the scorer–logger contract rather than define deformation as a separate metric.
- Upper-arm results reuse forearm calibration parameters and therefore provide benchmark-level comparative scores, not independent region-specific biomechanical validation.
C. Benchmark Evaluation
Benchmark evaluation separates task completion from correct-region and safety-gated success, revealing method-specific differences in contact quality and robustness.
- 72.9% task-level completion is achieved by State Machine, but its median peak force is 19.48 N and stable contact remains distinct from completion.
- 56.4% safety-gated success is reported for State Machine at γ′ = 0.35, below nominal success but above VoxPoser throughout the gate sweep.
- 27.9% gated success remains unchanged for VoxPoser across γ′ from 0.20 to 1.00.
- 0.0% gated success remains unchanged for π0.5 across the γ′ sweep.
- C4 measures active-motion fraction, with higher values indicating fewer execution interruptions.
- Safety-gated deployment success requires original evaluator success, correct-region contact, and peak-force screening.
INITIAL RGB SNAPSHOT AND PERFORMS NO EXECUTION-TIME VISUAL
The evaluation contrasts task completion with contact quality, showing that light contact, continuous motion, or nominal success can still fail to produce physically valid interaction.
- 27.9%: VoxPoser achieves original evaluator success, despite producing the lightest and most stable conditional contact.Its contact statistics are C1 1.46 N and C5 0.21 N.
- 0.7%: π0.5 achieves near-zero success, with every recorded contact occurring on a non-target body region.Its scoreable contact set is limited to 30 force runs and 6 Push-margin runs.
- The pooled physics-aware scores use available scoreable runs, so valid sample counts vary by metric and method.
- State Machine’s heavy force tail distinguishes nominal completion from stable contact.Its median peak force is 19.48 N, and it has the second-worst C5 score.
- π0.5’s C1 and C5 statistics represent off-target-region contact rather than instructed target-region contact.Figure 4(e) shows TCP trajectories that fail to lock onto the instructed region.
D. What Conventional Evaluation Misses
Region correctness, sustained progress, force safety, and contact location expose failures that task-level success alone can conceal. The methods therefore occupy distinct regimes of effective, unstable, or off-target interaction.
- 95.0%: State Machine and VoxPoser establish target-region contact on forearm runs, but VoxPoser’s original success on that set is only 33.3%.Reaching the instructed site therefore remains distinct from finishing the action.
- 68/101: VoxPoser’s dominant failure is insufficient progress among region-correctness failures, despite high motion activity of C4 0.883.C4 and C3 together show that motion continuity does not determine effectiveness.
- Figure 4’s valid ratio measures scoreable-data coverage, not closed-loop success.Failure-mode percentages are conditioned on region-correctness failures, using denominators of 38, 101, and 140 runs.
- Reaching and sustained-contact phases may require distinct control regimes rather than one value-map objective throughout.The proposed progress signals include effective stroke for Scrub and wrist displacement for Push.
- 72.9%: State Machine task success drops from 102/140 to 79/140 under the safety gate.The benchmark identifies this post-hoc physical screening as deployment-usable success rather than nominal completion.
- 0/140: π0.5 records no correct-region success and no gated success after 1/140 original success.Its failures are 106/140 no-target-contact runs plus 30/140 off-target-contact runs.
V. CONCLUSION
The benchmark makes task–contact dissociations measurable in robotic bathing, but its characterization remains bounded by the manikin setting, limited trials, and a narrow method set.
- 20 trials per task: the study targets benchmark characterization rather than population-level statistical estimation.
- The evaluation uses a medical-care manikin rather than human tissue and reuses forearm calibration for the upper-arm simulation.
- The evaluated methods are limited to a LLM-augmented state-machine reference, a spatial-reasoning planner, and one zero-shot Franka-DROID checkpoint.The instantiation is a single robotic-bathing scenario.
- Table VIII reports region-wise contact outcomes, with rankings evaluated separately within each body region.Its denominators are 60 forearm runs and 40 runs each for upper arm and abdomen.
- Task completion does not imply physically valid contact: State Machine can finish with an unstable force tail, VoxPoser can remain light without sustained completion, and π0.5 can remain off-target.The benchmark is positioned as a pre-deployment physical/safety screen for contact-rich care policies.