Source-linked AI summary
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
Dingyi Rong, Yue Shi, Chaofan Ma, Jiezhang Cao, Zongrui Wang, Zeyu Zhang, Yao Mu, Guangtao Zhai, Ning Liu
TL;DR
Robot demonstrations are costly to collect, while transferring abundant human manipulation videos across embodiments remains challenging. H2R-Bench evaluates this transfer across tasks and robot embodiments, finding a substantial gap between visual quality and embodied manipulation capability.
Problem
Robot demonstrations are expensive to collect, and evidence remains limited on preserving human manipulation evidence while adapting it to robot embodiments.
Method
H2R-Bench converts egocentric human demonstrations into specified-embodiment robot videos and evaluates goals, actions, contacts, embodiment, and video quality.
Results
Current video world models show a substantial gap between visual generation quality and embodied transfer capability across six manipulation families and two embodiments.
Takeaways & Limitations
H2R-Bench provides a systematic diagnostic framework for assessing whether video world models can bridge the human-to-robot embodiment gap.
Takeaways & Limitations
The benchmark evaluates visible transfer rather than physical executability or downstream policy performance, and covers only short clips, 120 sources, and two embodiments.
Abstract
from arXiv · showhide
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.
Introduction
H2R-Bench addresses the scarcity and expense of robot demonstrations by evaluating whether models can transform egocentric human manipulation videos into robot videos under specified embodiments. It diagnoses transfer through source-grounded task, action, contact, embodiment, and video-quality criteria.
- Motivation: Robot manipulation video collection is expensive, whereas egocentric human videos abundantly capture object affordances, hand-object contacts, manipulation intent, and physical state changes.This asymmetry motivates human-to-robot and cross-embodiment video generation as a bridge from abundant human demonstrations to scarce robot demonstrations.
- Evaluation gap: Existing evaluation protocols mainly assess visual fidelity, temporal consistency, text-video alignment, motion quality, or related physical and robotic properties rather than complete human-to-robot transfer.A valid transfer should preserve the source task intent, required actions, and functional interactions.
- Benchmark: H2R-Bench evaluates source-conditioned transformation of an egocentric human demonstration into a robot manipulation video under a specified embodiment.Each demonstration is paired with a human-verified structured task specification grounded in the source video.
- Evaluation dimensions: The benchmark scores goal completion, action-event completion, functional contact transfer, embodiment correctness, and task-agnostic video quality.Its weighted aggregate H2RCore emphasizes contact and embodiment while retaining a smaller contribution from video quality.
- Benchmark scope: H2R-Bench spans six manipulation families and two target embodiments and evaluates 11 representative video generation models.Its structured annotations cover task goals, action events, object-state evolution, and functional contact.
Related Work
Prior work has advanced controllable video generation and leveraged egocentric human videos for robot learning through cross-embodiment transfer and robot-oriented data construction. However, existing benchmarks cover only subsets of the source-relative requirements for faithful human-to-robot video transfer.
- Controllable video generators increasingly model temporal coherence, visual consistency, interaction, object dynamics, state transitions, and interaction outcomes.
- Egocentric human videos offer scalable supervision about task goals, object affordances, hand–object interactions, and state changes for cross-embodiment transfer.
- Prior transfer methods use visual representations, reward learning, affordance, trajectory, latent-action abstractions, and visual translation, editing, rendering, retargeting, or generative modeling.
- Human-to-robot transfer requires faithful source-relative realization of task goals, action events, functional contact, object response, and target embodiment, whereas existing benchmarks evaluate only subsets.
H2R-Bench
H2R-Bench evaluates whether models can transform egocentric human manipulation demonstrations into videos of the same task performed by a specified robot embodiment. It combines source-grounded transfer cases with five complementary metrics covering task completion, functional interaction, embodiment correctness, and general video quality.
- Benchmark formulation: Each instance pairs an egocentric human manipulation video with a text-specified target embodiment, requiring preservation of the task goal, actions, scene, objects, and visible interactions.The source video is treated as evidence of what happened rather than merely as a visual style reference.
- Dataset: The benchmark contains 240 transfer cases from 120 EgoDex test clips, pairing each source with a parallel-jaw gripper and a dexterous hand across six manipulation families.Source clips are selected for visible task-relevant entities, interactions, and state changes.
- Evaluation protocol: All models use identical source cases, embodiments, task specifications, and evaluation criteria through their strongest documented source-conditioning interfaces.Video-conditioned models receive full clips, while image-conditioned models receive ordered frames within interface limits.
- Evaluation protocol: The main setting specifies the target robot only in text, excludes target-robot reference images, and evaluates every model with uniformly sampled frames under the same evidence budget.Robot reference images are studied separately in an ablation, while generated videos retain native resolution, duration, and frame rate.
- Metrics: M1–M4 assess goal-state completion, action-event completion, functional contact transfer, and embodiment correctness, while M5 measures task-agnostic video quality.M1–M4 use three independent MLLM judges with a shared 0–4 rubric; M5 combines imaging, aesthetics, temporal stability, and motion smoothness.
- Metrics: The primary score weights contact transfer and embodiment correctness at 0.30 each, goal-state and action completion at 0.15 each, and video quality at 0.10.Contact and embodiment together account for 60% because valid transfer requires functional interaction and the requested robot morphology.
Experiments
Experiments evaluate 11 video generators on 240 matched human-to-robot transfer cases across six manipulation families and two target embodiments. Results show that embodiment and functional-contact transfer, rather than generic visual quality alone, determine success, with automated rankings closely aligned with human judgments.
- Evaluation setup: The benchmark evaluates 11 models on 240 transfer cases from 120 human videos, evenly distributed across six manipulation families and paired with both target embodiments.All models receive the same task and embodiment information through their native source-conditioning interfaces.
- Overall results: Seedance 2.0 ranks first, reaching H2RCore scores of 77.3 for the Parallel-Jaw Gripper and 84.6 for the Dexterous Hand.Wan2.7 follows with 76.5 and 83.1, while Kling-V3 scores 74.5 and 81.7.
- Metric diagnosis: Task recognition and visual polish do not ensure robot transfer: Veo 3.1 has M1 0.725 and M2 0.816, while HunyuanVideo 1.5-I2V leads M5 at 0.806 and 0.808 but remains weak on contact and embodiment.Veo 3.1’s M4 scores are 0.100 and 0.227; HunyuanVideo’s contact score stays near 0.185 and embodiment score near zero.
- Target embodiment: The Dexterous Hand improves transfer for 9/11 models, with average H2RCore gain +3.3 and contact-transfer gain +0.055 for all 11 models.Embodiment correctness increases by +0.047 for eight models, although Grok Imagine Video and Mitty-EPIC14B score lower with the hand.
- Reference images: Target-robot reference images are model-dependent: Wan2.7’s gripper H2RCore rises from 76.5 to 83.1, whereas Kling-V3 loses 13.5 points and Seedance loses 8.8.For Wan2.7, gripper contact transfer rises from 0.766 to 0.871 and embodiment correctness from 0.772 to 0.875.
- Evaluation validity: Video Quality correlates weakly with transfer: it spans 0.73–0.81, H2RCore spans 30.0–84.6, and their Spearman association is ρ = 0.14.Human and MLLM rankings nevertheless achieve macro-average Spearman correlation ρ = 0.883 across M1–M4.
Conclusion
H2R-Bench evaluates whether video world models can transfer egocentric human manipulation demonstrations into videos of specified robot embodiments. It is designed to help bridge the human-to-robot embodiment gap and transform abundant human observations into reliable robot-centric training resources.
- Benchmark scope: H2R-Bench evaluates human-to-robot transfer across six manipulation families, two target embodiments, and five complementary dimensions.The dimensions include task goals, action dynamics, functional contact, embodiment consistency, and video quality.
- Research objective: The benchmark targets video world models that transform abundant human manipulation observations into reliable robot-centric training resources.This goal is framed as bridging the human-to-robot embodiment gap.
A Dataset Construction and Annotation Quality Control … B.1 Diagnostic Failure Rates
H2R-Bench constructs a source-grounded, cross-embodiment benchmark from balanced human manipulation clips, with manually verified annotations and explicit validity controls. Its diagnostics show frequent failures in embodiment consistency, functional contact, and required action execution.
- A.1 Source Selection and Task Stratification: The benchmark selects EgoDex test clips across six physical-state manipulation families, retaining only cases with identifiable entities, state transitions, and visual interaction evidence.Parallel-jaw gripper and dexterous-hand conditions use the same human source evidence.
- A.1 Source Selection and Task Stratification: Its taxonomy groups manipulation tasks by the physical state change defining completion, rather than inheriting EgoDex activity labels.The evaluation dimensions are designed around evidence that the requested state change was reproduced.
- A.2 Clip Grounding and Structured Annotation: Qwen3.7-Plus generates structured annotations from each 5-second clip and 32 uniformly sampled frames, preserving contacts, support, release, and expected object responses.Every annotation is manually checked against the source video, and corrections are incorporated before evaluation.
- A.3 Annotation Acceptance Criteria: Cases are accepted only when fixed-schema fields for task family, goals, action events, interactions, and embodiment strategies are complete and internally consistent.This acceptance stage determines whether the source clip provides sufficient evidence for benchmark evaluation.
- A.3 Annotation Acceptance Criteria: Clips with occlusion, ambiguous object states, or unclear interactions are excluded when goals or required contacts cannot be assessed reliably.Minor ambiguities in accepted cases are recorded as notes without changing the official protocol or benchmark set.
- B.1 Diagnostic Failure Rates: On the 120-source main evaluation set, diagnostic failures are recorded when judge-averaged M2–M4 components fall below 0.5, with non-exclusive categories.One generation may simultaneously violate embodiment, contact, and action requirements.
- B.1 Diagnostic Failure Rates: Video-conditioned systems show lower observed failure rates, yet frequent failures remain in human-led manipulation, visible contact support, contact-region or mode correctness, end-effector matching, and required action events.The descriptive comparison cannot identify an interface effect because model and interface vary together; M3 independently treats source-scene or task-entity substitution as a hard failure.
B.2 Matched Source-Conditioning Ablation
The ablation tests whether nine sparse, chronologically ordered source frames can substitute for Seedance 2.0’s native video conditioning across matched human-to-robot transfer cases.
- Experimental setup: The study compares native video conditioning with nine uniformly sampled chronological frames across 24 matched sources, evaluating both target embodiments for 48 transfer cases.The sources include four examples from each task family in the main evaluation set.
B.3 Performance Across Task Families · B.4 Attribute-Based Analysis · B.5 Statistical Reporting
Performance varies systematically across manipulation families and task attributes, while statistical analyses show that reported comparisons are robust to bootstrap uncertainty and alternative metric weights. Deformable-object configuration is strongest on average, and aggregate rankings remain highly stable across weighting schemes.
- B.3 Performance Across Task Families: Deformable-object configuration (F4) is the strongest manipulation family for both target embodiments when averaged over models.The comparison uses the 120-source main evaluation set and separates results by manipulation family and target embodiment.
- B.3 Performance Across Task Families: The six families cover rigid transport, mechanism actuation, insertion and assembly, deformable configuration change, bulk-material transfer, and surface transformation.Parallel-Jaw Gripper and Dexterous Hand results are reported in separate column groups.
- B.4 Attribute-Based Analysis: Family-level averages do not reveal whether difficulty is linked to the interaction interface or the number of required actions.Figure S3 instead groups evaluations by direct manipulation versus tool use and by two–three versus four or more required events.
- B.5 Statistical Reporting: H2RCore percentile 95% confidence intervals are reported on the 120-source main evaluation set.The intervals are obtained through stratified bootstrap resampling of source tasks within each task family.
- B.5 Statistical Reporting: All displayed paired intervals exclude zero under the stated bootstrap procedure.The smallest separation is Seedance–Wan2.7 for the Parallel-Jaw Gripper target.
- B.5 Statistical Reporting: Kendall rank correlation with the default ranking ranges from 0.927 to 1.000 across alternative weighting schemes.The tested schemes include equal weights, a larger quality weight, and a transfer-only variant, indicating stable aggregate ordering.
B.6 Detailed Evaluation Protocol · B.7 Human–Automatic Agreement · B.8 Inter-Judge Agreement
The protocol evaluates generated videos with structured, frame-grounded MLLM judgments across task completion, contact transfer, embodiment correctness, and video quality. Human–automatic correlations are strong, while inter-judge agreement is assessed using pairwise Pearson correlations on the main evaluation set.
- B.6 Detailed Evaluation Protocol: Three MLLM judges inspect 25 uniformly sampled frames, return structured scores with evidence indices and rationales, and contribute equally to M1–M4.The judges are Gemini 3.5 Flash, Qwen3.7-Plus, and GPT-5.4, all run at temperature 0.
- B.6 Detailed Evaluation Protocol: M1 scores weighted final-state predicates from 0 to 4, while M2 scores weighted required action events and reports coverage at scores ≥3 and =4.These metrics distinguish clearly completed outcomes from partial completion and summarize visible event completion across the generated sequence.
- B.6 Detailed Evaluation Protocol: M3 evaluates source-grounded contact and object-response transfer across five dimensions, assigning zero credit when the generated scene or task entities are substantially substituted.The dimensions include contact-region transfer, contact establishment, manipulation-mode transfer, temporally supported object response, and embodiment-compatible strategy.
- B.6 Detailed Evaluation Protocol: M4 scores robotic presence, human-hand absence, embodiment category, end-effector correctness, and temporal structural consistency, with hard failures receiving zero.The end-effector distinction is deliberately weighted because parallel-jaw grippers and dexterous hands are central benchmark embodiments.
- B.6 Detailed Evaluation Protocol: M5 is task-agnostic and measures aesthetic quality, temporal stability, and motion smoothness without rewarding incorrect manipulation or embodiment.It uses decoded-frame image metrics, normalized LAION aesthetic prediction, consecutive-frame differences, and AMT-S interpolation quality.
- B.6 Detailed Evaluation Protocol: H2RCore combines goal, action, contact, embodiment, and video components with weights 0.15, 0.15, 0.30, 0.30, and 0.10, prioritizing contact and embodiment.Human-agreement analyses omit M5 and use the transfer subscore 100(0.20Sgoal + 0.20Saction + 0.30Scontact + 0.30Semb).
- B.7 Human–Automatic Agreement: 0.930 is the Pearson correlation between combined human and averaged MLLM transfer subscores, with metric-wise correlations of 0.791 for M1, 0.818 for M2, 0.880 for M3, and 0.877 for M4.Three human raters scored M1–M4 on the same random sample of 660 videos, using records with complete paired evaluations.
- B.8 Inter-Judge Agreement: Inter-judge agreement is measured as pairwise Pearson correlation between Gemini, Qwen, and GPT per-video scores for M1–M4 on the 120-source main evaluation set.M5 is excluded because it does not use MLLM judges.
C Model Descriptions and Implementation Setups … C.6 Prompt Examples
H2R-Bench evaluates 11 video generators using each model’s strongest documented source-conditioning interface, shared task-and-embodiment prompts, and native outputs. Generation inputs vary across video- and image-conditioned systems, but evaluation uses uniformly sampled evidence and excludes source annotations from generation.
- C Model Descriptions and Implementation Setups: H2R-Bench evaluates 11 representative video generators through publicly available hosted interfaces or local implementations using aligned, case-specific English prompts.The main comparison uses each model’s strongest documented source-conditioning interface.
- C.1 Computing Infrastructure: The local pipeline ran on Ubuntu 24.04.2 LTS with two 16-core Intel Xeon Gold 6144 processors, 251 GiB memory, and four NVIDIA RTX 4090 GPUs.The environment used NVIDIA driver 580.173.02, CUDA 13.0, cuDNN 9.19.0, Python 3.10.16, and PyTorch 2.11.0.
- C.2 Commercially Hosted Models: Commercial systems use native source-conditioning interfaces: video-conditioned models receive complete clips, while Veo 3.1 and Grok Imagine Video receive temporally ordered frames without robot reference images.The main setting specifies target morphology through the common prompt rather than a target-robot image.
- C.4 Visual Inputs and Output Profiles: Frame-conditioned interfaces preserve chronological order, using the first frame for single-image inputs and uniformly spaced endpoint-inclusive frames up to each interface’s budget.Veo 3.1 receives first and last frames, whereas Grok Imagine Video receives up to seven uniformly spaced frames.
- C.5 Output Handling: Outputs remain in native generation formats, while every input video is uniformly sampled into 25 frames for M1–M4 and M5 is computed on the generated video.This fixes the evaluation evidence budget across models while preserving interface differences at generation time.
- C.6 Prompt Examples: Prompts combine native visual conditioning with shared task, scene, contact, and embodiment constraints, while source annotations remain evaluation-only.Complete case prompts, annotation templates, judge prompts, and JSON schemas are included in the benchmark release.
D Limitations
H2R-Bench measures visible evidence of human-to-robot transfer rather than physical executability or downstream policy performance. Its scope is limited by the available sources, embodiments, short clips, and native-interface comparison.
- H2R-Bench evaluates visible transfer evidence, not physical executability or downstream policy performance.
- The benchmark’s 120 EgoDex sources and two target embodiments cover only part of manipulation-setting variation.
- Its current benchmark is limited to short clips, while longer demonstrations may require efficient spatiotemporal modeling.The passage identifies Yang et al. 2025 in connection with this potential extension.
- The native-interface comparison combines model capability with differences in source conditioning.
E Human Evaluation Interface
The human-evaluation interface presents source and generated robot videos side by side for embodiment verification and source-grounded M1–M4 scoring. Evaluation uses balanced ten-video task groups, a common 0–4 scale, separate annotator records, and hides automatic MLLM judgments.
- Evaluation setup: Each task group contains five shared source scenes from one model and task family under both target embodiments, yielding ten videos.The sidebar records task assignment and progress over balanced ten-video groups.
- Evaluation setup: Annotators inspect source and generated clips side by side, verify the requested embodiment, and score source-derived M1–M4 criteria on a common 0–4 scale.The interface presents the requested embodiment and expandable M1–M4 criteria.
- Scoring procedure: Scores are stored separately for each annotator, while automatic MLLM judgments are not displayed during human evaluation.The interface is designed to collect human M1–M4 scores independently of automatic judgments.