Source-linked AI summary
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
Simple AI, :, Yuteng Wei, Jinming Ma, Jiawei Wang, Weitao Zhou, Yushen Zuo, Ke Rui, Minglei Li, Jinhao Zhang, Zhikang Pan, Xiang Wang, Haoran Jia, Huan Du, Zicheng Zeng, Jun Ma, Guiyu Qin, Di Zhang, Xiaofei Li
TL;DR
Deployable manipulation policies need data that is both high-fidelity and scalable, but robot-free demonstrations have not been shown to replace real-robot post-training. HiFi-UMI demonstrates that policies post-trained solely on high-fidelity robot-free data match teleoperation across three backbones, with gaps within roughly 3 percentage points.
Problem
Robot-free manipulation data scales readily but has not been shown to eliminate the real-robot data traditionally used for deployment-oriented post-training.
Method
HiFi-UMI co-designs a portable UMI capture and validation system to produce accurate, synchronized, wide-field robot-executable demonstrations for policy training.
Results
−2.5, +3.1, and −0.6 percentage points: HiFi-UMI-only post-training matches in-domain teleoperation across three VLA and WAM backbones, reaching 85% on precision insertion.
Takeaways & Limitations
High-fidelity robot-free data can support deployable manipulation policies beyond a pre-training role.
Takeaways & Limitations
The zero-robot post-training evidence covers four tabletop bimanual tasks and three backbones, leaving generality to other tasks, embodiments, and distribution shifts untested.
Abstract
from arXiv · showhide
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the teleoperation baseline is collected in the evaluation scene and no HiFi-UMI trajectory is. Pre-training on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by a further 18.1 percentage points. We open-source HiFi-UMI-2K, 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations, each automatically reconstructed and validated through simulation replay, as a large-scale, high-fidelity resource for the robot-learning community.
1 Introduction
HiFi-UMI addresses the fidelity limitations that have confined robot-free UMI data to pre-training by designing a portable, high-fidelity, replayable capture and curation system. Its demonstrations support zero-robot post-training that matches in-domain teleoperation across three backbones, while pre-training on the same corpus further improves transfer and efficiency.
- Motivation: Robot-free UMI capture is scalable and inexpensive, but pose reconstruction, synchronization, and limited field of view restrict its reliability for deployable policies.These deficiencies have led the field to reserve robot-free data mainly for pre-training while retaining real-robot teleoperation for post-training.
- System: HiFi-UMI co-designs head-mounted stereo SLAM, native inter-gripper pose, shared GPIO triggering, and wide-angle sensing for 3 mm accuracy and microsecond alignment without external tracking.An automated engine reconstructs, replays, and validates demonstrations, retaining 96% of raw captures as robot-executable data.
- Zero-robot post-training: −2.5, +3.1, and −0.6 percentage points are the success-rate differences between HiFi-UMI-only post-training and in-domain teleoperation across three VLA and WAM backbones.Parity holds even though teleoperation data comes from the evaluation scene while no HiFi-UMI trajectory does, establishing zero-robot post-training under a baseline-favoring asymmetry.
- Resource: HiFi-UMI-2K is an open 2,000-hour subset of microsecond-synchronized, replayable, ultra-wide-FoV demonstrations produced for deployment-grade post-training without real-robot teleoperation.The same HiFi-UMI corpus supplies both pre-training and post-training for policies deployed directly on a real bimanual robot.
- Pre-training: 41% lower offline action error on ten unseen tasks and 18.1 percentage points higher real-robot success result from 4,000-hour pre-training on StarVLA-QwenPI.At matched post-training data, pre-training matches the scratch-initialized baseline with a quarter of the task data.
2 Related Work
Related work spans grounded teleoperation, scalable but weakly grounded video, and UMI-style robot-free demonstrations as an intermediate tier. HiFi-UMI targets this middle tier by addressing capture fidelity and testing robot-free demonstrations as deployment-relevant supervision.
- Data landscape: UMI-style data provides low-cost, action-grounded demonstrations without a robot body, occupying a middle tier between scalable human video and embodiment-specific teleoperation.FastUMI-100K and RDT2 show that this recipe can scale to corpora rivaling large teleoperation efforts.
- Data landscape: UMI demonstrations preserve actionable wrist-view geometry and relative end-effector motion without robot-specific collection, but fidelity, synchronization, retargetability, and geometric consistency remain limiting factors.These limitations motivate using higher-fidelity UMI data for deployment-relevant supervision rather than only scalable pre-training.
- Capture systems: The original UMI uses ORB-SLAM3, an IMU, one 155° fisheye camera per gripper, and side mirrors, while later systems add external headset or room-based tracking infrastructure.AirExo-2 attributes handheld limitations to visual-SLAM inaccuracies and restricted field of view; UMI also reports SLAM and scale-ambiguity failures.
- Capture systems: HiFi-UMI addresses these limitations with offline stereo SLAM, a shared GPIO hardware trigger, and non-parallel cameras for ultra-wide coverage.The paper states that prior handheld systems had not compared robot-free post-training against teleoperation on the same robot under fixed backbone, recipe, and deployment stack.
- Policy backbones: Manipulation foundation models comprise reactive VLAs and predictive WAMs, while recent systems increasingly combine continuous action heads, flow matching, video generation, or diffusion backbones.The selected backbones differ architecturally but consume the same supervision interface, enabling the paper’s comparison across model families.
3 Data Collection and Processing Pipeline
HiFi-UMI combines four fidelity-focused subsystems with hard data-validity gates and a six-stage processing pipeline to produce deployment-grade robot-free demonstrations. The processed data achieves 3 mm end-effector accuracy, sub-40 µs timing offsets, and approximately 96% cumulative basic-validity yield.
- System design: HiFi-UMI targets four jointly sufficient requirements: accurate scalable pose acquisition, manipulation-oriented gripper morphology, wide-coverage multimodal sensing, and online quality control.The system is organized around these requirements, with a six-stage pipeline converting raw captures into training-ready episodes.
- Pose acquisition: 3 mm end-effector accuracy is achieved using head-mounted stereo-inertial SLAM and head-relative fiducial-marker localization for globally consistent bimanual trajectories.Both hands are localized relative to the head in one camera frame, so inter-gripper relative pose is measured natively and accumulated drift is reduced.
- Multimodal sensing: About 200° camera coverage per hand reduces occlusion and improves gripper observability, while unified triggering synchronizes the six-camera, multi-IMU system.Each hand carries two non-parallel fisheye cameras, alongside stereo head cameras, hand and head IMUs, and gripper encoders.
- Quality control: 98% of reconstructed trajectories pass simulation replay validation, and the two serial gates yield approximately 96% cumulative basic-validity of raw captures.Replay discards trajectories that are kinematically or dynamically infeasible; reconstruction and replay validation each pass approximately 98% of captures reaching that stage.
- End-to-end fidelity: 3 mm end-effector accuracy, cross-sensor timing offsets below 40 µs, fewer than two dropped frames per hour, 98% trajectory-reconstruction success, and gripper-state error below 0.1° characterize the processed data.The 3 mm mean translational error is measured against base-station tracking ground truth over approximately 2 m of accumulated head-trajectory length, used only for evaluation.
4 Dataset and Release
HiFi-UMI provides a large, high-fidelity manipulation corpus with synchronized multimodal trajectories and reproducible composition tracking. HiFi-UMI-2K is openly released under CC BY 4.0 with faces masked from recordings.
- Dataset: Over 20,000 hours span more than 4.32 million episodes across 480+ scenes, captured with six cameras and exported with synchronized multimodal annotations.Each episode includes multi-view video, calibrated bimanual trajectories, gripper states, language annotations, and subtask boundaries.
- Dataset: Every retained episode satisfies collection-time and processing-time fidelity requirements, while corpus composition and experimental subsets remain measured, reproducible, and traceable.The corpus is characterized across tasks, scenes, objects, and manipulation attributes, with subsets linked to batch, device, and review history.
- Release: CC BY 4.0 permits redistribution, derivative use, and commercial use with attribution, and all released human faces are masked.HiFi-UMI-2K therefore contains no facial imagery.
5 Baselines and Training Setup · 5.1 StarVLA-QwenPI
The study evaluates high-fidelity UMI data across three standardized foundation-policy baselines, including the Qwen-based StarVLA-QwenPI implementation. StarVLA-QwenPI augments its π-style action head with action-side self-attention and uses receding-horizon inference over 20-step chunks.
- 5 Baselines and Training Setup: Three foundation-policy baselines—StarVLA-QwenPI, OpenPI-π0.5, and LingBot-VA—are evaluated under standardized tasks, action semantics, initialization, optimization, interfaces, and safety conditions.Agreement across backbones is intended as convergent evidence about the data rather than a pooled architecture comparison.
- 5 Baselines and Training Setup: All backbones use four wrist-camera views as policy inputs, while the head-mounted stereo pair is reserved for capture-time trajectory reconstruction.UMI and teleoperation variants also share prompts, camera selection, temporal offsets, tensor layout, normalization, and evaluation protocol within each backbone.
- 5 Baselines and Training Setup: All three policies share active bimanual end-effector semantics: 10 physical channels per arm and 20 channels for bimanual tasks.Translation is anchor-frame relative, orientation uses Rotation6D, and gripper targets remain absolute; tensorization remains backbone-specific.
- 5.1 StarVLA-QwenPI: StarVLA-QwenPI combines Qwen3-VL-4B-Instruct, with 36 transformer layers and hidden width 2,560, with a π-style conditional flow-matching DiT action head.The DiT blocks cross-attend to corresponding layer-wise Qwen features encoding multi-view observations and instructions.
- 5.1 StarVLA-QwenPI: The modified QwenPI path retains cross-attention to all 36 Qwen layers and adds attention-only self-attention residuals after every odd-numbered DiT block.This couples H = 20 action steps without repeating feed-forward computation, addressing the official path’s lack of action-side self-attention.
- 5.1 StarVLA-QwenPI: 8 explicit-Euler steps integrate the learned vector field at inference, producing H = 20 actions while the robot executes Hexec = 10 receding-horizon actions.Images are resized to 224×224, and action dimensions use training-split statistics for normalization.
5.2 OpenPI-π0.5 · 5.3 LingBot-VA
OpenPI-π0.5 combines visual-language modeling with a continuous action expert, while LingBot-VA predicts future visual states and recovers actions through inverse dynamics. Both baselines use chunked action prediction, with distinct conditioning, flow-matching, and deployment conventions.
- 5.2 OpenPI-π0.5: OpenPI-π0.5 combines a SigLIP So400m/14 visual encoder, an 18-layer Gemma-2B language model, and an 18-layer Gemma-300M action expert.Its streams use modality-specific parameters and Mixture-of-Transformers joint attention.
- 5.2 OpenPI-π0.5: OpenPI-π0.5 serializes discretized proprioceptive state with the instruction, while action tokens attend jointly to the full prefix and one another.The visual-language prefix contains images and the tokenized task instruction.
- 5.2 OpenPI-π0.5: OpenPI-π0.5 uses continuous flow matching by interpolating action chunks with Gaussian noise over flow time.The conditioning context is c = (o, q, ℓ), with xt = (1 − t)a + tϵ.
- 5.2 OpenPI-π0.5: OpenPI-π0.5 uses a sole fine-tuning objective and omits knowledge insulation, autoregressive subtask generation, and auxiliary text loss.Inference generates continuous action chunks with 10 Euler steps for receding-horizon execution.
- 5.3 LingBot-VA: LingBot-VA is a causal WAM that predicts future visual states before recovering their realizing actions from synchronized multi-view video-action history.It represents observations with VAE latents and conditions on task instruction ℓ.
- 5.3 LingBot-VA: LingBot-VA jointly trains future video-latent prediction and continuous action recovery with equal-weight video and action losses.A block-causal mask orders interleaved video-action chunks, while the first factor acts as inverse dynamics.
- 5.3 LingBot-VA: LingBot-VA inserts 20 active bimanual dimensions into a native 30-dimensional action tensor through a fixed channel map.Unused channels are masked from action loss, Rotation6D uses the first two rotation-matrix rows, and condition-specific normalization is reused at inference.
- 5.3 LingBot-VA: LingBot-VA deployment uses bounded rolling KV caching, guidance scales of 5/1 and 8/16 denoising steps, and an attention window of 24.It follows the WAM protocol in Sec. 5.7.2 for receding-horizon prediction.
5.4 Training Variants · 5.5 Data Processing and Episode Filtering · 5.6 Training and Optimization
The experiments compare real-robot and UMI-only training under matched architectures and deployment protocols, while processing UMI trajectories into backbone-native actions and filtering invalid episodes. Training uses backbone-specific optimization schedules, and checkpoints are selected using validation loss, action and gripper errors, and real-robot rollouts before blind evaluation.
- 5.4 Training Variants: UMI post-training uses only high-fidelity UMI demonstrations, directly testing whether robot-free data can replace the conventional real-robot anchor.The real-robot reference uses task-specific teleoperation demonstrations collected directly on the target robot.
- 5.4 Training Variants: UMI pre-training plus UMI post-training uses robot-free data in both stages and is instantiated only on StarVLA-QwenPI.This variant tests whether broad UMI pre-training provides reusable visual-motor priors for downstream UMI-only specialization.
- 5.4 Training Variants: Matched backbones, tensorization, observation interfaces, normalization, sampling, and controllers make data source and training stage the primary experimental differences.The comparison specifically tests whether high-fidelity UMI can replace real-robot post-training data.
- 5.5 Data Processing and Episode Filtering: UMI trajectories are converted into deployment-robot end-effector actions, packed into 20 active physical channels, and aligned with teleoperation tensor layouts.Hand poses are transformed from a shared world frame into robot end-effector frames, with backbone-specific future offsets and relative targets.
- 5.5 Data Processing and Episode Filtering: Episodes are removed for tracking, frame, timestamp, pose, action, execution, or gripper-state failures, while long episodes are sliced into coherent segments.The source-validity checks rarely trigger on the already high-fidelity UMI corpus.
- 5.5 Data Processing and Episode Filtering: Position, rotation, and gripper channels are normalized separately, with separate left- and right-arm statistics for bimanual tasks.Each condition uses statistics tied to its deployment profile under the backbone’s unchanged normalization convention.
- 5.6 Training and Optimization: 4,000 hours of multi-task UMI data correspond to 180,000 optimization steps, using an effective global batch size of 2,048 and a 3,000-step warm-up.StarVLA-QwenPI is optimized end to end with AdamW in mixed precision, gradient clipping at 1.0, and no accumulation.
- 5.6 Training and Optimization: Checkpoint selection combines held-out flow-matching loss, de-normalized per-dimension action error, gripper error, and real-robot rollout metrics before blind evaluation.Checkpoints are chosen with a fixed validation protocol rather than validation loss alone, then frozen for final evaluation.
5.7 Deployment Protocol
Deployment comparisons standardize task, robot, safety, workspace, velocity, and evaluation settings while preserving each backbone’s native execution interface. VLA policies use timestamped receding-horizon control, whereas LingBot-VA uses native block-causal streaming.
- Matched evaluation settings: Matched deployment settings standardize task definitions, initial states, robot action semantics, safety, workspace and velocity limits, and success criteria within each backbone.VLA and WAM backbones retain their native control cadence and chunk-consumption rules.
- VLA execution: VLA deployment aligns sensor streams to a common observation time and independently restores each predicted pose row from the synchronized query-time end-effector pose.This follows UMI’s latency-matching principle rather than recursively integrating targets from preceding commands.
- VLA execution: H = 20 and Hexec = 10 configure StarVLA-QwenPI, while OpenPI-π0.5 retains its native horizon and replanning interval across both post-training variants.Predicted targets receive backbone-specific timestamps and enter a latency-compensated command buffer.
- LingBot-VA execution: 12 actions execute after reset and 24 actions thereafter under LingBot-VA’s identical schedule for UMI- and teleoperation-post-trained variants.The protocol measures current end-effector poses at each chunk start and updates context with synchronized observations during execution.
6 Experiments
The experiments test whether high-fidelity UMI data can produce directly deployable real-robot manipulation policies without real-robot teleoperation data. They compare three policy backbones spanning VLA and WAM foundation-policy families to isolate data-source effects from architecture.
- Experimental question: The central experiment asks whether high-fidelity UMI data alone can support directly deployable real-robot manipulation policies.The comparison specifically removes reliance on real-robot teleoperation data as a post-training source.
- Experimental design: The study separates data-source effects from model architecture by evaluating three policy backbones.The backbones span both foundation-policy families: VLA and WAM.
- Experimental design: The evaluated policies are StarVLA-QwenPI, OpenPI-π0.5, and LingBot-VA.StarVLA-QwenPI and OpenPI-π0.5 are VLA policies, while LingBot-VA is a WAM.
6.1 Experimental Design
The study evaluates policies on a four-task stationary bimanual platform with physically matched capture and deployment interfaces, standardized rollout procedures, and controlled training-data comparisons. It contrasts HiFi-UMI and real-robot teleoperation data across three backbones, while isolating initialization effects for StarVLA-QwenPI.
- Platform and execution: The stationary bimanual platform uses two seven-joint force-controlled arms with the same gripper and four wrist cameras as HiFi-UMI, leaving only arm kinematics mismatched.Policies output end-effector pose targets at 125 Hz; inverse kinematics runs at 125 Hz and joint commands stream over EtherCAT at 1 kHz.
- Benchmark tasks: The benchmark covers four tasks spanning sustained contact and coverage, bimanual deformable-object coordination, precise constrained placement, and semantic category-conditioned manipulation.The tasks are Stain Wiping, Shirt Folding, Remote Insertion, and Produce Sorting.
- Data collection: 3,200 UMI trajectories per task provide approximately 10–20 hours of demonstrations, versus approximately 300 real-robot teleoperation trajectories requiring approximately 3–7 hours per task.UMI collection uses multiple operators across sites and conditions, while teleoperation occurs on the target robot in the evaluation environment and requires an operator plus assistant.
- Evaluation setting: No UMI trajectory comes from the evaluation scene, creating shifts in background, illumination, tabletop appearance, and visual context that teleoperation data does not face.This asymmetry tests transfer from diverse, multi-site UMI data to an unseen deployment scene.
- Training conditions: The study varies backbone and task-specific post-training source across all three backbones, while varying initialization only for StarVLA-QwenPI.Data-source comparisons change only demonstration origin; C1 versus C7 holds HiFi-UMI post-training fixed and changes only whether HiFi-UMI pre-training precedes it.
6.2 Can UMI-Only Post-Training Match Teleoperation-Based Post-Training?
Across VLA and WAM evaluations, HiFi-UMI-only post-training approximately matches teleoperation-based post-training without systematically reducing real-robot performance. Scaling UMI demonstrations improves low-data performance to 85.0% at 3,200 demonstrations, while task-level differences remain small and variable.
- Aggregate comparison: −2.5, +3.1, and −0.6 percentage points are the UMI-minus-teleoperation aggregate success-rate differences for StarVLA-QwenPI, OpenPI-π0.5, and LingBot-VA, respectively.The differences have no consistent direction across backbones.
- Aggregate comparison: 64.4% versus 64.1% aggregate success shows that UMI-only post-training remains comparable to teleoperation across the two VLA backbones.Within-backbone comparisons remain primary, and replacing teleoperation demonstrations does not systematically reduce performance.
- Interpretation: 3,200 UMI trajectories versus 300 teleoperation trajectories means the comparison evaluates practical data-production pipelines rather than equal sample sizes.The results support using high-fidelity UMI as the sole task-specific post-training source without a real-robot teleoperation anchor.
- Data scaling: 85.0% success is reached with 3,200 UMI demonstrations, up from 37.5% with 400 and 65.0% with 800 demonstrations.Further scaling to 1,600 demonstrations reaches 70.0%, while increasing from 3,200 to 6,400 episodes slightly decreases success from 85.0% to 82.5%.
- Aggregate comparison: 56.9% versus 57.5% aggregate success shows comparable WAM performance for UMI-only and teleoperation post-training.Under the separate WAM evaluation protocol, replacing teleoperation demonstrations does not result in systematic degradation in closed-loop deployment performance.
6.3 Does Large-Scale UMI Pre-Training Yield a Better Base Model?
Large-scale StarVLA-QwenPI pre-training on 4,000 hours of UMI data improves held-out action prediction, transfers to ten unseen tasks, and yields stronger real-robot post-training. Transfer varies by interaction family, reflecting uneven representation in the pre-training mixture.
- Scaling on held-out data: 4,000 hours of multitask UMI data are used to pre-train StarVLA-QwenPI, with fixed held-out chunks and Euler integration isolating training exposure.The mixture spans diverse scenes, objects, and manipulation skills; one pass ends at 180k steps.
- Scaling on held-out data: 61%: held-out action error falls over one pass through the corpus, following a pre-decay exposure-scaling fit with α = 0.268 and R2 = 0.993.The fit measures cumulative UMI action chunks processed globally and is not attributed to the final learning-rate decay.
- Transfer to unseen tasks: 41%: mean OOD error decreases across ten unseen tasks, and every task improves, with a smaller aggregate scaling exponent of α = 0.095.Utensil and tableware interactions improve fastest, granular transfer is intermediate, and cloth folding improves more slowly.
- Transfer to unseen tasks: Interaction-family transfer differs: rigid utensil-to-receptacle tasks reach the lowest final error, while garment folding remains most difficult.Object and receptacle placement exceeds one third of pre-training frames, whereas textile folding is below one percent; the analysis motivates broader deformable and granular-material coverage.
- Benefits for post-training: 18.1 percentage points: UMI pre-training raises aggregate StarVLA-QwenPI success across four benchmark tasks when initialization is the only controlled difference.The comparison uses identical 3,200 task-specific trajectories and post-training protocols; gains are particularly strong on wiping, folding, and insertion.
7 Discussion
HiFi-UMI achieves approximate aggregate parity with in-domain teleoperation as the sole task-specific post-training source, but the evidence remains limited in scope, statistical resolution, fidelity decomposition, and sample matching.
- Main finding: Across four evaluated tasks and three backbones, HiFi-UMI-only post-training achieves approximate aggregate parity with in-domain teleoperation.OpenPI-π0.5 and LingBot-VA retain publicly released pre-trained initializations, with no target-task teleoperation data introduced in the HiFi-UMI condition.
- Evaluation scope and generality: Zero-robot post-training covers four tabletop bimanual tasks and three backbones under scene-level distribution shift, leaving generality to other tasks, embodiments, and shifts untested.Pre-training scaling and downstream gains were measured only on StarVLA-QwenPI.
- Task-level statistical resolution: 40 rollouts per task–policy pair mean one additional success changes the estimated rate by 2.5 percentage points, limiting fine-grained task-level comparisons.The parity claim is therefore based on aggregate evidence across tasks, while per-task differences are treated as descriptive.
- Fidelity validation: Fidelity is validated jointly through trajectory accuracy, inter-gripper relative pose, synchronization, and field of view, without controlled degradation of individual factors.The study shows that high fidelity suffices but does not establish how much of each property a deployable policy requires.
- Data efficiency and transfer across post-training sources: Without pre-training, UMI-only post-training uses roughly ten times as many demonstrations as the teleoperation baseline, so the comparison is not sample-matched.The comparison reflects practical data-production pipelines rather than per-trajectory efficiency; whether pre-training gains continue as the corpus grows remains unresolved.
8 Conclusion
HiFi-UMI argues that fidelity, rather than the robot-free setting itself, limits demonstrations, and presents a portable capture pipeline addressing trajectory accuracy, relative pose, synchronization, and field of view. Policies trained solely on HiFi-UMI match teleoperation across three backbones, while large-scale pre-training improves unseen-task action error and real-robot success.
- Conclusion: Over 20,000 hours of HiFi-UMI data were generated using portable capture co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view.The conclusion frames fidelity as the limitation of robot-free demonstrations rather than the robot-free setting itself.
- Conclusion: Roughly 3 percentage points separate HiFi-UMI-only post-training from teleoperation across three VLA and WAM backbones, with differences of both signs within sampling noise.This result supports zero-robot post-training without requiring a teleoperated anchor.
- Conclusion: 85% success was reached on precision insertion despite scene shift and zero teleoperated data.The result demonstrates direct real-robot deployment from HiFi-UMI-only post-training under a shifted evaluation scene.
- Conclusion: 41% lower action error on ten unseen tasks and an 18.1 percentage-point real-robot success increase resulted from 4,000 hours of pre-training.The success improvement was measured at matched post-training data.