Source-linked AI summary
InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions
Sirui Xu, Samuel Schulter, Morteza Ziyadi, Xialin He, Xiaohan Fei, Yu-Xiong Wang, Liangyan Gui
TL;DR
Human-object interaction requires controllers that turn sparse intentions into physically coherent whole-body behavior while generalizing across a vast configuration space. InterPrior distills full-reference imitation into a goal-conditioned variational policy and finetunes it with physically perturbed data and reinforcement learning. The resulting controller improves robustness and competence beyond demonstrations, including unseen objects and interactions, while supporting interactive control and potential robot deployment.
Problem
Existing interaction motor priors must scale from sparse high-level goals to feasible whole-body behaviors across the large configuration space of human-object interaction.
Method
InterPrior combines large-scale imitation distillation into a goal-conditioned latent policy with physical perturbations and reinforcement-learning finetuning.
Results
InterPrior maintains natural whole-body coordination while improving robustness and competence, composing skills, recovering from failures, and generalizing to novel objects and interactions.
Takeaways & Limitations
The controller provides a reusable motion prior for diverse human-object interaction skills, with interactive control and embodiment flexibility demonstrated within the reported scope.
Takeaways & Limitations
InterPrior still fails on extremely thin or elongated unseen objects and can partially complete multi-goal chains when canonicalization creates large alignment discrepancies.
Abstract
from arXiv · showhide
Humans rarely plan whole-body interactions with objects at the level of explicit whole-body movements. High-level intentions, such as affordance, define the goal, while coordinated balance, contact, and manipulation can emerge naturally from underlying physical and motor priors. Scaling such priors is key to enabling humanoids to compose and generalize loco-manipulation skills across diverse contexts while maintaining physically coherent whole-body coordination. To this end, we introduce InterPrior, a scalable framework that learns a unified generative controller through large-scale imitation pretraining and post-training by reinforcement learning. InterPrior first distills a full-reference imitation expert into a versatile, goal-conditioned variational policy that reconstructs motion from multimodal observations and high-level intent. While the distilled policy reconstructs training behaviors, it does not generalize reliably due to the vast configuration space of large-scale human-object interactions. To address this, we apply data augmentation with physical perturbations, and then perform reinforcement learning finetuning to improve competence on unseen goals and initializations. Together, these steps consolidate the reconstructed latent skills into a valid manifold, yielding a motion prior that generalizes beyond the training data, e.g., it can incorporate new behaviors such as interactions with unseen objects. We further demonstrate its effectiveness for user-interactive control and its potential for real robot deployment.
1 University of Illinois Urbana-Champaign 2 Amazon † Equal Advising
InterPrior is a goal-conditioned generative controller for simulated humanoid-object interaction, supporting multiple goal types, failure recovery, and operator steering.
- The controller pursues long-horizon snapshot, trajectory, and contact goals in a physics-based simulator.
- InterPrior demonstrates recovery from unsuccessful grasps.
- The system supports steering by a human operator and application to humanoid robot embodiments.
1. Introduction
The paper targets scalable, physics-based human-object interaction control from sparse goals rather than fully specified trajectories. InterPrior combines broad imitation-based skill acquisition with RL finetuning to improve generalization, robustness, and embodiment flexibility.
- Motivation: An interaction motor prior should sample feasible loco-manipulation behaviors from sparse goals instead of mimicking deterministic, fully specified trajectories.
- Motivation: Existing generative controllers can expand motion coverage but are difficult to scale because of unstable optimization, discriminator mode collapse, and handcrafted task objectives.
- InterPrior: InterPrior scales task, skill, motion, and dynamics coverage within one physics-based human-object interaction controller.
- InterPrior: The method distills large-scale demonstrations into a natural goal-conditioned policy, then uses RL as a local optimizer anchored to that pretrained model.
- Contributions: The resulting controller supports failure recovery, goal execution across varied configurations, mid-trajectory command switching, re-grasps, and stability under perturbations.
2. Related Work
Related work progresses from kinematic and scenario-specific interaction animation toward physics-based control and latent motor priors. These approaches improve realism, diversity, or scalability, but remain limited by physical inconsistencies, planning fragility, sample inefficiency, or expert coverage.
- Physics-based interaction animation: Kinematic interaction methods often produce implausible contact drift and interpenetration, partly because datasets contain spatial or physical inconsistencies.
- Physics-based character animation: Physics-based controllers improve scalability through multi-clip trackers and reference planners but remain fragile when planned motions are dynamically unstable.
- Latent motor priors: Latent motor-prior methods distill skills into compact codes using VAEs, universal trackers, masked policies, or diffusion models, yet are often limited by expert coverage.
- Task coverage: Physics-based character control has expanded from simple actions such as striking or sitting to complex sports, games, carrying, and rearrangement tasks.
3. Methodology
InterPrior learns a physics-based human-object interaction policy from high-level goals rather than full references. Its methodology combines expert imitation, variational distillation, and reinforcement-learning post-training to produce diverse, physically simulated behavior and improve robustness beyond demonstrations.
- Task Formulation: The policy conditions on current human-object state, recent history, and high-level goals, then samples control signals to generate physically simulated interaction motions.Goals may come from users, kinematic generators, or motion-capture keypoints.
- Variational Distillation: The expert is distilled into a masked conditional variational policy that maps sparse goals and multimodal observations to a multimodal action distribution.The policy incorporates contact information and uses a structured latent space for skill embeddings.
- Expert Training: InterPrior trains a full-reference imitation expert with augmented data, physical perturbations, and shaped rewards for stable coordination and precise grasping.The expert receives complete future references and is trained with PPO using a composite reward.
- Variational Distillation: Latent sampling preserves behavioral diversity, while projecting latent vectors onto a hypersphere limits rare out-of-distribution draws and stabilizes skill learning.The decoder maps the latent and observation to actions, and an auxiliary head reconstructs masked goal entries during training.
- Post-Training: RL post-training improves the distilled policy by recovering from failures, exploring plausible unseen configurations, and preserving natural behavior from pretraining.A sparse goal reward activates when masked feature distance to the target falls below a threshold, avoiding a dense distance-based reward.
4. Experiments
InterPrior is evaluated across full-reference tracking, sparse goal following, long-horizon compositions, perturbations, and adaptation to novel interactions. Results indicate improved robustness and generalization, while strict reference tracking can retain lower position error.
- Evaluation setup: InterPrior is evaluated on full-reference tracking and sparse goal following across snapshot, trajectory, contact, and composed task specifications.The evaluation also includes long-horizon multi-goal chaining and random-initialization stress tests.
- Full-reference tracking: InterPrior improves success under thin-geometry interactions and initialization noise, while InterMimic retains lower position error through stricter reference tracking.InterPrior may deviate to re-align contact, trading strict tracking precision for interaction completion.
- Goal-conditioned tasks: InterPrior consistently improves goal-conditioned success and reduces errors, with the largest gains on long-horizon chaining and random-initialization tests.RL finetuning trains recovery from diverse initializations and under-covered intermediate states encountered during goal switching.
- Long-horizon behavior: InterPrior sustains minute-long multi-object interactions, transitions smoothly across skills, and self-corrects contact or balance drift.The reported skill transitions include approach, grasp, lift, and reposition.
- Generalization and transfer: A single InterPrior model generalizes zero-shot to unseen objects and interaction styles, including data from BEHAVE and HODome.It also transfers sustained object-conditioned interactions from IsaacGym to MuJoCo.
- Ablation and failure cases: InterPrior recovers from imperfect contact initialization and continues rollouts, but can fail on extremely thin or elongated unseen objects and some chained goals.Large canonicalization discrepancies can lead the policy to prioritize balance over precise goal completion.
5. Conclusion
InterPrior combines large-scale imitation distillation with reinforcement finetuning to produce a physics-based generative controller for human-object interaction.
- InterPrior combines large-scale imitation distillation with reinforcement finetuning in a physics-based generative motion controller.
Supplementary Material
The supplementary material documents the video, simulation setup, goal representation, methodology, implementation, additional experiments, limitations, and societal implications.
- The supplementary material describes the organization of the demo video and the simulation configuration.
- It explains how snapshot, trajectory, and contact goals are constructed using masks.
- It provides formulations for the reference-free hand reward, variational distillation and latent-shaping losses, and reinforcement-learning finetuning.
- Additional implementation details cover network architectures, training schedules, data augmentation, and techniques for G1 sim-to-sim experiments.
- Further experiments include kinematic HOI-generator integration, metric details, and failure cases.
- The supplementary material also examines system limitations and potential societal implications.
- The demo video renders snapshot, trajectory, and contact-conditioned behaviors for diverse object shapes without post-processing beyond camera selection and cropping.
- Experiments use IsaacGym with GPU PhysX, 30 Hz policy control, 60 Hz simulation stepping, and two internal substeps per control step.
C. Goal Formulation
InterPrior represents goals with masked future states and uses stochastic body-wise visibility during training, task-specific masks during inference, and long-horizon context for chained goals despite canonicalization limits.
- C. Goal Formulation: A goal state y_t matches the observation structure, while binary mask m_t marks which goal components are provided to the policy.
- C.1. Horizon for Goals: Short-horizon previews use offsets K = {1, 2, 4, 16}, while long-horizon snapshots use a sampled offset L ∈ [1, 128].
- C.2. Stochastic Mask Sampling during Training: Training masks randomly reveal partial future states rather than tying visibility to snapshot, trajectory, or contact tasks.
- C.2. Stochastic Mask Sampling during Training: Body-wise masking jointly hides or reveals positions, orientations, velocities, interaction vectors, and contact states for each rigid body.
- C.2. Stochastic Mask Sampling during Training: Independent rigid-body visibility follows a Bernoulli process whose masks persist across steps unless a reset occurs.
- C.2. Stochastic Mask Sampling during Training: With p_reset = 0.01, masks usually persist for multiple steps while occasional resets diversify visibility patterns.
- C.3. Task Definition for Inference: At inference, masks remain fixed for each task, except that multi-goal chaining resamples a mask at each sub-goal transition.
- C.3. Task Definition for Inference: Snapshot, trajectory, and contact control reveal different goal subsets, while chained goals use long-horizon context to help compensate for canonicalization artifacts.
D. Additional Details on Methodology
The methodology expands expert training, variational distillation, and reinforcement-learning post-training with rewards and losses that promote contact, stable latent representations, temporal smoothness, and masked-goal reconstruction.
- The framework comprises InterMimic+ expert training, variational distillation, and reinforcement-learning post-training.
- The reference-free hand reward uses fingertip-to-surface bearing alignment and distance-dependent weighting during upcoming interactions.
- This hand reward encourages all five fingers to maximize upcoming surface contact with the object.
- The scale loss regularizes the prior mean toward the unit hypersphere to prevent latent-mean collapse or explosion.
- Temporal consistency penalizes changes between consecutive prior distributions using squared 2-Wasserstein distance between Gaussians.
- Goal reconstruction trains a decoder head to predict masked future goal entries from the latent representation.
- The reconstruction objective encourages the latent to capture intent and context sufficient to infer hidden goal components from visible subsets.
- In practice, the goal reconstruction loss reconstructs the short future with k = 1.
D.3. InterPrior: Post-Training Beyond Reference
Post-training adds new behaviors and preserves the original prior by combining specialized reward shaping with distributed RL and distillation environments.
- Get-Up Training: The get-up behavior uses a new learnable token and an auxiliary reward for episodes initialized from fallen states.The reward targets pelvis elevation and torso reorientation toward an upright configuration.
- Get-Up Training: The get-up reward combines pelvis-height and torso-upright terms using a clipped linear shaping function.The pelvis target height is 0.7, and torso orientation is measured against the world-up direction.
- Distributed Training: Distributed training separates RL environments from distillation environments while synchronously updating shared policy parameters.RL environments optimize the post-training reward, whereas distillation environments retain the ELBO objective and expert supervision.
- Distributed Training: New-dataset states provide additional long-horizon goals and initializations while distillation regularizes the policy toward the original prior.This setup supports incremental acquisition of new object categories and interaction styles without retraining from scratch.
E. Implementation Details
The implementation uses PPO-based expert training and RL finetuning, with MLP-based imitation and distillation components alongside a Transformer prior.
- PPO Setup: Both expert training and RL finetuning use PPO with generalized advantage estimation, a clipped surrogate objective, Adam, and gradient clipping.The implementation keeps the PPO discount, GAE, clipping, and entropy settings specified in Table B.
- Full-Reference Imitation Expert: The InterMimic+ expert uses separate actor and critic MLPs with hidden sizes (1024, 1024, 512).The critic receives full observations and references and outputs a scalar value.
- Variational Distillation: InterPrior’s variational distillation encoder and decoder share an MLP backbone with hidden sizes (1024, 1024, 512).Its prior is a four-layer Transformer encoder with four attention heads and latent dimension 512.
F. Additional Experimental Results
Additional experiments define evaluation procedures, show diverse and adaptive behaviors, and identify limitations involving training coverage, object rigidity, dexterity, and training complexity.
- Evaluation Metrics: Trajectory-following errors are computed on unmasked components, while snapshot tasks report the minimum rollout-to-goal distance.The snapshot metric tests whether the policy reaches the target configuration without a time-aligned reference trajectory.
- Qualitative Comparisons: InterPrior achieves higher success under the same task goal in qualitative comparisons with baseline methods.The comparison is presented in Figure A.
- Diverse Behaviors: Given the same goal, InterPrior produces multiple valid yet distinct interaction trajectories, indicating a diverse learned latent space.Figure B illustrates behavioral diversity under identical goal conditions.
- Kinematic HOI Integration: InterPrior can follow InterDiff-generated targets using sparse wrist, feet, and object inputs while adaptively deviating from those targets.The integration converts generated pose sequences into snapshot and trajectory goals before policy inference.
- Limitations and Future Work: The model is bounded by training-data coverage and quality, often defaulting to conservative balanced strategies for highly corrupted or unseen interactions.It is tailored to rigid objects and can exhibit shallow interpenetrations, foot skating, and long-rollout object drops.
- Limitations and Future Work: The current hand and contact representation does not support fine-grained finger dexterity or in-hand manipulation.Future work includes richer hand models and broader dataset diversity.