Source-linked AI summary
Humanoid Locomotion and Manipulation: Current Progress and Challenges in Control, Planning, and Learning
Zhaoyuan Gu, Junheng Li, Wenlan Shen, Wenhao Yu, Zhaoming Xie, Stephen McCrory, Xianyi Cheng, Abdulaziz Shamsah, Robert Griffin, C. Karen Liu, Abderrahmane Kheddar, Xue Bin Peng, Yuke Zhu, Guanya Shi, Quan Nguyen, Gordon Cheng, Huijun Gao, Ye Zhao
TL;DR
Humanoid loco-manipulation must unify locomotion and manipulation despite complex contact, dynamics, and hardware constraints. This survey synthesizes model-based planning and control with learning-based methods, foundation models, and tactile sensing, finding complementary strengths across paradigms while identifying persistent limitations in data, transfer, and hardware.
Problem
Whole-body manipulation remains undeveloped, while dynamic environments, high-dimensional actions, scarce robot data, and hardware constraints challenge general humanoid loco-manipulation.
Method
The survey reviews model-based planning and control, reinforcement and imitation learning, foundation models, and tactile sensing for humanoid loco-manipulation.
Results
Model-based methods offer motion accuracy, reinforcement learning robustness, imitation learning versatility, and vision-language-action methods potential generalizability, though humanoid skill mastery remains unproven.
Takeaways & Limitations
Combining model-based control with learning, foundation models, and whole-body tactile sensing is a central direction for advancing contact-rich humanoid loco-manipulation.
Takeaways & Limitations
Battery capacity restricts humanoid operation time and the feasibility of sustained high-power full-body motion in untethered scenarios.
Abstract
from arXiv · showhide
Humanoid robots hold great potential to perform various human-level skills, involving unified locomotion and manipulation in real-world settings. Driven by advances in machine learning and the strength of existing model-based approaches, these capabilities have progressed rapidly, but often separately. This survey offers a comprehensive overview of the state-of-the-art in humanoid locomotion and manipulation (HLM), with a focus on control, planning, and learning methods. We first review the model-based methods that have been the backbone of humanoid robotics for the past three decades. We discuss contact planning, motion planning, and whole-body control, highlighting the trade-offs between model fidelity and computational efficiency. Then the focus is shifted to examine emerging learning-based methods, with an emphasis on reinforcement and imitation learning that enhance the robustness and versatility of loco-manipulation skills. Furthermore, we assess the potential of integrating foundation models with humanoid embodiments to enable the development of generalist humanoid agents. This survey also highlights the emerging role of tactile sensing, particularly whole-body tactile feedback, as a crucial modality for handling contact-rich interactions. Finally, we compare the strengths and limitations of model-based and learning-based paradigms from multiple perspectives, such as robustness, computational efficiency, versatility, and generalizability, and suggest potential solutions to existing challenges.
I. INTRODUCTION
Humanoid robots are suited to human-level whole-body loco-manipulation, but complex dynamics, safety, and unstructured environments remain challenging. This survey reviews model-based and learning-based methods, tactile sensing, foundation models, and open challenges.
- Motivation: Humanoid robots replicate human morphology and functionality to perform whole-body locomotion and manipulation in human-designed environments.The survey focuses on bipedal humanoids and whole-body motion rather than visual appearance alone.
- Motivation: Complex robot dynamics, safe human collaboration, and unstructured environments make simultaneous loco-manipulation difficult.Human-like morphology supports tasks such as carrying payloads over stairs and collaborative physical interaction, but does not remove these challenges.
- Survey scope: The survey synthesizes traditional contact planning, motion planning, and control methods alongside reinforcement learning, imitation learning, and foundation models.It examines these approaches from model-based and learning-based perspectives across humanoid loco-manipulation.
- Model-based methods: Model-based methods remain foundational, increasingly combining whole-body or centroidal model predictive control with local task-space whole-body controllers.These methods trade physical-model fidelity against the quality and speed of motion generation and control.
- Learning-based methods: Learning-based methods expand motor-skill acquisition, but pure reinforcement learning is inefficient for high-dimensional humanoid tasks and sim-to-real transfer remains difficult.Imitation learning benefits from abundant human data, whereas reinforcement learning commonly relies on simulation because real-world trial and error is costly.
- Emerging capabilities: Foundation models provide open-world reasoning and multimodal semantic understanding for humanoid task planning, while whole-body tactile sensing supports contact-rich interaction.The survey compares model-based and learning-based methods and outlines remaining challenges and potential solutions.
II. BACKGROUND
Humanoid background research spans bipedal locomotion, navigation, whole-body manipulation, and their integration under complex contact and dynamics. Progress has exposed complementary limitations in classical planning and control versus pure learning, motivating integrated approaches.
- Bipedal Locomotion: Bipedal locomotion has progressed from passive and quasi-static walking to dynamic walking, while learning methods now address running, jumping, stair climbing, and parkour.
- Navigation: Navigation has advanced from static obstacle avoidance on flat terrain to height-constrained spaces, dynamic social environments, rough terrain, and other challenging settings.
- Whole-body Manipulation: Whole-body manipulation uses any body surface for interaction, unlike robots commonly restricted to predefined end-effectors such as fingertips or foot soles.
- Whole-body Manipulation: Whole-body manipulation creates combinatorial planning complexity because the robot may form an effectively infinite number of contacts across perception, estimation, planning, and control.
- Whole-body Manipulation: Forceful, compliant control and whole-body sensing remain necessary because contact-rich coordination lacks a general framework for estimating state and reacting across body contacts.
- Open Challenges: Classical approaches face huge complexity, whereas pure learning lacks flexibility for contact adaptation; the survey therefore anticipates integrated methods combining both strengths.
D. Loco-manipulation
Humanoid loco-manipulation combines robot mobility and object movement, extending whole-body interaction to coordinated contact across all limbs and surfaces. The survey emphasizes multimodal and tactile sensing while reviewing model-based and learning-based approaches for contact-rich manipulation.
- Definition and Scope: Loco-manipulation simultaneously moves the robot through locomotion and objects through manipulation; whole-body loco-manipulation additionally uses all body surfaces for environmental interaction.
- Definition and Scope: Humanoids are harder than quadrupeds for loco-manipulation because their smaller support region and higher center of mass complicate dynamic balance.
- Definition and Scope: Whole-body loco-manipulation schedules contact across all limbs to support robust movement and safe object interaction in tasks including door opening, trolley pushing, bobbin rolling, and ladder climbing.
- Tactile Sensing: Tactile sensing complements vision and proprioception by providing detailed contact information, including when vision is occluded, and by estimating properties such as roughness, texture, and weight.
- Tactile Sensing: Tactile applications in humanoid loco-manipulation are organized around hands, foot soles, and whole-body sensing, with roles spanning manipulation, obstacle recognition, terrain classification, and balance control.
- Tactile Sensing on Hands: Dynamic contact tracking, stability monitoring, and interaction-outcome prediction are crucial for complex manipulation, yet model-based methods struggle with multi-contact dimensionality and human-level dexterity.
- Tactile Sensing on Hands: Model-free reinforcement learning incorporates tactile measurements directly into state spaces, while diffusion policies and tactile foundation-model integrations target more generalizable manipulation.
- Tactile Sensing on Hands: Tactile robot hands must balance delicate-manipulation dexterity with heavy-lifting payload capacity, a trade-off motivating modular short-term designs and unified multimodal hands as a long-term goal.
B. Tactile Sensing on Feet
Foot tactile sensing can provide contact and terrain information that vision and proprioception estimate indirectly, but humanoid deployment remains technically demanding. Future systems must support richer terrain inference, multimodal fusion, and dynamic contact reasoning.
- Foot tactile sensing can estimate ground reaction forces and terrain properties that vision and proprioception cannot accurately provide indirectly.
- Existing ankle force/torque and load-cell sensors lack accurate contact-patch location, force distribution, and detailed terrain information.
- Only a few humanoid tactile-foot studies address terrain classification, ground-slope recognition, and pressure-shape reconstruction.
- Humanoid feet face larger impulse and shear forces, while robust sensors must withstand varied terrains and distribute computing and power effectively.
- Future work should estimate richer terrain properties, define terrain-complexity metrics, and fuse tactile, proprioceptive, and visual sensing for terrain-aware locomotion.
- Whole-body tactile systems improve contact awareness for compliance, physical interaction, balance, collision avoidance, and large-object manipulation, but dexterous loco-manipulation remains challenging.
A. Search-based Contact Planning
Humanoid multi-contact planning spans search, optimization, and learning approaches that trade off feasibility, dynamics integration, computation, and adaptability. Contact-implicit planning unifies contact and motion decisions but remains computationally difficult online.
- A. Search-based Contact Planning: Search-based methods explore contact configurations and can produce stable, task-efficient contact-mode sequences, with whole-body feasibility checked during or after search.
- A. Search-based Contact Planning: Search-based planning struggles to cover the exploration space within limited online time budgets and can produce high-variance solutions.
- A. Search-based Contact Planning: Optimization-based contact planning can simultaneously incorporate whole-body motion and contact interactions by integrating dynamics directly into planning.
- A. Search-based Contact Planning: CITO acceleration uses strategies including warm starts, hierarchical decomposition, and specialized optimization algorithms, but real-time humanoid deployment remains challenging.
- A. Search-based Contact Planning: Learning-based planners can improve computation efficiency and predict centroidal dynamics and contact sequences for dynamic humanoid loco-manipulation in under 0.1 s.
- A. Search-based Contact Planning: Future approaches should combine search, optimization, and learning while improving computational efficiency, contact prediction, and robustness to transient contact changes.
V. MODEL PREDICTIVE CONTROL FOR LOCO-MANIPULATION
MPC formulates loco-manipulation as finite-horizon optimal control over states, inputs, and constraint forces. Model choice determines the trade-off between computational efficiency, dynamic accuracy, and whole-body interaction capability.
- MPC optimizes state trajectories and control inputs over a finite future horizon, with costs, dynamics, equality constraints, inequality constraints, and contact forces.
- The OCP formulation can become linear convex MPC or nonlinear MPC depending on the selected dynamics models, costs, and constraints.
- Simplified dynamics models enable lightweight, high-frequency online planning, including object carrying and rough-terrain locomotion under simplified interaction assumptions.
- LIPM lacks contact-interaction and loco-manipulation dynamics, requiring lower-level whole-body control for balancing and manipulation.
- Simplified models improve efficiency but reduce accuracy and whole-body planning capability, whereas whole-body models better represent versatile interactions at greater computational cost.
- Mixed-fidelity models use different abstraction levels across the horizon to retain near-term accuracy and extend planning farther into the future, but require careful model coordination.
D. NMPC Speed-up
NMPC speed-up methods exploit problem structure, linearization, warm starts, and sampling to improve online solvability. Loco-manipulation additionally requires models that account for object dynamics and uncertain interactions.
- D. NMPC Speed-up: NMPC structure exploitation identifies interaction, repetitive, symmetric, and block-diagonal patterns to improve solvability and efficiency.
- D. NMPC Speed-up: Direct multiple shooting and direct collocation can reduce sparse NLP complexity from O(N^3) to O(N), while DDP and iLQR scale linearly over the horizon.
- D. NMPC Speed-up: Successive linearization converts nonlinear dynamics into piece-wise affine models formulated as large, sparse quadratic programs solvable online.
- D. NMPC Speed-up: Warm starts reuse previous solutions or offline gait libraries, reducing online computation to inexpensive interpolation in the latter case.
- D. NMPC Speed-up: Sampling-based MPPI reduces computational demands through search-space reduction and parallelization, but high-dimensional contact-implicit tasks remain difficult.
- Object-aware planning must address free-floating, articulated, or actuated objects whose interaction forces depend on object dynamics and states.
3) Interaction with a dynamic environment or deformable objects:
Dynamic environments and deformable objects make loco-manipulation difficult because their interaction dynamics are time-varying, difficult to model, and require adaptive sensing and planning. Whole-body control addresses these tasks through contact-constrained dynamic-task formulations, while model choice and computational efficiency remain central trade-offs.
- Dynamic surfaces and human interactions require sensor feedback to predict environmental motion and adaptively replan loco-manipulation.Human collaboration additionally requires anticipating intentions; force feedback can trigger movements, but future force evolution is difficult to predict.
- Deformable-object manipulation requires modeling flexibility, elasticity, and force-induced deformation, so task-specific simplifications are often necessary.
- Whole-body control generates joint torques, constraint forces, and generalized accelerations for desired dynamic tasks, typically using full-order dynamics.
- Humanoid WBC must satisfy underactuated, contact-constrained dynamics because environmental contact provides balance, mobility, and manipulation.
- Dynamic tasks can be represented linearly in generalized accelerations, external forces, and joint torques as equalities, inequalities, or cost terms.
- MPC commonly supplies WBC with centroidal and end-effector trajectories, which can be converted into joint-space tasks through whole-body inverse kinematics.
- Teleoperation generates posture, walking-direction, and grasp targets through visual interfaces, with retargeting used to account for morphology or motion feasibility.
- Whole-body control includes closed-form and optimization-based approaches, both of which can incorporate multiple dynamic tasks and resolve task conflicts.
B. WBC in Closed Form
Whole-body control solves inverse dynamics under humanoid underactuation and contact constraints. Closed-form methods are efficient, whereas optimization-based methods provide greater flexibility for inequalities and conflicting tasks; loco-manipulation extends these choices to interaction wrenches and unified robot-object dynamics.
- B. WBC in Closed Form: Closed-form inverse-dynamics control solves desired generalized acceleration but depends on constraint-force measurement or analytical elimination.
- B. WBC in Closed Form: Operational-space control uses humanoid redundancy to prioritize multiple tasks, such as interaction-force generation below whole-body balance.
- B. WBC in Closed Form: Closed-form WBC is computationally efficient and straightforward, but handles inequality tasks such as joint limits and obstacle avoidance less naturally.
- C. WBC through Optimization: Optimization-based WBC allows dynamic tasks, including inequalities, to be added or removed modularly.
- C. WBC through Optimization: Optimization-based WBC resolves task conflicts through strict hierarchies or soft weights and is often formulated as a quadratic program.
- C. WBC through Optimization: Hierarchical QP solves prioritized subproblems sequentially, while weighted QP uses one optimization with soft constraints and is faster than hierarchical QP.
- D. WBC for Loco-manipulation: Loco-manipulation WBC maintains motion, instantaneous balance, and contact stability while treating static or quasi-static interactions as external wrenches.
- 1) Interaction as an External Wrench:: Pre-optimization separates contact-wrench distribution from torque computation, but rotational-momentum non-holonomy requires additional body-orientation regulation.
E. Challenges in Numerical Optimization
Numerical optimization improves humanoid planning and control but remains limited by dimensionality, local optimality, infeasibility, and unmodeled uncertainty. Learning-based methods offer novel and versatile skills, yet RL still faces substantial reward, data, and sim-to-real barriers.
- Contact-explicit optimization converges faster and simplifies formulation, but the curse of dimensionality limits problem complexity, preview horizon, and discretization resolution.
- Contact-explicit methods require users to specify contact-mode sequences, limiting their ability to generate complex motions.
- Contact-implicit formulations remove strict contact-sequence dependence but introduce nonsmooth complementarity conditions and severe computational challenges.
- Nonconvex nonlinear optimization generally provides only local optimality guarantees and can fail to find feasible solutions when motion must depart from local candidates.
- Deterministic optimization often omits stochastic state estimates and future contacts, limiting robustness when solutions transfer to the real world.
- Search-based and sampling methods such as MPPI address uncertainty and local minima, with computational parallelization used to improve expediency.
- Optimization robustness remains constrained by infeasibility and difficult, task-dependent weight tuning in high-dimensional objectives.
- Learning-based methods use computation and data to develop novel behaviors, with versatility and generalization defined as central skill-learning goals.
2) Improving Learning Efficiency:
Learning efficiency depends on selecting among exploration, curriculum, sim-to-real, and demonstration strategies. Humanoid RL faces especially severe transfer difficulties, while IL trades scarce robot data for abundant but morphologically mismatched human data and teleoperation-specific constraints.
- Curriculum learning improves RL efficiency by progressing from simple tasks to increasingly difficult and complex tasks.
- Curiosity mechanisms promote exploration by encouraging visits to unexplored states without requiring an explicit reward design.
- Humanoid loco-manipulation has steeper sim-to-real challenges than quadruped locomotion because humanoids have higher DoFs and unstable dynamics.
- Domain randomization varies mass, friction, and actuator dynamics to train policies intended to remain robust in the real world.
- System identification fits simulation parameters to real trajectories, whereas domain adaptation fine-tunes a simulator-trained policy using real-world data.
- A systematic sim-to-real solution remains elusive because existing approaches are case-specific and physics engines struggle with accurate contact dynamics under real-time requirements.
- Imitation learning uses expert demonstrations, with sources including policy execution, teleoperation, motion capture, and human videos.
- Robot data are directly applicable but scarce, while human data are abundant but have significant morphological differences from humanoid robots.
2) Approaches to Learning from Robot Experience Data:
Robot-experience learning methods provide reliable skill acquisition, but collecting diverse loco-manipulation data is costly and difficult to scale. Human motion data offers a more accessible alternative, yet retargeting, missing interaction sensing, and simulation-to-real gaps remain substantial challenges.
- Robot data: Imitation learning from paired robot observations and actions includes Behavior Cloning, Inverse Reinforcement Learning, Action Chunking Transformers, and diffusion policies.These methods address policy distillation, reward reconstruction, multimodal actions, and distribution shifts from compounding errors.
- Robot data: Robot-experience imitation learning remains reliable for attaining expert-level performance, while teleoperation is a prominent method for collecting humanoid data.High-quality data collection nevertheless demands considerable effort and resources.
- Robot data: Teleoperation and other robot-data collection approaches are costly, tedious, and potentially unsafe when deployed on hardware.These constraints make large-scale loco-manipulation datasets difficult to obtain.
- Human data: Human motion can be obtained through motion capture, video reconstruction, animation, or motion generation, then retargeted from a human skeletal model to the robot.Motion capture provides diverse interactions but is expensive to scale, whereas Internet videos are more accessible but noisier and less physical.
- Human data: Human-data imitation lacks tactile and force measurements, and many highly agile or interaction-rich results remain confined to simulation.Real-world transfer is further limited by privileged simulator information and the embodiment gap between humans and humanoids.
- Human data: Internet-scale human data broadens humanoid loco-manipulation capabilities, but much of simulation’s progress has yet to be realized on real robots.Affordable capable hardware and high-fidelity simulators could accelerate this direction.
D. Skill Learning: Hybrid Methods
Hybrid methods combine demonstrations, reinforcement learning, model-based guidance, and trajectory augmentation to improve learning of complex humanoid skills. Their predefined references accelerate learning but constrain emergent behavioral diversity, motivating implicit skill representations for multi-task composition.
- Hybrid methods: Hybrid methods combine imitation and reinforcement learning, including teacher-student policies that distill privileged-observation behavior into deployable partial-observation students.The student can achieve similar performance using onboard observations.
- Hybrid methods: Model-based MPC can generate reference motions or imitation rewards, while offline trajectories avoid the training-time cost and occasional infeasibility of online MPC.Trajectory augmentation instead learns residual task-specific forces or modifications on top of references.
- Hybrid methods: Imitation and trajectory augmentation expedite learning but rely on predefined trajectories, limiting their ability to learn emergent and diverse behaviors.This trade-off motivates representations that support flexible skill composition.
- Hybrid methods: Hybrid methods provide effective guidance for learning complex and robust behaviors and have achieved stronger efficiency, versatility, and performance than single-method approaches.The survey reports successful humanoid hardware deployments.
- Implicit skill representations: Implicit skill representations include hierarchical mixtures of experts, low-dimensional motion latents, goal-conditioned policies, and learned world models.These structures select or blend skills, encode goals, or generate imaginary transition data for more efficient multi-task learning.
- Implicit skill representations: Versatile multi-task policies require structured representations of skill motion, task goals, or environment dynamics.Latent spaces, goal conditioning, and world models are identified as promising approaches.
F. Learning for Humanoid Loco-manipulation
Learning-based humanoid loco-manipulation remains less mature than locomotion and tabletop manipulation because stable contacts, precise forces, sensing, and real-world transfer are difficult. Progress depends on better benchmarks, hardware, data, and representations that capture objects, modalities, and intentions rather than poses alone.
- Current capabilities: Learning-based loco-manipulation often oversimplifies physical interactions, making stable contacts and precise contact forces difficult and limiting sim-to-real demonstrations.Only a few studies have transferred such skills to real hardware.
- Current capabilities: Hierarchical reinforcement learning manages distinct skills for tasks such as falling recovery and ball kicking, while imitation learning through teleoperation has supported autonomous loco-manipulation progress.Teleoperation is not autonomous but remains an important intermediate data-collection step.
- Benchmarks and hardware: Humanoid loco-manipulation needs large-scale systematic benchmarks spanning simulation and real-world environments, with defined tasks and evaluation metrics.Existing simulation benchmarks cover only a subset of human skills.
- Benchmarks and hardware: Affordable capable humanoids, standardized hardware platforms, open-source systems, and diverse robot prototypes could accelerate evaluation and cross-embodiment generalization.The survey identifies several open-source humanoid efforts as valuable contributions.
- Data and generalization: Robot skill learning is bottlenecked by a lack of high-quality large-scale data, while the value of scaling data remains debated.The survey frames data quality and availability as a central trade-off.
- Data and generalization: General-purpose learning requires human data to capture intentions, manipulated-object motion, and multimodal sensing rather than joint poses alone.Proposed remedies combine video or physics-based synthetic data with limited real-world multimodal data, including tactile sensing.
1) FMs for selecting low-level skills (task planning):
Foundation models can support humanoid task planning by selecting skills or generating intermediate representations, while robot foundation models directly connect multimodal inputs to control. Their deployment remains constrained by embodied knowledge, data, inference cost, safety, and the absence of established scaling laws.
- FM task planning: Pre-trained LLMs and VLMs can select skills, goals, or task graphs for humanoid and related robots, but fixed skill sets limit flexibility.Intermediate representations such as code, rewards, poses, and contacts provide more adjustable motion generation.
- Robot foundation models: Robot Foundation Models process multimodal observations and task descriptions to produce robot actions, aligning semantic knowledge with physical control behavior.Hierarchical designs combine high-level language or vision-language reasoning with multitask sensorimotor policies.
- Robot foundation models: Language-conditioned visuomotor policies have demonstrated broad manipulation skills, but successful implementations require stable robot dynamics and substantial high-quality data.These resource requirements constrain current deployments.
- Humanoid challenges: Humanoid loco-manipulation remains especially challenging for robot foundation models because humanoid dynamics are inherently unstable.Recent systems have extended robot foundation models mainly to humanoid upper-body applications.
- Robot foundation models: End-to-end vision-language-action models tokenize robot observations and actions, enabling direct action outputs without a separate trainable low-level policy.Alternative token outputs can represent diffusion processes or map learned tokens to actions through diffusion heads.
- Humanoid challenges: Autoregressive transformers are computationally inefficient over long sequences, motivating exploration of efficient high-capacity alternatives such as state-space models.Input and output representation choices are also central to robot foundation model training.
- Humanoid challenges: Foundation models face high onboard inference cost, unsafe action generation, bias concerns, and a missing robotics scaling law for coordinating model, compute, and data growth.Cloud-based high-level decisions and lower-frequency control are among the proposed responses to inference constraints.
- Outlook: The survey expects increasing use of vision-language-action models and robot foundation models trained with real, synthetic, or combined data.The intended outcome is improved understanding of robot actions and the physical world with less additional robot data.
IX. ADDITIONAL DISCUSSIONS
The survey compares model-based and learning-based humanoid methods across performance dimensions and skills, while discussing hardware constraints and future deployment opportunities. It highlights complementary strengths in robustness, versatility, and accuracy, alongside persistent limitations in energy, reliability, and benchmarking.
- Model-based Methods Versus Learning-based Methods: The comparison evaluates model-based and learning-based methods across sensor integration, robustness, accuracy, real-time feasibility, versatility, generalizability, and humanoid skills.The evaluated skills include locomotion, loco-manipulation, and whole-body multi-contact control.
- Model-based Methods Versus Learning-based Methods: Learning-based methods provide superior robustness through reinforcement learning and versatility through imitation learning, whereas model-based methods offer greater motion accuracy.
- Hardware Limitations for Humanoid Robots: High-power-density actuators, strong lightweight limbs, and accurate sensing remain hardware requirements for humanoid loco-manipulation.These requirements constrain hardware performance and motivate continued advances in humanoid physical capabilities.
- Hardware Limitations for Humanoid Robots: Quasi-Direct Drive actuators use less than 10:1 gear ratios and support dynamic, accurate motion through backdrivability and high force-control bandwidth.
- Hardware Limitations for Humanoid Robots: Remote actuation can reduce rotational inertia and torque demand for agile maneuvers, but belt drives and four-bar transmissions can require maintenance, modeling, or reduced joint range.Specialized mechanisms may address these drawbacks while preserving range of motion.
- Hardware Limitations for Humanoid Robots: Battery capacity restricts untethered operating time and high-power full-body motion, despite reported operation times of 4hrs for Apollo and 2hrs for Unitree’s G1.Continued reliance on tethered power in laboratories indicates that battery technology remains a limitation.
- Potential Applications and Future Directions: Future humanoid applications span emergency response, industrial work, households, and hospitals, but societal adoption requires physical and mental reliability.The survey identifies robust and safe mobility and manipulation as core technical challenges for acceptance.
- Potential Applications and Future Directions: Foundation models, whole-body tactile sensing, and their integration with planning and control are presented as promising directions toward open-world understanding and generalized humanoid agents.The survey reports that the effectiveness of VLA methods in mastering humanoid skills remains unconvincingly demonstrated.