Source-linked AI summary
Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven Locomotion
Jian Zhou, Xingyu Zhang, Rui Ma, Yu Cao, Shane Xie, Zhi-qiang Zhang
TL;DR
Muscle-driven locomotion must achieve both physiological plausibility and adaptability to changing body capacity and disturbances. The paper regulates a fixed phase-dependent reflex controller with four state-dependent residual parameters, yielding more human-like nominal locomotion and robustness without retraining.
Problem
Muscle-driven locomotion must simultaneously provide physiological plausibility and adapt to changes in musculoskeletal capacity and external disturbances.
Method
A reinforcement-learning policy produces four residual parameters that regulate selected gains and thresholds of a fixed phase-dependent reflex controller.
Results
Residual-Reflex RL generated more human-like joint kinematics, ground reaction forces, and gait consistency across 0.8, 1.0, and 1.2 m/s, while remaining effective under weakness and perturbations without retraining.
Takeaways & Limitations
Regulating existing neuromuscular control mechanisms provides a compact formulation for improving physiological plausibility and adaptability in muscle-driven locomotion.
Abstract
from arXiv · showhide
Muscle-driven locomotion provides a physically grounded approach to generating realistic human movement. However, achieving both physiological plausibility and adaptability to changes in musculoskeletal capacity and external disturbances remains a fundamental challenge. To address this limitation, we propose a Reflex-Informed Neuromuscular Reinforcement Learning framework for muscle-driven locomotion. Within this framework, a fixed phase-dependent reflex controller serves as the underlying neuromuscular control mechanism, while the reinforcement learning policy produces four biomechanically meaningful residual parameters to modulate key reflex gains and thresholds associated with hip swing, knee support, and ankle propulsion according to the current state. Experimental results demonstrate that the proposed framework generates physiologically plausible locomotion with improved kinematic accuracy and dynamic consistency, as well as better bilateral symmetry and stride-to-stride consistency under nominal walking conditions. The learned policy remains robust under muscle weakness and external perturbations without retraining.
1 Introduction
The paper addresses the tension between physiological plausibility and adaptability in muscle-driven locomotion by regulating existing neuromuscular mechanisms with reinforcement learning. Its framework uses four residual parameters to modulate a fixed reflex controller and evaluates nominal walking, weakness, and perturbation conditions.
- Muscle-driven locomotion must combine physiological plausibility with adaptation to musculoskeletal and environmental changes.
- Structured neuromuscular controllers produce meaningful locomotion but are usually fixed, whereas direct muscle-action reinforcement learning adapts online through a high-dimensional redundant action space.
- The framework reformulates reinforcement learning as state-dependent regulation of existing neuromuscular control mechanisms rather than direct muscle-action generation.
- Its policy modulates key hip-swing, knee-support, and ankle-propulsion reflex pathways through four compact, interpretable residual parameters.
- Experiments assess physiological plausibility under nominal walking and adaptability under muscle weakness and external perturbations.
2 Related Work
Prior work establishes muscle-driven simulation as a mature basis for realistic locomotion, but existing controllers do not simultaneously provide physiological plausibility and adaptability. Reflex-based methods offer structured sensory control, while reinforcement learning improves adaptation but can produce less human-like gait.
- Musculoskeletal simulation explicitly models muscle activation, muscle–tendon dynamics, skeletal dynamics, and body–environment interaction for human locomotion studies and character animation.
- Reflex-based controllers map muscle sensory feedback to stimulation and have generated stable, speed-adaptive, and robust locomotion, but commonly rely on manually designed or offline-optimized parameters.
- Direct muscle-control reinforcement learning generates activations online, while related work improves realism or exploration through rewards, model-based optimization, embodied feedback, or learned reflex networks.
- Existing approaches cannot simultaneously provide physiological plausibility and adaptability: structured controllers struggle with changes, while learned muscle control can deviate from human kinematics and dynamics.
3 Method
The method combines a fixed phase-dependent reflex controller with muscle-driven dynamics and a reinforcement-learning policy that regulates reflex pathways from musculoskeletal observations. The controller uses sensory feedback, phase-specific topology, and Hill-type muscle mechanics to generate locomotion.
- The framework combines a fixed phase-dependent reflex controller, a muscle-driven simulation, and four-dimensional residual actions produced from the current musculoskeletal state.
- The character uses 18 Hill-type muscle–tendon units representing nine bilateral muscle groups in a sagittal-plane musculoskeletal model.
- 3.2.2 Muscle-driven dynamics.: Neural stimulation is filtered into activation, Hill-type muscle mechanics compute muscle–tendon force, and muscle forces are transformed into generalized joint moments for forward dynamics.
- 3.3 Reflex control.: The reflex controller uses delayed muscle length, velocity, force, and posture feedback to compute stimulation through phase-dependent pathways.
- 3.3.2 Phase-dependent topology.: A five-phase gait controller shares parameters across all phases, stance phases, or swing-related phases, with sparse directed feedback graphs selecting active pathways.
3.4 Residual Neuromuscular Regulation
The framework uses a four-dimensional residual action to regulate a fixed neuromuscular controller rather than directly generating muscle stimulation. These residuals target gait-critical reflex pathways involved in swing, support, and propulsion.
- 3.4 Residual Neuromuscular Regulation: The policy regulates a fixed neuromuscular controller through a four-dimensional residual action instead of directly generating muscle stimulation.Each action component modulates one selected reflex parameter around its nominal value.
- 3.4 Residual Neuromuscular Regulation: Table 1 summarizes the selected residual-reflex actions and their corresponding controller pathways.
- 3.4 Residual Neuromuscular Regulation: The residual parameters modulate reflex pathways associated with hip swing, knee support, and ankle propulsion.The selected parameters were chosen for their roles across major lower-limb joints and locomotor functions.
3.5 Policy Learning with MPO
Policy learning uses an off-policy actor–critic framework with MPO, where the actor maps musculoskeletal observations to residual actions and the critic evaluates state–action pairs. The reward combines target-speed tracking with penalties for loading, excitation variation, recruitment, joint limits, instability, and rapid residual changes.
- 3.5 Policy Learning with MPO: MPO trains an actor–critic policy whose actor receives musculoskeletal observations and outputs four-dimensional residual actions.The critic evaluates the corresponding state–action pair.
- 3.5 Policy Learning with MPO: The observation includes muscle lengths, velocities, forces, excitations, activations, foot positions, and joint states with their velocities.
- 3.5 Policy Learning with MPO: The reward combines target-speed tracking with penalties for excessive loading, excitation variation, muscle recruitment, joint-limit loading, instability, and rapid residual changes.These terms jointly shape locomotion behavior during policy optimization.
4.1 Evaluation Protocol
The evaluation compares reinforcement-learning and offline-optimized reflex controllers across nominal walking, plantarflexor weakness, and external push perturbations. Nominal gait is assessed using task performance, human-reference kinematics and GRF, bilateral symmetry, and stride-to-stride repeatability.
- 4.1 Evaluation Protocol: Four approaches are compared: E2E-RL, DEP-RL, Residual-Reflex RL, and a reflex controller with parameters optimized offline using CMA-ES.The CMA-ES baseline keeps its optimized reflex parameters fixed during evaluation.
- 4.1 Evaluation Protocol: The protocol evaluates nominal walking, plantarflexor weakness, and external push perturbations to assess gait quality and adaptability.The same character and simulation settings are used across controllers unless weakness is introduced.
- 4.1 Evaluation Protocol: Nominal walking is tested at 0.8, 1.0, and 1.2 m/s using speed, stride, human-reference kinematics, vertical GRF, symmetry, and repeatability metrics.Kinematic and GRF scores measure agreement with human-reference ranges, while lower symmetry and variation measures indicate more consistent gait.
- 4.1 Evaluation Protocol: Bilateral symmetry uses RMS left–right waveform differences, while stride repeatability uses coefficients of variation for stride time and stride length.Lower values indicate more symmetric or repeatable gait cycles.
4.2 Nominal Gait Generation
Under nominal walking conditions, Residual-Reflex RL generated sustained locomotion with the strongest human-reference kinematics, ground-reaction-force consistency, bilateral symmetry, and stride repeatability among the compared controllers.
- Task performance: All three controllers completed the walking tasks, but Residual-Reflex RL achieved the most accurate speed tracking at 0.8, 1.0, and 1.2 m/s.Its absolute speed errors were 0.01, 0.05, and 0.01 m/s, respectively.
- Joint kinematic plausibility: Residual-Reflex RL achieved the highest joint-kinematic score at every target speed, averaging 88.0% across 0.8, 1.0, and 1.2 m/s.Its largest improvements were in knee and ankle coordination.
- Gait consistency: Residual-Reflex RL maintained narrower and more human-reference-aligned gait-cycle distributions, supporting more consistent bilateral joint coordination across successive cycles.The qualitative gait sequence also showed upright posture and smooth stance-to-swing transitions.
- Dynamic plausibility: Residual-Reflex RL consistently produced the highest vertical-GRF scores and smoother human-like double-support loading across target speeds.E2E-RL showed irregular loading, while DEP-RL produced larger impact peaks and stride-to-stride variation.
- Gait consistency: Residual-Reflex RL achieved lower left–right waveform differences and lower stride-time and stride-length variability than E2E-RL and DEP-RL.At 1.2 m/s, its stride-time and stride-length coefficients of variation were 0.70% and 0.74%.
4.3 Adaptation to Plantarflexor Weakness
The trained Residual-Reflex RL policy adapted to soleus weakness without retraining, preserving stable, physiologically plausible locomotion while reducing speed as force capacity declined.
- Survival comparison: Residual-Reflex RL completed the full 25 s evaluation at every tested weakness level, whereas the fixed CMA-ES controller failed under every weakened model.The CMA-ES controller remained stable only at nominal strength.
- Speed adaptation analysis: Realized speed decreased from 1.20 m/s at nominal strength to 0.95 m/s at 83% soleus strength, producing a slower but stable gait.The controller reduced speed rather than preserving the command at the expense of stability.
- Kinematic preservation: The policy preserved human-reference kinematics across weakness levels, achieving an 89.0% joint-kinematic score even at 83% soleus strength.This preservation occurred with the same trained policy and no retraining.
- Residual-action modulation: All four residual channels remained active, with coordinated redistribution across swing-related and stance-support pathways rather than single-parameter compensation.The soleus-related a3 channel showed the largest RMS magnitude across the tested weakness range.
- Residual-action modulation: The four-dimensional residual action provided a compact adaptive interface for online neuromuscular regulation across weakened musculoskeletal models.Channel a3 modulates soleus force feedback for ankle support and propulsion, while a4 modulates vasti feedback for knee support.
4.4 Recovery from External Pushes
Residual-Reflex RL recovered from backward torso pushes more effectively than the fixed CMA-ES reflex controller, sustaining walking under single pushes and some repeated disturbances.
- Adaptation mechanism: The recovery results show that the learned policy adapts online to external disturbances without retraining, unlike fixed offline-optimized reflex parameters.The same 1.2 m/s policy was retained for the perturbation experiment.
- Perturbation survival: Residual-Reflex RL completed 25 s after both 75 N and 100 N single pushes, while CMA-ES failed after approximately 4.5–4.6 s under every perturbation condition.Under repeated pushes, Residual-Reflex RL completed 75 N trials but terminated at 20.1 s for repeated 100 N pushes.
- Recovery process: Residual-Reflex RL reorganized foot placement and resumed periodic walking after single pushes, whereas the CMA-ES controller failed to regain a stable gait.Repeated 100 N pushes exposed the learned controller’s recovery boundary at 20.1 s.
5 Discussion
The discussion attributes the method’s combined plausibility and adaptability to reinforcement learning over a fixed reflex structure rather than direct learning of muscle actions.
- Main findings: Residual-Reflex RL improved joint kinematics, human-like ground-reaction forces, and gait consistency over end-to-end muscle-control policies under nominal walking.The same trained policy also remained effective under muscle weakness and external disturbances without retraining.
- Interpretation: The fixed phase-dependent reflex controller supplies neuromuscular topology, while reinforcement learning regulates selected reflex parameters through residual actions.This formulation avoids directly learning muscle actions in a high-dimensional, redundant action space.
- Adaptability: Continuous residual regulation allows the controller to adjust reflex parameters when muscle capacity changes or external disturbances occur, instead of relying on one fixed parameter set.This contrasts with the fixed parameters of the CMA-ES-optimized reflex controller.
6 Conclusion
The framework combines a fixed phase-dependent reflex controller with residual reinforcement learning to preserve neuromuscular structure while adapting reflex behavior online. Experiments showed improved physiological plausibility and robustness under nominal walking, muscle weakness, and external perturbations.
- The framework learns four biomechanically meaningful residual parameters to regulate a phase-dependent reflex controller instead of directly generating muscle excitations.This compact interface preserves the underlying neuromuscular control structure while enabling online adaptation.
- Residual-Reflex RL generated more human-like joint kinematics, ground reaction forces, and gait consistency than end-to-end muscle-control policies across walking speeds of 0.8, 1.0, and 1.2 m/s.The comparison was conducted under nominal walking conditions in a Hyfydy-based musculoskeletal simulation.
- The same trained policy remained effective under plantarflexor weakness and external perturbations without retraining while preserving physiologically plausible locomotion.
- Integrating reinforcement learning with an existing neuromuscular control structure simultaneously improves physiological plausibility and adaptability in muscle-driven locomotion.The framework is also described as applicable to muscle-driven character animation, computational human-locomotion studies, and bio-inspired locomotor control.