Source-linked AI summary

Learning coordinated badminton skills for legged manipulators

Yuntao Ma, Andrei Cramariuc, Farbod Farshidian, Marco Hutter

arXiv:2505.22974v2cs.ROcs.LG

TL;DR

Legged mobile manipulators must coordinate locomotion, arm swings, and perception to play badminton dynamically. The paper develops unified whole-body reinforcement-learning control with a camera-noise-informed perception model, and demonstrates autonomous human play with active perception and swings up to 12.06 m s−1. The results support legged manipulation for dynamic sports tasks, while the current swing and sensing designs remain limited.

  • Problem

    Coordinating rapid locomotion, accurate arm motion, and perception remains difficult for legged robots performing dynamic badminton interactions.

  • Method

    A unified reinforcement-learning controller uses all robot degrees of freedom, multi-target training, asymmetric actor-critic learning, and a perception model incorporating the hardware EKF and camera noise.

  • Results

    The robot autonomously plays badminton with humans using onboard perception, achieves swings up to 12.06 m s−1, and completes ten consecutive shots in an outdoor rally.

  • Takeaways & Limitations

    The system demonstrates that tightly coupled whole-body control and perception can support legged mobile manipulation in dynamic human sports.

  • Takeaways & Limitations

    The system uses configurable swing rules and relies heavily on an EKF with a single off-the-shelf stereo camera, limiting swing diversity and sensing richness.

Abstract

from arXiv · show

Coordinating the motion between lower and upper limbs and aligning limb control with perception are substantial challenges in robotics, particularly in dynamic environments. To this end, we introduce an approach for enabling legged mobile manipulators to play badminton, a task that requires precise coordination of perception, locomotion, and arm swinging. We propose a unified reinforcement learning-based control policy for whole-body visuomotor skills involving all degrees of freedom to achieve effective shuttlecock tracking and striking. This policy is informed by a perception noise model that utilizes real-world camera data, allowing for consistent perception error levels between simulation and deployment and encouraging learned active perception behaviors. Our method includes a shuttlecock prediction model, constrained reinforcement learning for robust motion control, and integrated system identification techniques to enhance deployment readiness. Extensive experimental results in a variety of environments validate the robot's capability to predict shuttlecock trajectories, navigate the service area effectively, and execute precise strikes against human players, demonstrating the feasibility of using legged mobile manipulators in complex and dynamic sports scenarios.

INTRODUCTION

Badminton exposes the difficulty of coordinating whole-body locomotion, arm motion, and perception in dynamic settings. The paper addresses this challenge with unified reinforcement-learning control that enables a legged manipulator to track and strike shuttlecocks using onboard perception.

  • Badminton requires coordinated footwork, upper-limb motion, hand-eye coordination, and shuttlecock trajectory prediction for successful returns.
  • Legged badminton robots must balance rapid locomotion with accurate arm movements despite many degrees of freedom and dynamic task demands.
  • Commercial cameras constrain frame rate, angular resolution, field of view, and transmission delay, motivating perception-aware control that keeps the target visible.
  • Existing athletic-robot methods largely rely on simplified models, local feedback, static-scene locomotion, or decoupled legged manipulation.
  • The proposed unified RL controller jointly controls base locomotion and arm actions across all degrees of freedom to track timed end-effector targets.
  • Real-world camera noise is incorporated into training to support active perception, while the deployed policy uses onboard perception and computation across laboratory, industrial, and outdoor environments.
  • The quadrupedal manipulator achieves swing velocities up to 12.06 m s−1 while tracking and intercepting shuttlecocks against human opponents.

RESULTS

The robot demonstrated coordinated badminton skills across simulation and hardware, combining shuttlecock perception, whole-body motion, and racket control. It maintained stability, learned active perception behaviors, adapted locomotion to target distance, and achieved repeated human rallies despite performance limits at difficult locations and time constraints.

  • Human-game validation: The robot maintained long human-play rallies, including a 10-shot streak, while returning varied shuttlecock trajectories and repositioning toward court center after hits.The perception module registered trajectories after 0.357 s on average, leaving 0.654 s until the target interception height.
  • Spatial robustness: Velocity tracking degraded beyond 2.5 m laterally and diagonally and beyond 2.0 m longitudinally from court center, while hardware tests confirmed stable net-crossing hits without exceeding arm-current constraints.The difficult regions included shuttlecocks approaching from behind or overhead, where maintaining visual contact was harder.
  • Racket control: 12.06 m s−1 was the robot’s peak executed swing velocity, although tracking accuracy diminished above commanded velocities of 10 m s−1.At a commanded 12 m s−1 swing, the robot achieved mean velocity of 10.8 m s−1 and mean position error of 0.117 m.
  • Active perception: The proposed policy reduced perception error relative to ground-truth-trained control and matched the FOV-reward baseline while avoiding explicit FOV rewards.The FOV-reward policy achieved similar perception error but used less energy-efficient base attitude and locomotion control.
  • Active perception: The robot learned pitch maneuvers that reduced shuttle angular velocity in the camera frame and extended visual tracking by 0.10 s before a hit.It pitched down, then up to retain the shuttlecock in view, before adjusting posture for the racket swing.
  • Whole-body coordination: Locomotion adapted to target distance: nearby swings used minimal reorientation, medium targets used irregular four-leg gaits, and 2.2 m targets induced high-frequency galloping and 1.0 m arm extension.Reduced leg regularization encouraged more dynamic leg movements, while a 0.4 s deadline remained physically unreachable and caused a missed hit without excessive limb motion.

DISCUSSION

The system demonstrates coordinated badminton play by coupling whole-body locomotion, manipulation, and onboard perception, while adapting gait and perception-related behavior to demanding conditions. Hardware experiments show sustained human rallies, but performance remains bounded by swing-command design and camera-based perception limits.

  • System capabilities: The robot plays badminton using onboard perception through coordinated legged locomotion, manipulation, multi-target training, and asymmetric actor-critic reinforcement learning.The learned behaviors include follow-through after hits and active perception without explicit training heuristics.
  • Experimental validation: The system achieved ten consecutive shots in a human rally across environments while responding to varied shot angles, speeds, and landing locations.
  • Perception and transfer: The training pipeline aligns noisy perception in simulation and hardware by using a consistent perception model and the same EKF, encouraging active perception behaviors.This design addresses the information gap between privileged training and deployment observations.
  • Gait adaptation: For targets 2.2 m away, the robot reached the target in 1.6 s using a galloping-like gait, while nearby targets required little foot lifting and unreachable targets triggered stepping.Under tight time constraints, the policy balanced safety with target-tracking accuracy.
  • Limitations and extensions: Current swing rules constrain racket control to configured height, velocity, orientation, interception heights of 0.9-1.4 m, and one racket side.The authors identify adaptive high-level commands and more diverse swing motions as extensions.
  • Limitations and extensions: Returns landing behind the robot have lower success rates because maintaining the shuttlecock within the camera field of view becomes harder during walking.A wider-FOV camera or actuated camera pitch joint could mitigate this limitation.
  • Limitations and extensions: The current state estimator relies heavily on an EKF applied to a single off-the-shelf stereo camera, motivating additional sensing modalities for more intense gameplay.Suggested additions include torque, sound, RGB, depth, event-based cameras, and human pose estimation.

MATERIALS AND METHODS

The method trains a whole-body visuomotor policy in a high-fidelity simulation with constrained reinforcement learning, asymmetric actor-critic training, and hardware-derived perception modeling. The training design improves value estimation and policy learning while reusing perception components during deployment.

  • Policy training: The whole-body maneuvering policy is trained with reinforcement learning in a high-fidelity simulator containing identified robot dynamics and hardware-specific constraints.The simulation models manipulator transmission, joint actuators, and system-identification parameters.
  • Policy training: Six swing targets per episode train end-effector tracking and post-swing follow-through, while asymmetric actor-critic training keeps episode-level information available only to the critic.Time-based rewards support the consecutive-swing training setup.
  • Critic design: The critic receives enhanced sensor data and preemptive knowledge unavailable during deployment to improve value-function estimation.These observations include noiseless state information and opponent-dependent task variables.
  • Ablation results: Preemptive knowledge improved both learning and final end-effector tracking performance compared with symmetric privileged training.Symmetric training induced unnecessary time-dependent actor behaviors that complicated value estimation.
  • Ablation results: Removing additional critic observations increased value-function errors and worsened learning outcomes.With enhanced observations, critic predictions closely matched discounted trajectory returns.
  • Perception modeling: The simulated perception model represents detection probability and measurement noise as functions of shuttlecock distance, robot angular velocity, and camera field of view.The model was regressed from hardware-collected data, and the same EKF and shuttlecock trajectory prediction module were used during training and deployment.
  • Reward design: Swing rewards activate for one timestep and score racket position, orientation, and velocity at the intended swing time.

F. Leg joint torques

The training and deployment pipeline combines perception modeling, shuttlecock prediction, constrained reinforcement learning, and hardware system identification to support robust badminton control. Leg torque usage remained within limits, while N-P3O prevented current-limit violations during deployment.

  • Perception and prediction: The deployment pipeline reused hardware-matched perception, filtering, and prediction components during training to reflect real-world conditions.The pipeline integrated a perception noise model, EKF, shuttlecock trajectories, and target-offset approximation into the learning loop.
  • Deployment limitation: Perception error levels depend on testing-site lighting and ambient color, creating an expected error difference when deploying the trained policy in a new environment.The authors report that learned active perception behavior would still qualitatively transfer and decrease perception error in different environments.
  • Perception and prediction: A narrower camera field of view improved angular resolution, reduced measurement noise, and supported more accurate shuttlecock state estimation.Synchronized camera timestamps were also used to compensate for base angular velocity during swing preparation.
  • Sim-to-real preparation: Hardware model parameters were optimized with CMA-ES using joint-position trajectories to reduce sim-to-real mismatch for the arm dynamics.The optimization targeted friction, damping, and armature parameters because accurate manipulator torque measurements were unavailable.
  • Constrained control: N-P3O enforced the 8 A arm current constraint, whereas the unconstrained baseline violated it even with soft over-current penalties.The N-P3O-trained policy never violated the constraint during hardware deployment.
  • Leg joint torques: Leg torque limits were modeled as soft reward penalties because hardware treated them as soft constraints, and observed leg torque usage remained within limits across swing phases.The reported distributions were computed across all target positions on the court using 199,650 trajectories per target-position set.

SUPPLEMENTARY MATERIALS

The supplementary materials specify the perception, trajectory, reward, control, deployment, and hardware-modeling components used to train and evaluate the badminton policy. They also report perception-loop timing, state-dependent action entropy, and a simulation-only extension to humanoid robots.

  • Deployment and hardware: Hardware interception requires the predicted trajectory to cross both the 1.55 m and ground-level service-area rectangles, with swing time based on a configurable height crossing.The supplementary deployment procedure also models the DynaArm belt transmission and reports a perception-loop mean of 0.375 s, ranging from 0.217 s to 0.517 s.
  • Trajectory sampling: Pre-sampled shuttlecock trajectories reduce training computation while covering varied interception and landing locations across the court.The nominal evaluation flight uses distribution means and shifts the starting position to assess different court locations.
  • Training rewards: The reward design combines single-step interception-point, racket orientation, and swing-velocity tracking with perception-error and impact regularization.Interception tracking is activated at the swing timestep, while impact penalties discourage excessive vertical link acceleration and stomping.
  • Policy observations: The policy receives robot-available actor observations, while critic-only observations provide additional information to reduce value-estimation error.The observation design distinguishes deployable actor inputs from critic information representing complete MDP information.
  • Action distribution: State-dependent action standard deviation decreases action entropy near the swing, producing a minor reward improvement during training.The evaluation used 16 independent swings to examine the entropy trend leading up to the swing.
  • Interception estimation: The interception estimate is computed by linearizing shuttlecock trajectory prediction around state-estimation error, reducing the cost of full trajectory prediction.Higher-order Δt terms are rejected because their magnitudes are small.
  • Potential extensions: The framework was also applied to Unitree G1 humanoids in simulation, while hardware deployment remains a notable challenge.The humanoid experiment is presented as an early indication of potential future directions rather than a hardware validation.
Loading 2505.22974v2…