Source-linked AI summary
AM-Bench: A Modular Simulation Suite and Benchmark for Aerial Manipulation Policy Learning
Yutong Wang, Dongjae Lee, Xiaofeng Guo, Yuanzhu Zhan, Yufei Jiang, Bavin Saravanan, Muqing Cao, Jia Xie, Chenyang Mao, Sebastian Scherer, Junyi Geng, Guanya Shi
TL;DR
Ground-focused manipulation benchmarks provide limited support for evaluating the coupled control, embodiment, and policy challenges of aerial manipulation. AM-Bench introduces a modular multirotor simulation benchmark spanning these factors, and uses controlled studies and real-world validation to characterize their interactions. The benchmark provides a foundation for system-level analysis, while remaining scoped to selected aerial-manipulation morphologies and policy paradigms.
Problem
Existing benchmarks provide limited system-level evaluation for multirotor aerial manipulation, whose performance depends on disturbances, coupled floating-base dynamics, constrained actuation, embodiment, control, and policy design.
Method
AM-Bench factorizes multirotor aerial manipulation into modular task, embodiment, policy-interface, low-level-control, disturbance, and actuator-constraint components across 12 representative tasks.
Results
AM-Bench characterizes high-level policy, policy–control interface, and embodiment-dependent behavior, with real-world validation of modeled effects and a hardware test of the learning pipeline.
Takeaways & Limitations
The benchmark supports controlled system-level studies of how embodiment, control, disturbances, and policy choices interact in dynamics-critical aerial manipulation.
Takeaways & Limitations
The benchmark focuses on single-robot multirotor aerial manipulators with rigid arms and currently emphasizes imitation-learning and vision-language-action baselines rather than reinforcement learning.
Abstract
from arXiv · showhide
Standardized benchmarks have played a central role in advancing robot manipulation learning, yet most focus on ground-supported manipulation systems, which limits their applicability to dynamics-critical domains such as aerial manipulation (AM). AM presents distinct system-level challenges, including environmental disturbances, coupled dynamics between the manipulator and floating base, and constrained degrees of freedom. Consequently, task performance depends jointly on robot embodiment, low-level control, and high-level policy design. We introduce AM-Bench, a modular simulation suite and benchmark for multirotor-based AM policy learning. AM-Bench includes representative embodiments spanning underactuated, fully actuated, and overactuated systems, 12 tasks across contact, transport, and constrained interaction, configurable aerodynamic disturbances and actuator saturation, standard low-level controllers, and baseline policy-learning algorithms. Unlike prior manipulation benchmarks that primarily emphasize end-to-end policy performance, AM-Bench enables system-level evaluation of how embodiment, control, disturbances, and policy choices interact. We demonstrate its diagnostic value through three simulation studies spanning high-level policies, policy--control interfaces, and embodiments, together with real-world validation of modeled effects and a hardware test of the learning pipeline.
1 Introduction
AM-Bench addresses the limited system-level evaluation of multirotor aerial manipulation by modularizing task, embodiment, policy, control, disturbance, and actuator factors. It supports controlled studies of their interactions rather than only end-to-end policy performance.
- Motivation: Existing benchmarks offer limited system-level evaluation for multirotor aerial manipulation, where disturbances, coupled base–manipulator dynamics, and mechanical constraints complicate control.AM systems operate in elevated or near-building environments and face wind, aerodynamic effects, noisy state estimation, limited degrees of freedom, workspace, and wrench capabilities.
- AM-Bench: AM-Bench organizes aerial manipulation evaluation around task environments, robot embodiments, high-level policy interfaces, low-level control, disturbances, and actuator constraints.The suite supports underactuated, fully actuated, and overactuated platforms, common visuomotor interfaces, standard controllers, and modeled aerodynamic effects.
- Benchmark scope: Its 12 tasks represent recurring capabilities in inspection, installation, harvesting, transport, and constrained-contact maintenance.These capabilities remain difficult to execute reliably with current aerial manipulation methods.
- Evaluation role: The benchmark enables controlled studies of individual system components and their interactions, including high-level policies, policy–control interfaces, and embodiment-dependent behavior.The studies are supported by real-world validation and are intended to isolate design factors rather than provide exhaustive conclusions across all aerial manipulation scenarios.
2 Related Work
Aerial manipulation requires coordinated hardware design, low-level control, whole-body planning, and high-level decision-making because floating bases must handle gravity and contact wrenches through flight dynamics. Existing benchmarks support standardized robot-learning evaluation, but aerial-manipulation coverage remains limited and often emphasizes hardware capabilities or navigation rather than policy learning.
- Aerial manipulators must compensate for gravity and contact wrenches solely through flight dynamics, motivating adaptive and robust control strategies.
- Underactuated multirotors couple horizontal acceleration to attitude, so lateral manipulation motion can perturb end-effector orientation.Whole-body planning and control coordinate base and manipulator trajectories to address this coupling.
- Robot-learning benchmarks enable standardized and reproducible evaluation across imitation, reinforcement, and meta-reinforcement learning, while expanding task, horizon, and sensing complexity.
- Recent dynamics-critical benchmarks evaluate perception and control alongside policy learning, but aerial benchmarks primarily target navigation, while one AM benchmark focuses mainly on hardware prototype capabilities.
3 Simulation Suite
AM-Bench is a modular simulation suite that evaluates multirotor aerial-manipulation systems across configurable tasks, embodiments, policy interfaces, controllers, disturbances, and actuator constraints. Its 12 tasks span instantaneous interaction, object transport, and articulated or constrained contact, while supported platforms cover underactuated, fully actuated, and overactuated designs.
- Simulation Suite: AM-Bench factorizes aerial-manipulation evaluation into task environment, robot platform, high-level policy interface, low-level control, and disturbances or actuator constraints.
- Simulation Suite: The suite contains 12 representative tasks organized into instantaneous interaction, object transport, and articulated-object or constrained-contact skills.
- Simulation Suite: Object-transport tasks test visual target grounding and adaptation to payload-induced dynamics, while toss ball additionally requires fast whole-body motion.
- Simulation Suite: Constrained-contact tasks restrict aerial motion through articulation or low-dimensional manifolds and require compliance with trajectories while compensating for destabilizing contact wrenches.
- Simulation Suite: AM-Bench spans physically motivated underactuated, fully actuated, and overactuated multirotors, including UA-Quad, UA-Hexa, FA-Hexa, and Omni-Hexa.
- Simulation Suite: High-level policies can use RGB and proprioceptive observations with either end-effector targets or base-and-arm joint targets as action interfaces.
- Simulation Suite: The suite combines independent manipulator position control with wrench-based multirotor control, producing thrust-and-torque inputs for underactuated platforms and 6-DoF wrenches for fully or overactuated platforms.
- Simulation Suite: End-effector targets can be mapped through whole-body MPC or decoupled IK tracking, with the latter using geometric PID or L1 adaptive base control.
4 Experiments
The experiments assess high-level policies, policy–control interfaces, and robot embodiments, using AM-Bench to expose performance differences across these system factors. Real-world experiments further test modeled effects and hardware deployment of the learning pipeline.
- High-Level Visuomotor Policy: The study evaluates visuomotor policies across 12 aerial-manipulation tasks using specialist imitation-learning baselines and pretrained vision-language-action policies.All methods use FA-Hexa, the EE target interface, and IK-PID control; each task has 80 scripted demonstrations and 30 evaluation rollouts per method.
- High-Level Visuomotor Policy: 28.1 and 46.4 percentage points are the macro-average success-rate gains from zero-shot to multi-task fine-tuning for π0 and π0.5, respectively.Task-specific specialization adds 8.3 points for π0 and 3.1 points for π0.5.
- High-Level Visuomotor Policy: Failures remain concentrated in frame assembly, rotate valve, and toss ball, despite meaningful intermediate behaviors such as grasping.The reported bottlenecks concern precise end-effector control and contact-rich interactions.
- Policy–Control Interface: The EE target interface consistently achieves higher task success rates and generally lower tracking error, tilt use, and saturation than direct base-and-arm commands.EE target control with whole-body MPC performs best among the evaluated configurations because the controller handles redundancy, coordination, and constraints.
- Robot Embodiment: Underactuated platforms require roll and pitch for lateral translation, whereas FA-Hexa translates without base tilting and Omni-Hexa exhibits lower saturation.These embodiment differences appear under identical EE objectives with the same scripted policy and IK-PID controller.
- Robot Embodiment: With matched base and arm properties, UA-Hexa couples lateral end-effector motion to base attitude changes, while FA-Hexa provides attitude-independent lateral translation.The comparison isolates rotor tilt as the embodiment difference in the rotate valve task.
- Robot Embodiment: Across selected cases, the EE target interface reveals embodiment-dependent differences in tracking, tilt use, and rotor saturation.Real-world experiments compare modeled ground effects and instantiate the collection, diffusion-policy training, and deployment pipeline on hardware.
5 Conclusion
AM-Bench provides a modular benchmark for system-level evaluation of multirotor aerial-manipulation policy learning. Experiments characterize policy, interface, and embodiment effects, while real-world tests validate selected modeled effects and hardware deployment.
- 5 Conclusion: AM-Bench integrates tasks, embodiments, disturbances, low-level controllers, and high-level interfaces for controlled evaluation of aerial-manipulation system factors.The benchmark supports studies of imitation-learning and vision-language-action policies, policy–control interfaces, and embodiment-dependent behavior.
6 Limitations
AM-Bench has scope and fidelity boundaries that limit the generality of its current conclusions. The benchmark excludes several morphologies, emphasizes IL and VLA baselines, and does not model the full range of real-world dynamics and disturbances.
- 6 Limitations: The benchmark covers single-robot multirotor aerial manipulators with rigid arms, excluding cooperative systems, cable-suspended payloads, continuum or compliant manipulators, and soft grippers.Supporting these morphologies would require different dynamics models and evaluation protocols.
- 6 Limitations: Current studies emphasize IL and VLA baselines, while reinforcement-learning evaluation and broader sim-to-real validation remain future work.The paper cites reward-design and sim-to-real-transfer challenges as reasons RL is not emphasized.
- 6 Limitations: The simulator models selected aerodynamic effects and reduced-order actuator dynamics but omits full motor–propeller, ESC, and complete real-world disturbance dynamics.Higher-fidelity models and their effects on learned-policy robustness and task success remain to be evaluated.
A.1 Task Environment
The task environment section documents the task suite, supported platforms, action and observation interfaces, and configurable observation components for AM-Bench.
- Task Suite: Table 5 records each task’s success criteria, binary subtask criteria, and domain-randomization configuration.
- Robot Platforms: AM-Bench supports four aerial manipulator platforms spanning underactuated, fully actuated, and overactuated multirotor systems.The platforms are UA-Quad, UA-Hexa, FA-Hexa, and Omni-Hexa.
- Robot Platforms: An EE-only floating end-effector model provides an oracle baseline for task execution under the shared end-effector target interface.
- Action Space: Policies can command either end-effector targets or the drone base pose and manipulator joint targets simultaneously.The base-and-arm interface has action dimension 8 + na, where na is the number of manipulator joints.
- Observation Space: Observations combine proprioceptive states with RGB images from end-effector- and base-mounted cameras, and components can be selectively enabled for ablations.
A.4 Low-Level Control Implementation Details
The low-level implementation maps policy targets into whole-body motion and wrench commands while modeling controller choices, kinematic constraints, actuator behavior, and environmental disturbances.
- Inverse Kinematics: The IK solver jointly optimizes aerial base pose and manipulator joints so the end-effector tracks a desired pose with smooth, safe whole-body motion.It is warm-started from the previous solution and solved with Levenberg–Marquardt.
- Base Tracking: The geometric base controller uses measured pose and velocity states to produce a desired body wrench, with gravity and feedforward terms included.For underactuated platforms, the commanded body force is restricted to the thrust direction.
- Adaptive Control: The L1 controller augments a nominal wrench with adaptive force and torque estimates obtained from state prediction and filtered online adaptation.
- Model Predictive Control: Whole-body MPC uses a prediction horizon of H = 32 shooting intervals and a total horizon duration of 0.8 s.Its costs weight tracking, state regulation, control effort, and control-rate penalties under base, arm, and end-effector constraints.
- Disturbances and Actuators: The disturbance model combines aerodynamic drag and wind, ground and near-wall proximity effects, per-rotor thrust limits, and optional actuator transients.Proximity effects are applied at rotor level before corrected forces and torques are reconstructed into a body wrench.
B Experiment Details
AM-Bench evaluates aerial manipulation through modular system components spanning metrics, data collection, policy training, dynamics, control, embodiment, and physical constraints.
- Evaluation metrics: Success rate averages task-success rollouts, while subtask completion averages binary criteria across evaluation rollouts.Diagnostic metrics are averaged across successful rollouts and reported as mean ± standard deviation when uncertainty is shown.
- Evaluation metrics: End-effector and base tracking errors are normalized by full arm reach for dimensionless comparison across robot morphologies.Tracking metrics are reported as N/A when the corresponding commanded reference trajectory is undefined.
- Evaluation metrics: Tilt use measures maximum roll-pitch tilt magnitude, while rotor saturation uses pre-proximity thrust after allocation, limiting, and enabled actuator dynamics.The logged thrust signal precedes ground- and near-wall-effect stages.
- Data and policies: Task datasets contain 80 successful demonstrations per task, recorded at 120 Hz and downsampled to 20 Hz for policy training and evaluation.Recorded modalities include 384 × 384 end-effector images, measured poses, gripper states, and absolute end-effector commands.
- Data and policies: ACT and DP are trained separately per task, with ACT using relative action chunks and DP using a CLIP-conditioned one-dimensional UNet denoiser.Adaptation levels include zero-shot, multi-task fine-tuning, and multi-task-to-single-task fine-tuning.
- System modeling and control: AM-Bench models coupled base-manipulator dynamics, underactuated thrust constraints, multiple embodiments, control allocation, and modular physical disturbances and actuator constraints.End-effector tracking depends on both translational and rotational control, while underactuated systems use cascaded position and attitude control.
D.1.1 Diffusion Policy Data Scaling
The diffusion-policy data-scaling study evaluates success with increasing demonstrations per task and finds stronger aggregate performance at the largest tested budget, with task-dependent variation.
- Data scaling: The study compares DP training budgets of 10, 20, 40, and 80 demonstrations per task.Per-task success rates are summarized over 30 evaluation rollouts.
- Data scaling: 80 demonstrations per task provide the strongest aggregate performance among the evaluated budgets.Each policy is evaluated on 30 rollouts.
- Data scaling: Macro-average success increases with demonstration count, although the effect varies across tasks.Press button, peg-in-hole, lemon harvesting, and NDT benefit substantially, whereas pull lever and push slider show non-monotonic trends.
D.3 Actuator Dynamics Identification and Vertical Tracking
The actuator study identifies physical FA-Hexa response characteristics, evaluates their simulation effects on vertical tracking, and compares simulated behavior with real-world observations.
- Actuator identification: 47.5 ms is the identified physical FA-Hexa rotor-response time constant, supplemented by a normalized-speed rate cap of 0.75 s^-1.The rate cap is adopted because the logs do not show a repeatable hard acceleration plateau.
- Vertical tracking: Motor-delay modeling increases high-frequency base-z tracking RMSE by 19%, while no rotor reaches the 23 N thrust limit.The comparison holds controller, embodiment, and other simulation settings fixed across actuator configurations.
- Vertical tracking: The vertical-tracking evaluation compares no motor dynamics, identified motor delay, acceleration limiting, and their combination on two sinusoidal references.The references span 0.18–0.71 Hz at low frequency and 0.25–1.50 Hz at high frequency.
- Real-world validation: The hardware experiment executes end-effector vertical references through the EE target interface, with IK generating whole-body targets and PID tracking the base trajectory.Simulation-to-real end-effector-z RMSE is reported over two shaded altitude intervals.
- Real-world validation: Physical and simulated systems exhibit similar control-input saturation patterns when wrench-space limits are reached.The comparison uses x-axis step responses under two limit configurations.