Source-linked AI summary

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

Zhengyi Luo, Ye Yuan, Tingwu Wang, Chenran Li, Fernando Castañeda, Sirui Chen, Zi-Ang Cao, Jiefeng Li, David Minor, Qingwei Ben, Jinhyung Park, David Sami, Zi Wang, Xingye Da, Runyu Ding, Cyrus Hogg, Lina Song, Edy Lim, Eugene Jeong, Tairan He, Haoru Xue, Wenli Xiao, Simon Yuen, Jan Kautz, Yan Chang, Umar Iqbal, Linxi "Jim" Fan, Yuke Zhu

arXiv:2511.07820v4cs.ROcs.AIcs.CVcs.GReess.SY

TL;DR

Humanoid controllers have remained small and behavior-specific despite scaling gains in other foundation models. SONIC scales motion tracking across model capacity, motion data, and compute, then couples it with a planner and shared token space. The resulting policy supports general whole-body behaviors, unseen-motion transfer, multimodal interfaces, and VLA-driven loco-manipulation, while retaining limitations under extreme conditions and lacking formal safety and energy-efficiency treatment.

  • Problem

    Humanoid control has not yet demonstrated the scaling gains seen in other foundation models, with existing controllers limited in size, behaviors, and training compute.

  • Method

    SONIC uses large-scale physics-based motion tracking with dense motion-capture supervision, a real-time kinematic planner, and a universal token space for heterogeneous motion inputs.

  • Results

    Performance improves consistently with data, model capacity, and compute, while policies generalize to unseen motions and support robust whole-body control, multimodal interfaces, and VLA-driven loco-manipulation.

  • Takeaways & Limitations

    Large-scale motion tracking is presented as a practical foundation for broad, transferable humanoid whole-body control without per-task reward engineering.

  • Takeaways & Limitations

    The system lacks formal safety and energy-efficiency treatment, and may lose balance under more extreme conditions or highly dynamic motions.

Abstract

from arXiv · show

Despite the rise of billion-parameter foundation models trained across thousands of graphical processing units (GPUs), similar scaling gains have not been shown for humanoid control. Current neural controllers for humanoids remain modest in size, target a limited set of behaviors, and are trained on a handful of GPUs. We show that scaling model capacity, data, and compute yields a generalist humanoid controller capable of natural, robust whole-body movements. We position motion tracking as a scalable task for humanoid control, leveraging dense supervision from diverse motion-capture data to acquire human motion priors without manual reward engineering. We build a foundation model for motion tracking by scaling along three axes: network size (1.2M to 42M parameters), dataset volume (100M+ frames from 700 hours of motion capture), and compute (21k GPU hours). Beyond demonstrating the benefits of scale, we further show downstream utility through a real-time kinematic planner that bridges motion tracking to tasks such as navigation, enabling natural and interactive control, as well as a unified token space that supports virtual reality (VR) teleoperation and vision-language-action (VLA) models with a single policy. Through this interface, we demonstrate autonomous VLA-driven whole-body loco-manipulation requiring coordinated hand and foot placement. Scaling motion tracking exhibits favorable properties: performance improves steadily with compute and data diversity, and learned policies generalize to unseen motions, establishing motion tracking at scale as a practical foundation for humanoid control.

1. Introduction

Humanoid control has not matched the scaling of other foundation models because existing approaches target narrow tasks and depend on task-specific rewards. SONIC instead treats motion tracking as a scalable foundation, then extends one policy across interactive control, heterogeneous interfaces, and whole-body applications.

  • Motivation: Existing humanoid controllers are small, task-specific networks trained with limited compute and manually engineered rewards.These designs do not generalize naturally across behaviors such as locomotion, dancing, recovery, and teleoperation.
  • Scalable foundation: Motion tracking supplies dense frame-by-frame supervision from large, diverse human motion-capture datasets without per-task reward engineering.Available motion data spans walking, running, dancing, sports, and object interactions.
  • Scalable foundation: SONIC scales physics-based motion tracking to 100 million frames using 128-GPU training while maintaining real-time performance across diverse human behaviors.The framework is presented as a universal tracker rather than a controller built for one behavior.
  • Downstream control: A real-time kinematic motion-generation system converts motion tracking into interactive, goal-directed control such as locomotion and game-like character control.This planner provides a bridge from tracking to downstream applications.
  • Unified interfaces: A universal token space maps VR teleoperation, vision-language-action models, and generative motion sources into one shared control interface.The interface accommodates heterogeneous robot, human, and generated motion inputs.
  • Applications: SONIC demonstrates natural teleoperation, multimodal video, text, and music control, plus VLA-driven whole-body loco-manipulation requiring coordinated hand and foot placement.The proposed framework connects large-scale motion tracking to foundation-model-based humanoid control.

2. Results

SONIC scales motion tracking into a general-purpose humanoid controller that generalizes across diverse motions, interactive controls, and real-world deployments. A real-time planner and universal token space extend this capability to multimodal teleoperation and autonomous whole-body loco-manipulation.

  • Motion Tracking: 99.6% success with 23.8 mm MPJPE-L on unseen test-content motions, versus 98.0% and 27.7 mm for the 1.2M-parameter model.Scaling data, model size, and compute consistently improved success on both unseen content and held-out repetitions, with the largest gains on out-of-distribution motions.
  • Motion Tracking: 98.5%/99.2%/97.2% success on test-content/test-repetition/PHUMA exceeded BeyondMimic’s 82.0%/85.4%/73.8% and Any2Track’s 61.6%/69.4%/78.5%.SONIC also achieved 23.7 mm MPJPE-L, a 42% reduction relative to BeyondMimic’s 40.9 mm under the reported cross-dataset protocol.
  • Motion Tracking: 98.5% overall survival versus 43.0% for the OpenHomie specialist across 0–5 m/s velocity commands.SONIC maintained near-100% stability up to approximately 4 m/s, while OpenHomie’s survival rate fell below 20% beyond approximately 2.5 m/s.
  • Motion Tracking: 99.2% real-world success versus 100% in simulation across 124 diverse motion sequences.Overall MPJPE-L was 25.7 mm in the real world versus 22.3 mm in simulation; the largest gap occurred at the feet.
  • Interactive Whole-Body Control: A real-time kinematic planner enabled responsive navigation, boxing, squatting, kneeling, crawling, and smooth transitions between motion segments.The same large-scale dataset trained both the planner and tracking policy, supporting arbitrary navigation velocity, direction, style, and pelvis-height control from 0.3 m to 0.8 m.
  • Multimodal and VLA Control: The universal token space supported video, text, music, and VR control, while VLA control achieved 75% average success across five loco-manipulation tasks.The VLA experiments included coordinated hand grasping and precise foot placement within single action sequences.

3. Materials and Methods

SONIC is evaluated as a scalable whole-body motion-tracking framework using diverse motion-capture data, held-out splits, and multiple control interfaces. Its method combines a universal multi-encoder token space, auxiliary consistency objectives, and a generative kinematic planner, with design choices validated through quantizer and encoder comparisons.

  • Study Design: The study varies data, model size, and compute in simulation, then evaluates policies on held-out motion sets and real-world sequences.The evaluation includes test-content, test-repetition, PHUMA, and real-world trials; scaling curves use six independent evaluations per configuration.
  • Humanoid Motion Dataset: The motion dataset contains diverse behaviors, participants, categories, mirrored clips, and a publicly released subset with annotations and actor information.The collection includes locomotion, daily activities, gestures, combat, object manipulation, tool use, stylistic variations, and role-play.
  • Universal Humanoid Motion Tracking: Specialized encoders map robot, human, and hybrid commands into a shared quantized token space used by a common robot-control decoder.The design supports gamepad control, VR and whole-body teleoperation, video-based teleoperation, and multimodal inputs through motion generators.
  • Universal Humanoid Motion Tracking: Consistency and reconstruction objectives align heterogeneous encoders while training the policy to control the robot and reconstruct robot motion commands.The token loss aligns pairwise encoder outputs, while reconstruction provides auxiliary supervision and functions as a human-to-robot retargeting loss for human motion inputs.
  • Generative Kinematic Motion Planner: The kinematic planner generates motion in latent space by autoregressively in-betweening current and target keyframes, using iterative masked-token prediction.The planner encodes continuous motion into latent tokens and progressively finalizes high-confidence token predictions during inference.
  • Validation of Key Design Choices: FSQ is selected to avoid codebook collapse, while capacity and multi-encoder consistency are tested through quantizer and alignment ablations.FSQ outperformed VQ-VAE on test-content, increased capacity improved performance, and removing consistency losses increased cross-encoder divergence eightfold.

Supplementary Materials for

The supplementary materials describe SONIC’s architecture, training setup, reward design, domain randomization, and quantization choices. They also identify the universal control policy’s hyperparameter and architecture specifications.

  • The supplementary materials document SONIC’s network architecture, training hyperparameters, reward terms, and domain-randomization settings.
  • Table S1 reports the encoder-decoder configuration, universal-token quantizer, robot-control decoders, layer dimensions, activations, and latent sizes.
  • Tables S2–S4 specify PPO settings, reward components, and randomized physical and motion-command parameters.
  • FSQ is used to produce the universal token instead of VQ-VAE because it avoids codebook collapse and auxiliary commitment-loss or EMA updates.FSQ also provides straight-through gradient estimation compatible with joint PPO optimization.

S1.2. Deployment Architecture

SONIC deploys its policy, planner, and management stack onboard a Unitree G1 using a multi-rate architecture and GPU-accelerated inference. Shared motion data structures and safety mechanisms support pre-recorded, planned, streamed, and interactive control.

  • SONIC ran on a 29-joint Unitree G1, outputting desired joint angles directly to the built-in joint-level PD controller.Inference and management ran onboard the Jetson Orin GPU.
  • The deployment stack used concurrent loops for input, planning, control, and command writing, with matched rates and latest-data-wins communication.The control loop operated at 50 Hz, while the command writer operated at 500 Hz.
  • A source-agnostic MotionSequence unified pre-recorded clips, planner trajectories, and streamed joint, SMPL, VR, or token data for the control loop.
  • TensorRT with CUDA Graph acceleration produced 1–2 ms policy forward passes and approximately 12 ms motion-generation calls on Jetson Orin.
  • Startup and safety procedures included a three-second transition to standing, operator readiness confirmation, a joint-velocity watchdog, and fallback when required encoder data was unavailable.
  • The planner generated trajectories at up to 10 Hz, using up to 64 frames of approximately two seconds conditioned on recent context and operator commands.

S1.3. Kinematic Planner Details

The generative kinematic planner produces short-horizon whole-body motion by in-betweening current states and target keyframes. It supports navigation, expressive skills, and interactive manipulation using motion data without retraining.

  • Navigation control: Navigation targets use selected style-matched clip segments placed at target root trajectories, yielding natural and smooth motions aligned with specified styles.
  • Entertainment tasks: Boxing supports expressive target segments and motion layering, allowing upper-body specifications while the planner generates the lower body.
  • Interactive manipulation modes: Interactive squatting and kneeling retrieve keyframes online by desired height, while a single clip can generate the full distribution of transitional motions for a skill.

S1.4. VR Teleoperation and VLA Integration Details

SONIC supports full-body and lightweight VR teleoperation interfaces and connects their motion representations to VLA models. These interfaces provide inputs for navigation, manipulation, and whole-body loco-manipulation.

  • VR-Based Whole-Body Teleoperation: The full-body VR interface uses a PICO headset, two ankle trackers, and handheld controllers to provide SMPL whole-body pose estimates.
  • VR-Based 3-Point Teleoperation: The lightweight 3-point VR interface omits ankle trackers and outputs head and wrist poses, hand angles, waist height, and locomotion commands.
  • 3-Point VLA Integration: A GR00T N1.5 VLA was fine-tuned on 300 teleoperated mobile pick-and-place demonstrations, producing the same control-signal format for the planner and hybrid policy.
  • Whole-Body VLA Integration: For whole-body tasks, the VLA predicts a 78-dimensional action containing a 64-dimensional universal motion token and 14 hand-joint angles.

S2.1. Motion Dataset Samples

Random dataset samples illustrate the breadth of behaviors represented in the large-scale humanoid motion dataset. The training data spans 33 motion categories.

  • Random samples show distinct motion sequences retargeted to the Unitree G1 humanoid.
  • The dataset includes locomotion, dance, sports, combat, gestures, and object interactions.
  • The training data covers 33 motion categories.

S2.2. Qualitative Analysis of Success and Failure Motions

The tracker successfully handles several unseen motions, including dynamic dances and martial-arts movements, but struggles with motions requiring sustained or complex ground contact. These failures occur for extreme floor-level or static poses outside the training distribution.

  • Successful out-of-distribution motions: Unseen hip-hop dance, stage bow, sword lunge, and roundhouse kick motions were successfully tracked.
  • Failure motions: Zombie crawl failed because its extreme floor-level motion exceeds the training distribution.
  • Failure motions: Cross-legged sit failed as a static pose requiring sustained ground contact far from the training distribution.
  • Failure motions: The failures indicate that sustained or complex ground-contact motions remain challenging for the current system.

S2.3. Real-World Evaluation Motions

Real-world evaluation on the Unitree G1 demonstrated successful tracking across 124 diverse motion sequences and resilience to a strong external disturbance. The supplementary scaling comparison shows that specialist locomotion performance can saturate or degrade with additional compute, unlike SONIC’s reported improvements.

  • Real-world evaluation: 124 diverse motion sequences were evaluated on the real Unitree G1 robot, with all shown examples successfully tracked.Examples include hip-hop dance, stage bow, high jump, kick, crouch walk, and grovel.
  • Robustness test: An approximately 11 kg (25 lb) object dropped from above head height was absorbed while the robot maintained balance and continued tracking.No recovery module or policy adaptation was used.
  • Scaling comparison: OpenHomie’s velocity tracking performance peaked at 8 GPUs with 0.18 m/s tracking error before worsening at larger compute scales.At 32 GPUs, the reported error was 0.29 m/s and survival was 91.2%.
  • Scaling comparison: OpenHomie’s survival rate was 91.2% when scaling to 32 GPUs.
  • Evaluation protocol: The OpenHomie evaluation measured mean velocity tracking error and survival rate over n=80 runs across commanded velocities and directions.

S2.6. Latent Space Alignment of the Multi-Encoder Design

The multi-encoder design aligns robot, human, and hybrid motion inputs in a shared token space using consistency losses. Supplementary materials also document the broader interface demonstrations and robustness results associated with the system.

  • Latent space alignment: Robot motion, human SMPL poses, and hybrid teleoperation commands are mapped into a shared token space.
  • Latent space alignment: Removing the consistency losses caused an 8× increase in cross-encoder divergence.The result supports the necessity of these losses for cross-encoder alignment.
  • Supplementary demonstrations: Supplementary Movies S1 to S5 demonstrate motion tracking, interactive control, multimodal interfaces, teleoperation, and foundation-model-driven loco-manipulation.
  • Supplementary demonstrations: The demonstrations include robustness under an approximately 11 kg (25 lb) external perturbation without recovery-module or policy adaptation support.

S3.2. Movie S2: Interactive Motion Control

The universal tracking policy supports interactive whole-body control across planner-driven locomotion, combat, posture changes, teleoperation, text, and music-conditioned motion.

  • Latent Alignment: Consistency losses align tokens from matching frames across encoders, whereas removing them causes cross-encoder alignment to break down.
  • Stylized Locomotion: The kinematic planner replans at 10 Hz, enabling responsive transitions among walking, running, stealth, injured, and sprinting styles.
  • Interactive Boxing: SONIC executes continuous boxing sequences, including jabs, hooks, stance shifts, and reactive upper-body movements.
  • Posture and Crawling: Arbitrary pelvis heights from 0.3–0.8 m support squatting, kneeling, crawling, confined-space locomotion, and downstream teleoperation behaviors.
  • Video Teleoperation: Video-driven pose estimation at ≥60 fps lets the humanoid reproduce walking, running, turning, jumping, dancing, and floor motions in real time.
  • Multimodal Motion: Text and music modalities use the same universal control policy without retraining, including language-commanded actions and rhythm-conditioned choreography.

S3.4. Movie S4: VR Teleoperation

SONIC supports both full-body and lightweight VR teleoperation, with the kinematic planner generating lower-body motion from upper-body inputs. The same universal token interface also enables autonomous VLA-driven whole-body manipulation tasks.

  • Full-Body VR Control: Full-body VR streams from a headset, ankle trackers, and controllers drive walking, sidestepping, crouching, and loco-manipulation with low latency.
  • 3-Point VR Teleoperation: The 3-point interface uses only a headset and handheld controllers, while the kinematic planner generates the lower body for mobile manipulation and navigation.
  • VLA Control: A GR00T N1.5 VLA model drives SONIC through the universal token interface across five autonomous tasks requiring object interaction, navigation, and foot placement.
  • Data Availability: The quantitative source data for main-text figure panels are provided separately in the machine-readable Data S1 spreadsheet.
Loading 2511.07820v4…