Source-linked AI summary
GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
GigaBrain Team, Angen Ye, Axiang Sun, Can Jin, Chenxi Cheng, Chong Shi, Dengke Shang, Dingqian Zhang, Guan Huang, Guangqiang Wang, Guangqing Ding, Guo Li, Hangcong Li, Hengyu Zhong, Hongtao Lu, Jianbo Qin, Jiming Mao, Jing Zhu, Jindi Lv, Jingzhi Cui, Junjie Xie, Junyi Bao, Kai Liu, Lei Yuan, Limin Long, Lv Feng, Mingming Yu, Peng Li, Pengfei Yi, Qi Li, Qianli Zhang, Qingfang Li, Qitang Hu, Rui Zhang, Shaoyan Sun, Shibo Sun, Shiying Duan, Tenghui Chen, Tianze Liu, Weijie Ke, Wenyao Xue, Xiaofeng Wang, Xiaoyu Tian, Xinyu Liu, Xinze Chen, Yang Wang, Yankai Wang, Yejun Zeng, Yifan Li, Yifei Nie, Yilong Li, Yilong Liu, Yongchao Feng, Yumeng Wang, Yun Ye, Zhichao Liu, Ziheng He, Zonghai Yang, Zheng Zhu
TL;DR
Existing VLA systems leave open how to scale heterogeneous embodied data and improve generalization across tasks and robot embodiments. GigaBrain-0.7 addresses this with a three-system architecture and one-stage joint alignment, reporting improved foundation, instruction-following, and post-training capabilities while retaining limitations under demanding out-of-distribution shifts.
Problem
Current VLA systems face scaling challenges because robot data differ substantially across embodiments, action spaces, and execution conditions.
Method
GigaBrain-0.7 uses a three-system architecture and one-stage multi-embodiment VLA pretraining that jointly optimizes understanding, prediction, and action.
Results
GigaBrain-0.7 reports improved foundation capabilities, language-conditioned instruction following, and post-training task success across diverse robot embodiments.
Takeaways & Limitations
The coordinated design supports task adaptability and completion across home and industrial scenarios on Maker H01 and mainstream robot embodiments.
Takeaways & Limitations
Generalization degrades when object identity, scene configuration, and task composition shift simultaneously, especially in complex out-of-distribution manipulation.
Abstract
from arXiv · showhide
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including $π_{0.5}$, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.
1. Introduction
GigaBrain-0.7 targets the challenges of scaling VLA systems across heterogeneous robot data, embodiments, and long-horizon tasks. It combines one-stage alignment with a three-system architecture coordinating understanding, prediction, action, and experience-driven improvement.
- Heterogeneous robot data vary across embodiments, action spaces, and execution conditions, making alignment and contextualization central scaling challenges.
- GigaBrain-0.7 jointly trains a pretrained VLM backbone, continuous-action expert, and multimodal understanding objectives in one stage.This avoids the optimization fragmentation associated with multi-stage training pipelines.
- System 1 generates actions, System 2 performs visual-language understanding and hierarchical planning, and System 3 predicts future observations and task-related values.
- The system coordinates heterogeneous experience through unified processing, Soft Knowledge Insulation, and interfaces for offline and online experience reinforcement.Soft Knowledge Insulation attenuates action gradients entering the VLM backbone rather than blocking them completely.
- Larger-scale pretraining consistently reduces validation loss and makes challenging real-robot behaviors increasingly reliable, with results indicating emergent embodied capabilities.The evidence combines scaling studies, System 3 ablations, and out-of-distribution evaluations.
2. Related Work
VLA architectures differ in how they represent actions and couple continuous action generation to pretrained vision-language models. Recent work increasingly combines joint training, richer context, and task-specific post-training, while world models support data generation and unified world-action modeling.
- VLA Architectures: Existing VLA architectures comprise autoregressive VLM-as-Actor, serial VLM–action-expert, and parallel Mixture-of-Transformers families.These families differ in action representation and in how action generation interacts with the pretrained VLM.
- VLA Architectures: Parallel Mixture-of-Transformers models retain continuous Action Experts while coupling them more deeply to vision-language backbones.π0 established block-wise causal interaction with flow-matching action generation, a lineage extended by π0.5 and π0.7.
- Training Strategies: Training recipes range from staged adaptation to joint optimization of vision-language and action objectives within one VLA pretraining stage.GigaBrain-0.7 jointly optimizes general vision-language supervision, hierarchical task prediction, discrete action representations, and continuous action generation.
- Training Strategies: GigaBrain-0.7 scales the joint-training direction to heterogeneous multi-embodiment data while unifying multiple prediction and action objectives.Its approach follows earlier single-stage formulations but expands the training scope across embodiments and supervision types.
- Post-training: Post-training specializes pretrained manipulation representations to target embodiments, tasks, and deployment conditions, with reinforcement-learning methods extending beyond demonstrations.Recent methods use environment feedback, offline experience, advantage information, or corrective trajectories for further adaptation.
- World Models: World models support embodied learning through scalable data generation, unified world-action modeling, and predictive or evaluative signals for policy learning.Prior work uses generated interaction trajectories, future video prediction, robot-action prediction, and combined video-action diffusion.
3. Data Curation
GigaBrain-0.7’s data curation combines heterogeneous embodied trajectories with large-scale vision-language data, then standardizes representations and applies quality control for joint training. Simulation and world-model generation complement physical collection by broadening long-tail interaction coverage.
- Corpus Composition: The corpus combines embodied trajectories with VLM image–text/question-answering data to support instruction understanding, perception, reasoning, and continuous control.Embodied data provide state–action sequences and robot-control supervision, while VLM data preserve visual-semantic capabilities.
- Corpus Composition: 37,256.98 hours of embodied data span real robots, UMI demonstrations, EGO human demonstrations, simulation, and world-model-generated data.The real-robot subset alone covers 16 robot types, 1,810,101 episodes, 2,077,837,071 frames, and 20,535.65 hours.
- Human Demonstrations: 8,251.83 hours of UMI data provide near-manipulation-viewpoint human demonstrations, while 2,862.36 hours of EGO data broaden object, scene, and interaction coverage.These sources complement robot trajectories with human manipulation viewpoints, semantic understanding, and temporal modeling of manipulation processes.
- Synthetic and Simulated Data: 5,607.14 hours of simulation and generated data comprise 1,453.92 simulation hours and 4,153.22 world-model-generated hours.Physics-based simulation offers configurable task and scene control, while world-model generation derives data from real interaction samples.
- VLM Data: 271,976,674 VLM image–text/question-answering samples strengthen spatial understanding, object localization, affordance reasoning, and embodied visual reasoning.The corpus includes 15,234,327 self-collected samples and covers captioning, VQA, multi-image reasoning, grounding, and point prediction.
- Processing Pipeline: A unified processing pipeline converts heterogeneous sources into standardized representations and applies multi-stage quality control before training.It includes LeRobot v3.0 conversion, unified robot state and action representations, camera mapping, instruction rewriting, subtask annotation, and anomaly filtering.
- Synthetic and Simulated Data: World-model generation expands real-interaction diversity, whereas physics-based simulation controls task and scene configurations, together broadening long-tail interaction coverage.Their combination adds experience beyond what can be collected efficiently from physical robots alone.
4. Model Architecture
GigaBrain-0.7 uses three coordinated systems to connect task understanding, future prediction, evaluation, and continuous robot control. Its action-alignment design combines multimodal context, temporal memory, and embodiment-aware action generation.
- Three-System Architecture: System 2 interprets observations and decomposes instructions, System 3 predicts future states and task progress, and System 1 generates actions.System 2 produces subtasks; System 3 supplies future observations and value estimates to System 1.
- World Simulation: The world-simulation layer gives System 1 both a predicted subgoal image and a scalar progress signal beyond the current observation.The subgoal image represents the expected short-horizon physical state, while the value estimate evaluates task progress.
- Three-System Architecture: System 1 fuses observations, visual memory, subtasks, robot state, Robot ID, and predictive signals into action chunks for closed-loop execution.The action-alignment layer integrates semantic planning, predictive information, temporal context, and embodiment-specific states.
- Action Alignment: The Mixture-of-Transformers couples vision-language and action-expert streams through joint attention while keeping their feed-forward parameters separate.The vision-language stream preserves causal self-attention, while the action expert attends bidirectionally to both streams.
- Temporal Context: Temporal-Spatial Blocks causally aggregate recent frames, then retain only enriched current-frame tokens so temporal context does not naively expand high-level context length.This addresses task-state ambiguity from occlusions and repeated visual configurations.
- Action Alignment: Soft Knowledge Insulation attenuates action-learning gradients entering the vision-language backbone while preserving its autoregressive supervision.Continuous-action learning can therefore adapt representations toward embodied control without fully blocking action gradients.
5. Model Training
Training combines joint vision-language-action pretraining, world-value-model construction, task-specific conditioning, and experience-driven refinement. The recipe supports heterogeneous embodiments and data while coupling semantic, hierarchical, and continuous-action objectives.
- Training Pipeline: Training proceeds through joint Systems 1–2 pretraining, System 3 world-value-model training, frozen-System-3 post-training, and offline and online reinforcement learning.System 3 is frozen during downstream VLA post-training while Systems 1 and 2 are optimized.
- VLA Pre-Training: Systems 1 and 2 jointly optimize semantic understanding, hierarchical task prediction, and continuous action generation across the full heterogeneous corpus.The corpus includes real-robot, UMI, EGO, simulation, generated trajectories, and vision-language data.
- VLA Pre-Training: The NTP pathway predicts language, subtasks, discrete actions, and vision-language outputs, while the FM pathway predicts continuous action chunks.Both pathways are optimized throughout pretraining rather than introducing continuous action generation only later.
- VLA Pre-Training: Soft Knowledge Insulation attenuates flow-matching gradients before they enter the vision-language backbone while retaining full autoregressive supervision.The Action Expert receives full flow-matching gradients, whereas the VLM adapts through a reduced gradient contribution.
- Multi-Embodiment Optimization: Multi-embodiment optimization maps heterogeneous states and actions into a shared semantic structure while preserving embodiment-specific validity information.Validity masks exclude unavailable state dimensions and invalid action dimensions from supervision and noise injection.
- World-Model Training: System 3 is extended through robot-centric video pretraining and value learning to provide future-state prediction and task-progress estimation.The predicted video horizon matches System 1’s action-chunk horizon, and the final frame becomes the subgoal image.
- Experience-Driven Reinforcement Learning: Offline reinforcement learning weights higher-progress experience more strongly, improving behavior initialization while reducing exploration required for online learning.Successful and higher-progress segments receive greater training weight, while failed or lower-progress behaviors are down-weighted.
6. Experiments
The experiments establish evaluation protocols across two substantially different robot embodiments, then identify a robust System 1 architecture and characterize how data scale and composition affect performance. Larger and more heterogeneous pretraining data consistently improves validation and task behavior, especially for complex manipulation.
- Evaluation setup: Evaluations use AgileX PiPER/PiPER-X and Maker H01, whose differing degrees of freedom, interfaces, cameras, and workspaces test cross-embodiment generalization.
- Backbone comparison: PaliGemma2 is selected because it remains competitive on fruit manipulation and is the only evaluated backbone with nonzero shirt-folding success.
- Architecture comparison: Dual Stream delivers the strongest real-robot performance, while final-layer coupling has the lowest latency and multi-layer coupling adds cost without improving task success.
- Data scaling: Larger training sets consistently achieve lower validation loss, with improvement continuing across the evaluated pretraining-data scales.
- Data scaling: Complex clothes folding depends more sharply on data scale than fruit manipulation, reflecting its broader intermediate-state coverage and sustained deformable-object interaction.
- Data composition: Combining robot, UMI, and EGO data produces the strongest evaluated results, while human-data benefits persist after task-specific post-training.
6.4. Temporal Context Analysis
Temporal context helps disambiguate visually similar manipulation states, while System 3 combines future-state prediction and value conditioning to guide progress. These signals improve task scores, completion efficiency, and difficult-task success, although foundation transfer still degrades under broader distribution shifts.
- Temporal context: Single-frame policies can repeat actions when similar visual states occur at different execution stages, whereas temporal context helps the policy advance to the next task stage.
- System 3: System 3 predicts future visual observations and state values, using value predictions to estimate the advantage of candidate action trajectories.
- System 3 ablation: 88.3% average score and 75 s completion time are achieved by +SubImage+Value on clothes folding, versus 68.3% and 107 s for Base despite 100% success for all variants.
- System 3 ablation: 80% success is achieved by +SubImage+Value on gift wrapping, compared with 0% for Base; its task score reaches 96.7%.
- System 3 ablation: 55% success and an 87.5% average task score are achieved by the combined configuration on cube sorting, while value-only conditioning has the shortest observed completion time at 80 s.
- Interpretation: System 3 provides predictive and progress-aware guidance that remains useful when binary success saturates and is more pronounced on challenging gift wrapping.
- Limitations: Foundation transfer weakens when object identity, scene configuration, and task composition shift simultaneously, with degradation on both evaluated embodiments.
6.7. Post-Training Evaluation
Post-training evaluation shows consistent gains over the preceding GigaBrain model across language-following and complex-manipulation tasks on AgileX PiPER and Maker H01. The evaluations span fine-grained instruction grounding, sustained interaction, deformable objects, tool use, and multi-stage execution.
- Multi-task language following: Post-training evaluates six language-conditioned tasks covering color, target, and directional discrimination on AgileX PiPER and Maker H01.
- Multi-task language following: GigaBrain-0.7 consistently improves over the preceding GigaBrain model and matches or exceeds π0.5 across reported PiPER language-following tasks.
- Multi-task language following: 84.2% average success is achieved on Maker H01, up from 69.6% for GigaBrain-0.1, with gains on button push and spoon grasping.
- Instruction grounding: Post-trained policies ground color, object identity, and relative spatial relations in target selection across both embodiments.
- Complex manipulation: Complex-manipulation evaluation covers sustained contact, multi-stage execution, deformable objects, tool-mediated interaction, sorting, organization, food tasks, and cleanup.
- Complex manipulation: GigaBrain-0.7 consistently improves over GigaBrain-0.1 and remains competitive with or stronger than π0.5 across completed complex-manipulation evaluations.
- Complex manipulation: On Maker H01, gains are larger across deformable-object, food-related, and multi-stage cleanup skills, while several contact-rich tasks remain unsaturated.
6.8. Benchmark Evaluation
GigaBrain-0.7 is evaluated across embodied vision-language understanding, simulation, and experience-driven real-robot improvement. It shows strong performance across varied embodiments and task settings, while staged reinforcement learning substantially improves execution success.
- Embodied Vision-Language Evaluation: GigaBrain-0.7 shows strong embodied vision-language capability before MiMo data are added, achieving the strongest overall and Spatial Understanding averages among compared models.MiMo-Embodied covers spatial understanding and affordance reasoning across 13 benchmarks.
- Embodied Vision-Language Evaluation: The final MiMo-enhanced checkpoint reaches an overall MiMo score of 0.5704, but this is treated as a continued-pretraining diagnostic rather than a held-out result.The later training mixture is no longer disjoint from the MiMo evaluation set.
- Simulation Benchmark Evaluation: GigaBrain-0.7 achieves the strongest overall RoboTwin 2.0 performance and ranks first under the challenging Hard setting, retaining stronger performance after domain randomization.The evaluation uses the official multi-task Co-Train protocol and contrasts clean Easy with domain-randomized Hard settings.
- Simulation Benchmark Evaluation: Among publicly available models, GigaBrain-0.7 achieves the strongest EBench overall success rate and aggregate task score, outperforming representative baselines including π0 and π0.5.EBench covers mobile, long-horizon, dexterous, and precise interaction beyond fixed-base bimanual manipulation.
- Simulation Benchmark Evaluation: GigaBrain-0.7 leads all four evaluated RoboColiseum dimensions: instruction following, spatial reasoning, robustness, and general manipulation.Together, the three simulation regimes cover complementary task distributions and evaluation settings.
- Experience-Driven Reinforcement Learning: Average success rises from 30.0% with SFT to 57.5% after offline RL and 100% after online RL across four real-robot tasks.Offline gains are largest where supervised policies encounter substantial failures, while online rollouts target remaining difficult states with human corrections.
7. Conclusion and Future Work
GigaBrain-0.7 scales heterogeneous embodied experience through a three-system architecture that coordinates understanding, prediction, and action. The paper reports strong generalization, robustness, and post-training performance across diverse tasks and embodiments, while identifying further scaling and closed-loop learning as future priorities.
- Conclusion: GigaBrain-0.7 is pretrained on over 37,000 hours of embodied trajectories spanning 16 robot morphologies, with joint multimodal understanding, prediction, and action optimization.Its one-stage VLA training covers hierarchical task prediction, discrete action supervision, and continuous action generation.
- Conclusion: System 3 is separately pretrained for future-state prediction and task-progress estimation, then conditions task-specific post-training and inference.The outputs provide future-state and progress-related guidance for policy improvement.
- Conclusion: Across real-robot, embodied vision-language, and simulation evaluations, GigaBrain-0.7 demonstrates strong out-of-the-box generalization, post-training performance, and robustness across diverse tasks and embodiments.Experience-driven offline and online reinforcement learning further improves the policy using rollout experience and corrective feedback.
- Future Work: Future work targets larger heterogeneous datasets, stronger cross-source alignment, improved long-horizon prediction, and more scalable closed-loop learning from autonomous experience and human correction.These directions are intended to support more general, robust, and continually improving real-world robot intelligence.
A. Simulation Benchmark Leaderboard Snapshots
The appendix preserves complete simulation benchmark leaderboard snapshots to complement the representative comparisons in the main text. Each snapshot includes source and access-date information because online leaderboards can change as submissions are added.
- Leaderboard Snapshots: Complete leaderboard snapshots include additional submissions not shown in the main-text tables, preserving leaderboard status at manuscript preparation.The main text focuses on representative publicly available methods for concise and reproducible comparison.
- Leaderboard Snapshots: Online leaderboards may change as new submissions are added, so each benchmark snapshot provides its source URL and access date.This records the temporal status of the reported comparisons.
A.1. RoboTwin 2.0
The RoboTwin 2.0 appendix reports the official leaderboard protocol and performance under clean Easy and domain-randomized Hard settings. GigaBrain-0.7 is evaluated using the official Co-Train protocol used for the comparison.
- RoboTwin 2.0: The RoboTwin 2.0 leaderboard reports evaluation under clean Easy and domain-randomized Hard settings.The snapshot presents protocol information together with performance for both conditions.
- RoboTwin 2.0: GigaBrain-0.7 is evaluated under the official Co-Train protocol used for the comparison in Table 9.This situates the reported model performance within the benchmark’s official evaluation procedure.
A.2. EBench
The EBench materials document an official leaderboard snapshot reporting overall task success rate and aggregate task score. The record includes submissions beyond the representative main-text comparison and distinguishes one reproduced result from the online snapshot.
- Qwen-RobotManip remains in the complete official leaderboard record despite not appearing in the representative public model comparison.
- The InternVLA-A1 EBench result is reproduced from Qwen-RobotManip rather than taken from the online leaderboard snapshot.The main-text table marks this result with a superscript asterisk.
- The complete official EBench leaderboard reports overall task success rate and aggregate task score for all listed submissions.The snapshot reflects submissions available at manuscript preparation.
A.3. RoboColiseum
RoboColiseum reports four complementary capability leaderboards, with separate snapshots documenting instruction following, spatial reasoning, robustness, and general manipulation. The materials provide the official comparison source and its access date.
- RoboColiseum separates evaluation into instruction following, spatial reasoning, robustness, and general manipulation leaderboards.Figures 30–33 correspond to these four dimensions in the comparison described for Tab. 11.
- Instruction Following: The instruction-following leaderboard evaluates whether robot policies follow language-conditioned instructions.
- Spatial Reasoning: The spatial-reasoning leaderboard evaluates spatial reasoning capability.
- Robustness: The robustness leaderboard evaluates policy robustness under the corresponding RoboColiseum protocol.
- General Manipulation: The general-manipulation leaderboard evaluates general manipulation capability.