Source-linked AI summary

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu, Yu Sun, Jing Liu, Dianhai Yu, Yanjun Ma, Jingzhou He, Zhongjun He, Dou Hong, Qiwen Liu, Shuohuan Wang, Junyuan Shang, Zhenyu Zhang, Yuchen Ding, Jinle Zeng, Jiabin Yang, Liang Shen, Ruibiao Chen, Weichong Yin, Siyu Ding, Dai Dai, Shikun Feng, Siqi Bao, Bolei He, Yan Chen, Zhenyu Jiao, Ruiqing Zhang, Zeyu Chen, Qingqing Dang, Kaipeng Deng, Jiajun Jiang, Enlei Gong, Guoxia Wang, Yanlin Sha, Yi Liu, Yehan Zheng, Weijian Xu, Jiaxiang Liu, Zengfeng Zeng, Yingqi Qu, Zhongli Li, Zhengkun Zhang, Xiyang Wang, Zixiang Xu, Xinchao Xu, Zhengjie Huang, Dong Wang, Bingjin Chen, Yue Chang, Xing Yuan, Shiwei Huang, Qiao Zhao, Xinzhe Ding, Shuangshuang Qiao, Baoshan Yang, Bihong Tang, Bin Li, Bingquan Wang, Binhan Tang, Binxiong Zheng, Bo Cui, Bo Ke, Bo Zhang, Bowen Zhang, Boyan Zhang, Boyang Liu, Caiji Zhang, Can Li, Chang Xu, Chao Pang, Chao Zhang, Chaoyi Yuan, Chen Chen, Cheng Cui, Chenlin Yin, Chun Gan, Chunguang Chai, Chuyu Fang, Cuiyun Han, Dan Zhang, Danlei Feng, Danxiang Zhu, Dong Sun, Dongbo Li, Dongdong Li, Dongdong Liu, Dongxue Liu, Fan Ding, Fan Hu, Fan Li, Fan Mo, Feisheng Wu, Fengwei Liu, Gangqiang Hu, Gaofeng Lu, Gaopeng Yong, Gexiao Tian, Guan Wang, Guangchen Ni, Guangshuo Wu, Guanzhong Wang, Guihua Liu, Guishun Li, Haibin Li, Haijian Liang, Haipeng Ming, Haisu Wang, Haiyang Lu, Haiye Lin, Han Zhou, Hangting Lou, Hanwen Du, Hanzhi Zhang, Hao Chen, Hao Du, Hao Liu, Hao Zhou, Haochen Jiang, Haodong Tian, Haoshuang Wang, Haozhe Geng, Heju Yin, Hong Chen, Hongchen Xue, Hongen Liu, Honggeng Zhang, Hongji Xu, Hongwei Chen, Hongyang Zhang, Hongyuan Zhang, Hua Lu, Huan Chen, Huan Wang, Huang He, Hui Liu, Hui Zhong, Huibin Ruan, Jiafeng Lu, Jiage Liang, Jiahao Hu, Jiahao Hu, Jiajie Yang, Jialin Li, Jian Chen, Jian Wu, Jianfeng Yang, Jianguang Jiang, Jianhua Wang, Jianye Chen, Jiaodi Liu, Jiarui Zhou, Jiawei Lv, Jiaxin Zhou, Jiaxuan Liu, Jie Han, Jie Sun, Jiefan Fang, Jihan Liu, Jihua Liu, Jing Hu, Jing Qian, Jing Yan, Jingdong Du, Jingdong Wang, Jingjing Wu, Jingyong Li, Jinheng Wang, Jinjin Li, Jinliang Lu, Jinlin Yu, Jinnan Liu, Jixiang Feng, Jiyi Huang, Jiyuan Zhang, Jun Liang, Jun Xia, Jun Yu, Junda Chen, Junhao Feng, Junhong Xiang, Junliang Li, Kai Liu, Kailun Chen, Kairan Su, Kang Hu, Kangkang Zhou, Ke Chen, Ke Wei, Kui Huang, Kun Wu, Kunbin Chen, Lei Han, Lei Sun, Lei Wen, Linghui Meng, Linhao Yu, Liping Ouyang, Liwen Zhang, Longbin Ji, Longzhi Wang, Meng Sun, Meng Tian, Mengfei Li, Mengqi Zeng, Mengyu Zhang, Ming Hong, Mingcheng Zhou, Mingming Huang, Mingxin Chen, Mingzhu Cai, Naibin Gu, Nemin Qiu, Nian Wang, Peng Qiu, Peng Zhao, Pengyu Zou, Qi Wang, Qi Xin, Qian Wang, Qiang Zhu, Qianhui Luo, Qianwei Yang, Qianyue He, Qifei Wu, Qinrui Li, Qiwen Bao, Quan Zhang, Quanxiang Liu, Qunyi Xie, Rongrui Zhan, Rufeng Dai, Rui Peng, Ruian Liu, Ruihao Xu, Ruijie Wang, Ruixi Zhang, Ruixuan Liu, Runsheng Shi, Ruting Wang, Senbo Kang, Shan Lu, Shaofei Yu, Shaotian Gong, Shenwei Hu, Shifeng Zheng, Shihao Guo, Shilong Fan, Shiqin Liu, Shiwei Gu, Shixi Zhang, Shuai Yao, Shuang Zhang, Shuangqiao Liu, Shuhao Liang, Shuwei He, Shuwen Yang, Sijun He, Siming Dai, Siming Wu, Siyi Long, Songhe Deng, Suhui Dong, Suyin Liang, Teng Hu, Tianchan Xu, Tianliang Lv, Tianmeng Yang, Tianyi Wei, Tiezhu Gao, Ting Sun, Ting Zhang, Tingdan Luo, Wei He, Wei Luan, Wei Yin, Wei Zhang, Wei Zhou, Weibao Gong, Weibin Li, Weicheng Huang, Weichong Dang, Weiguo Zhu, Weilong Zhang, Weiqi Tan, Wen Huang, Wenbin Chang, Wenjing Du, Wenlong Miao, Wenpei Luo, Wenquan Wu, Xi Shi, Xi Zhao, Xiang Gao, Xiangguo Zhang, Xiangrui Yu, Xiangsen Wang, Xiangzhe Wang, Xianlong Luo, Xianying Ma, Xiao Tan, Xiaocong Lin, Xiaofei Wang, Xiaofeng Peng, Xiaofeng Wu, Xiaojian Xu, Xiaolan Yuan, Xiaopeng Cui, Xiaotian Han, Xiaoxiong Liu, Xiaoxu Fei, Xiaoxuan Wu, Xiaoyu Wang, Xiaoyu Zhang, Xin Sun, Xin Wang, Xinhui Huang, Xinming Zhu, Xintong Yu, Xinyi Xu, Xinyu Wang, Xiuxian Li, XuanShi Zhu, Xue Xu, Xueying Lv, Xuhong Li, Xulong Wei, Xuyi Chen, Yabing Shi, Yafeng Wang, Yamei Li, Yan Liu, Yanfu Cheng, Yang Gao, Yang Liang, Yang Wang, Yang Wang, Yang Yang, Yanlong Liu, Yannian Fu, Yanpeng Wang, Yanzheng Lin, Yao Chen, Yaozong Shen, Yaqian Han, Yehua Yang, Yekun Chai, Yesong Wang, Yi Song, Yichen Zhang, Yifei Wang, Yifeng Guo, Yifeng Kou, Yilong Chen, Yilong Guo, Yiming Wang, Ying Chen, Ying Wang, Yingsheng Wu, Yingzhan Lin, Yinqi Yang, Yiran Xing, Yishu Lei, Yixiang Tu, Yiyan Chen, Yong Zhang, Yonghua Li, Yongqiang Ma, Yongxing Dai, Yongyue Zhang, Yu Ran, Yu Sun, Yu-Wen Michael Zhang, Yuang Liu, Yuanle Liu, Yuanyuan Zhou, Yubo Zhang, Yuchen Han, Yucheng Wang, Yude Gao, Yuedong Luo, Yuehu Dong, Yufeng Hu, Yuhui Cao, Yuhui Yun, Yukun Chen, Yukun Gao, Yukun Li, Yumeng Zhang, Yun Fan, Yun Ma, Yunfei Zhang, Yunshen Xie, Yuping Xu, Yuqin Zhang, Yuqing Liu, Yurui Li, Yuwen Wang, Yuxiang Lu, Zefeng Cai, Zelin Zhao, Zelun Zhang, Zenan Lin, Zezhao Dong, Zhaowu Pan, Zhaoyu Liu, Zhe Dong, Zhe Zhang, Zhen Zhang, Zhengfan Wu, Zhengrui Wei, Zhengsheng Ning, Zhenxing Li, Zhenyu Li, Zhenyu Qian, Zhenyun Li, Zhi Li, Zhichao Chen, Zhicheng Dong, Zhida Feng, Zhifan Feng, Zhihao Deng, Zhijin Yu, Zhiyang Chen, Zhonghui Zheng, Zhuangzhuang Guo, Zhujun Zhang, Zhuo Sun, Zichang Liu, Zihan Lin, Zihao Huang, Zihe Zhu, Ziheng Zhao, Ziping Chen, Zixuan Zhu, Ziyang Xu, Ziyi Liang, Ziyuan Gao

arXiv:2602.04705v1cs.CL

TL;DR

Existing autoregressive multimodal systems remain largely text-centric at output, limiting multimodal interaction. ERNIE 5.0 addresses this with a from-scratch unified autoregressive MoE model, elastic training, and multimodal post-training, achieving strong cross-task performance while supporting scalable deployment. A remaining boundary is a moderate gap on the most challenging reasoning benchmarks compared with Gemini 3-Pro.

  • Problem

    Existing autoregressive systems mainly support multimodal understanding while remaining text-centric in output, limiting broader multimodal interaction.

  • Method

    ERNIE 5.0 trains all modalities from scratch in a unified autoregressive MoE framework, adds elastic sub-model training, and uses multimodal reinforcement learning post-training.

  • Results

    Across diverse perception, reasoning, understanding, and generation benchmarks, ERNIE 5.0 matches or outperforms specialized baselines; elastic configurations retain near-full performance with 53.7% activated parameters and 35.8% total parameters.

  • Takeaways & Limitations

    ERNIE 5.0 provides a unified trillion-level autoregressive foundation model with flexible deployment configurations for multimodal understanding and generation.

  • Takeaways & Limitations

    A moderate performance gap persists on the most challenging reasoning benchmarks compared with Gemini 3-Pro.

Abstract

from arXiv · show

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio. All modalities are trained from scratch under a unified next-group-of-tokens prediction objective, based on an ultra-sparse mixture-of-experts (MoE) architecture with modality-agnostic expert routing. To address practical challenges in large-scale deployment under diverse resource constraints, ERNIE 5.0 adopts a novel elastic training paradigm. Within a single pre-training run, the model learns a family of sub-models with varying depths, expert capacities, and routing sparsity, enabling flexible trade-offs among performance, model size, and inference latency in memory- or time-constrained scenarios. Moreover, we systematically address the challenges of scaling reinforcement learning to unified foundation models, thereby guaranteeing efficient and stable post-training under ultra-sparse MoE architectures and diverse multimodal settings. Extensive experiments demonstrate that ERNIE 5.0 achieves strong and balanced performance across multiple modalities. To the best of our knowledge, among publicly disclosed models, ERNIE 5.0 represents the first production-scale realization of a trillion-parameter unified autoregressive model that supports both multimodal understanding and generation. To facilitate further research, we present detailed visualizations of modality-agnostic expert routing in the unified model, alongside comprehensive empirical analysis of elastic training, aiming to offer profound insights to the community.

1 Introduction

ERNIE 5.0 unifies multimodal understanding and generation through from-scratch autoregressive training, sparse modality-agnostic routing, elastic sub-model training, and scalable post-training. Across diverse benchmarks, it matches or exceeds specialized baselines while supporting efficiency–performance trade-offs.

  • Motivation and contribution: ERNIE 5.0 trains text, image, audio, and video from scratch under a unified autoregressive framework for multimodal understanding and generation.Its design avoids augmenting a pre-trained language model with modality-specific components.
  • Unified modeling: Heterogeneous modalities are mapped into a shared token space and optimized with a unified Next-Group-of-Tokens Prediction objective.Modality-agnostic routing dispatches tokens from different modalities to a shared expert pool.
  • Elastic training: Elastic training samples sub-models with varying depth, width, and routing sparsity while jointly optimizing them with the full model in one backpropagation process.This produces multiple capacity–efficiency configurations from a single pre-training run.
  • Post-training: Unified multimodal reinforcement learning combines supervised fine-tuning with techniques addressing sampling bias, sparse rewards, and entropy collapse in ultra-sparse MoE training.The pipeline includes a unified verifier system and scalable stability techniques.
  • Evaluation: ERNIE 5.0 matches or outperforms specialized baselines across perception, reasoning, understanding, and generation benchmarks.Routing top-k reduction yields over 15% decoding speedup with minor accuracy loss, while elastic training retains near-full performance using 53.7% activated and 35.8% total parameters.

2 Architecture

ERNIE 5.0 uses one shared autoregressive backbone to model language, image, video, and audio. Visual and audio tokenizers convert multimodal inputs into a unified sequence trained with a shared prediction objective.

  • Unified architecture: ERNIE 5.0 integrates language, image, video, and audio in a single autoregressive framework for understanding and generation.The architecture combines a shared backbone with visual and audio tokenizers.
  • Unified architecture: All modalities are converted into a unified token sequence and trained under a shared Next-Group-of-Tokens Prediction objective.The objective supports cross-modal interactions with end-to-end optimization.

2.1 Unified Autoregressive Backbone with Ultra-Sparse Mixture-of-Experts

The unified backbone addresses cross-modal differences by combining shared autoregressive objectives with sparse capacity scaling. Its design supports both understanding and generation while using modality-agnostic expert routing.

  • Motivation: Heterogeneous modalities differ in token semantics, temporal structure, and optimization dynamics, making naive parameter sharing prone to instability and degradation.The challenge is especially pronounced when understanding and generation are modeled together.
  • Unified objective: ERNIE 5.0 projects text, image, video, and audio into a shared sequence and optimizes them with Next-Group-of-Tokens Prediction.Text uses next-token prediction augmented by multi-token prediction, while vision and audio use grouped-token objectives.
  • Sparse scaling: Sparse Mixture-of-Experts scaling increases model capacity while controlling training and inference costs.The architecture is intended to capture diverse multimodal knowledge efficiently.
  • Sparse scaling: Modality-agnostic routing conditions expert selection on unified token representations rather than modality identifiers.This supports shared processing across modalities without modality-specific routing decisions.
  • Understanding and generation: The shared autoregressive backbone formally integrates multimodal understanding and generation despite their different representational demands.Understanding emphasizes abstract semantics, whereas generation requires fine-grained perceptual detail.

2.2 Visual Modeling

ERNIE 5.0 uses hybrid visual representations for understanding and a unified autoregressive NFSP paradigm for image and video generation. Attention-based merging compresses CNN and ViT features into compact representations for multimodal tasks.

  • Vision tokenization: Image is treated as a single-frame video, allowing visual understanding and generation to share one design philosophy.The hybrid representation combines global semantic information with local perceptual details.
  • Vision tokenization: NFSP formulates image generation as Next-Scale Prediction and extends it to video with Next-Frame Prediction over time.The visual autoregressive process models spatial and temporal dimensions in discrete token space.
  • Vision tokenization: Visual tokenizers use adversarial and semantic supervision to improve distributional fidelity, semantic consistency, learnability, and stability.These objectives support effective autoregressive modeling in the unified backbone.
  • Vision tokenization: Progressive tokenizer switching begins with a low-bit tokenizer and gradually increases the visual vocabulary during training.The tokenizer series uses progressively increasing bit numbers.
  • Visual understanding: Attention-based patch merging combines CNN perceptual features with ViT semantic features after aligning their representations.For images it groups 4 adjacent patches, while video groups 16 patches across 4 neighboring frames before attention and pooling.
  • Visual understanding: The attention-based module outperforms CNN-only and ViT-only baselines across benchmarks without noticeable computational overhead.Gains are particularly pronounced in document, chart, and general visual understanding tasks.
  • Visual generation: NFSP separates intra-frame multiscale prediction from inter-frame temporal prediction, while Uni-RoPE encodes token time and spatial coordinates.Historical-token corruption during training improves robustness to error accumulation in long visual sequences.
  • Visual generation: A cascaded diffusion refiner addresses the token-budget challenge of high-resolution image and video generation on top of the autoregressive backbone.The backbone first generates low-resolution samples.

2.3 Audio Modeling

ERNIE 5.0 represents audio as hierarchical discrete codec tokens and models them with a depth-wise autoregressive architecture for understanding and generation. Understanding aggregates residual-level embeddings, while generation uses Next-Codec Prediction for coarse-to-fine hierarchical prediction.

  • Audio signals are converted into hierarchical discrete codec tokens that capture high-level semantics and fine-grained acoustic details.The codec tokenizer uses Residual Vector Quantization to represent different levels of granularity.
  • For understanding, embeddings from multiple residual levels are additively combined into a single audio token representation.Each residual-level code uses a level-specific embedding before the embeddings are summed.
  • A depth-wise autoregressive architecture predicts across codec dimensions instead of flattening multi-codebook tokens into one long sequence.This structured prediction design avoids the prohibitive sequence length caused by flattening multi-codebook audio tokens.
  • For generation, Next-Codec Prediction produces hierarchical audio tokens across transformer layers in a coarse-to-fine sequence.After each prediction, the code embedding is fed back to condition subsequent residual-level predictions.

3 Pre-Training

ERNIE 5.0 uses a large, filtered multimodal corpus and a staged pre-training recipe to learn unified representations across modalities. Its elastic training strategy jointly optimizes models with different depth, expert width, and routing sparsity, reducing the need for post-hoc compression and supporting deployment under varied resource constraints.

  • Pre-Training Data: ERNIE 5.0 is trained simultaneously on text, images, videos, and audio from the beginning under a unified multimodal paradigm.The corpus includes paired and interleaved multimodal sequences with metadata and captions.
  • Pre-Training Data: Trillions of text tokens and multimodal instances are filtered, deduplicated, and decontaminated to balance scale with high-fidelity semantic content.Quality controls remove low-quality and unsafe material while protecting benchmark integrity.
  • Elastic Training: Elastic depth, width, and sparsity respectively vary active layers, available experts, and top-k experts per token.The framework samples these configurations during training to support different compute, memory, and latency constraints.
  • Training Recipe: Pre-training progressively extends context length from 8K to 32K and 128K tokens while adjusting learning-rate schedules for stable optimization.The initial stage uses WSD, while mid-training switches to cosine annealing from 1 × 10−4 to 1 × 10−5.
  • Elastic Training: Elastic training jointly optimizes a full model and sampled sub-networks with varying depth, expert width, and routing sparsity in one pre-training run.This replaces a static architecture with an elastic super-network trained under the same autoregressive objective.
  • Elastic Training: Selecting parameter subsets along layer number and expert dimensions produces deployable models with different configurations on demand.Representation-dimension elasticity is described as orthogonal and left for future extension.

4 Post-Training

ERNIE 5.0’s post-training pipeline combines supervised fine-tuning with unified multimodal reinforcement learning, addressing efficiency, stability, and sparse-reward challenges in ultra-sparse MoE training. It introduces complementary techniques for rollout efficiency, entropy-collapse mitigation, and hard-query learning.

  • Post-training pipeline: SFT provides instruction-following and long-chain-of-thought capabilities before unified multimodal reinforcement learning.The UM-RL stage merges training across heterogeneous multimodal inputs.
  • Enhancing Rollout Efficiency with Unbiased Replay Buffer: Rollout generation exceeds 90% of RL training time, while long-tail response lengths leave GPUs idle and underutilized.U-RB addresses this bottleneck by enforcing data ordering while preparing future batches during long-tail generation.
  • Enhancing Rollout Efficiency with Unbiased Replay Buffer: U-RB accelerates rollout generation by separating high-throughput inference and training pools while preserving the order of iteration-assigned data groups.This design permits trajectories from earlier inference runs to resume into the appropriate training group.
  • Stabilizing Training with Mitigated Entropy Collapse: Training–inference mismatch and early overfitting to easy queries contribute to entropy collapse, especially under ultra-sparse MoE routing.The reported mechanisms include numerical inconsistency between separate engines and pruning of low-entropy responses by sequence-level truncated importance sampling.
  • Stabilizing Training with Mitigated Entropy Collapse: Multi-granularity importance-sampling clipping and well-learned positive-sample masking stabilize RL by calibrating updates and shifting gradient budget toward harder samples.WPSM masks redundant positive signals, while modality-sensitive trust-region modulation supports more balanced exploration and exploitation.
  • Boosting Sample Efficiency with Hint-based Learning: AHRL injects partial think sketches into hard queries, improving sample efficiency when sparse rewards provide insufficient gradient signals.The revealed hint fraction is gradually reduced during training, transitioning the model toward full self-exploration and helping prevent training stalls.

5 Infrastructures

ERNIE 5.0’s infrastructure addresses the memory, communication, tokenizer, attention, and reinforcement-learning challenges of unified trillion-parameter multimodal training. The resulting systems support memory-constrained pre-training, heterogeneous attention, and stable high-throughput RL.

  • 5.1 Large-Scale MoE Training: Ultra-sparse MoE training creates heavy inter-node communication and memory pressure, addressed with hybrid parallelism and fine-grained memory control.The configuration combines tensor, pipeline, expert, data, and context parallelism; FP8 activations, adaptive offloading, sub-batches, and defragmentation further support memory sufficiency.
  • 5.2 Tokenizer–Backbone Disaggregation: Tokenizer–backbone disaggregation separates modality-specific tokenizers onto dedicated, horizontally scalable compute nodes to reduce load imbalance.The separation lets tokenizers and the backbone use parallelization strategies suited to their different workloads.
  • 5.1 Large-Scale MoE Training: Memory-control techniques ensure the feasibility and reliability of ERNIE 5.0 pre-training in memory-constrained scenarios.These techniques include FP8 mixed precision, adaptive activation offloading, sub-batch computation, and automatic memory defragmentation.
  • 5.3 FlashMask for Flexible Multimodal Attention: FlashMask supports heterogeneous multimodal attention masks and delivers up to 200% operator-level speedup over FlexAttention.It also provides more than 20% end-to-end training acceleration and an 80% improvement over the Megatron-LM solution when integrated with context parallelism.
  • 5.4 Scalable and Disaggregated RL Infrastructure: The disaggregated RL infrastructure coordinates training, inference, environment interaction, and reward evaluation asynchronously for high-throughput, deterministic post-training.A unified FP8 execution engine reduces training–rollout numerical mismatch, while an unbiased replay buffer mitigates sequence-length bias from asynchronous completion.

6 Evaluations

ERNIE 5.0 is evaluated across broad text, vision, audio, and multimodal benchmarks, with analyses of modality-agnostic routing and elastic training. It shows strong, balanced performance across many text-centric capabilities, while a moderate gap remains on the most challenging reasoning benchmarks.

  • Evaluation Scope: Evaluations span factual knowledge, reasoning, mathematics, coding, multilingual understanding, instruction following, agent tasks, and vision and video capabilities.The benchmark suite covers both pre-trained and post-trained models across language and visual tasks.
  • Pre-trained Models: ERNIE 5.0-Base achieves consistently strong and well-balanced pre-training performance across knowledge, reasoning, mathematics, coding, and multilingual benchmarks.Reported strengths include advantages on knowledge-intensive tasks, best results across several general reasoning evaluations, state-of-the-art coding performance on LiveCodeBench v6 and CRUXEval, and gains on MMMLU and INCLUDE.
  • Pre-trained Models: The unified architecture and elastic pre-training strategy jointly yield strong generalization across diverse text-centric tasks.The paper presents this performance as a foundation for subsequent post-training and deployment.
  • Post-trained Models: Post-training improves factual recall, answer calibration, instruction following, and agent capabilities while preserving strong general reasoning and knowledge performance.ERNIE 5.0 achieves best performance on MultiChallenge and Multi-IF, near-top IFEval results, and competitive or leading results on ACEBench and BrowseComp-zh.
  • Post-trained Models: ERNIE 5.0 remains competitive across reasoning, mathematics, and coding, but retains a moderate gap on the most challenging reasoning benchmarks versus models such as Gemini 3-Pro.The reported profile emphasizes robust and balanced capability development rather than aggressive optimization for extreme reasoning or competition-style tasks.

ERNIE 5.0-Base

ERNIE 5.0-Base delivers balanced multimodal performance across understanding and generation tasks, while its unified routing and elastic training support specialization and deployment flexibility.

  • Multimodal Understanding: ERNIE 5.0-Base performs strongly across visual reasoning, document understanding, and general VQA without instruction tuning.Instruction tuning further improves explicit reasoning, compositional understanding, and visual-language alignment.
  • Multimodal Understanding: ERNIE 5.0 achieves competitive or superior results across reasoning, document understanding, and general VQA benchmarks while maintaining balanced performance.Instruction tuning primarily refines reasoning and alignment rather than compensating for perceptual capacity.
  • Image and Video Generation: ERNIE 5.0 matches leading systems on image generation, producing high-aesthetic images with semantic alignment and fine-grained visual details.On GenEval, it performs on par with Nano-Banana Pro and Qwen-Image and comparably to GPT-Image and Seedream 4.0.
  • Image and Video Generation: ERNIE 5.0 achieves the best VBench-Semantic performance and remains competitive on overall video quality, visual fidelity, and temporal consistency.It performs on par with HunyuanVideo-1 and Wan2.1-14B-0725, including before post-training.
  • Audio Understanding and Generation: ERNIE 5.0 provides stable audio understanding across speech, environmental sound, acoustic scene, and knowledge-grounded interaction benchmarks.It maintains robust ASR across languages and acoustic conditions and captures audio semantics beyond speech content.
  • Audio Understanding and Generation: ERNIE 5.0 preserves speech-generation content competitively without task-specific TTS optimization, although specialist TTS systems achieve stronger results.On SEED-TTS, it performs comparably to Qwen3-Omni on Chinese and English test sets.
  • Expert Routing: Modality-agnostic routing develops both shared and modality-specific expert activation patterns across layers and tasks.Deeper layers show increasing overlap among modalities, consistent with a shift toward more unified semantic representations.
  • Expert Routing: Normalized entropy reveals modality-dependent load balancing: text routing remains stable, while visual understanding is less balanced at boundary layers.Middle layers maintain relatively uniform routing for visual understanding, whereas generation and audio tasks follow different trends.

7 Conclusion

The conclusion presents ERNIE 5.0 as a unified trillion-level autoregressive model combining multimodal understanding, generation, scalable routing, elastic deployment, and stable post-training.

  • Conclusion: ERNIE 5.0 integrates text, image, video, and audio understanding and generation under a shared next-group-of-tokens objective.The report identifies it as the first realization, to the authors’ knowledge, of this combination in a unified trillion-level autoregressive framework.
  • Conclusion: Ultra-sparse MoE routing, elastic training, and multimodal reinforcement learning provide scalable modeling, flexible deployment configurations, and stable post-training.The conclusion reports competitive and balanced performance across modalities and characterizes unified elastic pre-training as a pathway toward scalable foundation models.

8 Contributors

The report includes a contributor list spanning the authors credited across the conclusion and contributor section.

  • Contributors: The contributor section lists the report’s authors, continuing across two passages.The passages contain the credited names rather than a separate contribution-role breakdown.
Loading 2602.04705v1…