Source-linked AI summary
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-AI, Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chenchen Ling, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyu Hou, Chenhao Xu, Chenze Shao, Chong Ruan, Conner Sun, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Donghao Li, Dongjie Ji, Erhang Li, Fang Wei, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong Dai, Guangbo Hao, Guanting Chen, Guoai Cao, Guolai Meng, Guowei Li, Han Yu, Han Zhang, Hanwei Xu, Hao Li, Haofen Liang, Haoling Zhang, Haoming Luo, Haoran Wei, Haotian Yuan, Haowei Zhang, Haowen Luo, Haoyu Chen, Haozhe Ji, Hengqing Zhang, Honghui Ding, Hongxuan Tang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, J Yang, JQ Zhu, Jia Luo, Jia Song, Jia Yu, Jialiang Huang, Jialu Cai, Jian Liang, Jiangting Zhou, Jiasheng Ye, Jiashi Li, Jiaxin Xu, Jiewen Hu, Jieyu Yang, Jin Chen, Jin Yan, Jingchang Chen, Jingli Zhou, Jingting Xiang, Jingyang Yuan, Jingyuan Cheng, Jingzi Zhou, Jinhua Zhu, Jiping Yu, Joseph Sun, Jun Ran, Junguang Jiang, Junjie Qiu, Junlong Li, Junmin Zheng, Junxiao Song, Kai Dong, Kaige Gao, Kang Guan, Kexing Zhou, Kezhao Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Wang, Leyi Xia, Li Zhang, Liang Zhao, Lihua Guo, Lingxiao Luo, Linwang Ma, Linyan Zhu, Litong Wang, Liyu Cai, Liyue Zhang, Longhao Chen, MS Di, MY Xu, Max Mei, Miaojun Wang, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingming Li, Mingxu Zhou, Minmin Han, Ning Wang, Panpan Huang, Panpan Wang, Peixin Cong, Peiyi Wang, Peng Zhang, Qiancheng Wang, Qihao Zhu, Qingyang Li, Qinyu Chen, Qiushi Du, Qiwei Jiang, Rui Tian, Ruifan Xu, Ruijie Lu, Ruiling Xu, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runqian Chen, Runqiu Yin, Runxin Xu, Ruomeng Shen, Ruoyu Zhang, Ruyi Chen, SH Liu, Shanghao Lu, Shangmian Sun, Shangyan Zhou, Shanhuang Chen, Shaofei Cai, Shaoheng Nie, Shaoqing Wu, Shaoyuan Chen, Shengding Hu, Shengyu Liu, Shiqiang Hu, Shirong Ma, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, Shuying Yu, Songyang Zhou, Tao Ni, Tao Yun, Tian Jin, Tian Pei, Tian Ye, Tianle Lin, Tianran Ji, Tianyi Cui, Tianyuan Yue, Tingting Yu, Tun Wang, W Zhang, WL Xiao, Wangding Zeng, Wei An, Weilin Zhao, Wen Liu, Wenfeng Liang, Wenjie Pang, Wenjing Luo, Wenjing Yao, Wenjun Gao, Wenkai Yang, Wenlve Huang, Wenqing Hou, Wentao Zhang, Wenting Ma, Xi Gao, Xiang He, Xiangwen Wang, Xianzu Wang, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaokang Zhang, Xiaotao Nie, Xiaowen Sun, Xiaoxiang Wang, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xingchen Liu, Xingkai Yu, Xingyou Li, Xinyu Yang, Xinyu Zhang, Xu Chen, Xuanyu Wang, Xuecheng Su, Xueyin Chen, Xuheng Lin, Xuwei Fu, YC Yan, YQ Wang, YW Ma, Yanfeng Luo, Yang Zhang, Yanhong Xu, Yanru Ma, Yanwen Huang, Yao Li, Yao Li, Yao Xu, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Qian, Yi Shao, Yi Yu, Yichao Zhang, Yifan Ding, Yifan Shi, Yijia Wu, Yiliang Xiong, Yiling Ma, Ying He, Ying Tang, Ying Zhou, Yingjia Luo, Yinmin Zhong, Yishi Piao, Yisong Wang, Yixiang Zhang, Yixiao Chen, Yixuan Tan, Yixuan Wei, Yiyang Ma, Yiyuan Liu, Yonglun Yang, Yongqiang Guo, Yongtong Wu, Yu Wu, YuKun Li, Yuan Cheng, Yuan Ou, Yuanfan Xu, Yuanhao Li, Yuduan Wang, Yuehan Yang, Yuer Xu, Yuhan Wu, Yuhao Meng, Yuheng Zou, Yukun Zha, Yunfan Xiong, Yupeng Chen, Yuping Lin, Yuqian Cao, Yuqian Wang, Yushun Zhang, Yuting Yan, Yutong Lin, Yuxian Gu, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuxuan Zhou, Yuyang Zhou, Yuzhen Huang, ZF Wu, Zehao Wang, Zehua Zhao, Zehui Ren, Zekai Zhang, Zhangli Sha, Zhe Fu, Zhe Ju, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zheren Gao, Zhewen Hao, Zhibin Gou, Zhicheng Ma, Zhigang Yan, Zhihong Shao, Zhixian Huang, Zhixuan Chen, Zhiyu Wu, Zhizhou Ren, Zhongyu Wu, Zhuoshu Li, Zhuping Zhang, Zian Xu, Zihao Wang, Zihua Qu, Zihui Gu, Zijia Zhu, Zilin Li, Zipeng Zhang, Ziwei Xie, Ziyi Gao, Ziyi Wan, Zizheng Pan, Zongqing Yao
TL;DR
Ultra-long contexts are limited by vanilla attention’s quadratic computational cost, motivating efficient support for long-horizon reasoning and related tasks. DeepSeek-V4 combines hybrid CSA-HCA attention, mHC, Muon, and infrastructure optimizations; at one million tokens, DeepSeek-V4-Pro uses 27% of DeepSeek-V3.2’s inference FLOPs and 10% of its KV cache, while evaluations report strong open-model performance.
Problem
Vanilla attention’s quadratic computational complexity creates a bottleneck for ultra-long contexts and reasoning processes, while long-horizon scenarios require efficient ultra-long-context support.
Method
DeepSeek-V4 combines hybrid CSA-HCA attention, mHC residual connections, Muon optimization, and infrastructure optimizations for efficient training and inference.
Results
At one-million-token context, DeepSeek-V4-Pro requires 27% of DeepSeek-V3.2’s single-token inference FLOPs and 10% of its KV cache, while DeepSeek-V4-Pro-Max shows strong results across core evaluations.
Takeaways & Limitations
The series natively supports one-million-token contexts, establishing a foundation for future test-time scaling and research on long-horizon tasks.
Takeaways & Limitations
The architecture remains relatively complex, and the principles underlying some training-stability techniques remain insufficiently understood.
Abstract
from arXiv · showhide
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) Manifold-Constrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability. We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances their capabilities. DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, redefines the state-of-the-art for open models, outperforming its predecessors in core tasks. Meanwhile, DeepSeek-V4 series are highly efficient in long-context scenarios. In the one-million-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2. This enables us to routinely support one-million-token contexts, thereby making long-horizon tasks and further test-time scaling more feasible. The model checkpoints are available at https://huggingface.co/collections/deepseek-ai/deepseek-v4.
1. Introduction
DeepSeek-V4 targets the efficiency bottleneck of ultra-long-context processing with architectural and optimization changes, supporting million-token contexts. Its models combine strong benchmark performance with substantially lower long-context inference costs.
- The quadratic complexity of vanilla attention constrains test-time scaling and creates a bottleneck for ultra-long contexts and reasoning processes.
- DeepSeek-V4 introduces hybrid CSA-HCA attention, mHC residual connections, and the Muon optimizer to improve long-context efficiency, modeling capability, convergence, and training stability.
- 27% of single-token inference FLOPs and 10% of KV cache are required by DeepSeek-V4-Pro versus DeepSeek-V3.2 at one-million-token context.The comparison uses equivalent FP8 FLOPs and applies to the one-million-token setting.
- 10% of single-token inference FLOPs and 7% of KV cache are required by DeepSeek-V4-Flash versus DeepSeek-V3.2 at one-million-token context.
- 32T tokens train DeepSeek-V4-Flash and 33T tokens train DeepSeek-V4-Pro, after which both models natively support one-million-token contexts.
- DeepSeek-V4-Pro-Max significantly outperforms leading open-source models on SimpleQA and Chinese-SimpleQA, while showing a marginal lead on educational-knowledge evaluations.
2. Architecture
DeepSeek-V4 retains the Transformer, DeepSeekMoE, and MTP foundations while adding mHC, hybrid CSA-HCA attention, and Muon optimization.
- mHC strengthens conventional residual connections in the retained Transformer architecture.
- Hybrid CSA-HCA attention improves long-context efficiency through compressed sparse and heavily compressed attention.
- Muon is employed as the optimizer, while DeepSeekMoE remains the feed-forward architecture with only minor adjustments from DeepSeek-V3.
- The Multi-Token Prediction configuration remains identical to DeepSeek-V3.
2.1. Designs Inherited from DeepSeek-V3
DeepSeek-V4 inherits DeepSeekMoE and Multi-Token Prediction from earlier DeepSeek models, with a modified expert-affinity activation function.
- DeepSeek-V4 uses DeepSeekMoE for feed-forward networks, with fine-grained routed experts and shared experts.
- The expert-affinity activation changes from Sigmoid(·) to Sqrt(Softplus(·)).
- DeepSeek-V4 adopts the same Multi-Token Prediction strategy and objectives as DeepSeek-V3 without modification.
2.2. Manifold-Constrained Hyper-Connections
mHC extends residual connections with dynamically generated mappings constrained for stable signal propagation. Its design addresses numerical instability observed when stacking conventional Hyper-Connections.
- Standard Hyper-Connections: Standard Hyper-Connections expand residual-stream width by a factor of n_hc while keeping inner-layer input and output size d.
- Standard Hyper-Connections: HC provides a complementary residual-width scaling axis with minimal computational overhead, but stacking multiple layers frequently causes numerical instability.
- Manifold-Constrained Residual Mapping: mHC constrains the residual mapping B_l to the manifold of doubly stochastic matrices, the Birkhoff polytope.
- Manifold-Constrained Residual Mapping: The constraint bounds the mapping spectral norm by 1, making residual transformation non-expansive and improving forward-pass and backpropagation stability.
- Manifold-Constrained Residual Mapping: Input and output mappings A_l and C_l are constrained to be non-negative and bounded using a Sigmoid function.
- Dynamic Parameterization: The three linear mappings are dynamically generated from input-dependent and input-independent components after flattening and normalizing X_l.
- Manifold-Constrained Residual Mapping: The residual mapping is projected onto the doubly stochastic manifold by exponentiation followed by iterative column and row normalization.
2.3. Hybrid Attention with CSA and HCA
DeepSeek-V4 uses interleaved Compressed Sparse Attention and Heavily Compressed Attention to reduce attention cost for long contexts. CSA compresses and sparsely selects KV entries, while HCA applies heavier compression without sparse attention.
- Compressed Sparse Attention: CSA compresses every m tokens into one KV entry, then applies sparse attention so each query attends only to k compressed entries.It uses a lightning indexer to select the top-k compressed KV entries for core attention.
- Heavily Compressed Attention: HCA consolidates every m′ (≫m) tokens into one KV entry and does not employ sparse attention.Its heavier compression reduces the sequence length to 1.
- Shared KV and output projection: Both CSA and HCA use shared-KV multi-query attention and grouped output projection after KV compression.Grouped projection reduces the cost of projecting many core-attention head outputs back to the hidden state.
- Other details: RMSNorm is applied to query heads and compressed KV entries immediately before core attention to avoid exploding attention logits and potentially improve training stability.The normalization is used in both CSA and HCA.
- Efficiency discussion: Approximately 2% of the BF16 GQA8 baseline KV cache is needed by DeepSeek-V4 in the one-million-token setting.The baseline uses head dimension 128.
2.4. Muon Optimizer
DeepSeek-V4 uses Muon for most modules to improve convergence and training stability, while retaining AdamW for specified parameter groups. Its update computation uses hybrid Newton-Schulz iterations and omits QK-Clip because RMSNorm prevents exploding attention logits.
- Optimizer choice: Muon updates most DeepSeek-V4 modules because it provides faster convergence and improved training stability.The optimizer is used for the majority of modules.
- Basic configurations: AdamW remains responsible for embeddings, prediction heads, selected mHC parameters, static biases, gating factors, and RMSNorm weights.All other modules are updated with Muon.
- Hybrid Newton-Schulz iterations: Hybrid Newton-Schulz uses 10 iterations: eight rapid-convergence steps followed by two stabilization steps.The first stage uses coefficients (3.4445, −4.7750, 2.0315), and the final stage uses (2, −1.5, 0.5).
- Attention-logit stability: DeepSeek-V4 omits QK-Clip because RMSNorm on attention queries and KV entries prevents attention logits from exploding.This architectural normalization is presented as the reason QK-Clip is unnecessary in Muon.
3. General Infrastructures
DeepSeek-V4 improves infrastructure efficiency by overlapping communication and computation in MoE execution, while providing fused, reproducible, and optimized kernels. These designs target bandwidth, orchestration, development, and determinism bottlenecks across training and inference.
- Fine-Grained Communication-Computation Overlap in Expert Parallelism: Fine-grained Expert Parallelism fuses communication and computation into a pipelined kernel to reduce inter-node bandwidth and latency bottlenecks.Experts are split into waves so communication, computation, and result transfer can proceed concurrently.
- Fine-Grained Communication-Computation Overlap in Expert Parallelism: 1.50–1.73× speedup is achieved for general inference workloads, rising to 1.96× for latency-sensitive RL rollouts and high-speed agent serving.The scheme was validated on NVIDIA GPUs and HUAWEI Ascend NPUs against non-fused baselines.
- Hardware and Communication-Compute Balance: 6.1 TFLOP/s of compute can have its communication hidden by each GBps of interconnect bandwidth once the required balance point is reached.The analysis argues that additional bandwidth then yields diminishing returns because computation remains the bottleneck.
- Kernel Development Infrastructure: TileLang replaces most fine-grained Torch ATen operators with fused kernels while balancing development productivity and runtime efficiency.It supports rapid prototyping of attention variants and optimized deployment kernels.
- Reproducibility and Kernel Libraries: Batch-invariant deterministic kernels preserve bitwise alignment across training and inference with minimal performance overhead.The attention implementation uses dual kernels with matched accumulation order, and mHC wall-time overhead is 6.7% of the overlapped 1F1B pipeline stage.
4. Pre-Training
DeepSeek-V4 uses longer, more diverse training data, hybrid attention configurations, and Muon-based optimization while addressing trillion-parameter training instability. Base-model evaluations report broad gains, with Flash improving over V3.2 at lower parameter cost and Pro providing the strongest foundation-model results.
- Model Configurations: DeepSeek-V4 interleaves CSA and HCA after initial layers, using distinct compression rates and sparse-attention configurations for Flash and Pro.Flash uses m=4, m′=128, and attention top-k 512; Pro uses m=4, m′=128, and attention top-k 1024.
- Optimization and Training Setup: Muon optimizes most parameters, while AdamW is retained for embeddings, prediction heads, and RMSNorm weights in both models.Pro training progressively extends sequence length from 4K to 16K, 64K, and 1M.
- Training Stability: Anticipatory routing and SwiGLU clamping address recurring loss spikes associated with MoE-layer outliers and routing instability during trillion-parameter training.The reported methods decouple routing updates from backbone updates and constrain anomalous numerical values.
- Evaluation Results: DeepSeek-V4-Flash-Base surpasses DeepSeek-V3.2-Base across most evaluations despite using substantially fewer activated and total parameters.The advantage is especially evident in world knowledge and challenging long-context tasks.
- Evaluation Results: DeepSeek-V4-Pro-Base establishes near-universal dominance over DeepSeek-V3.2-Base and DeepSeek-V4-Flash-Base across knowledge, reasoning, coding, and long-context capabilities.The evaluation framework covers world knowledge, language understanding and reasoning, coding and mathematics, and long-context processing.
5. Post-Training
DeepSeek-V4’s post-training develops domain-specific specialists and consolidates them with on-policy distillation, while supporting multiple reasoning-effort modes and tool-oriented agent workflows. Evaluations report strong performance across knowledge, reasoning, coding, agent, and million-token long-context tasks.
- Post-training pipeline: The post-training pipeline replaced mixed reinforcement learning with on-policy distillation after specialist development.Specialists are trained with fine-tuning and domain-specific reinforcement learning before consolidation.
- Reasoning effort: Three reasoning modes use distinct length penalties and context windows during reinforcement learning, producing different reasoning output lengths.DeepSeek-V4-Pro and DeepSeek-V4-Flash each support three reasoning-effort modes.
- Reward modeling and tools: The post-training system adds rubric-guided generative reward modeling for hard-to-verify tasks and an XML-based tool-call schema with a special token.The actor network is optimized as both a generator and evaluator, while XML tool calls reduce escaping failures and tool-call errors.
- Infrastructure: DeepSeek-V4 introduces targeted optimizations for reinforcement learning and on-policy distillation on million-token sequences, supported by the DSec sandbox platform.Dsec is designed for large-scale agentic post-training and evaluation workloads.
- Evaluation results: Higher reasoning effort improves knowledge results, and maximum effort outperforms high effort on the most challenging tasks.The comparison is reported for both knowledge benchmarks and challenging reasoning tasks.
- Evaluation results: DeepSeek-V4-Pro-Max outperforms prior open models across reasoning benchmarks, while DeepSeek-V4-Pro shows strong coding, agent, knowledge, and long-context results.Reported results include comparable performance to GPT-5.4 on coding competitions, strong agent evaluations, and better MRCR performance than Gemini-3.1-Pro while remaining behind Claude Opus 4.6.
5.4. Performance on Real-World Tasks
DeepSeek-V4 is evaluated on practical search, writing, coding, and enterprise productivity tasks using benchmark comparisons and human assessments. Results show advantages over prior or competing models across several real-world settings, with some remaining limitations against frontier systems.
- Chinese Writing: DeepSeek-V4-Pro achieves a 62.7% overall win rate versus 34.1% for Gemini-3.1-Pro on Chinese functional writing tasks.The evaluation covers common daily writing queries with concise prompts.
- Chinese Writing: DeepSeek-V4-Pro records 60.0% and 77.5% win rates against Gemini-3.1-Pro for instruction following and writing quality in Chinese creative writing.The two evaluation axes are instruction following and writing quality.
- Search: DeepSeek-V4-Pro outperforms DeepSeek-V3.2 across objective and subjective search question-answering categories.The largest gains occur in single-value search and planning-and-strategy tasks.
- Search: Agentic search consistently outperforms retrieval-augmented search, particularly on complex tasks, while remaining cost-efficient.Agentic search iteratively invokes search and fetch tools within a predefined thinking budget.
- Enterprise Productivity: DeepSeek-V4-Pro-Max achieves a 63% non-loss rate against Opus-4.6-Max on diverse Chinese white-collar tasks.Human evaluations cover task completion, instruction following, content quality, and formatting aesthetics across analysis, generation, and editing.
- Coding: DeepSeek-V4-Pro significantly outperforms Claude Sonnet 4.5 and approaches Claude Opus 4.5 on a 30-task internal R&D coding benchmark.The tasks span feature development, bug fixing, refactoring, and diagnostics across PyTorch, CUDA, Rust, and C++.
6. Conclusion, Limitations, and Future Directions
The paper concludes that DeepSeek-V4 combines architectural and infrastructure changes to support efficient million-token contexts and strong open-model performance. It identifies architectural complexity and incomplete understanding of training-stability techniques as limitations, while outlining further efficiency and agentic-task work.
- Conclusion: DeepSeek-V4 combines hybrid CSA-HCA attention and infrastructure optimization to support million-token contexts and efficient long-horizon processing.The conclusion also connects this foundation to future test-time scaling, online learning, and long-horizon tasks.
- Conclusion: DeepSeek-V4-Pro-Max substantially outperforms prior open-source models on knowledge benchmarks and delivers reasoning close to frontier proprietary models.The paper also reports competitive agent capabilities, while DeepSeek-V4-Flash-Max maintains comparable reasoning performance with a cost-efficient architecture.
- Limitations: The architecture is relatively complex because many preliminarily validated components and techniques were retained to reduce development risk.Future work aims to distill the design into more essential components without sacrificing performance.
- Limitations: The principles underlying Anticipatory Routing and SwiGLU Clamping remain insufficiently understood despite their effectiveness in mitigating training instabilities.The authors plan more foundational work on training stability and predictive internal metric monitoring.
- Future Directions: Future directions include new dimensions of model sparsity, low-latency deployment techniques, and continued investigation of long-horizon multi-round agentic tasks.The stated goal is improved computational and memory efficiency alongside more responsive long-context interaction.
A.1. Author List
The author list is organized by business function, with research and engineering contributors listed separately from business and compliance contributors. Names marked with an asterisk denote former team members.
- Author List: Authors are listed alphabetically by first name, and an asterisk marks individuals who have departed from the team.This convention applies to the listed contributors.
- Research & Engineering: The research and engineering contributors include a large group spanning the listed names from Anyi Xu through Hao Li and beyond.The supplied list continues across multiple passages.
- Business & Compliance: Business and compliance contributors are listed separately from the research and engineering group.Their names continue across the supplied business and compliance passages.
B. Evaluation Details
The evaluation details identify comparisons for agentic search, retrieval-augmented search, Chinese writing, and complex instruction-following tasks. The supplied captions specify the compared systems and evaluation domains, while the accompanying passages describe the reported outcomes.
- Search: Table 9 compares agentic search with retrieval-augmented search for DeepSeek-V4-Pro.The agentic-search evaluation concerns the model's search capability under the described tool-use setup.
- Search: Table 10 reports the mean cost comparison between agentic search and retrieval-augmented search, with most agentic-search tool calls executed in parallel.The table is specifically identified as a cost comparison for DeepSeek-V4-Pro.
- Chinese Writing: Table 13 analyzes DeepSeek-V4-Pro against Gemini-3.1-Pro in Chinese creative writing.The evaluation uses instruction following and writing quality as its two axes.
- Writing: Table 14 compares DeepSeek-V4-Pro with Claude-Opus-4.5 on complex instruction following and multi-turn writing.The supplied passages identify the comparison scope but do not provide the table's numerical results.