Source-linked AI summary

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

DeepSeek-AI, Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenhao Xu, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Erhang Li, Fangqi Zhou, Fangyun Lin, Fucong Dai, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, Hao Li, Haofen Liang, Haoran Wei, Haowei Zhang, Haowen Luo, Haozhe Ji, Honghui Ding, Hongxuan Tang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, Jialiang Huang, Jiashi Li, Jiaxin Xu, Jiewen Hu, Jingchang Chen, Jingting Xiang, Jingyang Yuan, Jingyuan Cheng, Jinhua Zhu, Jun Ran, Junguang Jiang, Junjie Qiu, Junlong Li, Junxiao Song, Kai Dong, Kaige Gao, Kang Guan, Kexin Huang, Kexing Zhou, Kezhao Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Wang, Liang Zhao, Liangsheng Yin, Lihua Guo, Lingxiao Luo, Linwang Ma, Litong Wang, Liyue Zhang, M. S. Di, M. Y Xu, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingxu Zhou, Panpan Huang, Peixin Cong, Peiyi Wang, Qiancheng Wang, Qihao Zhu, Qingyang Li, Qinyu Chen, Qiushi Du, Ruiling Xu, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runqiu Yin, Runxin Xu, Ruomeng Shen, Ruoyu Zhang, S. H. Liu, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shaofei Cai, Shaoyuan Chen, Shengding Hu, Shengyu Liu, Shiqiang Hu, Shirong Ma, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, Songyang Zhou, Tao Ni, Tao Yun, Tian Pei, Tian Ye, Tianyuan Yue, Wangding Zeng, Wen Liu, Wenfeng Liang, Wenjie Pang, Wenjing Luo, Wenjun Gao, Wentao Zhang, Xi Gao, Xiangwen Wang, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaokang Zhang, Xiaotao Nie, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xingkai Yu, Xingyou Li, Xinyu Yang, Xinyuan Li, Xu Chen, Xuecheng Su, Xuehai Pan, Xuheng Lin, Xuwei Fu, Y. Q. Wang, Yang Zhang, Yanhong Xu, Yanru Ma, Yao Li, Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Qian, Yi Yu, Yichao Zhang, Yifan Ding, Yifan Shi, Yiliang Xiong, Ying He, Ying Zhou, Yinmin Zhong, Yishi Piao, Yisong Wang, Yixiao Chen, Yixuan Tan, Yixuan Wei, Yiyang Ma, Yiyuan Liu, Yonglun Yang, Yongqiang Guo, Yongtong Wu, Yu Wu, Yuan Cheng, Yuan Ou, Yuanfan Xu, Yuduan Wang, Yue Gong, Yuhan Wu, Yuheng Zou, Yukun Li, Yunfan Xiong, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Z. F. Wu, Z. Z. Ren, Zehua Zhao, Zehui Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhibin Gou, Zhicheng Ma, Zhigang Yan, Zhihong Shao, Zhixian Huang, Zhiyu Wu, Zhuoshu Li, Zhuping Zhang, Zian Xu, Zihao Wang, Zihui Gu, Zijia Zhu, Zilin Li, Zipeng Zhang, Ziwei Xie, Ziyi Gao, Zizheng Pan, Zongqing Yao, Bei Feng, Hui Li, J. L. Cai, Jiaqi Ni, Lei Xu, Meng Li, Ning Tian, R. J. Chen, R. L. Jin, S. S. Li, Shuang Zhou, Tianyu Sun, X. Q. Li, Xiangyue Jin, Xiaojin Shen, Xiaosha Chen, Xinnan Song, Xinyi Zhou, Y. X. Zhu, Yanping Huang, Yaohui Li, Yi Zheng, Yuchen Zhu, Yunxian Ma, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, Dongjie Ji, Jian Liang, Jianzhong Guo, Jin Chen, Leyi Xia, Miaojun Wang, Mingming Li, Peng Zhang, Ruyi Chen, Shangmian Sun, Shaoqing Wu, Shengfeng Ye, T. Wang, W. L. Xiao, Wei An, Xianzu Wang, Xiaowen Sun, Xiaoxiang Wang, Ying Tang, Yukun Zha, Zekai Zhang, Zhe Ju, Zhen Zhang, Zihua Qu

arXiv:2512.02556v1cs.CL

TL;DR

Open-source models need better efficiency, reasoning, and agent performance to address their gap with proprietary systems. DeepSeek-V3.2 combines sparse attention, scalable RL post-training, and large-scale agentic task synthesis; it achieves comparable reasoning performance with GPT-5, stronger agentic results, and gold-medal-level outcomes for its high-compute variant. The model still trails frontier systems in world knowledge, token efficiency, and some complex tasks.

  • Problem

    Open-source models are limited by inefficient long-context attention, insufficient post-training computation, and weaker support for complex tool-use tasks.

  • Method

    DeepSeek-V3.2 combines DSA, a scalable RL protocol, and large-scale agentic task synthesis, while extending the DeepSeek-V3.1-Terminus architecture through continued training.

  • Results

    DeepSeek-V3.2 achieves comparable reasoning performance with GPT-5, advances open-model agentic capability, and its Speciale variant reaches gold-medal performance in the 2025 IMO and IOI.

  • Takeaways & Limitations

    The framework bridges computational efficiency with advanced reasoning and supports more robust, generalizable open-model agents.

  • Takeaways & Limitations

    DeepSeek-V3.2 still lags frontier models in world knowledge, token efficiency, and complex-task performance.

Abstract

from arXiv · show

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 are as follows: (1) DeepSeek Sparse Attention (DSA): We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity while preserving model performance in long-context scenarios. (2) Scalable Reinforcement Learning Framework: By implementing a robust reinforcement learning protocol and scaling post-training compute, DeepSeek-V3.2 performs comparably to GPT-5. Notably, our high-compute variant, DeepSeek-V3.2-Speciale, surpasses GPT-5 and exhibits reasoning proficiency on par with Gemini-3.0-Pro, achieving gold-medal performance in both the 2025 International Mathematical Olympiad (IMO) and the International Olympiad in Informatics (IOI). (3) Large-Scale Agentic Task Synthesis Pipeline: To integrate reasoning into tool-use scenarios, we developed a novel synthesis pipeline that systematically generates training data at scale. This methodology facilitates scalable agentic post-training, yielding substantial improvements in generalization and instruction-following robustness within complex, interactive environments.

1. Introduction

DeepSeek-V3.2 targets widening open–closed model gaps through efficient attention, increased post-training computation, and large-scale agentic task synthesis. It reports comparable reasoning performance with frontier models and stronger open-model agent capabilities.

  • Open-source models face an apparent widening performance gap relative to proprietary systems.
  • Vanilla attention constrains long-sequence efficiency, while insufficient post-training computation limits performance on hard tasks.
  • DSA reduces computational complexity while preserving performance in long-context scenarios.
  • The scalable RL framework allocates more than 10% of pre-training cost to post-training computation.
  • The agentic synthesis pipeline generates over 1,800 environments and 85,000 complex prompts to improve generalization and instruction following in tool-use contexts.
  • DeepSeek-V3.2 achieves similar performance to GPT-5 across multiple reasoning benchmarks and advances open-model agentic capabilities at substantially lower costs.

2. DeepSeek-V3.2 Architecture

DeepSeek-V3.2 introduces DSA into the DeepSeek-V3.1-Terminus architecture to select sparse key-value entries, reducing attention computation for long contexts. Continued pre-training and sparse adaptation preserve benchmark performance while improving long-context efficiency.

  • DeepSeek Sparse Attention: DSA is the only architectural modification relative to DeepSeek-V3.1-Terminus, introduced through continued training.
  • DeepSeek Sparse Attention: The DSA prototype combines a lightning indexer with fine-grained token selection.
  • DeepSeek Sparse Attention: The indexer scores preceding tokens for each query, and selection retrieves only the top-k key-value entries.
  • DeepSeek Sparse Attention: DSA is instantiated under MLA using MQA so each latent key-value vector is shared across all query heads.
  • Continued Pre-Training: The model starts from a 128K-context DeepSeek-V3.1-Terminus checkpoint and undergoes continued pre-training followed by post-training.
  • Continued Pre-Training: Training uses a dense warm-up to initialize the indexer, followed by sparse training that adapts all model parameters to DSA.
  • Evaluation: DeepSeek-V3.2-Exp shows similar standard-benchmark performance while significantly improving long-sequence computational efficiency.
  • Inference Costs: DSA reduces core attention complexity from O(L^2) to O(Lk), where k is much smaller than sequence length L.

3. Post-Training

DeepSeek-V3.2 combines specialist distillation with mixed reinforcement learning, while adding mechanisms intended to stabilize RL scaling and support reasoning in agentic tasks. Its post-training also includes large-scale synthesized agent environments and persistent reasoning context for tool calls.

  • Post-Training Pipeline: DeepSeek-V3.2 retains the post-training pipeline of DeepSeek-V3.2-Exp, combining specialist distillation with mixed RL training.The pipeline uses sparse attention during post-training and covers specialist models across multiple domains.
  • Specialist Distillation: Specialists cover mathematics, programming, logical reasoning, general agents, agentic coding, and agentic search in thinking and non-thinking modes.Specialist models are fine-tuned from the same pretrained DeepSeek-V3.2 checkpoint and trained with large-scale RL computation.
  • Mixed RL Training: GRPO merges reasoning, agent, and human-alignment training into one RL stage while using outcome, length, and language-consistency rewards.The unified stage is described as balancing performance across domains and avoiding catastrophic forgetting associated with multi-stage training.
  • Scaling GRPO: Group-relative advantages are computed by subtracting each group’s mean outcome reward from the corresponding output reward.Reward models score each response in a sampled group before advantage normalization.
  • RL Stabilization: Unbiased KL estimation and off-policy sequence masking improve RL stability by correcting KL gradients and masking highly divergent negative-advantage sequences.The masking threshold is controlled by a divergence hyperparameter, while empirical observations associate sequence masking with improved stability in otherwise unstable scenarios.
  • Agentic Post-Training: An automatic synthesis agent produces 1,827 task-oriented environments, while retained reasoning content supports subsequent reinforcement-learning stages for tool-use trajectories.The resulting dataset contains 1,827 environments and 4,417 tasks; retaining reasoning addresses token inefficiency from repeatedly re-reasoning after tool calls.

4. Evaluation

DeepSeek-V3.2 is evaluated across reasoning, coding, search, tool-use, and synthesized-agent tasks, with results indicating strong performance and several important efficiency and scope trade-offs.

  • Main Results: DeepSeek-V3.2 achieves similar reasoning performance to GPT-5-high, remains slightly below Gemini-3.0-Pro, and uses substantially fewer output tokens than K2-Thinking.The reported performance gains are associated with increased RL-training computation exceeding 10% of pre-training cost.
  • Main Results: DeepSeek-V3.2 significantly outperforms open-source models on SWE-bench Verified and Terminal Bench 2.0 in code-agent evaluations.Terminal Bench 2.0 reports 46.4 with Claude Code and 39.3 with Terminus in non-thinking mode; SWE-bench Verified robustness tests range from 72 to 74.
  • Main Results: Reported evaluations remain bounded by deployment and benchmarking constraints, including 128K context limits and stricter token constraints for official DeepSeek-V3.2.DeepSeek-V3.2-Speciale has significantly inferior token efficiency to Gemini-3.0-Pro.
  • Main Results: DeepSeek-V3.2 narrows the open-versus-closed tool-use gap but remains below frontier models.Tau2-bench category scores are 63.8 for Airline, 81.1 for Retail, and 96.2 for Telecom; redundant self-verification can exceed the 128K context limit.
  • Results of DeepSeek-V3.2-Speciale: DeepSeek-V3.2-Speciale surpasses Gemini-3.0-Pro across multiple benchmarks by leveraging increased reasoning tokens.It reaches gold-medal thresholds in IOI 2025, IMO 2025, ICPC WF 2025, and CMO 2025 under the reported evaluation settings.
  • Synthesis Agentic Tasks: Synthetic agentic tasks are challenging for both the synthesis model and frontier closed-source models, with accuracies of 12% and at most 62%, respectively.The evaluation samples 50 instances from the general synthesized agentic tasks.
  • Synthesis Agentic Tasks: Large-scale RL on synthetic agentic tasks improves DeepSeek-V3.2-SFT on Tau2Bench, MCP-Mark, and MCP-Universe, whereas code-and-search-only RL does not.The experiment excludes long chain-of-thought and other RL data by using non-thinking mode.
  • Context Management of Search Agent: Context management scales BrowseComp test-time compute by extending reasoning trajectories, with Summary reaching 364 average steps and Discard-all scoring 67.6.Summary achieves performance improvement of up to 60.2 but has relatively low overall efficiency; Discard-all is comparable to parallel scaling with fewer steps.

5. Conclusion, Limitation, and Future Work

DeepSeek-V3.2 combines computational efficiency with advanced reasoning and agent capabilities, while remaining limited in world knowledge, token efficiency, and complex-task performance relative to frontier models.

  • Contributions: DeepSeek-V3.2 bridges computational efficiency and advanced reasoning through DSA, increased compute, and large-scale agentic task synthesis.The framework preserves long-context performance, reaches comparable reasoning performance with GPT-5, and improves tool-use proficiency.
  • Contributions: DeepSeek-V3.2-Speciale achieves gold-medal performance in the IMO and IOI.
  • Limitations and Future Work: Due to fewer total training FLOPs, DeepSeek-V3.2 has less world knowledge breadth than leading proprietary models.The authors plan to address this gap by scaling pre-training compute.
  • Limitations and Future Work: DeepSeek-V3.2 typically requires longer generation trajectories to match models such as Gemini-3.0-Pro.Future work targets improving the intelligence density of its reasoning chains.
  • Limitations and Future Work: Complex-task performance remains inferior to frontier models, motivating further refinement of the foundation model and post-training recipe.

A. MHA and MQA Modes of MLA

MLA supports both MHA and MQA modes, with a transformation between them illustrated in Figure 7.

  • For DeepSeek-V3.1-Terminus, MHA is used for training and prefilling, whereas MQA is used for decoding.
  • Figure 7 illustrates the MHA and MQA modes of MLA and the transformation between them.

B. Cold Start Template

The cold-start data system uses explicit formats for reasoning, tool descriptions, and tool calls, including tool execution within the thinking process.

  • The reasoning data system prompt requires the model to output its reasoning process within <think></think> tags.
  • {TOOL-DESCRIPTIONS} and {TOOLCALL-FORMAT} are replaced with task-specific tools and the designed tool-call format.
  • The model executes tool calls during the thinking process.

C. Non-thinking DeepSeek-V3.2 Agentic Evaluation

Non-thinking DeepSeek-V3.2 remains competitive with thinking mode in agentic evaluation, although its performance is slightly lower.

  • Terminal Bench scores in Table 9 use the Claude Code framework, while non-thinking mode scores 39.3 with the Terminus framework.
  • Non-thinking mode performs slightly worse than thinking mode but remains competitive.

D. Evaluation Method of IOI, ICPC World Final, IMO, and CMO

The evaluations use contest-constrained generation and task-specific solution-selection procedures across IOI, ICPC, IMO, and CMO.

  • All competitions cap maximum generation length at 128k, prohibit tools and internet access, and follow official time and attempt limits.
  • IOI evaluation samples 500 candidate solutions per problem before multi-stage filtering, including checks for sample-case validity and length constraints.
  • ICPC evaluation generates 32 candidate solutions per problem and applies the same filtering criteria used for IOI.
  • IMO and CMO evaluations iteratively generate, verify, and refine solutions until perfect self-evaluation or the maximum revision cap is reached.

E. Author List

The author list comprises research and engineering, data annotation, and business and compliance contributors, with names presented alphabetically by first name.

  • Research and engineering contributors are listed in the author materials.
  • Data annotation contributors are separately identified in the author materials.
  • Business and compliance contributors are separately identified in the author materials.
  • Authors are listed alphabetically by their first name.
Loading 2512.02556v1…