Source-linked AI summary

Kimi K3: Open Frontier Intelligence

Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, M. C., Jianfeng Cai, Xinyuan Cai, Peizhou Cao, Yuxuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Guanduo Chen, Guangyu Chen, Guanzheng Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kexin Chen, Peng Chen, Ruijue Chen, Wentao Chen, Xin Chen, Yang Chen, Yanru Chen, Yifei Chen, Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen, Zhirong Chen, Dazhi Cheng, Yean Cheng, Jialei Cui, Jingbing Cui, Anqi Dai, Jiaqi Deng, Hao Ding, Rui Ding, Shaofeng Ding, Mengfan Dong, Mengnan Dong, Yuhao Dong, Yuxin Dong, Angang Du, Chenzhuang Du, Dikang Du, Jusen Du, Yulun Du, Yu Fan, Jing Feng, Qiulin Feng, Yichen Feng, Kelin Fu, Qiang Fu, Fuxuan Gao, Hongcheng Gao, Jingyue Gao, Tong Gao, Weijia Gao, Shangyi Geng, Jie Gong, Linhu Gong, Shengao Gong, Xiaochen Gong, Qizheng Gu, Yicheng Gu, Shuhao Guan, Haiqing Guo, Shiqi Guo, Xiang Guo, Zhengyan Guo, Beixi Hao, Wenxin Hao, Xiaoru Hao, Dailan He, Haotian He, Lehan He, Qi He, Weiran He, Xinran He, Xinyi He, Yibo He, Yunjia He, Chao Hong, Tiange Hong, Hao Hu, Jiaxi Hu, Ruikun Hu, Weiming Hu, Yangyang Hu, Zhenxing Hu, Liang Hua, Jinbin Huang, Ke Huang, Ruiyuan Huang, Siying Huang, Weixiao Huang, Yan Huang, Zhengjie Huang, Zhiqi Huang, Yulong Hui, Chaobo Jia, Yutong Jiang, Zhejun Jiang, Zuoyou Jiang, Wenyi Jin, Xinyi Jin, Yu Jing, Huanjun Kong, Guokun Lai, Aidi Li, Cheng Li, Chengyuan Li, Cong Li, Fang Li, Guanyu Li, Haoyang Li, Jia Li, Junxiong Li, Lei Li, Letian Li, Lincan Li, Weihong Li, Wentao Li, Xintong Li, Yang Li, Yishen Li, Yiwei Li, Yuxiao Li, Zhaowei Li, Zhaoxi Li, Zheming Li, Zhengxiao Li, Zhiyuan Li, Jiawei Lin, Xiaohan Lin, Yibo Lin, Zichao Lin, Ziyan Lin, Bill Liu, Boxiao Liu, Chuan Liu, Liang Liu, Shaowei Liu, Shudong Liu, Shuran Liu, Tianwei Liu, Weizhou Liu, Yangyang Liu, Yanming Liu, Yibo Liu, Yipeng Liu, Zhengying Liu, Zhiheng Liu, Enzhe Lu, Haoyu Lu, Linqiang Lu, Tingzhan Lu, Zhiyuan Lu, Aotian Luo, G. Luo, Junyu Luo, Yifan Luo, B. Lyu, Wenzhou Lyu, Shaoguang Mao, Yuan Mei, Xin Men, Minqing Ni, Yixuan Niu, Siyuan Pan, Shujun Peng, Zhangyang Qi, Ruoyu Qin, ZeChao Qin, Zeyu Qin, Haiquan Qiu, Jianxin Qiu, Jiezhong Qiu, Bowen Qu, Yuhao Qu, Zeyu Shang, Youbo Shao, Han Shen, Jincheng Shi, Juanfeng Shi, Lidong Shi, Shengyuan Shi, Wingchun Siu, Pengwei Song, Xiaoxi Song, Jianlin Su, Yunfeng Su, Zhaochen Su, Lin Sui, Jingsong Sun, Junyao Sun, Shaoning Sun, Shuzhe Sun, Tongyu Sun, Yujun Sun, Yunpeng Tai, Chuning Tang, Heyi Tang, Sirui Tang, Zecheng Tang, Chaoran Tian, Rongpeng Tian, Yu Tian, Wei Tu, Chensi Wang, Chuang Wang, Chunjie Wang, Dinglu Wang, Feng Wang, Hailong Wang, Haiming Wang, Hao Wang, Hao Wang, Huaqing Wang, Hui Wang, Jiayi Wang, Jinglong Wang, Jinhong Wang, Jiuzheng Wang, Linian Wang, Shaobo Wang, Shenzhi Wang, Shuyi Wang, Si Wang, Siyuan Wang, Tianfu Wang, Wenjue Wang, Xingran Wang, Xinmei Wang, Xinyuan Wang, Xusheng Wang, Yalin Wang, Yangkun Wang, Yao Wang, Yaoyu Wang, Yejie Wang, Yiqin Wang, Yucheng Wang, Yuzhi Wang, Zhaoji Wang, Zhaowei Wang, Zhengtao Wang, Zhenhao Wang, Zhongsheng Wang, Zifan Wang, Chu Wei, Ming Wei, Shouxin Wei, Zichen Wen, Fan Wu, Haoning Wu, Rucong Wu, Wenhao Wu, Xiaoxue Wu, Yingcong Wu, Yongqi Wu, Yuxin Wu, Zijian Wu, Xinglang Xian, Chenxuan Xiang, Yuye Xiang, Bocheng Xiao, Chenjun Xiao, Xin Xiao, Jin Xie, Xiaotong Xie, Yifeng Xie, Zhe Xie, Bowei Xing, Yiming Xiong, Baosheng Xu, Boyu Xu, Jiale Xu, Jianfan Xu, Jing Xu, Jinjing Xu, L. H. Xu, Qingtao Xu, Shuyao Xu, Suting Xu, Tiantian Xu, Tianxiang Xu, Weixin Xu, Xinran Xu, Yangchuan Xu, Ye Xu, Yueni Xu, Ziyao Xu, Haonan Xue, Junjie Yan, Yaoyao Yan, Fan Yang, Guangyao Yang, Hao Yang, Junwei Yang, Ruoyu Yang, Wenjie Yang, Xiaofei Yang, Xinyu Yang, Yi Yang, Yiling Yang, Ying Yang, Yuchen Yang, Zhen Yang, Zhilin Yang, Zian Yang, Zuhao Yang, Haotian Yao, Dan Ye, Haoran Ye, Wenjie Ye, Zhanbo Ye, Bohong Yin, Haoxiang Yin, Xietong Yin, Chengzhen Yu, Haozhen Yu, Longhui Yu, Shengnan Yu, Shuying Yu, Tianxiang Yu, Enming Yuan, Mengjie Yuan, Tongtian Yue, Wei Yue, Yang Yue, Dunyuan Zha, Haobing Zhan, B. H. Zhang, Dehao Zhang, Fei Zhang, Hao Zhang, Haoyuan Zhang, Huanyu Zhang, Jiapei Zhang, Jiaxuan Zhang, Jin Zhang, Kaiyi Zhang, Miaozhen Zhang, Puqi Zhang, Qinglei Zhang, Rong Zhang, Rui Zhang, Shaoshuai Zhang, Shiyi Zhang, Xiaobin Zhang, Xiaoyun Zhang, Y. Zhang, Yangkun Zhang, Ye Zhang, Yichi Zhang, Yikun Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Zijing Zhang, Bin Zhao, Chenguang Zhao, Feifan Zhao, Jinglun Zhao, Jinxiang Zhao, Shuai Zhao, Wenshuo Zhao, Xiangyu Zhao, Xuanle Zhao, Yikai Zhao, Zijia Zhao, Haozhi Zheng, Huabin Zheng, Ruihan Zheng, Shaojie Zheng, Tengyang Zheng, Haofeng Zhong, Lei Zhong, Longguang Zhong, M. Zhou, Qiankang Zhou, Runjie Zhou, Ruozhang Zhou, Xinyu Zhou, Yiqiao Zhou, Zaida Zhou, Jinguo Zhu, Liya Zhu, Xinhao Zhu, Yangjunfeng Zhu, Yuxuan Zhu, Zhen Zhu, Chen Zhuang, Weiyu Zhuang, Xinxing Zu

arXiv:2607.24653v2cs.CLcs.LG

TL;DR

Frontier models increasingly scale through both pretraining and test-time computation, motivating broader evidence on open models. Kimi K3 combines a large native multimodal MoE architecture with reinforcement learning across multiple domains and reasoning efforts, achieving frontier-level performance across coding, agentic, knowledge, reasoning, and vision tasks while trailing the strongest proprietary systems.

  • Problem

    As test-time computation becomes a second scaling axis for language models, open frontier systems require evaluation across broad reasoning and agentic capabilities.

  • Method

    Kimi K3 combines a 2.8T-parameter native multimodal MoE architecture with KDA, Attention Residuals, Stable LatentMoE, refined training, and multi-domain reinforcement learning.

  • Results

    Kimi K3 achieves frontier-level performance across coding, agentic, knowledge, reasoning, and vision tasks, trailing the strongest proprietary systems but outperforming other evaluated models.

  • Takeaways & Limitations

    Kimi K3 establishes a new open frontier and releases its weights for broader research, deployment, and innovation.

  • Takeaways & Limitations

    In cybersecurity evaluation, unsolved tasks expose a remaining gap to human-level capability, including failures in exploit completion, strategy selection, debugging, and verification.

Abstract

from arXiv · show

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.

1 Introduction

Kimi K3 is a 2.8T-parameter native multimodal MoE model with 104B activated parameters and a 1M-token context window, combining architectural, post-training, and infrastructure advances for open-frontier intelligence. It trails the strongest proprietary systems overall but consistently surpasses other evaluated open and proprietary models, and its full weights are released.

  • Pre-training: Kimi K3 has 2.8T total parameters, 104B activated parameters, and a context window of up to one million tokens.Its native multimodal architecture uses Kimi Delta Attention, Gated MLA, Attention Residuals, and Stable LatentMoE to scale information flow across sequence length, depth, and width.
  • Post-training: Reinforcement learning spans long-horizon coding, general agents, general reasoning, and knowledge tasks across multiple reasoning-effort levels for 1M-context test-time scaling.Training includes verifiable search, professional knowledge work, software engineering, kernel optimization, multimodal reasoning with vision-in-the-loop tool use, and persistent assistant workflows.
  • Infrastructure: Infrastructure contributions include KDA systems co-design, MoonEP and memory-efficient 2.8T-parameter MoE pre-training, and co-located RL with resumable sandboxes for million-token trajectories.KDA support includes fused kernels, context parallelism, and state-aware prefix caching; MoonEP provides perfectly balanced expert execution with static computation shapes and zero-copy communication.
  • Results and release: Kimi K3 trails Claude Fable 5 and GPT-5.6 Sol overall but consistently leads the other open and proprietary models evaluated across coding, agentic, knowledge, reasoning, and vision benchmarks.The full model weights are released for research, deployment, and further innovation.

2 Model Architecture · 2.1 Hybrid Attention · 2.2 Attention Residuals

Kimi K3 scales information flow across sequence length, depth, and width using hybrid KDA–Gated MLA blocks and Attention Residuals. Its bounded-decay KDA and blockwise AttnRes designs support efficient long-context mixing while reducing memory, communication, and inference overhead.

  • 2.1 Hybrid Attention: KDA provides long-context token mixing through recurrent delta-rule updates with channel-wise forget gates, while Gated MLA retains unrestricted global token interaction.KDA applies channel-wise decay before its delta-rule update; MLA compresses key–value states into latent vectors and reconstructs them during attention.
  • 2.1 Hybrid Attention: Each backbone block contains 3 KDA layers followed by 1 Gated MLA layer, with an additional Gated MLA layer at the end for final global attention.This establishes a repeated 3:1 local-to-global attention pattern while ensuring the final layer performs global attention.
  • 2.1 Hybrid Attention: Kimi K3 bounds KDA log-decay with a scaled sigmoid using gmin = −5, preventing unbounded reciprocal decay and enabling dense Tensor Core computation for all causal tiles.Every retention factor satisfies αh t,j > e−5 ≈6.7 × 10−3, keeping 16-token-tile cumulative decay within (−80, 0).
  • 2.1 Hybrid Attention: Gated MLA applies NoPE to its layers and uses an input-dependent, channel-wise full-rank output gate, letting KDA provide positional mixing while MLA supplies global content access.The full-rank gate matches KDA’s updated output-gate parameterization and modulates channels read from global attention.
  • 2.2 Attention Residuals: Attention Residuals let each layer selectively retrieve representations from all preceding layers instead of compressing prior information into one standard residual state.Full AttnRes uses layer-specific pseudo-queries, normalized keys, and softmax attention over earlier layer outputs.
  • 2.2 Attention Residuals: For L < 100, full AttnRes has affordable O(L2d) arithmetic but requires O(Ld) memory and cross-stage communication to keep all layer outputs alive.Block AttnRes addresses this practical overhead by aggregating layers into block-level representations.
  • 2.2 Attention Residuals: Block AttnRes reduces memory and communication overhead from O(Ld) to O(Nd), bounds inference-time state, and significantly reduces inference-time cost through online-softmax merging.The method combines parallel inter-block results with sequential intra-block partial sums.
  • 2.2 Attention Residuals: N ≈8 recovers most of AttnRes’s benefit across model scales; Kimi K3 uses 8 blocks with 12-layer size, producing 9 total blocks including the embedding layer.The ninth counted block is a partial final block.

2.3 Stable LatentMoE · 2.4 Native Vision · 2.5 Per-Head Muon

These sections present Stable LatentMoE for stabilizing extreme sparse routing, native multimodal processing through a jointly trained vision pathway, and per-head Muon for attention projections. Together, they address activation growth, expert-load balance, vision integration, and optimizer treatment at Kimi K3’s scale.

  • 2.3 Stable LatentMoE: Stable LatentMoE combines pre-up-projection RMSNorm, SiTU-GLU, and Quantile Balancing to address routed-branch activation explosion and load imbalance.The motivating failures arise from nearly four consecutive routed-path matrix multiplications and balancing a routed expert pool of 896 per layer.
  • 2.3 Stable LatentMoE: RMSNorm between routed-expert aggregation and the up-projection reduces sensitivity to scale variation before the routed branch joins the shared branch.Kimi K3 fixes the number of full-width shared experts to Ns = 2 in every layer.
  • 2.3 Stable LatentMoE: SiTU-GLU uses smooth soft caps on both the Swish gate’s linear factor and the up branch to control large-value growth while preserving SwiGLU’s local response.Kimi K3 sets β1 = 4 for the gate branch and β2 = 25 for the up branch.
  • 2.3 Stable LatentMoE: Quantile Balancing sets expert biases from router-score quantiles to target equal loads without changing mixture weights or router gradients.For m tokens, n experts, and Top-k routing, the target load is q := mk/n tokens per expert; inference uses a frozen final bias.
  • 2.4 Native Vision: Kimi K3 processes text, images, and videos through one shared backbone in a single context, enabling iterative code-based inspection and refinement of visual outputs.Rendered outputs and their generating code occupy the same token stream, supporting vision-in-the-loop behavior.
  • 2.4 Native Vision: MoonViT-V2 is trained from scratch with next-token prediction, and its architecture uses 27 transformer layers, roughly 0.4B parameters, RMSNorm, and bias-free projections.Visual inputs pass through MoonViT-V2 and a lightweight MLP projector into the language model.
  • 2.5 Per-Head Muon: Per-head Muon partitions attention-projection momentum matrices by head and applies Newton–Schulz orthogonalization separately to each block.This differs from full-matrix orthogonalization, which treats all attention heads as one coupled block.

3 Pre-Training

Kimi K3’s pre-training combines curated multimodal data, dedicated scaling-law retuning, and native joint language–vision optimization. Progressive context extension supports contexts up to 1 million tokens, while scaling studies report a 2.5× efficiency gain over Kimi K2.

  • Training data: Kimi K3 is pre-trained on curated Web Text, Code, Mathematics, and Knowledge corpora plus a large-scale vision corpus.Vision data includes captions, interleaved image–text documents, OCR, perception, video, and visual coding data; text domains use filtering, quality scoring, deduplication, and domain-specific sampling.
  • Scaling laws: 2.5× gain in scaling efficiency over Kimi K2 is reported by Kimi K3’s fitted scaling-law curves.The studies retune batch size, learning rate, tokens-per-parameter ratio (TPP), and model shape using held-out OOD validation data.
  • Multimodal training: Native multimodal training jointly optimizes language and vision from the start, interleaving visual and textual tokens under a single next-token prediction objective.This avoids post-hoc vision-encoder alignment and trains the shared backbone to learn unified multimodal representations.
  • Long-context training: Up to 1 million tokens are supported through progressive context extension during pre-training and cooldown.The curriculum grows from 8K to 64K tokens during pre-training and from 256K to 1M tokens during cooldown, concentrating costly long-sequence computation in a small fraction of training.

4 Post-Training

Kimi K3’s post-training uses a three-stage pipeline combining SFT, domain- and reasoning-effort-specific RL, and MOPD consolidation. It further develops long-horizon agent training through configurable environments and verify-in-the-loop Autonomous Execution Tasks.

  • Post-training pipeline: The pipeline progresses from SFT cold-start initialization to specialized RL experts and Multi-Teacher On-Policy Distillation consolidation.MOPD unifies domain-specific policies across varying reasoning-effort levels.
  • Reinforcement learning: RL spans general, agentic, and coding domains, training one expert per domain at each reasoning-effort level.The general domain includes vision, reasoning, faithfulness, search, and knowledge-work tasks.
  • Reasoning Effort RL: Per-problem budget control penalizes trajectories whose token usage exceeds a scaled threshold, targeting reasoning-effort and token efficiency.The budget measures thinking tokens for general tasks and cumulative usage for agentic tasks.
  • Agentic environments: A configurable white-box RL environment varies tools, prompts, context management, skills, memories, and subagents to reduce overfitting to a single agent harness.Different harness configurations are dynamically composed for task groups and domains.
  • Agentic environments: Autonomous Execution Tasks train long-horizon agent intelligence through verify-in-the-loop optimization without reference trajectories or predefined procedures.Each task provides an initial state, constrained goal, tool-based action space, execution budgets, and an independent verifier.

5 Infrastructure

Kimi K3’s infrastructure is co-designed for hybrid KDA attention, 3T-class sparse multimodal workloads, and million-token agentic execution. It combines KDA-specific parallelism and caching with balanced expert execution, memory management, and large-scale sandboxed workloads.

  • KDA systems: KCP preserves KDA’s recurrent-state dependence with fixed-size synchronization and achieves linear compute scaling.It decomposes each segment into a cumulative transition and locally generated state, requiring only fixed-size all-gather synchronization.
  • Balanced multimodal training: MoonEP3 achieves perfect expert-parallel load balance by assigning every rank exactly S × K tokens and using dynamically planned redundant experts.A balanced plan exists with at most E/R redundant experts per rank, while static computation shapes eliminate per-layer MoE host synchronization.
  • Memory-efficient training: A unified activation manager composes recomputation, quantization, and offload policies at tensor granularity without coupling them to model code.Each backward-saved tensor uses a pluggable storage backend, and policies are declared through lightweight annotations.
  • Agentic workloads: 51,219,741 sandboxes were created across 1,505,678 images during Kimi K3’s training and evaluation.These workloads support the infrastructure requirements of million-token agentic execution.
  • Deployment: Kimi K3’s KDA-aware prefix cache jointly manages recurrent-state and conventional cache types to retain and reuse million-token prefixes efficiently.It decouples fine-grained prefix hashing from coarse physical allocation and checkpoints KDA states at sparse hash endpoints; a 2,800-token match can resume at B = 2,560.

6 Evaluations

Kimi K3 demonstrates frontier-level performance across reasoning, coding, agentic, vision, cyber, and external evaluations, while generally trailing the strongest proprietary models. Its strongest results include agentic orchestration, coding, multimodal tasks, vulnerability discovery, and cost-efficient deployment, with persistent gaps in research-level reasoning and hardened-target exploit completion.

  • Overall evaluation: Kimi K3 closely trails Claude Fable 5 and GPT-5.6 Sol while consistently outperforming other evaluated open and proprietary models across the benchmark suite.The suite spans reasoning and knowledge, coding, agentic, and vision capabilities.
  • Reasoning & Knowledge: 93.5% on GPQA Diamond demonstrates competitive graduate-level reasoning, but HLE-Full scores of 56.0% and 43.5% with and without tools respectively reveal a research-level gap.Kimi K3 scores 23.4% on CritPt and trails the leading models there as well.
  • Coding: 77.8% on ProgramBench is the best score, while 42.0% on SWE-Marathon leads Claude Fable 5 by 7 points and 88.3% on Terminal-Bench 2.1 nearly matches GPT-5.6 Sol’s 88.8%.Kimi K3 ranks second on FrontierSWE with 81.2% as of July 16, 2026.
  • Agentic: 91.2% on BrowseComp, 95.0% F1 on DeepSearchQA, and 94.5% on MCPMark-Verified exemplify state-of-the-art agentic performance, despite third place on GDPval-AA v2 at 1,686.Other reported results include 76.2% on ResearchRubrics, 30.8% on AutomationBench, 34.8% on SpreadsheetBench 2, 33.4% on τ 3-Banking, and 94.6% criterion pass rate on Harvey Lab-AA.
  • Vision: 97.8% on Math-Vision and 41.0% pass@5 on ZeroBench-main with Python tools show strong multimodal gains; Kimi K3 also leads OmniDocBench at 91.1%.Without Python tools, Math-Vision is 94.3% and ZeroBench-main is 23.0% pass@5, where it ties Claude Fable 5.
  • Cybersecurity: Approximately 70% of human-reviewed vulnerability findings were confirmed genuine, including 16 previously unknown vulnerabilities across six projects, while exploit-suite performance was 38.9% versus GLM-5.2’s 22.2%.The model remains behind human experts and frontier cyber models on hardened-target exploit completion, including arbitrary code execution on 0 of 41 tasks.
  • External evaluations: 57.1 on Intelligence Index v4.1 ranks Kimi K3 fourth of 580 models, while 74.7% on the Vals Index ranks it second of 39 and 1,678 Elo ranks it first of 99 on WebDev Arena.These external evaluations place Kimi K3 behind Claude Fable 5 on the Intelligence and Vals indices but ahead of other evaluated models.
  • Efficiency: 91.2% on BrowseComp costs $2.03 per task, half GPT-5.6 Sol’s cost, while GDPval-AA v2 is within 50 Elo of GPT-5.6 Sol at 13% lower cost.Kimi K3 is also reported at 2.6× cheaper than Claude Fable 5 on GDPval-AA v2 and roughly half Claude Fable 5’s cost on AA-Briefcase.

7 Case Studies

Kimi K3 demonstrated broad capabilities across GPU optimization, compiler development, chip design, research coding, knowledge work, and multimodal video production. These case studies show autonomous execution of complex technical workflows, including long iterative processes and end-to-end implementation and verification.

  • GPU kernel optimization: Kimi K3 substantially improved performance across four representative GPU kernels in independently configured sandboxes with up to 24 hours per task.The kernels were AttnRes, DeepSeek Sparse Attention (DSA), KDA, and MLA with head dimension 512, tested on an NVIDIA Hopper GPU and an alternative-vendor GPGPU.
  • GPU compiler development: Kimi K3 developed MiniTriton5, a compact Triton-like compiler with custom frontend, layout, MLIR optimization, and PTX code-generation layers.It also built a dual-mode tensor library with a PyTorch-like interface and shared DSL compiler and runtime.
  • Chip design: In a single 48-hour autonomous run, Kimi K3 built, optimized, and verified an inference-chip prototype using open-source EDA tools and the Nangate45 standard-cell library.The prototype followed a nano-model architecture with hybrid KDA and NoPE-MLA attention, Block AttnRes, sigmoid-based MoE routing, one shared expert, and group-wise INT4 quantization with group size 128.
  • Coding for research: In about two hours, Kimi K3 reproduced I–Love–Q universal relations by reviewing more than 20 papers, evaluating over 300 equations of state, and writing more than 3,000 lines of Python.It cross-validated results, identified inconsistencies in published formulas, implemented the numerical pipeline, and produced an interactive HTML dashboard, versus a typical one to two weeks for an experienced researcher.
  • Knowledge work: Kimi K3 produced an interactive AI ASIC industry research website from 87 quarterly reports and 99 original PDFs totaling more than 11,000 pages.The work involved more than 120 refinement rounds, over 2,800 web searches, and over 1,100 terminal queries; a second case analyzed 391 gravitational-wave events using more than 20 concurrent subagents and produced seven scientific visualizations.
  • Video editing and motion design: Kimi K3 created a 3Blue1Brownstyle motion-graphics explainer of its architecture and edited a teaser video from 56 source clips.The workflow included clip selection, motion-matched cuts, frame-accurate beat synchronization, audio processing, and multiple revision rounds; comparable production would typically take an experienced editor one to two days.

8 Conclusion

Kimi K3 is an open 2.8-trillion-parameter Mixture-of-Experts model with native vision and a 1-million-token context window. It delivers frontier-level performance across multiple demanding task domains while remaining behind the strongest proprietary models.

  • Conclusion: Kimi K3 is an open 2.8-trillion-parameter Mixture-of-Experts model with native vision capabilities and a 1-million-token context window.It is built on Kimi Delta Attention and Attention Residuals.
  • Conclusion: Kimi K3 delivers frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks.
  • Conclusion: Although gaps to the strongest proprietary models remain, Kimi K3 establishes a new open frontier within everyone’s reach.

A Contributions

The contributions section presents the contributors in alphabetical order by last name, spanning the extensive author list from Bai through Zhu.

  • The contributor list is explicitly ordered alphabetically by last name.
  • The list begins with contributors whose surnames include Bai, Bao, Cai, Cao, and Chen.
  • The alphabetical list continues through surnames including Wang, Xiang, Xiao, Xu, Yang, Zhang, Zhao, Zheng, Zhou, and Zhu.

B Details of Sigmoid Tanh Unit GLU

SiTU-GLU bounds the SwiGLU product while preserving Swish’s approximately linear response near zero and vanishing negative tail. It smoothly caps both gate and up branches to control large positive activations and prevent either branch from dominating.

  • Design goal: SiTU-GLU bounds the SwiGLU product while preserving Swish’s approximately linear response around the origin and vanishing negative tail.The design goal retains the characteristic scalar response shape.
  • Branch construction: SiTU caps the Swish linear factor as β1 tanh(Wgx/β1) while retaining the sigmoid factor, primarily controlling large positive activations.The sigmoid continues driving negative gate responses toward zero, so the negative tail is preserved.
  • Branch construction: Kimi K3 applies the same construction to the up branch as β2 tanh(Wux/β2), preventing either branch from dominating the product.Both gate and up branches are smoothly capped.

Local and limiting behavior For a scalar z near the origin, the scaled tanh satisfies

SiTU-GLU matches SwiGLU to first order near the origin and recovers it pointwise as β1, β2 →∞. Its tanh–Sigmoid construction bounds outputs, while smooth capping with β1 = 4 and β2 = 25 preserves nonzero gradients away from saturation boundaries and improves training behavior.

  • Local and limiting behavior: SiTU-GLU matches SwiGLU to first order around the origin and recovers it pointwise as β1, β2 →∞.
  • Bounded output: Every output coordinate is bounded because | tanh(z)| < 1 and 0 < Sigmoid(z) < 1.
  • Bounded output: With β1 = 4 and β2 = 25, smooth capping preserves nonzero gradients away from saturation boundaries and gives better training behavior than hard clamping.

C Derivation of Quantile Balancing

This appendix derives Quantile Balancing as an exact dual-coordinate method for maximum-score balanced expert assignment. The resulting routing uses frozen expert biases with ordinary fixed Top-k selection at inference, preserving train–inference consistency.

  • Assignment formulation: Quantile Balancing starts from maximum-score assignment where each token selects k experts and every expert serves exactly mk/n tokens.The assignment uses binary variables xi,j indicating token–expert assignments.
  • Dual derivation: The linear relaxation is exact because the bipartite b-matching polytope is integral, enabling a dual formulation with token- and expert-side equality multipliers.Minimax exchange follows from linearity and convex feasible sets.
  • Quantile updates: Alternating coordinate minimization sets both token and expert multipliers using the same (1 − k/n)-th quantile along their respective axes.The token update uses the kth-largest adjusted score, while the expert update uses the (mk/n+1)-th-largest adjusted score.
  • Routing implementation: At optimum, selected experts are the Top-k entries of si − β*, so deployment needs only frozen expert biases and no quantile computation.Token thresholds are batch-dependent intermediate variables and are discarded; the equivalent routing bias is b = −β.
  • Related methods: QB is an exact coordinate-minimizer counterpart to sign-based loss-free balancing, while inequality-constrained BIP updates add clipping that suppresses over-selected experts and slows equilibration.Unlike Expert Threshold routing, the resulting scheme keeps fixed Top-k selection rather than allowing a variable number of experts per token.

D Histogram-Based Quantile Estimation

Kimi K3 replaces impractical exact whole-step quantile computation with per-expert histograms that aggregate locally and recover global quantiles after one all-reduce. The estimator bounds quantile error by bin width while keeping communication independent of the number of tokens.

  • Method: Per-expert histograms replace exact quantiles over O(mn) margins with fixed-cost distribution summaries during training.The method histograms the required biases r_i,j := α_i − s_i,j rather than retaining all margins.
  • Binning range: Biases fall within [b_min − 1, b_max + 1], which is partitioned into B uniform bins.This range follows from sigmoid router scores and the cutoff being a biased score.
  • Accumulation and recovery: Each rank accumulates per-expert counts locally across micro-batches, then one all-reduce produces pooled histograms for identical quantile recovery.The target rank is exact because each expert histogram counts every token once.
  • Properties: B = 1000 limits quantile error to a few 10^-3, with no measurable residual load imbalance.The true and estimated quantiles lie in the same bin, so error is bounded by the bin width.
  • Properties: Communication is one integer all-reduce of nB values per layer per step, independent of m.In the stated configuration, this communication is below 1% of the truncated reference quantity.

E MoonEP General Upper Bound Proof

The section proves a general upper bound M(I) ≤ E/R for redundant experts in expert-parallel planning and shows it is essentially tight. The proof constructs balanced token assignments, while an adversarial router output attains ⌈E(R −1)/R2⌉≈E/R.

  • General upper bound: M(I) ≤ E/R always, where M(I) minimizes the maximum redundant experts assigned to any rank.This is Theorem 1’s general upper bound for every router output I.
  • Tightness: ⌈E(R −1)/R2⌉≈E/R redundant experts are necessary for some router outputs, establishing near-tightness of the upper bound.Theorem 2 routes no tokens to rank 0 while evenly distributing all tokens across experts on the other R −1 ranks.
  • General upper bound: A plan exists in which every expert-parallel rank receives exactly S × K tokens, with each rank’s remote tokens sourced from only one other rank.The construction repeatedly matches underloaded and overloaded ranks, beginning with each rank’s local tokens.
  • Tightness: For the adversarial construction, rank 0 must receive S × K entirely remote tokens involving at least SK distinct experts, forcing the stated ceiling lower bound.A matching filling procedure with expert-wise preferential migration achieves the same value, so equality holds.

F Chat Template

The Kimi K3 chat template is designed for extensibility, low alignment cost, and unified generation, using backward-compatible messages and natural-language options rather than template revisions or dedicated syntax. It organizes assistant outputs into explicit reasoning, response, and tool channels while supporting dynamically expanding toolsets and unambiguous tool-call association.

  • Design goals: The template prioritizes extensibility, low alignment tax, and unified generation through backward-compatible message formats, minimal supervised data, and a single model-wide template.This design allows a lightly fine-tuned pretrained model to proceed directly to reinforcement learning.
  • Messages and zones: Messages encode request inputs and scoped options, while dynamically loaded tools expand the available toolset through additional tool-declare messages without rebuilding prior context.Global options include tool declarations and reasoning effort; input messages cover system, user, assistant, and tool roles.
  • Channels: Assistant messages use think, response, and tools channels for reasoning traces, user-visible answers, and tool calls, with thinking or instruct mode selected by the generation prefix.Kimi K3 supports only preserved thinking in thinking mode.
  • Tool calling: Tool calls pair tool and index attributes with matching ordered tool-result messages, while arguments preserve raw strings and compactly serialize other JSON types.This makes free-form text such as code a first-class value rather than escaped JSON.
  • Reasoning effort and options: Option messages express tool_choice, response_format, and thinking-effort as short natural-language instructions, enabling new options with little or no additional training.Reasoning effort is represented by a global thinking-effort message stating a natural-language level rather than changing the generation prefix or exposing a token budget.
Loading 2607.24653v2…