Source-linked AI summary
Kimi K2.5: Visual Agentic Intelligence
Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, S. H. Cai, Yuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Cheng Chen, Guanduo Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kefan Chen, Liang Chen, Ruijue Chen, Xinhao Chen, Yanru Chen, Yanxu Chen, Yicun Chen, Yimin Chen, Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen, Zhirong Chen, Ziwei Chen, Dazhi Cheng, Yean Cheng, Minghan Chu, Jialei Cui, Jiaqi Deng, Muxi Diao, Hao Ding, Mengfan Dong, Mengnan Dong, Yuxin Dong, Yuhao Dong, Angang Du, Chenzhuang Du, Dikang Du, Lingxiao Du, Yulun Du, Yu Fan, Shengjun Fang, Qiulin Feng, Yichen Feng, Garimugai Fu, Kelin Fu, Hongcheng Gao, Tong Gao, Yuyao Ge, Shangyi Geng, Chengyang Gong, Xiaochen Gong, Zhuoma Gongque, Qizheng Gu, Xinran Gu, Yicheng Gu, Longyu Guan, Shuhao Guan, Yuanying Guo, Xiaoru Hao, Dailan He, Tianhong He, Weiran He, Wenyang He, Yibo He, Yunjia He, Chao Hong, Hao Hu, Jiaxi Hu, Yangyang Hu, Zhenxing Hu, Ke Huang, Ruiyuan Huang, Weixiao Huang, Zhiqi Huang, Chaobo Jia, Tao Jiang, Zhejun Jiang, Xinyi Jin, Yu Jing, Guokun Lai, Aidi Li, C. Li, Cheng Li, Fang Li, Guanghe Li, Guanyu Li, Haitao Li, Haoyang Li, Jia Li, Jingwei Li, Junxiong Li, Lincan Li, Mo Li, Weihong Li, Wentao Li, Xinhang Li, Xinhao Li, Yang Li, Yanhao Li, Yiwei Li, Yuxiao Li, Zhaowei Li, Zhaoxi Li, Zheming Li, Weilong Liao, Jiawei Lin, Xiaohan Lin, Yibo Lin, Zhishan Lin, Zichao Lin, Cheng Liu, Chenyu Liu, Hongzhang Liu, Liang Liu, Shaowei Liu, Shudong Liu, Shuran Liu, Tianwei Liu, Tianyu Liu, Weizhou Liu, Xiangyan Liu, Yangyang Liu, Yanming Liu, Yibo Liu, Yuanxin Liu, Zhengying Liu, Zhongnuo Liu, Enzhe Lu, Haoyu Lu, Zhiyuan Lu, G. Luo, Junyu Luo, Tongxu Luo, Yashuo Luo, Long Ma, Shaoguang Mao, Yuan Mei, Xin Men, Fanqing Meng, Zhiyong Meng, Yibo Miao, Minqing Ni, Kun Ouyang, Siyuan Pan, Bo Pang, Yuchao Qian, Ruoyu Qin, Zeyu Qin, Jiezhong Qiu, Bowen Qu, Zeyu Shang, Youbo Shao, Tianxiao Shen, Zhennan Shen, Juanfeng Shi, Lidong Shi, Shengyuan Shi, Feifan Song, Pengwei Song, Tianhui Song, Xiaoxi Song, Hongjin Su, Jianlin Su, Zhaochen Su, Lin Sui, Jinsong Sun, Junyao Sun, Tongyu Sun, Flood Sung, Yunpeng Tai, Chuning Tang, Heyi Tang, Xiaojuan Tang, Zhengyang Tang, Jiawen Tao, Shiyuan Teng, Chaoran Tian, Pengfei Tian, Bowen Wang, Chensi Wang, Chuang Wang, Congcong Wang, Dingkun Wang, Dinglu Wang, Dongliang Wang, Feng Wang, Hailong Wang, Haiming Wang, Hao Wang, Hengzhi Wang, Huaqing Wang, Hui Wang, Jiahao Wang, Jinhong Wang, Jiuzheng Wang, Kaixin Wang, Linian Wang, Qibin Wang, Shengjie Wang, Shuyi Wang, Si Wang, Wei Wang, Xiaochen Wang, Xinyuan Wang, Yao Wang, Yejie Wang, Yipu Wang, Yiqin Wang, Yucheng Wang, Yuzhi Wang, Zhaoji Wang, Zhaowei Wang, Zhengtao Wang, Zhexu Wang, Zifan Wang, Zihan Wang, Zizhe Wang, Chu Wei, Ming Wei, Chuan Wen, Zichen Wen, Chengjie Wu, Haoning Wu, Junyan Wu, Rucong Wu, Wenhao Wu, Yuefeng Wu, Yuhao Wu, Yuxin Wu, Zijian Wu, Chenjun Xiao, Jin Xie, Xiaotong Xie, Yuchong Xie, Bowei Xing, Boyu Xu, Jianfan Xu, Jing Xu, Jinjing Xu, L. H. Xu, Lin Xu, Suting Xu, Weixin Xu, Xinbo Xu, Xinran Xu, Yangchuan Xu, Yichang Xu, Yuemeng Xu, Zelai Xu, Ziyao Xu, Junjie Yan, Yuzi Yan, Guangyao Yang, Hao Yang, Junwei Yang, Kai Yang, Ningyuan Yang, Xiaofei Yang, Xinlong Yang, Xinyu Yang, Ying Yang, Yi Yang, Yi Yang, Zhen Yang, Zhilin Yang, Zonghan Yang, Haotian Yao, Dan Ye, Haoran Ye, Wenjie Ye, Zhuorui Ye, Peng Yebo, Bohong Yin, Chengzhen Yu, Longhui Yu, Tao Yu, Tianxiang Yu, Enming Yuan, Mengjie Yuan, Xiaokun Yuan, Yang Yue, Weihao Zeng, Dunyuan Zha, Haobing Zhan, Dehao Zhang, Hao Zhang, Jin Zhang, Puqi Zhang, Qiao Zhang, Rui Zhang, Xiaobin Zhang, Xiaoyun Zhang, Y. Zhang, Yadong Zhang, Yangkun Zhang, Yichi Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yushun Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Chenguang Zhao, Feifan Zhao, Jinxiang Zhao, Shuai Zhao, Xiangyu Zhao, Xuanle Zhao, Yikai Zhao, Zijia Zhao, Huabin Zheng, Ruihan Zheng, Shaojie Zheng, Tengyang Zheng, Junfeng Zhong, Longguang Zhong, Weiming Zhong, M. Zhou, Runjie Zhou, Xinyu Zhou, Zaida Zhou, Jinguo Zhu, Liya Zhu, Xinhao Zhu, Yuxuan Zhu, Zhen Zhu, Jingze Zhuang, Weiyu Zhuang, Ying Zou, Xinxing Zu
TL;DR
Multimodal agentic systems face questions about effective vision–text training and limits from sequential execution. Kimi K2.5 jointly optimizes text and vision and introduces Agent Swarm for parallel task decomposition, achieving state-of-the-art results across agentic and other benchmarks while reducing latency.
Problem
Existing agentic models rely on sequential tool execution that scales inference time linearly, while multimodal training strategies for a fixed vision–text budget remain an open design question.
Method
Kimi K2.5 jointly optimizes text and vision and uses Agent Swarm with reinforcement-learned orchestration to decompose heterogeneous tasks into concurrently executed sub-problems.
Results
Kimi K2.5 achieves state-of-the-art performance across coding, vision, reasoning, and agentic benchmarks, including 74.9% on BrowseComp with Discard-all context management.
Takeaways & Limitations
Joint vision–text intelligence and parallel agent execution support scalable, general-purpose agentic systems for complex workloads.
Abstract
from arXiv · showhide
We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of techniques such as joint text-vision pre-training, zero-vision SFT, and joint text-vision reinforcement learning. Building on this multimodal foundation, K2.5 introduces Agent Swarm, a self-directed parallel agent orchestration framework that dynamically decomposes complex tasks into heterogeneous sub-problems and executes them concurrently. Extensive evaluations show that Kimi K2.5 achieves state-of-the-art results across various domains including coding, vision, reasoning, and agentic tasks. Agent Swarm also reduces latency by up to $4.5\times$ over single-agent baselines. We release the post-trained Kimi K2.5 model checkpoint to facilitate future research and real-world applications of agentic intelligence.
1 Introduction
Kimi K2.5 advances general-purpose agentic intelligence through joint text-vision optimization and Agent Swarm, a framework for dynamically orchestrating parallel agents. It reports strong performance across agentic and frontier benchmarks, including state-of-the-art visual-to-code and internal software-engineering evaluations.
- Joint Optimization of Text and Vision: Kimi K2.5 jointly optimizes text and vision, finding that early vision fusion with lower ratios can outperform late visual-token addition under a fixed total token budget.The paper presents joint text-vision optimization as enhancing both modalities while avoiding conflict.
- Joint Optimization of Text and Vision: Architecturally, MoonViT-3D combines native-resolution variable-resolution image processing with lightweight temporal compression for video understanding.It uses NaViT packing for images and groups consecutive video frames in fours for patch-level temporal averaging.
- Joint Optimization of Text and Vision: Text-only SFT activates visual reasoning and tool use, while human-designed visual trajectories hurt generalization; joint reinforcement learning then covers text and vision tasks.The reported findings are attributed to strong vision-text alignment established during joint pretraining.
- Parallel Agent Orchestration: Kimi K2.5 introduces Agent Swarm, using Parallel-Agent Reinforcement Learning with sub-agent creation and task delegation to address sequential agents’ latency and scalability limits.The framework optimizes tool execution with verifiable rewards while incorporating interfaces for sub-agent creation and task delegation.
- Overview: Kimi K2.5 achieves state-of-the-art visual-to-code generation and strong internal software-engineering results while scaling specialized-agent diversity and parallelism.The model is presented as a unified architecture integrating vision and language, thinking and instant modes, and chats and agents.
2 Joint Optimization of Text and Vision
Kimi K2.5 jointly optimizes text and vision through large-scale mixed-token pre-training, zero-vision SFT, outcome-based visual RL, and joint multimodal RL. This approach activates robust visual capabilities while also improving textual performance.
- Joint Pre-training: Kimi K2.5 extends Kimi K2 through joint pre-training on approximately 15 trillion mixed visual and text tokens, enhancing both modalities simultaneously.The model is natively multimodal rather than vision-adapted after linguistic training.
- Joint Pre-training: Early fusion with a lower vision ratio outperforms later high-ratio vision injection under a fixed total vision-text token budget.Ablations varying vision ratio and injection timing found that vision ratio has minimal impact on final multimodal performance, contrary to conventional wisdom [8] [21].
- Zero-Vision SFT: Zero-vision SFT activates visual and agentic capabilities using only text SFT data and programmatic IPython image operations.The method supports diverse cross-modal reasoning behaviors, while text-vision SFT performed much worse on visual and agentic tasks in preliminary experiments.
- Multimodal Reinforcement Learning: Outcome-based visual RL corrects failures where visual inputs are ignored by training on grounding, chart and document understanding, and vision-critical STEM tasks.The resulting trajectories support rejection-sampling fine-tuning and richer subsequent multimodal reasoning traces.
- Multimodal Reinforcement Learning: 84.7% → 86.4% on MMLU-Pro, 84.3% → 86.4% on GPQA-Diamond, and 56.7% → 58.9% on LongBench v2 after outcome-based visual RL.Visual RL improved textual performance while enhancing basic visual capabilities and complex agentic behaviors.
3 Agent Swarm
Agent Swarm addresses the limits of sequential single-agent execution by dynamically decomposing complex tasks, instantiating specialized subagents, and scheduling subtasks in parallel. PARL learns orchestration decisions and trains with rewards and critical-step constraints that encourage effective, feasible parallelization.
- Agent Swarm: Agent Swarm replaces sequential reasoning and tool calling with dynamic task decomposition, specialized subagent instantiation, and parallel subtask scheduling for complex tasks.Sequential execution creates bottlenecks as information gathering and multi-branch reasoning expand, while a single agent can exhaust practical reasoning depth and tool-call budgets.
- PARL Reward Training: PARL learns whether, when, and how to parallelize through environmental feedback rather than assuming parallelism is always beneficial.Its reward design combines task-level performance with auxiliary terms that discourage serial collapse and guide valid task decomposition.
- Architecture and Learning Setup: PARL uses a trainable orchestrator with frozen subagents instantiated from fixed intermediate policy checkpoints, avoiding end-to-end co-optimization to reduce credit-assignment ambiguity and training instability.The decoupled architecture separates orchestration from subagent policies in response to sparse and noisy outcome-based rewards.
- Critical Steps as Resource Constraint: Critical steps measure parallel execution by the longest-running branch, incentivizing balanced decompositions that reduce maximum execution time rather than merely creating more subtasks.Constraining training and evaluation with critical steps makes excessive subtask creation provide little benefit unless it shortens the critical path.
- Prompt Construction for Parallel-agent Capability Induction: Synthetic prompts target wide search across independent information sources and deep search across reasoning branches with delayed aggregation to induce parallel-agent capability.The prompt suite stresses sequential agentic execution using broad information gathering and multi-branch reasoning patterns.
4 Method Overview
Kimi K2.5 builds on the Kimi K2 MoE foundation with a native-resolution multimodal architecture, staged text-vision pre-training, and post-training methods for joint multimodal and agentic optimization. Its design includes shared image-video representations, token-efficient reinforcement learning, and decoupled multimodal training that reaches 90% of text-only training efficiency.
- Architecture: Kimi K2.5 combines the Kimi K2 MoE language model with a MoonViT-3D vision encoder and an MLP projector for native-resolution multimodal processing.The architecture follows design principles established in Kimi-VL [54].
- Architecture: MoonViT-3D shares parameters and an embedding space across images and videos by packing patches from up to four consecutive frames into a single spatiotemporal sequence.This lets the same attention mechanism operate across spatial and temporal dimensions.
- Pre-training: Pre-training proceeds through standalone ViT training, joint text-vision training, and long-context mid-training over approximately 15T tokens, extending language and multimodal capabilities.The joint stage uses additional vision-text tokens at 4K sequence length, while the third stage incorporates higher-quality data and long-context activation.
- Post-training: Post-training combines synthesized instruction data, joint text-vision reinforcement learning, and PARL within a Unified Agentic Reinforcement Learning Environment.The instruction-tuning data uses specialized domain pipelines, human annotation, prompt engineering, and multi-stage verification.
- Training efficiency: 90% multimodal training efficiency relative to text-only training is achieved through DEP, which balances load while decoupling optimization of the vision encoder and main backbone.K2.5 inherits Kimi K2’s parallel strategy.
5 Evaluations
Kimi K2.5 is evaluated across reasoning, coding, agentic, vision, video, and computer-use benchmarks against proprietary and open-source baselines, achieving competitive or state-of-the-art results across these domains. Agent Swarm further improves benchmark performance, reduces execution time, and provides proactive context management through parallel orchestration.
- Agentic Capabilities: 74.9% on BrowseComp with Discard-all context management exceeds GPT-5.2’s 65.8%, while Kimi K2.5 scores 60.6% without context management.Kimi K2.5 also leads evaluated models on DeepSearchQA, FinSearchCompT2&T3, and Seal-0.
- Comprehensive Results: Kimi K2.5 achieves competitive or state-of-the-art performance across reasoning, coding, agentic, vision, video, and computer-use benchmarks against proprietary and open-source baselines.Table 4 summarizes the comprehensive comparison across these capability domains.
- Computer-Use Capability: 63.3% OSWorld-Verified success using only GUI actions exceeds Qwen3-VL-235B-A22B at 38.1% and Operator at 42.9%, while approaching Claude Opus 4.5 at 66.3%.This result demonstrates strong real-world computer-use capability without external tools.
- Agent Swarm Performance: 78.4% on BrowseComp gives Agent Swarm a 17.8% absolute gain over single-agent K2.5 and surpasses GPT-5.2 Pro at 77.9%.WideSearch improves from 72.7% to 79.0% Item-F1, while the in-house swarm benchmark gains 16.7%.
- Execution Time Savings via Parallelism: 3× ∼4.5× lower execution time lets Agent Swarm reach target WideSearch performance faster than a single-agent baseline.The efficiency advantage scales with task complexity as the target Item-F1 increases from 30% to 70%.
- Agent Swarm as Proactive Context Management: Agent Swarm outperforms Discard-all in BrowseComp efficiency and accuracy by preserving orchestrator-level coherence while bounding subagent contexts and retaining essential coordination signals.This proactive context management differs from reactive truncation or history-discarding strategies.
6 Conclusions · A Contributors · B Pre-training
Kimi K2.5 demonstrates general agentic intelligence through joint text–vision optimization and parallel agent execution, while the paper also documents its contributors and vision-to-text pre-training analysis.
- 6 Conclusions: Joint text–vision optimization across pre-training and reinforcement learning provides cross-modal alignment and visual–text reasoning, while Agent Swarm concurrently executes heterogeneous sub-tasks to reduce latency and improve complex agentic workloads.These mechanisms jointly support scalable, general agentic intelligence.
- 6 Conclusions: Agent Swarm enables parallel execution of heterogeneous sub-tasks, reducing inference latency while improving performance on complex agentic workloads.The framework dynamically supports concurrent agent execution as part of Kimi K2.5’s agentic design.
- A Contributors: The authors are listed alphabetically by last name.The contributor section contains the full author list and identifies an affiliation marker for The University of Hong Kong.
- A Contributors: The contributor listing spans the paper’s extensive author roster, including the Kimi K2.5 project attribution.The list includes the † affiliation marker for The University of Hong Kong.
- B Pre-training: Early fusion with lower vision ratios tends to yield better results across vision and language tasks under a fixed vision–text token budget.Figure 9 compares 10:90, 20:80, and 50:50 vision-to-text ratios.
B.1 Joint-Training … C Infra
The study finds that early fusion yields more stable multimodal training than introducing vision later, while Kimi K2.5 is trained on curated text and vision corpora using scalable parallel infrastructure.
- B.1 Joint-Training: Early fusion maintains healthier, more stable text performance, whereas mid- and late-fusion training shows a temporary text-performance dip after vision data is introduced.The dip-and-recover behavior is attributed to modality domain shift disrupting established linguistic representations; early exposure supports unified representations and smoother gradients.
- B.2 Text data: The pre-training text corpus covers Web Text, Code, Mathematics, and Knowledge, with correctness and quality validation plus targeted data experiments.Most processing pipelines follow Kimi K2 [53].
- B.2 Text data: Code data emphasizes repository-level reasoning, real-world development patterns, and code-related documents to strengthen complex and agentic coding capabilities.The expanded sources include cross-file repositories, issues, reviews, commit histories, and retrieved code documents.
- B.3 Vision data: The multimodal corpus spans caption, interleaving, OCR, knowledge, perception, video, and agent data for modality alignment, long-context comprehension, and visual understanding.Caption data [49] [19] supports alignment with limits on synthetic captions, while interleaving data [81] [32] supports multi-image comprehension and longer contexts.
- B.3 Vision data: A specialized multimodal problem-solving corpus targets STEM reasoning by converting retrieved and crawled informational material into structured academic problems from K-12 through university levels.In-context learning [11] reformulates content lacking explicit query formats.
- B.3 Vision data: Agentic and temporal understanding data combines GUI screenshots, action trajectories, human demonstrations, diverse-source video, and fine-grained grounding annotations.The grounding data includes bounding boxes, point-based references, and contour-level annotations.
- C Infra: Training uses H800 GPU clusters with 8×400 Gbps RoCE interconnects, combining 16-way pipeline parallelism, 16-way expert parallelism, and ZeRO-1 data parallelism.The strategy supports node counts that are multiples of 32 and overlaps expert all-to-all communication with computation under interleaved 1F1B scheduling.
C.1 Data Storage and Loading · D Unified Agentic Reinforcement Learning Environment · E Evaluation Settings
The paper supports multimodal training with native-format object storage and a flexible, deterministic, scalable data pipeline. Its unified Agentic RL environment standardizes diverse tasks, composes specialized components, scales asynchronous rollouts, and supports evaluation across white-box and black-box settings.
- C.1 Data Storage and Loading: The data infrastructure retains visual data in native format on S3-compatible storage and supports dynamic shuffling, blending, tokenization, loss masking, and sequence packing.These capabilities enable adjustable data ratios as training requirements evolve.
- C.1 Data Storage and Loading: Stochastic augmentation covers visual and textual modalities while preserving 2D spatial coordinates and orientation metadata during geometric transformations.
- C.1 Data Storage and Loading: Deterministic seed and worker-state management makes interrupted training resume with the same data sequence as an uninterrupted run.
- C.1 Data Storage and Loading: Tiered caching scales data-loading throughput across distributed clusters while regulating object-storage request frequency, alongside platform-wide dataset governance and cross-cloud synchronization.The platform also supports registration, visualization, statistical analysis, and lifecycle governance.
- D Unified Agentic Reinforcement Learning Environment: The Agentic RL framework exposes a standardized Gym-like interface [10] with pluggable toolset, judge, prompt-diversification, and instruction-following modules.These components dynamically compose with core agent loops to support flexible environment customization and model generalization.
- D Unified Agentic Reinforcement Learning Environment: Up to 100,000 concurrent agent tasks run as asynchronous coroutines that can recursively trigger sub-task rollouts, enabling Parallel-Agent RL and Agent-as-Judge [31].The Rollout Manager provides fine-grained control, including partial rollout.
- D Unified Agentic Reinforcement Learning Environment: Token-in-Token-out inference with recorded log probabilities supports train-inference mismatch correction, while LLM Gateway records black-box rollout requests and responses under a custom protocol.Monitoring, profiling, visualization, and verification tools support efficiency and correctness in highly parallel asynchronous execution.
- E Evaluation Settings: Evaluation Settings specify configuration details and testing protocols for all benchmarks reported in Table 4.
E.1 General Evaluation Protocol
Unless explicitly stated otherwise, Kimi-K2.5 experiments use a standardized configuration with a 256k-token context length.
- E.1 General Evaluation Protocol: Kimi-K2.5 experiments generally use a 256k-token context length unless explicitly stated otherwise.This is the specified context-length setting in the general evaluation protocol.
E.2 Baselines
Baselines were evaluated using their respective high-performance reasoning configurations, with modality-specific settings for DeepSeek-V3.2 and Qwen3-VL-235B-A22B. GPT-5.2 vision scores may be conservative, and unstable API access prevented some costly benchmarks.
- Baselines used high-performance reasoning settings: extended thinking for Claude Opus 4.5, maximum reasoning effort (xhigh) for GPT-5.2, and high thinking for Gemini 3 Pro.
- DeepSeek-V3.2 used thinking mode for text-only benchmarks, while Qwen3-VL-235B-A22B used thinking mode for vision benchmarks.
- Approximately 10% of GPT-5.2-xhigh vision evaluations produced no output after three retries; these failures were counted as incorrect, making scores conservative lower bounds.The failures occurred during vision and multimodal benchmarks.
- Some high-cost benchmarks, including WideSearch, were skipped because stable access to the GPT-5.2 API was unavailable.
E.3 Text Benchmarks
Text-benchmark evaluation uses extended reasoning budgets and repeated sampling for high-complexity tasks, while LongBench v2 standardizes contexts and reports GPT5.2-high because GPT5.2-xhigh often violates the required output format.
- Reasoning Benchmarks: High-complexity reasoning benchmarks use a maximum completion budget of 96k tokens to support sufficient reasoning depth.The benchmarks include HLE-Full, AIME 2025, HMMT 2025, GPQA-Diamond, and IMO-AnswerBench.
- Reasoning Benchmarks: AIME 2025 and HMMT 2025 results are averaged over 64 runs, while GPQA-Diamond uses 8-run averaging to reduce stochastic-path variance.These aggregation settings are reported as Avg@64 and Avg@8, respectively.
- LongBench v2: LongBench v2 standardizes contexts to approximately 128k tokens using the truncation strategy from [9] and reports GPT5.2-high for format compliance.GPT5.2-xhigh frequently produces free-form question–answer responses instead of the required multiple-choice format.
E.4 Image and Video Benchmarks · E.5 Coding and Software Engineering
The image/video benchmarks use standardized sampling and protocol-specific evaluation settings, while coding benchmarks rely on constrained agent frameworks, non-thinking configurations, and repeated runs for stability.
- E.4 Image and Video Benchmarks: Image and video evaluations average results over three independent runs using Avg@3.
- E.4 Image and Video Benchmarks: ZeroBench multi-step reasoning uses constrained step-wise generation with a 24k maximum token budget per step.
- E.4 Image and Video Benchmarks: MMMU-Pro follows its official protocol by preserving input order and prepending images to text sequences.
- E.4 Image and Video Benchmarks: Video evaluations sample 128 uniform frames at 896 resolution for short videos and 2048 uniform frames at 448 resolution for long videos.Short-video benchmarks include VideoMMMU, MMVU, and MotionBench; long-video benchmarks include Video-MME, LongVideoBench, and LVBench.
- E.5 Coding and Software Engineering: Terminal Bench 2.0 uses the default Terminus-2 framework and JSON parser, evaluated in non-thinking mode because thinking-mode context handling is incompatible with its conversation state.
- E.5 Coding and Software Engineering: SWE-Bench variants use an internal framework with six repository-editing tools and tailored prompts, with peak performance achieved in non-thinking mode.The variants are Verified, Multilingual, and Pro.
- E.5 Coding and Software Engineering: Coding results are averaged over five independent runs using Avg@5 to improve stability across environment initialization and nondeterministic test ordering.CyberGym is reported at difficulty level 1, while PaperBench uses the CodeDev setting.
E.6 Agentic Evaluation … F Visualization
The evaluation sections define tool-enabled protocols for agentic, computer-use, and Agent Swarm tasks, including context management, sampling, step budgets, and one-shot settings. The visualization section illustrates long-form video analysis and tool-augmented visual reasoning through qualitative examples.
- E.6 Agentic Evaluation: Agentic evaluations equip Kimi-K2.5 with web search, code interpretation, and web browsing tools across HLE and multiple search benchmarks.The evaluated benchmarks include BrowseComp, WideSearch, DeepSearchQA, FinSearchComp T2&T3, and Seal-0.
- E.6 Agentic Evaluation: Context management is generally disabled, with overlong tasks counted as failures; HLE retains prior reasoning while truncating older tool results, whereas BrowseComp also tests full-history and discard-all settings.For Seal-0 and WideSearch, results are averaged over four independent runs, while other agentic benchmarks use single runs unless specified otherwise.
- E.7 Computer-Use Evaluation: Computer-use experiments use a one-shot protocol with max_steps_per_episode = 100, temperature = 0 for OSWorld-Verified, and temperature = 0.1 for WebArena.The agent context contains the last three history images, complete thought history, and task instruction; WebArena uses corrected evaluation scripts and GPT-4o as judge.
- E.8 Agent Swarm Configuration: Agent Swarm adds create_subagent and assign_task tools, enabling specialized agents to be defined and dispatched concurrently for parallel execution.The configuration explicitly encourages launching multiple agents concurrently when possible.
- E.8 Agent Swarm Configuration: Agent Swarm imposes benchmark-specific step budgets: BrowseComp gives the orchestrator 15 steps and each sub-agent 100, while WideSearch gives both 100 and In-house Bench gives the orchestrator 100 and each sub-agent 50.These limits apply to aggregate tool invocations and environment interactions.
- E.9 GDPVal: GDPVal-AA results cite Artificial Analysis and use official leaderboard metrics reported as of January 28, 2026.The supplied passage identifies the metric provenance and cutoff date but does not provide the table’s scores.
- F Visualization: Qualitative visualizations show Agent Swarm analyzing 24 hours of Black Myth: Wukong gameplay across 32 videos and Kimi K2.5 solving visual reasoning tasks with tool-augmented methods.The video analysis uses parallel agents for frame extraction, temporal event analysis, and key-moment identification; visual reasoning examples include maze solving, pie-chart analysis, and spot-the-difference detection.