Source-linked AI summary

Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning

ByteDance Seed, :, Jiaze Chen, Tiantian Fan, Xin Liu, Lingjun Liu, Zhiqi Lin, Mingxuan Wang, Chengyi Wang, Xiangpeng Wei, Wenyuan Xu, Yufeng Yuan, Yu Yue, Lin Yan, Qiying Yu, Xiaochen Zuo, Chi Zhang, Ruofei Zhu, Zhecheng An, Zhihao Bai, Yu Bao, Xingyan Bin, Jiangjie Chen, Feng Chen, Hongmin Chen, Riwei Chen, Liangqiang Chen, Zixin Chen, Jinsong Chen, Siyan Chen, Kaiyuan Chen, Zhi Chen, Jin Chen, Jiecao Chen, Jinxin Chi, Weinan Dai, Ning Dai, Jiahui Dai, Shihan Dou, Yantao Du, Zhengyin Du, Jianhui Duan, Chen Dun, Ting-Han Fan, Jiazhan Feng, Junda Feng, Ziyuan Feng, Yuwei Fu, Wenqi Fu, Hanjie Fu, Hao Ge, Hongyi Guo, Mingji Han, Li Han, Wenhao Hao, Xintong Hao, Qianyu He, Jerry He, Feng He, Wen Heng, Zehua Hong, Qi Hou, Liang Hu, Shengding Hu, Nan Hu, Kai Hua, Qi Huang, Ziyue Huang, Hongzhi Huang, Zihao Huang, Ting Huang, Wenhao Huang, Wei Jia, Bin Jia, Xiaoying Jia, Yuhua Jiang, Haobin Jiang, Ziheng Jiang, Kaihua Jiang, Chengquan Jiang, Jianpeng Jiao, Xiaoran Jin, Xing Jin, Xunhao Lai, Zheng Li, Xiang Li, Liyi Li, Hongkai Li, Zheng Li, Shengxian Wan, Ya Wang, Yunshui Li, Chenggang Li, Niuniu Li, Siyu Li, Xi Li, Xiao Li, Aoyan Li, Yuntao Li, Nianning Liang, Xinnian Liang, Haibin Lin, Weijian Lin, Ye Lin, Zhicheng Liu, Guanlin Liu, Guanlin Liu, Chenxiao Liu, Yan Liu, Gaohong Liu, Juncai Liu, Chundian Liu, Deyi Liu, Kaibo Liu, Siyao Liu, Qi Liu, Yongfei Liu, Kang Liu, Gan Liu, Boyi Liu, Rui Long, Weiqiang Lou, Chenwei Lou, Xiang Luo, Yao Luo, Caiping Lv, Heyang Lv, Bole Ma, Qianli Ma, Hongzhi Ma, Yiyuan Ma, Jin Ma, Wenchang Ma, Tingting Ma, Chen Mao, Qiyang Min, Zhe Nan, Guanghan Ning, Jinxiang Ou, Haojie Pan, Renming Pang, Yanghua Peng, Tao Peng, Lihua Qian, Lihua Qian, Mu Qiao, Meng Qu, Cheng Ren, Hongbin Ren, Yong Shan, Wei Shen, Ke Shen, Kai Shen, Guangming Sheng, Jinlong Shi, Wenlei Shi, Guang Shi, Shuai Shuai Cao, Yuxin Song, Zuquan Song, Jing Su, Yifan Sun, Tao Sun, Zewei Sun, Borui Wan, Zihan Wang, Xiaohui Wang, Xi Wang, Shuguang Wang, Jun Wang, Qinlong Wang, Chenyuan Wang, Shuai Wang, Zihan Wang, Changbao Wang, Jiaqiang Wang, Shihang Wang, Xuwu Wang, Zaiyuan Wang, Yuxuan Wang, Wenqi Wang, Taiqing Wang, Chengzhi Wei, Houmin Wei, Ziyun Wei, Shufa Wei, Zheng Wu, Yonghui Wu, Yangjun Wu, Bohong Wu, Shuang Wu, Jingqiao Wu, Ning Wu, Shuangzhi Wu, Jianmin Wu, Chenguang Xi, Fan Xia, Yuqiao Xian, Liang Xiang, Boren Xiang, Bowen Xiao, Zhen Xiao, Xia Xiao, Yongsheng Xiao, Chao Xin, Shulin Xin, Yuwen Xiong, Jingjing Xu, Ziwen Xu, Chenyin Xu, Jiayi Xu, Yifan Xu, Wei Xu, Yufei Xu, Shikun Xu, Shipeng Yan, Shen Yan, Qingping Yang, Xi Yang, Tianhao Yang, Yuehang Yang, Yuan Yang, Ximing Yang, Zeyu Yang, Guang Yang, Yifan Yang, Xuesong Yao, Bairen Yi, Fan Yin, Jianian Yin, Ziqiang Ying, Xiangyu Yu, Hongli Yu, Song Yu, Menghan Yu, Huan Yu, Siyu Yuan, Jun Yuan, Yutao Zeng, Tianyang Zhan, Zheng Zhang, Yun Zhang, Mofan Zhang, Wang Zhang, Ru Zhang, Zhi Zhang, Tianqi Zhang, Xinyi Zhang, Zhexi Zhang, Sijun Zhang, Wenqiang Zhang, Xiangxiang Zhang, Yongtao Zhang, Yuyu Zhang, Ge Zhang, He Zhang, Yue Zhang, Renjie Zheng, Ningxin Zheng, Zhuolin Zheng, Yaowei Zheng, Chen Zheng, Xiaoyun Zhi, Wanjun Zhong, Cheng Zhong, Zheng Zhong, Baoquan Zhong, Xun Zhou, Na Zhou, Huan Zhou, Hang Zhu, Defa Zhu, Wenjia Zhu, Lei Zuo

arXiv:2504.13914v3cs.CL

TL;DR

Seed1.5-Thinking addresses the challenge of building reliable, broadly capable reasoning models by combining reinforcement-learning techniques for data, stability, reward modeling, and infrastructure. It achieves strong results across mathematics, coding, science, and non-reasoning tasks, including 86.7 on AIME 2024, 55.0 on Codeforces, 77.3 on GPQA, and an 8.0% rise in positive user feedback over DeepSeek R1.

  • Problem

    Reinforcement-learning reasoning models require stable training and evaluation methods that support both reasoning and diverse non-reasoning tasks.

  • Method

    Seed1.5-Thinking combines chain-of-thought and mixed RL data with VAPO/DAPO stability methods, scalable asynchronous infrastructure, and specialized reward modeling.

  • Results

    86.7 on AIME24, 55.0 on Codeforces, 77.3 on GPQA, and an 8.0% rise in positive user feedback over DeepSeek R1 demonstrate strong performance across reasoning and non-reasoning tasks.

  • Takeaways & Limitations

    The results support generalized reasoning capabilities that extend beyond benchmark reasoning tasks to diverse user scenarios.

Abstract

from arXiv · show

We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 86.7 on AIME 2024, 55.0 on Codeforces and 77.3 on GPQA, demonstrating excellent reasoning abilities in STEM and coding. Beyond reasoning tasks, the method demonstrates notable generalization across diverse domains. For instance, it surpasses DeepSeek R1 by 8% in win rate on non-reasoning tasks, indicating its broader applicability. Compared to other state-of-the-art reasoning models, Seed1.5-Thinking is a Mixture-of-Experts (MoE) model with a relatively small size, featuring 20B activated and 200B total parameters. As part of our effort to assess generalized reasoning, we develop two internal benchmarks, BeyondAIME and Codeforces, both of which will be publicly released to support future research. Model trial link: https://www.volcengine.com/experience/ark.

1 Introduction

Seed1.5-Thinking combines reinforcement-learning advances in data, algorithms, and infrastructure to achieve strong reasoning and non-reasoning performance. It reaches competitive results in mathematics and science while improving human preference feedback on diverse user scenarios.

  • Seed1.5-Thinking achieves strong performance across both reasoning and non-reasoning tasks.
  • Mathematical Reasoning: 86.7 on AIME 2024 matches o3-mini-high and surpasses o1 and DeepSeek R1, while remaining behind o3 and Gemini Pro 2.5.BeyondAIME was constructed to provide a more discriminative mathematics evaluation than AIME 2024.
  • Science: 77.3 on GPQA is close to o3-level performance, with the gain attributed largely to mathematical-training generalization rather than more science-specific data.
  • Non-reasoning Tasks: 8.0% overall rise in users’ positive feedback over DeepSeek R1 demonstrates improved handling of intricate scenarios in human-evaluated non-reasoning tasks.
  • The training effort targets chain-of-thought data, reinforcement-learning stability, and scalable infrastructure for reasoning models.The infrastructure uses asynchronous streaming rollouts, prioritized sample pools, mixed precision, and automatic fault recovery; iteration cycles are 3× faster than synchronous frameworks.

2 Data

The RL training data is divided into verifiable and non-verifiable problems. Verifiable STEM, coding, and logic tasks provide definitive correctness signals, while non-verifiable tasks require preference-based quality assessment.

  • RL training data consists of verifiable problems with definitive answers and non-verifiable problems without definitive answers.
  • The model’s reasoning ability primarily comes from verifiable problems and can generalize to non-verifiable problems.
  • Verifiable data primarily covers STEM questions, coding tasks with unit tests, and logic reasoning tasks amenable to automated verification.

STEM Data

The paper constructs and filters diverse training and evaluation data for STEM, coding, logic, and non-reasoning tasks, while introducing BeyondAIME to address limitations of AIME-based evaluation.

  • STEM Data: The STEM training set contains 100k problems after cleaning and augmentation, with correctness evaluated by the model-based Seed-Verifier.
  • Coding Data: Coding data emphasizes challenging competitive-programming tasks with specifications, unit tests, checker scripts, and difficulty filtering.
  • Coding Data: An offline code sandbox executes and assesses generated programs during RL, with offline results strongly correlated with official verdicts.
  • Logic Data: Logic data covers 22 tasks with configurable generators and verifiers, enabling difficulty adjustment during training.
  • Non-verifiable Data: Non-verifiable data spans creative writing, translation, knowledge QA, and role-playing, using reward-based filtering to remove low-diversity or overly simple prompts.
  • BeyondAIME: BeyondAIME contains 100 expert-curated integer-answer problems at least as difficult as AIME’s hardest questions, addressing AIME’s small annual sample size.

3 Reward Modeling

The paper develops distinct reward-modeling approaches for verifiable and non-verifiable tasks, including reasoning-based verification to improve judgment robustness. These methods aim to provide reliable reward signals for RL across varied response types.

  • Reward modeling defines the RL objective, so precise and reliable response rewards are essential; verifiable and non-verifiable problems use distinct methods.
  • Reasoning-based verification provides generalized judgments beyond rule-based reward systems, though its thinking process consumes significant GPU resources.
  • Verifiable Problems: Seed-Verifier evaluates whether a model answer is mathematically equivalent to a reference answer using human-written principles and LLM capabilities.
  • Verifiable Problems: Seed-Thinking-Verifier generates a detailed reasoning path to compare reference and model answers and produce more nuanced judgments.
  • Verifiable Problems: Seed-Thinking-Verifier significantly alleviates reward hacking and inconsistent predictions, while handling corner cases more accurately than Seed-Verifier.
  • Non-verifiable Problems: For non-verifiable problems, pairwise generative rewards compare two responses using YES/NO probabilities, improving RL stability in mixed verifiable and non-verifiable training.

4 Approach

Seed1.5-Thinking’s approach combines curated supervised and reinforcement-learning data with methods addressing exploration, reward, value estimation, and domain interference. The framework integrates verifiable and preference-scored data while targeting stable long-CoT optimization.

  • Data: 400k SFT instances include 300k verifiable and 100k non-verifiable problems to initialize reinforcement learning.The SFT stage produces more readable outputs, fewer hallucinations, and reduced harmfulness than starting RL from a base model.
  • Data: Long-CoT responses are created through iterative model synthesis, human annotation, and rejection sampling.The workflow first accumulates high-quality cold-start samples before training a reasoning model with long chain-of-thought.
  • Data: The unified RL framework combines verifiable, reward-model-scored, and hybrid data from multiple domains.Verifiable data uses direct correctness checks, general data uses human-preference reward models, and hybrid data combines both signals.
  • Reinforcement Learning: Length-adaptive GAE sets λ_policy = 1 − 1/(αl) to distribute temporal-difference errors more uniformly across sequence lengths.The framework also decouples value and policy GAE parameters to support efficient and stable training.
  • Reinforcement Learning: Online Data Distribution Adaptation addresses interference caused by differing difficulty levels, reward hacking, and other factors when merging domains.It transforms the stationary prompt distribution into an adaptive distribution during reinforcement learning.

5 Infrastructures

The infrastructure combines hybrid-engine execution, streaming rollouts, distributed parallelism, workload balancing, memory optimization, and automatic configuration. These components target scalability, efficiency, and reduced idle time during large-scale RL training.

  • Training System: The hybrid engine co-locates models to prevent GPU idle time when switching between training and generation.This architecture is used for Seed1.5-Thinking’s training workload.
  • Training System: SRS streaming rollout decouples model evolution from runtime execution and adjusts on-policy versus off-policy sampling through α.α denotes the proportion of samples generated on-policy by the latest model version, while 1 − α uses versioned snapshots.
  • Training System: The system combines FP8 policy networks, TP, EP, and SP to manage precision, expert assignment, and MoE token imbalance.A kernel auto-tuner selects CUDA configurations using real-time load monitoring.
  • Distributed Training: TP, EP, CP, and FSDP are composed across attention and MoE layers for distributed training.Attention layers use TP and CP, whereas MoE layers use EP.
  • Distributed Training: KARP rearranges sequences within mini-batches to balance effective sequence lengths across micro-batches and data-parallel ranks.This addresses workload imbalance and low training efficiency caused by differing sequence lengths.
  • Distributed Training: Layer-wise recomputation, activation and optimizer offload, AutoTuner, and ByteCheckpoint support memory efficiency, configuration selection, and elastic resumption.ByteCheckpoint resumes training from different distributed configurations with minimal overhead.

6 Experiment Results

Seed1.5-Thinking shows strong results across mathematics, science, coding, and human-evaluated non-reasoning tasks, while performance varies by benchmark. Ablations also examine initialization effects and algorithm-ranking consistency across model scales.

  • Auto Evaluation Results: Table 2 reports mathematics, coding, science, and general-knowledge results using task-specific response averaging protocols.Mathematical results average 32 responses, GPQA averages 8, Codeforces reports avg@8 and pass@8, and other tasks average 1 response.
  • Auto Evaluation Results: 86.7 on AIME 2024 matches OpenAI’s o3-mini-high, while performance on AIME 2025 and BeyondAIME remains below o3-level results.Seed1.5-Thinking scores 77.3% on GPQA and nearly matches Gemini 2.5 Pro on Codeforces, but trails o3-mini-high on coding.
  • Human Evaluation Results: 8.0% overall win ratio is reported for Seed1.5-Thinking in human evaluations against DeepSeek R1 across diverse non-reasoning scenarios.The evaluations cover dimensions including coherence, relevance, creativity, and adherence to human preferences.
  • Effects of pre-train models: RFT initialization saturates faster during training but ultimately achieves lower performance than training without RFT.The comparison is presented as an ablation on pretrained-model initialization.
  • Effects of pre-train models: RL algorithm rankings remain consistent across model sizes and architectures, suggesting Qwen-32B can serve as a proxy for algorithm investigation.The comparison includes Qwen-32B and Seed-150B-MoE, which differ in both architecture and size.

7 Related Work

The paper situates Seed1.5-Thinking within the shift toward test-time scaling and large-scale reinforcement learning for complex reasoning. It presents the model and organizes its contributions around data, RL algorithms, and infrastructure.

  • Related Work: Test-time scaling methods use extended chain-of-thought and reinforcement learning to improve mathematical and coding performance.The related-work discussion names OpenAI o1 and DeepSeek R1 as examples of this paradigm.
  • Related Work: Seed1.5-Thinking is presented as a state-of-the-art-level model developed through advances in data, RL algorithms, and RL infrastructure.These three aspects structure the paper’s account of how its performance is achieved.

8 Conclusion

Seed1.5-Thinking is presented as a reasoning model with strong performance across reasoning and non-reasoning tasks. The conclusion highlights advanced reinforcement learning and reports results on AIME24, AIME25, and Codeforces.

  • Seed1.5-Thinking achieves strong performance across both reasoning and non-reasoning tasks.
  • 86.7% on AIME24, 74.0% on AIME25 and 55.0% on Codeforces are reported for Seed1.5-Thinking.
  • The model uses advanced reinforcement learning techniques to improve thinking ability stably and reliably.

9 Contributions and Acknowledgments

This section lists team members and presents a verifier case study comparing Seed-Verifier with Seed-Thinking-Verifier. The latter handles complex answers through step-by-step analysis and generalizes across domains.

  • Contributions and Acknowledgments: The listed names are sorted alphabetically by last name, and an asterisk marks members who have departed.
  • Contributions and Acknowledgments: Table 5 compares Seed-Verifier and Seed-Thinking-Verifier in a case study.
  • Contributions and Acknowledgments: Seed-Thinking-Verifier provides accurate judgments on complex answers through step-by-step analysis and can generalize to almost any domain.

B Case Study on Creative Writing

The creative-writing case studies show examples in Chinese and English, separating each example into the user prompt, chain of thought, and final response.

  • Case Studies: Tables 6, 7, and 8 present Chinese and English examples demonstrating the model’s creative-writing proficiency.
  • Case Studies: Each creative-writing example contains the original user prompt, the model’s chain of thought, and the final response.
  • Case Studies: Table 6 is labeled as Case 1 on Creative Writing.
  • Case Studies: Table 7 is labeled as Case 2 on Creative Writing.
  • Case Studies: Table 8 is labeled as Case 3 on Creative Writing.
Loading 2504.13914v3…