Source-linked AI summary
InternAgent-1.5: A Unified Agentic Framework for Long-Horizon Autonomous Scientific Discovery
Shiyang Feng, Runmin Ma, Xiangchao Yan, Yue Fan, Yusong Hu, Songtao Huang, Shuaiyu Zhang, Zongsheng Cao, Tianshuo Peng, Jiakang Yuan, Zijie Guo, Zhijie Zhong, Shangheng Du, Weida Wang, Jinxin Shi, Yuhao Zhou, Xiaohan He, Zhiyin Yu, Fangchen Yu, Qihao Zheng, Jiamin Wu, Mianxin Liu, Chi Zhang, Shaowei Hou, Shuya Li, Yankai Jiang, Wenjie Lou, Lilong Wang, Zifu Wang, Jiong Wang, Wanghan Xu, Yue Deng, Dongrui Liu, Yiheng Wang, Wenlong Zhang, Fenghua Ling, Shufei Zhang, Xiaosong Wang, Shuangjia Zheng, Xun Huang, Siqi Sun, Shuyue Hu, Peng Ye, Chunfeng Song, Bin Wang, Conghui He, Yihao Liu, Xin Li, Qibin Hou, Tao Chen, Xiangyu Yue, Bin Wang, Liang He, Dahua Lin, Bowen Zhou, Bo Zhang, Lei Bai
TL;DR
Autonomous scientific systems need to support cross-disciplinary reasoning, heterogeneous computational and empirical work, and sustained refinement over long discovery cycles. InternAgent-1.5 addresses this with coordinated Generation, Verification, and Evolution subsystems backed by deep research, solution refinement, and long-horizon memory. Evaluations report leading scientific reasoning performance and strong results across algorithmic and empirical discovery workflows.
Problem
Existing AI4S systems have domain-specific architectures, uneven foundational capabilities, limited cross-search optimization, and restricted support for extended research cycles.
Method
InternAgent-1.5 unifies Generation, Verification, and Evolution with deep research, solution refinement, and long-horizon memory for computational and empirical discovery.
Results
InternAgent-1.5 achieves leading agentic reasoning performance and performs competitively on both algorithm discovery and empirical discovery tasks.
Takeaways & Limitations
The framework provides a general substrate for cross-disciplinary scientific workflows spanning hypothesis construction, methodological evaluation, evidence-driven refinement, and extended operation.
Takeaways & Limitations
Future work must strengthen coupling between computational reasoning and experimental validation and accelerate transition from hypotheses to verifiable results.
Abstract
from arXiv · showhide
We introduce InternAgent-1.5, a unified system designed for end-to-end scientific discovery across computational and empirical domains. The system is built on a structured architecture composed of three coordinated subsystems for generation, verification, and evolution. These subsystems are supported by foundational capabilities for deep research, solution optimization, and long horizon memory. The architecture allows InternAgent-1.5 to operate continuously across extended discovery cycles while maintaining coherent and improving behavior. It also enables the system to coordinate computational modeling and laboratory experimentation within a single unified system. We evaluate InternAgent-1.5 on scientific reasoning benchmarks such as GAIA, HLE, GPQA, and FrontierScience, and the system achieves leading performance that demonstrates strong foundational capabilities. Beyond these benchmarks, we further assess two categories of discovery tasks. In algorithm discovery tasks, InternAgent-1.5 autonomously designs competitive methods for core machine learning problems. In empirical discovery tasks, it executes complete computational or wet lab experiments and produces scientific findings in earth, life, biological, and physical domains. Overall, these results show that InternAgent-1.5 provides a general and scalable framework for autonomous scientific discovery.
1. Introduction
InternAgent-1.5 addresses limitations in autonomous scientific systems with a unified framework for computational and empirical discovery. It combines coordinated discovery stages, foundational capabilities, and long-horizon operation to support reasoning and iterative scientific workflows.
- Motivation: Existing AI4S systems are often domain-specific, uneven in foundational capabilities, limited in search integration, and unable to sustain extended research cycles.These limitations hinder unified cross-disciplinary reasoning, heterogeneous dry-lab and wet-lab coverage, broad search-history reuse, and long-term iterative operation.
- Framework: The framework treats Algorithm Discovery and Empirical Discovery as complementary domains requiring unified architecture, strong foundational capabilities, long-horizon optimization, and computational-experimental operation.Algorithm Discovery transforms objectives into formal-system solutions, whereas Empirical Discovery transforms observations into generalizations about the physical world.
- Framework: InternAgent-1.5 organizes end-to-end discovery into Generation, Verification, and Evolution subsystems supported by deep research, solution refinement, and long-horizon memory.The architecture covers hypothesis construction, methodological evaluation, and evidence-driven refinement across computational and empirical tasks.
- Evaluation: InternAgent-1.5 attains leading performance on agentic reasoning benchmarks and supports stable iterative refinement across extended discovery cycles.Its benchmark evaluation covers GAIA, GPQA, HLE-full, and FrontierScience, while long-horizon memory supports continued operation.
- Evaluation: Beyond benchmarks, InternAgent-1.5 performs competitively in algorithm discovery and empirical discovery, extending from benchmark reasoning to practical scientific workflows.The paper positions these tasks as evidence that the unified framework operates across computational and empirical settings.
2. InternAgent-1.5
InternAgent-1.5 unifies scientific discovery through coordinated Generation, Verification, and Evolution subsystems supported by deep research, solution refinement, and long-horizon memory. Its design combines cross-disciplinary reasoning, graph-augmented search, and structured memory for sustained, continuously improving discovery.
- System Architecture: The system automates hypothesis formulation, methodological evaluation, and evidence-driven refinement through an integrated, iterative workflow.Generation formulates hypotheses and plans, Verification evaluates them computationally or empirically, and Evolution updates knowledge, strategies, and memory from the evidence.
- Foundational Capabilities: Three foundational capabilities support the architecture: deep research for Generation, solution refinement for Verification, and long-horizon memory for Evolution.Together, these capabilities support literature-based construction, methodological evaluation, iterative refinement, and continuity across extended discovery cycles.
- System-Level Outcome: Across its capabilities, InternAgent-1.5 is designed to maintain continuity, adaptability, and scalability for reliable, continuously improving scientific discovery.The system overview describes sustained operation across extended discovery horizons through coordinated subsystems and persistent contextual information.
- Deep Research: The cross-disciplinary knowledge graph structures documents, concepts, methods, datasets, empirical settings, and problem statements with typed relations.It converts flat scientific text into a map where cross-field dependencies emerge as paths and supports retrieval from heterogeneous sources.
- Solution Refinement: Graph-Augmented Monte Carlo Search replaces rigid tree structures with a dynamic solution graph that aggregates information across exploration paths.The approach addresses isolated trajectories, unused search history, and limited composition of promising ideas during experimental design.
- Long-Horizon Memory: Structured Cognitive Memory maintains procedural, episodic, and semantic information across discovery cycles, enabling short-term refinement, mid-term adaptation, and long-term conceptual development.Retrieved strategic priors guide coherent reasoning and reduce redundant execution, while retrieved episodes help avoid unsuccessful directions and exploit successful ones.
3. Experiments
InternAgent-1.5 is evaluated through cross-disciplinary benchmarks, autonomous algorithm development, and scientific discovery tasks spanning computational and empirical domains. It achieves leading results on several agentic reasoning benchmarks and demonstrates applications across multiple scientific fields.
- Evaluation scope: The evaluation covers cross-disciplinary benchmarks, autonomous algorithm development, and scientific mechanism discovery across Earth, life, biological, and physical sciences.The benchmark suite includes SGI-Bench, GAIA, HLE, FrontierScience, and GPQA-diamond, while discovery experiments include scientific and AI algorithm tasks.
- Agentic reasoning benchmarks: 37.74% on SGI-Bench Deep Research surpasses Gemini-3-pro’s 18.48% by 19.26 percentage points, while Idea Generation reaches 58.11% versus GPT-5’s 55.40%.These are the best reported results on both SGI-Bench tracks.
- Agentic reasoning benchmarks: InternAgent-1.5 outperforms the cited GAIA baselines, including Mirothinker at 80.8% and Manus at 73.30%, and reaches 61.54% on Level 3 questions.The passage attributes this performance to an iterative workflow combining knowledge planning, collection, and refinement.
- Agentic reasoning benchmarks: 40.87% accuracy on HLE text-only and 40.00% on the full benchmark are reported as the best overall results among compared systems.The improvements are described as consistent across most HLE sub-domains.
- Agentic reasoning benchmarks: InternAgent-1.5 achieves the best overall FrontierScience results in both Olympiad at 77.20% and Research at 12.00%, with particularly strong Chemistry and Physics performance.It outperforms all cited baselines, including DeepSeek-V3.2-Thinking and Mirothinker-v1.5.
- Agentic reasoning benchmarks: 87.37% average accuracy on GPQA-diamond is reported as state-of-the-art, with particularly strong results in Chemistry and Physics.The benchmark evaluates expert-written biology, chemistry, and physics questions requiring deep scientific reasoning.
3.3. Results for Algorithm Discovery Tasks
InternAgent-1.5 consistently outperforms domain-specific baselines across scientific algorithm tasks and supports end-to-end computational and empirical discovery workflows. Results span chemistry, physics, biology, Earth science, target discovery, and laboratory-guided design.
- InternAgent-1.5 consistently achieves superior performance across six scientific algorithm task types compared with prior work and state-of-the-art domain-specific baselines.The evaluation covers algorithmic and empirical scientific workflows across multiple domains.
- Chemical and Molecular Analysis: 36.6 R2 on AutoRYP exceeds LoRA-finetuned LLaMA3 at 27.6 and Dolphin at 31.8, while Energy MAE reaches 0.114 versus VisNet at 0.158.These results cover Suzuki-Miyaura reaction analysis and Molecular Dynamics.
- Physics and Engineering Systems: 0.00318 RMSE on AutoPower improves over SenseFlow’s 0.00473, and AutoTSF reaches 0.423 MAE versus DLinear’s 0.4382.The results are reported on the IEEE 39-Bus and ETTh1 datasets, respectively.
- Biological and Genomic Prediction: 0.143 MSE on AutoTPPR outperforms GEARS at 0.197, while AutoEAP reaches 0.91 Pearson correlation versus DeepSTARR at 0.65.The tasks evaluate transcription prediction and enhancer activity prediction.
- Empirical Discovery Workflows: InternAgent-1.5 autonomously integrates multi-omics evidence, generates mechanistic hypotheses, and produces experiment-ready recommendations for target discovery.The workflow narrowed 125 candidate genes to GPR160 and separately developed mechanistic and validation guidance for ARG2.
- Empirical Discovery Workflows: The system coordinates dry-lab predictions with wet-lab procedures and refines molecular candidates using Synthetic Accessibility, Tanimoto Similarity, and LogP.Applied to a DprE1 inhibitor template, it proposed bioisosteres based on piperidinopyrimidine scaffolds and simulated hit-to-lead optimization.
3.5. Effectiveness of Structured Cognitive Memory
Structured Cognitive Memory supports sustained improvement by combining episodic, semantic, and procedural memory across different timescales of autonomous research. Ablation analyses associate these components with smoother adaptation, greater novelty and continuity, and more efficient planning.
- Structured Cognitive Memory is designed to enable continuous operation and sustained improvement across diverse scientific discovery tasks.The subsystem is evaluated through Task-Episodic Memory, Semantic-Knowledge Memory, and Strategy-Procedural Memory.
- Task-Episodic Memory: With TEM active, performance trajectories rise smoothly, whereas removing it produces irregular or stagnant progress and repeated ineffective strategies.Retrieved episodes provide evidence from earlier trials for short-horizon adaptation and hypothesis refinement.
- Semantic-Knowledge Memory: SKM preserves semantic continuity and measurable novelty while helping objectives avoid saturated conceptual regions after multiple exploration batches.Figure 13 illustrates iterative objective evolution from an initial seed.
- Strategy-Procedural Memory: The full system achieves higher success rates and more coherent plans than the SPM-ablated baseline, with fewer unnecessary branches and redundant tool calls.Without SPM, plans become longer and fragmented, with error propagation and imprecise tool-call parameters.
- TEM, SKM, and SPM provide complementary support for short-term adaptation, long-term knowledge accumulation, and efficient reasoning execution.Together, these roles support sustained improvement across discovery cycles.
4. Related Work
Related work shows a progression from end-to-end research automation and tool-using deep-research agents toward systems with increasingly explicit memory mechanisms. These lines of work motivate InternAgent-1.5’s focus on autonomous reasoning, dynamic tool use, and long-horizon adaptation.
- AI Scientist systems automate hypothesis generation and experiment design, while later search-based procedures broaden exploration of methodological space.AlphaEvolve represents another approach to scientific discovery through broader algorithmic search.
- Deep Research agents extend language models from retrieval-augmented generation to adaptive, tool-driven workflows involving planning, iterative retrieval, and external tools.WebGPT and Toolformer established early examples of web and API integration for tool-mediated reasoning.
- Agent-memory research spans token-level retention, parametric accumulation of experience, latent structured trajectories, and short-term interaction memory.These mechanisms target long-horizon reasoning, continual adaptation, and interaction with complex environments.
5. Conclusion
InternAgent-1.5 presents a unified architecture that integrates generation, verification, and evolution for end-to-end scientific discovery. Comprehensive evaluations report strong reasoning, algorithmic, computational, and empirical capabilities, while future work targets tighter computational–experimental coupling and faster validation.
- InternAgent-1.5 integrates generation, verification, and evolution with deep research, solution refinement, and long-horizon memory for cross-disciplinary scientific workflows.The architecture is intended to maintain consistent information flow across discovery stages.
- Comprehensive evaluations report strong structured scientific reasoning, competitive algorithmic solutions, extended experimental optimization, and multi-step computational and empirical workflows.Across algorithmic and empirical domains, outputs align with established scientific principles and reproduce findings observed in real studies.
- Future work will strengthen computational reasoning–experimental validation coupling and accelerate the transition from generated hypotheses to verifiable results.These directions are presented as ways to improve the efficiency and reliability of cross-disciplinary discovery.
Lead Authors
The listed lead author is Shiyang Feng.
- Shiyang Feng is listed as a lead author.
- Runmin Ma is listed among the lead authors.
- Xiangchao Yan is listed among the lead authors.
Core Authors
The core-author list includes Yue Fan, Yusong Hu, and Songtao Huang, among others.
- Yue Fan is included in the core-author list.
- Yusong Hu is included in the core-author list.
- Songtao Huang is included in the core-author list.
- The list also includes Shuaiyu Zhang, Zongsheng Cao, and Tianshuo Peng.
Scientific Directors
The scientific-director listing includes Wenlong Zhang, Fenghua Ling, Shufei Zhang, Xiaosong Wang, and Shuangjia Zheng, followed by additional contributors.
- Wenlong Zhang and Fenghua Ling appear in the scientific-director listing.
- Shufei Zhang and Xiaosong Wang also appear in the listing.
- Shuangjia Zheng and Xun Huang are listed among the contributors.
- Siqi Sun, Shuyue Hu, and Peng Ye are included in the extended author listing.
- The extended listing also includes Chunfeng Song, Bin Wang, Conghui He, and Yihao Liu.
Main Affiliations
The authors are affiliated with six academic and research institutions, including Shanghai Artificial Intelligence Laboratory and Fudan University.
- Shanghai Artificial Intelligence Laboratory is listed as affiliation 1.
- Fudan University and Lingang Laboratory are listed as affiliations 2 and 3.
- The Chinese University of Hong Kong, East China Normal University, and Nankai University are listed as affiliations 4, 5, and 6.
A.2. Earth Science example
The Earth Science example reviews the AMOC, its global significance, and conflicting evidence about weakening versus abrupt collapse. It synthesizes observational, paleoclimate, and modeling evidence while emphasizing substantial methodological and modeling uncertainties.
- Definition and significance: The AMOC is a planetary-scale conveyor transporting heat northward through warm surface waters and returning cold, dense deep water southward.It is also described as important for global climate and biogeochemical systems.
- Observed change: 15–20% weakening since the mid-20th century is reported, with some analyses placing the AMOC in its weakest state in over a century.The evidence is presented as an observed or reconstructed decline rather than a settled collapse forecast.
- Mainstream projections: Climate models consistently project gradual 21st-century weakening, while abrupt collapse before 2100 remains assessed as very unlikely but not impossible.The IPCC consensus is described as robust across models and emissions scenarios, while confidence in the collapse assessment has changed from high to medium.