Source-linked AI summary
Atria Dawn: The Dawn of Agentic Superintelligence
Honglin Guo, Tao Gui, Yicheng Chen, Guanting Dong, Qiming Ge, Yuyang Hu, Zixian Huang, Jiajie Jin, Alexander Lam, Yining Li, Jiahang Lin, Yanjiang Liu, Xinyu Lu, Haijun Lv, Junlin Shang, Qisheng Su, Guoqiang Wang, Rui Wang, Zhecan Wang, Hao Xiang, Xinchen Xie, Shuhao Xing, Xiaoyu Xing, Wanghan Xu, Xinyu Yang, Yajie Yang, Chengfeng Zhao, Haoran Zhao, Ruojun Zhou, Yunhua Zhou, Yicheng Zou, Kun Cai, Qiye Cai, Xinmeng Che, Haodong Chen, Jiabei Chen, Jiahao Chen, Jiayi Chen, Yujia Chen, Lizhi Cui, Youheng Dai, Xin Deng, Yi Dong, Shihan Dou, Chenya Gu, Xu Guo, Ding Han, Feiyang Hao, Haotan He, Jie Hou, Binze Hu, Zijian Hu, Junhao Huang, Huicheng Jiang, Jiazhen Jiang, Shufan Jiang, Jiahao Kuang, Bowen Lai, Bo Li, Jiaqiang Li, Peng Li, Qilong Li, Zhuoqun Li, Jiaxiang Liu, Shuainan Liu, Tong Liu, Yi Liu, Zhonghang Lu, Jianwen Luo, Yanyi Luo, Huijie Lv, Ningsheng Ma, Zerun Ma, Houcheng Min, Chengjun Pan, Qiyuan Peng, Xiaoxuan Peng, Jianmin Qian, Jiantao Qiu, Wanying Ren, Huayu Sha, Jifei Shan, Zixin Shang, Bing Shao, Zhuohui Sheng, Jiayang Shi, Yang Shu, Aierpanjiang Simayi, Sirui Song, Yuxiao Song, Zhe Sun, Zhichao Sun, Wenzhe Tan, Wenhui Tian, Zhongbo Tian, Hanchen Wang, Pengbo Wang, Rui Wang, Yiding Wang, Yuhui Wang, Zhiheng Xi, Caijun Xu, Chao Xu, Yongfeng Xu, Xiaolei Yang, Zhixiong Yang, Qian Yao, Shihong Yi, Yuankai Ying, Jia Yu, Dingbo Yuan, Hao Yuan, Junjie Yuan, Bo Zhang, Caixian Zhang, Qiuyinzhe Zhang, Jiyuan Zhao, Penghao Zhao, Ying Zhao, Pujun Zheng, Xiaoxue Zhong, Xiaohao Zhou, Xinyu Zhou, Dongsheng Zhu, Guanru Zhu, Yulun Zhu, Yaojie Lu, Tao Ji, Hongyu Lin, Yutao Zhu, Pengfei Cao, Guoxiu He, Xianpei Han, Ben He, Zhicheng Dou, Kang Liu, Qi Zhang, Le Sun, Jun Zhao, Ji-Rong Wen, Xuanjing Huang, Yu-Gang Jiang, Bowen Zhou
TL;DR
Atria Dawn Preview addresses how agentic systems can support scientific and engineering work while participating in the development of successor AI systems. It uses a Verifiable Experience Pipeline and broad benchmark evaluation, achieving leading results across several tasks. Its development record shows agents proposing and executing work while humans retain final judgment over research direction, with oversight remaining necessary as autonomy advances.
Problem
As agents participate in developing AI systems, task performance alone does not determine who identifies worthwhile problems, selects methods, interprets uncertain results, or directs research.
Method
Atria Dawn Preview is trained through a Verifiable Experience Pipeline linking tool-mediated trajectories to executable environments, artifacts, and externally verified outcomes, then evaluated across 16 benchmarks.
Results
Atria Dawn achieves the highest reported score on five benchmarks and the second-highest on three others, while agents frequently propose approaches and execute changes during development.
Takeaways & Limitations
The development record indicates a shift toward project-level partnership, with human effort concentrating on evaluation, selection, and steering research direction.
Takeaways & Limitations
Successful task execution does not ensure that agents can anticipate which tasks or training strategies will produce meaningful capability gains, so research judgment remains necessary.
Abstract
from arXiv · showhide
As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-and-development process behind this model as a case study of human--AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback. These observations indicate a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. Progress toward more autonomous AI research must therefore advance both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development.
1. Introduction
Atria Dawn Preview is introduced as an agentic model for scientific and engineering workflows, alongside an examination of how humans and agents divide responsibilities in AI R&D. The study finds that agents often initiate and execute work, while humans evaluate, select, and steer research.
- Atria Dawn Preview is a foundation agentic language model designed for scientific research and engineering workflows, trained through a Verifiable Experience Pipeline and evaluated across 16 benchmarks.The pipeline connects tasks and agent trajectories to artifacts and externally checked outcomes.
- The paper analyzes 769 task records from 56 participants alongside agent logs to study responsibility distribution during Atria Dawn’s development.
- Agents frequently initiate approaches and execute changes, while humans concentrate on evaluation, selection, and steering the direction of inquiry.
- More autonomous AI research requires not only task competence but also identifying worthwhile directions, designing informative experiments, learning from uncertain outcomes, and deciding when to redirect effort.
2. Atria Dawn Preview
Atria Dawn Preview combines a verified tool-mediated training pipeline with broad evaluations and concrete demonstrations across research, software, CAD, and professional deliverables. It achieves leading results across multiple benchmarks, while the case studies show that artifact production and execution checks are distinct from broader validation.
- Atria Dawn Preview: The Verifiable Experience Pipeline links tool-mediated trajectories to executable environments, artifacts, feedback, and externally checked outcomes, while curation removes incomplete or invalid examples.Failure analysis guides later task construction and environment refinement.
- Evaluation Summary: Atria Dawn achieves the highest reported score on five benchmarks and the second-highest on three others across 16 evaluations spanning research, digital work, software engineering, and related agentic tasks.Rankings use available entries in each benchmark row, with missing results not treated as zero.
- Evaluation Summary: Atria Dawn leads four general-agentic benchmarks—AutomationBench (53.8), BFCL v4 (77.0), DeepSearchQA (96.0), and BrowseComp (92.5)—and ranks second on three workspace-oriented benchmarks.Its margins over the runner-up are 4.1 points on AutomationBench and 2.9 points on BFCL v4.
- Evaluation Summary: On remaining general-agentic tasks, Atria Dawn stays competitive, while GDPval (1583) and JobBench (50.3) retain the clearest headroom for transferring general ability to long-horizon professional work.
- Evaluation Summary: In coding, Atria Dawn reaches the highest reported CyberGym score (86.5), leads MLE-bench Lite with HumanRank 86.2, and remains competitive on SWE-bench Pro and Terminal-Bench 2.1.The results indicate particular strength in vulnerability analysis and security tasks.
- Selected Capabilities and Case Studies: Selected cases demonstrate weather-model training, iterative CUDA optimization, sourced literature analysis, MiniOS persistence checks, and professional report generation across concrete tool environments.These demonstrations are selected cases rather than estimates of average task success, and their checks differ by artifact and workflow.
3. The Evolution of Human–AI Collaboration
AI agents are moving from task execution toward research partnership, while humans retain higher-level judgment over goals, exploration, and intervention. This emerging division of labor expands agent responsibility but leaves a gap between project-level autonomy and self-sustaining recursive improvement.
- 3.1. Evolving Human–AI Roles: AI agents increasingly design workflows and revise experimental plans, shifting from task runners toward research partners within human-defined goals.Human involvement correspondingly shifts toward higher-level judgment and targeted intervention.
- 3.2. Autonomy Still Needs Human Judgment: Human researchers remain important because they guide exploration and make targeted interventions as agents assume more planning and execution responsibility.Their contributions include pruning unproductive research directions before substantial resources are committed.
- 3.1. Evolving Human–AI Roles: The proposed progression runs from models as research objects, through task-level runners, to agents that plan and iterate as research partners.A speculative next stage raises questions about recursive self-improvement and the future role of humans.
- 3.2. Autonomy Still Needs Human Judgment: Agents can execute tasks and formulate research plans, but these abilities do not ensure they can identify worthwhile tasks, training strategies, or next research directions.Meaningful recursive self-improvement requires connecting capability gains with reliable discovery, experimentation, and interpretation of uncertain evidence.
4. Human–AI Collaboration in Atria Dawn R&D
In Atria Dawn’s development, AI handled substantial execution, proposal generation, and revision, while humans retained responsibility for feasibility judgments, final choices, and feedback. AI-assisted work frequently expanded what participants considered feasible, but rising agent activity did not imply reduced human oversight.
- 4. Human–AI Collaboration in Atria Dawn R&D: The daily median ratio of agent actions to human prompts rose from 11.0 to 28.5, but the increase was uneven across participants.The ratio measures logged agent actions divided by human prompts over the preceding seven days, with 21–22 valid participants per day.
- 4.1. AI Expands the Range of Feasible Tasks: 33.2% of 455 completed AI-assisted tasks were reported as infeasible without AI under fixed scope, quality, and resource constraints.The 151 tasks came from 27 of 56 participants, indicating the result was not confined to a few heavy users.
- 4.2. Role Distribution in Method Proposal and Selection: Human final selection covered 85.5% of method or parameter decisions, while “AI proposes, human selects” was the most common pattern at 55.4%.Humans also made the final choice for 93.4% of goals or scope and 81.9% of acceptance criteria.
- 4.2. Role Distribution in Method Proposal and Selection: AI proposal share varied from 16.9% for goals or scope to 55.4% for methods or parameters, while AI final-choice share remained between 6.1% and 9.2%.Decision type changed who generated options more than who selected among them.
- 4.2. Role Distribution in Method Proposal and Selection: For the 151 tasks rated infeasible without AI, humans selected the final goal in 144 cases, or 95.4%.This stronger dependence on AI for task feasibility therefore did not correspond to greater AI autonomy in final goal decisions.
- 4.3. Feedback and Revision Across Humans and Agents: 76.0% of 588 tasks with recorded difficulties moved forward through human intervention, compared with 23.0% recovering autonomously and 1.0% remaining unresolved.Most human help supplied context or diagnosis rather than directly editing or taking over the work.
- 4.3. Feedback and Revision Across Humans and Agents: 95.9% of 627 main AI outputs entered deliverables in some form, while agents made revisions after human feedback in 75.4% of substantively revised tasks.Only 1.0% of outputs were not adopted, and direct human editing occurred in 19.2% of substantively revised tasks.
5. Open Challenges
The section identifies unresolved technical and governance challenges for sustained recursive self-improvement, including generating valuable research directions, learning from experience, and preserving meaningful human oversight.
- 5. Open Challenges: AI can perform predetermined R&D tasks and draft plans, but sustained recursive self-improvement remains unresolved.The section frames the gap as moving beyond task-level competence toward improving the research process itself.
- 5.1. Technical Challenges: 55.4% of method and parameter decisions involved agent proposals, compared with 16.9% of goal or scope decisions.Most proposals were variations within researcher-fixed directions, while new workflows still came from researchers prompting agents toward them.
- 5.1. Technical Challenges: Models struggle to prioritize directions and identify preliminary findings worth scaling, especially when compute is limited.This challenge can matter more than optimizing individual experiments.
- 5.1. Technical Challenges: Communication among working agents and systematic low-cost exploration are proposed to address homogeneous research directions.Researchers’ questions, preferences, and trade-offs broadened exploration, while anticipation is not treated as a permanent human advantage.
- 5.1. Technical Challenges: Experience from failed runs often remained session-bound, requiring researchers to provide diagnoses and carry lessons into later tasks.Existing workflows, skills, and records can accumulate experience without altering the model’s basic research capability.
- 5.2. Human Oversight and Safety: Humans retained final authority, making 85.5% of method and parameter decisions and 93.4% of goal or scope decisions.Meanwhile, the daily median number of agent actions per human prompt rose from 11.0 to 28.5 over four weeks, lengthening chains that one person may not inspect closely.
6. Conclusion
Atria Dawn Preview combines a large agentic foundation model with training grounded in executable environments and externally verified outcomes. Its evaluation and development record show competitive broad capability alongside a collaboration pattern in which agents propose and execute while humans retain judgment.
- 6. Conclusion: Atria Dawn Preview is built on a 744-billion-parameter mixture-of-experts foundation and trained through a Verifiable Experience Pipeline.The pipeline grounds tasks in real execution environments and checks outcomes against external signals.
- 6. Conclusion: Across 16 benchmarks spanning research, digital work, software engineering, and other agentic tasks, Atria Dawn shows competitive performance.The evaluation covers tool use, search and research, workspace and professional tasks, terminal engineering, ML engineering, and cybersecurity.
- 6. Conclusion: AI was used in 96.5% of 739 tasks with definitive involvement responses, proposed 64.6% of 567 methods and decisions, and humans made 85.5% of final choices.Among 588 tasks with recorded difficulties, 76.0% advanced through human intervention involving context, clarification, diagnosis, or method adjustments.
A. Contributions
The listed contributors are ordered alphabetically by last name.
- A. Contributions: Contributor listing follows alphabetical order by last name.
- A. Contributions: The ordering rule applies to the contributor list.
- A. Contributions: Last names determine contributor-list order.
A.2. Core Contributors
The paper’s core contributors are listed alphabetically by last name.
- A.2. Core Contributors: Core contributors are listed in alphabetical order by last name.
- A.2. Core Contributors: The contributor sequence is determined by last names.
- A.2. Core Contributors: The listed names form the core-contributor roster.
A.3. Contributors
The paper credits a large group of contributors, whose names are listed across the contributor passages.
- The contributor list spans two passages and includes numerous researchers.
A.4. Advisors
The paper identifies a separate group of advisors involved in the work.
- The advisor list names Yaojie Lu, Tao Ji, Hongyu Lin, and additional advisors.
B. Evaluation Protocols
The evaluation protocols specify benchmark versions, execution environments, resource limits, tools, scoring procedures, and run configurations across 16 benchmarks.
- The protocols cover benchmark versions, execution budgets, tools, and scoring procedures for 16 benchmarks.
- Search and research evaluations use shared tools, large interaction budgets, context windows, and multiple independent scoring passes.DeepResearch Bench II uses identical search tools and a four-hour limit, with three scoring passes.
- AgentCompass evaluations use specified datasets, compute allocations, execution limits, coding environments, and an LLM judge.Workspace-Bench-Lite and Workspace-Bench use 2 CPUs, 8 GiB, and a two-hour limit; GDPval uses 4 CPUs, 8 GB, and four hours.
- JobBench uses 65 tasks with a 7,200-second timeout, up to two attempts, concurrent workers, and mean task scoring.API failures and two failed attempts receive zero scores.
- Software and terminal benchmarks impose controlled environments, timeouts, interaction caps, repeated runs, and restrictions against information leakage.SWE-bench Pro disables network access and removes Git history; Terminal-Bench reports average pass rates over four runs.
- Other evaluations define benchmark-specific harnesses, datasets, retrieval tools, weighting schemes, judges, seeds, and repeated scoring passes.The protocols include MLE-bench Lite, CyberGym, AutomationBench, BFCL v4, SkillsBench, Banking Knowledge, three search benchmarks, and DeepResearch Bench II.