Source-linked AI summary

AI for Auto-Research: Roadmap & User Guide

Lingdong Kong, Xian Sun, Wei Chow, Linfeng Li, Kevin Qinghong Lin, Xuan Billy Zhang, Song Wang, Rong Li, Qing Wu, Wei Gao, Yingshuo Wang, Shaoyuan Xie, Jiachen Liu, Leigang Qu, Shijie Li, Lai Xing Ng, Benoit R. Cottereau, Ziwei Liu, Tat-Seng Chua, Wei Tsang Ooi

arXiv:2605.18661v2cs.AI

TL;DR

AI-assisted research increasingly spans the full research lifecycle, but evidence remains limited on whether generated artifacts are scientifically reliable and meaningfully validated. This paper maps AI across four research phases and finds that systems produce artifacts more effectively than they verify novelty, faithfulness, reproducibility, and scientific meaning, supporting human-governed collaboration.

  • Problem

    Evaluation remains a central bottleneck because no single metric captures research quality across the lifecycle’s diverse artifacts and processes.

  • Method

    The paper presents an end-to-end analysis organized into four epistemological phases and eight stages spanning Creation, Writing, Validation, and Dissemination.

  • Results

    Across the lifecycle, AI increasingly produces research artifacts but remains less reliable at verifying their novelty, faithfulness, reproducibility, and scientific meaning.

  • Takeaways & Limitations

    The most credible deployment paradigm is human-governed collaboration, with AI assisting routine research tasks while researchers retain scientific judgment and responsibility.

  • Takeaways & Limitations

    The paper is a structured snapshot through its search cutoff, and AI-generated research outputs require independent verification before scholarly use.

Abstract

from arXiv · show

AI-assisted research is crossing a threshold: fully automated systems can now generate research papers for as little as $15, while long-horizon agents can execute experiments, draft manuscripts, and simulate critique with minimal human input. Yet this productivity frontier exposes a deeper integrity problem: under scientific pressure, even frontier LLMs still fabricate results, miss hidden errors, and fail to judge novelty reliably. Studying developments through April 2026, we present an end-to-end analysis of AI across the complete research lifecycle, organized into four epistemological phases: Creation (idea generation, literature review, coding & experiments, tables & figures), Writing (paper writing), Validation (peer review, rebuttal & revision), and Dissemination (posters, slides, videos, social media, project pages, and interactive agents). We identify a sharp, stage-dependent boundary between reliable assistance and unreliable autonomy: AI excels at structured, retrieval-grounded, and tool-mediated tasks, but remains fragile for genuinely novel ideas, research-level experiments, and scientific judgment. Generated ideas often degrade after implementation, research code lags far behind pattern-matching benchmarks, and end-to-end autonomous systems have not yet consistently reached major-venue acceptance standards. We further show that greater automation can obscure rather than eliminate failure modes, making human-governed collaboration the most credible deployment paradigm. Finally, we provide a structured taxonomy, benchmark suite, and tool inventory, cross-stage design principles, and a practitioner-oriented playbook, with resources maintained at our project page.

1 Introduction

AI-assisted research is expanding from local writing and coding support toward end-to-end lifecycle operation, but its ability to generate plausible artifacts still outpaces its ability to verify novelty, correctness, faithfulness, and scientific meaning. The paper organizes this landscape into four phases and eight stages, synthesizes its tools and benchmarks, and identifies capability boundaries and governance challenges.

  • Motivation: AI systems are beginning to operate across the research lifecycle, with The AI Scientist generating complete papers at roughly $15 each and FARS producing 100 papers in 228 hours.FARS consumed 11.4 billion tokens and averaged one paper every 2.3 hours.
  • Problem: AI-generated ideas, code, manuscripts, and reviews can appear plausible while remaining unreliable in novelty, faithfulness, executability, correctness, or scientific meaning.The passages note that ideas may weaken after implementation and code may run while implementing the wrong algorithm.
  • Problem: Research must be analyzed as a lifecycle because errors introduced in ideas, experiments, or claims can amplify downstream when outputs lack preserved evidence or provenance.The lifecycle links ideas to experiments, claims, manuscripts, revisions, and public-facing summaries.
  • Scope and framework: The paper presents an end-to-end analysis organized into four epistemological phases and eight stages spanning Creation, Writing, Validation, and Dissemination.Creation covers idea generation, literature review, coding and experiments, and tables and figures; the other phases cover paper writing, review and revision, and public dissemination.
  • Central findings: AI performs best on structured, grounded, and externally checkable tasks, while capability drops sharply on open-ended work requiring novelty, implicit domain knowledge, long-horizon reasoning, or scientific judgment.Across stages, artifact generation outpaces verification: systems can produce plausible outputs faster than they can establish that those outputs are correct, faithful, or meaningful.
  • Contributions: The work contributes a unified taxonomy, a synthesis of tools and benchmarks, and cross-cutting capability boundaries involving faithfulness, judgment, reproducibility, provenance, governance, generalization, and cognitive ownership.The synthesis tracks evolution from prompt-based assistance to retrieval-augmented, agentic, fine-tuned, and hybrid workflows.

2 Preliminaries

The paper organizes AI-assisted research into four functional phases spanning eight interconnected stages, while emphasizing recurring methodological families and feedback loops across the lifecycle. Its literature corpus is concentrated in Creation, indicating uneven research maturity and publication availability across phases.

  • 2.1 Research lifecycle: The research lifecycle comprises eight interconnected stages grouped into Creation, Writing, Validation, and Dissemination, covering production, manuscript organization, scrutiny, and broader communication.Creation produces evidence and artifacts; Writing organizes them into a manuscript; Validation challenges and refines claims; Dissemination communicates them to broader audiences.
  • 2.1 Research lifecycle: The lifecycle is iterative rather than strictly linear: reviewer critiques can trigger new Creation experiments, while dissemination can expose ambiguities or errors requiring Writing revisions.These feedback loops are especially important for AI-assisted workflows because unchecked errors can propagate across stages.
  • 2.2 Methodological families: Five recurring methodological families organize AI systems across stages: prompt engineering, retrieval-augmented generation, training-free agentic methods, training-based methods, and hybrid approaches.The families are neither mutually exclusive nor strictly chronological; they describe how systems elicit, ground, specialize, and orchestrate LLM behavior.
  • 2.3 Literature scope and methodology: The literature corpus spans all four phases but concentrates most heavily on Creation, particularly literature review, coding, and experiment automation, followed by Writing, Validation, and Dissemination.This imbalance reflects both research maturity and publication availability: Creation tools are more often benchmarked and open-sourced, whereas dissemination tools are frequently commercial and workflow-specific.
  • 2.4 Developments and timeline: Across the full lifecycle, the field is constrained not only by model capability but also by orchestration, evaluation, reliability, and governance.The paper therefore frames progress as a cross-stage systems problem rather than a capability-only problem.

3 Phase 1: Creation

Phase 1, Creation, covers idea generation, literature review, evidence production, and visual representation, and currently has the richest tool ecosystem but uneven maturity. AI is strongest when ideas and synthesis are externally grounded, while implementation feasibility, citation fidelity, research-level code, and scientific judgment remain difficult.

  • Phase overview: Creation spans idea generation, literature review, coding and experiments, and tables and figures, answering what the contribution is and what evidence supports it.These stages are denoted S1–S4.
  • Phase overview: Creation has the richest tool ecosystem and broadest benchmark coverage, but maturity is uneven across its four stages.Idea generation has extensive tooling; literature review is improving; coding performance drops on genuinely novel research code; and tables and figures remain underdeveloped.
  • S1 Idea Generation: 100+ NLP researchers rated LLM-generated ideas significantly higher in novelty than human ideas (p < 0.05), but apparent novelty often fails to translate into executable, valuable research.IdeaBench reports novelty above 0.6 for many LLMs but feasibility below 0.5, while HindSight evaluates whether ideas retain impact after implementation and testing.
  • S1 Idea Generation: Multi-agent ideation can improve novelty, with VirSci reporting 5.24 versus 4.94 for a single-agent AI Scientist baseline, but extra critique rounds yield diminishing returns and generated ideas may cluster narrowly.A SIGDIAL 2025 study found three critique–revision rounds often sufficient, while the Artificial Hivemind study identified narrow clustering as a possible structural limitation.
  • S2 Literature Review: Retrieval-augmented and structure-aware literature systems can produce reasonable surveys and improve selected dimensions toward expert performance, but citation grounding remains a bottleneck.ScholarCopilot reports only 40.1% top-1 citation accuracy, showing that plausible synthesis remains easier than reliably linking claims to correct sources.

4 Phase 2: Writing

AI-assisted writing is now mainstream, spanning local editing to end-to-end paper generation, but reliable systems preserve researcher control while autonomous outputs remain limited in argumentation, rigor, and scientific judgment. Evaluation therefore emphasizes manuscript quality and disclosure over unreliable detection alone.

  • Writing adoption: AI modification appears in up to 17.5% of computer science abstracts and 13.5% of biomedical abstracts, while more than half of researchers report seeking AI writing help.These imperfect measurements together indicate that AI writing assistance is embedded in research practice.
  • Semi-automated assistance: Semi-automated tools increasingly follow an “AI writes with you” paradigm, handling polishing, citations, drafting, and revision while researchers retain control of the argument.This approach is most credible when AI augments rather than replaces framing, interpretation, and defense of a contribution.
  • Fully automated generation: Fully automated systems can produce complete paper-like artifacts, but their outputs often remain constrained by shallow argumentation and weak experimental rigor.The supplied discussion presents end-to-end generation as feasible but insufficient for reliably producing publishable research.
  • Fully automated generation: 5.36 on the ICLR scale was reported for CycleResearcher-generated papers, below the accepted-paper average of 5.69.The gap indicates that publishable work still requires argumentative depth, experimental rigor, and reviewer anticipation beyond surface fluency.
  • Section-specific systems: 79% of the time, human experts preferred APRES-revised papers, showing that rubric-guided systems can improve manuscript components without generating the entire paper.Other section-specific systems target Future Work and introduction generation.
  • Assessment and governance: 26.89% reduction in Proxy MAE relative to individual human reviewers was reported by CycleReviewer for score prediction, while detection remains unreliable because of unacceptable false positives.Major venues are therefore shifting toward requiring authors to declare AI use, while quality evaluation examines correctness, citations, coherence, completeness, novelty, and style.
  • Assessment and limitations: AI can increase publication output while complex-language AI-assisted papers may be less likely to be accepted, illustrating a productivity–quality divergence.Fluent text can conceal unresolved problems with grounding, citation support, experimental sufficiency, and contribution-level reasoning.

5 Phase 3: Validation

Phase 3 validates manuscripts through adversarial peer review and rebuttal-driven revision, where AI assists with critique, opinion synthesis, reviewer matching, and response drafting but remains unreliable at independent judgment, robustness, evidence generation, and accountable follow-through.

  • Phase 3: Validation: Validation covers peer review and rebuttal with revision, asking whether a manuscript meets the field’s epistemic standards.Reviewers assess unsupported claims, methodological flaws, missing comparisons, unclear writing, and novelty.
  • Peer review: AI systems can summarize manuscripts, draft reviews, investigate weaknesses, synthesize opinions, and support reviewer matching, but disagreement often produces diluted compromises instead of defensible judgments.Reviewer matching supports allocation rather than replacing expert judgment, yet requires conflict handling, expertise modeling, and human oversight.
  • Peer review: 0.42 Spearman correlation from Stanford Agentic Reviewer is comparable to 0.41 human–human correlation, but consistency does not prevent inflated scores, bias, shallow critiques, or acceptance misclassification.ReviewAgents was trained on 37,403 papers and 142,324 reviews, illustrating scale without resolving judgment quality.
  • Peer review: 15.8% of ICLR 2024 reviews were estimated to be AI-assisted, while prompt injections can raise review scores and alter rankings, making robustness and governance central deployment concerns.At least one AI-assisted review reached 49.4% of submissions, and a major 2026 conference rejected 497 papers for AI-use policy violations.
  • Rebuttal and revision: 55.7–57.6% acceptance rates followed score improvement after rebuttal versus 7.8–12.4% with unchanged scores, yet AI cannot generate missing evidence or ensure revision commitments are fulfilled.ICLR 2025 authors averaged 11.8 commitments per paper, with approximately 25% unfulfilled in camera-ready versions.

6 Phase 4: Dissemination

Dissemination converts validated manuscripts into audience-adaptive artifacts, including posters, slides, videos, social media, project pages, and interactive agents. AI substantially lowers production costs for predictable formats, but public-facing and agentic outputs require author responsibility and evaluation of correctness, reproducibility, and boundary awareness.

  • Phase 4: Dissemination: Dissemination creates independent knowledge artifacts adapted to audiences, media, and interaction modes rather than merely reproducing the paper.Outputs include visual narratives, oral presentations, multimodal explanations, public-facing summaries, project pages, and interactive agents.
  • Posters: Poster automation is shifting from direct paper summarization toward editable, design-aware production across heterogeneous source modalities.Recent systems support interactive editing, unified manipulation, quality-aware generation, and inputs spanning multiple domains.
  • Videos and talks: Video remains among the hardest Paper2X formats because systems must synchronize slides, subtitles, speech, and temporal or avatar-based presentation while preserving fidelity and concision.Current systems therefore work best as first-draft generators of synchronized presentation assets.
  • Social media and web: Public-facing dissemination is a distinct trust problem because overclaims, omitted caveats, or misleading comparisons can shape perceptions when outputs are read without the paper.AI assistance is most credible for audience-specific drafts, style variants, and claim-checking, while authors retain final messaging and factual responsibility.
  • Interactive agents: Agentic dissemination should be evaluated for answer quality, tool correctness, reproducibility, error handling, and boundary awareness, while future systems increasingly expose executable interfaces and reusable workflows.No large-scale adoption study has yet established whether these systems are effective in practice.
  • Phase 4: Dissemination: AI substantially lowers dissemination costs when the input paper is complete and the target format has a predictable structure.The phase increasingly includes interactive agents or tools alongside posters, slides, videos, project pages, and social media posts.

7 Cross-Cutting Analysis

AI research systems increasingly combine sequential, agentic, multi-agent, and tool-based designs, but their coverage remains concentrated in Creation and Writing rather than Validation and Dissemination. Evaluation is consequently shifting toward multidimensional, execution-grounded, long-horizon, and lifecycle-level protocols because artifact production is easier than verifying scientific validity.

  • System architectures: Current end-to-end systems concentrate on Creation and Writing, while substantive Validation and Dissemination remain comparatively underrepresented.Recent systems increasingly combine search, role separation, tool use, persistent memory, and self-improvement, extending beyond predominantly sequential 2024 pipelines.
  • System architectures: Sequential pipelines are interpretable and easy to implement, but errors can propagate from weak ideas and incorrect code into misleading results and polished unsupported claims.Each stage produces an artifact that becomes the next stage’s input, creating phase-boundary risk.
  • Evaluation: Evaluation is moving from narrow output metrics toward multidimensional, process-aware, execution-grounded, and increasingly lifecycle-oriented assessment.Different research artifacts require distinct criteria, including novelty, feasibility, citation accuracy, semantic correctness, review consistency, and fidelity.
  • Evaluation: ρ = 0.41 human–human correlation shows that expert evaluation is indispensable but expensive, slow, noisy, and difficult to scale.Automated evaluators provide useful but imperfect review-style signals: cleReviewer reports a 26.89% reduction in Proxy MAE, while Stanford Agentic Reviewer reaches ρ = 0.42 versus human ρ = 0.41.
  • Evaluation: Execution-guided evaluation checks whether claims are supported by generated evidence, implementation, and experiments rather than relying on textual judgment alone.PaperBench decomposes papers into individually gradable subtasks, while execution-guided search uses empirical feedback to improve discovery workflows.
  • Evaluation: No existing benchmark evaluates the complete research lifecycle with human-equivalent rigor while preserving traceability from ideation through dissemination.Long-horizon benchmarks remain limited, and practical evaluation must also account for delayed feedback, failed attempts, evolving evidence, computational cost, and reproducibility.

8 Conclusion

AI-assisted research spans the complete lifecycle, but systems remain better at producing plausible artifacts than verifying their scientific meaning. The paper therefore advocates human-governed collaboration, with AI reducing mechanical friction while researchers retain responsibility for scientific judgment and final accountability.

  • Lifecycle framework: The framework organizes AI-assisted academic research into four phases—Creation, Writing, Validation, and Dissemination—and eight lifecycle stages.The stages span ideation, literature review, coding and experiments, tables and figures, paper writing, peer review, rebuttal and revision, and Paper2X dissemination.
  • Reliability boundary: AI systems increasingly produce research artifacts, but plausible outputs can conceal failures in scientific meaning and verification.Examples include ideas weakening after execution, misrepresented evidence, incorrect algorithms, shallow arguments, missed methodological flaws, and unfulfilled revision promises.
  • Deployment principles: Human-governed AI-assisted research is the most credible path forward: AI should reduce mechanical friction while researchers retain judgment, interpretation, experimental design, argumentation, and final responsibility.Future systems should preserve provenance, use retrieval and execution grounding, and support human checkpoints at phase boundaries.

Responsible Use and Limitations

AI-assisted research tools are presented as aids for responsible human-governed collaboration, not replacements for human scientific judgment. Current systems are most reliable for structured assistance, while humans remain responsible for novelty, interpretation, verification, authorship, and accountability.

  • Responsible Use: The work informs responsible use of AI-assisted research tools rather than endorsing the replacement of human scientific judgment with full automation.It frames deployment around human-governed collaboration.
  • Reliable Assistance: Current systems are most reliable for retrieval, drafting, coding, visualization, review support, and dissemination.These uses are characterized as assistance rather than autonomous scientific judgment.
  • Human Responsibility: Humans retain responsibility for novelty, interpretation, verification, authorship, and accountability.The passage assigns these scientific and ethical responsibilities to people.
  • Limitations: The paper should be read as a structured snapshot because the field evolves rapidly.Its scope is bounded by the state of the field at the time of the search.

Appendix · A Auto-Research Tool Inventory

The appendix provides a comprehensive inventory of all surveyed works, organized by research stage.

  • A Auto-Research Tool Inventory: The appendix inventories all surveyed works and organizes them by stage.

A.1 Phase 1: Creation

Phase 1 spans idea generation, literature review, and coding and experiments, with AI systems increasingly using retrieval, multi-agent collaboration, external signals, and tool-mediated workflows. However, evaluations show persistent limits in novelty assessment, hallucination control, research-level coding, and autonomous scientific reliability.

  • Idea Generation: Idea-generation systems use internal knowledge, external signals, and multi-agent collaboration to improve novelty, quality, feasibility, or diversity against baselines.Reported approaches include iterative refinement, knowledge-graph reasoning, paper-anchored generation, and critique-revision protocols.
  • Idea Generation: Idea evaluation remains unreliable: LLM novelty negatively correlated with impact (ρ=−0.29), and SoundnessBench identified optimism bias in distinguishing sound from flawed ideas.The inventory also reports limits of LLM-as-Judge novelty assessment.
  • Literature Review: Literature-review systems combine retrieval, academic databases, graph traversal, multi-agent generation, citation verification, and evidence-constrained synthesis.Examples range from citation-fidelity and retrieval-precision benchmarks to structured survey generation and trusted-repository routing.
  • Coding & Experiments: Coding and experiment tools range from software-engineering agents to paper-to-code systems, but benchmarks expose substantial gaps in research-level autonomous coding.SWE-bench Pro’s best score was 23% on 1,865 enterprise problems, and SWE-EVO’s best score was 25%.

A.2 Phase 2: Writing

Phase 2 covers AI-assisted paper writing, spanning collaborative drafting, citation and editing assistance, fully automated generation, quality evaluation, and AI-use detection. The inventory reports substantial automation and evaluation progress, but also highlights uncertainty around writing quality, citation reliability, and scientific integrity.

  • Semi-Automated Writing Assistance: Writing assistance ranges from collaborative workflows and in-editor tools to citation support, revision systems, and multi-agent Overleaf plugins.Examples include CoAuthor, PaperDebugger, ScholarCopilot, XtraGPT, and PaperMentor.
  • Fully Automated Paper Generation: $2–13/paper and 84% cost reduction were reported for Agent Laboratory, while AI Scientist generated papers for $15/paper across 3 ML subfields.Agent Laboratory received a 3.5–4.0 score, and CycleResearcher reported 5.36 on the ICLR scale versus 5.24 for preprints and 5.69 for accepted papers.
  • Writing Quality and AI Detection Assessment: Writing evaluation includes expert preference, AI judging, introduction benchmarks, section-generation benchmarks, and citation-predictive rubrics.APRES reported 79% expert preference; CycleReviewer reported a 26.89% MAE reduction versus individual human reviewers; Stanford Agentic reported ρ = 0.42 versus human ρ = 0.41.
  • Writing Quality and AI Detection Assessment: Detection and integrity research reports up to 17.5% of CS papers as AI-modified, near-zero false-positive watermarking under controlled conditions, and risks from indirect data poisoning.The inventory also notes that AI expands research impact while contracting focus, and that 57% of researchers use AI in peer review.

A.3 Phase 3: Validation

Phase 3 surveys AI-assisted peer review across review generation, meta-review, matching, adversarial robustness, bias, detection, policy, and consistency assessment. The inventory shows expanding automation alongside measurable gains and persistent concerns about factual grounding, bias, manipulation, and detector reliability.

  • Automated Review Generation: AI review systems generate strengths/weaknesses analyses, multi-LLM reviews, meta-review syntheses, and factually grounded critiques, with DeepReviewer reporting an 88.21% win rate versus GPT-o1.MARG produces 3.7 good comments per paper, 2.2× over baseline; DeepReviewer reports 64% accept/reject accuracy.
  • Meta-Review & Reviewer Matching: Studies extend beyond generation to full review-lifecycle simulation, expertise-based reviewer matching, and analyses of social, authority, and reviewer bias.AgentReview simulates the full review lifecycle, while RATE performs expertise-based matching via profile distillation.
  • Adversarial Attacks & Bias Analysis: Peer-review systems remain vulnerable to adversarial manipulation and bias, including universal adjective triggers, score inflation to ∼8, and 95.8% misclassification of rejected papers as acceptable.A 5% manipulation rate flips 12% of outcomes, while AI-assisted reviews represented 15.8% of ICLR reviews and increased borderline scores by 4.9 percentage points.
  • Detection & Policy: Detection and policy research examines 788,984 AI-written reviews, widespread academic use, and enforcement failures, including all 5 SOTA detectors misclassifying LLM-polished reviews.A Nature survey reports 57% of 1,600 academics use AI in peer review; 497 papers were rejected, ∼2% of submissions.
  • Review Consistency and Bias Assessment: Consistency benchmarks can approach human agreement, with Stanford Agentic reporting ρ = 0.42 versus human ρ = 0.41, while other work targets grounding and graph-based quality improvements.ReViewGraph reports a +15.73% average improvement, and ClaimCheck identifies gaps in factual basis.

A.4 Phase 4: Dissemination

Dissemination tools now span paper-to-poster, slide, video, website, and social-media workflows, alongside benchmarks for fidelity, comprehension, coherence, interactivity, and correspondence. The inventory includes multi-agent systems, retrieval-augmented conversion, interactive editing, audience conditioning, and narrated presentation generation.

  • Paper2Video: Paper-to-video methods apply top-down decomposition, paper–video training pairs, and agentic presentation generation, with PresentEval reporting near-human-level performance.Paper2Video uses 101 paper–video pairs and reports +10% PresentQuiz accuracy; PresentAgent is evaluated on the PresentEval benchmark and approaches human-level results.
  • Paper2Web & Social Media: Paper2Web systems generate multimedia-rich academic homepages, jointly automate posters, videos, and blogs, and evaluate interactivity in LLM-generated scientific web apps.Paper2Web covers 10,716 papers, while ResearchStudio- automates multiple formats jointly and I-WebGenBench assesses interactivity.
  • Fidelity and Adoption Assessment: Dissemination benchmarks assess presentation content, design, coherence, comprehension, end-to-end narrated-video quality, audience conditioning, and fine-grained paper–slides–video correspondence.PPTEval covers 10,448 presentations; PresentQuiz reports +10% over humans on comprehension, and PresentEval reports near-human-level narrated-video quality.

A.5 Cross-Phase: End-to-End Systems

This section provides a comprehensive inventory of end-to-end and cross-phase systems. It notes that some evaluation information may be uncertain.

  • The section catalogs end-to-end and cross-phase systems in a comprehensive inventory.
  • Some evaluation information in the inventory may be uncertain.

B Survey Coverage Comparison & Taxonomy Analysis

The eight-stage framework extends prior taxonomies by covering the complete research lifecycle and grouping stages by epistemological function. It also makes feedback loops and cross-stage dependencies explicit, showing how review, revision, and dissemination can redirect research while keeping evidence, claims, and communication aligned.

  • Framework distinctions: The framework analyzes AI auto-research across the complete lifecycle, grouping eight stages by epistemological function rather than task name or autonomy level.This organization distinguishes Creation, Writing, Validation, and Dissemination as four functional phases.
  • Coverage comparison: Compared with AI4Research, the framework newly elevates S4 (Tables & Figures), S7 (Rebuttal & Revision), and S8 (Dissemination) as independent lifecycle stages.AI4Research defines Comprehension, Survey, Discovery, Writing, and Review, overlapping with S1–S3, S5, and S6.
  • Coverage comparison: Unlike autonomy-level taxonomies, the framework specifies where systems operate in the research lifecycle, while allowing each stage to be instantiated at different autonomy levels.The autonomy axis remains complementary rather than competing with lifecycle placement.
  • Feedback structure: Separating S6 (Peer Review) from S7 (Rebuttal & Revision) makes the review–response loop explicit, unlike four-part structures that combine review without a distinct feedback stage.This separation clarifies how critique leads to revision rather than treating review as an isolated downstream step.
  • Cross-stage dependencies: Review and dissemination can redirect work to S3 for experiments, S4 for revised figures or tables, and S5 for manuscript restructuring, exposing ambiguities and propagating errors across stages.The framework emphasizes alignment among evidence, claims, critique, and public-facing communication in AI-assisted workflows.
Loading 2605.18661v2…