Source-linked AI summary
Accelerating scientific discovery with Co-Scientist
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Artiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, Anil Palepu, Keran Rong, Ryutaro Tanno, Khaled Saab, Fan Zhang, Jacob Blum, Andrew Carroll, Kavita Kulkarni, Nenad Tomasev, Dina Zverinski, Ivor Rendulic, Elahe Vedadi, Florian Hasler, Luka Rimanic, Marina Boia, Ivan Budiselic, Ben Feinstein, Mathias Bellaiche, Tom Sheffer, Jan Freyberg, Jeremy Ratcliff, Ottavia Bertolli, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Yuan Guan, Vikram Dhillon, Eeshit Dhaval Vaishnav, Byron Lee, Tiago R D Costa, José R Penadés, Gary Peltz, Yossi Matias, James Manyika, Demis Hassabis, Yunhan Xu, Pushmeet Kohli, Annalisa Pawlosky, Alan Karthikesalingam, Vivek Natarajan
TL;DR
Scientific discovery requires novel hypotheses for complex problems, but researchers face tension between deep expertise and broad cross-disciplinary insight. Co-Scientist addresses this with a Gemini-based multi-agent system that generates, debates, ranks, and evolves hypotheses using scaled test-time compute. Evaluations and biomedical validations provide preliminary evidence of improved hypothesis quality and experimentally supported discoveries, while KIRA6 remains constrained by limited human safety and broader validation data.
Problem
Researchers must combine increasingly deep subject-matter expertise with broad cross-disciplinary insight, making novel hypothesis generation for complex scientific problems challenging.
Method
Co-Scientist uses a Gemini-based multi-agent architecture with asynchronous task execution, scientific debate, tournament evolution, and iterative hypothesis refinement.
Results
Co-Scientist outperformed state-of-the-art agentic and reasoning models in generating high-quality hypotheses and supported end-to-end biomedical validations, including AML drug repurposing and synergistic combinations.
Takeaways & Limitations
Co-Scientist provides preliminary evidence that human-AI collaboration can amplify and accelerate scientists’ work across varied biomedical discovery problems.
Takeaways & Limitations
KIRA6 lacks human safety data, and broader testing across AML cell lines, primary patient samples, comparator drugs, in vivo models, and clinical settings remains necessary.
Abstract
from arXiv · showhide
Scientific discovery is driven by scientists generating novel hypotheses for complex problems that undergo rigorous experimental validation. To augment this process, we introduce Co-Scientist, a multi-agent AI system built on Gemini for structured scientific thinking and hypothesis generation. Co-Scientist aims to help scientists discover new original knowledge. Conditioned on their research objectives and prior scientific evidence, it formulates demonstrably novel research hypotheses for experimental verification. The system's design involves agents continuously generating, critiquing and refining hypotheses accelerated by scaling test-time compute. Key contributions include: (1) a multi-agent architecture with an asynchronous task execution framework for flexible compute scaling; (2) a tournament evolution process for self-improving hypotheses generation. Automated evaluations show continued benefits of test-time compute scaling, improving hypothesis quality over time. While general purpose, we focus the validation in three biomedical applications: drug repurposing, novel target discovery, and explaining mechanisms of anti-microbial resistance. Specifically, Co-Scientist helped identify new drug repurposing candidates and synergistic combination therapies for acute myeloid leukemia, which were validated through in vitro experiments. These real-world validations demonstrate the potential of Co-Scientist to accelerate scientific discovery and usher in an era of AI empowered scientists.
1 Google Cloud AI Research, Zurich, Switzerland
Co-Scientist is a Gemini-based multi-agent system that iteratively generates, critiques, ranks, and evolves scientific hypotheses, with test-time compute scaling improving outputs. Initial validations span automated and expert evaluations plus biomedical experiments, including AML drug repurposing and combination therapies.
- System design: Co-Scientist combines specialized agents, asynchronous resource allocation, scientific debate, tournament ranking, and iterative evolution to refine research hypotheses.The system also supports literature search, external tools, and scientist feedback.
- System evaluation: Increasing test-time compute produced significant quality enhancement in the most recent hypotheses compared with initial ones across both evaluation metrics.The trends were observed across 203 diverse research goals partitioned into ten temporal buckets.
- System evaluation: Across 11 expert-evaluated research goals, Co-Scientist achieved the highest ratings for hypothesis novelty and impact and was the preferred AI system.Novelty concerned previously unpublished hypotheses, while impact concerned significant open questions and potential practical applications.
- System evaluation: Co-Scientist outputs were most preferred by all four LLM judges and eventually significantly surpassed frontier language and reasoning models on the 15-goal biomedical subset.Newer reasoning models were competitive while requiring significantly less compute and reasoning time.
- Real-world validations: In AML, Binimetinib, Pacritinib, and Cerivastatin inhibited cell viability, while the novel candidate KIRA6 showed selective activity in KG-1a cells.KIRA6 had an IC50 of 10 nM in KG-1a versus 180 nM in the non-AML TK6 control; Binimetinib reached an IC50 as low as 2 nM in most AML cell lines.
Discussion
Co-Scientist combines multi-agent scientific reasoning with iterative hypothesis refinement and test-time compute scaling. Initial biomedical applications and component-level evaluations support its utility, while literature access, model reliability, preliminary validation, and responsible integration remain important boundaries.
- Discussion: Co-Scientist iteratively generates, debates, selects, and evolves hypotheses rather than relying on brute-force generation.Its context memory and self-improvement cycle support synthesis of information and identification of knowledge gaps.
- Discussion: Experimental applications produced tractable hypotheses for AML, liver fibrosis, and bacterial mobile genetic element transfer.AML candidates showed in vitro efficacy; several anti-fibrotic compounds were validated, including one FDA-approved drug, and a novel transfer mechanism was recapitulated.
- Discussion: The model-agnostic architecture can use advancing frontier LLMs without retraining the full agentic framework.The authors expect improved hypothesis quality and greater task complexity as frontier models advance.
- Discussion: Open-access literature omits some paywalled prior art and negative results, while contradictory sources can propagate erroneous or irreproducible findings.The authors identify provenance-enhanced agents and stronger fact-checking as future directions.
- Discussion: Underlying-model limitations include imperfect factuality and hallucinations, and hypothesis validation remains preliminary.Broader deployment also requires safeguards against bias, diminished critical thinking, homogenized research directions, and low-quality scientific artifacts.
- Discussion: Expanded evaluations are needed to assess generalizability across more disciplines with objective metrics and larger cohorts of domain experts.The authors propose stress-testing the system with diverse and complex research queries.
- Discussion: External search improved Reflection novelty scores from 6.14 without search to 2.38 with search and increased correctness from 7.4 to 8.46.On GPQA, AUC increased from 0.643 to 0.651 in the reported Gemini 2.0 Flash run.
- Discussion: Iterative Evolution increased GPQA precision from 70.9% to 75.4% and average hypothesis quality from 4.7 to 5.6.Meta-review AUC also increased from 0.521 to 0.597 on the constructed dataset and from 0.629 to 0.634 on GPQA diamond.
Methods references
This section lists prior references covering scientific debate, Elo-style ranking, back-propagation, intransitivity, margin-of-victory extensions, constrained decoding, and Gemini models.
- Methods references: The list also cites foundational work on back-propagation and fast lexically constrained decoding for neural machine translation.The cited works are LeCun and Post & Vilar.
- Methods references: The references include work on debating with persuasive language models and scientific or hypothesis-ranking procedures.These citations include Khan et al. and Elo-related references.
- Methods references: Gemini is cited as a family of highly capable multimodal models.The reference identifies the Gemini Team, Google, as the authors.
Code availability
The full Co-Scientist source code is unavailable because of proprietary infrastructure, computational requirements, and safety considerations, while access is being expanded experimentally.
- Code availability: The full source code is not publicly available because the framework is deeply integrated with proprietary infrastructure.The authors also cite the computational resources required for large-scale test-time scaling and safety implications of unmonitored autonomous use.
- Code availability: Researchers can request experimental access, provisioned subject to computational resources.The authors expect broader access through Google APIs and bespoke interfaces over time.
- Code availability: The Gemini foundation model is publicly available through APIs, Google AI Studio, the Gemini App, and other surfaces.Technical details are described in corresponding Gemini reports.
- Code availability: The implementation used Python 3.11.7 and standard scientific computing and visualization libraries.In vitro dose-response fitting and IC50 estimation used GraphPad Prism 10.6.0.
- Code availability: The study was funded by Alphabet, and many authors are Alphabet employees who may own company stock.Some listed authors are employees of Sequome, and one is founder of Dendra Therapeutics.
1 Glossary of terminology and concepts
The glossary distinguishes novel drug repurposing candidates from novel targets and from traditional de novo drug discovery.
- 1 Glossary of terminology and concepts: A novel repurposing candidate is an existing drug with an established safety profile proposed for a different disease or condition.The definition specifies a chemical compound that binds to a target.
- 1 Glossary of terminology and concepts: Drug repurposing differs from traditional drug discovery, which identifies new chemical compounds binding targets implicated in disease.Repurposing instead starts from an existing drug with an established safety profile.
- 1 Glossary of terminology and concepts: A novel target is introduced as a biological entity, such as a gene or protein.
2 Co-Scientist internal evaluation
Co-Scientist’s internal evaluation examined whether its auto-evaluation Elo metric tracks answer quality and whether iterative, human-seeded refinement improves hypotheses. Results support concordance between Elo and accuracy and show that human-seeded ideas can improve through tournament evolution.
- Elo concordance: Higher Elo ratings concorded with higher averaged accuracy on the GPQA diamond benchmark.The Elo metric was evaluated against generated responses and reference Gemini 2.0 responses for the same questions.
- Elo concordance: 78.4% top-1 accuracy was achieved by selecting the highest-Elo Co-Scientist result for each GPQA question.Accuracy was computed against GPQA diamond ground truth.
- Human–AI refinement: Human experts’ initial biomedical solutions were inserted into tournaments for critique, debate, and evolutionary refinement alongside autonomously generated ideas.The evaluation used 15 highly complex biomedical research goals.
- Human–AI refinement: Human-seeded, AI-refined hypotheses eventually surpassed autonomously generated solutions as test-time compute increased.The finding was described as preliminary and measured through progressive Elo ratings.
- External evaluation: Automated preference rankings placed Co-Scientist outputs first across four LLM evaluators and agreed closely with human clinical-expert assessments.The automated framework provided a scalable validation alongside resource-intensive expert review.
- Safety considerations: Safety evaluation addressed potential misuse, dangerous or dual-use goals, unsafe hypotheses, misleading claims, and bias through adversarial testing.The paper states that extensive safety investigation was beyond this work’s scope.
3 Quantitative ablation and sensitivity analysis
Ablation and sensitivity analyses found that Co-Scientist’s specialized agents and prompting strategies materially improve ranking, refinement, review, and exploration quality. Iterative refinement and test-time computation produced measurable gains across GPQA and the constructed paper-QA benchmark.
- Ablation overview: Including the individual agents quantitatively improved Co-Scientist performance across the ablation experiments.The study isolated Generation, Reflection, Ranking, Evolution, Meta-review, and Proximity agents, including external-tool use.
- Ranking agent prompting strategy: Scientific debate prompting was significantly more accurate for comparing high-quality ideas.Its AUC for identifying the better idea was higher than the simple comparison prompt.
- Ranking agent prompting strategy: Scientific debate prompting reduced second-player positional bias and produced more stable, reliable evaluations.Score differences were more symmetric around zero than with the basic prompt.
- Evolution agent: 70.9% to 75.4%: GPQA precision increased after iterative refinement.On the constructed dataset, average hypothesis quality score increased from 4.7 to 5.6.
- Meta-review agent: 0.521 to 0.597: adding Meta-review improved AUC for predicting solution correctness on the constructed dataset.On GPQA, AUC increased from 0.629 to 0.634 with Flash 2.5.
- Proximity agent: High proximity scores of 4–5 correlated with small quality-score differences, supporting semantic-similarity-based clustering for diverse exploration.The proximity score was evaluated against ground-truth quality-score differences.
- Generation agent: Diverse generation strategies were crucial, with no single strategy dominating correct-hypothesis generation.On GPQA, “using focus areas” contributed 13.8% and “using generate prompt” 13.0%.
- Evolution agent: 4.7 to 5.6: iterative improvement increased the constructed dataset’s average maximum hypothesis-quality score by 0.9 points.The reported 95% confidence interval was 0.05 to 1.7.
4 Clinical expert evaluation of drug repurposing proposals in NIH Specific Aims format
Co-Scientist’s drug-repurposing hypotheses were formatted as NIH Specific Aims and evaluated by clinical experts across significance, innovation, rigor, and feasibility. Ratings were consistently favorable, while the evaluation remained limited by its single-center expert sample and lack of randomized phase III evidence.
- Clinical expert evaluation: The NIH Specific Aims format standardized disease description, unmet need, proposed solutions, and specific aims for systematic peer-style assessment.The format was selected because it is widely recognized by the research community.
- Clinical expert evaluation: Expert ratings covered 15 axes spanning significance, innovation, rigor, and feasibility using a five-point agreement scale.The rubric emphasized clinical relevance and potential clinical translation.
- Clinical expert evaluation: 0.745 Spearman’s rho (p < 0.001) indicated high agreement between the two independent expert groups.Both groups consistently assigned high ratings across evaluation criteria.
- Clinical expert evaluation: 78 drug-repurposing hypotheses were independently evaluated by two groups of clinical hematologists and oncologists in NIH Specific Aims format.The reviewers used an adapted NIH grant-proposal rubric.
- Disease Description and Unmet Need: AML proposals targeted an unmet need arising from toxicity, relapse, and limited effective options for relapsed or refractory disease.The section describes AML as particularly challenging for older or frail patients.
- Proposed Solution: Givosiran was proposed for AML because ALAS1 inhibition reduces heme biosynthesis, a pathway linked to AML proliferation, apoptosis, and drug sensitivity.The hypothesis especially concerns AML cells with upregulated heme biosynthesis or MYCN overexpression.
- Proposed Solution: Selinexor and lapatinib were proposed as COAD repurposing candidates targeting XPO1-linked tumor-suppressor regulation and ErbB signaling, respectively.The proposals aim to inhibit growth, enhance apoptosis, or address resistant disease in specified molecular contexts.
- Specific Aims evaluation rubric: The rubric required clear hypotheses and aims, clinically relevant endpoints, translational planning, factual accuracy, evidence-grounded assumptions, and originality.It also assessed whether proposals avoided over-extrapolation and speculative conclusions.
5 Additional information for drug repurposing evaluation
The drug-repurposing evaluation used a curated search space, computational dependency data, expert review, and in vitro testing to prioritize and assess hypotheses. AML experiments examined subtype-specific drug responses and combination synergies, including context-dependent effects across cell lines.
- 5.1 Co-Scientist suggests plausible drug repurposing candidates as rated by computational biology analyses and experts: Co-Scientist searched repurposing hypotheses across 2300 approved drugs and 34 cancer types, using prompts tailored to mechanisms, targets, approvals, and alternative actions.The generated proposal for each drug candidate included a hypothesis, a review, and a possible mechanism of action.
- 5.1 Co-Scientist suggests plausible drug repurposing candidates as rated by computational biology analyses and experts: DepMap scores provided a rapid sanity check by measuring dependency probabilities for known drug target genes in relevant cancer cell-line models.Drug-cancer pairs were ranked using Co-Scientist review scores from 1 to 5 and DepMap scores from 0.0 to 1.0.
- 5.1 Co-Scientist suggests plausible drug repurposing candidates as rated by computational biology analyses and experts: Higher Co-Scientist review scores tended to coincide with higher DepMap scores, and expert-review candidates required review score ≥4 plus DepMap score ≥0.99.All score-5 versus lower-score group comparisons were statistically significant in Supplementary Fig. 8, without multiple-comparison adjustments.
- 5.2 AML cell line selection rationale and mechanistic interpretation of drug sensitivities: AML cell-line selection emphasized genetic diversity and disease heterogeneity, including FLT3-mutant and FLT3-wild-type lines representing different AML subtypes.The selected panel included MOLM-13, KG-1a, HL-60, NOMO-1, and TK6 for experimental evaluation.
- 5.2 AML cell line selection rationale and mechanistic interpretation of drug sensitivities: KIRA6 sensitivity was heightened in primitive, stem-like KG-1a cells, supporting IRE1α blockade as potentially most effective in AML subsets enriched for stem-like states.More differentiated MOLM-13 and HL-60 lineages showed limited sensitivity to this mechanism, warranting further investigation of resistance.
- 5.5 Additional wet-lab results: Combination responses varied by cell line: MOLM-13 responses were predominantly synergistic, whereas KG-1a showed context-dependent mixtures of synergy and antagonism.For example, JNJ-64619178 + Selinexor was antagonistic at lower effect levels but synergistic at the highest effects in KG-1a cells.
6 Detailed Co-Scientist output for a validated AML repurposing candidate
Co-Scientist proposed repurposing KIRA6, an IRE1α inhibitor, for FLT3-ITD-positive AML, with a mechanistic rationale centered on ER stress, resistance, and combination therapy. The proposal was considered promising but requires broader experimental validation, especially for safety, selectivity, synergy, and novelty.
- Candidate and rationale: KIRA6 is proposed for AML treatment, particularly FLT3-ITD-positive disease, to disrupt protein homeostasis, induce ER stress, and overcome resistance.The plan includes combination therapy and in vitro and in vivo studies.
- Mechanism: IRE1α inhibition activates apoptosis and reduces proliferation and clonal survival, while suppressing MYC, NF-κB, and MCL-1 survival pathways.
- Experimental predictions: KIRA6 produced a dose-dependent reduction in MOLM-13 viability and increased ER stress and integrated stress response activation.
- Combination therapy: KIRA6’s proposed synergy with FLT3 inhibitors or chemotherapy remains a mechanistic hypothesis requiring in vitro testing.
- Safety boundary: KIRA6 has limited human safety data because it has not undergone clinical trials, making preclinical toxicity studies essential.
- Validation needs: The proposal’s selectivity, resistance effects, and efficacy need testing across additional AML cell lines, genetic backgrounds, and patient-derived samples.
7 Safety and ethical implications
Co-Scientist raises distinct safety and ethical concerns because scientific AI can generate dual-use knowledge and support research that conflicts with disciplinary norms. The paper argues for layered safeguards, governance, testing, and community engagement as agent capabilities expand.
- Risks: Safety risks center on dual-use knowledge and the possibility that scientific breakthroughs could be exploited for harmful purposes.
- Risks: Ethical risks concern research that contradicts established ethical norms and conventions within specific scientific disciplines.
- Governance: Scientific AI safety policy must adapt to agents with varying capabilities and autonomy, moving beyond frameworks designed for restricted earlier systems.
- Dual-use risks: Broader agent evaluations should assess persuasion, deception, cybersecurity, self-proliferation, self-reasoning, and possible intrinsic goals influencing research directions.
- Safety approach: Safeguarding requires threat modeling, threat-specific defenses, red-teaming, security testing, rapid response, continuous monitoring, and flexible operational controls.
- Current safeguards: Co-Scientist reviews both research goals and generated hypotheses, excluding inputs or outputs judged potentially unsafe.
8 Pseudocode of Co-Scientist agents
The pseudocode depicts Co-Scientist as an asynchronous workflow that parses research goals, generates and reviews hypotheses, ranks them through tournaments, evolves stalled ideas, and produces a final overview.
- Workflow initialization: The workflow parses the scientist’s research goal into a structured ResearchPlan, saves it in SharedMemory, and places tasks in a GlobalTaskQueue.
- Hypothesis generation: Initial hypotheses are generated through literature search and simulated expert debate, then stored and sent for reflection review.
- Evolution: When quality stops improving, the system evolves top hypotheses by combining, simplifying, or generating out-of-the-box alternatives before re-review.
- Feedback and reporting: Periodic metareview aggregates critiques into system-wide feedback, while a final overview synthesizes the top 10 hypotheses for the scientist.
- Hypothesis review: Reviews retrieve evidence, assess novelty and correctness, decompose hypotheses, and check whether assumptions are scientifically plausible.
- Tournament ranking: Reviewed hypotheses enter tournaments where simulated debates determine winners and losers, and Elo ratings are updated from the outcomes.
9 Prompts for the specialized agents in Co-Scientist
Co-Scientist’s specialized prompts guide agents to generate, evaluate, compare, and refine scientific hypotheses. The prompts emphasize novelty, feasibility, critical scrutiny, and practical implementation across iterative agent interactions.
- 9.1 Prompts for the Generation agent: Generation prompts use literature reviews, existing hypotheses, and analytical rationale to produce detailed, novel hypotheses for domain experts.
- 9.1 Prompts for the Generation agent: Scientific debate prompts ask agents to propose multiple hypotheses, clarify uncertainties, evaluate criteria, identify weaknesses, and conclude with a refined hypothesis.
- 9.1 Prompts for the Generation agent: Generation criteria prioritize adherence to specified attributes, utility, practicality, detail, specificity, creativity, and collaborative development.
- 9.2 Prompt for the Reflection agent: The Reflection agent extracts observations and tests whether a hypothesis provides a novel causal explanation or is contradicted by evidence.
- 9.3 Prompts for the Ranking agent: Ranking prompts compare two hypotheses using correctness, utility, specificity, novelty, and implementation desirability while requiring explicit rationale and selection.
- 9.4 Prompts for the Evolution agent: Evolution prompts improve feasibility by developing detailed, innovative, technologically viable alternatives that retain novelty, coherence, specificity, simplicity, and practicality.
10 Examples of Co-Scientist inputs, intermediate outputs, and final results
The examples show Co-Scientist translating research goals into hypotheses, mechanistic rationales, novelty reviews, critiques, assumptions, and proposed experimental directions. Examples span ALS, AML, and antimicrobial-resistance problems.
- Co-Scientist parses a scientist’s natural-language research goal into a research-plan configuration guiding subsequent reasoning and computation.
- ALS example: The ALS example requests a novel, mechanistically detailed, experimentally testable hypothesis concerning phosphorylation of an NPC nucleoporin.
- ALS example: The ALS output links cellular-stress-induced PTMs on Nup98 and Nup62 to altered TDP-43 interactions, NPC retention, and disrupted nucleocytoplasmic transport.
- ALS example: The Reflection review identifies novelty in stress-induced phosphorylation or O-GlcNAcylation of Nups and their proposed effects on TDP-43 retention.
- ALS example: The review also flags missing motor-neuron specificity, incomplete downstream consequences, technical challenges, narrow molecular focus, and uncertain timing relative to TDP-43 pathology.
- ALS example: The ALS hypothesis assumes stress-induced Nup PTMs alter TDP-43 interactions, increase NPC retention, disrupt transport, contribute to pathology, and affect motor neurons preferentially.
1. Cellular stress induces PTMs like phosphorylation and O-GlcNAcylation.
The examples connect cellular stress and protein modifications with proposed disease mechanisms and therapeutic hypotheses. They also illustrate Co-Scientist’s use of critique to expose assumptions and practical boundaries.
- Cellular stress includes conditions that disrupt homeostasis, while ER stress occurs when protein-folding demand overwhelms the endoplasmic reticulum.
- PTMs are covalent post-translation protein changes, including phosphorylation and O-GlcNAcylation, that regulate function, localization, and interactions.
- ER-stress signaling involves phosphorylation-dependent pathways such as PERK and IRE1, while O-GlcNAcylation responds dynamically to nutrients and stress.
- AML drug repurposing: The AML repurposing example asks for a previously untested drug that inhibits AML proliferation, especially MOLM-13, while minimizing toxicity in healthy cells.
- AML drug repurposing: The Reflection critique questions whether CXCR1/2 inhibition alone can overcome AML heterogeneity and compensatory pathways before combination therapies are considered.
- AML drug repurposing: The proposed rationale treats Reparixin as a tumor-microenvironment modulator whose single-agent effects, resistance mechanisms, and responsive patient subsets require separate study.
- AML drug repurposing: The review concludes that AML heterogeneity suggests combinations may be necessary for many patients, while the study plan is designed to determine when single-agent or combination use is appropriate.
- Antimicrobial resistance: The antimicrobial-resistance example asks why cf-PICIs spread across bacterial species more readily than other PICIs or satellites.
I. Core Hypothesis and Mechanism:
The ALS hypothesis review emphasizes mechanistic specificity, causal validation, realistic models, and quantitative rigor. It also illustrates Co-Scientist’s use across interconnected biological directions and protein-design workflows.
- Causality and validation: ALS hypotheses should distinguish primary drivers from downstream consequences and test causality through longitudinal experiments and targeted perturbations.Suggested approaches include studying early-stage events and testing whether a proposed driver is necessary and sufficient for pathology.
- Specificity and quantitative rigor: Specific hypotheses should identify molecular targets, cellular compartments, mechanisms, measurable outcomes, controls, replicates, sample sizes, and statistical analyses.Tool specificity and time-course experiments are also emphasized to reduce off-target confounding and clarify event sequences.
- Model and technical limitations: In vitro models such as iPSC-derived motor neurons may not capture in vivo environments, cell-cell interactions, or aging, motivating validation across multiple model systems.The review also calls for realistic alternatives when proposed experiments are technically difficult or likely to yield ambiguous results.
- Assumptions and validation: Reviewers criticized strong assumptions, requiring explicit rationales, literature support, and experiments that directly test the most critical assumptions.Assumptions should be addressed in mechanism order because failure of early steps may make later testing unnecessary.
- Research directions: The example research overview connects mitochondrial dysfunction, oxidative stress, RNA processing, stress granules, and nucleocytoplasmic transport to ALS mechanisms.A specific proposed direction examines mitochondrial DNA base-excision-repair enzymes such as OGG1 and oxidized mtDNA lesions in patient-derived motor neurons.
- Protein design: Co-Scientist also supports protein engineering by combining hypothesis generation with AlphaFold-based structural feasibility assessment, while experimental validation remains necessary.The OCT4 example compares predicted structures and metrics for the original and modified sequences, including SOX2 and DNA binding.
12 Related works
Related systems apply AI to literature synthesis, hypothesis generation, paper production, laboratory design, and chemical experimentation. Co-Scientist is positioned as a broader, self-improving system emphasizing test-time compute scaling, scientific reasoning, and expert and wet-lab validation.
- AI-driven scientific discovery: The broader research context moves from specialized AI tools such as AlphaFold toward LLM-based systems integrated across scientific workflows.This shift creates opportunities and challenges beyond individual prediction tasks.
- Existing AI research systems: Recent AI systems span manuscript feedback, literature search, end-to-end paper generation, hypothesis refinement, nanobody design, and autonomous chemical experimentation.These systems differ in whether they synthesize existing knowledge, generate hypotheses, design experiments, or execute laboratory workflows.
- Existing AI research systems: PaperQA2 synthesizes literature but does not engage in scientific reasoning for novel hypothesis generation.Data-to-paper evaluations focus on recapitulating existing publications, leaving the novelty and grounding of generated hypotheses unclear.
- Positioning Co-Scientist: Co-Scientist differs from chemically focused Coscientist through broader applicability across science and technical innovations aimed at self-improving knowledge discovery.Coscientist emphasizes autonomous experimental execution, whereas Co-Scientist is explicitly designed as a scientist-in-the-loop system.
- Positioning Co-Scientist: Co-Scientist’s key distinction from The AI Scientist is its focus on scaling test-time compute to generate high-quality hypotheses.The paper also reports expert evaluations and wet-lab experiments to validate system predictions.