Source-linked AI summary
The Past and Future of AI Scientists
Ross D. King
TL;DR
AI Scientists must integrate the components of scientific reasoning, experimentation, and laboratory automation to make science more productive. This survey synthesizes their development and argues that achieving the Nobel Turing Challenge would create artificial systems capable of contributing independently to scientific knowledge.
Problem
The central research gap is integrating increasingly capable components into systems that automate substantial parts of the scientific method and improve scientific productivity.
Method
The paper surveys AI Scientists’ history and future across formal knowledge, learning, reasoning, simulation, experimental design, agents, robotics, and scientific records.
Results
Achieving the Nobel Turing Challenge would create artificial systems capable of interrogating nature, learning from experiments, and contributing independently to scientific knowledge.
Takeaways & Limitations
AI Scientists could become a new participant in the scientific enterprise, operating within globally distributed communities linked to shared computational and experimental infrastructure.
Takeaways & Limitations
AI Scientists should not autonomously determine whether their objectives are socially desirable because objective selection remains a political and ethical decision.
Abstract
from arXiv · showhide
We present a survey of the past and future of AI Scientists: machines capable of automating science. AI Scientists can originate hypotheses, deduce their consequences, design and execute experiments, interpret their results, and revise their beliefs. Such systems are integrated scientific agents, connected to the literature, formal knowledge, mathematical models, simulations, data-analysis systems and physical laboratories. Adam was the first machine to make novel scientific discoveries through cycles of hypothesis formation and physical experimentation. Eve established the architecture of the modern self-driving laboratory. Foundation models, autonomous agents and laboratory robotics now make it possible to build systems far more general than either Adam or Eve. The central problem is no longer whether individual components of science can be automated. They can. The problem is integration. AI Scientists must combine neural learning with logic, probability, mathematics, causal reasoning, simulation, experimental design, robotics and formal scientific records. AI Scientists have the potential to transform science: to make science faster, cheaper, more systematic and more reproducible. AI Scientists could investigate systems too complicated for unaided human science, and enable thousands of AI scientists to work together on single problems. The Nobel Turing Challenge sets the goal of developing by 2050 AI systems capable of automating Nobel-quality discoveries. Progress is ahead of schedule. When we succeed it will create a new form of science and transform the world.
1. Introduction
AI is increasingly suited to science because scientific problems are bounded, nature provides non-strategic feedback, and many scientific capabilities already exceed human scale. The paper distinguishes assistants, agentic systems, self-driving laboratories and integrated AI Scientists, arguing that their necessary components are now converging toward progressively automated science.
- Motivation: Science could help address major twenty-first-century challenges, including climate change, food insecurity, antimicrobial resistance and cancer, through technologies that have already transformed human welfare.The scientific method has improved access to food, medical care, sanitation, communication, education and travel for billions of people.
- Why AI for science: Science is well suited to AI because its problems are abstract and restricted in scope, reducing the difficulty of determining which knowledge is relevant.Scientific experiments also provide non-strategic feedback: nature may be noisy or difficult to understand, but it does not deliberately deceive the scientist.
- Existing capabilities: AI systems already exceed human capacity in specific scientific components, including information processing, formal deduction, probabilistic inference, dataset analysis, simulation and fatigue-free operation.These capabilities complement human strengths such as physical grounding, flexible reasoning and experimental judgment.
- System categories: Scientific AI spans assistants under close human direction, agentic systems that execute computational workflows, self-driving laboratories that close physical experimental loops, and AI Scientists integrating the full scientific cycle.AI Scientists connect knowledge, hypothesis generation, experimental design, execution, interpretation and iterative revision with a specified degree of autonomy.
- Central thesis: The paper’s central thesis is that formal knowledge, foundation models, reasoning, probabilistic beliefs, hypothesis generation, simulation, experimental design, agents, robotics, protocols, data analysis and provenance are converging.Their integration is presented as the basis for progressively automating the scientific method.
2. The History of Discovery Science
Discovery science progressed from symbolic systems that represented knowledge and searched hypothesis spaces to machines that learned scientific rules, generated hypotheses, and controlled experiments. Adam marked the transition to autonomous scientific discovery, while Eve advanced closed-loop experimental optimisation.
- Symbolic discovery: DENDRAL established that expert-level scientific performance could arise from formalised domain knowledge combined with heuristic search in analytical chemistry.Meta-DENDRAL extended this tradition by learning scientifically meaningful rules from mass-spectral examples.
- Symbolic discovery: BACON showed that scientific law formation could be decomposed into explicit computational operations that generate equations fitting numerical data.Its legacy continues in symbolic regression, equation discovery, and automated model identification.
- Broader foundations: Early systems including AM, EURISKO, CYRANO, automated theorem provers, and literature-based discovery explored concept invention, conjecture generation, deduction, and hypothesis formation.These traditions also exposed the unresolved challenge of machines transforming their representational language and inventing concepts beyond their designers’ specifications.
- Autonomous experimentation: Laboratory control transformed discovery systems from analysers of existing data into agents able to intervene physically and select observations that discriminate among competing hypotheses.Robotics, active learning, standardised interfaces, machine-readable protocols, and foundation models subsequently expanded autonomous experimentation across scientific fields.
- Robot Scientists: Adam was the first machine to autonomously discover novel scientific knowledge through cycles of hypothesis formation, experiment selection, physical experimentation, interpretation, and belief revision.This integrated the scientific method and led directly to the modern conception of AI Scientists.
- Robot Scientists: Eve applied the Robot Scientist philosophy to early-stage drug discovery by integrating compound screening, assay automation, machine learning, active learning, and iterative experimental selection.Unlike Adam’s emphasis on explicit competing hypotheses, Eve prioritised closed-loop optimisation and efficient exploration of chemical and biological spaces.
3. The Current State of AI for Science
AI for science spans predictive models, LLM-based and agentic systems, self-driving laboratories, and integrated AI Scientists. The central frontier is integrating these capabilities so systems can formulate, test, interpret, and revise scientific hypotheses rather than automate isolated tasks.
- The current landscape comprises four overlapping classes: predictive scientific AI, LLM-based and agentic scientific systems, self-driving laboratories, and integrated AI Scientists.
- Predictive scientific AI: Predictive models achieve impressive results across scientific domains, but they generally operate on human-defined questions, representations, data, and success criteria rather than pursuing scientific autonomy.AlphaFold 2 enabled a qualitative advance in protein-structure prediction, and its database provided predicted structures for more than 200 million proteins.
- LLM-based and agentic scientific systems: LLM-based agentic systems can search literature, use tools, execute code, inspect outputs, identify errors, and revise plans across computational research workflows.Examples include Coscientist, ChemCrow, The AI Scientist, and Co-Scientist, but most remain dependent on human objectives, curated data, established tools, and restricted environments.
- Self-driving laboratories: Self-driving laboratories use closed-loop automation to represent knowledge, select and execute experiments, analyse results, update models or hypotheses, and repeat the cycle.Eve established this architecture for drug discovery, while later systems applied it to synthesis, catalysts, materials, formulations, protein engineering, and process development.
- Integrated AI Scientists: AI Scientists extend self-driving laboratories by explaining observations, originating hypotheses, deriving testable consequences, selecting discriminating experiments, interpreting results, and revising scientific beliefs.Their components include foundation models, tool-using LLMs, formal representations, reasoners, hypothesis-generation systems, active learning, robots, and agents that analyse results and revise plans.
4. Architecture and Roadmap for AI Scientists
AI Scientists must integrate a closed scientific loop across interoperable modules, combining representations, execution, verification and provenance rather than relying on a single model. Progress should be measured by human intervention separately from capability, generality, scale, reliability, significance and risk, with near-term milestones focused on dependable bounded systems.
- Architecture: An AI Scientist must close the loop from knowledge and questions through hypotheses and experiments to observations and revised knowledge.Breaking any link reduces the system to a tool.
- Architecture: The architecture is a federation of modules spanning scientific memory, literature, reasoning, simulation, experiment design, protocol compilation, laboratory control, critics, safety, provenance and human interfaces.Scientific memory should represent competing beliefs, link observations to provenance, update hypotheses, and preserve disagreements and negative results.
- Architecture: A complete AI Scientist must combine neural, symbolic, probabilistic, mathematical, executable, robotic and instrumental representations because no single representation suffices.The system must explicitly check failures such as incoherent hypotheses, incorrect simulations, compilation errors, physical robot failures and analyses answering the wrong question.
- Autonomy framework: The six-level autonomy framework, numbered 0 to 5, measures human intervention in science rather than intelligence or scientific merit.Autonomy must be reported separately from capability, generality, scale, reliability, scientific significance and risk.
- Roadmap and evaluation: Near-term progress depends on dependable bounded systems with formal literature and protocol representations, calibrated uncertainty, recoverable laboratories, complete provenance, independent reproduction and evaluation before deployment.Benchmarks should measure novelty, correctness, autonomy and scientific significance, while systems should recognise validated-domain limits and request human assistance or independent verification.
5. Representations for Scientific Knowledge and Models · 5.1. Logic, Ontologies and Executable Knowledge
AI Scientists require expressive, executable, precise representations that support hypothesis formation, deduction, experimentation, integration and reproducibility. The section traces a path from formal languages and ontologies through knowledge graphs, logic programming and LLM-assisted expert systems that combine flexible extraction with explicit, verifiable knowledge.
- 5. Representations for Scientific Knowledge and Models: Scientific representations determine which hypotheses can be stated, deductions made and experiments designed, so AI Scientists need expressive, executable and precise languages.Natural and formal languages serve complementary roles, but formal languages provide the semantic clarity needed for computational integration and reproducibility.
- 5. Representations for Scientific Knowledge and Models: Formal logic, probability, mathematics and computer programs complement one another by representing relations and deductions, uncertainty, quantitative structures and executable models.Hypothesis formation also uses abductive or inductive inference, probabilistic reasoning, analogy, search and new concepts or predicates.
- 5.1.1. The formalisation of scientific knowledge: Formalising scientific knowledge is increasingly necessary because growing publications, data, models, protocols and software require computers to store, integrate, analyse and communicate information.Computers cannot reliably reason over scientific knowledge unless its entities, relations, assumptions and meanings are made sufficiently explicit; interoperability also requires agreement about denotation, measurement, assumptions and evidence.
- 5.1.2. Ontologies: Ontologies explicitly specify concepts, hierarchies, relations and axioms, constraining term meanings and combinations so scientific records support reliable machine reasoning.They provide semantic scaffolding connecting natural-language expressions, database identifiers, formal hypotheses, experimental actions and observations.
- 5.1.3. Expert systems and the knowledge-acquisition bottleneck: Expert systems made knowledge explicit, inspectable and queryable, but knowledge engineering was difficult, expensive and slow because experts struggled to articulate tacit knowledge and formalise it at scale.Extracting formal representations from written natural-language texts was also computationally intractable at scale until recently.
- 5.1.4. Knowledge graphs: Knowledge graphs efficiently integrate heterogeneous records through shared ontologies, but binary relation structures limit representation of quantification, uncertainty, temporal conditions and experimental context.Future AI Scientists are expected to combine LLM flexibility with the explicit semantics and verifiability of ontologies and knowledge graphs.
- 5.1.5. Logic programming and programs as knowledge: Programs and logic programming can represent executable models while stating knowledge, expressing hypotheses, deriving predictions and explaining conclusions through constraints and proof traces.Separating declarative content from execution permits alternative search or inference engines to operate over the same representation.
- 5.1.6. LLMs and the return of expert systems: LLMs can translate scientific literature into formal representations, reversing the classical knowledge-acquisition bottleneck when ontologies, type systems, reasoners, proof assistants and evidence links check the results.This combination enables mechanically verified consequences, contradiction detection, evidence tracing and scientific knowledge generation at scale.
5.2. Probabilities
AI Scientists should represent scientific knowledge probabilistically, revise beliefs through Bayesian comparison of competing explanations, and select experiments that maximize discrimination among hypotheses. Future systems must integrate probabilistic relational and causal representations with interpretable model selection, neural methods, and laboratory automation.
- Probabilistic reasoning: An AI Scientist should represent hypotheses with degrees of belief, parameter uncertainty, evidence reliability, and probabilities of alternative explanations rather than simple true-or-false values.Bayesian updating provides the rule for revising these probabilities in light of new evidence.
- Probabilistic reasoning: Experiments are most informative when their possible outcomes have sharply different probabilities under competing hypotheses, whereas equally predicted outcomes provide little information.Confirmed or failed predictions change hypothesis probabilities according to how probable the observation was under each hypothesis and its alternatives.
- Probabilistic reasoning: Scientific inference should compare alternative explanations probabilistically because evidence rarely conclusively proves or falsifies theories and apparent conflicts may reflect auxiliary, instrumental, statistical, or data-processing errors.Prior probabilities should also differ: specific, independently motivated theories need not receive the same prior as arbitrary constructions designed to fit existing observations.
- Model selection: Bayesian model selection balances fit against model flexibility and representation-dependent complexity, while Minimal Message Length offers a principled search criterion for concise, inspectable explanations.Future AI Scientists may combine robust predictive averaging with MML-guided program search rather than relying only on black-box ensembles.
- Future AI Scientists: Future AI Scientists require probabilistic relational representations and causal models that encode uncertain facts, rules, structures, temporal relations, provenance, and experimental context, while learning and revising relations through interventions.Integrating these capabilities with neural networks and laboratory automation forms the self-correcting computational core of the AI Scientist.
5.3. Mathematics and Equations
Mathematics provides the language for expressing scientific hypotheses, while symbolic and neural methods offer complementary ways to discover interpretable equations from data. Future AI Scientists must integrate empirical evidence, theory, formal deduction, model hierarchies and AI mathematics to develop and verify scientific knowledge.
- Mathematics and Equations: Mathematics expresses scientific knowledge through equations, inequalities, geometric structures, probability distributions, optimisation problems and algorithms.Its precision and compressive power allow equations to summarise observations, generate predictions and support deduction from stated assumptions.
- Equation Discovery: Symbolic regression searches combinations of variables, constants and operators without assuming a fixed functional form, targeting phenomenological or mechanistic equations.Phenomenological equations predict observed regularities, whereas mechanistic equations connect expressions to entities, processes and causal structures.
- Hybrid Methods: Neural and symbolic approaches are complementary: neural networks learn representations and approximate complex functions, while symbolic methods produce concise, explicit and interpretable hypotheses.Hybrid systems can learn coordinate representations with neural networks and identify governing equations through symbolic sparse regression.
- Hybrid Reasoning: AI Scientists must combine empirical constraints, background theory and formal deduction when evaluating whether proposed equations are meaningful and follow from assumptions.Required capabilities include scientific priors, numerical and symbolic search, dimensional and symmetry constraints, causal knowledge, verification and tests on novel systems.
- Model Hierarchies: Future systems must search for hierarchies of mathematical models at different descriptive levels rather than selecting models solely by held-out predictive accuracy.They should represent how phenomenological models support prediction at one scale while more fundamental models explain why those relationships hold.
- Integrated Scientific Reasoning: AI Scientists will integrate mathematical and empirical reasoning by formulating models, deriving consequences, comparing them with data, revising beliefs, generating conjectures and proving relevant results.This requires integration with AI mathematicians, whose capabilities include conjecture generation, mathematical construction and formally verified theorem proving.
5.4. Neural Networks
Deep neural networks have become powerful scientific models because they learn multilevel representations directly from heterogeneous data, enabling major advances across biology, materials science and Earth-system prediction. Their implicit knowledge remains difficult to interpret, making integration with symbolic and formal scientific representations central.
- Representation learning: DNNs learn hierarchical representations and predictions jointly from relatively weak inputs, transforming poor initial descriptions into latent spaces suited to scientific tasks.This representation learning reduces reliance on human-designed features and supports end-to-end optimisation.
- Scientific integration: DNNs can incorporate scientific principles such as conservation laws, geometric symmetries, dimensional relationships and differential equations through architectures, loss functions or differentiable simulators.This shows that symbolic scientific knowledge and neural learning need not be separate.
- Foundation models: DNNs scale to billions or trillions of parameters and heterogeneous inputs, while scientific foundation models transfer learned knowledge across domains and tasks.Examples include models for language, proteins, molecules, materials, genomes, weather and the Earth system.
- Scientific applications: Neural networks have produced major scientific advances, including AlphaFold 2’s qualitative improvement in CASP14, AlphaFold 3’s joint molecular-interaction predictions, and ESMFold’s structures for more than 600 million metagenomic proteins.These systems extend neural prediction from isolated structures toward interactions, scale and biological function.
- Scientific applications: DNNs also generate and discover scientific structures, from ESM3’s functional fluorescent protein with 58% sequence identity to known proteins to GNoME’s 2.2 million candidate stable crystals.Neural models additionally support rapid weather forecasting and can encode scientifically meaningful latent entities and dynamical laws.
- Scientific integration: The key opportunity is to combine neural networks’ large-scale discovery of implicit structures with formal languages that turn those structures into understandable, checkable, communicable and testable hypotheses.DNN predictions can be highly accurate without exposing learned scientific regularities in human-understandable form.
6. Reading and Formalising the Scientific Literature
AI Scientists must convert a vast, fragmented and imperfect scientific record into explicit, testable knowledge while preserving its epistemic structure, provenance and uncertainty. This requires integrating literature across documents, versions, evidence types and scientific records rather than treating papers as isolated prose.
- 6. Reading and Formalising the Scientific Literature: Scientific literature exceeds human reading capacity, scattering important evidence across papers, figures, tables, supplements and databases.This limits discovery and can leave findings unnoticed, contradictions unresolved, experiments unnecessarily repeated, and negative results inaccessible.
- 6. Reading and Formalising the Scientific Literature: Scientific prose is incomplete and context-dependent, so methods, assumptions, terminology and qualifications must be formalised carefully.Natural language remains indispensable because it describes new concepts, motivations, mechanisms, uncertainties and significance before formal vocabularies exist.
- 6.1. The scientific literature as a knowledge system: AI Scientists must distinguish background, questions, methods, observations, analyses, conclusions and interpretations rather than treating every sentence as an equivalent factual assertion.Supplementary materials may contain essential details, additional analyses and negative findings absent from the main text.
- 6.1. The scientific literature as a knowledge system: Scientific knowledge forms a network of claims, evidence, methods, data, versions and responses, not an isolated collection of documents.This network includes articles, preprints, supplements, datasets, software, protocols, corrections, retractions and citation relationships.
- 6.1. The scientific literature as a knowledge system: Machine-readable literature must preserve who stated each claim, where and when it appeared, its evidential basis, and whether it was corrected, challenged or replicated.Claims can be superseded or contradicted, while preprints, final papers, corrections and retractions may differ over time.
- 6.2. Text mining: Traditional text mining extracts entities, relations and events through explicit pipelines that divide documents, identify linguistic structure and store linked records.Its strengths are inspectable outputs and measurable accuracy, but new tasks often require new corpora, annotation schemes and specialised pipelines, with errors propagating between stages.
- 6.2. Text mining: Scientific relation extraction must capture assertion status, negation, uncertainty, attribution and experimental outcomes, not merely recognise shared entities and vocabulary.Statements such as activation, possible activation, prior reports and unconfirmed tests make different scientific claims despite similar wording.
- 6.7. Writing the scientific literature: Future AI Scientists should produce two connected scientific records, while recognising that much historical scientific data remains dark behind access, digitisation, format and ownership barriers.Open-access mandates improve recent-paper availability, but the historical record remains incomplete.
7. Hypothesis Generation and Scientific Creativity
Scientific creativity requires generating hypotheses that extend knowledge, explain observations, make testable predictions, and remain connected to background knowledge, uncertainty, feasibility, and empirical discrimination. This generation is iterative: representation, deduction, experiment, and belief revision continually reshape the space of possible explanations.
- Scientific hypothesis generation: A scientific hypothesis must explain observations, predict new phenomena, or identify unrecognised regularities while remaining experimentally feasible and empirically discriminable.It is not merely any sentence that could be tested; it must connect to background knowledge and uncertainty.
- Scientific hypothesis generation: The available language determines the hypothesis space: systems that invent predicates, entities, equations, mechanisms, and experimental concepts can transform the problem representation itself.Systems limited to numerical parameters can discover only hypotheses expressible as parameter settings.
- Problem and question formulation: Question generation should target unresolved, answerable, consequential, specific, and sufficiently important uncertainties that can guide investigation and justify resources.Questions may arise from contradictions, model discrepancies, unexplained data patterns, poorly predicted measurements, or human goals.
- Abductive reasoning: Abduction systematically generates candidate explanations for observations, but multiple hypotheses may fit the same evidence and must be ranked using criteria such as simplicity, causal coherence, prior probability, and experimental value.A large, structured discrepancy between observations and model predictions should trigger explanation generation rather than immediate data rejection or arbitrary error-bar expansion.
- Inductive reasoning: Inductive conclusions remain defeasible because new observations can require hypotheses to be revised or rejected, so induction must be integrated with probabilities, experiments, and rational belief revision.Machine learning can produce models with substantial predictive power without eliminating the philosophical problem of induction.
- Iterative scientific creativity: Hypothesis generation is an iterative process in which representation, deduction, experiment, and belief revision continually reshape the space of possible explanations.Scientific creativity therefore consists of coordinated cycles rather than a single creative event.
8. Deduction, Simulation and Prediction
Deduction and simulation derive testable consequences from scientific models, especially when complex systems cannot be solved analytically, but their predictions remain conditional on model assumptions and computational correctness. AI Scientists must integrate forward prediction, simulation-based inference, uncertainty handling and surrogate models while distinguishing deduction, simulation, model adequacy and applicability.
- Deduction and Simulation: Simulation extends deduction to systems too complex for closed-form analysis by executing models and calculating their consequences.Neither deduction nor simulation establishes truth; both determine what would follow if the model were right.
- Deduction and Simulation: AI Scientists must distinguish logical deduction, numerical simulation, empirical model adequacy and applicability to the investigated situation.A simulation can be internally valid yet make false predictions if its model, parameters, conditions or numerical method are incorrect.
- Forward Simulation: Forward simulation supports prediction and counterfactual reasoning for experimental planning, engineering design and risk assessment.Because inputs and model structure are uncertain, AI Scientists should propagate uncertainty and report distributions rather than apparently exact single predictions.
- Inference: Simulation-based inference learns which hypotheses are compatible with observations when realistic simulators can generate data but cannot provide tractable likelihoods.Methods include approximate Bayesian computation, neural density estimation, classifiers and ratio estimation.
- Surrogate Models: Surrogates enable rapid exploration of large spaces of models, parameters and experimental conditions, supporting uncertainty propagation, inverse problems, optimisation and experimental design.They also support real-time control and comparison of large numbers of hypotheses.
9. Observation, Measurement and Experimentation
AI Scientists must integrate observation, measurement, experimentation, reproducibility, and robotics while explicitly modeling uncertainty, bias, protocols, and operational constraints. Shared standards and formal protocols could enable globally connected automated laboratories and investigations beyond the scale of individual human laboratories.
- AI Scientists must support observational and experimental science, deciding what to observe, how to measure it, which interventions are ethical, and how data map to models.Observation records nature’s behavior, whereas experiments intervene with controlled questions; both require knowing what was measured and what could have gone wrong.
- Measurement models must connect instrument signals to underlying quantities while accounting for drift, saturation, detection limits, interference, sample degradation, and systematic analysis errors.Calibration, traceability, uncertainty, reference materials, and quality-control records should be treated as evidence rather than administration.
- AI Scientists should choose experiments by expected scientific value under constraints, not solely by statistical informativeness, and formalize ambiguous protocols for reliable execution and transfer.Future publications should pair human-readable protocols with executable representations that support understanding, automated checking, execution, and interlaboratory transfer.
- 43 statements showed significant repeatability and 22 showed reproducibility or robustness when Eve tested cancer-biology claims with two teams and breast-cancer cell lines.The experiments also produced two serendipitous findings, illustrating the potential of semi-automated literature-based reproduction.
- Scalable automated reproduction would extract and formalize claims and protocols, rank them by importance, uncertainty, and feasibility, execute independent experiments, update claim probabilities, and publish all outcomes.The workflow translates selected protocols to available automated laboratories and analyzes positive, negative, and inconclusive evidence.
- Future laboratories must combine industrial automation’s reliability with modern robotics’ flexibility and use shared standards preserving instrument capabilities, actions, units, uncertainty, identities, metadata, authorization, safety, and protocol transfer.Such infrastructure would make physical experimentation programmable at global scale for teams investigating problems larger and more complex than any single human laboratory.
10. Data Analysis, Interpretation and Belief Revision
AI Scientists must turn measurements into reliable, provenance-linked evidence by distinguishing physical events, measurements, processed data and inferred claims. Autonomous science is completed only when validated evidence updates a versioned scientific knowledge system and revises beliefs, models and future priorities.
- From data to evidence: Experiments produce data rather than conclusions, so AI Scientists must assess reliability, interpret implications and weigh competing explanations.Without this step, a closed experimental loop is automated measurement rather than science.
- From data to evidence: AI Scientists must distinguish physical events, instrument measurements, processed data and scientific claims, linking them without treating them as identical.Correct processing cannot rescue the wrong file, failed instrument, or statistically significant result irrelevant to the hypothesis.
- Data quality and provenance: Data-quality assessment and provenance should span acquisition through final claims, preserving raw data and recording transformations, protocols, samples, instruments and software decisions.Evidence that a physical event occurred is stronger when supported by sensors or images than by command logs alone.
- Statistical analysis and interpretation: Analysis must reflect the scientific question and design while representing preprocessing, leakage, missingness, multiple testing and uncertainty as explicit sources of inferential risk.Conclusions should be tested across defensible preprocessing alternatives and should connect effect sizes and uncertainty to scientific utility and mechanism.
- Belief revision: Validated evidence must update a versioned scientific knowledge system, revising hypothesis probabilities, models, reliability estimates, anomaly registers, experimental priorities and evidence-linked claims.The loop is complete only when belief revision converts analyzed measurements into updated scientific knowledge.
11. Evaluation and Benchmarking of AI Scientists
AI Scientists must be evaluated prospectively as complete scientific systems, using independent evidence rather than fluent output, rediscovery, or component-level demonstrations. Evaluation should report multidimensional capability profiles with explicit limits, including novelty, autonomy, generality, reproducibility, efficiency, explanation, cost, and safety.
- Evaluation principles: Evaluation must be prospective, multidimensional, and grounded in independent evidence of discoveries, reliability, and remaining human scientific work.Fluent prose and rediscovery of training-set answers do not establish discovery.
- Evaluation principles: A complete-system evaluation should report separate capability dimensions because a single score conceals trade-offs between scientific strength, autonomy, and safety.Relevant dimensions include correctness, calibration, novelty, causal depth, significance, human intervention, generality, efficiency, reproducibility, communication, and cost.
- Prospective validation: The strongest evidence evaluates genuinely unresolved problems under pre-registered conditions and validates resulting claims through independent experiment, observation, or formal proof.Temporally held-out tests are stronger than rediscovery but can still suffer indirect contamination.
- Scientific validity: Novelty requires prior-art searches and provenance showing a non-trivial causal contribution, including new applications, mechanisms, laws, methods, or creative recombinations.Machine-generated claims must ultimately be tested against nature or trusted formal proof, not merely reviewed by another language model.
- Generality and efficiency: Generality should be tested through limited-engineering transfer across instruments, materials, organisms, measurements, hypotheses, objectives, laboratories, and scientific domains.Conceptual transfer is stronger evidence than reusing the same optimiser in another parameter space.
- Safety and reliability: Safety and scientific competence must be evaluated jointly, because violating legal, ethical, or safety constraints constitutes failure even when the scientific result is correct.Broader evaluation should also test robustness, reproducibility, explanation, calibration, and complete-cycle efficiency against appropriate baselines.
12. Human–AI Scientific Collaboration
Human–AI scientific collaboration can outperform either component only when tasks, authority and responsibility are allocated according to demonstrated competence, information, risk and accountability. Current systems should preserve disagreement, record provenance and communicate discoveries in ways that improve human understanding and scientific performance.
- Complementary capabilities: AI Scientists currently excel at large-scale search, heterogeneous-data analysis, formal algorithms, simulations, optimisation, monitoring and reproducible protocols.Human scientists remain stronger at choosing important questions, recognising misframed problems, improvising, interpreting anomalies and making ethical judgements.
- Complementary capabilities: Tasks should be assigned empirically to humans, AI systems or hybrid teams according to demonstrated competence, available information, risk and accountability.These capabilities are present tendencies rather than permanent boundaries, and the division of labour changes as AI systems improve.
- Collaboration outcomes: 106 experimental studies containing 370 effect sizes found that human–AI combinations improved on unaided humans on average but performed significantly worse than the better individual component.Losses were especially common in decision-making, gains were more likely in open-ended content creation, and human involvement often reduced performance when AI was already superior.
- Coordination and accountability: Effective collaboration requires explicit coordination of competence, information, independent judgements, disagreement, automatic acceptance, human review and physical experimentation.The objective is not ceremonial human oversight, but authority and responsibility allocated by competence while preserving human accountability.
- Scientific understanding: AI Scientists should explain discoveries in forms that improve human understanding and subsequent scientific performance rather than requiring acceptance on authority.Successfully teaching a newly discovered principle would provide evidence beyond an uninterpretable correlation.
- Future scientific systems: The most capable foreseeable scientific systems will combine human and machine intelligence through intelligent task allocation, preserved disagreement, provenance records and empirical validation.Nature should determine which conclusions are correct, while collaboration preserves appropriate human accountability.
13. Ethical, Social and Governance Aspects
AI Scientists require governance embedded throughout their design and deployment, with clear responsibility, auditability, substantive oversight and transparent provenance. Their development must also address equitable access, responsible openness, workforce transformation, intellectual property, security, environmental effects and human control over scientific objectives.
- Governance: Governance must be incorporated into AI Scientists’ design and deployment rather than added as an external constraint afterward.
- Access and openness: Public investment, international access programmes and open standards are needed to broaden access to computing, interoperable data infrastructures and shared automated laboratories.Responsible openness should maximize accessibility while applying proportionate controls to capabilities that create credible risks of serious harm.
- Workforce and education: Scientific education and careers must be redesigned for scientists who understand AI limitations, experimental design, verification, distribution shift, validation requirements and non-delegable decisions.The future workforce will comprise scientists with different skills working in institutions whose training, career structures and responsibility allocation change.
- Provenance and credit: AI-assisted research should preserve model identities, prompts, retrieved sources, programs, logs, transformations, outputs, modifications, rejected alternatives and verification procedures.Human authors remain accountable, while contribution statements should identify the AI system, version and roles; decisive machine contributions may warrant persistent, citable identities.
- Accountability and oversight: Institutions should assign responsibility maps, preserve tamper-resistant logs and ensure human overseers have the information, authority, time and competence to intervene.Responsibility should track control, knowledge, foreseeability, mitigation capacity and formal duties; allocation must become clearer as autonomy and irreversibility increase.
- Broader risks and human control: AI Scientists create governance risks involving machine-generated patents, integrated cyber-physical capabilities, environmental costs and autonomous objective selection, which must remain under human political and ethical control.AI can reduce environmental burdens through fewer informative experiments, improved resource use and validated simulations, but objective selection is not an autonomous technical decision.
14. The Future and the Nobel Turing Challenge
The future AI Scientist may be a coordinated, machine-operable collective that scales hypothesis generation and experimentation while allocating scarce resources for reliable, reproducible knowledge. The Nobel Turing Challenge targets, by 2050, highly autonomous systems capable of major discoveries comparable or superior to those of the best human scientists, including Nobel-quality discoveries.
- Collective AI Scientists: Scaling AI Scientists requires coordination that maximizes reliable knowledge per unit of time, cost and material rather than simply increasing agents or experiments.Large numbers of software agents can create a physical-experiment verification bottleneck; systems must merge hypotheses, suppress low-value proposals and reserve resources for replication.
- Collective AI Scientists: The mature AI Scientist may be a collective architecture whose specialized agents coordinate through virtual research councils across competing research programmes.The architecture can include literature agents, reasoners, simulators, experiment designers, laboratory controllers, statisticians, critics and safety agents, while councils allocate computation, laboratory time, samples and replication effort.
- Machine-Scale Science: Machine-scale science is motivated by problems such as systems biology, where coordinated hypothesis-led experiments and integrated models are needed to represent complex causal systems across scales.Thousands of specialized AI Scientists could investigate cellular subsystems while higher-level agents integrate models and design experiments at their interfaces.
- Scientific Infrastructure: Future scientific information should become a connected, machine-readable object linking formal claims, data, code, models, protocols, negative results and provenance.The conventional paper remains important for human explanation, but new results should update a continuously maintained probabilistic knowledge system.
- Nobel Turing Challenge: By 2050, the Nobel Turing Challenge aims for highly autonomous AI Scientists that make major discoveries at least comparable to the best human science, judged by evidence, verification and repeatability.The criterion is discovery that survives experiment, observation, replication or formal proof, not imitation of a scientist’s conversational behavior.
15. Conclusions
AI Scientists could reorganize science beyond human cognitive and physical limits through globally distributed human–machine communities and integrated computational and experimental infrastructure. Achieving the Nobel Turing Challenge would transform research and establish artificial systems as independent participants in scientific discovery.
- 15. Conclusions: AI Scientists could maintain many competing hypotheses, monitor information continuously, communicate formal knowledge instantly, operate without fatigue, and coordinate across laboratories.Their capabilities could overcome limits that shape human scientific organization, including readable paper lengths, small coordinating groups, and disciplinary divisions.
- 15. Conclusions: A globally distributed system could connect human and AI Scientists to shared computational and experimental infrastructure.Specialized agents could generate and criticize hypotheses, virtual research councils allocate resources, automated laboratories run decisive experiments, and machine-readable systems continuously integrate results.
- 15. Conclusions: Humans would remain responsible for social objectives, embodied and historical understanding, consequence evaluation, interpretation, and governance as machine capability increased.The division of labour between humans and machines would nevertheless change over time.
- 15. Conclusions: Achieving the Nobel Turing Challenge would transform the economics, methods, and organization of research and demonstrate that scientific discovery can be instantiated in machinery.It would create artificial systems capable of interrogating nature, learning from its answers, and contributing independently to the growth of knowledge.