Source-linked AI summary
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence
Ling Yang, Zhenfei Yin, Yingcheng Wu
TL;DR
The paper addresses the gap between solving human-specified tasks and participating in the construction and revision of research problems and knowledge. It formulates Discovery Foundation Models around revisable research states, instantiates them with Zetema, and grounds them with GALILEO; the framework treats discovery as trainable and evaluable through validated progress and transferable improvement.
Problem
Existing pipelines reward competence after the research structure is fixed, limiting systems when discovery requires revising problems, representations, explanations, or evaluators.
Method
The paper defines DFMs and a Discovery Process over revisable research states, instantiates Zetema, and formulates training and process-centered evaluation for discovery operations.
Results
The framework is empirically grounded by GALILEO, where external wet-lab outcomes revise subsequent discovery decisions across experimental rounds and are consolidated into a reusable design rule.
Takeaways & Limitations
Discovery is presented as a learnable, executable, and evaluable capability involving externally validated progress and transferable improvement beyond final-answer performance.
Takeaways & Limitations
GALILEO does not establish that the full general-purpose DFM problem is solved because its domain, modalities, interfaces, and objectives remain substantially structured, with cross-domain transfer open.
Abstract
from arXiv · showhide
Foundation models have progressed from learning and reasoning over existing knowledge, to increasingly learning through action, tool use, and outcome feedback. We argue that the next frontier is a further transition: from solving and acting within problems specified by humans to participating in the process by which new problems, representations, explanations, and knowledge are created. We refer to this capability as Discovery Intelligence. We formulate Discovery Foundation Models (DFMs) as general-purpose model systems for open-ended discovery. A DFM operates over a revisable research state and supports seven coupled capabilities spanning problem discovery, formulation, representation construction, hypothesis formation, intervention, evidence-grounded revision, and continual discovery improvement. We instantiate this framework with Zetema, which couples explicit research-state dynamics, verification and experimental gating, external grounding, and cross-task Discovery Skill evolution. We further ground the framework with GALILEO, a real therapeutic-discovery system in which Dry-Lab reasoning, robotic and hands-on Wet-Lab experimentation, external biological evidence, and iterative hypothesis and design revision form a closed physical discovery loop. We then formulate a unified approach to capability formation and process-centered evaluation, enabling discovery behavior to be trained, improved, and measured beyond final-answer performance. Together, these components establish discovery as a learnable, executable, and evaluable capability of foundation-model systems. We view this shift as a broader progression in intelligence scaling: from learning over existing knowledge, to learning from action outcomes, and ultimately to participating in the construction, testing, and revision of the structures through which new knowledge is discovered. Code: https://github.com/Gen-Verse/DFM-Plans
1. Introduction
The paper argues that foundation models remain limited when research requires revising the problem structure itself, and introduces Discovery Foundation Models for open-ended discovery. It instantiates this framework with Zetema and grounds it in GALILEO, while proposing training and evaluation over research-state transitions.
- Motivation: Foundation models are strong at solving supplied tasks but weaker when progress requires changing the question, variables, representation, or evaluator.Open-ended discovery may require identifying what is unknown, reformulating the problem, and selecting interventions that distinguish competing explanations.
- Discovery Intelligence: Discovery Intelligence concerns constructing, testing, and revising the process through which a partially understood world becomes validated knowledge.It extends beyond hypothesis generation to deciding which parts of knowledge production are fixed inputs and which become objects of model action and revision.
- DFM Framework: DFMs identify valuable unknowns, formulate researchable problems, revise representations, form testable explanations, design interventions, update research states from evidence, and improve across tasks.These capabilities distinguish open-ended discovery from optimization over a predefined task.
- Zetema: Zetema operationalizes the framework through explicit research-state revision, branching and rollback, verification-gated actions, external grounding, and validated cross-task Discovery Skill evolution.Its organization makes the capability operational without requiring one monolithic model or maximal autonomy.
- GALILEO: GALILEO demonstrates a real Dry-Lab/Wet-Lab loop where physical biological feedback revises scientific decisions and is distilled across five optimization rounds into a reusable design rule.The case connects multi-omics-informed target nomination and peptide design with robotic synthesis, multimodal phenotyping, and orthogonal hands-on assays.
- Capability Formation and Evaluation: Training targets intermediate research decisions, while evaluation measures externally validated progress and future discovery improvement under matched resources and retrieval controls.The formulation covers problem formulation, representation, hypothesis construction, intervention, falsification, and verification.
2. From Generalist Problem Solving to Discovery Intelligence
The paper distinguishes ordinary generalist problem solving within fixed research structures from Discovery Intelligence, which must revise those structures when evidence exposes their inadequacy. Science provides a capability-forming environment because incomplete specifications, costly interventions, delayed outcomes, and external observations make these decisions observable.
- Fixed Research Structures: Most current systems optimize within a human-supplied question, variable set, objective, tool interface, and evaluator rather than revising that structure.The paper identifies this as the specific limitation motivating Discovery Intelligence, not a lack of scientific knowledge or reasoning.
- Structural Ceiling: A fixed representation cannot express omitted variables, and an evaluator that ignores a property cannot recover it through more search or samples.Automation may therefore pursue a misframed question more efficiently without recognizing the misframing.
- Evidence and Intervention: Scientific discovery exposes these structural decisions because systems must act under partial observability and revise after outcomes contradict predictions.Perturbations, counterexamples, boundary tests, simulations, or replications can separate explanations left equivalent by observational data.
- Capability Formation: Capability-forming environments differ from additional text pretraining because they require consequential research decisions under incomplete specification and observations the model cannot freely choose.Such environments may include codebases, formal systems, causal simulators, robotic platforms, laboratories, or human-mediated processes.
- Evaluator Incompleteness: Evaluator incompleteness can itself become part of the research state when leakage, simulator artifacts, non-reproducibility, or plausibility create apparent progress without stronger knowledge.The paper also notes that AI may change which problems are pursued, not only how quickly they are solved.
- Discovery Transitions: Discovery spans framing, modeling, and grounding-and-revision transitions that turn unknowns into researchable problems, mechanisms into expressible representations, and evidence into updated research objects.The seventh capability transfers validated experience across tasks as Discovery Skill, while the first six operate within an investigation.
3. Discovery Foundation Models: Defining Discovery Intelligence
Discovery Foundation Models are defined by the research structures they can construct and revise, rather than by a particular architecture or autonomy level. They extend scientific AI from solving fixed tasks toward evidence-grounded discovery processes that can change problems, representations, interventions, and validation.
- 3.1. Problem Setting and Formal Definition: Unlike conventional tasks with supplied problems, representations, objectives, tools, and evaluators, discovery allows these structures to change during an investigation.
- 3.1. Problem Setting and Formal Definition: DFMs are general-purpose model systems that identify valuable unknowns, formulate problems, construct and revise representations, generate testable explanations, design interventions, learn from external evidence, and improve across tasks and domains.
- 3.3. System Boundary: Scientific claims require domain-appropriate external evidence, transferable operations require validation beyond one episode, and attribution must record decisive human interventions or unequal representations.
- 3.2. Discovery Capabilities: The seven coupled capabilities comprise finding and formulating problems, constructing representations, forming hypotheses, designing interventions, revising from evidence, and continually improving discovery behavior.
- 3.3. System Boundary: A DFM is evaluated as an integrated model system containing policy, research state, memory, tools, environment, validation mechanisms, and human oversight, with roles defined functionally rather than architecturally.
- 3.4. Relation to Existing Scientific AI Paradigms: Search or physical execution alone is insufficient: DFM status depends on recognizing inadequate search spaces or evaluators and revising the relevant research structure accordingly.
4. Discovery Process: Operationalizing Discovery Intelligence
The Discovery Process treats discovery as revision over an evolving research state rather than a fixed sequence of textual stages. It operationalizes progress through problem formulation, representation, competing explanations, interventions, evidence-based revision, and validated knowledge.
- Discovery Process: Discovery advances from valuable unknowns through researchable problems, representations, explanations, interventions, external evidence, and validated knowledge, while evidence can return the investigation to earlier objects.The research state records both epistemic progression and the revision paths that produce it.
- 4.1. Problem Discovery and Formulation: Problem formulation determines which evidence counts, which interventions are feasible, and where resulting knowledge is expected to apply.A selected unknown is framed using the object, scope, scale, conditions, observables, and unresolved relation or mechanism.
- 4.2. Representations and Competing Explanations: Representations are revised by adding latent variables, removing misleading proxies, changing scale, separating regimes, revising ontologies, or translating formal structure.A representation is useful only when it changes downstream prediction, intervention, regime separation, mechanism compression, or transfer.
- 4.2. Representations and Competing Explanations: Candidate explanations encode mechanisms, assumptions, validity ranges, predictions, and falsifiers, and they are useful only when feasible conditions can make them disagree.If hypotheses remain observationally identical under every feasible action, the bottleneck may be the representation or measurement interface rather than hypothesis generation.
- 4.3. Interventions and Evidence: Interventions are chosen for expected uncertainty reduction while accounting for feasibility, cost, risk, statistical power, measurement quality, and inconclusive outcomes.Experiments, simulations, code execution, ablations, counterexamples, alternative measurements, and replications can reveal shared false assumptions or unreliable measurements, not only confirm hypotheses.
- 4.3. Interventions and Evidence: Evidence-grounded revision updates implicated problems, representations, hypotheses, intervention plans, or uncertainties only after provenance, calibration, protocol fidelity, leakage, and replication are assessed.Failure attribution distinguishes theory failure, measurement error, implementation error, protocol deviation, confounding, noise, environmental shift, and simulator misspecification.
- 4.4. Episode Output and Candidate Discovery Lessons: A completed episode yields both domain knowledge and candidate lessons about the discovery process, but lessons require preserved states, alternatives, actions, observations, and provenance for validation.Zetema is introduced as the mechanism for maintaining this state and controlling which lessons influence future discovery.
5. Zetema: A System Instantiation of Discovery Intelligence
Zetema instantiates Discovery Intelligence as an explicit, inspectable, and revisable research-state system coupling within-task reasoning, evidence-based gating, experimentation, and cross-task skill evolution. GALILEO demonstrates the narrower but concrete case of physical evidence changing research actions and yielding transferable discovery behavior, while broader general-purpose transfer remains open.
- 5. Zetema: A System Instantiation of Discovery Intelligence: Zetema couples research-state revision, evidence-based action gating, external experimentation, and cross-task Discovery Skill evolution through functional interfaces rather than a mandatory neural architecture.Foundation models, tools, simulators, verifiers, laboratories, and human researchers can implement different parts of the organization.
- 5.1. Research-State Dynamics: Every scientifically consequential operation in Zetema leaves an inspectable state transition, preserving problems, representations, hypotheses, evidence, interventions, uncertainties, budgets, and reusable experience.This state makes revision, branching, attribution, and later training possible.
- 5.2. Research Operations: Zetema can split problems, introduce variables, change scale, replace representations, construct adversarial explanations, revise assumptions, request replication, terminate branches, and maintain parallel interpretations of failure.Branches can remain associated with diagnostic actions until evidence separates missing-variable, dataset-shift, or implementation-error accounts.
- 5.4. Research World Model and Experimental Gating: World-model error affects resource selection: conservative screening can suppress valid interventions, while overconfident screening can repeatedly favor actions implied by misspecification.Zetema retains disagreement, calibration, validity ranges, and out-of-distribution signals; high uncertainty can trigger bounded pilots and repeated prediction failures update screening.
- 5.7. Empirical Case Study: A Real Dry-Lab/Wet-Lab Discovery Loop: GALILEO places biological outcomes inside the decision loop, where physical feedback updates target beliefs, motif-policy weights, assay choices, mechanism hypotheses, and subsequent interventions.The capability claim belongs to the coupled system of policies, tools, experimental environments, validators, and human execution, not to the robotic platform alone.
- 5.7. Empirical Case Study: A Real Dry-Lab/Wet-Lab Discovery Loop: GALILEO-LRC and GALILEO-SLC followed distinct orthogonal evidence routes, with organoid IC50 values of 18.922 𝜇M and 0.736 𝜇M, respectively, while later experiments refined their mechanisms.The LRC branch involved selective LRRC8A/LRRC8C current blockade; the SLC branch involved mitochondrial localization and evidence toward citrate-export-dependent metabolic collapse.
- 5.7. Empirical Case Study: A Real Dry-Lab/Wet-Lab Discovery Loop: Across five optimization rounds, weak motif branches were pruned, abandoned organizations were reopened under revised hypotheses, and productive modules were consolidated into the reusable Amphiphilic Balance Grammar.ABG transferred episode-level physical feedback into a more portable discovery operation, distinguishing continual discovery improvement from static candidate ranking.
- 5.7. Empirical Case Study: A Real Dry-Lab/Wet-Lab Discovery Loop: GALILEO demonstrates evidence-driven state revision in a real physical workflow but does not establish that the full general-purpose DFM problem or broad cross-domain transfer has been solved.Its domain, molecular modality, experimental interfaces, and scientific objectives remain substantially structured.
6. Capability Formation: Training Discovery Operations
Capability formation trains discovery as state-conditioned research behavior rather than final-answer production, using trajectories, interactive environments, multiple learning signals, verification, world models, and validated skill transfer.
- Capability Formation: Training Discovery Operations: Discovery training targets state-conditioned research operations and explicit state transitions, not merely polished scientific answers.Records preserve available operations, external observations, and how evidence changes the research state.
- 6.1. Training Data and Interactive Research Environments: Useful trajectories retain malformed questions, missing variables, misleading representations, failed replications, contradictory evidence, abandoned branches, and attribution errors.Failure sources must remain visible because low statistical power and non-identifying experiments require different updates.
- 6.1. Training Data and Interactive Research Environments: Historical and synthetic trajectories provide complementary signals but require provenance because retrospective records omit alternatives and generated environments inherit their assumptions.Provenance is treated as part of each trajectory rather than flattened into one demonstration format.
- 6.1. Training Data and Interactive Research Environments: Interactive environments make the model choose observations or tests while exposing external outcomes across causal worlds, simulators, codebases, robotic platforms, and laboratories.They should hide meaningful structure while retaining evaluable consequences and, in simulators, counterfactual alternatives.
- 6.1. Training Data and Interactive Research Environments: Training curricula should vary uncertainty and realism from static data through simulation, digital twins, shadow-mode planning, and supervised physical experiments.Difficulty-aligned co-evolution should vary hidden mechanisms and uncertainty while preserving independently evaluable outcomes.
- 6.2. Learning Objectives and Scientific Feedback: The system is represented through a research policy, state updater, Research World Model, process verifier, and Discovery Skill selector that receive distinct training signals.Training records connect states, operations, external observations, resulting states, and process, validity, cost, preference, or transfer labels.
- 6.2. Learning Objectives and Scientific Feedback: Learning objectives combine stage-weighted discovery imitation, state-transition learning, preference optimization, process verification, world-model prediction, and reinforcement learning from scientific feedback.Discovery roles include finding, formulation, representation, hypothesis, intervention, and revision; preferences can encode testability, information gain, evidence, cost, robustness, or risk.
- 6.2. Learning Objectives and Scientific Feedback: The verifier filters malformed or unsupported transitions, while the Research World Model predicts observations, state changes, costs, and risks before expensive interaction.Scientific feedback rewards externally supported knowledge progress, information gain, validation quality, and transfer while accounting for cost and risk.
7. Capability Evaluation: A Process-Centered Protocol
The evaluation protocol measures discovery through externally validated knowledge progress and future capability improvement, while profiling research operations, efficiency, transfer, and controls against memorization and resource confounds.
- 7. Capability Evaluation: A Process-Centered Protocol: Discovery progress is evaluated on two separate axes: externally validated current knowledge progress ΔKval and improvement in future discovery capability ΔCfuture.The outcome must be assessed using evidence the proposing system does not control.
- 7. Capability Evaluation: A Process-Centered Protocol: Future capability concerns recognizing malformed questions, constructing useful representations, choosing discriminating interventions, responding to counterevidence, and allocating resources to bottlenecks.Additional factual knowledge alone is insufficient.
- 7. Capability Evaluation: A Process-Centered Protocol: Final outcomes are paired with process profiles covering representation quality, non-identifying interventions, calibration, and responses to contradictory evidence.These profiles reveal which discovery operations have actually formed when systems reach the same conclusion.
- 7. Capability Evaluation: A Process-Centered Protocol: The protocol separately evaluates problem discovery, formulation, representation, hypotheses, interventions, and revision using task-specific criteria for validity and informativeness.It tests whether selected unknowns expose useful structure, representations improve scientific operations, hypotheses contain discriminative predictions, interventions produce information, and revisions attribute failures appropriately.
- 7.3. Efficiency and Resource-Matched Evaluation: Discovery efficiency reports resource allocation across compute, search, tools, simulations, expert time, physical experiments, replication cost, and action risk.Components of total cost remain separate because these resources are not interchangeable.
- 7.3. Efficiency and Resource-Matched Evaluation: Comparisons must match relevant resources and distinguish genuine efficiency from selecting trivial unknowns or terminating difficult investigations.Additional memory, verification opportunities, or compute can otherwise masquerade as capability gains.
- 7.5. Continual Improvement and Transfer: Continual evaluation measures transfer across held-out episodes, controlling for fact memory, trajectory retrieval, near duplicates, benchmark-specific prompts, and extra inference compute.Within-domain, cross-task, and cross-domain transfer test whether operations generalize selectively rather than replaying content.
- 7.5. Continual Improvement and Transfer: Transfer should be reported as an episode-wise curve revealing sample efficiency, saturation, interference, and persistence across increasingly distant mechanisms.Recursive updates require stronger independence because they can alter how later progress is measured.
8. Analysis: Research Horizons and Grounding Regimes
Discovery operators apply across digital, simulation-grounded, physical, and recursive settings, but each horizon changes the requirements for grounding, provenance, action consequence, validation, and responsibility.
- 8.1. Digital Discovery: Digital environments enable fast branching, hidden tests, reproducible evidence, and large-scale measurement of research-policy decisions.Their limitation is that supplied objectives, formalizations, and evaluators can remain unexamined, while self-confirming loops require independent tests.
- 8. Analysis: Research Horizons and Grounding Regimes: The four research horizons differ in grounding and validation burden rather than forming a capability hierarchy or maturity ladder.A digital system may show stronger formulation and revision than a robot executing a fixed protocol.
- 8.2. Simulation-Grounded Discovery: Simulation-grounded settings expose latent dynamics while supporting repeatable interventions, counterfactual access, and evaluation against known hidden mechanisms.Their main analytical risk is simulator dependence, including implementation regularities and assumptions that fail in target worlds.
- 8.2. Simulation-Grounded Discovery: Simulator disagreement can distinguish predictive accuracy from intervention quality and trigger measurement, representation revision, or escalation to physical testing.Two systems may fit observations equally while differing in their ability to identify latent mechanisms.
- 8.3. Physical Discovery: Physical discovery adds provenance, execution noise, scarcity, irreversibility, and institutional constraints, making failure attribution more consequential.Calibration, contamination, drift, operator intervention, and protocol deviation can alter what a result means.
- 8.3. Physical Discovery: Physical infeasibility can force revision before execution by requiring redesigned observables, narrowed claims, or a return to simulation.Feasibility is therefore part of research-state reasoning rather than a separate engineering concern.
- 8.3. Physical Discovery: Near-term physical DFMs are collaborative: models propose, verification filters, experts authorize, laboratories execute, and independent groups replicate.Audited investigations may provide stronger evidence than many automated runs with weak attribution.
- 8.4. Recursive Discovery: Recursive discovery revises infrastructure such as skills, tools, simulators, memory, evaluators, permissions, and workflows, increasing validation demands as effect radius grows.Updates must be tested inside integrated loops because stricter verification or biased retrieval can suppress useful exploration.
9. Discussion: Epistemic Boundaries, Governance, and Recursive Risk
Discovery systems require explicit epistemic status, independent validation, auditable provenance, scoped action authority, and reversible controls because recursive updates can alter future evidence and evaluation.
- 9. Discussion: Epistemic Boundaries, Governance, and Recursive Risk: Validation, provenance, and responsibility become part of the epistemic architecture as action authority and update scope increase.A novel output, plausible explanation, or model confidence is not itself a scientific discovery, validated theory, or certainty.
- 9.1. Epistemic Status and Independent Validation: Research states should distinguish speculative, observationally supported, intervention-tested, independently replicated, and validity-range-limited claims.Novelty, importance, and validity are separate judgments.
- 9.1. Epistemic Status and Independent Validation: Independent validation should vary relevant axes such as data, model family, institution, instrument, protocol, or analysis rather than relying on nominally separate reviewers or replications.This breaks self-confirmation loops in which one system proposes, tests, interprets, and evaluates its own conclusion.
- 9.1. Epistemic Status and Independent Validation: Evidence requirements should scale with claim impact, using different validators for speculative ideas, resource allocation, laboratory changes, and infrastructure updates.Formal proofs, held-out computation, controlled interventions, conceptual replications, and independent laboratory replications support different claim types.
- 9.1. Epistemic Status and Independent Validation: Preregistered expectations should remain distinguishable from evidence-triggered revisions, and process improvements should be tested on independently selected tasks.This applies to new skills, tools, evaluators, and simulators as well as scientific hypotheses.
- 9.2. Provenance and Auditability: The auditable object is a linked research state containing versions, inputs, outputs, conditions, analyses, rejected explanations, failed interventions, human edits, and deviations.Relational provenance connects observations to protocols, revisions to evidence, and skills to supporting episodes and validation tests.
- 9.3. Action Authority and Human Responsibility: Capability and action authority are separate: permission should attach to each action and environment, progressing from observation and proposal through simulation, approval, and execution.Permissions can be narrowed when calibration degrades or the environment leaves the demonstrated competence regime.
- 9.4. Recursive Update Risk: Recursive updates can compound distortions because skills, evaluators, simulators, and memory policies shape which evidence and strategies appear later.Performance may improve while genuine discovery capability narrows when difficult cases are removed, internal agreement is rewarded, or alternatives are suppressed.
10. Conclusion
The paper defines Discovery Foundation Models as systems that add research-structure construction, evidence-based revision, and transferable discovery improvement to scientific problem solving. Zetema and GALILEO operationalize this framework, while the authors frame discovery as measurable without claiming general autonomous discovery is solved.
- Discovery requires identifying worthwhile unknowns, constructing researchable formulations and representations, designing discriminating evidence, revising after disagreement, and transferring validated lessons.
- Zetema instantiates the framework with revisable research state, a Research World Model, action gating, Dry-Lab and Wet-Lab grounding, and cross-task Discovery Skill evolution.
- GALILEO grounds the physical formulation by using external wet-lab outcomes to revise later decisions and consolidate a reusable design rule across experimental rounds.
- The proposed capability-formation and evaluation view measures externally validated episode progress alongside transferable improvement under matched resources and retrieval controls.
- The framework operationalizes DFM claims through observable formulation, representation, intervention, revision, and transfer decisions backed by external evidence the system cannot control.