Source-linked AI summary

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

Zijian Wang, Hanqi Li, Ziyue Yang, Zijian Hu, Shenghan Zuo, Yunzhe Zhang, Da Ma, Danyu Luo, Chenrun Wang, Jing Peng, Tiancheng Huang, Sijia Guo, Huayang Wang, Zichen Zhu, Senyu Han, Yilu Cao, Bo Chen, Xin Chen, Kai Yu, Lu Chen

arXiv:2606.18874v3cs.AI

TL;DR

Automated scientific workflows often leave the reasoning connecting evidence, methods, experiments, and claims implicit and difficult to audit. Xcientist externalizes research synthesis and experimental validation into inspectable structures, preserving traceable trajectories across several research tasks. The evaluation demonstrates traceable method evolution, while showing that universal autonomous discovery ability remains unestablished.

  • Problem

    Automated research lacks sufficiently inspectable links connecting evidence, methods, procedures, failures, and scientific claims.

  • Method

    Xcientist externalizes research synthesis and experimental validation into explicit structures that can be inspected, constrained, reproduced, and revised.

  • Results

    Across memory, graph forecasting, and multiscale PINN tasks, Xcientist preserved traceable method evolution from mechanism design through validation and bounded repair.

  • Takeaways & Limitations

    AI scientists should be evaluated not only by final artifacts but also by whether synthesis and validation remain attributable, inspectable, and scientifically accountable.

  • Takeaways & Limitations

    The case-based evaluation demonstrates traceable evolution across several domains but does not establish universal autonomous discovery ability.

Abstract

from arXiv · show

AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside model inference. Here we introduce Xcientist, a research harness that externalizes research synthesis and experimental validation into inspectable, contract-governed processes. Xcientist organizes literature evidence, idea states, implementation plans, ablation records and repair traces as persistent research artifacts, so that generated mechanisms can be grounded, executed, tested and revised without losing their evidential basis. We identify claim drift as a failure mode of automated research, where runnable artifacts no longer support the mechanism originally claimed. Across training-free memory systems, graph-structured traffic forecasting and multi-scale physics-informed neural networks, Xcientist preserves traceable trajectories from problem formulation to mechanism design, validation and bounded revision. These results suggest that AI scientists should be evaluated not only by their final artifacts, but by whether their synthesis and validation processes remain attributable, inspectable and scientifically accountable.

b Paper-graph grounding Structured ideation Validated results Targeted repair

XCIENTIST connects research synthesis and experiment validation through a Paper Graph Infrastructure. It grounds literature review, structured ideas, targeted repairs and validated results while preserving links among motivations, mechanisms and experimental evidence.

  • Paper-graph grounding: The Paper Graph Infrastructure connects research synthesis and experiment validation.It supports literature review, idea generation, validation-resource retrieval and staged validation contracts.
  • Structured ideation: XCIENTIST converts paper-graph evidence into structured ideas, targeted repairs and validated results.The framework preserves links between motivations, mechanisms and experimental evidence across these stages.
  • Validated results: Three representative tasks demonstrate the framework’s progression from evidence grounding to ideation, repair and validation.This progression is organized through connected research synthesis and experiment-validation processes.

1 Introduction

Xcientist addresses the auditability gap in AI-driven research by externalizing research synthesis and experimental validation into inspectable, controllable, and reproducible structures. It grounds downstream reasoning in explicit evidence and governs execution through contracts that require validated artifacts and recorded repairs.

  • Motivation: Scientific research requires an inspectable chain connecting assumptions, evidence, procedures, failures, and conclusions.The paper identifies auditability as necessary for results to be reproducible, controllable, and open to challenge.
  • Problem: AI scientists can generate ideas, code, and experiments end to end, but their intermediate judgments often remain hidden in transient prompts or inaccessible model weights.This hidden basis makes automated research decisions difficult to audit.
  • Contribution: Xcientist transforms research synthesis and experimental validation from latent model competencies into explicit external structures that can be inspected, controlled, and reproduced.The research harness operationalizes the two capabilities underpinning scientific judgment rather than delegating them to latent statistical patterns.
  • Research synthesis: A structured evidence graph makes relationships among methods, baselines, and datasets explicit and queryable, supporting traceable research trajectories and verifiable gaps under novelty and feasibility constraints.These structures ground knowledge selection and reasoning in inspectable evidence rather than opaque model weights.
  • Experimental validation: Contract-based execution decomposes implementation and evaluation into discrete steps whose required inputs, permitted operations, deliverables, and acceptance criteria are checked before progression.Each repair loop is recorded, making experimental validation externally governed.
  • Evaluation: Xcientist is evaluated across training-free memory systems, graph-structured spatio-temporal forecasting, and multi-scale physics-informed neural networks while exposing how mechanisms are formed, implemented, diagnosed, and repaired.The validation emphasizes process transparency rather than only final artifacts.

2 XCIENTIST Architecture Overview

XCIENTIST is a three-layer research harness that connects explicit literature evidence to executable validation, repair, claim auditing and human inspection. Its ideation-validation-evolution loop makes research trajectories inspectable and shifts evaluation from final artifacts to accountable processes.

  • Three-layer architecture: XCIENTIST comprises Paper Graph Infrastructure, Research Harness and System User Interface layers that preserve the chain from evidence to claims and inspection.The layers provide the evidence substrate, governed research workflow and inspectable interface, respectively.
  • Paper Graph Infrastructure: Paper Graph Infrastructure parses full-text papers into schema-bound records organized as a method-evolution graph, grounding synthesis in explicit literature evidence.The records include problems, contributions, components, limitations, innovations, baselines, datasets and experimental relations.
  • Research Harness: The Research Harness converts evidence into constrained states for structured review, mechanism ideation, staged implementation, ablation, repair and claim auditing.Idea candidates are searched, scored and fused under novelty, feasibility and verifiability criteria before validation and reporting.
  • Ideation-validation-evolution loop: The ideation-validation-evolution loop grounds candidate ideas in evidence, translates them into executable plans, tests checkable artifacts, repairs defects and bounds claims.Ideas are not treated as final model outputs; they are revised when validation exposes defects before being written as scientific claims.
  • System User Interface and accountability: The interface records runs, artifacts, messages, traces and approvals while exposing workflow lanes and inspection surfaces for review, intervention and audit.Together, the layers shift evaluation from a final generated artifact to the research trajectory that produced it.

3 Results

Across three task domains, XCIENTIST evaluates research evolution through inspectable synthesis, validation, ablation-driven repair and quantitative testing rather than final performance alone. The results show directional design improvement, interpretable mechanisms and explicit limits against strong baselines.

  • Experimental scope: XCIENTIST evaluates research synthesis and validation across training-free memory, graph-based traffic forecasting and multi-scale physics-informed neural networks.The domains stress mechanism formation, incorporation of domain structure and proposal generation under theoretical constraints.
  • Training-free memory system: The memory-system trajectory improved overall textual F1 from 0.306 to 0.391 while reducing average token length from 2844.1 to 1017.2.This corresponds to a 64.2% reduction in output length and reflects a directional performance-cost improvement.
  • Training-free memory system: The converged memory system formed an interpretable pipeline from atomic notes and staged retrieval to reranking and an evidence bundle with defined functional roles.Its versions progressively emphasized provenance, atomic-first evidence, a fixed bundle contract and scalable reading.
  • Graph-based traffic forecasting: Graph-design repair preserved diffusion complementarity: removing the coverage cell increased clean MAE by (0.0042) and worsened block-40% degradation from (1.060) to (1.101).The result reversed the negative orthogonal-projection finding and supported targeted ablation-driven repair rather than continued scaling.
  • Multi-scale physics-informed neural networks: Across six PINN versions, average rank improved from 6.33 in v1 to 2.33 in v6, with final mean relative L2 errors of 0.0672 ± 0.0118 on heat1d_multiscale and 0.4307 ± 0.0139 on pinnacle_heat.The final design used a frozen coarse PINN reference and projected fine residual complements, but remained behind the strongest external baselines on heat2d_multiscale.

4 Discussion

XCIENTIST makes automated scientific reasoning inspectable by externalizing evidence, ideas, implementation, validation, ablations, and repairs into governed research artifacts. The discussion identifies claim drift as a central failure mode and notes that the case-based evaluation demonstrates traceable method evolution while retaining important limitations.

  • Core contribution: XCIENTIST addresses attribution by preserving traces from proposed mechanisms through runnable artifacts and final validation.A plausible proposal and runnable implementation can still fail scientifically when the final result cannot be traced to the originally claimed mechanism.
  • Core contribution: The harness represents research as governed state transitions spanning evidence, idea states, implementation contracts, validation, ablations, and repairs.The paper graph stores methods, baselines, datasets, limitations, and experimental relations as queryable objects rather than implicit model knowledge.
  • Case studies: The case studies show traceable method evolution, including compact memory-system design and repair of an inert orthogonal-projection mechanism in traffic forecasting.The memory system used immutable atomic notes, bounded enrichment, and deterministic slotted retrieval, while traffic forecasting moved from input-space filtering to proposal-space intervention.
  • Failure mode: Claim drift arises when proposal mechanisms are not structurally preserved through implementation and validation, including semantic, experimental, and mechanistic drift.Mechanistic drift occurs when numerical gains cannot be attributed to the claimed component because controls are inadequate or absent.
  • Limitations: Evaluation remains limited by paper-graph quality, case-based coverage, runnable repositories, available datasets, and meaningful benchmark protocols.The evaluation demonstrates traceable method evolution across several domains but does not establish universal autonomous discovery ability.

5 Methods · Appendix

XCIENTIST externalizes automated research through a paper-graph evidence substrate and a coupled harness spanning literature synthesis, structured ideation, executable validation, reporting and claim auditing. Its methods preserve source anchoring, explicit contracts and inspectable transitions from mechanisms to validated evidence.

  • 5 Methods: XCIENTIST converts full-text papers into schema-bound evidence records and a method-evolution graph that supports literature review, ideation, validation and reporting.The framework couples a Paper Graph Infrastructure with a research harness and an ideation-validation-evolution loop.
  • 5 Methods: The paper graph is built from approximately 50,000 computer-science papers indexed through Semantic Scholar and parsed with MinerU.MinerU combines rendered full text with structured content lists to capture method components, baselines, datasets, experimental conditions and limitations.
  • 5 Methods: Graph records store textual fields as keyword, summary, insight and quote tuples, anchoring retrieval and synthesis to source passages.The graph represents core methods, baselines and datasets with typed methodological, comparison and evaluation edges.
  • 5 Methods: Literature synthesis uses DeepSurvey to transform a user topic or seed set into a reusable analysis substrate through bounded graph and citation expansion.Expanded candidates are filtered using title- and other criteria described by the literature-synthesis procedure.
  • 5 Methods: Idea generation searches over structured idea states containing problem framing, mechanisms, supporting components, expected advantages, risks and preliminary validation plans.Multiple Idea Taste Modes explore the same research problem under different encoded preferences.
  • 5 Methods: Experiment validation turns structured ideas into executable evidence under validator supremacy, separation of concerns and workspace encapsulation.Independent validators enforce explicit contracts, while planning, execution and verification remain distinct and experiments use self-contained projects.
  • 5 Methods: The reporting module converts validated artifacts into scientific prose through a write-audit-repair loop that checks code, ablations, iteration reports and literature evidence.It produces blog_idea.md as a composition contract containing the project overview, architecture, outline and citation table.

A XCIENTIST Architecture Implementation … A.3.4 Idea Fusion

XCIENTIST builds an evidence-grounded research harness that parses and structures literature, supports traceable synthesis, and guides controlled idea generation. Its ideation process combines diverse search preferences through constrained fusion and accepts revisions only after re-evaluation.

  • A.1 Paper Graph Infrastructure: XCIENTIST decomposes papers into method entities linked to shared baselines and datasets, exposing experimental structure for method-evolution analysis.The graph is built from nearly 50K source papers and represents method contributions rather than treating papers as monolithic nodes.
  • A.1.2 Full-text Structured Parsing: Full-text parsing and schema-bound extraction capture problems, contributions, components, limitations, future work, baselines, datasets, and experimental relations as traceable evidence fields.Each field uses keywords, summary, insight, and a verbatim quote, while baseline and dataset entities are resolved against Semantic Scholar references.
  • A.1.4 Graph Construction with Entity Resolution: The resulting heterogeneous graph contains approximately 72K core nodes, 250K baseline nodes, 63K dataset nodes, 385K total nodes, and 1.15M typed edges.Comparison and evaluation edges retain metrics, context, summaries, and supporting quotes; globally shared entities are merged through canonical Semantic Scholar Paper IDs.
  • A.1.5 Integration with Downstream Modules: The graph supplies evidence for literature synthesis and validation by tracing method lineages and specifying baselines, datasets, metrics, and component-level ablations for validation contracts.Its structured innovations, limitations, future-work, component, baseline, and dataset fields support downstream review and executable experiment planning.
  • A.2 Literature Review: DeepSurvey turns graph-retrieved, full-text evidence into structured survey analyses through staged retrieval, keynotes, clustering, multi-perspective comparison, citation assignment, and refinement.Graph expansion follows citation and reference edges with hybrid filtering, while source attribution and citation verification constrain generated claims to assigned evidence.
  • A.2.5 Clustering and Multi-perspective Relation Analysis: The literature-review substrate compares methods through relation graphs, comparison tables, guided Q&A, code analysis, and localized citation constraints.These analyses capture technical lineage, mechanisms, experiments, limitations, implementation details, and research gaps while monitoring source attribution.
  • A.3 Idea Generation: XCIENTIST frames ideation as a staged, cross-mode search that moves from evidence and problem understanding to hypothesis construction, refinement, and selection.Idea Taste Modes alter preferences over the same grounded problem, favoring distinct combinations of novelty, transfer, feasibility, implementation ease, and experimental defensibility.
  • A.3.3 Memory-Guided MCTS: Idea fusion selects a host thesis, retains only components that strengthen its single core mechanism, records synthesis decisions, and accepts repairs only when re-evaluation improves the composite score.Candidates come from different taste modes but share a prepared root context; the repair loop preserves the prior draft when the revised version does not improve.

A.4 Experiment Validation · A.4.1 Preparation: Building the Execution Surface · A.4.2 Master Scheduling and Convergence Judgment

XCIENTIST frames experiment validation as a convergent, validator-backed process that converts structured proposals into executable code, empirical evidence, and inspectable reports. Preparation and scheduling enforce contract-based progression, state-dependent iteration, repair, and strict convergence criteria.

  • A.4 Experiment Validation: A.4 Experiment Validation: Validation transforms structured research proposals into self-contained executable code and gathers evidence through staged benchmark execution and component-level ablation.Only validator-backed execution artifacts constitute valid progress, and validation decisions are recorded in inspectable structured reports.
  • A.4 Experiment Validation: A.4 Experiment Validation: Validator supremacy requires independent verification to confirm explicit contracts before any stage can declare completion.Planning, execution, and verification are treated as separate concerns, while claim acceptance or rejection rests on empirical evidence rather than model output alone.
  • A.4.2 Master Scheduling and Convergence Judgment: A.4.2 Master Scheduling and Convergence Judgment: The master scheduler selects actions from validated workspace state rather than following a fixed linear script.Its priority checks cover self-contained code, implementation validation, standard-experiment artifacts, and ablation science.
  • A.4.1 Preparation: Building the Execution Surface: A.4.1 Preparation: Preparation builds the execution surface through repository acquisition, environment construction, dataset staging, model staging, and synthesis under explicit contracts.Each contract specifies goals, input paths, permitted write roots, required outputs, and a done condition.
  • A.4.1 Preparation: Building the Execution Surface: A.4.1 Preparation: A step executor advances only after validator pass, routing failed validation feedback to worker repair until contract satisfaction or a repair-round limit.Strict sequential gating prevents downstream phases from beginning with ambiguous prerequisites.
  • A.4.1 Preparation: Building the Execution Surface: A.4.1 Preparation: Synthesis produces prepare_target_inventory.json and prepare_idea.md as authoritative machine-readable and human-readable handoff artifacts.The inventory records verified repository paths, dataset files, model identifiers, environment variables, and benchmark entrypoints.
  • A.4.2 Master Scheduling and Convergence Judgment: A.4.2 Master Scheduling and Convergence Judgment: Scheduler artifacts record decisions, phase status, blocking issues, and evidence paths, supporting auditability and resumption after interruption.Convergence requires validator-backed passes for every required phase, no unresolved blocking issues, and correction of all canonical components.

A.4.3 Code Enablement and Ablation Exposure

The implementation phase converts research handoff documents into self-contained runnable code through ordered, verified steps. It also requires per-component ablations that isolate marginal contributions without changing other components.

  • Implementation workflow: The planning layer specifies ordered implementation steps with goals, input paths, permitted write roots, and verification commands.Each step translates the research proposal encoded in handoff documents into runnable code under project/.
  • Implementation workflow: The plan must end with an integration smoke test running the full experiment command on the real prepared dataset and model.This test checks the implementation chain in an end-to-end execution.
  • Ablation exposure: Code must support disabling every canonical proposal component individually while preserving the behavior of all other components.Each ablated variant must document its method context, enabling isolation of each component’s marginal contribution.

A.4.4 Standard and Ablation Science … B.2 Full-text Keynote Extraction

XCIENTIST validates methods through standard and component-wise ablation experiments, integrates validated evidence across iterations, and produces audited reports grounded in source code, citations, and generated figures. Its interface and runtime substrate preserve inspectable execution state, while the literature case study combines graph expansion with full-text keynote extraction to enrich and diagnose evidence.

  • A.4.4 Standard and Ablation Science: Standard science compares the complete method with a baseline, while ablation science disables one canonical component per ordered experiment step to measure marginal contributions.Both lanes share worker-validator infrastructure, and ablation plans contain exactly one step for each proposal component.
  • A.4.5 Iteration Integration and Final Evidence: The scheduler grounds each iteration in validator reports and decision history, then emits a canonical ablation results file when converged or records the best available progress at the iteration limit.The final file records per-component results, metric values, confidence levels, and method-context analyses.
  • A.5 Report Writing: Report writing transforms validated workspace evidence into a source-faithful narrative through contracted exploration, section-by-section composition, bidirectional citation constraints, and explicit graph placeholders.Exploration creates blog_idea.md with project structure, narrative arc, and justified candidate citations; composition consults assigned source fragments and papers.
  • A.5.3 Quality Analysis and Source Fidelity: Auditing scores reports across six categories, verifies technical claims and citations against source materials, and applies a twenty-point penalty when papers or citation verification are missing.Source Fidelity checks function names and parameter defaults against the code, while Research Integrity verifies cited PDFs and attributed claims.
  • A.5.4 Iterative Refinement and Convergence: Refinement repeats verification and repair until the quality score exceeds 90 or a maximum of three iterations is reached, preserving correct code structure while fixing critical mismatches.The pipeline also generates aligned figures from graph method files and resumes interrupted workflows from workflow_status.json checkpoints.
  • A.5.7 Integration with the Research Harness: As the harness’s terminal stage, report writing consumes source code, validated ablations, and iteration reports, producing a scored, audited article whose claims, citations, and figures are traceable to research artifacts.The frontend–backend stack supplies stable interaction, orchestration, intermediate-artifact preservation, and runtime-trace support rather than an independent scientific contribution.
  • A.6 System Interface: The system exposes topic-centric review, ideation, and experiment lanes over snapshot–stream state, backend lifecycle orchestration, isolated run contexts, persistent artifacts, and normalized events for inspection and recovery.Its deliberately single-host deployment and local filesystem storage improve operational simplicity but constrain distributed scaling and cross-host artifact management.
  • B Literature Review Case Study: DeepSurvey turns AutoSurvey into an inspectable evidence substrate by graph-expanding seed papers, hybrid-filtering the corpus to approximately 100 papers, and extracting full-text keynotes that reveal errors hidden by abstracts.The keynote condenses 1015 words to 105 words and identifies overgeneralization, rather than missing citations, as the dominant error mode.

B.3 Clustering and Multi-perspective Relation Analysis

DeepSurvey clusters retrieved papers by thematic research questions and analyzes each cluster through complementary views of citation relations and structured comparison. The synthesis highlights a central tradeoff between explicit evolutionary modeling and flexible iterative refinement in survey generation.

  • Clustering: DeepSurvey groups retrieved papers into thematic clusters organized around primary research questions.The clustering supports local comparison and layered analysis within each research theme.
  • Multi-perspective relation analysis: Within each cluster, relation graphs capture typed citation relationships, while comparison tables align methods across shared dimensions.The relation graph models foundation, extension, and substitution links; the comparison table supports systematic horizontal comparison.
  • Method comparison: Survey-generation methods differ mainly between explicit hierarchical structures that model evolution and iterative outline refinement that favors flexibility and efficiency.Knowledge trees and citation graphs provide stronger organization and lineage modeling at higher cost, whereas iterative refinement is simpler and more adaptive but less explicit about evolution.
  • Design implication: The resulting synthesis identifies explicit evolutionary modeling versus iterative refinement as the central design choice for planning mechanisms.This distinction is intended to inform the Idea Agent’s search for more effective planning mechanisms.

B.4 Code Repository Analysis … C.6 Component-grounded Novelty Checking in MCTS Ideation

The paper traces DeepSurvey’s code-grounded survey construction and XCIENTIST’s evidence-grounded idea generation from implementation analysis and citation control through diagnosis, mechanism design, MCTS refinement, and component-based novelty checking. Across these stages, persistent evidence, artifacts, and validation constraints shape both survey outputs and research ideas.

  • B.4 Code Repository Analysis: DeepSurvey finds AutoSurvey evolving from monolithic single-pass systems toward modular, multi-agent, iterative frameworks built around hierarchical pipelines.The dominant architecture combines RAG, dynamic planning, specialized agent collaboration, and iterative refinement.
  • B.4 Code Repository Analysis: PyTorch v2.0.0+ with Hugging Face Transformers forms the standard stack, while LangChain and LlamaIndex dominate orchestration; deep understanding, evaluation alignment, long-context handling, and domain generalization remain gaps.These code-grounded findings expose implementation patterns and engineering constraints that paper prose alone cannot convey.
  • B.5 Outline-driven Drafting and Citation Enforcement: DeepSurvey generates an eight-section hierarchical outline with scoped descriptions, assigned paper subsets, and citation anchoring that constrains each section’s evidence.After each paragraph, citation verification checks consistency against the assigned set and triggers retry generation when errors are detected.
  • B.5 Outline-driven Drafting and Citation Enforcement: DeepSurvey’s outline stays focused on automated survey generation and organizes evidence around a four-paradigm taxonomy culminating in the skeleton-versus-flesh trade-off, unlike AutoSurvey’s irrelevant content.The taxonomy comprises one-shot, iterative, multi-agent, and structure-first paradigms.
  • B.6 Multi-granularity Agentic Refinement: After drafting, DeepSurvey coordinates keynote reader, reviewer, and reviser roles through explicit action plans, global memory, and paragraph-level citation-claim validation.The reviewer checks whether cited papers support claims such as “AutoSurvey achieves near-human performance.”
  • C Idea Generation Case Study; C.1 Literature Grounding; C.2 Analysis of Field Structure: XCIENTIST converts a partially validated training-free memory research state into an evidence-grounded idea by reconstructing context from prior candidates, ablations, and a mature-idea anchor.Literature retrieval and advanced analysis then structure the field into methods, consensus, open problems, and evaluation gaps.
  • C.3 Problem Diagnosis; C.4 From Evidence to Root Idea: The diagnosis identifies similarity-first retrieval, brittle admission and refinement thresholds, duplicate or contradictory memory growth, and redundancy- or order-sensitive evidence packing as key failure modes.The resulting idea retains immutable span-grounded atomic notes and shifts retrieval toward family-scoped, reliability-aware selection.
  • C.5 Monte-Carlo Tree Search; C.6 Component-grounded Novelty Checking in MCTS Ideation: MCTS refines the root idea, which scores 3.74 in simulation, by prioritizing repairs for validation gaps, brittle single paths, rare-regime failures, and silent failures while component retrieval grounds novelty checks.Paper-graph components provide reuse and revision references, novelty comparison, and cross-domain mechanism-transfer principles such as budget control and quality-latency routing.

C.7 Idea Fusion

XCIENTIST fuses five candidate ideas into a family-calibrated training-free memory system, retaining mature infrastructure while selecting and refining a thesis-bearing mechanism. Conflict resolution favors bounded, ephemeral reliability adjustment, and local repair raises the draft score from 3.99 to 4.04.

  • Idea fusion: The fusion preserves shared infrastructure, including two_stage_retrieval, modular_atomic_note_enricher, per_slot_quota_- and_decay, and note_attached_micro_handles.These components form the common backbone across the five candidate ideas.
  • Idea fusion: The fused draft selects the moonshot inventor candidate as host and proposes Analytic Family-Calibrated Rank-and-Pack for Training-Free Atomic-Note Memory.Its core mechanism is CalibratedBayesianFamilyRankAndPack, supported by ShrinkageAnalyticNullEstimator.
  • Conflict resolution: Conflict resolution uses a small write-time prior and computes the main family-reliability adjustment from the bounded retrieved pool, avoiding drift-heavy mutable family state.The fusion also rejects adding a router/controller layer.
  • Local repair: 3.99 to 4.04: The first local repair improves the fused-draft score by narrowing the mechanism and making the analytic null estimator thesis-bearing.Deterministic packing remains the mature one-pass substrate with explicit degraded-mode behavior, while later boundary alternatives do not improve the score.

D Experiment Validation Case Study … E.1 Evidence Inputs and Report Framing

XCIENTIST validates a training-free slotted evidence retrieval system through an inspectable workflow linking an idea contract, runnable implementation, matched-condition experiments, and bounded claim reporting. The report-generation stage uses structured artifacts rather than free-form prompting, preserving the distinction between headline evidence and diagnostic materials.

  • D Experiment Validation Case Study: The case study validates training-free slotted evidence retrieval for scalable LLM-agent memory using immutable notes, deterministic enrichment, and capacity-capped retrieval.The idea contract defines four canonical components, including slotted evidence retrieval.
  • D.1 From Idea Contract to Runnable System: XCIENTIST converts the idea contract into a self-contained implementation under project/ and completes all nine implementation steps with validator-backed PASS verdicts.The package includes modules for immutable notes, metadata extraction, sparse handles, two-stage retrieval, and duplicate suppression.
  • D.2 Standard Science Validation: Standard validation compares an all-components-disabled baseline with the full method under the same LoCoMo subset, embedding model, and answer generator.The matched conditions use all-MiniLM-L6-v2 embeddings and gpt-4o-mini generation.
  • D.2 Standard Science Validation: The standard science phase passes validator checks for required improvement-ratio, absolute-improvement, and prediction criteria.The supplied passage states that the overall improvement ratio exceeds its threshold and the absolute improvement exceeds its minimum margin.
  • D.3 Controlling Claim Boundaries: XCIENTIST retains component-ablation files as diagnostic artifacts but does not promote incompatible ablation percentages as headline results.The replacement evaluation table contains only Baseline and Full Method rows, so headline claims remain bounded by compatible evidence.
  • D.4 Validation Outcome: The validation process records the path from idea to implementation to metric claim through concrete files, commands, validators, and reproducible artifacts.Its conclusion concerns inspectability and traceability, not merely superior method performance.
  • E Report Writing Case Study: XCIENTIST turns implementation artifacts, configurations, experimental outputs, figure placeholders, and literature references into a structured technical report.The report organizes material into problem framing, architecture, retrieval, validation, efficiency, ablation, and future directions.
  • E.1 Evidence Inputs and Report Framing: Report generation uses source code, experimental results, generated figures, and supporting papers as structured evidence rather than a free-form writing prompt.The resulting report frames training-free deterministic slotted retrieval as improving long-term LLM-agent memory by replacing heavy memory mechanisms.

E.2 Implementation-Grounded System Narrative · E.3 Retrieval-Centered Technical Narrative

The report grounds its system narrative in concrete implementation artifacts and experimental outputs, then centers the technical explanation on retrieval rather than memory storage. It describes enriched notes, bounded candidate selection, and slotted evidence retrieval under explicit capacity constraints.

  • E.2 Implementation-Grounded System Narrative: The report constructs a problem narrative contrasting heavy canonicalization and learned memory-evolution layers with immutable atomic notes and slotted retrieval.This frames the system as a deterministic alternative rather than an abstract idea.
  • E.2 Implementation-Grounded System Narrative: Concrete components make the architecture inspectable, including AtomicMemoryStore, ModularAtomicNoteEnricher, and UltraSparseFacetHandleIndex.The report also names SlottedEvidenceReranker and DedupMinorityAwareProvenanceAdjudicator.
  • E.2 Implementation-Grounded System Narrative: XCIENTIST converts implementation artifacts, experimental outputs, and literature references into a structured report.This evidence-to-narrative transformation is summarized in Table 18.
  • E.3 Retrieval-Centered Technical Narrative: The central technical narrative focuses on retrieval rather than memory storage.This emphasis organizes the explanation around the read path and evidence selection mechanism.
  • E.3 Retrieval-Centered Technical Narrative: Notes are enriched at write time with embeddings, entity and temporal tags, QA keywords, and context descriptors.These enrichment fields support the later retrieval process.
  • E.3 Retrieval-Centered Technical Narrative: The read path builds a bounded candidate pool through ANN search and lexical screening.Candidate generation is therefore presented as a constrained retrieval stage.
  • E.3 Retrieval-Centered Technical Narrative: Slotted evidence retrieval assigns candidates into support, provenance, temporal, and conflict slots under explicit capacity constraints.The mechanism is presented as the main technical contribution of the retrieval-centered narrative.

E.4 Experimental Evidence Integration · E.5 Post-Generation Audit · F Key Instructions and Prompt Contracts

XCIENTIST integrates quantitative experimental evidence into report generation, then audits claims against implementation, data, figures, and citations. Its prompt contracts enforce grounded extraction, bounded idea revision, contract-governed execution, component-complete validation, and source-fidelity checks.

  • E.4 Experimental Evidence Integration: 0.391 F1 versus 0.306 for the baseline represents a 27.8% relative improvement, while token usage falls 64.2% from 2,844 to 1,017 tokens per query.The quality report checks these reported values rather than leaving them as unverified prose.
  • E.5 Post-Generation Audit: The post-generation audit verifies source-code fidelity, parameter alignment, citation authenticity and completeness, experimental consistency, figure placement, and writing quality.It checks functions, key parameters, embedding models, cited papers, and article-body references.
  • E.5 Post-Generation Audit: The audit identified no mismatches requiring substantive repair, treating report writing as controlled conversion of artifacts into claims followed by verification.Claims are checked against code, data, figures, and citations.
  • F Key Instructions and Prompt Contracts: Paper-graph extraction contracts require concise factual summaries, independent insights, verbatim quotes, representative keywords, relation-to-core links, and high recall.They distinguish standalone core contributions from component modules and require strict JSON outputs.
  • F Key Instructions and Prompt Contracts: Idea-generation contracts preserve the mature idea’s topic and mechanism axis, apply localized repairs within explicit boundaries, and reject unsupported paradigm shifts.They prioritize diagnosed mechanism bottlenecks over evaluation tooling and favor repairing existing rules, objectives, representations, or training contracts.
  • F Key Instructions and Prompt Contracts: MCTS contracts require exactly one edit operator per child, explicit defect targeting, concrete algorithmic intervention, evidence references, and structured risk and experiment fields.Referee scoring rewards mechanism-level edits and minimal falsification protocols while penalizing feature dumping, scope drift, and unjustified complexity.
  • F Key Instructions and Prompt Contracts: Code and science contracts require real prepared data and bindings, baseline-versus-full-method comparisons, one ablation per canonical component, raw evidence, and validator-reviewed convergence.They prohibit synthetic substitutes, hidden dependencies, fake smoke tests, and incomplete or reordered component coverage.
  • F Key Instructions and Prompt Contracts: Source-fidelity contracts verify names, parameters, paths, architecture, experimental values, citations, references, and figures, then require minimal repairs only for confirmed critical issues.Auditors must flag fabricated statistics, unsupported conclusions, and missing evidence while preserving correct content and authorial intent.
Loading 2606.18874v3…