Source-linked AI summary

ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents

Shuhan Xue, Jianyuan Zhong, Ziyuan Nan, Wenbin Li, Zhaochen Yu, Jinchao Ding, Qiang Gao, Pengyu Zhan, Yuntong Zhang, Tian Cheng, Zhenfei Yin, Yingcheng Wu, Ling Yang

arXiv:2609.17523v1cs.AIcs.CL

TL;DR

Scientific agents need ways to turn researcher collaboration into sustained improvements across tasks. ScienceBuddy addresses this with an interactive workspace and recursive-in-recursive self-improvement, coupling evaluated harness evolution with model reinforcement learning. Across case studies, harness refinement and model learning improve their respective metrics and broaden scientific task performance.

  • Problem

    Scientific collaboration reveals how work should be conducted and assessed, but correcting answers within a conversation does not establish improvement across tasks.

  • Method

    ScienceBuddy combines an interactive research workspace with inner harness adaptation under a fixed task model and outer reinforcement learning under the evolving harness.

  • Results

    Overall single-attempt test accuracy increases from 42.2% to 73.3% across held-out scientific tasks.

  • Takeaways & Limitations

    The case studies support complementary procedural and model-learning mechanisms within a system returned to researchers for continued interaction.

Abstract

from arXiv · show

We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports researchers in carrying out scientific tasks while transforming their requests, feedback, and execution evidence into tasks and evaluation rubrics for continual learning. At its core is recursive-in-recursive self-improvement, a paradigm that couples harness evolution with model reinforcement learning: the inner recursion improves the harness with the model fixed, while the outer recursion trains the model under the improved harness. Harness evolution shapes training experience, and model learning creates new opportunities for harness adaptation. We present case studies of researcher interaction, harness refinement, and model learning, with the benchmark cases spanning four scientific task families. By releasing ScienceBuddy as a research product, we make this paradigm available to the scientific community and take a step toward discovery intelligence: scientific AI that advances through sustained collaboration with researchers and evolves alongside the research it supports. Website: http://science-buddy.io

A. Scientific workspace and interaction

ScienceBuddy provides an inspectable workflow for scientific analysis, connecting researcher interaction with tools, data, execution, and persistent artifacts. Its workspace supports iterative planning, execution, checking, and reporting while converting interaction evidence into tasks and fixed evaluation criteria.

  • The workspace connects model instructions, reusable skills, and context with scientific tools, data, and execution.
  • Researchers can inspect interaction outputs and intermediate results while refining analyses through a persistent workspace.
  • Requests, replies, actions, observations, and artifacts provide diagnostic evidence for scientific tasks and fixed evaluation criteria.

B. Recursive-in-recursive self-improvement

ScienceBuddy couples harness evolution with model reinforcement learning in two recursions. The inner process improves procedures with the task model fixed, while the outer process trains the model under the selected harness.

  • The inner recursion keeps the task model fixed while bounded harness edits are proposed and retained only after improved paired evaluation.
  • The outer recursion uses the selected harness for environment calibration and fresh on-policy rollouts that support model updates.
  • Execution under the evolving harness supplies evidence for further refinement, coupling procedural adaptation with model learning.

1. Introduction

ScienceBuddy addresses how researcher collaboration can produce sustained improvements in scientific agents. It combines an interactive workspace, interaction-grounded tasks and rubrics, and recursive co-adaptation of harnesses and models, illustrated across four scientific task families.

  • 1. Introduction: Case studies cover literature reading, database judgments, protocol troubleshooting, and gene and variant assessment.
  • 1. Introduction: The released product returns the updated model–harness system to researchers for renewed interaction and adaptation.
  • 1. Introduction: ScienceBuddy releases an interactive workspace connecting researcher dialogue, executable analysis, inspectable artifacts, and a pluggable harness.
  • 1. Introduction: The framework derives executable tasks and evaluation rubrics from researcher requests, clarifications, execution records, and artifacts.
  • 1. Introduction: Recursive-in-recursive self-improvement couples evaluated harness edits with rubric-supervised model reinforcement learning.

2. ScienceBuddy

ScienceBuddy is an interactive scientific research workspace that combines researcher interaction, scientific tools, execution environments, and an extensible agent harness. Its recursive improvement workflow converts collaboration evidence into evaluated tasks, refines the harness, and supplies training experiences for continual model improvement.

  • ScienceBuddy: ScienceBuddy integrates evidence access, computational analysis, methodological guidance, and iterative researcher feedback in one conversational scientific workflow.Researchers can submit questions and data, inspect analyses, and refine work through subsequent exchanges.
  • ScienceBuddy: The workspace provides 224 tools across 22 functional modules, with Python, R, and Bash execution plus biomedical databases and a local data lake.Its coverage spans genomics, molecular and cancer biology, pharmacology, bioimaging, literature retrieval, and database queries.
  • ScienceBuddy: The harness is separated from interaction, execution, and workspace infrastructure, enabling alternative harnesses and consistent evaluation of revised problem-solving procedures.Editable harness components include instructions, skills, and selected context-management procedures, while surrounding infrastructure remains fixed.
  • ScienceBuddy: Collaboration records become self-contained tasks and evaluation rubrics: instructions preserve requirements and inputs, while rubrics assess scope, methods, evidence, and artifacts.Historical answers and researcher approval are not automatically treated as scientific ground truth.
  • ScienceBuddy: The post-training pipeline uses rubric-qualified trajectories for SFT and fresh on-policy rollouts with rubric scores as RL rewards.Input and runtime checks establish executability, while rubric checks and a fixed judge assess scientific requirements.
  • ScienceBuddy: Inner recursion revises the harness with task-model parameters fixed, accepting bounded edits when candidates improve normalized rubric scores without violating constraints.Parent and candidate harnesses use identical development tasks, seeds, budgets, and frozen rubrics; previous successes help detect regressions.
  • ScienceBuddy: Outer training uses the selected harness on new tasks, while environment evolution varies inputs, analysis conditions, or computational dependencies to produce fresh RL rollouts.Development, policy-training, and held-out test tasks are separated, and the held-out set never informs editing or selection.

3. Scientific Workspace and User Experience

ScienceBuddy integrates multimodal scientific inputs, evidence retrieval, computation, and execution inspection in a persistent workspace. Researchers can carry context across image-based requests while reviewing evidence, gaps, and tool traces.

  • Workspace overview: ScienceBuddy combines multimodal input, long-context reasoning, and researcher interaction in a shared scientific workspace.Its interfaces support multiple scientific domains, while current tools and data specialize in biomedicine.
  • Multimodal analysis: Researchers can connect uploaded diagrams to target analysis and structured evidence tables while keeping retrieved evidence and data gaps visible.The example distinguishes a retrieved PDE4/rolipram fragment from CD40 and AHR searches that returned no matches.
  • Workspace overview: The Chat and Compute panels place scientific material, responses, and execution history in one inspectable view.Researchers can review both the generated analysis and the records of computation supporting it.
  • Continuing sessions: Three successive image-based requests show retained dialogue and execution context connecting interpretation, retrieval, and synthesis across a continuing session.The middle response uses model knowledge without a new database query, whereas other traces record repeated searches and literature or protein queries.
  • Inspection controls: Researchers can inspect uploaded images, switch to Trajectory, and open tool events to review inputs, outputs, and metadata without starting a new conversation.The demonstrated event revisits an earlier HMGCR lookup within the same session.

4. Case Studies

The case studies examine researcher-guided task construction, coupled harness–model improvement, and independent harness or model learning. Across four scientific task families, the reported evaluations show improved validation, test accuracy, first-response accuracy, and problem coverage.

  • 4.1. Researcher interaction: Researcher requests produce ordered JAK1 study plans and evidence-linked ARL4C presentation specifications with objectives, criteria, and required artifacts.The specifications preserve gene-specific scope, sequence association before mechanistic follow-up, and link presentation claims to supporting panels and comparisons.
  • 4.2. Two-Cycle Recursive-in-Recursive Dynamics: Across three co-evolution cycles, harness refinement and reinforcement learning each continue improving their respective validation or reward metrics.Harness validation rises 38.9%→44.4%, 34.4%→46.7%, and 61.1%→70.0%; mean training reward rises 33.3%→38.8%, 44.1%→60.5%, and 57.8%→69.8%.
  • 4.2. Two-Cycle Recursive-in-Recursive Dynamics: Overall single-attempt test accuracy increases from 42.2% to 73.3%, with gains across all four scientific task families.33.3% of test problems transition from incorrect to correct, compared with 2.2% transitioning from correct to incorrect.
  • Case-study scope: The benchmark cases span literature reading, database judgments, protocol troubleshooting, and gene and variant assessment from LAB-Bench and Biomni-Eval1.Holding one component fixed isolates harness changes or model-learning effects.
  • 4.3. Harness Adaptation with a Fixed Model: With model weights fixed, harness adaptation raises separate validation accuracy from 31.1% to 51.1%, a gain of 20 percentage points.Across 24 adaptation batches, the best observed first-response accuracy reaches 75.0%.
  • 4.4. Model Learning: With the harness and attempt budget unchanged, reinforcement learning increases problem coverage from 48.3% to 67.8%, a gain of 19.5 percentage points.Coverage measures the fraction of test problems solved at least once within four attempts.

5. Related Work

Related work provides foundations for persistent procedural adaptation, recursive self-improvement, and learning from interaction. ScienceBuddy distinguishes its approach by nesting repeated harness adaptation with repeated task-model learning under a fixed reflector.

  • Persistent experience in agents: Prior systems adapt prompts, procedures, harness code, contextual playbooks, or reusable execution skills from interaction and trajectory evidence.Examples include Reflexion, GEPA, ACE, Meta-Harness, and PILOT.
  • Recursive self-improvement: Recursive self-improvement systems evolve agent archives, editable meta-procedures, or data and update directives for parameter adaptation.The cited examples include Darwin Godel Machine, Hyperagents, and SEAL.
  • ScienceBuddy’s distinction: ScienceBuddy studies a nested dependency between repeated harness adaptation and repeated task-model learning, with a fixed reflector.Its task model returns to the next inner process after learning under the active harness.
  • Learning from interaction: Interaction-learning approaches extract evaluative or directive signals from post-action states, while other systems jointly adapt environments, policies, and reward models.ScienceBuddy instead connects user-guided procedural revision with task-verification-supervised policy trajectories.

6. Conclusion

ScienceBuddy is an interactive workspace where researcher collaboration informs both working procedures and model learning. Its case studies report complementary gains from harness revision and model learning, supporting continued study of their coordination.

  • Conclusion: ScienceBuddy connects evaluated harness refinement with rubric-supervised model updates and returns the updated system to further scientific interaction.The workspace combines evidence access, computational analysis, and methodological guidance in a conversational workflow.
  • Conclusion: Researcher requests and follow-up requirements define scientific tasks and assessment criteria for continual improvement.These interactions provide the basis for studying how collaboration informs both procedures and model learning.
  • Conclusion: The case studies show complementary contributions: harness revision improves first-response accuracy with the task model fixed, while model learning expands problem coverage under the initial harness.Together, the findings support studying coordination across continued researcher collaboration.

Appendix S1. Implementation Details

This appendix connects the case studies to the scientific workspace and recursive-in-recursive framework, distinguishing fixed-model harness evolution from fixed-harness model learning.

  • The appendix explains how the case studies map onto the workspace and recursive-in-recursive framework.
  • It distinguishes harness evolution with a fixed model from model learning under a fixed harness.

S1.1. Datasets and Environments

The benchmark comprises 895 tasks across four scientific task families and 16 subtopics, with separate adaptation, validation, and model-learning evaluation roles.

  • 895 tasks span LitQA2, DbQA, ProtocolQA, and GWAS across 16 subtopics.The collection contains 96 LitQA2, 511 DbQA, 108 ProtocolQA, and 180 GWAS tasks.
  • The harness study uses adaptation conversations to revise procedures, then compares initial and selected harnesses on a separate validation set.The model-learning study instead compares two checkpoints on the same panel under the initial harness, using four attempts per problem.

S1.2. Harness Evolution

Harness evolution uses bounded, feedback-driven procedural edits, with information boundaries separating diagnosis from optimization and validation selecting the final harness.

  • The reported adaptation run updates a fixed Qwen3.5-4B task model every 12 conversations, producing 24 updates across 288 conversations.A fixed Qwen3.8-27B helper supports bounded simulation and feedback interpretation, while GPT-6 Astra proposes edits.
  • Adaptation is restricted to instruction and scoped-skill text while execution, context handling, checks, interfaces, and budgets remain fixed.H0 has no added instruction or skill entries, whereas H24 contains four instruction entries and nine scoped skills.
  • 51.1% correct versus 31.1% shows the selected harness outperforming the initial harness on validation tasks.The selected harness is chosen using adaptation-set performance, while validation scores do not guide revision or checkpoint selection.
  • The simulator and interpreter use bounded feedback, while GPT-6 Astra diagnoses unmet criteria and proposes one constrained procedural edit.Private reference answers remain outside policy and proposer inputs; proposed edits undergo component validation and execution preflight.
  • Diagnostic feedback supports harness revision, whereas separately evaluated trajectory reward supports policy optimization.The interpreter’s ternary score is not substituted for the trajectory reward used in GRPO optimization.

S1.3. Reinforcement Learning

The reinforcement-learning case trains the task model under a fixed initial harness, evaluates coverage with pass@4, and embeds this step within the broader recursive framework.

  • The model-learning case keeps H0 fixed throughout training and evaluation, comparing checkpoints on identical problems with four attempts each.This isolates model learning under a common procedural interface and does not invoke a new harness-adaptation phase.
  • 67.8% versus 48.3% pass@4 indicates higher problem coverage after model learning under H0.Coverage counts a problem once when at least one of four attempts succeeds, unlike first-response accuracy in the harness study.
  • During each outer stage, a frozen prior policy generates grouped trajectories under the selected harness, with records linking inputs, tokens, probabilities, task versions, and harness identity.The model-only case instantiates this procedure with H0, while full recursive cycles can incorporate selected harnesses and validated environments.
  • GRPO assigns each trajectory a reward-based group-relative advantage, averages at the token level, and excludes researcher messages, tool outputs, instructions, and deterministic harness actions from optimized tokens.Groups with identical rewards have zero policy-gradient advantage.
  • The objective clips policy ratios and applies a sampled KL surrogate relative to a frozen reference policy.The sampled quantity is evaluated on behavior-policy tokens and is not asserted to be an exact KL divergence under an updated policy.
  • Trajectory rewards are computed after rollout from outputs and task-relevant evidence, without backfilling later evaluation information into earlier contexts.The resulting checkpoint returns to harness re-evaluation and deployment in the full recursive framework.
Loading 2609.17523v1…