Source-linked AI summary

Traceable Trust for action-ready artificial intelligence in bioscience

Huayu Xin, Yizhi Cai, Mukilan Deivarajan Suresh, Gavin Michael Farrell, Iwona Gajda, Charlie Harrison, Conor Houghton, Mato Lagator, Yang Lu, Virginia Portillo, Reyer Zwiggelaar, Sebastian Lobentanzer

arXiv:2608.17997v1cs.CYcs.AI

TL;DR

AI outputs increasingly shape bioscience workflows, but the conditions for using them to guide experiments require reviewable handling. This paper proposes Traceable Trust, whose cases show that making evidence and decisions visible can support challenge, reconstruction, and revision.

  • Problem

    The paper addresses which conditions should govern using an AI output to guide a real experimental step, where decisions may otherwise lack records for reconstruction or audit.

  • Method

    Traceable Trust defines an assessment-and-design framework for the boundary between an AI output and its implementation in a research workflow.

  • Results

    The cases show that making evidential and decision-making processes visible can support challenge, reconstruction, and revision as part of scientific work.

  • Takeaways & Limitations

    Traceable Trust offers researchers, journals, institutions, and infrastructures a common language for documenting trust when AI outputs shape scientific work.

  • Takeaways & Limitations

    The framework still requires empirical validation.

Abstract

from arXiv · show

Artificial intelligence (AI) is becoming part of the working infrastructure of the biosciences. AI models can predict biomolecular structures, design proteins, rank variants, annotate images, recommend strains and optimise experimental conditions. We argue that the decision to use an AI output to guide laboratory action is a key juncture for trustworthy research and should follow a defined, reviewable process. We propose Traceable Trust as a proportionate assessment-and-design framework for this output-to-action boundary. It asks what evidence supports the output, what capability is being claimed, what agency has been delegated, what threshold authorises action, who can override it and how outcomes inform later decisions. We illustrate the framework through three case studies spanning ecosystem resources, project design and laboratory action. Together, the cases show how trust can be documented where AI outputs begin to shape scientific work.

1. Assessing trust at the output-to-action boundary

Traceable Trust assesses whether an AI output is sufficiently supported to justify a specific laboratory action at the output-to-action boundary. It distinguishes action readiness from AI-readiness and creates a proportionate record that enables later reconstruction, challenge and revision.

  • Reviewability: Without a reviewable process, AI outputs may be followed, rejected or modified without records that support later reconstruction or audit.The framework addresses this gap by focusing on the transition at which an AI output is allowed to influence laboratory decision-making.
  • Output-to-action boundary: Trust is assessed at the boundary between an AI output and its implementation in a research workflow.The relevant outputs include suggestions, predictions, rankings and plans, while actions include repeating measurements, ordering variants, committing liquid-handler runs or changing DBTL rounds.
  • Action readiness: Action readiness concerns whether a particular AI output can be justified as the basis for a research action, whereas AI-readiness concerns preparation for effective AI use.The two can diverge when predictions fall outside validated domains, are insufficiently calibrated, receive excessive automation permissions or lack complete records of negative outcomes.
  • Traceable Trust framework: Traceable Trust is a framework for assessing and designing the transition from AI output to laboratory action.It aims to provide enough evidence for another competent researcher to reconstruct, challenge and revise the decision.
  • Traceable Trust framework: Its six components specify the evidence needed to justify a particular output-to-action decision, while outcome records support subsequent review.The broader framing considers model performance, available evidence, context, actors, required checks and anticipated and recorded consequences.

2. Traceable Trust as an assessment-and-design framework

Traceable Trust is a proportionate framework for deciding whether an AI output can support defined experimental action. It makes capability, delegated agency, thresholds, authorisation and revision visible at the output-to-action transition.

  • Framework purpose: Traceable Trust assesses whether an AI output can support a defined experimental action at the laboratory’s point of considering action.The transition record marks the review point between AI output and authorised laboratory work.
  • Six components: The framework uses six elements: evidential provenance, capability claim, delegated agency boundary, action threshold, authorisation and override, and outcome-based revision.These elements capture evidence, intended capability, permissions, safeguards, accountability and how later decisions incorporate outcomes or exceptions.
  • Capability and agency: Traceable Trust separates what a system has been validated to do from what the laboratory allows it to influence.Capability may involve classification, ranking, generation, optimisation or planning, while agency ranges from read-only advice to resource commitment.
  • Delegated agency: Delegated agency has four levels: advisory, coordinative, executive and resource-committing, with stronger validation and explicit approval often required as agency increases.Executive use initiates software or instrument steps under constraints, whereas resource-committing use consumes samples, reagents or instrument time or shifts research direction.
  • Action-level explainability: Action-level explanations identify the evidence, signal or rationale behind an output, making agentic decisions easier to reproduce, contest and revise.Examples include tracing a qPCR repeat flag to its file, well, rule or model signal and recording why a DBTL candidate was chosen.

3. Case studies

Three case studies apply Traceable Trust across the ecosystem, project-design and experimental-action levels. They show how OSAI, AAC and AI-guided DBTL contribute different evidence and records at the output-to-action transition.

  • Overview: The cases span reusable AI resources, project design and laboratory action, showing what each approach contributes to Traceable Trust.OSAI operates at the ecosystem level, AAC at the project-design level, and AI-guided DBTL at the experimental-action level.
  • OSAI: OSAI strengthens evidential provenance by helping researchers find, reuse, compare and maintain resources behind AI workflows.Its recommendations are mapped to more than 300 ecosystem components, with guiding implementation pathways.
  • OSAI: OSAI improves the resource base for local decisions, while teams must add situated capability, agency, thresholds, safeguards, authorisation and outcome records.A possible extension is transition-oriented metadata connecting ecosystem components to action-level assessments.
  • AAC: AAC makes an AI system’s planned role explicit before laboratory connection, contributing capability claims, agency boundaries, action thresholds and authorisation routes.In the qPCR example, the system drafts repeat lists and reagent estimates while human operators retain control of automated runs and reported results.

4. Aligning the case studies

The three case studies place Traceable Trust at complementary layers, tracing a reviewable path from reusable resources to action-level evidence. Together, they show that output-to-action transitions draw on traceable resources, pre-action roles and thresholds, and outcomes that inform the next decision.

  • Complementary layers: OSAI supports AI-ready resources, AAC defines project scope and agency before deployment, and AI-guided DBTL laboratories select new physical experiments.The cases span ecosystem resources, project design, and laboratory action.
  • Complementary layers: The three cases trace the path from reusable resources to action-level evidence across Traceable Trust’s six components.Table 2 summarises how each case contributes to the six components.
  • Action thresholds and authorisation: OSAI provides supporting resource evidence while local teams define laboratory thresholds and authorise specific wet-lab actions.Its resource links connect reusable outputs to later records.
  • Outcome and revision records: AI-guided DBTL laboratories use thresholds to select candidates, controls, replicates, and uncertain options, then use failed, invalid, negative, and low-yield results to inform the next decision.Approval is recorded for each round, especially when automated platforms consume scarce resources.
  • Action thresholds and authorisation: AAC defines design-stage approval, while runtime records show who approved, rejected, or changed recommendations before laboratory action.This makes roles and authorisation reviewable at the project and runtime stages.

5. Evaluating Traceable Trust

Traceable Trust adapts evidence-based trustworthy AI assurance to the output-to-action boundary and proposes six practical indicators spanning its six components. The indicators support proportionate, reviewable records that preserve enough credible information for action review while keeping routine work manageable.

  • Framework scope and indicators: Traceable Trust focuses trustworthy AI assurance on the output-to-action boundary and proposes six practical indicators mapped across six components.Some indicators test individual components, while others assess the transition record as a whole.
  • Framework scope and indicators: The indicators assess provenance completeness, decision reconstruction time, threshold compliance, override learning, failure capture and user burden.They cover evidence and workflow versions, action rationale, predefined quality controls, human changes, unsuccessful outcomes and manageable record-keeping.
  • Proportionate implementation: Application should be proportionate: qPCR repeat recommenders may need simple records, whereas protein-design or DBTL platforms usually require richer technical and experimental information.Richer records may include model versions, candidate generation, batch design, instrument logs and failed constructs.
  • Proportionate implementation: The components provide common questions, but evidential depth and practical weight vary with the action, risk and user; partial records support only limited action-readiness claims.Quantitative measures and qualitative review may both be useful, while compliance scores remain secondary to decision reconstruction.

6. Implications for journals, research institutions and infrastructures

Traceable Trust coordinates reviewable evidence, permissions, authorisation and outcomes across journals, laboratories, institutions and infrastructures as AI outputs begin guiding scientific action. It also provides a basis for studying how action-ready arrangements are assembled, governed, interrupted and revised.

  • For journals: Journals can require authors to identify AI output-to-action transitions, document provenance and delegated agency, and align action-readiness claims with available evidence.Methods sections and supplementary information can record thresholds and outcome records for AI outputs that materially changed experimental actions.
  • For laboratories: Laboratories can classify delegated agency, set action thresholds before initial runs, and connect outcome records to subsequent model or workflow decisions.High-value transitions include repeat recommendations, variant selection, microscopy triage, plate scheduling and automated protocol changes.
  • For research institutions and universities: Institutions and universities need domain-specific governance and training that let scientists question AI outputs, explain reliance, and record how evidence, uncertainty and outcomes shaped decisions.Embedding Traceable Trust in supervision, methods teaching and research-integrity training connects responsible use, sensitive-data safeguards and academic integrity to reviewable evidence and authorisation.
  • For research infrastructures: Research infrastructures can bridge AI-ready and action-ready resources by linking catalogues, planned component roles, execution records and wet-lab context.The practical challenge is interoperability across these systems while minimising duplicate data entry.
  • Beyond bioscience and participant-facing research: Beyond bioscience and in participant-facing research, trustworthy AI action requires transparent records, accountable governance, and routes to question or contest data use.In biobanking, technology that manages consent may reduce trust when autonomy, social relationships or institutional responsibility are obscured.
  • For science and technology studies and responsible innovation research: For science and technology studies and responsible innovation research, Traceable Trust makes action-ready arrangements an empirical object for examining delegated agency, authority, failures and revision.Such studies can test whether trace records support scientific judgement or become paperwork, while critical evaluation and human oversight address overestimated understanding of complex processes.

7. Limitations and next steps

The framework remains to be empirically validated, with comparative traceability metrics and testing across laboratories differing in automation, data sensitivity, staff capacity and regulatory exposure. Next steps include resolving implementation, governance and burden questions while addressing the limits of the case selection and anticipating responsibilities before traceability becomes an after-the-fact audit trail.

  • Validation: The framework requires empirical validation and a comparative metric for traceability practices across laboratories.Validation should span laboratories with different levels of automation, data sensitivity, staff capacity and regulatory exposure.
  • Scope and burden: Validation must attend to burden because excessive form-filling can undermine scientific practice.The perspective also excludes deliberate misuse and dual-use oversight, which require separate AI and biosecurity treatment.
  • Research questions: Open research questions concern automating transition records, defining task-specific action thresholds, journal treatment of recommendations, and retaining failures for trustworthy model updates.Examples include qPCR repeats, variant synthesis, microscopy triage and DBTL optimisation.
  • Case selection and implementation: The case selection is limited because OSAI, AAC and DBTL represent ecosystem-level, design-stage and literature-derived laboratory cases, respectively.Future work could compare lightweight transition records with existing documentation and test an openly available implementation across bioscience and other scientific and engineering fields.
  • Stewardship and accountability: Further work should develop stewardship for updating and verifying the framework with researchers, journals, institutions and infrastructure providers.It should also examine who benefits from traceability, carries the recording burden and can reinterpret a trace when something goes wrong.
  • Anticipatory governance: These limitations support anticipatory governance that considers responsibilities, burdens and failure modes before traceability becomes an after-the-fact audit trail.This is especially relevant as AI systems begin shaping experimental action before their institutional forms are settled.

8. Conclusion

Traceable Trust defines the conditions under which AI outputs can guide real experimental steps in bioscience. It makes the evidential and decision-making process supporting trust visible and reviewable as part of scientific work.

  • Motivation: The framework responds to a pressing trust question as bioscience advances in AI performance, reporting and openness.The central issue is the conditions under which AI outputs can guide real experimental steps.
  • Framework contribution: Traceable Trust offers researchers, journals, institutions and infrastructures a common language for assessing when AI outputs can guide experimental action.The framework focuses on the transition from an AI output to a real experimental step.
  • Assessment dimensions: The framework examines the evidence behind an output, the claimed capability in context, delegated agency, action thresholds, authority to approve or stop, and subsequent learning.These questions define the conditions under which an output can guide action and how outcomes inform later decisions.
  • Reviewability: Making the output-to-action transition visible supports challenge, reconstruction and revision of the evidential and decision-making process.This process can become a reviewable part of scientific work.

Funding

The authors report funding and institutional support from BBSRC, ELIXIR-STEERS, BioComputingUP, AIBIO-UK, and Responsible AI UK.

  • Mukilan Deivarajan Suresh received funding from the BBSRC NLD DTP (BB/T008695/1).
  • Gavin Michael Farrell is funded by ELIXIR-STEERS (101131096), supported through Silvio Tosatto’s BioComputingUP lab at the University of Padova.
  • Charlie Harrison and Reyer Zwiggelaar are supported by BBSRC through AIBIO-UK (BB/Y006933/1).
  • Virginia Portillo is supported by Responsible AI UK (EP/Y009800/1).
Loading 2608.17997v1…