Source-linked AI summary

AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI Agents

Pengze Li, Cui Tao

arXiv:2608.28345v1cs.AI

TL;DR

Healthcare AI agents lack a unified way to report architecture, execution, provenance, and governance, despite the importance of these details for responsible interpretation. AGENT-O addresses this gap with a modular ontology, semantic Agent Card, and reporting-completeness workflow; its corpus assessment found evaluation and benchmark procedures more consistently reported than runtime architecture, governance, and provenance.

  • Problem

    Health-oriented AI-agent reporting lacks a unified framework for architecture, execution, paper evidence, and governance, while these details matter for reproducibility, safety, accountability, and generalization.

  • Method

    AGENT-O combines a modular OWL 2 ontology, ontology-conformant semantic Agent Card, and workflow for assessing publication-level reporting completeness.

  • Results

    Evaluation and benchmark procedures were reported more consistently than runtime architecture, governance, and provenance across the assessed corpus.

  • Takeaways & Limitations

    AGENT-O provides a reusable framework for structured agent representation and reporting-gap identification.

  • Takeaways & Limitations

    AGENT-O is a reusable framework rather than a complete or universal ontology, and reporting-completeness scores assess publication content rather than agent performance or clinical validity.

Abstract

from arXiv · show

AGENT-O is a modular ontology framework that defines a semantic Agent Card for representing health-oriented AI agent systems and supports assessment of reporting completeness in scientific publications. AGENT-O was developed as an OWL 2/RDF ontology covering runtime, models, workflow, tools, clinical use, evaluation, provenance, governance, and reporting assessment. Evaluation included ontology inventory, OWL-RL reasoning, three SHACL suites, 12 SPARQL competency queries, three cases, and model-assisted reporting-completeness assessment of 279 papers across five dimensions. The ontology contained 1,962 RDF triples and 1,922 Protege axioms, with 252 active classes, 198 active object properties, and 51 datatype properties. All SHACL suites conformed on example graphs, all competency queries returned prespecified evidence, and all 279 papers were scored. Incomplete reporting was highest for runtime/architecture (84.6%), governance/safety (82.8%), and provenance/reproducibility (78.1%), compared with evaluation (25.8%) and benchmark-process alignment (29.8%). AGENT-O supported semantic Agent Card representation and reporting assessment while revealing an evaluation-specification gap: evaluation and benchmark procedures were reported more consistently than runtime architecture, governance, and reproducibility. AGENT-O provides a reusable ontology, semantic Agent Card profile, and reporting-completeness workflow for structured reporting and gap identification, but does not assess agent quality or deployment readiness.

Background and Significance

Healthcare AI agents operate through models, workflows, tools, actions, and human interactions, but existing reporting resources do not represent these integrated systems. AGENT-O addresses this gap with a reusable semantic framework and reporting-completeness workflow.

  • Existing health AI reporting frameworks generally emphasize data, model development, performance, and clinical evaluation rather than agent operation.
  • Transparent reporting must specify an agent’s operational boundaries, accessed tools, authorized actions, and points for human review, override, or escalation.These details support assessment of reproducibility, safety, accountability, and generalizability conditions.
  • Existing model cards, datasheets, and related resources were not designed to represent an integrated agent with runtime actions, artifacts, human interactions, and governance constraints.Consequently, publications may leave agent design, operation, provenance, and governance insufficiently specified.
  • No single existing resource unified representation of an agent’s architecture, execution process, paper-based evidence, and governance-relevant information.
  • AGENT-O represents health-oriented AI agents across components, workflows, execution context, models, clinical use, evaluation, provenance, and governance.Its semantic Agent Card is an ontology-conformant reporting profile, and its workflow identifies whether publications provide information needed for reproducible, queryable annotation.
  • AGENT-O is intended as a reusable framework rather than a universal ontology and was evaluated for structured representation, publication-level gaps, and semantic interoperability.

Scope and Design Principles

AGENT-O is a modular OWL 2 ontology designed to represent agents and the reporting elements needed to describe them consistently. Its design emphasizes reproducibility-oriented detail, reuse of established concepts, explicit boundaries, and queryable governance, provenance, evidence, and reporting gaps.

  • AGENT-O uses a modular OWL 2 design to represent agent systems and the reporting elements needed for structured, consistent description.
  • Its scope covers identity, runtime architecture, model roles and deployments, workflow, planning, reasoning, memory, tools, execution, outputs, and proposed actions.
  • The ontology also covers evaluation, provenance, governance, clinical context, health-data interoperability, reporting assessment, and evidence.
  • An Agent Card is an ontology-conformant report describing an AgentSystem and assembling its relevant reporting elements.
  • The design separates system entities from publication descriptions and assessment activities while keeping governance, provenance, evidence, and reporting gaps queryable.
  • Development combined conceptual modeling, alignment with established resources, semantic validation, competency queries, case and corpus assessment, and feedback-driven refinement.

Data sources

AGENT-O drew on agent literature, established biomedical and governance resources, and a curated review-derived corpus. The corpus contained 279 records but was not intended to exhaustively cover the literature.

  • The framework’s initial concept inventory came from definitions and descriptions of AI agents in survey and review papers.
  • Biomedical, provenance, software, reporting, and governance ontologies and standards supplied concepts for reuse or alignment.These included PROV-O, DUO, FHIR RDF, IAO, SWO, MCRO, NIST AI RMF, and WHO guidance.
  • Reviewed resources contributed semantics for software, model reporting, governance, health AI principles, statistical methods, and human–AI processes.
  • The curated corpus comprised 278 extracted paper documents plus the prespecified AgentArena case, for 279 records.
  • The corpus was derived from review literature rather than an independent systematic bibliographic-database search and was not intended to exhaustively cover health-oriented AI-agent research.

Ontology design

AGENT-O links six semantic domains into an ontology for representing runtime systems, evaluation, provenance, governance, clinical context, and reporting assessment. Its distinctions and external mappings are designed to preserve interoperable, queryable boundaries.

  • AGENT-O represents agent systems through six linked domains spanning runtime and models, evaluation, provenance, governance and safety, clinical context, and reporting assessment.
  • The core module distinguishes functional model roles, versioned model specifications, and runtime deployments, while linking execution runs to deployments.Generated outputs are distinct from proposed actions, and the three model classes are disjoint.
  • Evaluation, provenance, governance, and clinical modules connect studies, artifacts, policies, risks, human review, use contexts, data permissions, and operational safeguards.
  • AGENT-O separates model-level intended use, system-level clinical use, manuscript statements, and assessment records from the external workflows that produce them.
  • External alignment maps AGENT-O to established provenance, data-use, clinical-data, information-artifact, software, reporting, risk-management, and safety resources.
  • Runtime, evaluation, provenance, clinical, and governance properties make system structure, execution configuration, evidence, lineage, oversight, and safety information queryable.

Evaluation Methods

AGENT-O was evaluated through complementary technical-validation, semantic-query, case-based, and corpus-level reporting-completeness procedures. These procedures tested ontology integrity, inferential behavior, representability, and publication-level reporting.

  • Evaluation design: Four complementary procedures assessed formal integrity, OWL-RL reasoning with SHACL validation, SPARQL competency queries, and case- and corpus-level reporting completeness.The first three procedures assessed technical integrity and inferential behavior, while coverage evaluation tested representability and publication completeness.
  • Technical validation: Independent parsing checked modules, profiles, alignments, shapes, and example graphs for namespaces, labels, domains, ranges, alignment terms, and consistency.RDF serialization and Protégé axiom counts were reported separately.
  • Semantic validation: OWL-RL reasoning materialized subclass, subproperty, domain, range, equivalence, inverse-property, and provenance inferences before competency-query execution.Twelve SPARQL queries tested model, clinical-use, governance, evidence, provenance, assessment, and reporting-scope relations.
  • Constraint validation: Three SHACL suites tested architecture, governance, and reporting profiles against corresponding example graphs, rather than global ontology completeness.The suites covered model-layer distinctions, oversight and risk controls, report structure, evidence, assessor provenance, and completeness records.
  • Case evaluation: Three cases tested whether benchmark, diagnostic, and multi-agent decision-support reports could be represented across runtime, evaluation, provenance, governance, and benchmark alignment.The cases included AgentArena, MedAgent-Pro, and a multi-agent medical decision consensus matrix system.

Results

AGENT-O passed its technical validation checks and represented all three evaluated cases, while the corpus assessment produced completeness scores for all 279 papers. Reporting gaps were concentrated in runtime architecture, governance and safety, and provenance rather than evaluation and benchmark-process alignment.

  • Ontology inventory: 1,962 RDF triples and 1,922 Protégé axioms comprised the evaluated release, with 252 active classes, 198 active object properties, and 51 datatype properties.The active inventory spanned six modules, alongside 183 external mappings and three application profiles.
  • Technical validation: All 25 Turtle files parsed successfully, with no deprecated predecessor namespaces, missing labels or active-property domains and ranges, or undefined alignment terms.Automated validation found no parse failures or the listed consistency problems.
  • Semantic validation: All architecture, governance, and reporting SHACL suites conformed on their corresponding example graphs, and all 12 competency queries returned prespecified evidence.Returned evidence covered model structure, clinical use, FHIR inputs, governance, model-card alignment, provenance, and assessment records.
  • Case results: All three cases were representable but differed in reporting completeness across benchmark alignment, medical diagnosis, and multi-agent decision support.AgentArena was partial for provenance and governance; MedAgent-Pro was partial for governance; the consensus case satisfied all five dimensions.
  • Corpus results: 63.7/100 mean and 67.5/100 median completeness scores were returned for all 279 papers; 72 papers scored 50–64.9 and 130 scored 65–79.9.The workflow produced no failed or unscored cases.
  • Reporting gaps: 84.6% runtime/architecture, 82.8% governance/safety, and 78.1% provenance/reproducibility were incomplete, versus 25.8% evaluation and 29.8% benchmark-process alignment.The corpus therefore documented evaluation and benchmark procedures more completely than runtime architecture, governance, and provenance.

Discussion

AGENT-O combines a modular ontology, semantic Agent Card profile, and reporting-completeness workflow for representing health-oriented AI agent systems and identifying documentation gaps. Its corpus assessment found evaluation and benchmark procedures reported more completely than runtime architecture, governance, and provenance, while scores assess publication content rather than system quality or deployment readiness.

  • Framework and implications: AGENT-O combines a modular ontology, semantic Agent Card profile, and reporting-completeness workflow.The framework represents health-oriented AI agent systems and supports structured, queryable reporting assessment.
  • Framework and implications: The ontology represents runtime, models, workflow, tools, clinical use, evaluation, provenance, governance, and evidence.It distinguishes model roles from specification and deployment, clinical use from manuscript statements, and assessment representation from the external workflow.
  • Reporting findings: Evaluation and benchmark procedures were reported more completely than runtime architecture, governance, safety, and provenance.Missing details included system configuration, human review, fallback, uncertainty, privacy, security, compliance, and traceability.
  • Limitations and future work: AGENT-O is a reusable framework rather than a complete or universal ontology, and evolving architectures, clinical uses, and governance requirements require continued refinement.Its selective alignments support interoperability but do not replace or fully capture the semantics of aligned resources.
  • Framework and implications: Reporting-completeness scores assess publication content, not agent performance, clinical validity, fairness, safety, or deployment readiness.This distinction supports checklists, curation schemas, reviewer aids, repository metadata, and governance registries without conflating documentation with performance.
  • Limitations and future work: The corpus was derived from two review-associated inventories rather than an independent systematic search, which may affect observed reporting-gap frequencies.The label-blinded GPT-5.1 workflow used source verification and deterministic checks, but labels were not adjudicated by multiple human reviewers.

Conclusion

AGENT-O provides a reusable ontology, semantic Agent Card profile, and workflow for representing health-oriented AI agent systems and assessing publication-level reporting completeness. The corpus showed stronger reporting of evaluation and benchmark procedures than runtime architecture, governance, and provenance, while the framework supports structured annotation and cross-study comparison without implying agent performance or deployment readiness.

  • Conclusion: AGENT-O provides a reusable ontology, semantic Agent Card profile, and workflow for representing health-oriented AI agent systems.It assesses publication-level reporting completeness.
  • Conclusion: Evaluation and benchmark procedures were reported more completely than runtime architecture, governance, and provenance.This is the paper's reported evaluation–specification pattern.
  • Conclusion: Making system specifications, clinical use, evidence, and governance constraints queryable supports structured annotation and cross-study comparison.The framework does not imply agent performance or deployment readiness.
Loading 2608.28345v1…