Source-linked AI summary

A Semantic Model of Genetic Evidence: A Step Toward Bridging the Basic-Science-Clinic Gap

Michael Bouzinier, Dmitry Etin

arXiv:2609.04509v1cs.DBcs.AIq-bio.GN

TL;DR

Existing standards do not capture the fine-grained structure of basic and pre-clinical genetic evidence needed for literature-based decision-making. The paper introduces a FHIR-aligned semantic model with structured credibility and conditional dimensions, then pilots it across six genetics papers. The resulting framework is presented as a feasibility-study increment toward trustworthy, AI-ready infrastructure, with explicit limits on generalizability and clinical deployment.

  • Problem

    Existing standards provide limited support for the heterogeneous, fine-grained genetic evidence from basic and pre-clinical research needed for structured clinical interpretation.

  • Method

    The paper develops a three-class genetic-evidence model aligned structurally with FHIR and ontology standards, with dimensional credibility decomposition and conditional activation rules.

  • Results

    The model was applied to six genetic-evidence papers through a human–AI annotation workflow under curator supervision.

  • Takeaways & Limitations

    The work offers a reference data model and validation schema for representing genetic evidence toward trustworthy, AI-ready variant-interpretation infrastructure.

  • Takeaways & Limitations

    Six papers and one annotator do not support claims about inter-annotator agreement or generalizability, and the pilot is not a clinical deployment.

Abstract

from arXiv · show

Scientific and clinical decision-making depends on evidence from the primary literature, but existing standards for representing that evidence (FHIR Evidence, ECO, SEPIO, and the GA4GH Genomic Knowledge Standards) are oriented toward clinical-trial workflows, evidence codes, or single-variant assertions, and do not capture the fine-grained, domain-specific structure of claims in basic and pre-clinical research. We introduce a semantic model for scientific evidence with three core classes, specialize it for genetics, align it structurally to FHIR Evidence with a SEPIO-anchored credibility decomposition, and attach a compact dimensional vocabulary whose conditional-activation rules are validated by a SHACL schema for the implemented constraints. Using clinical variant interpretation as the driving use case, we evaluate the model through a human-AI annotation pilot over six genetics papers, yielding 28 evidence items and 95 source-anchored assertions, with a workflow that keeps curator-authored reference annotations distinct from AI-drafted annotations. Treating the pilot as a feasibility study rather than a benchmark, we argue that the model is a useful increment toward trustworthy, AI-ready infrastructure for variant interpretation: a reference data model and validation schema for representing genetic evidence.

1 Introduction

Existing evidence standards support interoperability but do not fully represent fine-grained basic-science and pre-clinical genetic evidence. This paper introduces a genetics-focused semantic model, aligned with FHIR and ontology standards, and demonstrates it through a supervised annotation pilot.

  • Motivation: Current standards emphasize clinical trials, high-level assertions, or variant representations, leaving limited support for fine-grained basic-science evidence.This gap constrains representation of heterogeneous evidence needed for automated reasoning and AI-ready infrastructure.
  • Motivation: Clinical use of literature-derived genetic evidence requires standardized representations of relevance, specificity, and confidence.These dimensions concern how precisely findings map to patients, disease contexts, or biological hypotheses and whether evidence is robustly usable.
  • Contributions: The paper contributes a domain-analysis-driven genetic-evidence model with three core classes, genetic specializations, a compact vocabulary, and SHACL validation.The model is designed to make source-anchored scientific claims suitable for structured annotation and automated conformance checking.
  • Evaluation: The evaluation applies the model to six genetic-evidence papers through extraction tooling and a human–AI annotation workflow under curator supervision.The workflow separates curator-authored reference annotations from AI-drafted annotations.
  • Relation to existing standards: The model structurally generalizes FHIR Evidence from clinical-trial workflows to basic science while aligning dimensions with ECO, SEPIO, IAO, and OBI.Genetics is developed as the first specialization, with term-level crosswalks documenting verified mappings and remaining gaps.

3 A Semantic Model of Genetic Evidence

The model represents genetic evidence through core classes, dimensional annotations, and explicit conditional requirements. It separates evidence variables from claims and credibility, while using structured facets and validation rules to support consistent annotation.

  • Model structure: The model combines core classes, a roughly twenty-axis dimensional vocabulary, and conditional activation rules for semantically appropriate dimensions.Conditional requirements keep the vocabulary compact while preserving expressive coverage.
  • Core classes: ScientificEvidence records source-anchored object-level claims, and one paper may produce multiple evidence items for distinct evidential claims.Examples include case–control associations and in vivo functional assays concerning the same variant.
  • Core classes: EvidenceVariable represents named quantities such as allele frequency, LOD score, or odds ratio, while EvidenceAssertion serves as a meta-predicate for decision rules.The pilot records object-level assertions but does not instantiate EvidenceAssertion in its reasoning-layer role.
  • Dimensional vocabulary: Six dimensions are always required: Knowledge Domain, Method, Target Type, Resolution, Credibility, and Phenotype Scale.Together they identify the evidence type, acquisition method, subject, resolution, and contextual credibility.
  • Credibility: Credibility is decomposed into an overall ordinal rating and separate facets for power or sample size, replication, multiple-testing control, population-stratification control, and ascertainment.Separate facets allow clinical and research users to weigh methodological factors differently.
  • Conditional activation: Conditional dimensions are activated by values of other dimensions, such as requiring inheritance mode for relevant pedigree evidence but not for polygenic scores.The explicit conditions make requirement status transparent to curators and automated validators.

4 Illustrative Examples from the Literature

The six-paper pilot spans variant-level association, functional genetics, and polygenic-score evidence, showing both broad coverage and representational strain for derived whole-genome scores.

  • Scope: Six papers cover Mendelian genetics, population association, case–control resequencing, functional genetics, and genome-wide polygenic scores.Four publications were annotated manually and two with AI assistance.
  • Duerr et al. 2006: Duerr’s study yields seven GeneticEvidence items spanning discovery, replication, transmission, fine-mapping, negative results, and supporting functional or mouse-model evidence.Six of the seven items fit the current schema without strain.
  • Duerr et al. 2006: The cited anti-p40 antibody trial reveals a translational-scope gap because the model lacks a fitting genetic Knowledge Domain for drug-response evidence.Three related candidate extensions remain provisional under the two-papers rule.
  • Inouye et al. 2018: Inouye’s evidential target is a derived whole-genome score with epidemiological measurements rather than a single variant.The study combines three component scores and evaluates performance in UK Biobank participants.
  • Inouye et al. 2018: The Inouye annotation produces four evidence items, but forced placeholder assignments expose six candidate extensions for score targets, resolution, composition, measurements, cohorts, and derived artifacts.The model records these limitations explicitly rather than leaving them implicit.

5 Human–AI Annotation Workflow and Evaluation

The workflow separates curator-authored reference annotations from AI-drafted annotations while enforcing source anchoring, schema validation, explicit provenance, and conservative schema extension.

  • Evaluation scope: The pilot is a feasibility study, not a benchmark: six papers and one annotator cannot support inter-annotator agreement or generalizability claims.The evaluation demonstrates implementability of the workflow’s epistemic commitments rather than broad performance.
  • Annotation tracks: Two annotation tracks use curator-authored references and AI drafts, with both validated against the same SHACL shapes and curator authority over recorded annotations.AI drafts use the source PDF, schema, and manual annotations as exemplars before curator review.
  • Source anchoring: Every assertion carries a source span, and all 95 pilot assertions are source-anchored under a mandatory validation requirement.Each span identifies a page and a short searchable phrase from the source text.
  • Flags and provenance: AI drafts record uncertainties and assumptions for review, while manual annotations use disagreement, suggestion, and query flags to preserve distinct epistemic provenance.The two tracks therefore represent primary and secondary knowledge work rather than interchangeable annotations.
  • AI review: Nine AI-drafted items were rejected zero times: one was approved as drafted, three with non-blocking flags, and five were edited.The later Inouye draft required fewer edits than the earlier Duerr draft, but the papers and protocol versions differ.
  • Extension promotion: The extension rule promotes a representational gap only after independent use in a second paper; the pilot promoted Phenotype Scale and Variant Ascertainment.Fifteen other candidates remain unpromoted pending further evidence or curator decision.

6 Discussion

The discussion presents a compact, auditable model whose SHACL constraints and source anchoring support structured genetic evidence, while documenting coverage limits and unresolved representational gaps.

  • Model properties: Conditional activation keeps the vocabulary compact, while source-span anchoring makes every assertion independently auditable.SHACL shapes formalize implemented activation conditions and make them machine-checkable.
  • Extensions: Two candidate dimensions were promoted to first-class status, while fifteen further extensions remain unpromoted under the two-papers rule.The promoted dimensions are Phenotype Scale and Variant Ascertainment.
  • Coverage: Of ten evidence patterns, five are represented natively, three are documented forced fits, and two lie outside the model’s intended scope.The outside patterns are therapeutic-response evidence and cross-publication synthesis.
  • Limitations: The pilot does not support claims about inter-annotator agreement or generalizability and stops short of live clinical deployment.The authors therefore characterize it as a feasibility study rather than a benchmark.
  • Crosswalk limitation: The UMLS crosswalk maps some intersective GEM tokens to the nearest available single concept, which cannot express their full compound meaning.Statistical_genetics is anchored to Biostatistics with a narrower relation.
  • Crosswalk limitation: Concept-level post-coordination is identified as a future solution, but UMLS lacks an aggregate-level mechanism for compositional mappings.Composition remains a feature of individual source vocabularies rather than UMLS itself.
  • Resolution limitation: The current Resolution enumeration presupposes genomic coordinates even though genetic evidence may use genomic, transcript, or protein coordinates.The paper records variant_in_transcript as a curator-surfaced candidate value awaiting independent use.

8 Conclusion

The paper introduces a compact semantic model for genetic evidence, combining structured classes, conditional dimensions, SHACL validation, and provenance-aware annotation workflows. It presents these elements as a concrete step toward trustworthy, AI-ready infrastructure for variant interpretation.

  • The model combines core classes, conditional dimensional rules, and SHACL validation for genetic evidence.Three activation rules are currently machine-enforced in SHACL.
  • The workflow separates curator-authored reference annotations from AI-drafted annotations and requires independent confirmation before promoting candidate schema extensions.This preserves review authority while documenting possible extensions.
  • The contribution is positioned as a reference data model and validation schema for representing genetic evidence from biomedical literature.

Background and Related Work

The coverage analysis compares GEM with existing standards across genetic evidence patterns, showing that GEM represents several basic-science patterns natively while exposing specific forced fits and scope boundaries. The comparison is a coverage inventory rather than a ranking.

  • The matrix records coverage across standards without ranking them, because each standard was developed for different purposes and intended scopes.
  • Pedigree segregation: GEM represents pedigree segregation natively through Mode of Inheritance and Mendelian Segregation, although an activation-scope gap is logged as CE-DU4.The examples include X-linked recessive and autosomal recessive inheritance.
  • Case–control burden: GEM represents case–control burden natively with population, statistical-genetics, and gene dimensions, while structured cohort fields remain a candidate extension.Case and control counts are currently recorded as free text.
  • Common-variant association: GEM represents common-variant association through GWAS, Association Study, and Fine Mapping, while direction of effect for protective alleles remains a documented gap.The pattern is exercised by multiple Duerr evidence items.
  • Functional and animal-model evidence: GEM represents functional assays and animal models natively, using method and conditional organism or knockout dimensions, while resolution gaps remain.Jossin items expose protein-subdomain and organism-specific modeling requirements.
  • Relations and variant identity: GEM supports gene–gene relations and GA4GH-based variant identity, but transcript-coordinate resolution and related identifier bindings remain extension or implementation needs.The Nelson insertion exposed a forced fit across coordinate systems.

The annotation and review protocol suite

The protocol suite formalizes manual and AI-assisted annotation under a shared schema, with source anchoring, validation, explicit uncertainty flags, and curator-controlled review. The pilot preserves curator annotations as authoritative references rather than statistical gold standards.

  • A mode-agnostic protocol defines valid annotations, while autonomous and interactive specializations and a separate review protocol govern different workflows.Requirements include classes, dimensions, source anchoring, and flag semantics.
  • Manual track: Manual annotation converts highlighted passages and summary judgments into YAML conforming to the GeneticEvidence model.Normalization changes preserve original text and are recorded for later audit.
  • Validation: Manual annotations are validated against SHACL, and missing required or conditionally activated values must be resolved before corpus admission.
  • Reference annotations: The four manual annotations are authoritative references for evaluating AI drafts, but are not statistical gold standards because the corpus has one annotator.
  • AI-drafted track: AI drafts use the same schema and are reviewed assertion by assertion for source fidelity, schema conformance, and flag appropriateness.The curator, not the model, determines the recorded annotations and logs corrections.
  • Provenance: Every assertion must be independently auditable through a source span, and all 95 pilot assertions carry source spans by validation requirement.Paraphrase flags document cases where exact phrases could not be reconstructed.

Executable skills

Executable skills package annotation and review protocols so an AI assistant can apply the same rules as a human curator under supervision. Their stated purpose is reproducibility and reuse, not autonomy.

  • A skill is a version-controlled package of instructions and supporting resources that an AI assistant loads to perform a defined task repeatably.
  • Skills are the executable counterparts of written protocols, translating human task requirements into instructions an assistant can apply directly.
  • The mechanism is not specific to one model or vendor, although the accompanying skills are implemented as Claude Agent Skills.
  • The annotation and review skills package the corresponding protocols and apply them under curator supervision, with the curator retaining authority over every recorded annotation.The skills are published for reproducibility and reuse rather than autonomy.

AI-track review outcomes

The review compared two AI-drafted annotations against a curator protocol, distinguishing source fidelity and schema conformance from scientific adjudication. The drafts captured direct claims more reliably than synthetic conclusions or cited translational evidence.

  • Review process: Two AI-drafted annotations were reviewed assertion by assertion against REVIEW_PROTOCOL v1.0, with final adjudication by the human curator.The review assessed source fidelity and schema conformance rather than the underlying science.
  • Review outcomes: Two of Inouye’s three items were accepted without edits and one was edited, whereas four of Duerr’s six items were edited.The comparison is suggestive rather than controlled because the papers differ in genre and difficulty.
  • Review outcomes: All thirty-seven cited Inouye spans were verbatim and within the length limit, whereas the earlier Duerr draft had a different provenance record.
  • Omitted evidence: Reviewer-added items were higher-level or interpretive claims, including Inouye’s score-combination claim and Duerr’s cited anti-p40 antibody trial.Drafts were more reliable for specific, directly stated claims than for synthetic conclusions and cited translational evidence requiring cross-item interpretation.
  • Schema review: Existing-item edits primarily concerned schema-level judgments about knowledge domain, method placement, credibility calibration, and variant-ascertainment applicability.
  • Follow-up: Both reviews recommend a second pass because each paper gained a new item and several existing items were edited.The promotion record also requires a separate update after Variant Ascertainment was no longer exercised by Inouye.

Standards crosswalk

The crosswalk extends GEM’s structural alignment with FHIR into term-level mappings while documenting semantic gaps and preserving the most faithful external vocabulary matches. It also distinguishes object-level source claims from the EvidenceAssertion meta-predicate.

  • Term-level alignment: The standards crosswalk extends class-level GEM–FHIR alignment with term-level mappings for classes, properties, dimensions, and selected values.Mappings cover SO, HPO, NCBITaxon, ECO, SEPIO, and GA4GH VA, with identifiers verified against source vocabularies.
  • Mapping semantics: The crosswalk uses relation labels such as exact, close, narrower, broader, related, and none, plus structural labels for non-equivalent correspondences.
  • Mapping policy: Mappings prioritize semantic fidelity rather than a fixed vocabulary preference, retaining specialized exact matches when broader preferred-vocabulary terms would be less faithful.Animal Genetics maps exactly to an NCI Thesaurus term because MeSH has no corresponding descriptor.
  • Coverage caveat: The UMLS variant mapping is recorded as close because one concept merges population-level and entity-level senses, while its genomics-relevant entity sense lacks a generic entity-typed parent.
  • Assertion semantics: The assertion property records an object-level source-anchored claim, whereas EvidenceAssertion is a meta-predicate over a decision rule.The property aligns structurally more closely with GA4GH VA Statement and SEPIO Assertion than with FHIR’s human-readable Evidence.assertion.
  • Coverage gaps: Knowledge Domain, Phenotype Scale, Variant Ascertainment, Resolution, and Gene Relation have no single external equivalent in the reviewed crosswalk.These dimensions are characterized as GEM-specific epistemological axes rather than failed duplications of existing vocabularies.

Case reports: Duerr 2006 and Inouye 2018

The case reports test GEM on conventional variant evidence and on a whole-genome polygenic score. Duerr largely fits the schema, while Inouye exposes target, resolution, measurement, and cohort gaps requiring provisional extensions.

  • Duerr 2006: Duerr’s study identifies IL23R as an inflammatory bowel disease gene using GWAS, replication in a Jewish cohort, and family-based transmission testing in 883 nuclear families.
  • Duerr 2006: Six of Duerr’s seven items fit the schema without strain; the seventh, an anti-p40 antibody trial, reveals a translational-scope gap.The trial concerns drug-response evidence without a fitting genetic Knowledge Domain.
  • Duerr 2006: The Duerr annotation records rs11209026 in IL23R as a clinical, variant-level GWAS finding with high credibility and replication evidence.The discovery item includes corrected P = 1.56e-3 and OR = 0.26, followed by independent replication.
  • Duerr 2006: Three provisional Duerr extensions cover direction of effect, cross-item therapeutic synthesis, and variant-level activation of Mendelian dimensions.A separate negated-assertion extension was retracted because assertions can already express absence or presence.
  • Inouye 2018: Inouye developed and externally validated a coronary-artery-disease metaGRS from 1,745,180 variants in 482,629 UK Biobank participants.Its evidential target is a derived whole-genome score measured with epidemiological quantities such as hazard ratios and C-indices.
  • Inouye 2018: All four Inouye items require deliberate placeholder assignments because the current model targets variants rather than polygenic scores.The v1 annotation records these as forced fits and candidate extensions rather than leaving the mismatch implicit.
  • Inouye 2018: The Inouye extensions propose a score target, whole-genome aggregate resolution, target composition, epidemiological measurement targets, structured cohorts, and natural-versus-derived-artifact status.The score target is described as a derived composite predictor rather than a gene or variant.

Dimension coverage over the annotated corpus

Coverage analysis over 28 GeneticEvidence items shows strong population of applicable dimensions but exposes gaps in Knowledge Domain, Variant Ascertainment, Gene Relation, Knockout Type, and activation conditions.

  • Corpus coverage: 28 GeneticEvidence items were analyzed for dimension population and conditional-dimension applicability using the released YAML annotations.
  • Always-required dimensions: All always-required dimensions were populated by construction except Knowledge Domain, which was populated in 27 of 28 items.The exception is a cited anti-p40 antibody trial with no fitting genetic Knowledge Domain.
  • Conditional dimensions: Variant Ascertainment was populated in 14 of 18 applicable items, while the four Inouye items marked it not applicable for polygenic-score targets.
  • Conditional dimensions: Organism was populated in 8 of 8 applicable items and Measurement Target in 10 of 11, indicating generally complete coverage when their conditions hold.
  • Conditional dimensions: Gene Relation was populated in 1 of 11 applicable items and Knockout Type in 1 of 6, exposing enumeration and over-broad-activation gaps.
  • Activation conditions: Mode of Inheritance and Mendelian Segregation were each populated in 0 of 1 applicable items because the current activation condition was broader than the meaningful context.The single human-genetics-plus-gene item marked both dimensions not applicable.

UMLS crosswalk and illustrative OMOP CDM positioning

The UMLS crosswalk maps Genetic Evidence Model dimensions and values to semantic types and Metathesaurus concepts, while recording adjudicated gaps when no faithful mapping exists. Its mappings are generated by a harness and curator review, and remain distinct from the structural standards crosswalk.

  • Crosswalk scope: The UMLS crosswalk is distinct from the structural standards crosswalk and does not replace it.The structural alignment separately connects the model with FHIR Evidence and related vocabularies.
  • Crosswalk provenance: A concept identifier is asserted only after retrieval and confirmation against UMLS.The harness uses the UTS REST API or a locally indexed licensed copy, with curator adjudication recorded separately.
  • Crosswalk design: Each dimension axis is bound to a UMLS Semantic Type, and each value records an accepted Metathesaurus concept plus its relation to the GEM token.Relations include exact, close, narrower, broader, and related; curators may accept an out-of-scope concept with a recorded rationale.
  • Adjudicated gaps: Unmapped values are adjudicated conclusions supported by rejected near misses or a documented search protocol, not merely missing mappings.The relation none indicates that no faithful UMLS mapping was accepted.
  • Illustrative mappings: The illustrative mappings cover genetic, spatial, entity, research-activity, genetic-function, and molecular-function dimensions, alongside boolean or free-text dimensions without semantic types.Examples include GENE, VARIANT, POSITION, GWAS, inheritance, penetrance, expression, and regulation values.

Credibility tiers: definitions, the UMLS finding, and external alignment

GEM defines credibility as a defeasibility-based belief-revision policy for individual basic- and pre-clinical evidence items, separate from study design and other evidence facets. A UMLS sweep found no adequate proxy, while external alignments are conceptual and non-positional, with lossy FHIR export.

  • Tier definitions: GEM credibility tiers specify what would lead a curator to stop accepting an evidence item, rather than encoding study design or author-expressed certainty.The four levels range from very_high, accepted by default unless opposed by a stronger argument, to low, which is recorded pending independent replication.
  • Tier definitions: Medium evidence requires corroboration from an independent source or orthogonal evidence, whereas high evidence is questioned by comparable-strength contradiction.Very_high remains the default unless the opposing argument is stronger; low evidence is not relied upon.
  • Credibility decomposition: Credibility is anchored on SEPIO confidence, with cohort size, replication, multiple-testing control, ancestry control, and ascertainment informing—but not replacing—the overall judgment.These facets remain separate rather than being collapsed into one undifferentiated assessment.
  • UMLS finding: 174 UMLS queries returned 497 distinct concepts locally and 505 through the UTS REST API, but the four credibility tiers had no adequate UMLS proxy.The adequacy criteria required epistemic warrant, ordered scale points, disjoint levels, and domain-appropriate meaning.
  • UMLS finding: The nearest UMLS candidates failed because they represented measured-property degrees, study-design hierarchies, respondent confidence, or diagnostic likelihood rather than GEM defeasibility.The paper also rejects lexical post-coordination as adding no semantic content beyond the text string.
  • External alignment: External alignments are conceptual rather than equivalent or numerically convertible, and FHIR export is lossy because very_high is represented as high.GEM does not adopt GRADE’s rating rules, and its facets target individual evidence items rather than bodies of clinical evidence.
Loading 2609.04509v1…