Source-linked AI summary

The widening evaluation gap in medical large language model research 2023 to 2026

Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif

arXiv:2609.11770v1cs.CL

TL;DR

Medical model families advance faster than clinical evidence, raising whether evaluation remains current. The paper maps the literature, defines evaluation lag, and finds that lag widening is only partly offset by migration to newer systems, with randomized trials evaluating older models largely because of model selection.

  • Problem

    Model families are superseded every six to nine months, whereas randomized trials take one to two years from design to publication, creating concern that rigorous evidence may concern models no longer deployed.

  • Method

    The study builds a reproducible evidence map of indexed healthcare literature, defines evaluation lag as the interval from a study’s newest named model release to publication, and examines its variation by design and date.

  • Results

    Migration to newer systems offset 56.2% of the drift that the literature would otherwise accumulate, while randomized controlled trials evaluated models 4.62 quarters older than other empirical designs.

  • Takeaways & Limitations

    The gap between evaluation currency and research rigour reflects model selection, with current-model studies showing no design difference in the reported gradient.

  • Takeaways & Limitations

    The current-model randomized-trial subgroup is small, so the analysis does not establish that residual timeline effects are absent.

Abstract

from arXiv · show

Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing 45-fold; 2.5% used a randomised, controlled or prospective design. Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters. Because discontinued models age mechanically, we benchmarked this against a counterfactual holding model composition fixed: migration to newer systems offset only 56% of the drift (95% CI 50-65). Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19), yet among studies naming a model still under development no design differed from any other; 62% of randomised trials evaluated a discontinued family. Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.

1 Introduction

Medical language models are advancing faster than clinical evidence can be produced, creating a risk that rigorous studies assess superseded systems. The paper maps this temporal mismatch by defining evaluation lag and examining its variation across study designs and publication periods.

  • Motivation: Model families are superseded every six to nine months, whereas randomised trials take one to two years from design to publication.This divergence could leave clinical and regulatory decisions reliant on trials of systems no longer deployed.
  • Motivation: Existing reviews describe applications or fixed slices of the literature but do not measure the time between evaluated systems and publication.The paper identifies this temporal distance as the missing quantity.
  • Approach: The authors construct a reproducible evidence map of indexed generative-language-model research in healthcare from January 2023 through June 2026.Harvesting, screening, classification, and analysis run from deposited code so the map can be regenerated as the literature grows.
  • Study aims: The study defines evaluation lag as the interval between the newest model release named by a study and that study’s publication.It uses this measure to test whether lag widens, differs by design, and relates to model-reporting reproducibility.
  • Study aims: The study aims to characterise evidence-base growth and composition, quantify evaluation lag, assess model-specification precision, and locate divergence between research attention and trial-grade evidence.These aims connect temporal currency with study design, reporting, and the distribution of research effort.

2 Results

The healthcare literature expanded rapidly, but trial-grade evidence remained scarce while evaluation lag widened substantially. Randomised studies assessed older models mainly because they selected discontinued families, and model-version reporting was incomplete.

  • Evidence-base growth: 45-fold growth took the literature from 52 records in 2023-Q1 to 2,346 in 2026-Q2, while only 2.5% met the trial-grade design definition.Output increased at an incidence rate ratio of 1.272 per quarter, whereas 296 records were randomised, controlled, or prospective.
  • The evaluation gap: Mean evaluation lag rose from 1.33 quarters in 2023-Q1 to 6.08 in 2026-Q2, and newer-model migration offset only 56.2% of counterfactual drift.The observed slope was 0.329 quarters per quarter versus a fixed-composition counterfactual reaching 10.92 quarters.
  • The evaluation gap: Randomised controlled trials evaluated models 4.62 quarters older than other empirical designs, although the adjusted ordering weakened after Holm correction.The median-regression estimate was significant before correction, while the Holm-adjusted P value was 0.084.
  • The evaluation gap: Among studies naming actively developed families, no design differed materially from reference, while 62.1% of randomised trials named discontinued families.This indicates that the design gradient operates through model selection, although the current-model randomised subgroup is small.
  • Model reporting: Only 50.0% of randomised controlled trials specified a model version, leaving half of the randomised evidence unattributable to a determinate version.Comparative evaluations specified versions in 83.4% of cases, compared with 77.3% for preprints and 74.8% for observational studies.
  • Model landscape: Model use diversified sharply, with distinct families per quarter rising from 5 to 37 and the leading family’s share falling from 72.0% to 25.3%.The Herfindahl–Hirschman index declined 76%, while biomedical-specific models remained 1.63% of all mentions.
  • Attention–evidence gap: Scientific research support had the largest attention–evidence gap at 4.92, while medical dialogue summarisation had the smallest at 0.27.Benchmark validity and safety guardrails were also among the most evidence-poor themes, with 0.7% and 0.71% trial-grade shares.

3 Discussion

The evidence base is expanding but becoming less current, with the gap driven mainly by which models rigorous studies select rather than by research timelines alone. This creates a paired problem of outdated systems and incomplete model reporting, while model use is diversifying and trial-grade evidence remains uneven across domains.

  • Migration to newer systems offset only 56.2% of the evaluation-lag drift, showing that the literature is not keeping pace with model releases.The authors interpret this as a consequence of model selection rather than an irreducible cost of careful research.
  • Randomised trials specified model versions in only 50.0% of cases, below comparative evaluations at 83.4%, limiting reproducibility when proprietary systems disappear.A trial reporting neither version nor access date cannot be reproduced in principle because the evaluated system may no longer exist.
  • The field diversified substantially: model-family concentration fell 76% and the number of families in use rose sevenfold, largely through general-purpose systems.Purpose-built biomedical models never exceeded 2.6% of mentions in any quarter, so findings inherit the behaviour and release schedules of systems developed outside medicine.
  • The largest attention–evidence gaps occurred in scientific research support and agentic systems, where applications with greater autonomy had few trial-grade studies.Benchmark validity and safety guardrails ranked first and second among themes, yet were least often addressed with trial-grade designs.
  • The bibliometric analysis is limited by incomplete computer-science indexing, model-detection errors, uneven model naming, and lag estimates defined for only half the corpus.It describes what the literature evaluates and reports, not whether the evaluated systems are safe or effective.

4 Methods

This meta-research study maps indexed medical LLM literature using reproducible searches and deterministic metadata classification, then quantifies evaluation lag and its design differences.

  • Search: The study searched PubMed/MEDLINE for records published from 2023-Q1 through 2026-Q2 across fourteen clinical domains, deduplicating records by DOI or PMID.The technology block combined generic LLM terminology with named systems, and records could belong to multiple domains.
  • Scope and validation: Only PubMed was searched, and model-name matching was audited because family names can collide with clinical acronyms or anatomical abbreviations.Across eleven exposed families, 232 of 2,347 matches were flagged; flagged matches represented 2.0% of the corpus and an expected 2.5% of records with computable lag.
  • Characterisation: Design classes were assigned programmatically from PubMed publication types, with randomized, pragmatic, equivalence, clinical, and observational studies designated trial-grade.Publication quarter, country, venue, model families, and themes were also extracted using fixed deterministic rules and vocabularies.
  • Evaluation lag: Evaluation lag measures the quarters between publication and the latest release of any datable model family named by a record.The analysis used the latest available release for each family and then took the maximum across families named by the record.
  • Lag decomposition: The counterfactual holds baseline model-family composition fixed while allowing within-family lag to evolve, estimating how migration to newer systems offsets mechanical drift.The offset metric ranges from 0 for no change in studied systems to 1 for constant lag, with bootstrap confidence intervals over 2,000 record resamples.
  • Design differences: Design differences were estimated with median regression adjusted for publication quarter, using other empirical designs as the reference and Holm-adjusted contrasts.The analysis was repeated among studies naming actively developed families and separately reported a model-generation lag measure.

Data availability

Derived data supporting the study’s figures and statistics are openly deposited with the controlled vocabulary, release table, registration, and record index.

  • Data availability: The deposit contains derived tables, the controlled vocabulary, the hand-assigned model release table, study registration, and a record index for all 11,628 analysed records.The index includes PubMed identifier, DOI, publication quarter, design class, country, and detected model families.
  • Data availability: Users can reconstitute the corpus from PubMed using the record index and complete search strings, while titles and abstracts are not redistributed.A harvesting script performs the reconstruction unattended.

Code availability

The complete analysis pipeline and verification scripts are openly available and can regenerate the reported tables and statistics in a single pass.

  • Code availability: The deposited pipeline covers harvesting, deterministic characterisation, statistical analysis, figure generation, layout checks, and compliance checks.The archived version is available under the MIT licence.
  • Code availability: Executing the scripts regenerates every derived table and statistic, with verification scripts checking arithmetic closure, shared denominators, and figure layout.No generative model was used for screening, classification, extraction, or analysis.

Ethics

The study analysed published-report metadata without human participants, human material, or identifiable personal data, so ethics approval and informed consent were not required.

  • Ethics: The study analysed metadata from published reports and involved no human participants, human material, or identifiable personal data.Accordingly, no ethical approval or informed consent was required.

Use of generative artificial intelligence

No generative model was used to screen records, classify studies, extract data, perform analysis, or draft interpretive claims.

  • No generative model was used at any stage of screening, classification, extraction, analysis, or interpretive drafting.

Tables

Table 1 reports the principal estimates for the 11,628-record evidence map covering 2023-Q1 through 2026-Q2.

  • Table 1 reports principal estimates computed across 11,628 records first available between 2023-Q1 and 2026-Q2.Confidence intervals accompany estimated quantities, while evaluation lag and drift offset are defined by Eqs. (2) and (5).

Figures

The figures map the evidence base from study flow and workflow through growth, evaluation lag, model specification, landscape, thematic gaps, and geography.

  • Study flow and workflow: The evidence map contains 11,628 records, with study flow and fixed characterisation rules supporting analyses of composition, growth, lag, models, themes, and geography.Reviews and commentaries remain included, while trial-grade designs are highlighted for inference about clinical use.
  • Growth and composition: Corpus output grew from 52 records in 2023-Q1 to 2,346 in 2026-Q2, with an incidence rate ratio of 1.272 per quarter.
  • Evaluation, models, themes, and geography: The figures compare evaluation lag against a fixed-composition counterfactual and show design-specific lag distributions and adjusted differences.They also track model-version specification by year and design, model-family concentration, attention–evidence gaps, thematic communities, and geographic trial-grade shares.

Supplementary Information

The supplementary material documents the evidence-map protocol, search and characterisation procedures, analyses, reporting decisions, and limitations.

  • Study design: The study is a meta-research evidence map of published reports, not a treatment-effect synthesis.
  • Search and eligibility: PubMed/MEDLINE records were retrieved for 2023-Q1 through 2026-Q2 across fourteen domain blocks without design, language, or outcome restrictions.Reviews and commentaries were retained because study design was an analysis variable.
  • Characterisation: Eligibility and record characterisation used deterministic metadata rules, fixed patterns, and lookup tables, with no duplicate screening or duplicate extraction.
  • Analysis: Primary analyses covered evaluation lag, trial-grade share, version specification, model concentration, and the attention–evidence gap.Trends were modelled with negative binomial, median, and least-squares regression rather than pooled estimates.
  • Limitations: Key limitations include no risk-of-bias appraisal, PubMed indexing and metadata coverage, single-reviewer deterministic screening, regular-expression model detection, and a hand-assigned release table.

Supplementary Table S2 | Complete search strings

The supplementary search strategy combines fixed technology terms, fourteen clinical-domain blocks, and a 2023/01/01–2026/06/30 publication-date filter, with pooled records deduplicated by DOI or PMID. It also documents model-family release classifications, including 53 families divided into active and discontinued groups.

  • Search construction: Searches combine technology terms with separate clinical-domain blocks and the 2023/01/01–2026/06/30 publication-date filter.The domain queries were executed separately before pooling.
  • Search construction: Pooled search results are deduplicated using DOI, with PMID as the fallback identifier.
  • Clinical-domain coverage: Fourteen clinical domains span question answering, imaging, diagnostic reasoning, safety monitoring, decision support, agentic systems, patient-facing care, mental health, governance, documentation, and research support.The domain list also includes medical education and language translation.
  • Model-family classification: 53 model families were classified as 24 active and 29 discontinued using hand-assigned release quarters, with active status defined by a latest release in 2024-Q1 or later.Release quarters were assigned from developer announcements and model cards, and the authors identify them as the component most open to challenge.
  • Model-family classification: Family-status classification is defined at the family level rather than by study, so conditioning on status does not condition on the outcome.

Supplementary Table S4 | Controlled vocabulary

The controlled vocabulary applies fixed, case-insensitive rules to concatenated titles and abstracts, mapping model mentions and publication types into standardized analysis categories. Publication-type mappings use priority order, while additional patterns capture adaptation, agentic systems, benchmark validity, calibration, and deployment.

  • Classification rules: Fixed patterns and lookups classify records without using a generative model at screening, classification, or extraction.
  • Classification rules: Rules are applied case-insensitively to the concatenated title and abstract.
  • Model vocabulary: Model mentions are mapped into families and classes including closed generalists, open-weight generalists, biomedical specialists, biomedical multimodal systems, and extractive baselines.Examples include Claude, GPT, Gemini, LLaVA-Med, MedGemma, and BERT-family baselines.
  • Study-feature vocabulary: Additional vocabulary covers adaptation, agentic systems, benchmark validity, calibration, uncertainty, abstention, and deployment-related terms.
  • Publication-type mapping: Publication types are mapped to analysis classes by priority, with the first matching rule determining the classification.Randomized, pragmatic, and equivalence trials are grouped first as randomized controlled trials, followed by other clinical-trial and observational categories.
  • Publication-type mapping: Trial-grade classes comprise clinical trial, observational study, and randomized controlled trial.
Loading 2609.11770v1…