Source-linked AI summary

An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study

James Matheson, Betsy Castillo, Andrew Y. Shin, David Scheinker

arXiv:2608.20373v1cs.CL

TL;DR

The paper asks whether LLMs can accurately abstract clinical registry data from unprocessed EMR documents, where information is fragmented and question types differ in ambiguity. It uses a pilot-and-validation evaluation with consensus abstractor references and an ambiguity taxonomy, finding that accuracy declines as clinical reasoning demands increase. The results support category- and question-level evaluation rather than relying only on aggregate performance.

  • Problem

    Existing LLM evaluations often use clean, narrowly defined questions and aggregate metrics, leaving limited evidence about performance on fragmented, unprocessed EMR data across registry question types.

  • Method

    The study used a two-phase evaluation across two ACC/NCDR registries, with pilot-based document targeting, consensus abstractor references, and six categories ordered by ambiguity and clinical reasoning demand.

  • Results

    Accuracy fell from 96.1% on simple binary questions to 62.0% on questions requiring the most clinical reasoning to resolve ambiguity.

  • Takeaways & Limitations

    LLM registry-abstraction performance should be reported and interpreted across question and ambiguity categories rather than through aggregate accuracy alone.

  • Takeaways & Limitations

    The studies included relatively few patients, so per-question estimates were wide, and the reference standard reflected credentialed abstractors whose applicability to other registries requires evaluation.

Abstract

from arXiv · show

Objective: To evaluate large language model (LLM) performance on unprocessed electronic medical record (EMR) data for clinical registry abstraction. Methods: We evaluated LLM performance answering registry questions for the American College of Cardiology National Cardiovascular Data Registry (ACC NCDR). In a pilot study at an academic medical center, the model identified candidate data sources for each registry question and experienced abstractors used these results to define question-specific document sets. In a validation study at a second center with a second ACC NCDR registry, the LLM answered questions using the question-specific document sets. Before reviewing any output, two abstractors independently established the ground truth and assigned each question to one of six categories, ordered by the ambiguity and clinical reasoning required to resolve it: Medication/Event Flag, Binary Clinical Presence, Administrative, Quantitative Laboratory/Physiologic, Clinical Interpretation, and Event Timing. Results: The analytical sample comprised 9,430 abstractor answers reconciled to 4,715 consensus answers (501 pilot; 4,214 validation). In the pilot, candidate data sources per question averaged between 14.6 (SD 13.9) for demographics and 89.2 (SD 56.1) for history and risk factors. In validation, human inter-rater agreement was approximately 98\% while 87\% of LLM answers exactly matched consensus, 2\% partially, and 9\% did not. Mean question-level accuracy was 91.5\% (SD 13.4\%) across 157 questions with at least 20 answers, and declined as ambiguity increased, from 96\% for Medication/Event Flag to 62\% for Event Timing questions. Conclusions: LLMs answering clinical registry questions on unprocessed EMR data achieved far lower accuracy than human abstractors. LLM accuracy fell steadily as ambiguity and the level of required clinical reasoning increased.

Introduction

Existing LLM benchmarks often use clean, narrowly defined questions and aggregate metrics that may not reflect performance on fragmented, unprocessed EMR data. Registry abstraction offers a practical setting to evaluate variation across clinically distinct question types.

  • Motivation: Curated clinical benchmarks differ substantially from real-world documentation tasks because they use prespecified answers from a single clinical context.They commonly summarize performance with one overall metric, potentially obscuring variation across clinically distinct tasks.
  • Prior evidence: Prior studies found below-50% exact-match coding accuracy at one academic center and 69.7% success in a structured virtual EHR environment.The coding study was single-institution, while the virtual EHR tasks remained structured and artificial.
  • Motivation: These limitations motivate evaluation on real, unprocessed EMR documents containing fragmented, redundant, and occasionally conflicting information.Clinical registries provide a practical framework for this evaluation.
  • Research gap: Existing registry-style LLM studies often report strong aggregate accuracy without explaining performance differences across question types.The cited applications include oncology, pulmonary embolism, and cardiovascular report classification.

Methods

The study used a two-phase, multi-site design to evaluate registry abstraction from unprocessed EMR documentation. Credentialed abstractors established consensus references, and questions were classified by ambiguity and clinical reasoning demand before validation testing.

  • Study design: The pilot used the EPDI registry at Institution A, while validation used CathPCI v5.7.1 at Institution B across 3 and 25 de-identified patient records, respectively.The pilot included 168 questions and 501 question-patient observations; validation included 232 questions and 4,214 observations.
  • Reference standard: Two credentialed cardiac registry abstractors independently answered each question, reconciled disagreements, and produced a consensus reference standard.Inputs included clinical, procedure, consultation, and nursing notes; imaging reports; summaries; and available structured data.
  • Question classification: 157 questions were classified into six categories ordered from lowest to highest reasoning demand: Medication/Event Flag, Binary Clinical Presence, Administrative, Quantitative Laboratory/Physiologic, Clinical Interpretation, and Event Timing.The analytic category counts were 62, 41, 18, 10, 22, and 4, respectively.
  • LLM protocol: The same model configuration used Claude Sonnet 4.6 for more complex questions and Claude Haiku 3.5 for simpler questions in a secure, HIPAA-compliant cloud environment.Prompts included verbatim registry definitions and instructed the model to answer when confident or decline when unsure.
  • Outcomes: Primary accuracy measured exact or partial matches among non-abstained answers, with secondary analyses of question-level and category-level accuracy.Category differences were tested with Kruskal-Wallis and pre-specified pairwise Mann-Whitney U tests.

Results

Validation accuracy was high on average but varied substantially by question and ambiguity category. Performance was strongest for simple binary or flag questions and weakest for clinically interpretive and timing questions.

  • Overall validation: 89.6% weighted aggregate accuracy was observed across 4,214 validation observations, while mean question-level accuracy was 84.0% across all 232 questions.For 157 questions with at least 20 evaluable observations, mean accuracy was 91.5% (SD 13.4%; median 96.0%; range 12.0–100.0%).
  • Overall validation: 99% of observations produced a model answer, with 87% exact matches, 2% partial matches, 9% mismatches, and 1% abstentions.The model answered 4,171 of 4,214 observations and declined on 43.
  • Ambiguity gradient: Clinical Interpretation and Event Timing comprised 17% of the analytic set but represented 67% of questions below 80% accuracy, an approximately 13-fold enrichment.Only 5% of questions above 90% accuracy came from these two categories.
  • Ambiguity gradient: Five of the six questions at or below 60% accuracy were interpretive or timing questions, including cath-lab indication at 12% and arrival date/time at 20%.The sixth was an administrative name question at 48% accuracy.
  • Abstention and precision: Declining to answer increased from 0.2% for Medication/Event Flag questions to 4.9% for Event Timing questions.On questions below 80% accuracy, recall remained above 90% while precision fell below 60%.

Discussion

The study interprets ambiguity as a central determinant of registry-abstraction performance and argues that aggregate metrics can conceal clinically important weaknesses. The authors also identify limited sample size and registry-specific reference standards as important scope boundaries.

  • Interpretation: Overall accuracy fell from 96.1% on simple binary questions to 62.0% on questions requiring the most clinical reasoning to resolve ambiguity.The two most complex categories were about thirteen times over-represented among the lowest-accuracy questions.
  • Why ambiguity matters: Clinical vignette benchmarks present clean questions, whereas EMR abstraction requires reconciling repeated or conflicting facts distributed across many document types.A history or risk-factor question drew on an average of 89 candidate sources per patient in the pilot.
  • Evaluation implications: The authors recommend reporting performance by category and question because weighted accuracy can overstate performance when easy questions recur more often.Weighted accuracy exceeded question-mean accuracy by 5.6 percentage points in this study.
  • Evaluation implications: Event Timing questions were especially difficult because correct answers usually required reconciling timestamps across multiple records rather than reading one document.The discussion characterizes the scattered evidence as a needle-in-a-haystack problem.
  • Limitations: The studies included only 3 pilot and 25 validation patients, producing wide per-question estimates and limiting generalization beyond these registries and credentialed-abstractor references.Per-category human inter-rater reliability was not recorded.

Conclusion

Across two ACC/NCDR registries and institutions, LLM accuracy decreased as clinical ambiguity and required reasoning increased. The study supports reporting performance by ambiguity category rather than only in aggregate.

  • Accuracy by ambiguity: 96% accuracy on the simplest category declined to 62% on the most complex category as question ambiguity increased.The comparison spans Medication/Event Flag through Event Timing questions.
  • Reporting implication: The study recommends interpreting clinical abstraction performance across categories with different ambiguity levels rather than relying on aggregate metrics.This recommendation follows the observed accuracy gradient across categories.
  • Overall performance: 91.5% was the overall mean question-level accuracy across 157 CathPCI questions with at least 20 observations.Figure 1 summarizes medians, interquartile ranges, whiskers, and individual-question accuracy by category.
  • Accuracy by ambiguity: 5% of questions in the ≥90% accuracy tier but 67% in the <80% tier were Clinical Interpretation or Event Timing questions.These categories showed an approximately 13-fold enrichment in the lowest accuracy tier.
  • Pilot implications: Answerability in the pilot showed no consistent relationship with candidate-source volume, suggesting that question type mattered more than source volume.The pilot compared section-level answerable rates with mean candidate sources per question.

Supplementary Appendix

The appendix provides the paper’s title and frames the work as a multi-site prospective study of ambiguity in LLM-based clinical registry abstraction.

  • Appendix: The paper is titled “An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction.”It is described as a multi-site prospective study.

Supplementary Methods

The supplementary methods define the registries, ambiguity categories, prompting approach, and software used across the pilot and validation phases.

  • Registry background: The pilot used the ACC NCDR EPDI registry, while validation used CathPCI v5.7.1 across more than 1,600 US hospitals.The two phases used different ACC/NCDR registries.
  • Ambiguity taxonomy: Validation questions were assigned to six categories ordered from lowest to highest reasoning demand.The categories ranged from Medication/Event Flag and Binary Clinical Presence to Clinical Interpretation and Event Timing.
  • Prompt structure: Each prompt included a clinical-abstractor instruction, the verbatim ACC NCDR question definition, and a structured output format.The model was instructed to answer when confident and decline when not.
  • Software: Analyses used Python 3.10 with NumPy 1.24, Pandas 2.0, and SciPy 1.10.

Supplementary Material

The supplementary material documents pilot variability, retrieval contexts, record and document types, and ranked question-level validation performance. It includes detailed examples spanning accuracy from 76% to 100%.

  • Pilot detail: The pilot supplementary table disaggregates candidate-source counts and answerable rates across three patients, illustrating crosspatient variation in document footprints.The same question could have different source-data profiles across patients.
  • Document contexts: The document-context specifications list the EMR resources and document types used to route validation questions.Examples include patient resources, encounters, notes, discharge summaries, laboratory results, medications, and imaging reports.
  • Question and source coverage: The supplementary question inventory covers demographics, history and risk factors, laboratory results, medications, imaging, procedures, and timing fields.Examples include names and dates, comorbidities, laboratory values, medication administration, cardiac imaging, and procedure times.
Loading 2608.20373v1…