Source-linked AI summary
One note in three: a verified census of three deployed AI scribes, and the instrument that counted it
Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
TL;DR
Ambient AI scribes can introduce clinically consequential failures into signed notes, while adopters lack a standard way to measure them. This census audits three deployed products with adversarial verification and finds failures in one note in three, with reported rates strongly shaped by the measurement instrument.
Problem
Telephone notes can falsely document examinations, and adopters need standardized evidence on scribe errors for ongoing clinical documentation audits.
Method
The study ran 565 notes from 142 consultations through three commercial scribes, used broad model-based discovery, and adversarially verified important candidates with two model families.
Results
31.3% [27.0, 35.6] of notes carried a verified failure, while the instrument’s review instruction shifted candidate verification from 9.3% to 79.0%.
Takeaways & Limitations
Failure rates are joint properties of scribes and measurement instruments, so the released findings, prompts, model versions, and pipeline support reproducible evaluation.
Takeaways & Limitations
Discovery sensitivity for errors that no model-based pass proposed was not estimated, and adjudication covered only 33 of 618 findings with wide intervals.
Abstract
from arXiv · showhide
Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two models from different families, each told to refute what it could, and 618 survived. One note in three (31.3% [27.0, 35.6]) carries a verified failure, concentrated in allergy and medication information, invented patient identity, and history written up as examination on telephone consultations that can contain none. No product was given a patient record; setting aside the two classes a record would have prefilled, invented identity and dates, the rate is 24.8% [20.8, 29.0]. One failure mode did not fit our scheme, drawn from published scribe-error taxonomies: a treatment the clinician retracts, recorded as delivered care. Two clinicians adjudicated blind, disjoint samples: a physician author upheld 20 of 21 findings (95.2% [77.3, 99.2]) and an independent clinician, not an author, 12 of 12 ([75.8, 100]); both judged every sampled refusal genuine. A failure rate depends on the instrument as much as the scribes. With model, evidence and settings fixed, the review instruction alone moves the share of candidates verified from 9.3% to 79.0%, and the reviewing family moves it too: alone at that instruction the gentler flags 54.8% of notes against 27.8%. Between 28% and 97% of sampled notes carry a failure depending on the standard. Published audits disagree among themselves by a margin instrument differences alone can produce: omission is 54-86% of their errors against our 23.1%. We release all 618 findings with transcript-side evidence, every prompt and model version, and the re-runnable pipeline.
1 Introduction
The paper audits three deployed ambient scribes on identical consultations and shows that both verified failure rates and their interpretation depend on the counting instrument. It releases the evidence-quoted census, taxonomy, and re-runnable measurement pipeline.
- 565 notes from 142 consultations yielded 618 findings that survived adversarial verification after 13,678 candidate errors were proposed.The corpus combines recorded UK primary-care consultations, US ambulatory encounters, and authored scenarios.
- The counting instrument is a joint property of the scribes and review process, with the review standard accounting for almost the measured change in verification rate.The study separates the effects of review instruction, model family, and panel architecture.
- 9.3% of sampled candidates were verified under the strict instruction versus 79.0% under the lenient instruction, with the remaining machinery moving the count by one percentage point.The comparison keeps the model, evidence, and settings unchanged while replacing only the verification instruction.
- The paper contributes a verified cross-product census, a controlled decomposition of the counting instrument, and complete release of the instrument for re-running.The released materials include evidence and the components needed to reproduce the count.
2 Related work
Prior scribe audits establish recurring error classes but use heterogeneous human or vendor-defined standards, limiting direct comparison. This paper positions its contribution as a census released together with the instrument that produced its count.
- Published human-review audits report omission as 54-86% of errors, but none directly measures how its review standard contributes to that spread.Reported omission shares include 83%, 54%, and 76.3% [70.0, 83.3] across cited studies.
- Shared-corpus studies compare note-quality scores or clinician experience rather than producing a verified account of note failures.The cited designs include standardised cases with blinded raters and a large randomised deployment measured through surveys.
- Vendor evaluations also show detector-dependent error burdens, including 24.4% under a calibrated automated reviewer versus 6.2% under unaided clinician adjudication.These evaluations compare different detection sensitivities and do not provide the same measurement basis.
- Review standards are recognised as a moving part: validation methods are highly heterogeneous, and a scoping review says this limits cross-system comparison.Prior work includes consultation checklists and examples of evaluators conflating distinct error concepts.
- The paper releases an adversarially verified, evidence-quoted census on identical consultations together with a controlled decomposition of the counting instrument.The instrument weighs the review instruction, model family, and panel architecture and is intended to be re-run independently.
3 What three deployed scribes get wrong
The census audited three deployed scribes on a shared corpus using broad discovery, adversarial verification, and blinded clinician checks. Verified failures affected 31.3% of notes, concentrated in medication and allergy information, invented identity, and other clinically consequential documentation errors, while rates and product differences depended on the counting instrument.
- Corpus: 565 notes from 142 consultations were generated by three commercial scribes across recorded, benchmark, and authored scenarios.The corpus included 57 UK primary-care consultations, 45 US ambulatory encounters, and 40 authored scenarios.
- Discovery and verification: 13,678 candidate errors were proposed across twelve discovery passes, 5,898 cleared importance filtering, and 618 survived adversarial verification.Two skeptics from different model families reviewed the filtered candidates, with disagreements resolved by a tiebreak model.
- Human checks: 95.2% of the physician author’s sampled findings and 12 of 12 findings in the independent clinician’s sample were upheld.The physician author upheld 20 of 21 findings, while the independent clinician verified all 12 sampled findings.
- Overall failure rate: 31.3% [27.0, 35.6] of notes carried at least one verified failure, but the rate is a floor because missed candidates and undiscovered errors were excluded.The census counted verified findings rather than distinct underlying errors, and its two-stage instrument can only reduce the observed count.
- Comparability: Product-level rates cannot be interpreted as rankings because capture paths, note denominators, and review standards differed.Scribe A produced two notes per consultation, and lenient review narrowed the product gap while reversing the order of Scribes B and C.
- Failure patterns: The largest failure clusters involved allergy and medication state, invented patient identity, and errors such as omitted plans, false denials, corrupted vitals, and retracted treatments recorded as delivered.The allergy and medication cluster contained 111 findings, while invented identity contained 93; examples included an understated dose and a fabricated patient name.
4 What a failure rate means: the instrument, and the published audits
The paper shows that failure rates are jointly determined by scribes and the instrument used to count failures. Holding evidence and settings fixed, review standards produce large differences, while published omission rates are not directly comparable and may largely reflect instrument differences.
- Published audits: 23.1% versus 54–86%: published omission shares differ from this census, and instrument differences alone can produce much of that spread.Biro’s two products differed by 29 points under one review standard, while other studies used different instruments and sometimes different denominators.
- Review instrument: 79.0% versus 9.3%: changing only the review instruction moved candidate verification by 69.7 percentage points.The same model reviewed the same evidence; 902 candidates changed from refuted to verified, and none changed in the opposite direction.
- Review instrument: 7.6 times: the strictly nested lenient standard retained all 134 panel findings and added 889 on identical material.The resulting range is a measured range rather than a bound, because neither standard is established as the true rate.
- Review instrument: 1 percentage point: adding the second model family and tiebreak changed the strict rate from 9.3% to 10.3%.The paired difference was not distinguishable from no change, whereas the two families agreed on 86.2% of candidates and recorded dissent on the rest.
- Review instrument: 27.8% versus 54.8% of notes: at the strict instruction, the gentler reviewing family flagged roughly twice as many notes as the harsher skeptic.Under the lenient instruction, 96.5% of sampled notes were flagged, illustrating the broader standard-dependent range.
- Failure categories: 6.1% versus 24.6%: the panel verified omission candidates at a quarter of the fabrication rate, despite omission being the largest targeted hunt.The omission result reflects both the panel’s importance judgements and its treatment of restatements elsewhere in the note.
5 What this means in practice
The census turns failure rates into a buyer and clinician question about the instrument, the failure clusters it measures, and the checking burden it leaves behind.
- For the clinician who signs: Allergy and medication state, examination provenance, identity, and spoken diagnosis concentrate the high-importance failures.Plans written as done add 46 findings, 44 graded critical, from all three products.
- For the buyer: 28% to 97% of sampled notes are flagged when the review standard changes and nothing else does.At candidate level, the answer moves by a factor of eight and a half.
- For the buyer: A rate without its instrument cannot be compared with anyone else’s because thresholds can hide inside headline metrics.The paper releases every prompt, model version and verdict so the count can be interpreted and rerun.
- For the buyer: The released pipeline costs about $530 in model calls for 565 notes, while real-time audio capture remains the slow part.A run on real consultations sends complete transcripts and notes to two model providers through repeated discovery and review stages.
- The quality layer: 24.6% of notes containing an omission are caught by enumerate-then-check, versus 36.9% for the evolved prompt, with clean-note false-flag rates of 2.7% and 6.2%.Both methods operate on a benchmark rather than directly on scribe output; the evolved prompt costs roughly a tenth as much per note before amortisation.
6 Limitations
The study’s comparisons are controlled but bounded by simulated or authored encounters, missing clinical context, limited clinical scope, and incomplete human validation.
- Scope: The corpus is recorded, simulated or authored rather than live traffic, trading ecological validity for identical consultations and controlled ground truth.Real consultations may be longer, messier and more specialty-diverse than this corpus.
- Scope: No clinician-written comparator is counted on the same consultations, so the study cannot say whether scribe notes carry more failures than replacement notes.A useful comparator would need to be evaluated with the same instrument.
- Measurement: Capture paths and note templates differ by product, creating a confound and leaving per-note rates unnormalised for note length.Scribe A contributes two API-generated notes per consultation, while Scribes B and C hear replayed audio.
- Measurement: No product received demographics, a patient record, or an encounter date beyond its own application, limiting interpretation of identity and date findings.Removing identity and dates changes the headline from 31.3% [27.0, 35.6] to 24.8% [20.8, 29.0], but neither is an integrated-deployment rate.
- Scope: The census covers English UK primary-care and US ambulatory encounters, not specialty dictation, inpatient documentation, or other languages.Drug and code terminology plus biased or stigmatising language were deliberately not hunted and are not measured or zero.
- Measurement: The instrument’s strictness makes reported rates floors for what it missed, while its run-to-run variation was not measured.A lenient instruction verifies 79.0% of sampled candidates versus 10.3% under the panel, and one stochastic run cannot establish product differences between repeats.
- Validation: Human checks formally adjudicated only sampled findings, refusals and severity grades, with most adjudication performed by a physician author.The published precision is 95.2% [77.3, 99.2] on 21 sampled findings and 12 of 12 for an independent clinician’s disjoint sample.
7 Conclusion and release
The conclusion is that verified failures and their measurement instrument should be released together, enabling reproducible reassessment while preserving the study’s snapshot and scope boundaries.
- Conclusion: One note in three from three deployed scribes carries a failure surviving a refutation panel, concentrated in clinically consequential note locations.The paper treats the count as a joint property of scribes and instrument.
- Release: All 618 verified findings, transcript-side evidence, verdicts, ratings, prompts, model versions, manifests, clustering materials, and pipeline code are released.The release is tied to the same repository and DOI as the companion benchmark.
- Reproducibility: The census is a snapshot because products update continuously, but the published instrument can be rerun on later generations or local consultations.Per-product figures cannot be independently re-derived because products and notes are anonymised and unreleased.
- Next steps: A review-catch study, external calibration, terminology-grounded searches, and longitudinal reruns are the next measurements the census directly motivates.These studies would measure clinician detection, strictness gaps, currently unhunted classes, and trends over product generations.
- Study context: The corpus contains no real patient data, and the authors report no external funding.Consultations involve actors, simulated encounters or authored scenarios; no patients or clinical records were used.
- Study context: The authors build evaluation tooling commercially, but no author product was used in measurement and the released instrument is free.Following the paper’s evaluation advice does not require an author product.
A The census method in full
The census method combines a full taxonomy and controlled verification instrument, with product-specific note-generation differences documented for interpreting counts.
- Taxonomy and instrument: The appendix expands the taxonomy to all 17 discovered subcategories, including sizes, composition, clustering parameters, and stability.It also details the verification panel and controlled re-review under four review standards.
- Product inputs: Scribe A contributes two API-generated notes per consultation, whereas Scribes B and C contribute one from replayed audio.Its appendix columns are finding counts, while per-note rates accounting for template differences appear in Table 1.
Part 1 - The taxonomy
This section organizes the 618 verified findings by cluster and explains how those clusters were formed.
- The section presents the 618 verified findings one cluster at a time.
- The clusters are the organizing unit for describing the findings.
- The section focuses on both the findings and their construction.
A.1 All 17 clusters
The census groups 563 verified findings into 17 density-based clusters and leaves 55 unassigned, while the clusters span recurring documentation failures including omissions, invented facts, provenance errors, and retracted care recorded as delivered. Cluster counts must be read alongside consultation coverage because repeated findings within consultations can inflate totals.
- Cluster construction: 563 findings occupy 17 density-based clusters, while 55 remain unassigned.The unassigned findings are retained separately rather than folded into neighboring clusters.
- Cluster construction: 16 clusters fit the scheme’s tier-2 categories; the retracted-device group is reported separately.The retracted-device group could not be placed under any scheme category.
- Failure modes: The main clusters cover allergy and medication omissions, invented identity, dropped diagnoses or dates, examination provenance, and fabricated or distorted clinical information.
- Failure modes: Other clusters include phantom counselling, family-history misassignment, dropped work-absence plans, false symptom denials, alternative options recorded as orders, corrupted vitals, and downgraded safety-netting.
- Severity: 476 findings are critical, 140 supporting, and 1 peripheral among the 617 graded findings.One member of cluster 6 lacks a rubric grade.
- Interpretation: The severity distribution reflects two selection stages that remove lower-importance material before grading, not the severity mix of scribe errors generally.High-salience findings survive at 15.7%, versus 7.2% for low-salience findings.
- Interpretation: Cluster counts can overstate independent evidence when repeated findings occur within one consultation across products and templates.Four clusters rest on one or two consultations and are therefore presented as failure modes rather than rates.
A.1.1 One example from each cluster
Examples illustrate the 17 clusters through transcript-grounded discrepancies, including omissions, invented identity or dates, altered values, misassigned history, and retracted care recorded as delivered.
- Allergy and medication omissions can remove answered allergy status from a note when a drug is prescribed.
- Invented identity supplies a specific patient name even when the recording makes the name inaudible.
- A stated working diagnosis can disappear from the note despite being told to the patient.
- Examples include remote history written as examination, distorted diabetes values, invented calendar years, phantom counselling, and family-history misassignment.
A.2 The clustering recipe, in full
The clustering recipe discovers only the scheme’s third tier from verified-finding descriptions, using a fixed embedding and projection pipeline whose parameters are recorded for reproducibility.
- The recipe records the scheme definition version and hash so every parameter can be reproduced or contested.
- The clustering input is each verified finding’s description with a fixed failure-mode prefix.
- Cohere embed-v4.0 generates cached embeddings before UMAP projection to 15 dimensions with cosine distance and 15 neighbors.
- A 60% member threshold produced no member-vote placement for the retracted-device cluster.All 16 placed clusters were assigned by the labeling call rather than the member vote.
A.2.1 Seed stability, and what the parameter grid looked like
Cluster counts vary with projection seeds and parameter settings, so the reported grouping is only approximately stable. The parameter sweep did not identify a configuration meeting the prespecified target band.
- Seed stability: 17, 20 and 21 clusters arise from UMAP seeds 42, 43 and 44, with mean adjusted Rand index 0.727.The count is stable only within a few clusters, and boundaries are largely but not exactly reproducible.
- Parameter grid: 17 to 54 clusters and 33 to 119 unassigned findings occur across the 24-configuration parameter sweep.The sweep varied n_neighbors, min_cluster_size and min_samples.
- Parameter grid: No swept configuration met the target of 6 to 16 clusters with no cluster exceeding 45% of the pool.The criterion was defined before cluster contents were read.
Part 2 - The instrument
The instrument section examines what the verification panel does to the count, including its standards, surviving claim types, history, refusal audit, and clinician checks.
- The instrument: The verification panel is evaluated through its behavior, review standards, surviving claim types, refusal audit, and blinded clinician sittings.These components frame how the instrument produces and checks the reported count.
A.3.1 The panel, and how it behaves
An adversarial two-model panel reviews importance-filtered candidate findings using full notes and transcripts, with tiebreaks for disagreements. Its reviewers differ in strictness, and panel importance ratings predict survival.
- Panel procedure: 5,898 importance-filtered candidates were shown to two skeptics from different model families, each instructed to refute findings whenever defensible.Both keeping means verified, both refuting means cut, and a split goes to a final tiebreak call.
- Panel behavior: 90.6% and 81.3% of candidates were refuted by the harsher and gentler skeptics, respectively.The asymmetry motivates allowing a split to proceed to a third opinion rather than letting one dissent decide alone.
- Panel behavior: 4,663 findings were unanimously refuted, 812 were split to the tiebreak, and 423 were unanimously retained.These are the panel’s decision categories over the 5,898 candidates.
- Importance re-rating: 15.7% of panel-high, 8.1% of panel-medium, and 7.2% of panel-low candidates survived verification.The panel’s importance re-rating changed the discovery distribution to 2,014 high, 2,377 medium, and 1,507 low candidates.
- Reproducibility: Discovery prompts were preserved verbatim, while the panel prompt reordered the same sentences and recorded versioned definitions and hashes.Released manifests include exact model pins, routing, and settings.
A.3.2 The same candidates under four review standards
The study compares four review standards on the same candidate pool while holding evidence and, for the controlled contrast, model settings fixed. Verification rates vary sharply with the instruction and reviewing family, and the rates are bounded by the candidate pool and review design.
- Design: 1,295 candidates were sampled from 5,898 across 115 whole notes, preserving the note-level denominator for estimating notes with verified findings.The draw was stratified by product and source, with one note carrying no candidate retained in the denominator.
- Standards: The four standards comprise harsher-family solo review, gentler-family solo review, the full panel, and a newly run lenient review on the harsher model.The lenient run changed only the review instruction and closing question while keeping evidence sections identical.
- Candidate-level results: 9.34% of candidates were verified by the harsher skeptic alone under the strict instruction, versus 79.00% under the same model’s lenient instruction.The controlled contrast changes only the review standard.
- Controlled contrast: The lenient-versus-strict contrast produced a +69.7 percentage-point difference and an exact McNemar p = 5.9e-272.Against the full panel, lenient review was 7.63x; panel versus the harsher skeptic was +1.0 percentage point with p = 0.13.
- Note-level results: 54.8% of sampled notes carried a verified finding under the gentler strict review, compared with 27.8% under the harsher strict review.The full panel also yielded 27.8%, while the same model under lenient review yielded 96.5%.
- Failure mix: Omission comprised 22.3% to 33.9% of strict reviewers’ verified findings and 28.2% under lenient review, while misplaced text rose to 18.0% under leniency.The same candidate pool produced different failure-class mixes across standards.
- Scope: All four rates are shares of the discovery-assembled candidate pool, and no external ground truth adjudicates between standards.The physician review calibrates strict-panel refusals only; the lenient review uses one call per candidate from one model family.
A.3.3 What survives is what can be checked
What survives verification depends on how objectively checkable the claim is and on the review standard applied. Human checks support the published findings, while refusal handling and severity grading remain standard- and rater-dependent.
- Survival tracks how objectively checkable a claim is, not whether it is an omission.Omission is not the lowest-surviving category, contrary to a straightforward reading of the table.
- 37.3% to 6.25% survival fell when the corpus and model generation changed, while a later change in architecture, discovery instrument, and corpus raised survival to 21.2%.The comparison does not isolate those changes individually.
- 25.0% of sampled refusals were rejected because the fact appeared elsewhere in the note, a convention closer to misplaced text than omission.This share remained essentially unvalidated because the sitting answered one present-elsewhere item and abstained on two.