Source-linked AI summary

Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support

Jonathan A. Handler, Marlene I. Robles-Granda, Jacob E. Mefford, Jeremy S. McGarvey, Gregory S. Podolej, Colleen J. Klein, Matthew D. Dalstrom, William F. Bond

arXiv:2609.10421v1cs.CYcs.AI

TL;DR

ED revisit review is constrained by chart-review burden and commonly narrow time windows, motivating methods to identify potentially concerning pairs more efficiently. This exploratory study examined clinician and GPT-4 assessments of diagnosis pairs and derived a knowledge-graph algorithm for automated screening. The KGA achieved 83-100% positive predictive value in preliminary assessment, while further validation remains necessary.

  • Problem

    ED revisit quality review is limited by chart-review burden and commonly used 48-72-hour windows, potentially missing clinical failures outside those windows.

  • Method

    The study assessed diagnosis pairs from 99 randomly selected revisiting ED visits using clinician and GPT-4 ratings, then derived and preliminarily assessed an LLM-populated knowledge-graph screening algorithm.

  • Results

    The KGA achieved 83-100% positive predictive value for identifying pairs that at least one clinician judged to warrant further investigation.

  • Takeaways & Limitations

    The preliminary findings may support expanding revisit screening beyond traditional 72-hour windows without substantially increasing reviewer workload, pending further validation.

  • Takeaways & Limitations

    The study had a relatively small sample from one multihospital health system, limited rater agreement, and requires follow-up validation.

Abstract

from arXiv · show

Background: Emergency Department (ED) return visits are commonly reviewed for quality assurance, but are often limited (e.g., to revisits within 48-72 hours) to increase actionable finding yield while minimizing chart review burden. Those limitations may lead to missed quality improvement opportunities. Methods: We conducted an exploratory, retrospective study of randomly selected ED visits to a multihospital health system having an ED revisit within 1-14 days to the same health system. Given only each visit's primary diagnosis, raters (2-3 clinicians and GPT-4 large language model [LLM]) assessed characteristics of the diagnosis pairs, including the "target": whether a pair warranted further assessment. Informed by rater response analyses, an algorithm leveraging an LLM-populated knowledge graph ("KGA") was created to automatically screen for potentially concerning pairs, then preliminarily assessed. Results: 99 diagnosis pairs were included. GPT-4 responses poorly correlated to clinician raters, rating nearly all (94%) pairs as warranting follow-up (4.4-13.3 times more than clinicians). However, prompt engineering was minimal. Among clinician raters, revisit medical gravity was consistently significantly associated with the target, while a differential diagnosis/complication composite was significantly associated on unadjusted, but not adjusted (though less powered) analysis. The KGA achieved 83-100% positive predictive value for at least one clinician rater determining further assessment was warranted based on the diagnosis pair. Conclusion: These results can inform next steps for improving screening with LLMs like ChatGPT. Further research is warranted to validate this preliminary work's finding that the KGA may enable enhancing the scope and yield of screening without substantially increasing reviewer workload.

INTRODUCTION

ED revisit quality reviews are constrained by chart-review burden and low actionable yield, especially when restricted to short revisit windows. The study therefore explores broader, high-precision selection of records for manual review using clinician and LLM-informed approaches.

  • ED return-visit reviews support quality improvement but are limited by chart-review burden and low actionable yield.
  • 48-72-hour revisit windows are commonly used even though clinical failures can occur outside them.
  • The study sought factors associated with manual-review need and explored LLM-based screening plus a knowledge-graph algorithm.

Reviews

Reviewers evaluated diagnosis pairs using structured measures of diagnostic status, medical gravity, differential inclusion, complication inclusion, and need for further investigation. GPT-4 received separate prompts for these assessments, while the KGA used an LLM-populated knowledge graph.

  • Two emergency physicians and one urgent-care advanced practice provider independently assessed diagnosis pairs without encounter-specific information.
  • Reviewers scored diagnosis membership and medical gravity, including mild, typical, and severe forms of each diagnosis.
  • Reviewers rated whether the revisit diagnosis belonged in the index differential, represented a complication, and warranted further investigation.
  • GPT-4 was queried with separate prompts, with minor iterative prompt engineering producing wording differences from human assessments.
  • The KGA relied on an LLM-populated knowledge graph of potentially concerning diagnosis pairs.

Statistical methods

The analyses combined descriptive summaries with mixed-effects logistic regression to examine associations with the binary target while accounting for rater and encounter-pair variation.

  • Descriptive, target-response, and KGA-performance analyses were performed in Excel.
  • Mixed-effects logistic regression examined relationships with the binary target using rater and study-pair random intercepts.

RESULTS

After filtering, 99 encounter pairs were analyzed, with substantial variation between human and GPT-4 target ratings. In multivariable analysis, typical medical-gravity change and differential inclusion were significant predictors.

  • 99 encounter pairs remained after filtering correction.
  • GPT-4 marked 93.9% of pairs for further assessment, compared with 7.1%-21.2% among human raters.
  • The multivariate model included diagnosis membership, differential inclusion, typical medical-gravity change, and days between visits.
  • Only typical medical-gravity change and differential inclusion were statistically significant.

Stage 2 Analysis (Two Human Raters)

Stage 2 analysis examined whether diagnosis relationships influenced judgments about further review. The relationship composite was significant unadjusted but not adjusted, with exclusions reducing the analyzed sample.

  • Model specification: The revisit medical-gravity value was selected for additional analyses because it appeared more relevant than the change in medical gravity.Preliminary analyses found relatively similar odds ratios and significant p-values for revisit and delta medical gravity.
  • Relationship composite: The relationship composite used the larger of differential-inclusion or complication-inclusion values to represent potential misdiagnosis or complications.Pairs with both component values unavailable were excluded.
  • Association with the target: The relationship composite was significantly associated with the target in unadjusted analysis but not adjusted analysis.The adjusted analysis had fewer observations because of exclusions.
  • Association with the target: Reduced sample size may have produced a false conclusion of insignificance for the adjusted relationship-composite analysis.The unadjusted association also lost significance when restricted to the smaller adjusted-analysis dataset.

Stage 3 Analysis (Algorithm Performance)

The KGA used an LLM-generated knowledge graph and medical-gravity thresholds to flag diagnosis pairs for manual follow-up. In this preliminary evaluation, flagged pairs had high positive predictive value, including revisits beyond the traditional 48–72-hour window, but the study requires validation.

  • Algorithm design: The KGA selected diagnosis pairs potentially indicating misdiagnosis, delayed diagnosis, or an index-diagnosis complication, then filtered them by revisit medical-gravity index cutoffs.Pairs meeting the algorithm criteria were labeled positive for manual follow-up.
  • Performance: 28/99 visit pairs (28.3%) were flagged by at least one rater as warranting further investigation.These pairs were designated actual positives for evaluating algorithm performance.
  • Performance: At the lower MGI cutoff, 5/6 flagged pairs were true positives, yielding 83% PPV; at the higher cutoff, 4/4 were true positives, yielding 100% PPV.The higher cutoff produced perfect PPV in this preliminary sample.
  • Performance: At the higher cutoff, 4/4 true positives were revisits at 6–7 days that 48–72-hour screening would miss.At the lower cutoff, 4/5 true positives were similarly beyond the traditional window.
  • Limitations: The study had a relatively small sample from one multihospital health system, limited rater agreement, and a need for follow-up validation.Future work should also examine a broader set of inputs, such as revisit disposition.

DISCUSSION

The discussion places the study within efforts to automate costly ED quality review and interprets clinician associations, GPT-4 limitations, and the KGA’s preliminary high-precision performance. The authors conclude that further validation is needed before extending screening beyond traditional revisit windows.

  • Context: Quality-assurance chart review is costly, motivating computer-based screening approaches for adverse events and missed diagnostic opportunities.This study extends related work using diagnosis relationships and LLM-supported screening.
  • Clinician decision-making: Revisit medical gravity was most strongly and consistently associated with the target variable among clinician ratings.The relationship-composite association was significant unadjusted but not adjusted, with reduced power in the latter analysis.
  • Clinician decision-making: The reported associations may be stronger than their small odds ratios suggest because each odds ratio represents a one-point change on an approximately 100-point scale.The scale interpretation affects how the magnitude of associations should be read.
  • LLM performance: Most GPT-4 target responses were false positives, although the study used only light prompt engineering and separate prompts for visit pairs.The authors suggest that different models or prompt engineering may improve performance.
  • KGA implications: The KGA showed promise as a high-precision mechanism for expanding revisit analysis, with high PPV considered more relevant than sensitivity for this use.Its preliminary role is to identify potentially concerning pairs beyond restrictive revisit windows without substantially increasing reviewer workload.
  • KGA implications: The authors state that further research should validate whether the KGA can broaden screening scope and yield without substantially increasing reviewer workload.The conclusion is explicitly preliminary.

Data Availability Statement

Some study data are unavailable publicly because they contain protected health information and are governed by ethical, privacy, and institutional restrictions.

  • Data access: Some study data cannot be publicly shared because they are protected health information subject to ethical and privacy policies and regulations.Institutional review board and institutional approval or agreement limitations also apply.

Declaration of Interests

The declaration reports extensive financial, advisory, investment, and patent-related interests for Jonathan Handler, while the remaining authors declare no relevant interests.

  • Jonathan Handler reports leadership and shareholder interests in Keylog Solutions LLC and other healthcare and artificial-intelligence companies.
  • He also reports Pfizer funding, advisory roles, patents, and institutional-related stipends, food, or travel.
  • The remaining authors declare no relevant interests to disclose.

Tables

The paper presents three regression-analysis tables and a figure overview for the knowledge-graph algorithm, which was populated in part through repeated LLM queries about diagnosis relationships.

  • Table 1 reports the adjusted logistic regression analysis for whether revisit pairs warranted follow-up.
  • Table 2 reports the unadjusted analysis of the relationship composite and whether follow-up was warranted.
  • Table 3 reports adjusted and unadjusted logistic-regression analyses for whether follow-up was warranted.
  • The knowledge-graph algorithm overview is supported by a graph populated partly through repeated LLM queries about relevant diagnosis relationships, such as complications.
Loading 2609.10421v1…