Source-linked AI summary
Toward Workflow-Aware Benchmarking for Healthcare NLP Agents
Junyi Yao, Baichuan Li, Zihao Zheng, Jiayu Long
TL;DR
Healthcare-agent evaluations often underrepresent longitudinal state, interruptions, and handoffs because they focus on static questions or one-shot generation. The paper introduces a reproducible episode-level protocol spanning model, agent, and simulated-workflow behavior, with structured scoring for continuity, traceability, and escalation. It positions this protocol as an intermediate evaluation layer, not evidence of clinical outcomes or deployment value.
Problem
Healthcare-agent evaluation often relies on static question answering or one-shot generation that underrepresents longitudinal state, interruptions, and human handoffs.
Method
The paper specifies an episode-level protocol that separates model, agent, and simulated-workflow evidence using a reproducible five-field episode schema and scoring for workflow behaviors.
Results
The protocol provides four task templates and reproducible scoring for state continuity, evidence traceability, escalation, and handoff behavior across defined episodes.
Takeaways & Limitations
The protocol supplies an intermediate evaluation layer between static benchmarks and prospective workflow studies, while keeping deployment claims within the defined episodes.
Takeaways & Limitations
The protocol does not reproduce institutional implementation or patient outcomes and is intended as a testable measurement protocol rather than evidence of improved healthcare work.
Abstract
from arXiv · showhide
Large language model (LLM) agents are increasingly proposed for healthcare tasks such as clinical documentation, evidence retrieval, patient messaging, and care coordination. Yet many evaluations remain limited to static medical question answering or one-shot generation, under-representing longitudinal state, interruptions, and human handoffs. We introduce an episode-level evaluation protocol for healthcare NLP agents. The protocol separates evidence across model, agent, and simulated-workflow behavior; specifies a five-field episode schema; and defines annotation and scoring for state continuity, evidence traceability, and escalation decisions. It is instantiated as four task templates: documentation update, evidence retrieval, patient messaging, and triage handoff. The protocol does not claim to measure clinical outcomes or deployment value. Instead, it supplies a reproducible intermediate evaluation layer between static benchmarks and prospective workflow studies, with an explicit cost-sensitive treatment of missed versus unnecessary escalation.
I. INTRODUCTION
Healthcare-agent evaluation must extend beyond static prompt-response tests because clinical workflows are stateful, multi-step, and high stakes. The paper proposes a reproducible operational layer between static benchmarks and prospective workflow studies.
- Clinical workflows require agents to summarize records, retrieve evidence, maintain context, defer when needed, and support human handoffs.
- One-turn performance is insufficient evidence of practical usefulness in healthcare settings.
- Healthcare-agent failures often involve lost state, wrong evidence, un surfaced uncertainty, or exceeded workflow scope rather than isolated textual errors.
- The proposed benchmark layer sits between static prompt-response tests and prospective clinical workflow studies.
- The protocol distinguishes evidence at model, agent, and simulated-workflow levels and provides reproducible episodes, task templates, and cost-sensitive escalation scoring.
II. WHY HEALTHCARE AGENT EVALUATION IS DIFFERENT
Healthcare work unfolds over time, so useful agent evaluation must assess state, process, evidence use, interruptions, and escalation rather than fluent answers alone. Existing benchmarks cover important capabilities but lack a common logic for interpreting their claims.
- Healthcare tasks are sequential, can change goals mid-task, and may require clarification after new information arrives.
- Agents can produce polished language while omitting contraindications, retaining stale information, or retrieving irrelevant evidence.
- Workflow evaluation should track context preservation, evidence use, interruption recovery, and escalation at scope boundaries.
- Static QA and hallucination benchmarks test recall, explanation, coverage, and unsupported claims but rarely test deferral, missing-context queries, or cross-turn state.
- Emerging EHR-like, long-horizon, consultation-centered, and trajectory-based benchmarks improve realism, but the field lacks conceptual integration across evaluation layers.
IV. A THREE-LEVEL EVIDENCE PROTOCOL
The protocol organizes healthcare NLP agent assessment into three evidence levels and requires results to be reported at the level they actually support. This prevents strong episode scores from being promoted into deployment claims.
- Every evaluation result is assigned to model, agent, or simulated-workflow evidence rather than treated as a single undifferentiated benchmark outcome.
- The reporting rule prohibits promoting results to stronger deployment claims than their evidence level supports.
- Each evidence level is tied to a deployment question, a metric family, and the blind spot created by using that level alone.
A. Model Level
The model level evaluates foundational behavior such as factuality, calibration, fairness, robustness, and harmful errors. Static QA, explanation, and hallucination benchmarks are appropriate at this level, while agent execution requires separate assessment.
- The model level measures factuality, calibration, fairness, robustness, and harmful error patterns.
- Static QA, explanation, and hallucination benchmarks test medical knowledge, unsupported claims, and degradation under distributional or linguistic variation.
- Agent-level constructs such as planning, tool use, retrieval, memory, updating, and uncertainty handling are distinct from core model behavior.
C. Simulated Workflow Level
The simulated workflow level evaluates task-sensitive routing, continuity, traceability, and handoff behavior rather than treating healthcare agents as isolated response generators. It reports a performance vector whose claims remain limited to the defined episodes.
- Simulated workflows evaluate routing, interruption recovery, handoff completeness, and output reviewability, not clinical outcomes or deployment value.Prospective workflow studies with local governance and human participants are required for claims about time saved, workload, patient safety, or care outcomes.
- The protocol reports model competence, agent execution, and simulated-workflow performance as a vector rather than a single leaderboard score.Strong performance at the simulated-workflow level supports claims only about the defined episodes.
- Task classes require distinct evaluation emphases because documentation, retrieval, messaging, and triage have different primary risks.A universal leaderboard can obscure changing balances among correctness, grounding, escalation, and continuity.
- Triage shifts the dominant risk from factuality toward escalation and continuity, while evidence retrieval alone does not establish safety for patient-facing communication.Patient-facing communication also requires boundary testing and harmful-advice mitigation.
VI. EPISODE SPECIFICATION AND SCORING
The protocol defines auditable workflow episodes with structured state changes, constrained actions, reference decisions, and atomic scoring. Escalation is treated as a cost-sensitive decision whose trade-offs must be declared before evaluation.
- Each episode contains five required fields: initial context, state-changing event, allowed action space, reference decision profile, and adjudication rubric.The schema specifies whether agents may answer, retrieve, revise, ask, abstain, or escalate, making scoring criteria auditable.
- Two annotators label required facts, prohibited claims, admissible evidence, actions, and handoff fields, while a third reviewer adjudicates disagreements.Releases should retain final labels and disagreement types without assuming a uniquely correct reference response.
- State continuity, evidence traceability, and handoff completeness are scored from atomic annotations rather than global impressions.The measures respectively address retained or revised facts, claims linked to admissible sources, and required routing fields.
- Escalation is evaluated as a decision using sensitivity, specificity, and a configurable utility function.The cost matrix must be declared before evaluation and reviewed by the task’s clinical governance group.
- When missed escalation is judged more harmful, the utility assigns a higher cost to false negatives than false positives.This exposes the trade-off instead of silently rewarding excessive deferral or unsafe completion.
VII. ILLUSTRATIVE FAILURE ANALYSIS
The failure analysis illustrates why workflow episodes can reveal clinically salient errors that static benchmarks miss. A polished response may still be inadequate when worsening symptoms require escalation and complete handoff information.
- A patient message about worsening shortness of breath after a medication change may require escalation rather than a polished explanatory reply.The episode tests recognition of warning signals, retention of prior context, and inclusion of required handoff information.
VIII. DESIGN REQUIREMENTS AND BOUNDARIES
The protocol evaluates memory as governed state management across turns and states a boundary on what its workflow-level evidence can support. Its design emphasizes reproducible handling of retained, revised, and discarded information.
- The protocol operationalizes four requirements while setting a clear boundary on its claims.
- The protocol distinguishes retained, revised, and discarded facts when evaluating information across turns.This treats memory as a governance problem rather than only a context-window problem.
B. Evidence Traceability
The protocol makes evidence traceability and escalation decisions explicit, allowing distinct evaluation of retrieval, attribution, unsupported inference, abstention, and unnecessary escalation.
- The rubric links each claim to admissible sources, enabling separate scores for retrieval, attribution, and unsupported inference.
- Episodes specify appropriate actions and a pre-declared cost matrix for scoring abstention, escalation, and unnecessary escalation.This is particularly relevant to triage and patient-facing communication.
D. Institutional Realism
The protocol can model realistic workflow complications, but it does not reproduce institutional implementation or establish patient outcomes or deployment value.
- Episodes represent fragmented records, missing information, interruptions, and specialty-specific conventions.
- The protocol does not reproduce institutional implementation or patient outcomes, which require prospective evaluation.
- The paper should be read as a testable measurement protocol rather than evidence that any agent improves healthcare work.It identifies a clinically governed pilot as the immediate next step.