Source-linked AI summary

A Framework for Human Evaluation of Large Language Models in Healthcare Derived from Literature Review

Thomas Yu Chow Tam, Sonish Sivarajkumar, Sumit Kapoor, Alisa V Stolyar, Katelyn Polanska, Karleigh R McCarthy, Hunter Osterhoudt, Xizhi Wu, Shyam Visweswaran, Sunyang Fu, Piyush Mathur, Giovanni E. Cacciamani, Cong Sun, Yifan Peng, Yanshan Wang

arXiv:2405.02559v2cs.CLcs.AI

TL;DR

Automated metrics do not assess the clinical utility and accuracy needed for healthcare deployment, while human evaluation lacks established, healthcare-specific guidelines. This review applies a PRISMA-guided literature search and synthesizes evaluation practices into QUEST, a five-principle framework for human evaluation of healthcare LLMs.

  • Problem

    Automated metrics do not assess the clinical utility and accuracy needed for healthcare deployment, and human evaluation lacks established guidelines tailored to healthcare LLMs.

  • Method

    The study conducts a PRISMA-guided review of publications from January 2018 to February 2024 and synthesizes literature on human evaluation practices into proposed best practices.

  • Results

    The review categorizes evaluation methods into 17 dimensions grouped under QUEST: Quality of Information, Understanding and Reasoning, Expression Style and Persona, Safety and Harm, and Trust and Confidence.

  • Takeaways & Limitations

    QUEST provides guidelines for consistent, high-quality human evaluations intended to assess healthcare LLM safety, reliability, and effectiveness methodically.

  • Takeaways & Limitations

    Human evaluation approaches were inconsistent between trials.

Abstract

from arXiv · show

With generative artificial intelligence (AI), particularly large language models (LLMs), continuing to make inroads in healthcare, it is critical to supplement traditional automated evaluations with human evaluations. Understanding and evaluating the output of LLMs is essential to assuring safety, reliability, and effectiveness. However, human evaluation's cumbersome, time-consuming, and non-standardized nature presents significant obstacles to comprehensive evaluation and widespread adoption of LLMs in practice. This study reviews existing literature on human evaluation methodologies for LLMs in healthcare. We highlight a notable need for a standardized and consistent human evaluation approach. Our extensive literature search, adhering to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, includes publications from January 2018 to February 2024. The review examines the human evaluation of LLMs across various medical specialties, addressing factors such as evaluation dimensions, sample types and sizes, selection, and recruitment of evaluators, frameworks and metrics, evaluation process, and statistical analysis type. Drawing on the diverse evaluation strategies employed in these studies, we propose a comprehensive and practical framework for human evaluation of LLMs: QUEST: Quality of Information, Understanding and Reasoning, Expression Style and Persona, Safety and Harm, and Trust and Confidence. This framework aims to improve the reliability, generalizability, and applicability of human evaluation of LLMs in different healthcare applications by defining clear evaluation dimensions and offering detailed guidelines.

1. Introduction

Healthcare LLMs require human evaluation because automated metrics do not capture clinical utility, safety, and ethical reliability. This review addresses gaps in systematic coverage and standardization by proposing actionable evaluation guidance.

  • Automated metrics such as accuracy, F measure, and AUROC do not assess the clinical utility and accuracy required for healthcare deployment.
  • Prior reviews summarized evaluator characteristics and metrics but did not systematically synthesize human-evaluation dimensions or establish a standardized healthcare framework.
  • The study systematically reviews human evaluations of LLMs across medical domains, tasks, and specialties.
  • The authors synthesize best practices and actionable guidelines for reliable, valid, ethical, and standardized human evaluation.

2. Method

The review follows PRISMA and searches peer-reviewed English-language healthcare literature from January 2018 through February 2024. Studies were screened in two stages for relevance and methodological detail on human evaluation of LLM outputs.

  • The literature search followed PRISMA guidelines and covered publications from January 1, 2018, to February 22, 2024.
  • The review focused on peer-reviewed journal articles and conference proceedings published in English.
  • PubMed searches combined terms related to generative large language models, human evaluation, and healthcare.
  • The search returned 1191 results, while commentaries, summaries, and non-peer-reviewed preprints were excluded.
  • Two-stage screening used title-and-abstract review followed by full-text assessment emphasizing methodological detail and clinical applicability.

3. Results

The reviewed literature covers diverse healthcare applications of LLMs, reflecting uses across patient care, clinical practice, biomedical research, and education. The section emphasizes variation in methodologies and questionnaires.

  • The findings synthesize varied methodologies and questionnaires while identifying current practices, challenges, and areas for future research.

3.1 Healthcare Applications of LLMs

Human-evaluated healthcare LLM studies span clinical decision support, education, patient education, question answering, research, and diagnostic tasks. Clinical decision support is the largest application category, while specialized research performance can remain below expert standards.

  • Healthcare Applications of LLMs: Clinical decision support accounted for 31.9% of categorized tasks, the largest application category in the review.
  • Healthcare Applications of LLMs: Medical education and examination represented 24.8% of tasks, including evaluations on licensing examinations and higher-order medical biochemistry problems.
  • Healthcare Applications of LLMs: Patient education represented 19.6% of tasks, including vaccination answers, where one chatbot study reported performance exceeding medical students.
  • Healthcare Applications of LLMs: Patient-provider question answering accounted for 15% of tasks and included assessments of accurate, empathetic, and orthopedic information for patients.
  • Healthcare Applications of LLMs: In translational research, ChatGPT performed well in specific colorectal-cancer domains but generally fell short of expert standards, with limitations in specialized surgical research.
  • Healthcare Applications of LLMs: Studies evaluated diagnostic suggestions, medical evidence compilation, clinical-record clarity, and comparisons with clinician or gold-standard diagnoses.

3.2 Medical Specialties

The review found human-evaluated LLM studies spanning many medical specialties, with radiology most represented and varied specialty-specific evaluation needs.

  • Radiology led the reviewed specialties with 12 articles, followed by urology with 9 and general surgery with 8.
  • Plastic surgery, otolaryngology, ophthalmology, and orthopedic surgery each accounted for 7 articles, while psychiatry accounted for 6.
  • Radiology evaluations addressed report quality, accuracy, clinical utility, and interpretability in relation to radiological practice.
  • Urology studies used patient-centered methods including satisfaction surveys, feedback forms, and usability assessments for education and disease-management applications.
  • Emergency-medicine evaluations focused on time-critical decision-making and triage, including decision accuracy, timeliness, and resource utilization.

3.3 Evaluation Design

The review characterizes healthcare LLM evaluation as multifaceted, combining quantitative and qualitative methods, and proposes QUEST to organize 17 dimensions into five principles.

  • Evaluation Principles and Dimensions: QUEST groups 17 evaluation dimensions into Quality of Information, Understanding and Reasoning, Expression Style and Persona, Safety and Harm, and Trust and Confidence.
  • Evaluation Principles and Dimensions: Quality of Information covers accuracy, relevance, currency, comprehensiveness, consistency, agreement, and usefulness, while Understanding and Reasoning addresses prompt interpretation and logical processing.
  • Evaluation Principles and Dimensions: Expression Style and Persona measures clarity and empathy, whereas Safety and Harm includes bias, harm, self-awareness, fabrication, falsification, and plagiarism.
  • Evaluation Principles and Dimensions: Trust and Confidence captures user trust and satisfaction, with related strategies including Likert scales and presence-or-absence categories.
  • Evaluation Design: Evaluation designs considered sample number and variability, response consistency across similar queries, and comparisons using different prompts or time points.
  • Evaluation Checklist: The review notes that evaluator checklists should clarify evaluation criteria, although reporting was insufficient to determine whether training and examples aligned evaluators with study expectations.
  • Evaluation Checklist: Studies used concrete questions assessing comprehension, reasoning, clarity, empathy, bias, harm, recognition of limits, reliability, usefulness, and response consistency.

3.3.3 Evaluation Samples

The review examines how evaluation samples, human evaluators, and assessment procedures are selected for healthcare LLMs. It highlights tradeoffs among sample size, evaluator breadth, evaluation dimensions, blinding, and statistical comparison.

  • Evaluation samples: Most reviewed studies evaluated 100 or fewer LLM outputs, while one MTurk study evaluated 2,995 sentences seven times each.The larger sample accounted for variability in annotations associated with annotators’ reading age and language capabilities.
  • Evaluation samples: Sample variability supports evaluating LLM outputs across patient subpopulations characterized by symptoms, diagnoses, or demographic information.Reviewed studies used patient-specific information and clinician-prepared clinical vignettes to vary prompts.
  • Human evaluators: Most articles reported 20 or fewer human evaluators, and only three studies recruited more than 50 evaluators.Clinical or clinician-facing studies generally reported evaluator qualifications, specialties, experience, positions, and sometimes demographics.
  • Human evaluators: Non-expert evaluation can support larger-scale, lower-cost assessment, but reviewed studies generally used more evaluators across fewer dimensions than expert evaluations.This pattern suggests a potential tradeoff between evaluator quantity and evaluation breadth.
  • Application-level patterns: Across healthcare applications, patient-facing studies had larger sample sizes and evaluator counts, whereas clinical decision support had the lowest median evaluator count and the second-lowest median sample count.Sample-size variability across studies was high, and Figure 6 showed an inverse relationship between sample size and evaluator count.
  • Evaluation procedures: Human evaluations used tools including binary variables, Likert scales, comparisons with human or guideline benchmarks, and blinded assessments.Among 142 studies, 41 (29%) explicitly reported blinding, 20 (14%) reported unblinded evaluation, and 80 (56%) gave no blinding information.

3.4 Specialized Frameworks

Healthcare LLM studies use generic and specialized frameworks to assess service quality, understandability, learning outcomes, data quality, and information quality. These tools provide focused measures but do not collectively cover all QUEST dimensions.

  • PEMAT-P: PEMAT-P evaluates patient education materials for understandability and actionability using content, word choice, organization, layout, and design.
  • SOLO structure: SOLO structure assesses the complexity and depth of LLM responses across levels from pre-structural to extended abstract.
  • Data quality: Wang and Strong’s framework assesses medical-information data quality through accuracy, believability, completeness, conciseness, timeliness, and relevance.
  • SERVQUAL: SERVQUAL assesses ChatGPT’s medical-information service quality through reliability, responsiveness, assurance, empathy, and tangibles.It was applied with urologists and urological oncologists to evaluate information for kidney cancer patients.
  • Coverage limitation: Specialized questionnaires and frameworks measure selected properties such as factual consistency, harmfulness, coherence, accuracy, reliability, and readability, but lack comprehensive QUEST coverage.

3.5 QUEST Human Evaluation Framework

The QUEST framework standardizes healthcare LLM human evaluation through three phases: Planning, Implementation and Adjudication, and Scoring and Review. It adapts evaluation design to each use case while combining trained human ratings, consensus procedures, and statistical analysis.

  • Framework structure: QUEST organizes evaluation into Planning, Implementation and Adjudication, and Scoring and Review phases.
  • Planning: Planning defines model goals, tasks, stakeholders, and success criteria before selecting samples, checklists, and evaluators.
  • Planning: Dimensions and metrics should match the task and audience, with patient-facing applications emphasizing Understanding, Clarity, and Empathy when appropriate.
  • Implementation: Evaluator training supports consistent assessment, and QUEST can be implemented in crowd-sourced settings such as Mechanical Turk.
  • Adjudication: After output evaluation, statistical tests assess evaluator consistency and agreement; adjudication iteratively updates guidelines until consensus, such as Cohen’s kappa of 0.7 or above.
  • Scoring and Review: Final dimension scores use mean or median evaluator ratings and are compared with automatic metrics such as F1 and AUROC.

4. Case Studies

The case studies apply QUEST to emergency-department clinical note summarization and triage decision support. They illustrate use-case-specific planning, sampling, evaluator selection, tailored dimensions, training, adjudication, and statistical scoring.

  • Clinical Note Summarization: The clinical note summarization case aims to reduce physician documentation burden by extracting key information and producing accurate, consistently formatted summaries.
  • Clinical Note Summarization: The summarization evaluation uses a minimum sample of 130 notes and extracts demographic information where feasible to increase population variability.
  • Clinical Note Summarization: Clinicians prioritize Accuracy and Comprehensiveness, while using a binary Bias metric and AHRQ’s harm scale for Harm.
  • Clinical Note Summarization: The clinician-facing summarization use case recruits at least seven suitable evaluators and provides two hours of training before initial evaluation, adjudication, and guideline refinement.
  • ED Triage Decision Support: The triage case evaluates LLM decisions across emergent, non-emergent, and home-care categories using complaints, history, vital signs, and examination findings.
  • ED Triage Decision Support: A sample of 130 representative triage cases is independently assessed by three ED physicians and three nurses against expert-derived gold-standard decisions.
  • ED Triage Decision Support: Triage evaluators use a standardized questionnaire, with disagreements adjudicated before dimension-specific statistical analysis and final scoring.

5. Discussion

The discussion emphasizes that opaque LLM behavior and limitations of current evaluation practices motivate structured human evaluation. It also identifies scope, implementation, scalability, and outcome-linkage boundaries for QUEST and the review.

  • Rationale: Healthcare LLM evaluations face limited traceability, reliability, and trust because the models’ inner workings remain opaque.
  • Framework limitations: The proposed guidelines remain constrained by human-evaluation scale, reviewed sample size, measurement choices, proprietary models, and limited computational resources.
  • Review limitations: The review may omit post-February 22, 2024 research, non-English articles, studies missed by the search strategy, and non-healthcare evidence.
  • Framework limitations: QUEST implementation and standardization may vary across specialties and institutions because policies, infrastructure, resources, and trained personnel differ.
  • Framework limitations: QUEST may not capture every use case, focuses on human rather than automatic evaluation, and should be adapted or combined with specific frameworks.
  • Future research: Connecting LLM evaluation to actual patient outcomes or scientific understanding remains future work, while scalable low-resource human evaluation also remains unresolved.

Contributions

The study’s contributors jointly handled conceptualization, design, analysis, writing, review, and revision, with designated direction and organization roles. The paper also declares conflicting interests involving ownership and equity in several companies.

  • Contributions: T.Y.C.T. and S.S. conceptualized, designed, organized, analyzed, and contributed to writing, review, and revision.
  • Contributions: Y.W. conceptualized, designed, and directed the study and contributed to writing, review, and revision.
  • Conflicting interests: P.M., Y.W., and S.V. reported ownership and equity in multiple companies, while the other authors declared no potential conflicts.

Appendix

The appendix includes figures comparing and distributing LLMs, a result describing model-family representation, and a case-study table of human-evaluation questions for emergency-department triage.

  • Figures: Figure A1 presents the distribution of LLMs.
  • Results: Most reviewed studies applied OpenAI’s GPT family, whereas open-source models such as Meta’s Llama were not among the leading models.
  • Figures: Figure A2 presents a comparison of LLMs.
  • Case study: The emergency-department triage case study evaluates patient-management outputs for empathy, bias, harm, and trustworthiness.
Loading 2405.02559v2…