Source-linked AI summary
Evaluation of ChatGPT-Generated Medical Responses: A Systematic Review and Meta-Analysis
Qiuhong Wei, Zhengxiong Yao, Ying Cui, Bo Wei, Zhezhen Jin, Ximing Xu
TL;DR
Medical ChatGPT evaluations lack standardized methods, limiting consistency in the available evidence. This study systematically reviewed and meta-analyzed the literature, finding an overall integrated accuracy of 56% while documenting substantial methodological heterogeneity and incomplete reporting.
Problem
Medical ChatGPT evaluations lack standardized guidelines, with variation in study design, question formulation, and evaluation metrics.
Method
The study systematically reviewed literature on ChatGPT’s medical performance, assessed methodological quality, and synthesized objective accuracy studies through meta-analysis.
Results
ChatGPT showed an overall integrated accuracy of 56% in addressing medical queries.
Takeaways & Limitations
ChatGPT demonstrates considerable potential for application in healthcare.
Takeaways & Limitations
Study heterogeneity and insufficient reporting may affect the reliability of the results.
Abstract
from arXiv · showhide
Large language models such as ChatGPT are increasingly explored in medical domains. However, the absence of standard guidelines for performance evaluation has led to methodological inconsistencies. This study aims to summarize the available evidence on evaluating ChatGPT's performance in medicine and provide direction for future research. We searched ten medical literature databases on June 15, 2023, using the keyword "ChatGPT". A total of 3520 articles were identified, of which 60 were reviewed and summarized in this paper and 17 were included in the meta-analysis. The analysis showed that ChatGPT displayed an overall integrated accuracy of 56% (95% CI: 51%-60%, I2 = 87%) in addressing medical queries. However, the studies varied in question resource, question-asking process, and evaluation metrics. Moreover, many studies failed to report methodological details, including the version of ChatGPT and whether each question was used independently or repeatedly. Our findings revealed that although ChatGPT demonstrated considerable potential for application in healthcare, the heterogeneity of the studies and insufficient reporting may affect the reliability of these results. Further well-designed studies with comprehensive and transparent reporting are needed to evaluate ChatGPT's performance in medicine.
Introduction
ChatGPT has been studied across diverse medical applications, but evaluation remains inconsistent because standardized guidelines are lacking. This review addresses that gap by systematically examining performance, study methods, and evaluation metrics.
- ChatGPT has been evaluated across diverse medical fields, with reported accuracies ranging from 36% to 90%.
- The field lacks standardized guidelines for evaluating large language models, producing inconsistencies across studies.
- Existing evaluations differ in question sources, question-asking processes, and performance metrics.
- The study conducts a systematic review, meta-analysis, and methodological examination of research evaluating ChatGPT on medical questions.
- The review aims to characterize the current evidence and provide useful information for future studies.
Methods
The authors conducted a systematic literature search and independently screened, extracted, and appraised studies evaluating ChatGPT in medicine. Objective accuracy studies were synthesized using meta-analytic methods, with heterogeneity and robustness assessed statistically.
- The literature search covered ten databases on June 15, 2023, using “ChatGPT” without publication, language, or date restrictions.
- Two authors independently screened records and extracted data, with disagreements and extraction conflicts resolved through discussion or a third reviewer.
- Extracted characteristics included question sources, language, number and type of questions, ChatGPT version, inquiry mode, repetition, prompting, raters, metrics, and performance.
- Methodological quality was independently assessed with QUADAS-2 across question selection, index test, reference standard, and flow and timing.
- Objective medical-knowledge accuracy studies were meta-analyzed using confidence intervals, heterogeneity testing, subgroup analysis, sensitivity analysis, and publication-bias assessment.
Results
The review identified 3520 records, narrowed them to 60 studies, and included 17 in the meta-analysis. The included literature spanned publication types and medical disciplines, with no detected publication bias.
- 3520 records were identified, 60 studies were included in the systematic review, and 17 underwent formal meta-analysis.
- No publication bias was detected, with Egger’s test yielding a t-value of 0.02 and p-value of 0.984.
- 97% agreement was observed between the two reviewers for study inclusion and data extraction, with a third reviewer involved in 3% of studies.
- The 60 studies were published between January and May 2023 and included 37 articles, 16 letters, and seven short reports.
- The studies covered internal medicine, surgery, obstetrics and gynecology, specialty departments, clinical auxiliary departments, and other medical areas.
Question/Prompt Sources
The included studies used heterogeneous question sources, formats, languages, and quantities. Author-designed and examination-bank questions were common, while most studies used open-ended or English-language questions.
- 22 studies used author-designed questions, 21 used examination banks, and 13 used open websites.
- Among publicly sourced questions, 13 studies selected material developed after September 2021, while 11 did not report timing.
- 37 studies used English questions, one used English and Chinese, one used Dutch, one used Korean, and 20 did not report language.
- The median number of questions was 66, with an interquartile range of 25 to 195, although three studies did not provide exact counts.
- 36 studies used open-ended questions, 19 used multiple-choice questions, and five used both formats.
Conversation Process
The reviewed studies used varied ChatGPT versions, access modes, question-sequencing practices, repetition schedules, and prompts. Reporting was often incomplete, limiting comparability across conversation processes.
- 13 studies used GPT-3.5, 3 used GPT-3, 4 used GPT-4, and some combined multiple versions.
- 43 studies accessed ChatGPT through a web interface, none reported API use, and 17 did not report the access mode.
- 18 studies asked questions independently in new chat windows, while only one entered all questions in a single session.
- 14 studies asked each question once, whereas 17 repeated questions two to six times and 29 did not report repetition.Repetition occurred twice in 6 studies, three times in 7, five times in 3, and six times in 1.
- 40 studies simply prompted ChatGPT to answer, one used a laboratory-medicine-expert prompt, and 19 did not report prompting details.
Evaluation of ChatGPT's Performance
Studies evaluated ChatGPT with heterogeneous questions, raters, metrics, and methodological procedures, although formal risk-of-bias assessments were generally favorable. The meta-analysis estimated 56% integrated accuracy, with substantial heterogeneity and higher accuracy in internal medicine than surgery.
- Evaluation methods varied across accuracy rates, Likert scales, scoring systems, and metrics including reliability, safety, appropriateness, and accuracy.This diversity made it difficult to identify a standard quantitative metric for uniform assessment.
- Medical professionals evaluated responses in 39 studies, most often with two or three raters in 26 studies.
- The 17-study meta-analysis used objective multiple-choice questionnaires reporting total questions and correct answers.
- 56% integrated accuracy was estimated across the meta-analysis studies (95% CI: 51%-60%; I2 = 87%).
- 63% accuracy was observed in internal medicine versus 49% in surgery, with no significant differences in other subgroup categories.
Discussion
This review and meta-analysis assessed ChatGPT’s medical question-answering performance and found considerable potential, while substantial heterogeneity and incomplete reporting constrain interpretation. Accuracy differed across specialties, with higher performance in internal medicine than surgery.
- The study provides a comprehensive assessment of ChatGPT’s performance in answering medical questions.
- 56% overall integrated accuracy was reported for ChatGPT in addressing medical queries.
- 63% accuracy in internal medicine exceeded 49% in surgery.
- Internal medicine draws on synthesis of symptoms, laboratory results, and medical literature, whereas surgery emphasizes manual skills, tactile sensation, and real-world experience.
- Surgical scenarios may exceed text-based models’ capabilities because they require procedural nuance and tactile feedback, potentially producing less accurate responses.
- The authors propose standardized reporting and a 20-item checklist covering task generation, LLM version, conversation structure, and evaluation procedure.
Data Sharing
The extracted data used in the meta-analysis are available in the Supplementary Material, with additional data available on request.
- Meta-analysis data are available in the Supplementary Material.
- The paper distinguishes extracted meta-analysis data from additional data requested separately.
- Additional data are available on request.
National Knowledge Infrastructure; CBM: Chinese BioMedical Literature Database;
The paper identifies Chinese literature resources and presents figures describing ChatGPT’s medical assessments and performance.
- VIP is identified as the VIP Database for Chinese Technical Periodicals.
- Figure 2 presents the medical examinations and assessments undertaken by ChatGPT.
- Figure 3 presents ChatGPT’s performance using correct answers, total questions, 95% confidence intervals, and a fixed-effect model.