Source-linked AI summary
Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation
Jiacheng Xie, Xiaoting Tang, Yang Yu, Jinpu Li, Shouli Li, Congcong Jing, Yantao Yang, Zhiyong Zhao, Ziyang Zhang, Qilin Song, Guanghui An, Dong Xu
TL;DR
Complex or unresolved clinical presentations remain difficult to manage, creating uncertainty before diagnosis or treatment and burdening clinicians who integrate incomplete histories. The study evaluated LLMs and physicians on real-world TCM cases using blinded expert review across diagnostic and therapeutic dimensions. Cutting-edge general-purpose LLMs reached or exceeded average physician performance under blinded expert review, particularly in medical advice and treatment principles.
Problem
Complex or unresolved clinical presentations remain difficult to manage, creating uncertainty before diagnosis or treatment and burdening clinicians who integrate incomplete histories.
Method
The study evaluated LLMs and physicians on real-world TCM cases using blinded expert review across diagnostic and therapeutic dimensions.
Results
Cutting-edge general-purpose LLMs reached or exceeded average physician performance under blinded expert review, particularly in medical advice and treatment principles.
Takeaways & Limitations
The findings support potential LLM integration into clinical outputs while indicating that physician oversight and safeguards remain important for reliable deployment.
Takeaways & Limitations
High expert scores did not necessarily imply alignment with real therapeutic tasks, including prescription practices.
Abstract
from arXiv · showhide
Large language models (LLMs) are increasingly being explored for clinical applications, yet their assessment for real-world traditional Chinese medicine (TCM) practice remains limited We constructed a clinical case library comprising 349 de-identified outpatient cases from 62 hospitals and evaluated 16 LLMs and a comparator cohort of 60 practicing TCM physicians using 60 representative cases selected from this library. Model outputs and physician reports were anonymized and scored by five senior TCM experts across nine diagnostic and therapeutic dimensions. Cutting-edge general-purpose LLMs achieved higher expert scores than the physician comparators, particularly for medical advice, treatment principles and selected diagnostic tasks. However, prescription-level analyses revealed discrepancies in herb selection, dosage, and treatment strategy, and qualitative safety review identified hallucinations and undesirable template-driven outputs. These findings highlight the potential of LLMs for TCM decision support while underscoring the need for physician oversight, safety constraints and prospective clinical evaluation.
Introduction
Complex or unresolved clinical cases burden clinicians because diagnosis and treatment require integrating incomplete histories, diverse findings, and expanding literature under limited consultation time. LLMs have attracted interest because they can process free-text information and generate case-specific responses.
- Complex or unresolved cases can involve prolonged diagnostic pathways, repeated investigations, and extended uncertainty before appropriate care is reached.
- Clinicians must integrate incomplete histories, diverse clinical findings, and growing medical literature within limited consultation time.
- LLMs are being explored because they can process free-text clinical information, summarize dispersed knowledge, and generate case-specific responses.
5. Studies have shown that LLMs can perform well on medical examinations and clinical
This evaluation compares LLMs with physicians on real-world TCM cases across diagnostic and therapeutic dimensions. Cutting-edge general-purpose models generally outperformed physicians, including on more difficult cases, but performance varied substantially across models and prescriptions could diverge from clinical practice.
- Evaluation framework: The evaluation used real-world outpatient cases, physician comparators, blinded expert review, and prescription-level analysis across the practical sequence from case interpretation to treatment planning.The dataset was developed to support this evaluation and contains comprehensive patient-level information.
- Diagnostic performance: DeepSeek-R1 achieved the highest overall performance, exceeding the physician baseline by 0.667 points on average, followed by Claude Opus 4 and GPT-5.Claude Opus 4 and GPT-5 exceeded the baseline by 0.629 and 0.627 points, respectively.
- Diagnostic performance: Cutting-edge general-purpose LLMs generally outperformed physicians across multiple dimensions, especially Medical Advice and Usage, while performance remained heterogeneous across models.Several models exceeded the physician baseline by more than 1 point in Medical Advice and Usage; lower-performing models remained substantially below physicians.
- Case-level analyses: DeepSeek-R1 won against physicians in 58/60 cases, while Claude Opus 4 and Gemini 2.5 Pro each won in 57/60 cases.GPT-5 and GPT-o3 each won in 55/60 cases, reinforcing the leading models’ case-level advantage.
- Case-level analyses: Harder cases tended to favor LLMs relative to physicians, whereas easier cases more often showed comparatively stronger human performance.The relative LLM advantage was therefore maintained, and sometimes amplified, under more challenging clinical conditions.
Discussion
Cutting-edge general-purpose LLMs showed strong performance in structured TCM case evaluation, but their prescriptions, reasoning, and output styles remained incompletely aligned with physician practice. These gaps, together with limited sample scope and static evaluation, support decision-support use only with safeguards and clinician oversight.
- Overall performance: Cutting-edge general-purpose LLMs achieved higher expert scores than physician comparators across multiple clinical dimensions, especially medical advice, treatment principles, and selected diagnostic tasks.The study evaluated model outputs through blinded expert assessment across multidimensional diagnostic and therapeutic tasks.
- Model comparisons: TCM-specialized models showed no systematic advantage over general-purpose counterparts, while model performance remained highly heterogeneous across clinical dimensions.Some TCM-oriented models remained substantially below physicians overall, contrasting with the strongest general-purpose models.
- Safety and usability: Comparable quantitative performance did not ensure real-world clinical alignment because models exhibited hallucinations, repetitive or template-driven outputs, and substantial differences in explanatory style and formatting.Practical utility also depends on conformity with documentation norms, physician review, and suitability for patient-facing use.
- Prescription alignment: LLM prescriptions differed from physician practice in herb composition, ingredient selection, dosage, and treatment strategy, including more auxiliary ingredients, higher dosages, and occasional costly or protected substances.Physicians more often balanced therapeutic efficacy with affordability, producing more concise and cost-conscious prescriptions.
- Clinical reasoning: Physicians commonly incorporated follow-up, prescription adjustment, and further examinations, whereas LLMs produced more definitive static conclusions and less holistic integration of coexisting conditions.Model diagnoses more often resembled concatenated symptom-based judgments than holistic syndrome-based assessment.
- Implications and limitations: The findings support LLMs as potential TCM decision-support tools and efficient aids for information organization, diagnostic-highlight extraction, and structured-note drafting, but not as replacements for clinician judgment.The evaluation used a limited, potentially regionally biased sample, static case descriptions, and qualitative error analyses, so performance may not directly translate to clinical settings.
Methods
The study assembled representative real-world TCM outpatient cases and benchmarked general-purpose and TCM-specialized LLMs against practicing physicians under standardized conditions. Cases preserved detailed diagnostic, treatment, follow-up, and clinical-context information for structured evaluation.
- Case library: Each case integrated demographic, symptomatic, four-diagnostic-method, examination, laboratory, imaging, and pathology information.The structured records also documented prescriptions, administration methods, treatment duration, and detailed medical advice.
- Physician comparator: Sixty physicians from a recruited cohort reviewed assigned cases selected partly according to subspecialty expertise under the same formatting instructions as LLMs.Each physician evaluated six cases, and each case was independently evaluated by six physicians; responses underwent completeness and formatting quality control.
- Physician comparator: The final physician cohort represented varied experience, professional levels, sexes, and clinical specialties.The cohort had a mean practice duration of 13.47 years and included general TCM, pediatrics, Tuina, acupuncture, and multiple medical subspecialties.
- LLM benchmark: Sixteen publicly accessible LLMs comprised 12 general-purpose systems and 4 TCM-specialized systems selected for representative, reproducible evaluation.Some TCM-oriented models were excluded because weights, inference services, or reproducible deployment conditions were unavailable.
- LLM benchmark: All LLMs received the same 60 cases and standardized prompts requesting structured outputs across diagnostic and therapeutic dimensions.The input included complete de-identified clinical records, while outputs were evaluated across diagnosis, syndrome differentiation, treatment, and medical-advice dimensions.
Code availability
The study provides code and evaluation materials to support reproduction of its main and supplementary analyses.
- Code availability: Code and evaluation resources are available on GitHub, including prompts, model settings, deployment details, repeated-run configurations, and statistical scripts.The repository supports reproducible benchmarking of physicians and LLMs.