Source-linked AI summary
AI chatbots versus human healthcare professionals: a systematic review and meta-analysis of empathy in patient care
Alastair Howcroft, Amber Bennett-Weston, Ahmad Khan, Joseff Griffiths, Simon Gay, Jeremy Howick
TL;DR
Evidence comparing AI chatbots with human practitioners on empathy was previously unsynthesized despite growing clinical use of generative AI. This systematic review and meta-analysis found higher perceived empathy for ChatGPT than human practitioners in text-based interactions, while noting important evaluation and scope limitations.
Problem
Evidence comparing AI chatbot and human practitioner empathy had not been synthesized, limiting informed decisions about integrating AI into patient care.
Method
The review synthesized comparative studies of AI chatbots and human practitioners, pooling 13 studies while preventing double-counting across chatbot models and comparisons.
Results
ChatGPT showed significantly higher empathy ratings than human practitioners (SMD 0.87, 95% CI 0.54–1.20; P < .00001), roughly equivalent to a 2-point difference on a 10-point scale.
Takeaways & Limitations
In text-based interactions, AI chatbots are often perceived as more empathic than human practitioners across clinical contexts, with dermatology as a notable exception.
Takeaways & Limitations
The studies primarily assessed empathy through text-based interactions and proxy ratings rather than direct evaluations by patients receiving care.
Abstract
from arXiv · showhide
Background: Empathy is widely recognized for improving patient outcomes, including reduced pain and anxiety and improved satisfaction, and its absence can cause harm. Meanwhile, use of artificial intelligence (AI)-based chatbots in healthcare is rapidly expanding, with one in five general practitioners using generative AI to assist with tasks such as writing letters. Some studies suggest AI chatbots can outperform human healthcare professionals (HCPs) in empathy, though findings are mixed and lack synthesis. Sources of data: We searched multiple databases for studies comparing AI chatbots using large language models with human HCPs on empathy measures. We assessed risk of bias with ROBINS-I and synthesized findings using random-effects meta-analysis where feasible, whilst avoiding double counting. Areas of agreement: We identified 15 studies (2023-2024). Thirteen studies reported statistically significantly higher empathy ratings for AI, with only two studies situated in dermatology favouring human responses. Of the 15 studies, 13 provided extractable data and were suitable for pooling. Meta-analysis of those 13 studies, all utilising ChatGPT-3.5/4, showed a standardized mean difference of 0.87 (95% CI, 0.54-1.20) favouring AI (P < .00001), roughly equivalent to a two-point increase on a 10-point scale. Areas of controversy: Studies relied on text-based assessments that overlook non-verbal cues and evaluated empathy through proxy raters. Growing points: Our findings indicate that, in text-only scenarios, AI chatbots are frequently perceived as more empathic than human HCPs. Areas timely for developing research: Future research should validate these findings with direct patient evaluations and assess whether emerging voice-enabled AI systems can deliver similar empathic advantages.
Introduction
Empathic healthcare improves patient outcomes, while AI chatbots are increasingly entering patient care and may sometimes replace human practitioner roles. However, evidence comparing chatbot and human empathy remains heterogeneous, with no synthesis directly addressing the gap.
- Empathic healthcare is associated with improved quality of life and satisfaction, alongside reduced pain and psychological distress.
- AI chatbots increasingly support patient care through information provision, symptom monitoring, and emotional support.They can interact through text or speech and perform roles historically provided by human healthcare professionals.
- Twenty percent of UK general practitioners reportedly use generative AI for tasks such as writing patient correspondence.
- Prior studies disagree about whether AI responses are more empathic than human interactions, citing both higher perceived empathy and deficits in warmth or emotional understanding.
- A synthesis comparing AI and human empathy was lacking, limiting informed decisions about integrating AI into patient care.
Methods
The review searched broad healthcare and research databases for empirical comparisons of LLM-based AI chatbots with human HCPs. Eligible evidence was screened and assessed for bias before random-effects meta-analysis, with procedures to reduce double counting.
- Eligible studies empirically compared empathy between LLM-based conversational agents and human HCPs in authentic healthcare interactions.Hypothetical patient scenarios, scripted AI systems, opinion pieces, and reviews were excluded.
- The authors searched seven major databases plus trial registries, reference lists, and grey literature through November 2024.
- The search combined empathy, AI, and HCP terms using OR within categories and AND across categories.
- Reviewers independently screened records and resolved disagreements through discussion, with third-reviewer involvement when necessary.
- Data extraction covered study design, participants, settings, interventions, comparators, empathy measures, and key findings.
- Risk of bias was assessed with ROBINS-I across 15 studies, including 14 nonrandomized analyses and one randomized controlled trial.
- Random-effects models pooled standardized mean differences for GPT-3.5 and GPT-4 while excluding overlapping AI arms or datasets to prevent double counting.
Results
The review included 15 studies spanning varied health concerns, specialties, comparators, and patient-query sources. Most evaluations used text-based GPT systems and observer-rated, custom empathy measures.
- 15 studies met the inclusion criteria after 987 unique records underwent title, abstract, and full-text screening.
- Nine studies had moderate risk of bias and six had serious risk of bias.
- Fourteen studies were published in 2024, while one was published in 2023.
- The included studies covered general medical inquiries, mental health, autism-related questions, patient complaints, and multiple clinical specialties.
- Human comparators included surgeons, specialty physicians, nonspecialist physicians, advanced practice providers, nurses, and frontline staff.
- All but one study relied exclusively on text-based AI interaction, and most evaluated GPT-family models.
- Fourteen of 15 studies used unvalidated or custom empathy measures, while one used the validated CARE scale.
- Evaluators were observers rather than the people asking or answering inquiries, and included patient proxies, HCPs, laypeople, medical students, and psychology trainees.
Results of included studies
Table 1 summarizes the 15 included studies and reports their statistical results as presented in each study. Detailed study-specific metrics and model comparisons are provided separately in Appendix E.
- Table 1 summarizes the characteristics of all 15 included studies.
- Reported statistical results retain the formats presented in the individual studies, including sub-scores and overall means.
- Appendix E provides detailed study-specific results, including granular metrics and model comparisons.
Results of syntheses
Across the included comparisons, AI chatbots were usually rated as more empathic than human healthcare practitioners, although dermatology studies favored human responses. Pooled evidence for ChatGPT-3.5 and GPT-4 showed higher empathy ratings for AI, with GPT-4 producing a significant subgroup effect while GPT-3.5 results were mixed.
- Overall findings: 13 of 15 comparisons found a statistically significant empathy advantage for one or more AI chatbots, while both dermatology studies favored human dermatologists.The exceptions involved Med-PaLM 2 and ChatGPT-3.5 in dermatology.
- Overall findings: 0.87 SMD (95% CI 0.54–1.20; P < .00001) favored ChatGPT-3.5/4 over human practitioners across 13 pooled studies.The pooled analysis combined GPT-3.5 and GPT-4 studies while avoiding overlapping data.
- Other comparisons: Additional comparisons found higher empathy ratings for ChatGPT, Gemini, Le Chat, SPPEC, and ChatGPT-4 than for several human professional groups.These comparisons included physicians, nurses, rheumatologists, mental health professionals, and plastic surgeon or advanced-practice-provider responses.
- GPT-3.5 subgroup: GPT-3.5 showed mixed results: it outperformed clinicians in three of four settings, but its pooled advantage was not statistically significant.The dermatology comparison favored humans, with SMD −0.99 (95% CI: −1.52 to −0.46).
- GPT-4 subgroup: 1.03 SMD (95% CI 0.71–1.35; P < .00001) favored GPT-4 across nine studies, despite high heterogeneity (I2 = 87%).GPT-4 consistently outperformed human clinicians across the contributing studies.
- Subgroup comparison: The GPT-3.5 versus GPT-4 subgroup difference was not statistically significant (P = .16), so GPT-4 was not conclusively more empathic than GPT-3.5.GPT-4 showed more consistent results, whereas the combined analysis still favored both models over human practitioners.
Discussion
In text-based interactions, AI chatbots were often perceived as more empathic than human practitioners, but important limitations constrain interpretation and clinical relevance. Future work should test direct patient perceptions, voice interactions, prompt effects, transparency, and safety-preserving implementation.
- Findings: A 73% probability of superiority indicates ChatGPT was more likely to be perceived as empathic than human practitioners in text-based interactions.The result spans 13 studies, multiple specialties, and varied evaluation methods, corresponding to roughly two points on a 10-point scale.
- Limitations: Text-only evaluations omit non-verbal cues, although text communication remains relevant to some healthcare practice.Empathy commonly involves cues such as nodding and leaning forward, while text-based communication represents a relatively small portion of healthcare interactions.
- Limitations: Most studies used public-forum or one-off interactions, and 14 of 15 focused on GPT-3/3.5 or GPT-4, limiting generalisability to real-world clinical tools.Public posts may differ from private clinical interactions, while proprietary systems may have different architectures and training data.
- Limitations: Proxy raters, varied empathy measures, and incomplete statistics complicate pooled inferences about empathy.Evaluators included patient proxies, lay people, students, and healthcare professionals; some studies required approximations for meta-analysis.
- Limitations: The clinical relevance of statistically significant empathy differences remains uncertain because direct effects on patient outcomes were not established.The reported magnitude, approximately a 20% absolute difference, was considered potentially clinically relevant and warrants further investigation.
- Future research: A clinician-supervised “empathic enhancer” could refine patient-facing messages while retaining clinician control over core medical content.The proposed model aims to improve tone and empathy without compromising accuracy or safety, and should be evaluated through randomized trials.
- Future research: Future research should evaluate direct patient feedback and voice-enabled systems, including whether the text-based advantage persists in telephone consultations.The review also identifies prompt length, empathy instructions, model diversity, and disclosure of AI involvement as important research variables.
Conclusion
The review concludes that generative AI chatbots are often perceived as more empathic than human practitioners in text-based interactions, including across clinical contexts but with dermatology exceptions. It calls for voice-based evaluations, direct patient feedback, and rigorous validation and transparency.
- Conclusion: Generative AI chatbots, particularly GPT-4, are often perceived as more empathic than human practitioners in text-based interactions.The pattern is reported across clinical contexts, with notable exceptions in dermatology.
- Conclusion: Future research should assess voice-based interactions and direct patient feedback while using randomized trials to support clinical reliability.The conclusion also emphasizes rigorous validation and transparency.
Funding
The study received no specific grant funding from public, commercial, or not-for-profit sectors.
- Funding: No specific grant funding was received from public, commercial, or not-for-profit sectors.
Data availability
All data underlying the article are available in the main text and supplementary materials, with no additional datasets generated or analysed.
- Data availability: The main text and Appendices A–E provide the underlying data and supporting materials.The appendices include the search strategy, excluded studies, narrative syntheses, ROB assessments, and study-specific results.
Protocol and registration
The study followed PRISMA-P guidance and was prospectively registered with the Open Science Framework. It was conducted as a systematic review with meta-analysis, with eligibility criteria clarified after registration.
- The protocol followed Preferred Reporting Items for Systematic Reviews and Meta-Analysis Protocols (PRISMA-P) guidelines.
- The protocol was prospectively registered with the Open Science Framework on 28 October 2024.
- The study was conducted as a systematic review with meta-analysis rather than the scoping review described in the original protocol.
- Eligibility criteria were refined through clarifications incorporated after protocol registration.