Source-linked AI summary
The Ethics of ChatGPT in Medicine and Healthcare: A Systematic Review on Large Language Models (LLMs)
Joschka Haltaufderheide, Robert Ranisch
TL;DR
LLMs in healthcare offer potential benefits but raise ethical concerns, including documented harms and risks from misinformation, bias, opacity, and privacy challenges. This review examines those concerns and concludes that ethical guidance should define acceptable human oversight across varied healthcare applications and risk settings.
Problem
LLM adoption in healthcare raises ethical and social concerns, including documented real-world harms and potentially life-threatening patient consequences.
Method
The paper reviews ethical considerations of LLM use in healthcare and frames current implementation as a social experiment requiring iterative learning within ethical limits.
Results
The review identifies potential benefits alongside recurrent concerns about fairness, bias, non-maleficence, transparency, privacy, harmful misinformation, and inaccurate content requiring human oversight.
Takeaways & Limitations
Ethical guidance should define acceptable human oversight and validation according to users, application risks, healthcare settings, and acceptable performance and certainty thresholds.
Takeaways & Limitations
The review is limited by nascent evidence, substantial reliance on non-peer-reviewed preprints, and variation across settings, applications, and interpretations of LLMs.
Abstract
from arXiv · showhide
With the introduction of ChatGPT, Large Language Models (LLMs) have received enormous attention in healthcare. Despite their potential benefits, researchers have underscored various ethical implications. While individual instances have drawn much attention, the debate lacks a systematic overview of practical applications currently researched and ethical issues connected to them. Against this background, this work aims to map the ethical landscape surrounding the current stage of deployment of LLMs in medicine and healthcare. Electronic databases and preprint servers were queried using a comprehensive search strategy. Studies were screened and extracted following a modified rapid review approach. Methodological quality was assessed using a hybrid approach. For 53 records, a meta-aggregative synthesis was performed. Four fields of applications emerged and testify to a vivid exploration phase. Advantages of using LLMs are attributed to their capacity in data analysis, personalized information provisioning, support in decision-making, mitigating information loss and enhancing information accessibility. However, we also identifies recurrent ethical concerns connected to fairness, bias, non-maleficence, transparency, and privacy. A distinctive concern is the tendency to produce harmful misinformation or convincingly but inaccurate content. A recurrent plea for ethical guidance and human oversight is evident. Given the variety of use cases, it is suggested that the ethical guidance debate be reframed to focus on defining what constitutes acceptable human oversight across the spectrum of applications. This involves considering diverse settings, varying potentials for harm, and different acceptable thresholds for performance and certainty in healthcare. In addition, a critical inquiry is necessary to determine the extent to which the current experimental use of LLMs is necessary and justified.
I. INTRODUCTION
LLMs have rapidly attracted interest in medicine and healthcare, but their adoption raises ethical concerns and lacks a comprehensive systematic overview. This review maps researched applications and related ethical considerations.
- Healthcare LLM research spans clinical, educational, and research applications following ChatGPT’s public launch.
- Healthcare is especially exposed to ethical dilemmas because it is sensitive, heavily regulated, and governed by stringent professional and societal obligations.
- Key concerns include inadequate information, privacy risks from sensitive health data, and harmful gender, cultural, or racial biases.
- Case reports document actual, potentially life-threatening patient harm associated with ChatGPT.
- The review addresses the deficit of systematic overviews by mapping applications, ethical considerations, outcomes, opportunities, risks, benefits, and potential harms.
II. METHODS
The authors conducted a registered rapid-review-style synthesis of healthcare LLM literature, combining structured screening, extraction, quality appraisal, and meta-aggregation. The included evidence was organized into four application themes.
- A registered protocol guided searches of publication databases and preprint servers, with two-stage screening based on intervention, setting, and outcomes.
- Data were extracted with a self-designed form, coded in MaxQDA, and refined through independent coding and joint discussions.
- A meta-aggregative synthesis iteratively refined categories covering actors, values, device properties, arguments, recommendations, and conclusions until saturation.
- Quality appraisal used a hybrid approach distinguishing procedural quality control from critical assessment of comprehensiveness and validity.
- 796 database hits yielded 53 included records after duplicate removal and full-text assessment.The process included 738 title/abstract screenings and 158 full-text assessments.
- Four themes structured the results: clinical applications, patient support, support for health professionals, and public health perspectives.
A. Clinical applications
Clinical applications center on predictive analysis, diagnosis support, and triage, with potential efficiency and patient benefits tempered by risks from inaccurate, biased, and difficult-to-validate outputs.
- 1) Predictive analysis and risk assessment:: LLMs are proposed as clinical “co-pilots” that use patient information to flag concerns, predict diseases, and assess risk before or during clinical situations.
- 1) Predictive analysis and risk assessment:: Predicting health outcomes and patterns is described as potentially improving patient outcomes and benefiting patients.
- 1) Predictive analysis and risk assessment:: LLM-assisted triage of emergency-department notes could reduce length of stay and improve waiting-room time utilization.
- 1) Predictive analysis and risk assessment:: Clinical applications require close human oversight because inaccurate information could directly harm patients or provide dangerous rationales for decisions.
- 1) Predictive analysis and risk assessment:: Unstructured medical notes complicate accuracy prediction, fine-tuning, interpretability, and clinicians’ ability to identify inaccurate outputs.
2) Patient consultation and communication:
LLMs may improve patient-provider information exchange by bridging care settings, reducing communication barriers, and supporting engagement. Their integration remains uncertain and raises concerns about safety, equity, data protection, and human care.
- 2) Patient consultation and communication:: LLMs may bridge clinical and preclinical settings by facilitating informational exchange, self-management, community aids, and timely workflow support.
- 2) Patient consultation and communication:: Collecting patient information, providing additional information, translating language, and simplifying medical jargon may support decisions, satisfaction, engagement, and communication.
- 2) Patient consultation and communication:: Where, when, and how LLMs should be integrated into practice remains unclear in the reviewed dataset.
- 2) Patient consultation and communication:: These applications must address patient-data protection, safety, potentially unjust disparities, and the therapeutic relationship.
- 2) Patient consultation and communication:: Robust expert oversight is recommended to reduce incorrect information, while preserving the human dimension of care.
- 2) Patient consultation and communication:: Open communication and consent to technical mediation are required to promote trust but may be difficult to achieve.
3) Diagnosis:
LLMs are explored for diagnosis because they can analyze large amounts of unstructured data, but bias, opacity, hallucinations, and limited clinical reasoning raise risks requiring validation and human oversight.
- LLMs may support timely, efficient, and more accurate diagnosis by analyzing large amounts of unstructured data.
- Training-data bias can produce unfair treatment, unequal access, and harm to marginalized or vulnerable groups.
- Interpretability problems, hallucinations, and falsehood mimicry increase the risks associated with LLM-supported diagnosis.
- Opaque systems can hinder diagnostic justification, weaken professional authority, and erode trust, so generated data require clinical validation.
- Ethical evaluation should compare LLM diagnosis with existing alternatives because unaided diagnosis is also subjective and error-prone.
4) Treatment planning:
LLMs are investigated for personalized treatment recommendations and patient-facing information, while biases, privacy risks, and inaccurate recommendations complicate their ethical acceptability.
- Six studies examine personalized treatment regimens or treatment-decision support based on electronic patient information or history.
- Treatment-planning systems may perpetuate biases and stereotypes, disproportionately benefiting some groups while disadvantaging others.
- Entering patient data raises confidentiality, privacy, and data-security concerns, especially for commercial and publicly available models.
- Inaccurate treatment recommendations are identified as a potential source of harm.
- Patient-facing LLMs can improve access to comprehensible information, support health literacy, and facilitate autonomy and cross-lingual communication.
- Ethical acceptability varies by application area, with greater potential risks reported in pharmacology and mental health than in some other fields.
2) Symptom assessment and health management:
LLMs can provide personalized health guidance and support symptom assessment, triage, and emergency management, but their limited situational awareness may cause severe harm.
- LLMs can offer personalized guidance on lifestyle adjustments, symptom self-assessment, self-triage, and emergency-management steps.
- Although current systems can generate compelling responses, their common lack of situational awareness may lead to severe harm.
- Situational awareness requires tailoring responses to contextual criteria such as personal circumstances, medical history, or social situation.
- Most current LLMs cannot seek clarification by asking questions, limiting context-sensitive support.
- Automating medical reports, patient-interaction summaries, forms, and discharge summaries could streamline clinical workflows and save professionals time.
2) Research:
LLMs are explored across research support and public-health applications to reduce routine workload and expand access, but risks include distorted sources, weakened research integrity, misinformation, and unequal accessibility.
- Research applications: LLMs may summarize research text, evidence, and data; identify research targets; design studies; facilitate collaboration; and communicate results.
- Research applications: These capabilities could accelerate research and reduce routine workload, enabling more efficient research workflows.
- Research risks: Overreliance may cause deskilling and undue influence on research outcomes, prompting calls for vigilance, revalidation, and strict human oversight.
- Public health perspectives: Public-health applications include campaigns, outbreak monitoring through news and social media, and targeted communication strategies.
- Public health perspectives: LLMs may improve low-cost access to health information and health literacy in low-resource settings facing shortages of professionals or inequitable resource distribution.
- Public health risks: LLMs have dual-use potential: they may expand information access while enabling malicious actors to spread harmful health messages at unprecedented scale.
- Public health risks: An AI-driven infodemic could overwhelm recipients with imprecise, unclear, or false information, causing disorientation and potentially harmful behavior.
IV. DISCUSSION
The review finds broad experimentation with LLMs in healthcare, alongside recurrent ethical concerns and a need to define acceptable human oversight. It also questions whether current experimental uses are sufficiently necessary and justified.
- LLM research spans diverse healthcare applications, reflecting a vivid testing phase despite limited real-world clinical deployment.The review attributes the promise of these tools to data analysis, personalized information, decision support, reduced information loss, and improved accessibility.
- Recurrent concerns involve fairness, bias, non-maleficence, transparency, and privacy.
- LLMs can produce harmful misinformation or convincing but inaccurate content, creating severe challenges for clinical accuracy and reliability.Their statistical architecture and opacity make machine-generated outputs difficult to validate, while inaccurate information may directly harm patients or justify dangerous decisions.
- Erroneous outputs make human oversight and continual validation necessary, especially in the absence of professional guidelines or regulatory oversight.
- Ethical guidance should define acceptable human oversight and validation according to application, user context, potential harm, and uncertainty.
- The review calls for critical examination of whether current experimental LLM use is necessary and justified, because some research is driven by curiosity rather than methodological rigor.This concern is particularly acute when sensitive real patient data are used to explore system capabilities.
V. CONCLUSION
The review presents ethical examination of healthcare LLMs as an early starting point for further discussion. It highlights limitations arising from the field’s rapid technical development and the reliance of much source material on preprints.
- Ethical examination of LLMs in healthcare remains nascent and has difficulty keeping pace with rapid technical advancement.
- The review therefore serves as a starting point for further discussions rather than a definitive account of the field.
- A significant portion of the source material came from preprint servers and had not undergone rigorous peer review.
Declaration of Funding
The study was funded by the VolkswagenStiftung through the Digital Medical Ethics Network. The funder had no role in the study, and the authors reported no competing interests.
- The VolkswagenStiftung funded the study through the Digital Medical Ethics Network, grant 9B233.
- The funder played no role in study design, data collection, analysis, interpretation, or manuscript writing.
- All authors declared no financial or non-financial competing interests.