Source-linked AI summary
The opportunities and risks of large language models in mental health
Hannah R. Lawrence, Renee A. Schneider, Susan B. Rubin, Maja J. Mataric, Daniel J. McDuff, Megan Jones Bell
TL;DR
Rising mental health needs and inadequate care access motivate interest in large-scale LLM applications. This paper reviews LLMs for mental health education, assessment, and intervention, then examines risks and mitigation strategies. The literature shows promising capabilities but important safety, reliability, equity, and oversight requirements.
Problem
Mental health concerns are rising while access to care remains insufficient, creating a need for large-scale solutions.
Method
The paper summarizes research applying LLMs to mental health education, assessment, and intervention, and identifies associated opportunities, risks, and mitigation strategies.
Results
The reviewed literature reports promising applications across education, assessment, and intervention, including Med-PaLM 2 diagnostic accuracy ranging from 77.5 percent for diagnoses to 92.5 percent for diagnostic categories.
Takeaways & Limitations
Mental-health LLMs should be fine-tuned, tested for evidence-based practice, designed for equity and safety, and developed with people who have lived experience and domain expertise.
Takeaways & Limitations
LLM outputs remain unreliable and may diverge from clinicians or provide unsafe intervention advice, so human oversight is still needed.
Abstract
from arXiv · showhide
Global rates of mental health concerns are rising, and there is increasing realization that existing models of mental health care will not adequately expand to meet the demand. With the emergence of large language models (LLMs) has come great optimism regarding their promise to create novel, large-scale solutions to support mental health. Despite their nascence, LLMs have already been applied to mental health related tasks. In this paper, we summarize the extant literature on efforts to use LLMs to provide mental health education, assessment, and intervention and highlight key opportunities for positive impact in each area. We then highlight risks associated with LLMs' application to mental health and encourage the adoption of strategies to mitigate these risks. The urgent need for mental health support must be balanced with responsible development, testing, and deployment of mental health LLMs. It is especially critical to ensure that mental health LLMs are fine-tuned for mental health, enhance mental health equity, and adhere to ethical standards and that people, including those with lived experience with mental health concerns, are involved in all stages from development through deployment. Prioritizing these efforts will minimize potential harms to mental health and maximize the likelihood that LLMs will positively impact mental health globally.
1 Introduction
Mental health concerns are widespread and rising, while access to care remains insufficient. The paper presents LLMs as potential large-scale tools for mental health education, assessment, and intervention, and reviews opportunities, risks, and mitigation strategies.
- Half of all individuals may experience a mental health disorder during their lifetimes, and 1 in 8 people currently experience a mental health concern.
- Mental health concerns are increasing, but care access has not expanded enough to meet demand.
- In the United States, the average interval between symptom onset and treatment is 11 years.
- LLMs can process and generate language, organize complex concepts, and be fine-tuned for mental health applications.
- The paper synthesizes existing mental-health LLM research, identifies opportunities and risks, and recommends responsible development, testing, deployment, and risk mitigation.
2 Applications of LLMs to Mental Health
LLMs have been applied across mental health education, assessment, and intervention, with promising but uneven results. Evidence supports useful capabilities, while limitations in human comparability, reliability, personalization, and safety require caution and oversight.
- Education: LLMs can generate immediate mental health education, while domain-specific and general-purpose systems show varying levels of helpfulness, accuracy, and evidence-based alignment.Psy-LLM received moderate human ratings for helpfulness, fluency, relevance, and logic; other evaluations found ChatGPT responses near ceiling for clinical accuracy but also reported shortcomings relative to human answers.
- Education: Using ChatGPT to support behavioral-health training reduced provider content-development time by 37.5 percent.
- Assessment: Mental-health-specific models generally detected depression and suicidal ideation from social-media posts better than models pretrained in clinical-note or biomedical domains.
- Assessment: 77.5 percent of DSM-5 case studies received the correct diagnosis from Med-PaLM 2, increasing to 92.5 percent for the correct diagnostic category.
- Assessment: LLM assessments can diverge from clinicians: ChatGPT underestimated suicide risk, while Med-PaLM 2 showed 0.98 specificity, 0.30 sensitivity, and 20 percent comorbidity or modifier accuracy.
- Intervention: Chatbots trained in empirically supported treatments may reduce depressive and anxiety symptoms and stress while expressing empathy and sustaining therapeutic conversations.
- Intervention: Intervention chatbots may fail to personalize care, forget prior information, provide harmful advice, and respond inadequately to suicide risk.
3 Risks Associated With Mental Health LLMs
The paper argues that mental health LLMs require ethical, responsible development, testing, and deployment because they may reproduce inequities, produce unsafe or unreliable outputs, and operate without adequate transparency or human safeguards. Proposed safeguards include domain-specific fine-tuning and evaluation, clear competence limits, informed consent and confidentiality, explainability, and human involvement throughout the lifecycle.
- 3.1 Overview: Responsible development, testing, and deployment require identifying risks, mitigating them preemptively, and monitoring for new or unexpected harms.The paper notes that risks may differ across education, assessment, and intervention.
- 3.2 Perpetuating Inequalities, Disparities, and Stigma: Training LLMs on unrepresentative mental health data can perpetuate stigma, bias, and disparate performance across groups.The paper recommends representative data, mental-health-specific fine-tuning, and evaluation for toxic or discriminatory language.
- 3.8 Scaling LLMs: Scaling LLMs could expand mental health information, assessment, and treatment access, especially where providers are scarce or costs are substantial.The paper also identifies personalization and workforce training as possible routes to broader access, while noting implementation challenges.
- 3.3 Unethical Practices: LLMs should communicate competence limits, withhold outputs for tasks they cannot reliably perform, and undergo ongoing competence assessment.Users should also be educated about when LLM use is appropriate.
- 3.4 Insufficient Reliability: Repeated prompting can yield different responses, risking erosion of trust, misdiagnosis, and treatments poorly suited to an individual's concern.The paper uses repeated assessment of depressive symptoms as an example requiring a stable diagnostic conclusion despite varied phrasing.
- 3.5 Inaccuracy: Inaccurate or outdated training data can produce inaccurate mental health information, including biased representations or iatrogenic treatment options.Accuracy has multiple dimensions beyond factual correctness, so evaluation must remain ongoing during prompt fine-tuning.
- 3.6 Lack of Transparency and Explainability: Mental health LLMs should make their use, development, testing, data sources, and decision rationales apparent, while recognizing that generated explanations may be internally inconsistent.Explainability can help communicate assessment results or intervention justifications to patients.
- 3.7 Neglecting to Involve Humans: Anonymous mental health services carry added risk because unpredictable LLM content may become harmful or nontherapeutic, creating unresolved safety and liability concerns.The paper calls for legal and regulatory frameworks addressing safety and responsibility.
4 Conclusions
LLMs show promise for expanding mental health information and care, especially in education and assessment, but intervention requires greater caution and further research. Responsible development requires mental-health-specific fine-tuning, equity, safety, evidence-based practice, confidentiality, and stakeholder alignment.
- LLMs may expand access to mental health information and care, with education and assessment especially aligned with their strengths.The paper describes intervention as requiring greater caution than education and assessment.
- Further research should test treatment delivery and provider training, adaptation for youth and marginalized populations, rapport building, and high-acuity risk detection.
- Mental health LLMs should be fine-tuned for the domain and prioritize equity, safety, evidence-based practice, and confidentiality.The paper also emphasizes testing adherence to evidence-based practice and alignment with people with lived experience and mental health experts.
- Responsible development, testing, and deployment could expand access to evidence-based mental health information and services and improve mental health globally.The passage presents this as potential rather than an established outcome.
6 Conflicts of Interest
The authors disclose employment, compensation, equity, and shareholder relationships involving Google, Magnit, and several health-related companies.
- Several authors are Google employees who receive monetary compensation and equity in Alphabet.
- Two authors are Magnit employees receiving compensation and contracted for work at Google.
- Additional disclosures include authors’ shareholdings and consulting relationships with Meeno Technologies, The Orange Dot, Lyra Health, Trek Health, and Understood.