Source-linked AI summary

Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges

Yisong Chen, Yifan Gao, Sijing Yu, Chuqing Zhao, Yang Lu

arXiv:2608.18080v1cs.AI

TL;DR

Mental-health LLM research shows promise, but evidence remains constrained by limited validation, uncertain generalizability, and ethical risks. This systematic review synthesizes applications and methodological advances, finding broad use from population monitoring to individualized support while underscoring the need for safeguards.

  • Problem

    Evidence on LLM applications in mental health remains limited by small datasets, limited clinical validation, uncertain generalizability, and unresolved ethical concerns.

  • Method

    The review applies PRISMA 2020 to synthesize interdisciplinary LLM studies identified through seven scholarly databases and manual OpenReview checks.

  • Results

    The review finds LLM applications spanning social-media monitoring, clinical conversation, therapy support, psychoeducation, prompt engineering, and multimodal mental-health analysis.

  • Takeaways & Limitations

    Safe real-world deployment requires transparent, clinically validated evaluation frameworks, safeguards, and monitoring to address bias, reliability, and unintended harms.

  • Takeaways & Limitations

    Most studies rely on small, imbalanced, or convenience datasets, limiting generalizability and potentially introducing demographic and clinical biases.

Abstract

from arXiv · show

We present a review on the applications of large language models (LLMs) in health, e.g., social media analysis, clinical conversational agents, therapy support tools, prompt engineering, multimodal learning, and ethical considerations. We integrate findings from interdisciplinary studies utilizing diverse data sources such as social media posts, electronic medical records, and multimodal inputs to enable early detection of depression, suicide risk assessment, personalized therapy support, and psychoeducational content generation. Our review highlights advancements in LLM models and annotation strategies that enhance interpretability and clinical relevance, while we also emphasize the critical role of prompt engineering for domain adaptation. We also discuss emerging multimodal fusion techniques integrating text, speech, and sensor data for improved mental health diagnosis and monitoring. Finally, we address ongoing ethical, sociotechnical, and regulatory challenges, and advocate frameworks to ensure safe, equitable, and accountable deployment of LLMs in real-world mental health care.

1. Introduction

Mental-health conditions impose a substantial global disease burden, while stigma, financial barriers, and workforce shortages limit timely diagnosis and care. This review examines emerging LLM applications alongside methodological advances and persistent clinical, ethical, and sociotechnical challenges.

  • Depression, anxiety, and suicidal behavior contribute substantially to global disease burden, while stigma, financial barriers, and professional shortages constrain timely mental-health diagnosis and care.
  • Twitter and Reddit posts provide large, naturally occurring datasets in which people openly describe their mental-health experiences.
  • The mental-health LLM field remains nascent despite promising advances in detection, severity assessment, and conversational support.
  • Key unresolved challenges include small or imbalanced datasets, limited clinical validation, uncertain generalizability, bias, privacy, and accountability.
  • The review analyzes recent LLM applications in mental health and examines methodological innovations and ongoing challenges shaping their development.

2. Literature Review

The literature progresses from social-media-based psychological state modeling through multimodal deep learning and Transformers to LLM applications in data-scarce and high-stakes mental health tasks. Persistent gaps include limited generalizability, interpretability, reliability, multimodal integration, ethical safeguards, and real-world clinical evaluation, motivating priorities in clinical utility, privacy, and responsible deployment.

  • Foundations: Social media text contains linguistic and behavioral signals of depression, anxiety, and related conditions that support computational early detection and intervention.Early studies also revealed trade-offs between traditional machine-learning and newer deep-learning approaches.
  • Neural and Multimodal Methods: Multimodal deep-learning systems combined textual and visual features, while hierarchical networks and attention mechanisms improved depression classification accuracy and interpretability.These methods identified linguistically salient features in depression-related data.
  • Transformer Architectures: Transformers captured long-range dependencies and processed sequences in parallel, advancing speed, accuracy, and domain-specific mental health detection after pre-training and fine-tuning.Their self-attention mechanisms marked a turning point from recurrent architectures.
  • LLM Applications: LLMs extend mental health analysis with zero-shot and few-shot learning, synthetic training examples, medical-knowledge integration, and applications in data-scarce contexts.They are defined as Transformer-based architectures with billions of parameters.
  • High-Stakes Tasks: LLMs have been applied to high-stakes suicidality detection, including zero-shot identification of suicidal ideation and evidence extraction for pre-annotated suicide-risk labels.Broader evaluations also benchmarked GPT-3.5 and GPT-4 against domain-specific models for health-related text classification.
  • Gaps and Priorities: Research remains constrained by small, imbalanced datasets, limited generalizability, interpretability and reliability challenges, incomplete multimodal integration, ethical risks, and limited real-world clinical evaluation.The review prioritizes clinical utility, privacy, and responsible deployment, especially for depression and suicidality detection.

3. Research Method

The review used a PRISMA 2020-guided, systematic search and staged screening process to identify high-quality evidence on LLMs and Transformer-based architectures in mental health. It standardized data extraction and assessed methodological risk of bias, transparency, and reproducibility across included studies.

  • Search strategy: The review followed PRISMA 2020 and searched seven major scholarly databases, supplemented by manual checks on OpenReview.The databases were IEEE Xplore, ACM Digital Library, PubMed, Scopus, ScienceDirect, SpringerLink, and Google Scholar.
  • Study selection: A second screening round emphasizing methodological transparency, reproducibility, and clinical relevance produced a final evidence base of 92 high-quality studies.The PRISMA workflow for search, screening, and selection is illustrated in Figure 1.
  • Data extraction and quality assessment: A structured extraction form captured model architecture, input modality, mental-health domain, data sources, and evaluation metrics across studies.The review also assessed risk of bias using dataset representativeness, prompt-engineering or fine-tuning transparency, and reproducibility of reported results.
  • Data extraction and quality assessment: Risk-of-bias flags were assigned to studies with unclear methodology, unexplained proprietary data, or inconsistent performance metrics.The review also notes diverse data sources for conversational agents, including electronic medical records and counseling notes.

4. Applications

This section reviews three main applications of generative AI and LLMs in mental health: social-media analysis, clinical conversational agents, and therapy or decision-support tools. These applications support large-scale disorder detection and monitoring, interactive patient and clinician support, and routine clinical tasks.

  • Social-media analysis: LLMs support social-media analysis for large-scale detection and monitoring of mental-health disorders.The section identifies social-media analysis as one of three main application areas.
  • Clinical conversational agents: Clinical conversational agents provide interactive support for patients and clinicians.The section describes conversational agents as an application addressing mental-health needs through interaction.
  • Therapy or decision-support tools: Therapy and decision-support tools assist with routine clinical tasks.The section identifies therapy or decision-support tools as a third application area.

4.1 Social Media Analysis for Depression and Suicidal Ideation

Social media provides large-scale, real-world accounts of psychological distress, mental-health coping, and help-seeking, supporting analysis of depression and suicidal ideation. However, informal language, disclosure differences, demographic and cultural biases, and ethical concerns complicate interpretation and generalizability, while evolving annotation and modeling pipelines increasingly incorporate LLMs.

  • Data sources and value: Social media captures real-time, unfiltered indicators of psychological distress, including hopelessness, withdrawal, and suicidal thoughts, including among underserved or hesitant populations.
  • Data sources and value: Online mental-health communities also reveal peer support, informal help-seeking, and collective coping dynamics that complement clinical datasets.
  • Challenges: Informal language, variable self-disclosure, demographic and platform biases, and privacy and consent concerns limit interpretation and model generalizability.
  • Annotation strategies: Annotation approaches span expert clinical labeling, user self-disclosure, and LLM-assisted zero-shot detection, explanation extraction, and synthetic data generation.Expert annotation is clinically grounded but resource-intensive; self-disclosure scales more readily but can reflect uneven openness across user groups.
  • Modeling approaches: Modeling has progressed from interpretable traditional classifiers to deep learning, transformer architectures, semi-supervised expansion, and LLM-assisted pipelines supporting annotation, data augmentation, interpretation, and real-time analysis.This progression accommodates fine-grained tasks such as emotional-severity estimation and symptom-expression detection while addressing limited labeled data.

4.2 LLM-Powered Conversational Agents Across Medical Practice Stages

LLM-powered conversational agents support multiple stages of mental health care, including triage, symptom checking, therapy-session summarization, patient engagement, and treatment planning. They can provide personalized, accessible assistance while remaining tools for professionals rather than substitutes for clinical judgment.

  • Triage: LLM-powered conversational agents can improve patient triage by supporting confidence and judgment during decisions and processing large volumes of unstructured clinical information.Their triage role is part of broader clinical conversational-agent support across medical practice stages.
  • Symptom checking: Agents improve symptom checking through adaptable emotional support and personalized, empathetic responses tailored to users’ conditions.These capabilities extend traditional symptom checking in clinical and mental health contexts.
  • Patient engagement: LLM-powered conversational agents support patient engagement through accessible, customized, and interactive conversations, including psychoeducational content.Compared with conventional digital mental health interventions, these agents are reported to improve engagement by generating understanding and interactive conversations.
  • Session summarization: Fine-tuned LLMs summarize therapy sessions by identifying symptom-related dialogue, assigning clinical categories, and producing coherent, clinically relevant summaries for professionals.The summaries highlight content that can help health professionals support patients more efficiently.
  • Ethical oversight: Ethical guidance emphasizes that LLMs should assist mental health professionals, especially in low-resource settings, rather than substitute for professional healthcare providers.Ethical oversight is presented as part of the therapy-support workflow.
  • Decision aids and treatment planning: LLMs can assist treatment planning by extracting information from text and audio, deducing hidden influences on mental health, and providing customized information for professional decision-making.They support treatment refinement and planning while not taking over clinicians’ important judgment.

4.3 LLM Prompt Engineering Techniques for Mental Health Applications

Prompt engineering is presented as a lightweight, flexible approach for adapting general-purpose LLMs to domain-specific mental health tasks. The section describes prompting paradigms and applications including symptom extraction, therapeutic dialogue, chatbots, and mental health condition classification.

  • 4.3 LLM Prompt Engineering Techniques for Mental Health Applications: Prompt engineering systematically designs and adjusts input texts to elicit accurate, relevant LLM responses for mental health applications.It is described as an emerging technique for tailoring general-purpose LLMs to the mental health domain.
  • 4.3.1 Overview of Prompt Engineering in Mental Health: Compared with extensive fine-tuning, prompt engineering requires fewer labeled data and computational resources while remaining lightweight and flexible.The passage contrasts prompt engineering with fine-tuning for domain-specific mental health instruction.
  • 4.3.1 Overview of Prompt Engineering in Mental Health: Well-designed prompt templates can guide structured depressive symptom extraction and support psychotherapy simulation and psychoeducational content delivery.These examples illustrate how prompts can direct LLMs toward structured mental health tasks.
  • 4.3.2 Methods and Paradigms of Prompt Engineering: Prompt engineering supports automated symptom classification, simulated therapeutic dialogue, and question-answer chatbots in digital mental health.It enables efficient adaptation of foundation models without costly, time-consuming retraining.
  • 4.3.2 Methods and Paradigms of Prompt Engineering: Classical prompting paradigms include zero-shot, one-shot, and few-shot prompting, with curated examples helping ground one-shot and few-shot reasoning.Zero-shot prompting provides only the task description, whereas one-shot and few-shot prompting provide curated examples.
  • 4.3.3 Application Scenarios in Mental Health: LLMs can analyze unstructured sources, including social media posts and electronic medical records, for early detection and classification of depression and flourishing.The passage identifies depression diagnosis and flourishing classification as application scenarios.
  • 4.3.3 Application Scenarios in Mental Health: The section concludes that continued development is needed to adapt LLMs effectively for mental health despite ongoing challenges.It frames adaptation as an area for further exploration.

5. Innovations and Challenges

The section presents multimodal learning and LLMs as advances enabling richer mental-health representations, scalable interaction, monitoring, and moderation. It also emphasizes ethical and sociotechnical risks requiring accountable deployment frameworks.

  • Multimodal learning: Multimodal frameworks combine text, speech, physiological signals, and behavioral patterns to represent psychological states more comprehensively than single-modality approaches.They support scalable interaction, monitoring, and moderation in online platforms.
  • Multimodal learning: Multimodal fusion integrates diverse data sources for complex mental-health tasks, including more accurate diagnosis, symptom prediction, and personalized interventions.Fusion strategies include feature-level, decision-level, hybrid, and deep learning–based approaches.
  • Ethical and sociotechnical challenges: LLM-enabled mental-health platforms risk emotionally persuasive manipulation, inconsistent moderation, and misclassification of sensitive expressions as rule violations.These risks arise in social-media influence operations and autonomous moderation, especially where users may be vulnerable.
  • Ethical and sociotechnical challenges: LLMs interpreting mental-health disclosures may compromise privacy through social-network inference, behavioral profiling, and recontextualization of self-disclosed content.Mental-health communities encourage sharing personal struggles, emotions, and behavioral patterns, increasing the sensitivity of such data.
  • Ethical and sociotechnical challenges: Ethical frameworks identify illusion of understanding, false empathy, and unclear accountability for AI-generated advice as challenges to assigning therapeutic responsibility.Responsible deployment also requires attention to misinformation safeguards, community governance, disclosure dynamics, and long-term psychological implications.

6. Discussion

LLMs are reshaping mental health research and practice across population monitoring, individualized support, prompt engineering, and multimodal learning. Their wider deployment remains constrained by dataset limitations, reliability and interpretability concerns, ethical and regulatory challenges, and the need for representative data, adaptable tools, monitoring, and human oversight.

  • Applications: LLMs span social media analysis, clinical conversational agents, therapy support tools, prompt engineering, and multimodal learning across mental health applications.These applications extend support from population-level monitoring to individualized care.
  • Challenges and gaps: Small, imbalanced, and convenience datasets, poor annotation quality, and underrepresented demographics limit generalizability, robust evaluation, and equity in mental health care.These limitations are particularly associated with datasets drawn from social media platforms.
  • Ethical and regulatory considerations: Fairness, transparency, algorithmic bias, accountability, data privacy, and regulation require human oversight because LLMs complement rather than replace clinicians.Regulatory challenges are especially relevant to direct-to-consumer applications and wider clinical deployment.
  • Future work: Future progress depends on richer, more diverse datasets, more powerful and interpretable models, resilient and adaptable tools, and ongoing monitoring to address emerging issues.Datasets should represent different populations and real-world settings, while tools must adapt as environments evolve.

7. Conclusion

LLMs could transform mental health care through early detection, clinical conversation, therapy, psychoeducation, and multimodal, personalized interventions. Safe real-world deployment nevertheless requires addressing methodological, technical, social, ethical, and regulatory challenges, with human oversight.

  • Applications and contributions: LLMs support early detection of depression and suicidal ideation, clinical conversation, therapy, and psychoeducation across digital and traditional care settings.Applications draw on social media posts, electronic medical records, speech, and sensor data to enable more nuanced, personalized, and scalable interventions.
  • Methodological and technical challenges: Small, imbalanced, or convenience samples can introduce bias and limit generalizability, while reliability, explainability, and hallucinations remain unresolved concerns.These limitations are especially consequential across diverse populations and underrepresented subgroups.
  • Ethical and regulatory challenges: Fairness, privacy, transparency, and accountability should precede broad clinical use, with LLMs serving as supportive aids rather than replacements for human professionals.Ongoing expert oversight is essential for confidential and critical mental health interactions.
Loading 2608.18080v1…