Source-linked AI summary

Performance of a domain-specific large language model in answering patient questions in psychiatry

Alexander J. Hish, Arjun Nagendran, Scott N. Compton

arXiv:2608.22797v1cs.AI

TL;DR

This study asks whether a domain-specific LLM trained on patient education resources can answer psychiatric medication questions more safely and accurately than general-purpose chatbots. The authors developed MIND, compared it with ChatGPT and OpenEvidence on escitalopram questions, and evaluated responses using an automated rubric and psychiatrist ratings. MIND ranked highest across automated rubric domains, but psychiatrists preferred ChatGPT despite rating MIND more complete and similarly safe.

  • Problem

    The study addresses whether a domain-specific LLM can provide accurate, patient-centered answers to psychiatric medication questions while avoiding common chatbot inaccuracies and safety issues.

  • Method

    The authors developed MIND from authoritative patient education resources and compared its escitalopram responses with ChatGPT and OpenEvidence using automated rubric analysis and ratings from 10 psychiatrists.

  • Results

    MIND was rated highest across all automated rubric domains, while psychiatrist ratings favored MIND for completeness, found similar safety, and favored ChatGPT overall.

  • Takeaways & Limitations

    MIND generally provided accurate, complete, and safe escitalopram information, supporting domain-specific LLMs as a step toward safer psychiatric patient education.

  • Takeaways & Limitations

    MIND generally could not answer questions involving drug interactions because those questions were not covered by its training corpus.

Abstract

from arXiv · show

Background This study was designed to evaluate whether a domain-specific large language model (LLM) trained exclusively on patient education resources can answer questions about psychiatric medications, in a manner superior to LLM chatbots. We developed an LLM ("MIND") fine-tuned for clinical fidelity, trained on patient education resources from authoritative medical organizations. Methods We compared the responses of MIND, ChatGPT, and OpenEvidence to patient questions about escitalopram, using two methods: (1) computer analysis according to a rubric measuring accuracy, clarity, completeness, nuance, safety, and referral appropriateness; (2) ratings from N=10 board-licensed psychiatrists on similar metrics. Results When rated by rubric, MIND was rated highest in all domains (p<0.001). When rated by psychiatrists, ChatGPT was rated accurate more often than MIND with a negligible effect size (p=0.021, r=0.073); MIND was rated complete more often than ChatGPT with a small effect size (p<0.001, r=0.160); and MIND and ChatGPT were rated safe with the same frequency (p=0.955, r=0.002). The majority of psychiatrists preferred the responses generated by ChatGPT (57.6%) compared to MIND (42.4%, p=0.003). Conclusions MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time. However, despite MIND's ability to provide more complete responses, psychiatrists preferred ChatGPT's responses. MIND represents a step towards building safe LLM systems to enhance patient education in psychiatry.

1. Introduction

The study addresses whether a domain-specific LLM trained on patient education resources can provide accurate, patient-centered psychiatric medication information while avoiding chatbot inaccuracies and safety issues. This question is particularly relevant because psychiatric medication information is inconsistent, clinical encounters are time-limited, and patients may seek misleading online information.

  • No prior studies had examined whether a domain-specific LLM could answer psychiatric medication questions for patients.
  • Public information about psychiatric medications can be inconsistent or inaccurate, while general-purpose chatbots may omit safety considerations or introduce hallucinations.
  • Psychiatrists may have only 15–20 minutes for follow-up visits, with rising caseloads and workforce shortages limiting opportunities for thorough medication education.
  • Limited medication understanding, inadequate question time, and weak therapeutic alliance contribute to adherence problems and adverse outcomes.
  • A reliable, on-demand educational LLM could address common medication questions between visits and support evidence-based patient education.
  • The study evaluates whether an LLM trained exclusively on patient education resources can answer common psychiatric medication questions while avoiding chatbot inaccuracies and safety issues.Escitalopram was selected as the case study.

2. Methods

The study developed MIND using curated psychiatric medication education resources and compared it with ChatGPT and OpenEvidence on patient questions about escitalopram. Responses were assessed by an automated rubric and by board-certified psychiatrists using defined rating constructs.

  • MIND was fine-tuned for clinical alignment and fidelity using patient education resources curated from authoritative psychiatric and medical organizations.
  • MIND used retrieval-augmented generation by matching an embedded patient query to provider-approved references and conditioning generation on retrieved passages.
  • The escitalopram corpus included FAQ datasets, peer-reviewed medication guides, and standard clinical reference materials from seven authoritative sources.
  • Each model answered 50 English questions covering mechanisms, indications, dosing, side effects, interactions, special cases, and occasional patient vignettes.
  • Automated rubric: Automated evaluation rated accuracy, clarity, completeness, nuance, safety, and referral appropriateness on 1–5 scales.
  • Psychiatrist ratings: Ten board-licensed psychiatrists rated MIND and ChatGPT for accuracy, completeness, and safety on a 1–3 scale and selected preferred responses.

3. Results

MIND received the highest automated rubric ratings across all evaluated domains, but psychiatrist ratings showed a mixed comparison: ChatGPT was slightly more often accurate, MIND more often complete, and preferences favored ChatGPT.

  • Automated rubric: MIND was rated highest in all automated rubric domains: accuracy, clarity, completeness, nuance, safety, and referral appropriateness (all p<0.001).Effect sizes ranged from ε²=0.170 for referral appropriateness to ε²=0.610 for accuracy.
  • Psychiatrist ratings: 78.9% of psychiatrist ratings judged ChatGPT accurate versus 71.6% for MIND, a significant difference with negligible effect size (r=0.073).
  • Psychiatrist ratings: 70.3% of ratings judged MIND complete versus 56.3% for ChatGPT, a significant difference with small effect size (r=0.160).
  • Psychiatrist ratings: MIND and ChatGPT were rated safe with similar frequency: 83.9% for MIND and 83.4% for ChatGPT.
  • Preferences: 57.6% of psychiatrists preferred ChatGPT responses compared with 42.4% for MIND (p=0.003).Early-career psychiatrists showed a stronger ChatGPT preference, 71.7% versus 28.3% for MIND.

4. Discussion

MIND, a domain-specific LLM trained on patient-education resources, performed strongly on escitalopram questions, but psychiatrists generally preferred ChatGPT’s responses. MIND’s greater completeness coincided with longer responses and a possible accuracy–completeness tradeoff.

  • Study contribution: MIND was designed as a domain-specific LLM trained exclusively on patient-education resources.The study evaluated whether this approach could answer psychiatric medication questions with clinical fidelity.
  • Automated evaluation: MIND outperformed ChatGPT and OpenEvidence across accuracy, clarity, completeness, nuance, safety, and referral appropriateness in automated rubric ratings.These domains were directly relevant to patient education.
  • Psychiatrist evaluation: Psychiatrists found MIND more complete than ChatGPT, with no meaningful differences in accuracy or safety.The psychiatrist evaluation nevertheless differed from the automated rubric results.
  • Response characteristics: ChatGPT’s shorter responses were judged more accurate more often, whereas MIND’s longer responses were more complete, suggesting a possible accuracy–completeness tradeoff.The authors present this as a possible explanation rather than a demonstrated causal mechanism.

Supplement 1: Patient Education Resources

The patient-education corpus included resources from authoritative psychiatric, pediatric, government, and medical-information organizations. The listed materials covered anxiety, depression, and escitalopram treatment information.

  • AACAP resources included parent medication guides for anxiety disorders and depression.
  • The American Academy of Pediatrics and American Psychiatric Association supplied depression- and anxiety-related patient education materials.
  • Escitalopram resources came from the FDA, NAMI, and MedlinePlus.
  • UpToDate contributed patient education on medicines for depression and treatment options for adults, children, and adolescents.

Responses

The responses explain escitalopram’s serotonin-related mechanism, uses, and safety considerations, while addressing effects on the brain, personality, and child development. They also include treatment indications and monitoring considerations, although one response states that relevant information could not be found.

  • Mechanism: Escitalopram is described as an SSRI that blocks serotonin reuptake, leaving more serotonin available for signaling and potentially improving mood and anxiety.Some responses also describe its selectivity and minimal effects on other neurotransmitters or receptors.
  • Safety: The responses mention possible side effects and risks, including sleep changes, nausea, serotonin syndrome, bleeding, headaches, and fatigue, with use under medical supervision.One response states that escitalopram is not approved for headache treatment and that headache occurred more often than with placebo in clinical trials.
  • Effects on personality: The responses state that escitalopram does not fundamentally change personality and may instead produce modest trait shifts consistent with improved mental health.Other responses describe people feeling calmer, more balanced, or more like themselves as symptoms improve.
  • Child development: For children and adolescents, responses emphasize monitoring mood, appetite, weight, and growth, while noting that approved age ranges differ by indication.They distinguish observed appetite and weight effects from developmental findings in animal studies at very high doses.
  • Treatment indications: Escitalopram is presented as treating major depressive disorder and generalized anxiety disorder, with some responses also mentioning off-label anxiety-related uses.Responses describe improvement in depressive symptoms and, in clinical trials, greater improvement and longer time to relapse than placebo.
  • Safety: Responses warn that escitalopram can increase suicidal thoughts or behaviors in young people, particularly early in treatment or after dose changes, requiring close monitoring.One response reports 14 additional cases per 1,000 treated patients under 18 and 5 additional cases per 1,000 among those aged 18–24 compared with placebo.
Loading 2608.22797v1…