Source-linked AI summary
Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries
Phuong Anh Nguyen, Jill Noorily, Matthew Flathers, Haruka Notsu, Laura Ospina-Pinillos, Tommy Nguyen, Samantha Clark, Aoife Keane, Grace Thompson, John Torous
TL;DR
Conversational AI platforms now curate mental-health citations, but the sources they surface remain poorly characterized, especially beyond English and beyond crisis-focused evaluation. We audited three free products across prompts, questions, and languages, classifying their citations and validating the classifier against human coding. Citations concentrated on a small institutional core, platforms differed in consistency and source-type mix, source requests changed composition only modestly, and non-English queries generally received fewer and less language-appropriate citations.
Problem
Mental-health AI evaluation has concentrated on high-acuity safety tasks and English, leaving routine-query source routing and multilingual citation behavior poorly characterized.
Method
We audited three free consumer AI search products across 20 English questions, source-request prompts, and three translated questions in six additional languages, recording and classifying citations.
Results
The ten most-cited domains accounted for 43.6% of English citations; platforms varied sharply in consistency and source-type mix, while explicitly requesting sources changed composition only modestly.
Takeaways & Limitations
Consumer AI products act as mental-health information gatekeepers, and product choice shapes the evidence users are shown; multilingual routing remains constrained by available authoritative in-language material.
Takeaways & Limitations
This June/July 2026 snapshot covered free, logged-out tiers, and non-English estimates were based on only three questions per language.
Abstract
from arXiv · showhide
Online health information seeking is shifting from keyword search, where users consider a ranked list of links, to conversational systems that compose a single answer and curate its citations. Source evaluation therefore passes from user to platform, yet what these systems surface is poorly characterized. We audited three free consumer products (ChatGPT, Perplexity, Google AI Overview) on twenty English mental health questions under two prompt conditions, with a subset of three also translated into six further languages of varying resource tiers. We recorded 15,942 citations across 1,140 responses and 1,713 unique domains, then classified every citation with a nine-category organizational typology applied by a deterministic classifier validated against human coding. Citations were heavily concentrated: the ten most-cited domains accounted for 43.6% of English citations, and government, commercial health, and academic sources were closely matched at roughly 22% each. Platforms differed little in typical citation volume but sharply in consistency and in the source types they favored. Explicitly requesting sources shifted composition only modestly. Non-English queries surfaced fewer citations and were routed to language-appropriate resources at significantly lower rates. We release the typology, classifier, and annotated corpus as reusable instruments for auditing generative health search.
1 Introduction
Conversational AI systems increasingly compose health answers and curate their citations, shifting source-triage work from users to platforms. This study addresses limited characterization of what these systems surface across routine mental-health queries, prompting styles, and languages.
- Conversational systems compose answers and curate external citations, taking on an active gatekeeping role in mental-health information seeking.
- Online mental-health information seeking is increasingly important as distress and demand for accessible, privacy-sensitive guidance rise.
- Prior AI mental-health evaluation has focused mainly on high-acuity crises, suicide risk, and self-harm guardrails rather than routine psychoeducation.
- Evaluation is disproportionately anglocentric, despite lower native-language content and weaker cross-lingual retrieval creating dual scarcity for non-English health queries.
- The audit examines source characteristics, citation routing, prompting for sources, and multilingual behavior while releasing a typology, classifier, and annotated corpus.
2 Methods
The study audits three free consumer AI search products across platforms, prompts, questions, and languages, recording surfaced sources and classifying their organizational types. Repeated queries, URL standardization, and human validation support systematic comparison.
- The audit covered ChatGPT, Perplexity, and Google AI Overview across three factors: platform, language, prompt condition, and question.Each condition was administered five times by seven trained annotators, with source counts and identities as primary outcomes.
- Twenty English mental-health questions represented eight functional categories, while three questions were translated into six non-English languages.The translated set covered symptom or diagnosis, acute crisis, and trustworthy-source queries across three resource tiers.
- Two prompt conditions compared verbatim questions with questions appended by “List your Sources.”
- Annotators captured every presented citation channel, then standardized destination URLs to registrable domains for analysis.
- A nine-category ordered, rule-based classifier assigned citations using domain, host, and public-suffix information.Categories included government, academic, nonprofit health system, commercial health, nonprofit advocacy, encyclopedia, and social or video sources.
- Classifier assignments were validated through full review of the 100 most-cited domains and independent double-coding of 200 additional domains.
3 Results
The audit found concentrated and platform-specific citation behavior in English, modest effects from explicitly requesting sources, and lower or less localized citation behavior for many non-English queries.
- 15,942 citations across 1,140 responses covered 1,713 unique domains, while 7.3% of responses contained no citations.No-citation responses were most common for ChatGPT (14.5%) and least common for Perplexity (0.3%).
- 32.29 was ChatGPT’s mean English citation count versus 12.85 for Google AI Overview and 9.7 for Perplexity, although medians clustered at 10-12.Outlier responses, almost all from ChatGPT, drove the mean differences.
- 44.2 was the mean citation count for medication questions versus 9.7 for crisis-resource questions.The ten largest responses were all ChatGPT answers to medication or treatment-guideline questions.
- 43.6% of 10,959 English citations came from the ten most-cited sources, with government and commercial health each at 22.6% and academic or journals at 21.6%.
- Each platform led with a different source type: government or public for ChatGPT (26.4%), academic or journal for Perplexity (24.3%), and commercial health for Google AI Overview (25.7%).ChatGPT cited Wikipedia more, while Google AI Overview cited social or video sources more.
- Treatment-effectiveness questions drew 56.9% academic or journal citations, crisis questions 41.6% nonprofit or advocacy citations, and medication questions 33.7% commercial-health citations.
- 16.6 to 19.9 was the overall mean citation-count change after adding “List your Sources.”Source composition shifted only modestly, with the largest change being +2.1 percentage points for government or public citations.
- 9.2 versus 12.3 was the mean citation count for non-English versus English queries in the seven-language subset.Citation counts ranged from 11.2 for Spanish to 7.1 for Twi, while language-appropriate routing varied substantially by resource tier and measurement.
4 Discussion
Across the audited products, citations were concentrated, platform-dependent, only modestly changed by source requests, and less locally appropriate for non-English queries. These patterns create risks for users relying on conversational AI as a mental-health information gatekeeper.
- Citation concentration: Citations concentrated on a narrow institutional core, with ten destinations accounting for 43.6% of citations.The concentrated set included academic, government, commercial, and nonprofit health-system sources.
- Platform differences: Platforms differed little in citation volume but sharply in consistency, with similar questions sometimes returning a handful or up to a hundred sources.Perplexity presented a fixed number of source cards, whereas ChatGPT embedded uncapped inline hyperlinks.
- Platform differences: Source-type preferences varied by platform: Google AI Overview leaned toward commercial health and social/video, ChatGPT toward encyclopedias, and Perplexity toward nonprofit and advocacy sources.All three primarily drew on academic, government, commercial, and nonprofit health-system sources.
- Prompt effects: Explicitly requesting sources left citation volume essentially unchanged and shifted source composition by at most 2.1 percentage points toward government and academic sources.The shift was more institutional in direction but modest overall.
- Cross-lingual citations: Non-English citation behavior deteriorated, with English-language citations consistently prioritized and only 11.3% of Spanish citations coming from countries where Spanish is a national language.Localization tracked the supply of authoritative in-language material rather than language resource tier alone.
- Citation quality: Lexical similarity occasionally produced irrelevant citations, including Wikipedia’s mayonnaise page for an anxiety-medication query and DSM-Firmenich for depression diagnosis.These anomalous links appeared alongside legitimate citations and reflected shared terms or acronyms.
- Limitations: The audit’s practical significance is tempered by its single-point timing, free logged-out tiers, limited non-English query sample, unverified link destinations, and USA-based annotators.These boundaries limit generalization across product versions, account types, languages, and localization conditions.
5 Conclusion
The conclusion finds that consumer AI products act as non-neutral, non-uniform gatekeepers of mental-health information, with English citations concentrated in a small institutional core and multilingual localization constrained by available authoritative content.
- Conclusion: Roughly three-quarters of English citations came from academic, government, commercial-medical, and nonprofit health-system sources.Ten destinations alone accounted for 43.6% of citations.
- Conclusion: Each product led with a different source type, so product choice shaped the evidence users were shown.For non-English users, localization tracked the supply of authoritative in-language material.
- Conclusion: The released nine-category typology, deterministic classifier, and annotated corpus support comparable measurement across systems, languages, and time.The corpus contains 15,929 classified citations.
6 Generative AI Disclosure
The disclosure describes generative AI assistance in classifier drafting, manuscript copyediting, and appendix drafting, with all outputs reviewed and finalized by the authors.
- Generative AI Disclosure: Claude Opus 5 aided initial drafting of the rule-based classifier and copyediting of the main manuscript.The authors reviewed, corrected, and finalized the output.
- Generative AI Disclosure: Claude Fable 5 aided drafting of Appendix S1.The appendix output was reviewed, corrected, and finalized by the authors.
- Generative AI Disclosure: No generative AI was used to create or modify the study data.The disclosure distinguishes writing and classifier assistance from data production.
8 Funding Statement
The paper reports no specific grant from public, commercial, or not-for-profit funding agencies.
- Funding Statement: No specific grant supported this research.The statement covers public, commercial, and not-for-profit funding agencies.
- Funding Statement: The funding statement includes public-sector funding agencies.No specific grant was received from that sector.
- Funding Statement: The funding statement includes commercial and not-for-profit funding agencies.No specific grant was received from either sector.
10 Ethics Statement
The study did not involve human participants or identifiable human-subject data, so it was not classified as human-subjects research and did not require institutional review board approval.
- The study involved no human participants or identifiable human-subject data.
- Annotators served as data collectors rather than research participants.
- The study therefore did not require institutional review board approval.
Supporting Information
The paper is titled “Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries.”
- The paper audits citations in AI mental health information queries.
Appendix S1. An API-side source audit
The appendix audits the source surface of consumer AI search products while distinguishing the joint product interface from the underlying model. It also explains why API comparisons are suggestive rather than controlled decompositions of model and wrapper behavior.
- Appendix S1. An API-side source audit: The product surface combines a generative model with a retrieval, selection, and citation wrapper presented as one interface.
- Appendix S1. An API-side source audit: The consumer product is the appropriate gatekeeping-audit object because users act on what the interface displays.
- Appendix S1. An API-side source audit: The API audit approximates model behavior with the consumer interface removed, but product and API citations are exposed differently.
- Appendix S1. An API-side source audit: Comparison requires distinguishing sources formally cited by a system from sources merely named in generated prose.
- Appendix S1. An API-side source audit: The API comparison cannot confirm that its selected model serves the corresponding free product and is therefore suggestive, not a controlled model-wrapper decomposition.
- Appendix S1. An API-side source audit: Retrieval is nondeterministic, so repeated queries are treated as samples rather than fixed outputs.
Methods
The API-side methods queried one developer API per product family using a matched multilingual prompt instrument, repeated stateless requests, and stored complete response records for analysis and reproducibility.
- Models and interfaces: One developer API was queried for each audited product family: OpenAI, Google, and Perplexity.The selected models were gpt-5.4-mini, gemini-3.5-flash, and Sonar, respectively.
- Instrument: The prompt instrument covered twenty English mental health questions and three common questions in six additional languages, each in plain and source-request variants.The instrument spanned medication, diagnostic, symptom, treatment-efficacy, and crisis-resource topics.
- Instrument: The resulting instrument contained 76 prompts and preserved exact correspondence with the primary protocol, including an orthographic error in one Nepali item.
- Procedure: Each prompt was issued independently without conversation history and repeated five times per model to characterize retrieval nondeterminism.Provider-default sampling parameters were used without setting temperature or other hyperparameters.
- Data and reproducibility: Each full API response was stored with generated text, citation or grounding metadata, and request metadata for analysis and reproducibility.The data-collection software was released under the GNU Affero General Public License v3.0.
Results
The audit found substantial differences in citation behavior across API and consumer-product layers, especially for OpenAI, while explicit source requests had stronger effects at the API level. Non-English API responses contained fewer citations but often more language-appropriate or country-specific domains, with Nepali as an exception.
- Citation volume and coverage: 10,718 citations came from APIs versus 15,942 from products across the same 1,140 prompts.APIs averaged 9.4 citations per response, compared with 14.0 for products.
- Citation volume and coverage: 19.8% of API responses had no citation, compared with 7.3% of product responses, largely because gpt-5.4-mini searched in only 50.3% of requests.Responses without search still averaged over 1300 characters.
- Source range and concentration: gpt-5.4-mini cited 31 distinct domains overall, while ChatGPT cited 641 English domains; its ten most-cited destinations represented 97.4% of English citations versus 52.0% for ChatGPT.The other model-product pairs had much closer English domain counts and top-ten shares.
- Source-type composition: Government/public sources rose from 22.6% at the product level to 26.7% at the API level, while commercial health fell from 22.6% to 12.2%.Encyclopedia sources decreased from 3.9% to 0.4%.
- Effect of explicitly requesting sources: Requesting sources increased mean English citations from 8.3 to 13.1 at the API level and from 16.6 to 19.9 at the product level.The government/public share shifted by +11.6 percentage points at the API level versus +2.1 at the product level.
- Multilingual retrieval: Non-English APIs returned fewer citations but higher country-code and language-appropriate shares; Ukrainian and Spanish language-appropriate shares rose by 22.5 and 19.3 percentage points.Nepali diverged: its language-appropriate share fell by 10.0 points because gpt-5.4-mini returned no Nepali-language content and Sonar returned 7.4%.