Source-linked AI summary

Who Trusts AI with Their Emotions? Trust Formation and Sociodemographic Variation in LLM Use for Emotional Support

Natalia Amat-Lefort, Mert Yazan, Amanda Cercas Curry, Flor Miriam Plaza-del-Arco

arXiv:2608.21220v1cs.HC

TL;DR

Research on emotional-support AI lacks validated measures and large-scale evidence about how trust and adoption vary across user segments. Using 1,343 users from seven countries, the paper develops a seven-construct scale, models trust and benefits as pathways to use, and finds that system attributes and adoption logic differ across demographic groups. The findings support user-sensitive approaches to understanding and designing emotional-support AI, within the study’s Western sample and self-report boundaries.

  • Problem

    Research lacks validated psychometric instruments for affective LLM perceptions and large-scale, sociodemographically diverse evidence on trust and adoption across user segments.

  • Method

    The study develops and validates a seven-construct scale, tests an SEM linking system attributes to Trust and Perceived Benefits as mediators of Actual System Use, and conducts MGA across five demographic dimensions.

  • Results

    Privacy, Personalization, and Humanlikeness drive Trust, Perceived Bias degrades it, and trust formation and adoption logic vary across demographic dimensions.

  • Takeaways & Limitations

    Trust in emotional-support AI is not universal but depends on users’ sociodemographic characteristics and the values they bring to system interaction.

  • Takeaways & Limitations

    The seven-country sample covers only WEIRD contexts, leaving major user populations in East Asia, Sub-Saharan Africa, Latin America, and South Asia absent.

Abstract

from arXiv · show

Trust in AI for emotional support is not universal; it is shaped by who users are, where they come from, and what they value. Yet research in this area lacks validated psychometric instruments for assessing user perceptions in affective AI contexts and large-scale evidence on how trust formation varies across user segments. To address these gaps, we develop and validate a seven-construct psychometric scale, test a Structural Equation Model (SEM) linking system attributes to Trust and Perceived Benefits as mediators of Actual System Use, and conduct a Multi-Group Analysis (MGA) across five sociodemographic dimensions (gender, age, education, socioeconomic status, cross-national region), drawing on 1,343 active users from seven countries. We find that users experience empathy and anthropomorphism as a unified "Humanlikeness" construct, and that Privacy, Personalization, and Humanlikeness drive Trust while Perceived Bias degrades it. Notably, adoption logic diverges across groups: Privacy shapes women's trust more than men's, Anglosphere (UK, USA) users respond more positively to Humanlikeness than Europeans, and educated and higher-income users require Trust to engage, whereas older adults and lower socioeconomic groups bypass it entirely, relying on perceived practical benefits (e.g., 24/7 availability, non-judgmental support). Our findings extend technology acceptance theory and inform the equitable design of emotional support AI.

1. Introduction

Emotional-support LLM use raises measurement and equity gaps: existing tools inadequately capture social-affective and inclusion concerns, while evidence on demographic variation remains limited. This study develops a validated instrument and examines trust and adoption across user groups.

  • Motivation: Emotional-support LLMs are increasingly used as fluent, anthropomorphic, non-judgmental, always-available spaces for processing feelings and managing stress.Their use extends beyond productivity into sensitive interpersonal and wellbeing contexts.
  • Research gaps: More than half of conversational-agent survey instruments lacked cited validity evidence, and existing measures often omit empathy, warmth, bias, and fairness.These omissions limit assessment of user perceptions in affective human–AI dialogue.
  • Research gaps: Existing studies rarely provide large-scale, diverse evidence connecting demographics with trust, perception, and actual emotional-support adoption.Small, homogeneous samples leave joint effects of gender, age, education, socioeconomic status, and culture insufficiently understood.
  • Objectives: The study develops and validates a multidimensional psychometric scale covering social-affective, functional, and bias-related perceptions.Its objectives include a theoretically grounded instrument and a novel focus on perceived bias in emotional-support AI.
  • Objectives: The study fits an SEM linking system attributes, Trust, Perceived Benefits, and Actual System Use, then applies MGA across five sociodemographic dimensions.The dimensions are gender, age, socioeconomic status, education, and cross-national region.
  • Contribution: Across 1,343 users in seven countries, Privacy, Personalization, and Humanlikeness drive Trust, while Perceived Bias degrades it; trust formation varies by user identity.Humanlikeness integrates affective empathy and anthropomorphism as an emergent construct.

2. Literature Review and Hypotheses Development

The paper frames emotional-support LLM adoption as a layered process in which system attributes shape Trust and Perceived Benefits before Actual System Use. It motivates this framework through risks, validated measurement needs, and sociodemographic variation in trust and adoption.

  • Background: LLMs extend conversational-agent use from narrow scripts to open-ended emotional-support conversations, but reported risks include ineffective advice, stigma, sycophancy, and reliance that may discourage professional help-seeking.These concerns support treating emotional-support use as distinct from ordinary task-oriented AI adoption.
  • Conceptual framework: The framework models System Attributes, including empathy, anthropomorphism, personalization, privacy, and perceived bias, as influencing Actual System Use through Trust and Perceived Benefits.Actual System Use is operationalized as self-reported frequency of LLM use for emotional support.
  • Conceptual framework: Exploratory Factor Analysis later merged Cognitive Empathy, Affective Empathy, and Anthropomorphism into a single Humanlikeness construct.The initial model treated these dimensions as conceptually related but separate.
  • Conceptual framework: Perceived Benefits captures functional value, whereas Trust captures perceived safety and relational integrity, creating complementary utilitarian and relational pathways to adoption.The distinction is especially relevant when users disclose intimate feelings and face harms beyond ordinary task failure.
  • Hypotheses: Privacy is expected to support Trust and Perceived Benefits, while Perceived Bias is expected to reduce trustworthiness and usefulness through perceptions of unfair or stigmatizing system behavior.The Perceived Bias construct covers prejudiced assumptions, viewpoint imposition, and failure to acknowledge cultural limitations.
  • Sociodemographic perspectives: The study addresses limited demographic evidence by comparing emotional-support interactions across sociodemographic strata, using regional clusters treated as pragmatic approximations rather than purely cultural groupings.The clusters are Anglosphere, Southern Europe, and Western/Central Europe; the latter includes the Netherlands despite its plausible alternative cultural classification.

3. Methodology

The study uses a cross-sectional international survey to develop and validate a psychometric instrument, then tests structural relationships and demographic differences in emotional-support LLM use. Its analysis combines factor discovery, measurement validation, SEM, bootstrapped indirect effects, and invariant MGA comparisons.

  • Study design: The cross-sectional quantitative survey sampled active LLM users internationally and analyzed responses through EFA, CFA, SEM, and MGA.The design targets user perceptions of LLMs in affective contexts across seven countries.
  • Instrument development: The instrument began with literature-based item generation and included nine latent variables spanning Trust, system attributes, Perceived Benefits, and Actual System Use.Perceived Bias items were authored specifically for emotional-support dimensions such as cultural, gender, and viewpoint-related bias.
  • Instrument development: A 20-expert panel reviewed English items for clarity, relevance, and contextual appropriateness before pilot and translation workflows.The survey was translated into Spanish, French, Italian, German, and Dutch and reviewed by native-speaking experts for conceptual equivalence.
  • Sampling and data collection: Participants were recruited through a Qualtrics panel from seven countries selected to represent major ChatGPT visitor shares in North America and Europe.The study collected data through an online questionnaire with digital informed consent and GDPR-compliant anonymity assurances.
  • Sampling and data collection: Sociodemographic measures covered gender, age, education, and subjective socioeconomic status using the 10-step MacArthur social-status ladder.The ladder measures perceived social standing relative to others rather than objective income or education.
  • Data preparation: Quality control filtered suspected bots, duplicates, incomplete responses, straight-lining, speeders, and failed attention checks from N = 5,319 raw responses.The filters were applied alongside institutional ethics approval and embedded attention-check items.
  • Analysis: SEM tested paths from system attributes through Trust and Perceived Benefits to Actual System Use, with indirect effects estimated using 2,000 bootstrap iterations and bias-corrected 95% confidence intervals.MGA then compared structural paths across gender, age, region, socioeconomic status, and education after measurement-invariance testing.

4. Data Analysis and Results

The analysis validated a seven-construct measurement model and found that users experience empathy and anthropomorphism as one integrated Humanlikeness construct. Structural modeling showed that Privacy, Personalization, and Humanlikeness increase Trust and benefits, while Perceived Bias undermines both and indirectly reduces Actual System Use.

  • Measurement development: 1,343 users were split into EFA and CFA subsamples, while the full sample supported SEM estimation.The clean dataset comprised 671 participants for EFA and 672 for CFA; SEM used the entire sample.
  • Measurement development: EFA indicators for Anthropomorphism, Cognitive Empathy, and Affective Empathy loaded onto a unified Humanlikeness factor.The retained indicators showed loading values between .586 and .742, leading to consolidation into one construct.
  • Measurement validation: The measurement model showed strong fit, with CFI = .940, TLI = .934, RMSEA = .048, and SRMR = .038.All reported fit indices met the stated psychometric benchmarks.
  • Measurement validation: All retained indicators demonstrated convergent validity and internal consistency, with factor loadings from .656 to .876 and reliability coefficients exceeding .70.Cronbach’s alpha ranged from .766 to .926, Composite Reliability from .772 to .926, and Average Variance Extracted exceeded .50 across constructs.
  • Structural relationships: Privacy was the strongest positive predictor of Trust, followed by Personalization and Humanlikeness, while Perceived Bias significantly reduced Trust.The standardized effects were β = .486, β = .261, β = .153, and β = −.111, respectively.
  • Structural relationships: Perceived Benefits and Trust directly predicted Actual System Use, while Personalization and Privacy had the strongest positive indirect effects through these mediators.Direct effects were β = .476 for Perceived Benefits and β = .348 for Trust; indirect effects were βind = .239 for Personalization and βind = .235 for Privacy.

4.4. Multi-Group Analysis (MGA)

The MGA shows that trust formation and adoption pathways vary significantly across demographic and regional groups. Some groups engage through trust, while older, lower-SES, and less-educated users rely more on perceived practical benefits.

  • Model evaluation: The five MGA configurations achieved excellent overall fit, with RMSEA below .05 and CFI values ranging from .888 to .927.Measurement invariance was evaluated before comparing structural relationships across groups.
  • Overall group differences: Structural differences were significant for gender (p = .009), age (p = .001), region (p = .000), and education (p = .000), while SES was marginal overall (p = .080).SES path-level differences were therefore treated cautiously.
  • Gender differences: Privacy strengthened Trust more for women than men (C.R. = 2.44), while Humanlikeness strengthened Perceived Benefits more for women (C.R. = 2.37).The findings indicate gender-specific trust and benefit formation mechanisms rather than a uniform gender effect.
  • Age differences: For late-adulthood users, Trust did not significantly predict Actual System Use (β = .165), whereas Perceived Benefits had a strong positive effect (β = 1.179).Older adults therefore engaged through practical value such as availability, accessibility, and non-judgmental support rather than relational trust.
  • Regional differences: Humanlikeness predicted Trust more strongly in the Anglosphere (β = .303) than in Southern Europe (β = .044, non-significant) or Western/Central Europe (β = .139).Privacy also strongly predicted Anglosphere Trust (β = .576), while Perceived Bias damaged Trust most clearly in Southern Europe (β = −.332).
  • SES differences: Low-SES users showed the strongest Privacy-to-Trust path (β = .611), while Trust did not significantly predict their Actual System Use (β = .145, p = .252).These findings were exploratory because the overall SES test was non-significant and only the Privacy →Trust difference survived Bonferroni correction.
  • Education and adoption logic: Users without college degrees bypassed Trust (β = .077, non-significant) and relied on Perceived Benefits to drive use (β = 1.219).Across groups, educated, higher-SES, and younger users followed trust-mediated adoption, whereas older, lower-SES, and less-educated users followed benefits-driven adoption.

5. Discussion

The discussion presents a validated measurement foundation and a dual-pathway account of emotional-support AI adoption. It argues that users’ routes to engagement differ across demographic and cultural groups, with implications for design, governance, and future validation.

  • Measurement contribution: The study develops and validates a seven-construct instrument capturing social-affective and bias-related perceptions of LLM emotional support.The instrument includes a novel Perceived Bias scale and is intended to support related affective-computing and HCI research.
  • Technology acceptance theory: Technology acceptance is extended through two routes to actual use: a utilitarian Perceived Benefits pathway and a relational Trust pathway.These pathways can carry system attributes into use jointly and independently, and their dominance varies across users.
  • Humanlikeness: Cognitive empathy, affective empathy, and anthropomorphism load onto one latent Humanlikeness dimension rather than separate user-perceived constructs.The integrated scale reflects users’ holistic impressions of conversational flow and authenticity.
  • Perceived Bias: Perceived Bias independently degrades both Trust and Perceived Benefits, eroding relational and utilitarian foundations of adoption simultaneously.The study introduces a psychometric measure of subjective bias perception for LLM-based emotional support.
  • Sociodemographic variation: Privacy has a stronger influence on women’s Trust, while Anglosphere users respond more strongly to Humanlikeness than European users.These findings connect demographic variation in AI trust with privacy sensitivity and cultural responses to non-human social agents.
  • Design implications: Designers are advised to prioritize transparent privacy practices, calibrate Humanlikeness to users, and treat Perceived Bias as a first-class design concern.The discussion recommends demographic sensitivity checks and culturally pluralistic response generation.
  • Limitations: The sample is limited to Western WEIRD countries, so the reported cross-national patterns should not be generalized beyond Western contexts.The study also relies on self-reported use and measures perceived rather than objective bias; future work should add behavioral measures and systematic bias audits.

6. Conclusion

The conclusion addresses gaps in measurement and diverse evidence by combining a validated scale with structural and multi-group analyses. It finds that trust and benefits mediate adoption differently across user segments.

  • Conclusion: The study addresses missing validated instruments and large-scale sociodemographic evidence by analyzing 1,343 users across seven countries.It develops a seven-construct scale and tests Trust and Perceived Benefits as pathways to Actual System Use.
  • Conclusion: Privacy, Personalization, and Humanlikeness drive Trust, while Perceived Bias degrades both Trust and Perceived Benefits.Humanlikeness integrates cognitive empathy, affective empathy, and anthropomorphism into one construct.
  • Conclusion: Educated, higher-income, and younger users show trust-mediated adoption, whereas older and lower-socioeconomic groups bypass Trust and rely on perceived practical benefits.The conclusion presents this divergence as evidence that adoption logic is not universal across users.

Disclosure Statement

The authors report no potential conflict of interest.

  • No potential conflict of interest was reported by the authors.

Data Availability Statement

The study’s supporting data are available from the corresponding author upon request.

  • Supporting data are available from the corresponding author upon request.

Ethics Approval Statement

The study received ethics approval from the University of Amsterdam’s Economics and Business Ethics Committee.

  • Ethical approval was granted by the University of Amsterdam’s Economics and Business Ethics Committee (Approval No: EB-18813).

Author Contribution Statement

The authors divided responsibilities across conceptualization, analysis, data curation, methodology, supervision, project administration, and manuscript preparation, with all authors approving the final version.

  • Contributions covered conceptualization, data curation, formal analysis, investigation, methodology, supervision, project administration, and manuscript writing or editing.
  • All authors approved the final manuscript, critically revised it, and accepted accountability for all aspects of the work.

Appendix A. Psychometric instrument for emotional support AI

Appendix A presents the survey constructs, item codes, statements, and sources used to assess emotional-support chatbot perceptions, benefits, trust, use, and bias. Items removed during EFA or CFA purification are marked with gray strikethrough.

  • Survey instrument overview: Table A1 lists the survey constructs, item codes, statements, and sources.
  • Survey instrument overview: Gray-strikethrough items were removed during EFA or CFA purification to preserve structural integrity.
  • Perceived bias: The instrument measures perceived impartiality, cultural understanding, bias awareness, multiple perspectives, and avoidance of biased assumptions.
  • Trust, use, and outcomes: Additional items assess confidence, constructive challenges, alternative viewpoints, trust in support and information, overall trust, and frequent or routine use.
  • Perceived benefits: Benefit items assess practical advice, actionable well-being steps, 24/7 availability, affordability, overall value, emotional expression, and calmer feelings.
  • Trust, use, and outcomes: The questionnaire also asks whether the chatbot is among users’ first resources for mental well-being or emotional support.

Appendix B. Exploratory Factor Analysis Results

Appendix B reports rotated component matrices from the first split sample and documents the constructs and item-retention decisions used for later CFA.

  • The independent-variable analysis is presented in a rotated component matrix based on the first split sample of 671 participants.
  • Humanlikeness (HUM) combines anthropomorphism, cognitive empathy, and affect empathy.
  • The reported constructs also include perceived bias, privacy, personalization, perceived benefits, actual system use, and trust.
  • Items were retained for the CFA phase or dropped after EFA during CFA to improve discriminant validity.
  • The mediating and response-variable analysis is likewise presented in a rotated component matrix from the first split sample of 671 participants.
Loading 2608.21220v1…