Source-linked AI summary
All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers
Pablo A. Fonseca, Raquel Rodríguez-Carvajal, Rafael A. Calvo
TL;DR
LLM evaluations often overlook sustained, emotionally charged support conversations and differences among help-seekers. This paper introduces a personality-aware AAA evaluation using psychologically specified synthetic Advisees, four LLM Advisors and blind Auditors in an acute dementia-caregiving crisis. Profiles were recovered from dialogue, while all four models shared verbosity, talking more than listening, and premature problem-solving, without differing reliably in emotion stabilisation.
Problem
Existing evaluations mostly assess single-turn factual accuracy rather than sustained, emotionally charged conversations or differences among psychologically distinct users.
Method
The AAA framework uses psychologically specified synthetic Advisees, four LLM Advisors and blind Auditors to evaluate crisis-support interactions across emotional, psychological-needs and behavioural domains.
Results
Profiles were recovered from dialogue with band-score r = 0.78, while all four Advisors shared verbosity, talk-to-listen ratios above one and premature problem-solving, without reliable differences in emotion stabilisation.
Takeaways & Limitations
Personality-aware evaluation can extend beyond the Big Five to coping, self-efficacy, resilience and reactance while exposing shared support-process failures across models.
Takeaways & Limitations
The findings are a proof of concept based on one British dementia-caregiving scenario and seven profiles without within-cell replication, with emotional states inferred from conversation text.
Abstract
from arXiv · showhide
Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified profile, in an acute crisis: a caregiver learning of a relative's dementia diagnosis. Auditors blind to the profile prompt recovered the specified bands from dialogue alone with high agreement on every instrument (ICC(2,4) = 0.91; 0.79-0.96 by instrument; band-score r = 0.78), as expected for the Big Five but equally for coping style, coping self-efficacy, resilience and reactance, which the lexical approach never covered. Such evaluation therefore reaches beyond the Five Factor Model to motivational, regulatory and self-appraisal dispositions. The four models were not distinguishable on emotion stabilisation and failed alike, sharing three modes: verbosity, a talk-to-listen ratio above one, and problem-solving before the situation had been explored.
1 Introduction
LLM support is increasingly used during distress, but existing evaluations rarely test sustained, emotionally charged exchanges or account for users’ psychological profiles. The paper introduces a safe, role-separated, personality-aware framework to evaluate such interactions dynamically.
- Motivation: Single-turn factual-accuracy benchmarks do not capture how LLM support unfolds across sustained, emotionally charged conversations.
- Motivation: Acute stress narrows attention and reduces working-memory and executive resources, making listening and stabilisation important before information-giving or problem-solving.
- Research gap: Personality-aware evaluation requires synthetic help-seekers whose specified psychological profiles can be instantiated and recovered from dialogue.
- Contribution: The AAA framework separates synthetic Advisees, Advisors and Auditors so interactions can be simulated safely while each role’s contribution is isolated.
- Contribution: Seven synthetic Advisees with profiles spanning personality, coping, self-efficacy, reactance and resilience advised four widely available LLMs after a relative’s dementia diagnosis.
- Contribution: The framework evaluates emotional stabilisation, basic psychological needs and support behaviour, including a clinically grounded counter for problematic conduct.
2 Results
Study 1 tested whether specified psychological profiles could be reproduced in synthetic Advisees and recovered from sustained conversations, while Studies 2–4 evaluated emotional stabilisation and support across four Advisors. Profiles were recoverable with strong ensemble agreement, whereas Advisors were not distinguishable on emotional change and showed profile- and support-related differences in other outcomes.
- Study design: The four studies jointly tested profile reproduction, blind profile recovery, and Advisor performance across emotionally charged conversations.Seven synthetic Advisees were instantiated with four models; Study 1 used four independent Auditors, while Studies 2–4 used a single higher-capability Auditor.
- Profile recovery: ICC(2,4) = .906 for the four-Auditor ensemble, compared with ICC(2,1) = .706 for a single Auditor.Agreement was highest for NEO-FFI dimensions and lowest for reactance and CSES-8 self-efficacy, with CHIP and ResQ-Care between them.
- Profile recovery: GPT-4.1 remained the top instantiating model after each Auditor was removed, with r = .768 without GPT-4.1 versus .779 for the full panel.Lower rankings were unstable: Mistral-Medium and DeepSeek-V3.2 exchanged third and fourth place across panels.
- Profile recovery: GPT-4.1, Claude-Sonnet-4.5 and DeepSeek-V3.2 were not reliably separable on profile recovery, while all three exceeded Mistral-Medium.The reported paired differences were GPT-4.1 − DeepSeek-V3.2 ∆r = 0.064 [−0.039, 0.207], p = .31, and GPT-4.1 − Claude-Sonnet-4.5 ∆r = 0.035 [0.000, 0.071], p = .047.
- Emotion stabilisation: Perceived stress fell for every Advisor, but the reduction survived multiplicity correction only for GPT-4.1; positive affect increased pointwise for all four, with only Claude-Sonnet-4.5 surviving correction.Positive-affect change was reported descriptively because increasing it was not a prespecified stabilisation target.
- Emotion stabilisation: No omnibus Advisor difference was detected for perceived stress, positive affect, or negative affect, while Advisee identity explained 45–56% of emotional-change variance.The same seven Advisees appeared in all four arms, enabling direct comparisons blocked on Advisee.
- Basic psychological needs: DeepSeek-V3.2 ranked first in overall net support at 4.16, narrowly ahead of Claude-Sonnet-4.5 at 4.14.DeepSeek-V3.2 led autonomy and relatedness support, Claude-Sonnet-4.5 led competence support, and Mistral-Medium performed worst across all three need dimensions.
- Overall scoring: The composite averaged four equally weighted dimensions, with emotion handling defined as within-session change relative to achievable improvement.Intervention quality combined CEISStotal/36 with relational and technical global scores; behavioural conduct was computed separately as 1 − problematic behaviours per Advisor turn / 2.
3 Discussion
The AAA framework makes psychological profiles recoverable from synthetic help-seekers’ dialogue and reveals shared conversational failures across four Advisors, despite differences in process measures. Its findings support profile-aware, multi-turn auditing while remaining exploratory and bounded by the single crisis scenario and Auditor design.
- Framework: The AAA framework studies crisis conversations among Advisors, psychologically specified Advisees, and Auditors across four studies.Each role can be AI or human; here, the framework is applied to psychological-profile-informed mental-health crisis conversations.
- Shared failure modes: Across all four Advisors, word ratios exceeded one at 1.79–3.26, and models moved toward problem-solving before exploring the Advisee’s situation.Mean Advisor turns were roughly 280 to 590 words; rushing past emotional disclosure and boundary or scope overreach were rare in this scenario.
- Profile recovery: Band-score correlations reached r = 0.78, while four-Auditor agreement was excellent for NEO-FFI dimensions and good for resilience, coping style, coping self-efficacy, and reactance.Agreement values were ICC(2,4) = .964 for NEO-FFI, .870 for resilience, .827 for coping style, .817 for coping self-efficacy, and .791 for reactance.
- Profile recovery: Aggregating four independent judges lifts every construct into a usable range, making the panel the framework’s practical unit for auditing internal appraisals.The paper contrasts this with more economical scoring for dispositions that surface in speech.
- Advisor differences: Emotion outcomes did not differ reliably by Advisor, whereas process measures differed in crisis-intervention skill, relational quality, need support, false understanding, premature problem-solving, and reactance-producing expressions.For example, CEISS total differed with F(3, 18) = 30.61, Holm p < .001, partial η2 = .84.
- Scope and limitations: The findings are exploratory proof-of-concept evidence because the scenario was fixed to British family caregivers facing dementia diagnosis, profiles numbered seven, and there was no within-cell replication.Provider-level differences also confound model architecture, training data, and safety policies, while Studies 2 to 4 used a single Auditor vulnerable to a family-preference objection.
4 Methods
The AAA framework evaluates advice through separated Advisee, Advisor and Auditor roles, using synthetic help-seekers with prescribed psychological profiles in a fixed acute-crisis scenario.
- AAA framework: The framework separates Advisees receiving advice, Advisors providing it, and Auditors assessing its quality.
- AAA framework: Separating roles from whether participants are human or synthetic supports reusable software abstractions and transfers outcome measures across human–AI and human–human conversations.
- Crisis context: The study holds the precipitating event constant so interaction differences can be attributed to psychological profiles rather than different stressors.
- Advisee specification: The profile extends beyond the Five Factor Model to coping style, coping self-efficacy, resilience and reactance as dispositional motivational, regulatory and self-appraisal constructs.
- Advisee specification: Seven synthetic Advisees were specified on low, medium or high levels across personality, coping, self-efficacy, reactance and resilience-related dimensions.
- Advisee specification: Questionnaire-derived representative phrases were repurposed as prompting instructions to elicit specified Advisee responses.
4.3 The no-profile baseline
The no-profile baseline removes psychometric specifications and enactment instructions while matching the specified-profile conversations on the remaining Study 1 configuration.
- No-profile baseline: The baseline removed all five psychometric specification blocks and the instruction to enact the profile.
- No-profile baseline: 28 baseline conversations mirrored 28 specified-profile conversations, producing 56 Study 1 conversations in total.
- No-profile baseline: Recovery was compressed toward the scale midpoint at both extremes, with band-score correlations of r = 0.69–0.78 and in-band accuracy of 50–63%.
4.4 Advisor specification
Advisor specification used widely available models with minimal provider-layer safeguarding and instrument-specific Auditor access to conversation turns.
- Advisor specification: The study aimed to compare widely available models with minimal additional safeguarding so observed conduct would be attributable to the model rather than the provider product layer.
- Advisor specification: Four Advisor models were used: gpt-4.1, claude-sonnet-4-5, mistral-medium-2505 and DeepSeek-V3.2.
- Advisor specification: All Advisor models ran at temperature 0 with the provider default minimal system prompt, while Advisee instantiations ran at temperature 0.7.
- Auditor specification: Auditor access varied by instrument: PANAS and PSS used selected Advisee turns, TENS-Life and MITI used full dialogue, and CEISS used Advisor turns only.
- Auditor specification: Auditors received no profile prompt, demographic block or scenario briefing, and baseline and specified-profile conversations used identical condition-blind prompts.
- Auditor specification: The MITI implementation assigned at most one code per Advisor turn, making complex-reflection percentage a presence/absence indicator; only relational and technical globals entered the composite.
4.6 The behavioural counter
The behavioural counter operationalises crisis-support risks across relational, boundary and process behaviours, including verbosity and premature problem-solving.
- Behavioural counter: The counter targets Advisor behaviours that may be acceptable in ordinary conversation but are commonly counterproductive during acute crisis.
- Reactance-producing expressions: Reactance-producing expressions include false understanding, premature trust and false reassurance, which can threaten autonomy or bypass specific validation.
- Process failures: Process failures violate the listen-before-solving sequence by moving to information-giving or action planning before stabilisation, rapport and exploration.
- Process failures: Premature problem-solving jumps to solutions or action plans before the Advisee’s situation has been explored.
- Process failures: Rushing past emotion shifts topic or offers solutions without acknowledging an emotionally loaded disclosure or giving it space.
- Boundary and scope: Boundary failures include scope overreach and failure to close, respectively taking over the Advisee’s work or continuing after winding-down signals.
- Turn length: 89% of Advisor turns exceeded 160 words, making the verbosity count an indicator of near-uniform verbosity rather than a discriminating measure.
4.7 Scenario and task specification
The study uses scenarios to exploit multi-turn interaction and tailor support to psychological profiles. Scenario design is coordinated with persona specifications and informs which Auditors are relevant.
- Scenarios support the evaluation of multi-turn interactions.
- Scenario design controls how support is adjusted for different psychological profiles.
- The selected scenario concerns an Advisee receiving distressing news about a close relative.
4.8 Statistics and reproducibility
Analyses are exploratory because each Advisor was evaluated with seven Advisees without within-cell replication. The study therefore prioritizes exact, distribution-free inference, multiplicity control, and stringent inter-rater agreement estimates.
- Seven Advisees per Advisor and no within-cell replication make the analyses exploratory.Effect sizes with confidence intervals are primary, while p values are secondary.
- Exact and distribution-free methods are used throughout because n = 7.Confidence intervals for standardised mean differences are obtained by inverting the noncentral t distribution.
- Wilcoxon signed-rank tests provide distribution-free checks for pre–post contrasts.Normality was not rejected in any of twelve contrasts, but the Shapiro–Wilk test was recognised as low-power at n = 7.
- Between-Advisor comparisons use repeated-measures analyses blocking on Advisee, with Friedman’s rank test as a distribution-free check.The same seven Advisees appear in every arm.
- Holm correction controls familywise error within each pre–post instrument and between-Advisor family.
- Inter-rater reliability uses two-way random-effects intraclass correlations with absolute agreement.Systematic leniency differences between Auditors count against reliability rather than being partialled out.
4.9 Use of large language models in manuscript preparation
Large language models assisted with language editing and analysis-code writing, but the authors state that they did not generate scientific content or interpret results. The study used only synthetic conversations and did not require ethical approval.
- Large language models were used for language editing and assistance in writing analysis code.
- The authors state that no large language model generated scientific content, interpreted results, or drafted the arguments.
- The study involved no human or animal participants, human material, or human data.All Advisees were synthetic and conversations were generated between language models.
- The authors determined that ethical approval was not required and sought no institutional review.
Data availability
The study makes its conversation transcripts, Advisee profile specifications, and per-conversation Auditor scores available online. No human-participant data were collected.
- Conversation transcripts are available at the supplied DOI.
- Advisee profile specifications are included in the available study materials.
- Per-conversation Auditor scores are available with the study materials.The passage also states that no human-participant data were collected.
Code availability
The conversation-generation code, Advisee profiling, auditing, behavioural-counter computation, and statistical-analysis scripts are publicly available, while licensed instrument items are withheld and replaced with placeholders.
- Code availability: The code for generating conversations, instantiating Advisee profiles, running Auditors, computing the behavioural counter, and conducting statistical analyses is available online.The passage provides the repository DOI.
- Code availability: NEO-FFI and CHIP inventory items were removed from released Advisee prompt templates because the instruments are commercially licensed.The items are available from PAR and Multi-Health Systems; all other instrument content is included.
Competing interests
The authors report no financial or non-financial competing interests and no financial or advisory relationships with the evaluated model providers.
- Competing interests: No authors report financial or non-financial competing interests.
- Competing interests: No author has a financial or advisory relationship with any evaluated model provider.