Source-linked AI summary
AI emotional support is better only when chosen, but shifts preferences even when it is not
Yaoxi Shi, Cathy Mengying Fang, Guy LabanPattie Maes, Amit Goldenberg
TL;DR
People increasingly choose between human and AI emotional support, but prior research offered limited evidence when received support mismatched that choice. Across experimental and longitudinal studies, the paper finds that AI was rated more favorably when chosen, while personal conversations shifted future preferences toward AI.
Problem
Prior research offered limited evidence about emotional-support evaluations and preferences when the received partner mismatches the person’s initial choice.
Method
The paper combines experimental and longitudinal studies to examine how initial choice, partner congruence, and personal conversations shape emotional-support evaluations and preferences.
Results
AI was rated higher than humans only when initially chosen, while personal conversations increased AI preference from 25.4% to 31.9%.
Takeaways & Limitations
Emotional-support choices are path-dependent, with previous engagement substantially influencing future choice.
Takeaways & Limitations
It remains uncertain whether reported preferences translate into changes in real support seeking.
Abstract
from arXiv · showhide
People increasingly face a novel decision when seeking emotional support: human or AI. In existing studies, AI's empathic messages are rated as well as or better than humans'. But these studies either assigned the support source or honored people's choice. In real life, support is often incongruent with choice, as people want one source and receive the other. Across three experiments (N = 1,951), participants chose whether to share an emotional experience with a human or an AI, then were randomly assigned to a congruent or incongruent partner. AI support was rated as superior only among those who had chosen it. Yet regardless of congruence, interacting with AI increased willingness to choose it again. In a 28-day study with OpenAI (N = 981), daily conversations shifted preferences toward AI and away from humans, but only when conversations turned personal. Emotional support choices are thus path-dependent, progressively redirecting away from human connection.
AI and Emotional Support
AI often receives higher empathy ratings, but evaluations depend on awareness, beliefs, and preferences about the support source. This paper therefore examines how source choice, choice congruence, and prior exposure shape emotional-support evaluations and future preferences.
- Prior evidence: A meta-analysis estimated AI’s superiority in providing short empathic responses at Cohen’s d of 0.55.Earlier studies found AI empathic responses rated better than human responses, including those from doctors or professional support providers.
- Source awareness: Labeling responses as AI-generated rather than human-generated reduces participants’ satisfaction and perceived empathy.These findings indicate that empathy evaluations depend on beliefs and preferences about who or what provides the response, not only response content.
- Choice and access: People often choose an emotional-support source before receiving support, yet limited access can make support choice-incongruent in practice.Humans face availability and social frictions, whereas AI can offer an accessible alternative at almost no cost.
- Research goals: The paper’s second goal is to test how choice congruence and incongruence influence evaluations of interactions with human or AI supporters.This addresses situations in which people receive support from a source different from the one they desired.
- Research goals: The paper’s third goal is to examine path dependence in emotional-support choices, including whether experiencing one source shifts future preference toward it.AI exposure may occur incidentally because of its widespread availability and may update perceptions of AI’s empathic capabilities.
The Current Studies · Experimental Design
The paper combines survey, experimental, and longitudinal designs to study beliefs about human and AI emotional support, choice congruence, and preference shifts after interaction. Participants first evaluated and chose a support source; in Studies 2–3, they then experienced a randomly assigned human or AI conversation and reported post-conversation evaluations and preferences.
- The Current Studies: The studies addressed beliefs about human versus AI support, the effects of congruent versus incongruent choices, and path-dependent changes in later preferences.The program included a survey, two experiments, and a longitudinal study.
- The Current Studies: Study 1 measured beliefs about sharing emotions with humans and AI across support, nonjudgment, advice, and confidentiality before participants chose a partner.Participants recalled a recent significant emotional experience and made both binary and continuous preference assessments.
- The Current Studies: Studies 2–3 extended the design with real-time conversations to test how choice congruence affected interaction quality, beliefs, and subsequent support choices.Study 2 used ChatGPT and Study 3 used Claude.
- The Current Studies: Study 4, conducted with OpenAI, tracked daily AI interactions over a month to examine whether preference shifts generalized to naturalistic longitudinal conversations.Conversations covered personal, non-personal, and open-ended topics.
- Experimental Design: Before conversation, participants recalled a recent positive or negative emotional experience, rated beliefs about human and AI listeners, and selected a preferred partner.Beliefs covered emotional support, nonjudgment, confidentiality, and advice quality, including each factor’s importance.
- Experimental Design: In Studies 2–3, participants were randomly assigned to a human or AI text conversation whose partner matched or mismatched their initial choice.The five-minute conversations used GPT-4o in Study 2 and Claude Opus 4.6 in Study 3, with identical supportive instructions and structure.
- Experimental Design: After conversation, participants reassessed the four support dimensions, rated empathy, enjoyment, satisfaction, and continuation desire, and reported future human-versus-AI preferences.Post-conversation measures were designed to capture changes in beliefs and subsequent choices.
Pre-Conversation Beliefs and Choice · How do people perceive sharing emotional experiences with a human and an AI?
Before conversation, participants generally viewed humans as more emotionally supportive but more judgmental than AI, while AI was seen as more confidential. These beliefs varied systematically with initial choice and helped predict whether participants chose a human or AI partner.
- How do people perceive sharing emotional experiences with a human and an AI?: Participants in Studies 1–3 rated human and AI partners across 14 items clustered into emotional support, judgment, confidentiality, and advice quality.These dimensions were used to characterize pre-conversation beliefs and compare the two support sources.
- How do people perceive sharing emotional experiences with a human and an AI?: AI was perceived as more confidential than a human partner (t(1893) = 7.53; P < 0.001; Cohen's d = 0.17; mean difference, 0.39).Perceived advice ability did not differ significantly between AI and human partners (t(1893) = 0.33; P = 0.743; Cohen's d = 0.01; mean difference, 0.01).
- How do people perceive sharing emotional experiences with a human and an AI?: Initial choice corresponded to distinct belief profiles: AI choosers favored AI on judgment, confidentiality, and advice, while human choosers favored humans on emotional support, advice, and confidentiality.AI choosers still rated humans slightly higher in emotional support, whereas human choosers rated partners similarly on nonjudgment.
- How do people perceive sharing emotional experiences with a human and an AI?: Emotional support and nonjudgment were the primary predictors of conversation-partner choice, with emotional support showing the largest coefficient magnitude.The analysis entered human–AI difference scores for four perception dimensions into a logistic regression predicting binary partner choice.
- How do people perceive sharing emotional experiences with a human and an AI?: AI was chosen by 75.5% of Study 1 participants, 69.6% of Study 2 participants, and 63.2% of Study 3 participants.Participants selected whether to share the recalled emotional experience with an AI chatbot or a human partner.
- How do people perceive sharing emotional experiences with a human and an AI?: Participants reported greater comfort and willingness to share with AI than with a human partner (t(1893) = 15.80; P < 0.001; Cohen's d = 0.36; mean difference, 0.51 SD).These preferences were strongly associated with everyday emotion-sharing tendencies involving humans and AI.
Experienced Quality of the Conversation
The section examines how matching participants’ initial choice with their assigned human or AI conversation partner affects experienced conversation quality. Studies 2 and 3 tested whether choice congruence improves the sharing experience.
- Experienced Quality of the Conversation: The section investigates whether congruence between participants’ choice and their assigned partner influences conversation quality, for both human and AI interactions.The comparison concerns participants who interacted with a human versus an AI.
- Experienced Quality of the Conversation: In Studies 2 and 3, participants completed a 5-minute conversation with either an AI chatbot or a human listener, independent of initial preference.Study 2 used GPT-4o and Study 3 used Claude Opus 4.6 for the AI chatbot.
- Experienced Quality of the Conversation: The authors hypothesized that choice congruence would have a positive main effect on people’s sharing experience.This prediction was labeled H2a.
How does the congruence between choice and chat condition influence · Post-Conversation Perception and Choice
Choice–partner congruence improved conversation experiences mainly when participants received AI after choosing it, while AI support was rated more favorably than human support only among initial AI choosers. After conversation, AI increased perceived emotional support and nonjudgment, whereas human support perceptions were mixed, shifting future partner beliefs.
- How does the congruence between choice and chat condition influence: Choice–partner congruence positively affected empathy, desire to continue, enjoyment, and satisfaction, but these effects occurred among AI-assigned participants and not human-assigned participants.Congruence effects were significant for AI assignments (Ps < 0.001) and nonsignificant for human assignments (Ps > 0.40).
- How does the congruence between choice and chat condition influence: AI support was rated more favorably than human support only when participants had initially chosen AI, identifying initial preference as a moderator of AI superiority.The effect was driven primarily by participants who chose and received AI, rather than by objective differences in AI responses.
- How does the congruence between choice and chat condition influence: AI response empathy did not differ between choice-congruent and incongruent conversations, indicating that congruence altered subjective experience rather than the objective empathy provided.The comparison was t(525) = 0.02; P = 0.98; β = 0.0005; 95% CI, (−0.04, 0.04).
- Post-Conversation Perception and Choice: After a 5-minute interaction with AI, participants perceived more emotional support and nonjudgment than before, regardless of their initial choice.Emotional support increased by mean difference 1.07, while nonjudgment increased by mean difference 0.52.
- Post-Conversation Perception and Choice: Among initial human choosers, chatting with AI raised perceived AI emotional support to the pre-conversation level of initial AI choosers.The comparison was t(299.7) = -0.41; P = 0.685; Cohen's d = -0.04; mean difference, -0.06.
- Post-Conversation Perception and Choice: Human emotional-support perceptions among initial AI choosers remained similar to pre-conversation beliefs after chatting with a human.The comparison was t(257) = 1.87; P = 0.062; Cohen's d = 0.12; mean difference, 0.18; 95% CI, (-0.01, 0.38).
How does the interaction change people’s future choice on who to share emotions
A brief emotional-support conversation shifted people’s future sharing preferences toward the partner they had actually experienced. This path-dependent effect held regardless of participants’ initial preference for AI or humans.
- Future sharing preference: 70.1% of participants assigned to chat with an AI chose AI for future emotion sharing, versus 45.5% in the human condition.The result followed a 5-minute conversation about emotions and was statistically significant (z = 7.54; β = 0.51; 95% CI, (0.38, 0.65)).
- Future sharing preference: Among participants initially choosing AI, chatting with a human versus AI predicted greater likelihood of choosing a human in the future.The effect was significant (z = 6.64; s.e. = 0.11; P < 0.001; β = 0.70; 95% CI, (0.50, 0.91)).
- Future sharing preference: Among participants initially choosing a human, chatting with AI versus a human predicted greater likelihood of choosing AI in the future.The effect was significant (z = 4.05; s.e. = 0.15; P < 0.001; β = 0.61; 95% CI, (0.31, 0.90)).
- Future sharing preference: The choice of emotional-support partner was path-dependent: sharing emotions with a partner increased the likelihood of choosing that same partner later.The main effect of chat condition was consistent across participants who initially chose either a human or an AI.
Study 4: Examining Effects of Using AI on Future Preference in a Longitudinal Study · personal matters?
In a 28-day naturalistic experiment, daily AI conversations shifted preferences toward AI and away from humans for discussing personal issues, especially after personal or open-ended conversations. Emotionally engaging behaviors from both AI and participants predicted this shift.
- Study 4: Examining Effects of Using AI on Future Preference in a Longitudinal Study: Study 4 addressed earlier concerns that immediate post-interaction choices might reflect recency effects and that brief laboratory interactions might not generalize to everyday AI use.The longitudinal naturalistic design examined daily interactions over an extended period and included a broader range of personal topics.
- Study 4: Examining Effects of Using AI on Future Preference in a Longitudinal Study: Study 4 used a 28-day randomized trial in which participants conversed with ChatGPT daily about personal, non-personal, or open-ended topics.Participants were instructed to engage in at least a 5-minute chatbot conversation each day, with topic and modality experimentally assigned.
- Study 4: Examining Effects of Using AI on Future Preference in a Longitudinal Study: The topic manipulation produced more emotionally supportive conversations in the Personal condition, where Emotional Support & Empathy accounted for 43.1% of conversations versus 9.6% in Open-ended.This category represented a negligible share in the Non-personal condition.
- personal matters?: After 28 days, AI preference for discussing personal issues increased from 25.4% to 31.9% (t(980) = 3.88; P < 0.001; d = 0.12).Human preference shifted in the opposite direction, declining from 82.1%, although the supplied passage does not report the endpoint.
- personal matters?: AI preference rose by +11.6% in Personal (31.6% → 43.2%) and +9.5% in Open-ended (24.9% → 34.4%), but fell by −1.2% in Non-personal (20.2% → 19.0%).The corresponding human-preference changes were −10.3% in Personal, −5.4% in Open-ended, and −1.5% in Non-personal.
- personal matters?: Compared with Non-personal conversations, AI preference increased more in Open-ended (OR = 2.25, 95% CI, (1.55, 3.25), Tukey-adjusted P < 0.001) and Personal (OR = 3.03, 95% CI, (2.09, 4.38), Tukey-adjusted P < 0.001) conditions.Personal conversations also reduced human preference more than Non-personal conversations (OR = 0.56, 95% CI, (0.38, 0.83), Tukey-adjusted P = 0.010).
- personal matters?: Twelve of 13 conversational behaviors significantly predicted both outcomes: higher odds of choosing AI and lower odds of choosing a human.The analyses controlled for pre-study preference, conversation topic, and modality, with Benjamini–Hochberg correction across the 13 tests per outcome.
- personal matters?: AI validation of users’ feelings predicted choosing AI (OR = 1.98, 95% CI, (1.59, 2.48), P < 0.001), while AI empathy was the strongest negative predictor of choosing a human (OR = 0.55, 95% CI, (0.42, 0.71), P < 0.001).Participants’ sharing problems, seeking support, and self-disclosure of feelings also strongly predicted future preference for AI while negatively predicting choosing a human.
Discussion
The discussion identifies choice as a key moderator of AI emotional-support experiences and shows that interacting with AI can redirect future support preferences. This path dependence may shift people toward AI, while motivating designs that encourage continued contact with supportive humans.
- Drivers of choice: People who chose humans believed humans offered better emotional support, whereas people who chose AI viewed human and AI support as equally capable.AI-preferring participants also viewed AI as much more nonjudgmental, while human-preferring participants saw humans and AI as similarly nonjudgmental.
- Consequences of choice: AI was rated higher than humans only when participants had initially preferred and been assigned to AI, concentrating AI superiority among AI-preferring participants.The finding qualifies prior research suggesting that AI is generally superior in short empathic messages.
- Consequences of choice: Regardless of initial choice, interacting with AI about emotional or personal experiences increased willingness to choose AI for future sharing and updated beliefs about its support capacity.This result indicates that exposure to AI can alter subsequent preferences even when the interaction is incongruent with the participant’s initial choice.
- Implications and limitations: Because the limits of AI-driven preference shifts remain uncertain, systems could encourage users engaging in emotional interactions to reach out to supportive humans as well.The authors present this as a way to mitigate the path-dependent shift toward AI while acknowledging that people may continue to prefer human connection.
Limitations and future directions
The research is limited by its comparison with human strangers, structured support-seeking settings, and reliance on self-reported preferences. Future work should include closer human alternatives, naturalistic support pathways, actual behavior, and longer-term social and well-being consequences.
- Human alternatives: The lab studies compared AI only with human strangers, so results may differ when people choose between AI and friends, partners, family members, or therapists.Future research should examine how people choose and experience support when close others or professionals are available.
- Naturalistic pathways: 28 days of broader daily AI use increased participants’ preference for discussing personal matters with AI and decreased preference for human partners.The open-ended condition most closely approximated typical everyday use.
- Naturalistic pathways: It remains unclear whether effects from structured settings generalize to the diverse real-world pathways through which people receive AI emotional support.Emotional support may also emerge incidentally during task-oriented use of general-purpose AI platforms.
- Long-term consequences: Future research should examine downstream effects on social relationships and well-being over time as support-partner choices evolve.Naturalistic studies should also assess how people experience emotional support and how choice influences it.
- Behavioral validation: Preference-change measures relied on self-reports rather than actual daily support-seeking behavior, leaving it open whether reported preferences translate into behavior.Future work should track whom people actually turn to for support in real-world settings.
Ethical statement · Preregistrations · Participants
The studies received institutional review board approval, obtained informed consent, and were preregistered. Participants were recruited across Prolific and CloudResearch, with final samples of 229, 487, 457, and 981 after preregistered exclusions.
- Ethical statement: Studies 1–3 received Harvard University IRB approval, while Study 4 received joint OpenAI–MIT approval through Western Clinical Group IRB.The cited approvals were IRB24-1718, IRB25-0035, and WCG IRB #20243987.
- Ethical statement: Informed consent was obtained from participants in all studies.
- Preregistrations: All studies were preregistered, including Study 1 at aspredicted.org/7xhy-ds3c.pdf.
- Participants: Study 1 recruited 250 participants from Prolific and retained 229 after excluding 20 attention-check failures and 1 nonconsenting participant.Participants had Mean age = 40.1, SD = 15.0; 63.1% female; 77.9% White, and each was compensated $1.80.
- Participants: Studies 2 and 3 recruited U.S.-resident, fluent-English participants with approval rates of at least 90% through CloudResearch Connect.Real-time dyadic interaction required batch recruitment to maintain a sustained influx for prompt human-partner pairing.
- Participants: 457 participants remained in Study 3 after preregistered exclusions, including 194 in the human condition and 263 in the AI condition.Among them, 175 chose an AI partner and chatted with an AI, 109 chose AI but chatted with a human, 88 chose human but chatted with AI, and 85 chose and chatted with a human.
Procedure
Studies 1–3 elicited and measured participants’ emotional experiences, assessed perceptions of human and AI support, and then assigned participants to matched text conversations with either a human or an AI. Study 4 extended the design into a 28-day randomized trial varying ChatGPT’s voice modality and conversation topic.
- Studies 1–3: Studies 1–3 began with an attention check, followed by recalling and describing a vivid significant emotional experience in 4–5 sentences.Participants reported the experience’s valence and the intensity of happiness, sadness, fear, anger, excitement, pride, anxiety, and enthusiasm.
- Studies 1–3: Study 1 measured source preferences before perceptions, whereas Studies 2–3 reversed the order to test whether ratings reflected motivated perceptions after choice.Perceptions covered emotional support, judgment concerns, confidentiality, and advice quality.
- Studies 1–3: Studies 2–3 randomly assigned participants to a GPT-4o or Claude Opus 4.6 chatbot prompted as an empathetic listener, or to another participant assigned that role.Human participants were matched in real time and engaged in a 5-minute text-based conversation; the AI prompt matched the human instructions in content.
- Studies 1–3: Afterward, participants rated their assigned partner, conversation outcomes, emotional state, future human-versus-AI preferences, AI attitudes and use, and technical difficulties.Post-conversation measures included support, nonjudgment, confidentiality, advice quality, enjoyment, satisfaction, desire to continue, and positivity and negativity.
- Study 4: Study 4 was a one-month randomized controlled trial (n = 981, 28 days, > 300,000 messages) using a 3 × 3 factorial design of modality and conversation topic.Participants used GPT-4o for at least five minutes daily and were assigned to text-only, neutral professional voice, or engaging expressive voice conditions crossed with open-ended, non-personal, or personal prompts.
Main Dependent Variables.
The studies measured beliefs about emotional sharing, partner preferences, interaction quality, future choices, and typical sharing behavior with AI and humans. Measures used pre/post comparisons, partner-choice tasks, and standardized rating scales across Studies 1–4.
- Beliefs of Emotional Sharing with Human and AI: Beliefs about emotional sharing were assessed across perceived judgment, emotional support, advice, and confidentiality using 14 items.Participants rated assigned partners after interaction, enabling comparison with pre-conversation beliefs; ratings used a 7-point scale from 1 (strongly disagree) to 7 (strongly agree).
- Partner Choice: In Studies 1–3, participants chose between an AI companion and another participant and rated their comfort or desire to talk with each.Study 1 used a 0–10 comfort scale, whereas Studies 2 and 3 used a 1–7 desire-to-talk scale.
- Experienced Interaction Quality: Studies 2 and 3 measured experienced interaction quality through satisfaction, enjoyment, and willingness to continue, each rated from 1 (Not at all) to 7 (Very much).These items followed the interaction with the assigned partner.
- Other Measures: Studies 1–3 measured typical emotional-sharing behavior with AI and other people using the same 0–10 scale from 0 (Not at all) to 10 (All the time).Participants separately reported how often they usually shared emotional experiences with AI companions and with other people.
Data availability
The study’s data availability includes preprocessed data but excludes participants’ emotional experiences and chat messages.
- Preprocessed data exclude participants’ emotional experiences and chat messages.