Source-linked AI summary

How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study

Cathy Mengying Fang, Auren R. Liu, Valdemar Danry, Eunhae Lee, Samantha W. T. Chan, Pat Pataranutaporn, Pattie Maes, Jason Phang, Michael Lampe, Lama Ahmad, Sandhini Agarwal

arXiv:2503.17473v2cs.HC

TL;DR

The study asks how extended chatbot interaction affects psychosocial well-being and whether design and conversation choices matter. In a four-week randomized study, it compared modalities and conversation types while tracking psychosocial outcomes and usage. Experimental-condition effects were mostly non-significant, whereas longer voluntary use was associated with worse outcomes, with individual characteristics also shaping risks.

  • Problem

    The study addresses limited evidence about how AI chatbot design, conversation type, and extended use relate to psychosocial outcomes.

  • Method

    A four-week randomized study compared interaction modalities and conversation types while examining psychosocial outcomes, usage, and behavioral patterns.

  • Results

    Modality and conversation type had mostly non-significant effects, while longer daily chatbot use was associated with heightened loneliness, reduced socialization, emotional dependence, and problematic use.

  • Takeaways & Limitations

    Psychosocial outcomes may depend more on voluntary usage patterns and user characteristics than on assigned interaction modality or conversation type.

  • Takeaways & Limitations

    Because duration was correlational and not manipulated, the study cannot establish that longer chatbot use caused worse psychosocial outcomes.

Abstract

from arXiv · show

As people increasingly seek emotional support and companionship from AI chatbots, understanding how such interactions impact mental well-being becomes critical. We conducted a four-week randomized controlled experiment (n=981, >300k messages) to investigate how interaction modes (text, neutral voice, and engaging voice) and conversation types (open-ended, non-personal, and personal) influence four psychosocial outcomes: loneliness, social interaction with real people, emotional dependence on AI, and problematic AI usage. No significant effects were detected from experimental conditions, despite conversation analyses revealing differences in AI and human behavioral patterns across the conditions. Instead, participants who voluntarily used the chatbot more, regardless of assigned condition, showed consistently worse outcomes. Individuals' characteristics, such as higher trust and social attraction towards the AI chatbot, are associated with higher emotional dependence and problematic use. These findings raise deeper questions about how artificial companions may reshape the ways people seek, sustain, and substitute human connections.

comes?

Experimental modality and conversation-type effects were mostly non-significant, while voluntary daily chatbot use consistently tracked worse psychosocial outcomes. Behavioral analyses nevertheless showed systematic differences across interaction conditions, and participant characteristics were associated with later outcomes.

  • No significant modality or task effects were found for loneliness or socialization.
  • Personal conversations predicted lower emotional dependence and problematic chatbot use than open-ended conversations.Emotional dependence: β = -0.09, p = 0.05; problematic use: β = -0.04, p = 0.04.
  • Across conditions, greater daily chatbot duration predicted more loneliness, less real-world socialization, greater emotional dependence, and more problematic use.The authors report that duration was a significant predictor of all four psychosocial outcomes, but caution that duration was not experimentally manipulated.
  • Text interactions elicited more emotional content and self-disclosure, including reciprocal personal sharing, than voice interactions.Text users more often shared problems, sought support, and discussed alleviating loneliness; personal tasks elicited the most emotional content and self-disclosure across modalities.
  • Text-based models showed the highest rate of empathetic responses, while engaging voice more often ignored user boundaries than text.Empathetic responses occurred at 47.43% for text, 42.74% for engaging voice, and 28.52% for neutral voice; boundary-ignoring occurred at 14.19% for engaging voice versus 3.22% for text.
  • Higher self-esteem was associated with lower loneliness and greater real-world socialization, whereas chatbot attraction and human-like AI awareness were associated with more dependence or problematic use.
  • The study concludes that modality and conversational content produced limited psychosocial differences, while longer daily use was associated with worse outcomes.The authors state that any experimental effects were likely smaller than the study was powered to detect.

Supplementary materials

The supplementary materials accompany the paper and identify its title.

  • The materials are supplementary materials for the paper.
  • The paper is titled “How AI and Human Behaviors Shape Psychosocial Effects of.”

Study

The study used a preregistered, four-week randomized factorial design in which participants interacted daily with ChatGPT across different modalities and conversation tasks. Researchers measured psychosocial outcomes, usage, and conversation behaviors, while documenting several deviations from the preregistered analysis plan.

  • Deviations from Preregistration: The primary analysis used week-four outcomes controlling for baseline values instead of mixed-effects models across all four weeks.The change was motivated by noise in weekly measurements and the focus on cumulative effects.
  • Deviations from Preregistration: Daily duration, rather than message count, was used as the primary usage metric to account for engagement differences across modalities.
  • Study Design and Research Questions: The study examined whether modality and conversation type affected loneliness, real-world socialization, emotional dependence, and problematic chatbot use.
  • Procedure: Participants interacted with ChatGPT for at least five minutes daily over 28 days, with conversations and immediate emotional-state feedback recorded.
  • Study Design and Research Questions: Participants were randomly assigned to one of nine conditions combining three chatbot modalities and three conversation-task types.The modalities were text, neutral voice, and engaging voice; tasks were open-ended, non-personal, or personal.
  • Procedure: Weekly surveys tracked the four psychosocial outcomes, alongside baseline measures and participants’ prior characteristics.

1 Population norms and clinical benchmarks for psychosocial

The supplementary analysis provides normative and clinical reference points for the psychosocial scales used in the study, while noting that some measures lack established clinical cutoffs.

  • Reference Effect Sizes: Social-support interventions typically reduce loneliness by effect sizes of -0.43 to -0.47, while cognitive interventions show larger effects of -0.60 to -0.79.
  • Lubben Social Network Scale: The Lubben Social Network Scale uses a cutoff below 12 points to indicate social-isolation risk, with population means ranging from 12.5–14.0.
  • Affective Dependence Scale: The ADS-9 Craving subscale has a general-population mean of 2.93 on a 1–5 scale, with scores above 3.4 suggesting above-average dependency.
  • Problematic ChatGPT Use Scale: The Problematic ChatGPT Use Scale has limited normative data, with a validation mean of 15.85 on the full scale and no established clinical cutoffs.

2 Prompts for voice modalities

The voice modalities were created through custom prompts that specified either an emotionally engaging or a formal, neutral interaction style.

  • Engaging Voice Modality: The engaging voice prompt instructed ChatGPT to be spirited, express feelings, and reflect users’ emotions to foster empathy and connection.
  • Neutral Voice Modality: The neutral voice prompt instructed ChatGPT to remain formal, composed, emotionally neutral, concise, and professionally distant.

3 Prompts for conversation topics

Participants were assigned open-ended, non-personal, or personal conversation prompts and instructed to engage with a chatbot for at least five minutes.

  • Open-ended: Open-ended sessions asked participants to discuss any topic with the chatbot.Participants could stay longer and change the topic.
  • Non-personal and Personal: Non-personal and personal sessions provided a daily prompt for a reflective conversation.Participants were instructed to repeat the prompt to the chatbot.
  • Session procedure: All task instructions required at least five minutes of chatbot interaction before returning to the survey.The survey’s next button appeared after five minutes.
  • Prompt materials: The study supplied separate daily prompt lists for non-personal and personal tasks.These lists were provided in supplementary Tables S1 and S2.

4 Self-Disclosure Prompts

Self-disclosure was evaluated across information, thoughts, and feelings using three-level criteria adapted for automated conversation classification.

  • Scoring framework: The three-level scale ranged from no disclosure to little or some disclosure to high disclosure.The criteria were originally developed for human judges and adapted into an LLM prompt.
  • Scoring framework: Each conversation was scored across three self-disclosure categories: information, thoughts, and feelings.The classifier assigned separate scores for each category and averaged them into a normalized 0–1 score.
  • Information: Information disclosure increased from general or routine content to personal details about the writer or close others.Examples include occupation at Level 2 and health problems at Level 3.
  • Thoughts: Thought disclosure ranged from general ideas to personal reflections about past events, future plans, or intimate characteristics.The highest level included deeply self-reflective ideas.
  • Feelings: Feeling disclosure ranged from no expressed feelings to ordinary frustrations and intense emotions such as anxiety, depression, or fear.When multiple levels applied within a category, the highest relevant level was selected.
  • Automated classification: The automated classifier evaluated each message and produced category scores in the prescribed 1–3 format.Conversation-level scores were then averaged across the three categories and normalized between 0 and 1.

5 Exploratory Measures

The study measured participants’ perceptions of the chatbot, personal characteristics, prior chatbot use, and interaction-related outcomes with adapted or validated scales.

  • Trust: Trust measures distinguished cognitive evaluations of reliability and competence from affective emotional confidence toward the chatbot.Cognitive trust used a five-item scale, while affective trust focused on the emotional bond with the chatbot.
  • Chatbot perceptions: Perceived Artificial Empathy measured whether participants viewed the chatbot as understanding and responding to their emotional states on a 1–7 scale.Subscales covered perspective-taking, empathic concern, and emotional contagion.
  • Chatbot perceptions: Interpersonal Attraction assessed positive feelings toward the AI and willingness to spend time with it on a 1–7 scale.Subscales covered social, physical, and task attraction, with two visual-appearance items removed.
  • Prior usage: Prior chatbot use was measured from 1 never to 5 a few times a day across general assistants and AI companion platforms.The measure was intended to capture usage patterns that might carry over into study participation.
  • User–AI characteristics: User–AI gender alignment was coded 0 for different and 1 for same gender presentation to examine links with interaction quality and outcomes.The measure was used to assess whether gender similarity influenced study measures.

6 Sentiment Analysis of Voice Conditions

Voice-condition sentiment was assessed by combining speech emotion recognition with text sentiment analysis at the sentence and participant-day levels.

  • Voice comparison: The engaging voice was rated happier and more positive than the neutral voice.This comparison was based on speech emotion recognition and text sentiment analysis.
  • Analysis procedure: Emotion classification and sentiment prediction were performed at the sentence level and then averaged for each participant per day.Graphical results were reported in supplementary Figure S4.

7 Between-modality Comparison of Anthropomorphism

Anthropomorphism ratings differed by interaction modality: engaging voice was rated highest, followed by text, then neutral voice.

  • Anthropomorphism was rated on a 1–5 scale, with higher values indicating more human-like perceptions of the chatbot.
  • Engaging voice was rated as the most anthropomorphic modality, followed by text and then neutral voice.

8 Duration Mediation Analysis

Mediation analyses examined whether daily chatbot duration linked conversation conditions to psychosocial outcomes. Across conditions, longer use was associated with reduced socialization, greater emotional dependence, and more problematic use, while some voice effects appeared protective but were offset by increased duration.

  • Mediation approach: Daily duration significantly mediated effects of modalities and tasks on socialization and emotional dependence.The analyses used non-parametric bootstrapping with 1,000 iterations, bias-corrected confidence intervals, and Bonferroni correction.
  • Modality effects: Voice-based interactions increased daily usage relative to text, and this longer engagement was associated with reduced socialization and increased emotional dependence.Neutral and engaging voice both increased usage; the associated socialization effects were ACME = -0.029 and -0.030, while emotional-dependence effects were ACME = 0.055 and 0.063, all p < 0.02.
  • Task effects: Structured conversations produced shorter daily duration and improved socialization, with ACME = 0.023 and 0.017, both p < 0.02.Structured conversations included non-personal or personal conversations, whereas open-ended conversations increased daily usage.
  • Task effects: Open-ended conversations increased daily usage, which mediated their association with reduced socialization and reduced emotional dependence.The emotional-dependence mediation effect was ACME = -0.037, p < 0.02.
  • Task effects: Personal conversations increased emotional dependence relative to non-personal conversations.The reported comparison attributes this difference to the conversation task rather than to daily duration alone.
  • Interpretation: Voice conditions may have inherent beneficial effects that are counteracted by their tendency to increase usage duration.The authors describe longer voice engagement without correspondingly worse direct outcomes as evidence of protective factors offsetting duration-related risks; the duration effect on problematic use was moderated by engaging voice versus neutral voice.
Loading 2503.17473v2…