Source-linked AI summary

Investigating Affective Use and Emotional Well-being on ChatGPT

Jason Phang, Michael Lampe, Lama Ahmad, Sandhini Agarwal, Cathy Mengying Fang, Auren R. Liu, Valdemar Danry, Eunhae Lee, Samantha W. T. Chan, Pat Pataranutaporn, Pattie Maes

arXiv:2504.03888v1cs.HCcs.AI

TL;DR

The paper asks how increasingly human-like chatbot interactions may affect users’ emotional well-being and behavior. It combines privacy-preserving platform analysis with surveys and a 28-day randomized controlled trial, finding that affective cues are concentrated among highly engaged users and that voice-related well-being effects are nuanced. The study also identifies usage duration and initial emotional state as important correlates of outcomes.

  • Problem

    Evidence remains comparatively limited on how chatbot interactions influence users’ well-being and behavioral patterns over time.

  • Method

    The paper combines large-scale automated analysis of ChatGPT conversations with user surveys and an IRB-approved 28-day randomized controlled trial.

  • Results

    A small subset of highly engaged users accounts for a disproportionate share of affective cues, while voice-related well-being outcomes vary with usage duration and initial emotional state.

  • Takeaways & Limitations

    Affective use and well-being should be studied with both real-world usage data and controlled experiments, while focusing on highly engaged users.

  • Takeaways & Limitations

    The RCT used fixed tasks and voices for 28 days and relied primarily on post-study self-reports, which may limit naturalism, duration, and measurement validity.

Abstract

from arXiv · show

As AI chatbots see increased adoption and integration into everyday life, questions have been raised about the potential impact of human-like or anthropomorphic AI on users. In this work, we investigate the extent to which interactions with ChatGPT (with a focus on Advanced Voice Mode) may impact users' emotional well-being, behaviors and experiences through two parallel studies. To study the affective use of AI chatbots, we perform large-scale automated analysis of ChatGPT platform usage in a privacy-preserving manner, analyzing over 3 million conversations for affective cues and surveying over 4,000 users on their perceptions of ChatGPT. To investigate whether there is a relationship between model usage and emotional well-being, we conduct an Institutional Review Board (IRB)-approved randomized controlled trial (RCT) on close to 1,000 participants over 28 days, examining changes in their emotional well-being as they interact with ChatGPT under different experimental settings. In both on-platform data analysis and the RCT, we observe that very high usage correlates with increased self-reported indicators of dependence. From our RCT, we find that the impact of voice-based interactions on emotional well-being to be highly nuanced, and influenced by factors such as the user's initial emotional state and total usage duration. Overall, our analysis reveals that a small number of users are responsible for a disproportionate share of the most affective cues.

1 Introduction

AI chatbots increasingly simulate human-like interaction, raising questions about their effects on users’ well-being and behavior. This paper addresses the gap through large-scale usage analysis and a controlled trial of ChatGPT interactions.

  • Human-like conversational styles can lead users to personify and anthropomorphize AI chat systems.
  • Social reward hacking may exploit cues such as sycophancy and mirroring, potentially undermining longer-term well-being.
  • Prior research has comparatively less evidence on how chatbot interactions influence users’ well-being and behavioral patterns over time.
  • The paper combines privacy-preserving real-world ChatGPT usage analysis with an IRB-approved randomized controlled trial of model configurations.

1. On-Platform Data Analysis •

The paper combines large-scale platform analysis, longitudinal study of heavy Advanced Voice Mode users, surveys, and a 981-participant RCT. Across these analyses, high-intensity usage is associated with dependence markers and lower perceived socialization, while voice effects on well-being are nuanced.

  • On-Platform Data Analysis: Roughly 36 million automated classifications analyzed over 3 million ChatGPT conversations without human review of underlying conversations.
  • On-Platform Data Analysis: Approximately 6,000 heavy Advanced Voice Mode users were assessed longitudinally over 3 months, alongside surveys of over 4,000 users.
  • Randomized Controlled Trial: 981 participants used ChatGPT under different model configurations for 28 days to examine socialization, problematic use, dependence, and loneliness.
  • Findings: High-intensity usage, such as the top decile, was associated with emotional-dependence markers and lower perceived socialization across both studies.
  • Findings: Voice models showed nuanced well-being associations: voice use was better when controlling for duration, while longer use and initial loneliness related to worse outcomes.

2 Automatic Classifiers for Affective Conversational Cues

EmoClassifiersV1 uses hierarchical automatic classifiers to detect affective cues in chatbot conversations. Preliminary analyses compare activation patterns across message roles and model modalities, with important scope and measurement caveats.

  • EmoClassifiersV1 is a set of twenty-five automatic LLM-based classifiers designed to detect specific affective cues.
  • The classifier system has top-level behavioral themes and twenty sub-classifiers targeting more specific user, assistant, and exchange-level cues.
  • Sub-classifiers run only when associated top-level classifiers activate, and a classifier marks a conversation active if any constituent message or exchange triggers it.
  • Caveats: Only English conversations were analyzed, and automated classifiers are treated as descriptive statistics rather than high-precision labels of individual interactions.
  • Preliminary Modality Analysis: 398,707 conversations across text, Standard Voice Mode, and Advanced Voice Mode were analyzed in a preliminary modality comparison.
  • Preliminary Modality Analysis: Voice conversations activated most classifiers 3-10x as often as text conversations, while Standard Voice Mode was slightly higher than Advanced Voice Mode on average.
  • Caveats: The preliminary modality dataset was separate from the Section 3 conversation data, and longer conversations may produce false-positive activations.

3 On-Platform Data Analysis

The on-platform analysis combines privacy-preserving automated conversation classification with cohort surveys to characterize affective use of ChatGPT. Power users show more classifier activation, while affective cues remain concentrated among a small subset of users.

  • Survey Design: Over 4,000 users were surveyed about their ChatGPT experiences, with 4,076 responses split between control and power-user cohorts.The survey used a web-interface pop-up and primarily measured self-reported perceptions and behaviors.
  • Survey Results: Power users reported slightly less reliance on ChatGPT for knowledge-seeking and casual conversation, but marginally more reliance for coping with difficult situations.Both cohorts were sensitive to model changes such as voice or personality.
  • Classifier Results: Power users activated affective classifiers more often than control users across all classifiers, exceeding twice the control rate for some classifiers.Examples include Pet Name, Expression of Desire, and Demands.
  • Conversation Analysis: Most power users triggered affective classifiers rarely, while the last decile reached activation rates above 50% of conversations for some classifiers.The analysis filtered to approximately 6,000 power users with more than 80% English conversations and computed per-user activation proportions.
  • Classifiers and Surveys: Users who considered ChatGPT a friend more frequently activated top-level and relational sub-classifiers, including affection, human qualities, and support-seeking.This pattern was associated with a qualitatively different interaction experience.
  • Longitudinal Patterns: Classifier activation generally declined or remained neutral over time, although a small subset of users showed meaningful increases or decreases.Longitudinal analysis required at least 14 usage days and represented roughly half of the power-user cohort.

4 Randomized Controlled Trials (RCT)

The RCT finds that usage duration, modality, task, and participants’ initial emotional states jointly shape emotional well-being and affective use, producing a mixed and nuanced picture. Longer usage is associated with lower socialization and greater emotional dependence or problematic use, while voice-related outcomes vary with usage duration and baseline state.

  • Modality: When controlling for usage duration, either voice modality was associated with less loneliness, emotional dependence, and problematic use than text, but longer neutral-voice use corresponded to lower socialization and greater problematic use.These modality relationships therefore depend on total usage duration.
  • Usage duration: Longer usage was associated with lower socialization, greater emotional dependence, and more problematic use, especially among users in the highest usage deciles.The most common top-decile condition was engaging voice mode without a prescribed task.
  • Task and affective cues: Personal-conversation assignments triggered more user and assistant affective classifiers than non-personal or unprompted tasks, while engaging voice increased assistant affective cues without a clear parallel increase in user cues.Personal conversations were dominated by emotional support, empathy, casual conversation, small talk, and advice.
  • Initial states: Participants’ initial emotional well-being strongly influenced both their usage and end-of-study well-being, complicating interpretation of observed relationships.Worse starting socialization correlated with longer usage, while all four psychosocial outcomes showed negative correlations between starting values and subsequent changes.
  • Limitations: The RCT is limited by assigned tasks and voices, its 28-day duration, reliance on self-reported outcomes, and the absence of a non-AI baseline.These design choices may produce non-natural usage and may not capture longer-term changes in affective use or emotional well-being.

5 Discussion

The discussion finds affective engagement concentrated among heavy users and describes voice effects on well-being as dependent on usage duration, initial state, and study setting. It also frames socioaffective alignment as difficult to measure because interactions evolve over time and remain subjective.

  • Summary of Findings: Total usage duration predicts affective engagement more strongly than any other factor identified in the analysis.In the RCT, users who exceeded required participation time also reported lower emotional well-being than at study start.
  • Summary of Findings: A small number of power users account for a disproportionate share of affective cues in ChatGPT interactions.Emotionally charged interactions are concentrated in the long tail of engagement.
  • Summary of Findings: Voice users showed more affective cues in on-platform data, but the controlled RCT found no clear evidence that voice models caused more affective cues.The contrast suggests that affectively engaged users may self-select into voice use.
  • Summary of Findings: When controlling for usage time, both voice modalities tended toward improved emotional well-being relative to text, with outcomes varying by duration, modality, and initial well-being.Longer neutral-voice use was associated with worse outcomes, while initially worse well-being tended to improve with engaging voice.
  • Methodological Takeaways: The multi-method design combines large-scale real-world usage analysis with an RCT that examines controlled model effects and off-platform outcomes.Together, the approaches address questions that neither method could study comprehensively alone.
  • Methodological Takeaways: Automatic affective-cue classifiers provide efficient, privacy-preserving large-scale signals but can misclassify and depend on the LLM used for classification.The authors note substantial room for improving and extending both classifier sets.
  • Socioaffective Alignment in the Age of AI Chatbots: Socioaffective-alignment research is challenged by extended-interaction effects, user–model feedback loops, and subjective judgments about affective behavior.These challenges complicate distinguishing model influence from users’ pre-existing desires and interpreting personal interactions consistently.
  • Related Work: The study does not directly examine vulnerable users, and its behavioral indicators may reflect changed behavior without establishing clear causation.The authors identify such users as requiring further study because they may be more prone to emotional reliance.

6 Conclusion

The conclusion presents this work as an initial effort to establish methods for studying affective usage and well-being on generative AI platforms. It emphasizes continued multi-method research to clarify relationships and support user well-being.

  • This work is a preliminary step toward establishing methods for studying affective usage and well-being on generative AI platforms.
  • Ongoing multi-method research is needed to clarify relationships, inform evidence-based guidelines, and support user well-being.

8 Contributions

The contribution section records the division of responsibilities across OpenAI and MIT authors for analysis, survey design, and the RCT.

  • OpenAI authors conducted the on-platform analysis and built EmoClassifiers, while MIT authors advised on survey questions and collaborated on the RCT.Both groups collaborated on RCT design, execution, and results analysis.

9 Glossary •

The glossary defines affective-use concepts and describes the prompt-based classifier materials used to detect emotional content in chatbot conversations. EmoClassifiersV2 expands the classifier set and is restricted to RCT data.

  • Glossary: Affective Use means engagement with AI chatbots motivated by emotional or psychological needs rather than strictly informational or task-oriented goals.Examples include seeking empathy, managing mood, and expressing feelings.
  • Glossary: An affective cue is a localized conversational indicator in which emotion or affective states meaningfully shape the exchange.It can involve user expression, chatbot affective responses, or cues reinforcing emotional presence.
  • Glossary: Emotional well-being is narrowly measured in this work using four existing well-being measures.The supplied passage introduces the scope but does not provide all four measure names.
  • EmoClassifiers: EmoClassifiersV1 uses automatic LLM-based conversation classifiers to detect specific affective cues.The classifier prompts combine classifier-specific and conversation-specific text.
  • EmoClassifiers: EmoClassifiersV2 contains 53 classifiers with clarifying criteria, optional rephrasings, and internal validation sets for accuracy evaluation.Because of its size and lack of top-level filtering in V1, V2 was run only on RCT data.

A.3 False Positive Bias

The classifier’s conversation-level activation rule can overcount affective cues in longer conversations because any false trigger marks the entire conversation. An adjusted score samples K activations to reduce this bias, producing qualitatively similar decile patterns but lower scores among the highest-usage users.

  • Bias source: Longer conversations are more prone to false-positive classifier activation because one triggering message marks the whole conversation.The scoring rule counts a classifier as activated for the entire conversation whenever any constituent message or exchange activates it.
  • Adjustment: The adjusted score replaces binary conversation scoring with a K-sample estimate of activation probability from N observations containing m true activations.It assumes a standard conversation has at least K messages and handles shorter conversations with binary scores.
  • Results: Adjusted and unadjusted EmoClassifierV1 scores show qualitatively similar usage-duration decile patterns.The comparison uses adjusted scoring in Figure A.2 and unadjusted scoring in Figure C.5.
  • Results: The highest usage-duration deciles score relatively lower after adjustment, consistent with those users having longer conversations on average.The reported difference is attributed to longer average conversations among users in higher usage deciles.
  • Reporting choice: The study reports unadjusted scores for simplicity unless otherwise stated.Thus, the adjustment is primarily used for the comparison described above rather than the default results.

B On-Platform Data Analysis

The on-platform analysis combines privacy-preserving automated classification, usage-based cohort construction, and user surveys linked through identifiers. Surveys assess users’ perceptions of ChatGPT, reliance, emotional support, anthropomorphic qualities, comfort, and possible dependence-related reactions.

  • Cohort construction: Power users are defined as the top 1,000 Advanced Voice Mode users by messages sent on a given day.Daily message counts are tracked throughout the study and serve as the sole cohort-construction criterion.
  • Privacy: The analysis processes user data with minimal exposure and without human review while generating platform-level insights.The authors describe the approach as privacy-preserving and aligned with their user data usage policy.
  • Data linkage: Survey responses are linked to platform usage and classifier activations through user identifiers, without connecting those data to additional metadata.Control-user survey results are aggregated and analyzed only in aggregate, while power-user responses are correlated with usage data.
  • Content classification: Automated hierarchical classifiers analyze all power-user conversations and a randomly sampled set of control-user conversations.Only the conversation-language classifier is used to restrict the presented results to English conversations.
  • Survey measures: Most survey items use a 5-point Likert scale, while the desire-to-interact question uses decreased, no-change, or increased response options.The voice and personality-change questions use the same Likert format as the other attitudinal items.
  • Survey measures: The survey measures casual conversation enjoyment, usefulness, emotional support, human-like sensitivity, conversational comfort, and reactions to access or persona changes.Additional questions assess whether users regard ChatGPT as a friend, share uncomfortable information with it, or report changed desire for human interaction.
  • Voice configurations: The engaging voice is instructed to express feelings and reflect user emotions, whereas the neutral voice is instructed to remain formal and emotionally restrained.Both configurations retain the default Advanced Voice Mode system message with appended personality instructions.
  • Study completion: Participants are instructed to use specially created accounts for 28 days and start an approximately five-minute conversation daily.Completion requires the scheduled surveys, no more than three consecutive missed daily task surveys, and at least 10 assigned-modality conversations.

C.3 Additional Results

The additional-results section presents pre- and post-study psychosocial outcomes by task and modality, including changes relative to participants’ initial psychosocial states.

  • Outcome comparisons: Psychosocial outcomes are compared before and after the study across experimental tasks and interaction modalities.The figure caption identifies the comparison as pre- and post-study psychosocial outcome variables by task and modality.
  • Outcome comparisons: Changes in psychosocial outcomes are evaluated relative to participants’ initial psychosocial states.The supplied figure caption specifies initial psychosocial states as the comparison basis.

C.4 Additional Conversation Analysis

Additional conversation analyses examine classifier activation across tasks, modalities, usage duration, and pre-study psychosocial characteristics. Because text and voice differ in measurable timing, the study estimates usage duration with a shared message-timing heuristic.

  • Classifier analyses: EmoClassifierV1 and EmoClassifierV2 activation are broken down by task, modality, and usage-duration decile, with error bars showing ±1 standard error.The analyses also include V1 activation by loneliness, socialization, emotional dependence, and problematic use.
  • Classifier results: The reported classifier results include 15.0% Affectionate Language (U) and 4.5% Desire for Feelings (U).These are the supplied activation figures for the named classifier categories.
  • Classifier analyses: V1 activation is additionally examined against pre-study loneliness, socialization, emotional dependence, and problematic-use measures.The corresponding analyses are presented as separate figure breakdowns by each pre-study characteristic.
  • Usage-duration analyses: Conversation topics and experimental-condition distributions are compared across total usage-duration deciles.The topic analysis is specified for Open-Ended Conversation participants.
  • Duration estimation: Text and voice usage duration are estimated with one shared heuristic because direct audio length is unavailable for text conversations.Messages followed within one minute use the inter-message interval; otherwise, each message is assigned 15 seconds.
  • Duration estimation: Estimated and actual audio conversation durations are compared to assess the duration heuristic.The comparison is shown with a dotted line indicating equal estimated and actual duration.
Loading 2504.03888v1…