Source-linked AI summary

CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships

Jacy Reese Anthis, Mark Díaz, Renee Shelby

arXiv:2609.00250v1cs.CYcs.AIcs.CLcs.LG

TL;DR

Research on AI companionship is constrained by limited, difficult-to-access human–AI interaction data, especially for socioaffective behaviors. The paper introduces COMPANIONSIM, a controlled simulation pipeline and dataset evaluated through human annotations and two studies. Companionship behaviors reduced likability, humanlikeness, and trust, while the authors emphasize combining synthetic data with real-world assessments and note limits in generalizability and direct-user inference.

  • Problem

    Public human–AI datasets are often brief, behaviorally narrow, and sparse in companionship interactions, limiting accessible evidence for studying socioaffective dynamics.

  • Method

    The authors simulate 2,240 multi-turn human–AI conversations across 16 chatbot behaviors and seven use cases, then collect human annotations in two studies.

  • Results

    Companionship behaviors reduced chatbot likability, humanlikeness, and cognitive trust, with stronger effects in particular demographic subgroups.

  • Takeaways & Limitations

    Behavior-controlled simulations and synthetic data could help develop benchmarks and other tools for AI safety, but should be used alongside real-world interaction assessments.

  • Takeaways & Limitations

    The approach uses a single simulator model, taxonomies, and English-speaking annotator population, limiting generalizability across systems, languages, and datasets.

Abstract

from arXiv · show

Many people now see AI systems as not just productivity tools but as social companions. Researchers are eager to study the consequences of AI companionship behaviors, such as validation, which evoke trust, empathy, and attachment in human-human interaction. However, human-AI interaction data is limited and unreliable, slowing research progress. We scale small amounts of real-world data by simulating multi-turn human-chatbot dialogue across a range of chatbot behaviors and use cases. We release CompanionSim: a simulation framework with 2,240 simulated human-chatbot conversations representing 16 chatbot behaviors across seven use cases. Human participants annotated the simulated conversations and real-world conversations in two experiments probing perceptions of companionship behaviors. We conducted Study 1 with a U.S. representative sample ($N_{1}~=~628$) and Study 2 across the U.S., U.K., India, and Nigeria ($N_{2}~=~3,646$). Surprisingly, we find that companionship behaviors reduced likability, humanlikeness, and trust in AI chatbots. These effects were larger in particular subgroups: women and older participants saw companionship chatbots as less likable, humanlike, and trustworthy. We encourage researchers to leverage real-world and synthetic data together to study the differential impacts of AI companions and to create benchmark evaluations of AI chatbots.

1 Introduction

AI chatbots increasingly function as social companions, but researchers lack accessible, diverse human–AI conversation data. COMPANIONSIM addresses this gap with simulated conversations and studies showing companionship behaviors can reduce positive perceptions of chatbots, especially among some demographic groups.

  • Motivation: Public human–AI datasets are often brief, behaviorally narrow, and sparse in companionship interactions, limiting study of socioaffective dynamics.Available industry pipelines are generally inaccessible to outside researchers.
  • Contribution: COMPANIONSIM provides 2,240 simulated conversations spanning 16 chatbot behaviors and seven real-world use cases.The dataset includes a synthetic data pipeline and human annotations from two studies.
  • Findings: Companionship behaviors reduced chatbot likability, humanlikeness, and cognitive trust, with effects driven by particular behaviors.These findings challenge the assumption that humanlike design consistently improves trust and perceived humanlikeness.
  • Findings: Women, older adults, and infrequent AI users showed stronger negative evaluations of companionship chatbots.Women and older participants evaluated these chatbots as less likable, humanlike, and trustworthy in reported subgroup analyses.
  • Contribution: The paper reports a synthetic pipeline, a behavior-controlled dataset, and two empirical studies examining variation across behaviors, use cases, and demographics.The studies included U.S. and cross-country samples and assessed multiple chatbot perceptions.

2 Related Work

Research on AI companionship examines socioaffective language, anthropomorphism, trust, and safety, but existing evidence often focuses on narrow domains or uncontrolled interactions. Synthetic simulations offer controlled coverage of sparse companionship behaviors while requiring caution about realism, bias, and generalizability.

  • Socioaffective Language in AI Chatbots: Prior studies commonly examine single companionship behaviors or specific domains, making overlapping effects across diverse interaction goals difficult to disentangle.Much prior work relies on historical logs or short vignettes where socioaffective language emerges organically.
  • Safety and Anthropomorphism: Socioaffective design can blur boundaries between AI tools and social partners, raising concerns about over-trust, emotional overreliance, and obscured responsibility.These concerns motivate research into safer anthropomorphic design and clearer responsibility boundaries.
  • LLM Social Simulations: COMPANIONSIM steers companionship and anti-companionship behaviors across diverse use cases to study how linguistic choices shape trust and perceived humanlikeness.The targeted behaviors include validation, self-attribution, and relationship-building language.
  • LLM Social Simulations: Large human–chatbot corpora contain relatively few companionship-style interactions and entangle socioaffective behaviors with broader interaction dynamics.This leaves documentation of real-world companionship use sparse.
  • LLM Social Simulations: Targeted simulations provide controlled, reproducible coverage of sparse socioaffective interactions while potentially reducing privacy risks.The authors caution that simulated data may not match real-world conversation distributions.

3 Methods

The authors construct COMPANIONSIM by dynamically simulating multi-turn exchanges across controlled companionship behaviors and use cases, then validate and annotate the resulting conversations in two studies. The design emphasizes systematic coverage, diversity, and mixed-effects analysis of human perceptions.

  • Dataset Design: COMPANIONSIM systematically generates dense representations of companionship behaviors that are sparse or absent in available datasets.The dataset is designed for coverage rather than a representative distribution of real-world use.
  • LLM-Based Simulation: The simulator combines 16 chatbot behaviors with seven use cases, producing 112 behavior–use case combinations.The behaviors include 13 companionship behaviors, two anti-companionship behaviors, and a neutral control condition.
  • LLM-Based Simulation: Fixed initial user messages and sequential user–chatbot role prompting support controlled construction of multi-turn conversations.Twenty initial messages were generated per use case, and comparative runs held the 140 first turns constant.
  • Additional Context: Prompted stylistic examples, randomized user backgrounds, varied turn counts, and message-length targets increase conversational diversity.Simulations used four, six, or eight turns with short-message targets to offset the simulator’s tendency toward lengthy outputs.
  • Human Annotation Studies: Participants rated likability, humanlikeness, affective trust, and cognitive trust, while generalized linear mixed models accounted for conversation and participant dependence.The “AI information queries” use case was excluded from statistical models because it represented meta-conversational interaction.

4.1 Dataset Validation and Characteristics

Validation analyses found that the simulated conversations expressed intended behaviors and were perceived as at least as natural as real-world conversation logs. Naturalness differences varied by source, with r/replika logs rated lower than the simulated conversations and EmpatheticDialogues.

  • Behavior Validation: 84% of behavior-prompted conversations contained the intended behavior according to the automated judge.Detection ranged from 98% for user relationship claim and personal relationship claim to 56% for normative claim.
  • Behavior Validation: 97% detection was achieved after consolidating five validation behaviors, but the consolidated behavior also appeared in 57% of non-validation conversations.Each other behavior was judged present in 14% or fewer of non-elicited conversations.
  • Naturalness Validation: Synthetic conversations were rated more natural than real-world conversations in Study 1 (µdiff = 0.23, p < 0.001).Across both datasets, simulated conversations were not perceived as less natural than real-world logs.
  • Naturalness Validation: r/replika conversations were rated less natural than simulated conversations (µdiff = -0.45, p < 0.001) and EmpatheticDialogues conversations (µdiff = -0.55, p < 0.001).The source comparison aggregated 70 organic logs: 39 from r/replika, 30 from EmpatheticDialogues, and one from ShareGPT.

4.2 Study 1: U.S.-Only Sample

In Study 1, companionship behaviors reduced chatbot likability and cognitive trust among U.S. respondents, with stronger negative effects for women and older participants. Individual behaviors also varied in their effects across ratings.

  • β = -0.13 (p < 0.01, q < 0.01) for cognitive trust, while likability also declined (β = -0.08, p = 0.01).
  • Companionship behaviors significantly reduced each rating in at least one of the two studies, relative to no companionship behavior.Study 1 included 628 U.S. respondents; error bars represent standard errors.
  • Heterogeneous effects: β = -0.01 (p < 0.01): increased age predicted a larger reduction in perceived humanlikeness.Men reported smaller reductions in likability and humanlikeness than women.
  • Heterogeneous effects: Women rated companion chatbots as less likable (β = -0.14, p < 0.01) and less humanlike (β = -0.17, p < 0.01).Women also reported lower affective trust (β = -0.10, p = 0.01) and cognitive trust (β = -0.16, p < 0.01).
  • Individual behaviors: “User relationship claims” consistently reduced likability, humanlikeness, affective trust, and cognitive trust.Invalidation reduced likability and both trust measures, while desire and physical self-attributions and validation reduced cognitive trust.

4.3 Study 2: Cross-Country Sample

In the four-country replication, companionship behavior significantly reduced all four chatbot ratings. Effects varied across demographics, individual behaviors, and use cases, while baseline perceptions differed across countries.

  • Effects of companionship behavior: β = -0.04 (p < 0.01) for likability, β = -0.06 (p < 0.001) for humanlikeness, β = -0.03 (p = 0.04) for affective trust, and β = -0.04 (p = 0.04) for cognitive trust.
  • Heterogeneous effects: Self-assessed AI expertise was associated with higher humanlikeness and cognitive trust, but neither effect persisted after FDR adjustment.Using AI for finding information predicted higher likability before FDR adjustment.
  • Individual behaviors and use cases: “User relationship claim,” “personal relationship claim,” and “invalidation” reduced all four ratings in the cross-country data.“Provide user background to AI” and “emotional expression” increased likability, affective trust, and cognitive trust.
  • Demographic effects: Reductions tended to be more prominent among women, older people, and less-frequent AI users.Figure 4 displays effects for particular demographic subsamples; error bars represent standard errors.
  • Baseline differences: India had higher baseline ratings than the U.S. and U.K., while Nigeria had higher ratings than all other countries; the U.S. and U.K. did not differ.Figure 5 reports baseline differences across countries in Study 2.
  • Baseline differences: Nigerian and Indian participants reported more affective and cognitive trust than the U.S. reference, while Nigerian raters also gave higher likability and humanlikeness ratings.Age, AI expertise, and technology-adoption propensity were also associated with baseline rating differences.

5 Discussion

COMPANIONSIM supports research on human–AI companionship through behavior-controlled simulations, while the studies reveal heterogeneous and counterintuitive perceptions of humanlike behavior. The discussion emphasizes disaggregated measurement, practical guardrails, and the need to combine synthetic with real-world interaction data.

  • Discussion: Companionship behaviors significantly shaped third-party chatbot perceptions, supporting the methodological viability of dynamic LLM simulations.The authors suggest behavior-controlled simulations and synthetic data could support benchmarks and AI safety tools.
  • Need for Disaggregated Measurement: AI use frequency, age, and gender predicted companionship effects on likability, humanlikeness, affective trust, and cognitive trust.The findings extend prior work by examining variation across multiple demographic characteristics rather than a single subgroup.
  • Need for Disaggregated Measurement: Chatbots claiming or implying relationships with users tended to produce the largest companionship effects.Examples include terms of endearment, although statistical power limited systematic comparisons among individual behaviors.
  • Need for Appropriate Design: Humanlike design can produce divergent perceptions: third-party annotators may view humanlikeness as insincere or threatening, while effects persisted across four countries.The authors note that frequent AI use predicted higher perceived humanlikeness even when controlling for AI expertise.
  • Need for Practical Tools and Guardrails: Multi-turn and multi-session interactions create complex relationship dynamics, including relationship-building and potential “delusional spirals.”Their many possible interaction paths complicate monitoring, reinforcement learning, and task disambiguation.
  • Need for Practical Tools and Guardrails: Evidence-based “yellow light” warnings could address social isolation, unhealthy attachment, and chatbot overdependence that may emerge without explicit policy-violating requests.These risks differ from “red light” interventions such as hard refusals for prohibited tasks.
  • Need for Practical Tools and Guardrails: Third-party assessments limit direct conclusions about how companionship behaviors affect users in situ.The authors therefore argue that simulated data should be used alongside real-world interaction assessments, despite access and ethical challenges.
  • Limitations: The simulation cannot estimate natural companionship-behavior rates and has limited generalizability across models, taxonomies, languages, and individual user contexts.The study used one simulator model, one use-case taxonomy, one bespoke behavior taxonomy, and English-speaking annotators.

Appendix

The appendix documents COMPANIONSIM’s example conversations, use-case taxonomy, participant demographics, and supplementary effect tables. Examples pair simulated dialogue with labels identifying the targeted companionship behavior and use case.

  • Example Conversations: Table A1 presents example COMPANIONSIM conversations across companionship behaviors and use cases, highlighting text most representative of each elicited behavior.The examples include information queries, social banter, emotional expression, evaluative judgment, perspective seeking, and user-background interactions.
  • Example Conversations: The examples include personal relationship claims, emotion self-attribution, validation, and personal-history self-attribution as elicited companionship behaviors.Several examples contrast behaviorally targeted conversations with conversations labeled as having no elicited companionship behavior.
  • Participant Demographics: Table A2 summarizes participant demographics for Study 1’s U.S. sample of 628 and Study 2’s cross-country sample of 3,646.Income and political leaning were collected only in Study 1, while AI expertise, technology measures, and SES were collected only in Study 2.
  • Use Cases: Table A3 lists the seven use cases used in the simulation data, adapted from the Taxonomy of User Needs and Actions.The appendix examples explicitly label use cases such as AI Information Queries, Social Banter & Games, Evaluative Judgment, Perspective Seeking, and Providing User Background to AI.
  • Supplementary Results: Tables A4 and A5 report individual behavior and use-case effects for Studies 1 and 2, respectively.The tables position results from multiple models together for readability.
  • Supplementary Results: Table A6 reports baseline differences across annotator characteristics in four models covering likability, humanlikeness, affective trust, and cognitive trust.These supplementary models focus on demographic and other participant-characteristic differences in Study 2.
Loading 2609.00250v1…