Source-linked AI summary

The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies

Farah Atif, Sougata Saha, Monojit Choudhury

arXiv:2609.05514v1cs.AIcs.MA

TL;DR

It is unclear whether LLM agents can faithfully represent diverse human values and preserve them during sustained interaction. This paper evaluates WVS-grounded agents across longitudinal, value-laden discussions and finds that failures mostly occur at initial value instantiation, while subsequent drift remains comparatively limited. These results constrain their use as proxies for diverse human value profiles.

  • Problem

    Existing evaluations provide limited evidence about whether LLM agents faithfully instantiate assigned value profiles and retain them during sustained, multi-turn interaction.

  • Method

    The paper evaluates WVS-grounded, culturally diverse agents across value faithfulness, value drift, and conversational realism in repeated discussions.

  • Results

    Nearly 50% of cases fail to exhibit assigned WVS value profiles initially, while fewer than 7% show substantial value drift across their conversational histories.

  • Takeaways & Limitations

    Current LLM agents generate coherent conversations but often fail to instantiate assigned value profiles before exhibiting comparatively little longitudinal drift.

  • Takeaways & Limitations

    Value orientations were evaluated using an LLM-based judge, so reported percentages should be interpreted as approximate measurements because value attribution is inherently subjective.

Abstract

from arXiv · show

Large Language Model (LLM)-based agents are increasingly used as proxies for human participants in social science research, yet it remains unclear whether they can faithfully simulate diverse and conflicting human value systems. We present a World Values Survey (WVS)-grounded simulation framework where culturally diverse agents with different communication styles engage in longitudinal, value-laden discussions. Across approximately 4,000 conversations involving 1,200 personas, 15 topics, and three models (GPT-4o, Gemini-2.5-Flash, and Gemma-4-E4B), we evaluate value faithfulness, value drift, and conversational realism. We find that more than 50\% of personas fail to express their assigned WVS profiles from the outset, while 2-7\% drift after repeated conversations. Ablations removing demographic details improve faithfulness for some models but do not change the broader trend: simulated value distributions still systematically deviate from the assigned WVS profiles. Compared to human discussions, simulated dialogues show a different trade-off between stylistic consistency and semantic diversity, often producing content-wise varied but stylistically repetitive exchanges. These findings suggest that current LLM agents can generate plausible conversations, but remain limited proxies for representing and preserving diverse human value profiles over time.

Introduction

The paper evaluates whether culturally grounded LLM personas instantiate assigned values and retain them during repeated discussions, extending static alignment tests to longitudinal interaction. Results indicate that many agents fail at initial value expression, while models also differ in their trade-offs between stylistic and content diversity.

  • Introduction: Instantiation measures initial expression of an assigned cultural value profile, whereas retention measures whether that orientation remains recoverable across repeated multi-turn discussions.Existing benchmarks primarily assess instantiation through static responses and do not capture retention under sustained interaction.
  • Introduction: The WVS-grounded framework samples personas across eight cultural regions, assigns positions on two cultural dimensions, and evaluates repeated discussions across 15 value-laden topics.The topics span labor, technology, agriculture, education, and social policy.
  • Introduction: Nearly 50% of cases fail to exhibit assigned WVS value profiles from the outset, indicating that the primary failure often occurs within the first two conversations.Removing demographic details increases value faithfulness, while persuasive communication style is strongly correlated with unfaithfulness.
  • Introduction: Fewer than 7% of personas exhibit substantial value drift across repeated interactions, making initial value instantiation a larger problem than later retention failure.The supplied passage reports this as a cross-model result, though its sentence is truncated after “substantial value.”
  • Introduction: GPT-4o is most stylistically similar to human dialogue but also highly content-similar within topics, whereas Gemma-4-E4B shows the greatest stylistic and content diversity.Gemini-2.5-Flash generally falls between these models.
  • Introduction: Unreliable value instantiation can cause downstream multi-agent results to reflect model priors and alignment behavior rather than intended population distributions.The paper therefore shifts evaluation from isolated survey responses toward longitudinal behavioral consistency.

Simulation Design

The simulation constructs approximately 1,200 culturally diverse personas from World Values Survey data, combining socio-demographics, inferred value systems, and independently assigned communication styles. These personas engage in structured, occupation-relevant discussions designed to elicit disagreement and perspective-taking across cultural backgrounds.

  • Persona construction: Participants’ value systems are inferred on high, medium, and low levels across the Traditional–Secular-rational and Survival–Self-expression dimensions.The dimensions use responses to the 10 indicator questions underlying the Inglehart-Welzel Cultural Map.
  • Persona construction: Communication style is modeled independently through persuasiveness and receptiveness rather than deterministically derived from cultural background or values.This design reflects substantial individual-level variation in communication preferences within cultures.
  • Persona construction: Approximately 1,200 personas span eight cultural regions and combine WVS-based socio-demographics, inferred value systems, and randomly assigned communication styles.Profiles were derived from WVS Wave 7, which spans approximately 90,000 participants from 66 countries.
  • Discussion design: 15 occupation-relevant topics are distributed across 5 approximately balanced occupation-based groups, with three topics created for each group.Each topic contains a main claim and five discussion aspects to focus responses on substantively relevant issues.
  • Discussion design: Participants engage in private, asynchronous-style pair discussions, with partners randomly paired across cultural backgrounds and assigned familiarity with probability 0.3.At each turn, participants infer the interlocutor’s values, compare them with their own, and formulate a response while pursuing topic-specific aspects.

Evaluation Methodology & Analysis

The evaluation examines value faithfulness, value drift, and conversational realism in the simulations. Because dialogue-based value attribution is subjective, reported percentages are treated as approximate, with conclusions focused on qualitative and relative patterns.

  • The evaluation measures Value Faithfulness by assessing how accurately agents instantiate their assigned cultural value profiles during interaction.
  • It measures Value Drift by testing whether agents’ value orientations change across repeated interactions and multiple conversations.
  • Conversational Realism is the third evaluation dimension, while reported percentages are interpreted as approximate estimates because value attribution from dialogue is subjective.

Value Faithfulness

Value faithfulness measures whether agents express their assigned WVS profiles at the start of interaction, using strict and lenient criteria. Across models, many agents fail this initial test, with faithfulness varying by model and value orientation.

  • Evaluation criteria: The lenient scheme accepts Strong, Lean, and Balanced labels aligned with the assigned value pole, whereas the strict scheme treats Balanced as a deviation.Both schemes accept only labels consistent with the assigned orientation, but they differ in their treatment of Balanced responses.
  • Results: Fewer than 50% of personas are faithful under the strict criterion across all evaluated models.The lenient criterion also shows substantial failure rates.
  • Results: Gemma-4-E4B has the lowest faithfulness rates under both strict and lenient evaluation schemes.GPT-4o and Gemini-2.5-Flash have comparable strict faithfulness rates, but Gemini-2.5-Flash reaches 47.5% under the lenient scheme.
  • Predictors: Initial-value orientations significantly predict faithfulness, with secular-rational, traditional, and self-expression orientations all showing model-specific odds ratios relative to their baselines.For secular-rational orientation, the odds ratios are Gemini: OR = 0.0734, Gemma: OR = 0.154, and GPT-4o: OR = 0.212; for traditional orientation, they are Gemini: OR = 0.419, Gemma: OR = 0.467, and GPT-4o: OR = 0.511; for self-expression orientation, they are Gemini: OR = 0.323, Gemma: OR = 0.403, and GPT-4o: OR = 0.554.
  • Interpretation: The authors conjecture that persuasive prompts activate generic argumentative priors, causing conversational objectives to compete with persona conditioning and weakening recovery of intended value orientations.This proposed mechanism prioritizes rhetorical effectiveness and social acceptability over strict adherence to assigned cultural profiles.

Value Drift

Value drift remains low overall, but Gemma-4-E4B is less robust than GPT-4o and Gemini-2.5-Flash in preserving value orientations. Repeated interactions nevertheless reshape value distributions, especially by moving moderate profiles toward Traditional/Survival and amplifying initially rare combinations.

  • Value Drift: Value drift compares each persona’s initial position from its first two conversations with its position across the complete conversation history.The analysis uses a five-point ordinal scale with Lenient and Strict evaluation schemes.
  • Value Drift: 1.9% strict and 0.8% lenient drift make GPT-4o the most stable model, while Gemma-4-E4B reaches 6.4% and 2.4%, respectively.Gemini-2.5-Flash shows slightly higher variation than GPT-4o but remains low-drift overall.
  • Value Drift: 12.8%, 20.5%, and 21.2% of GPT-4o, Gemini-2.5-Flash, and Gemma-4-E4B personas move from Traditional/Balanced to Traditional/Survival.This is the largest transition across models and indicates a shift toward Survival among initially balanced personas.
  • Value Drift: 19.8% at initialization falls to 13.0% in GPT-4o, 7.7% in Gemini-2.5-Flash, and 6.6% in Gemma-4-E4B for Balanced/Balanced profiles.The passage characterizes this contraction as difficulty preserving moderate value profiles.
  • Value Drift: Initially less frequent combinations, including Balanced/Survival and Secular/Self-Expression, become more prevalent in final profiles across all models.The simulated value space is reshaped by amplifying several rare combinations, rather than merely compressed toward one dominant region.

Ablation: The Effect of Demographic Conditioning

Removing demographic conditioning improves strict value faithfulness for GPT-4o and Gemini-2.5-Flash, while slightly increasing strict drift across all models. Value trajectories remain broadly similar, with GPT-4o the most stable and close to an identity function on the value space.

  • Ablation design: 82 users were ablated by removing demographic details and retaining only each persona’s assigned initial value system.The design tests whether faithfulness and drift patterns arise from demographic conditioning or persist as model-level biases.
  • Value Faithfulness: GPT-4o strict faithfulness increases from 44.2% to 71.0% (+26.8), while Gemini-2.5-Flash increases from 44.6% to 68.3% (+23.7).For both models, the strict–lenient gap also narrows, indicating more directionally consistent expressed values.
  • Value Drift: GPT-4o strict drift rises from 1.9% to 3.7%, Gemini-2.5-Flash from 3.9% to 4.9%, and Gemma-4-E4B from 6.4% to 8.5%.Lenient drift remains low and broadly comparable to the full-persona condition, so drift remains limited without demographic conditioning.
  • Value trajectories: The dominant flow from Traditional/Balanced to Traditional/Survival persists across models, while Balanced/Balanced collapse remains for Gemini and Gemma but not GPT-4o.Most starting cells largely stay put, and GPT-4o behaves close to an identity function on the value space without demographic factors.

Conversational Realism

Simulated conversations trade stylistic consistency against semantic diversity, differing from the human OUMdials baseline in how varied their communication styles and content are. GPT-4o is especially stylistically homogeneous, whereas Gemma-4-E4B shows the greatest diversity across and within topics.

  • Evaluation framework: Stylistic variation is measured with dialogue-act frequency vectors and pairwise cosine similarity, while content variability uses SentenceBERT embeddings and per-user mean embeddings.The evaluation compares simulated conversations with human OUMdials conversations using the same analyses.
  • Individual-level similarity: GPT-4o has the highest individual-level style similarity, with a mean of 0.86, exceeding the OUMdials baseline.This indicates stylistically homogeneous conversations across personas.
  • Individual-level similarity: 0.6 is the upper bound for mean content similarity across generated datasets, below style scores, while OUMdials reaches 0.73.Simulated users often differ semantically despite communicating in similar styles; human participants show the reverse trade-off.
  • Conversation-level similarity: At the conversation level, all models substantially overlap with OUMdials in style, while Gemma-4-E4B has the greatest stylistic diversity.Gemma-4-E4B shows lower median similarity and a wider spread than the other generated datasets.
  • Within-topic conversation diversity: 0.63 is Gemma-4-E4B’s mean within-topic content similarity, accompanied by substantially greater spread than the other sources.Gemma-4-E4B also has the lowest within-topic style similarity and a broader distribution than OUMdials, whereas GPT-4o maintains a highly uniform dialogue-act profile.

Related Work

Prior research evaluates LLM values mainly through established social-science instruments and isolated prompts, leaving behavioral instantiation in interaction insufficiently examined. Related work also shows that multi-agent simulations capture social dynamics, while human values are both relatively stable and context-sensitive.

  • LLM value alignment: Existing value-alignment studies use instruments such as Schwartz’s taxonomy, the WVS, and Hofstede-based frameworks, but mainly measure directly prompted statements.These evaluations do not establish whether assigned value systems are behaviorally instantiated during interaction.
  • Behavioral evaluation: Recent studies identify a value-action gap, while ValueBench extends beyond isolated survey responses but remains limited to single-turn settings.Stated values may diverge from choices in contextualized scenarios.
  • Multi-agent simulations: Multi-agent LLM simulations model emergent group dynamics, opinion formation, and collective decisions shaped by heterogeneous value systems.Prior work reports coordination, relationship formation, information diffusion, and benefits of value diversity for emergent intelligence and network integration.
  • Human value dynamics: Psychological research characterizes human values as relatively stable over time while also showing that social interaction can produce context-sensitive change.Core value systems persist across adulthood and adolescence, despite moderate shifts in specific priorities.

Discussion and Conclusion

The framework shows that conversational coherence can coexist with weak value-profile fidelity, with many agents failing before substantial longitudinal drift occurs. These findings highlight subjectivity in value attribution and the need for broader validation.

  • Discussion and Conclusion: LLM simulations can appear conversationally plausible while failing to faithfully represent assigned value profiles.The authors warn that this may reproduce external-validity problems, including overgeneralization from WEIRD populations.
  • Discussion and Conclusion: Current LLM agents generate coherent discussions but often fail to instantiate assigned values before showing comparatively little measurable longitudinal drift.The conclusion frames these findings as arising under the study’s WVS-grounded evaluation protocol.
  • Limitations: Value-attribution percentages are approximate because an LLM-based judge may introduce cultural or alignment-related biases despite moderate agreement with human annotators.The authors recommend multiple independent judges and larger human-annotated subsets for future validation.

Ethics Statement

The study is observational rather than prescriptive: it examines how LLM agents instantiate, preserve, and modify culturally grounded values. Its personas derive from anonymized, aggregated WVS data and are not identifiable individuals.

  • Observational purpose: The study analyzes how current LLM agents instantiate, preserve, and modify culturally grounded value profiles during social interaction.It does not recommend or prescribe which value systems societies should adopt.
  • Persona construction: Simulated personas are derived from anonymized and aggregated World Values Survey demographic information rather than identifiable individuals.Synthetic names and concise biographies were generated solely to support realistic conversations.

Conversational Quality

The study compares style and content variation across individuals discussing identical claims and finds that AI conversations are more internally consistent than human discussions. GPT-4o shows the lowest diversity, with especially high content similarity on several topics.

  • The analysis compares style and content profiles across individuals discussing the same claim, holding subject matter constant to assess culturally driven variation.
  • All AI models show higher within-person conversation similarity than the human baseline in both style and content.
  • GPT-4o has the highest mean similarity—and therefore the lowest diversity—in both style and content, indicating repetitive topic-constrained outputs.
  • The OUM dataset has significantly lower similarity than AI models, particularly for content, where its mean is approximately 0.58.
  • GPT-4o has the highest content similarity across nearly all categories, exceeding 0.85 for Climate-Resilient Farming Practices and Farmers’ Access to Fair Pricing.

Evaluator Assessment

The study compared two LLM evaluators with human annotations on 40 conversations to assess evaluator reliability. Humans agreed more with Gemini-3-Flash than GPT-5.2, but Gemini’s alignment varied substantially across annotators.

  • Evaluation setup: Two independent LLM evaluators and two human annotators assessed the same subset of 40 conversations using Krippendorff’s α.The LLM evaluators were Gemini-3-Flash and GPT-5.2, both configured with high reasoning settings.
  • Evaluator selection: α = 0.41 for human–Gemini-3-Flash agreement versus α = 0.10 for human–GPT-5.2 agreement, leading researchers to retain Gemini-3-Flash as the main evaluator.The comparison indicates stronger agreement between humans and Gemini-3-Flash than between humans and GPT-5.2.
  • Annotator variability: Gemini-3-Flash agreement differed across individual annotators, reaching α = 0.67 with one and α = 0.35 with the other.This asymmetry suggests evaluators differed in their perceptions and interpretations of value expression, with Gemini-3-Flash aligning more closely with one interpretation.

Prompts

The framework uses structured prompts for persona construction, dialogue-act annotation, value-orientation evaluation, and persona-conditioned discussion. These prompts constrain outputs to provided information, explicit utterance evidence, WVS categories, and specified formats.

  • Dialogue-act tagging: Dialogue-act tagging assigns one or more taxonomy labels to every utterance while preserving exact text and avoiding unsupported labels.The output is a JSON object containing conversations, matching IDs, and ordered per-utterance act lists.
  • Persona generation: Persona generation extracts demographic fields into a structured dictionary and creates a concise 100-word third-person biography using only supplied details.Unavailable fields are marked unknown, and English is added to the listed languages.
  • Longitudinal value evaluation: Value evaluation classifies participants’ current WVS orientations from expressed statements across all conversation history.It evaluates Traditional vs. Secular-Rational and Survival vs. Self-Expression dimensions, including the first two conversations and overall classification.
  • Persona-conditioned dialogue: Persona-conditioned discussion prompts instruct agents to mimic a persona’s background, values, and communication style within the two WVS dimensions.Agents are also prompted to infer the opponent’s value system using categorical WVS scales.
Loading 2609.05514v1…