Source-linked AI summary

AI agents reshape consensus formation in human groups

Lin Chen, Ziyi Liu, Xia Hu, Yong Li

arXiv:2609.02122v1cs.CLcs.CYcs.SI

TL;DR

The paper asks how LLM agents reshape consensus formation when they participate in human groups rather than acting only as tools. Using repeated random pairwise communication in collaborative description games, it varies agent proportions and finds three consensus regimes, alongside shifts in semantic content and perceived ownership. The findings link agent-led influence to shared linguistic alignment, stable expression choices, and changing human adoption responses.

  • Problem

    Whether mixed human–AI groups form shared conventions, and whether those conventions resemble human or agent consensus, remains largely unknown.

  • Method

    The study varies LLM-agent proportions in repeated human–AI collaborative description games and measures lexical, conceptual, behavioral, and perceived influence.

  • Results

    Agent proportion produces three regimes: low proportions facilitate human-led consensus, intermediate proportions disrupt convergence, and high proportions restore strong agent-led consensus while transforming consensus content.

  • Takeaways & Limitations

    Agent composition, transparency, and interaction structure are first-order design variables because repeated interaction can reshape collective norms and their perceived ownership.

  • Takeaways & Limitations

    The abstract-stimulus referential communication game brackets content sensitivity, power asymmetries, and emotional investment, so generalization to richer contexts warrants caution.

Abstract

from arXiv · show

As large language model (LLM) agents shift from tools to participants in human groups, a fundamental question for collective behavior is how their growing presence reshapes consensus formation. Here we study mixed human-AI groups in a collaborative description game, in which shared conventions emerge through repeated rounds of random pairwise communication. Varying the proportions of LLM agents, we identify three distinct regimes of consensus formation: low agent proportions facilitate human-led consensus, intermediate proportions disrupt convergence, and high proportions restore strong consensus while shifting it toward agent-led conventions. Crucially, these regimes differ not only in the strength of convergence, but also in the semantic grounding and communicative form of the resulting consensus: human-led consensus is more concrete, holistic, and grounded in shared real-world analogies, whereas agent-led consensus is more abstract, less information-dense, and more geometrically segmented. Mechanistically, agent influence arises from a shared linguistic prior that places agents near one another in the expression space, combined with relatively stable expression choices across rounds; humans initially resist adopting expressions from partners perceived as AI but gradually yield to conformity pressure. These findings provide evidence that AI composition can shape the emergence, content, and perceived legitimacy of group norms, making agent proportion and transparency important design variables for human-AI systems.

Introduction

The paper asks whether LLM agents merely amplify human consensus dynamics or reshape the conventions produced by mixed groups. It addresses this question by studying repeated communication in hybrid human–AI groups, where consensus may differ in strength, authorship, and meaning.

  • Motivation: LLM agents are becoming social interlocutors in workplaces, online communities, and collaborative platforms, not only tools for individual users.
  • Research gap: Prior research established consensus in pure-human and pure-agent populations, but shared conventions in hybrid groups remained largely unknown.
  • Research gap: The paper frames consensus as multidimensional, encompassing convergence strength, whose expressions shape the convention, and the meaning those expressions carry.
  • Approach: The study uses a collaborative description game with random pairings and no externally imposed vocabulary, allowing conventions to emerge through repeated local adjustments.
  • Overview of findings: Varying agent proportions reveals low-proportion facilitation of human consensus, intermediate disruption, and high-proportion restoration of strong agent-led consensus with more abstract, geometrically organized descriptions.

Results

Across repeated description-game interactions, agent proportion produces non-monotonic consensus strength and changes both the direction of linguistic convergence and the semantic form of consensus. Agent influence progresses from lexical shaping toward conceptual dominance, supported by shared initial alignment and stable expressions, while human perceptions mediate adoption.

  • Experimental paradigm: 40 rounds of random pairings let participants iteratively develop shared conventions for describing the same abstract tangram.
  • Consensus regimes: Consensus strength rises 8.0% at 12.5% agents, falls 23.1% and 14.5% at 33.3% and 50%, then rebounds to 0.725 at 75%.The pure-human baseline is mean consensus strength = 0.695.
  • Convergence direction: At 12.5%, agents move toward human linguistic space more than humans move toward agents, whereas from 33.3% onward humans move toward agents significantly more.The reversal is significant at 33.3%, 50%, and 75%.
  • Convergence direction: Similar consensus strength at low and high agent proportions reflects different pathways: agent-facilitated human consensus versus agent-led consensus.
  • Lexical and conceptual influence: Agents contribute 60.7% of final words at 33.3% population share, while conceptual contribution reaches 100% at 75%, indicating lexical influence precedes conceptual dominance.
  • Consensus content: Agent-led consensus is less concrete and less propositionally dense than human-led consensus, with concreteness declining from 2.986 to 2.724 and idea density from 5.419 to 5.321.
  • Consensus content: Human-led descriptions favor compact holistic labels and real-world analogies, whereas agent-led descriptions are more abstract, sparser, and organized around geometric parts and static enumeration.
  • Mechanisms: Agents begin significantly closer to one another in expression space, creating a shared semantic anchor from which coordinated influence can emerge.

Discussion

LLM agents can reshape shared conventions through ordinary repeated participation, affecting consensus strength, ownership, and norms. These effects motivate designs that account for agent participation, transparency, interaction structure, and the experimental limits of the evidence.

  • Mechanism and significance: Agent influence emerges from repeated ordinary participation rather than persuasion or a designated mediator role.The paper attributes group-level shifts to the accumulation of locally rational adaptations, although no individual interaction is coercive or strategically manipulative.
  • Design implications: Different agent proportions produce non-monotonic consensus outcomes, so deployment goals may require different participation levels.The paper specifically distinguishes preserving human authorship of shared norms from using AI to coordinate groups.
  • Limitations: The evidence is constrained by an abstract referential communication game and fully mixed random-pairing network, limiting direct generalization to naturalistic social settings.The study brackets content sensitivity, power asymmetries, emotional investment, and structured networks with communities, opinion leaders, and information bottlenecks.
  • Design implications: Hybrid-system design should treat agent proportion, transparency, and interaction structure as first-order variables governing collective norms.This extends analysis beyond individual model capabilities to collective behavioral consequences in human social environments.
  • Design implications: AI coordination can benefit communities while preserving meaningful control over the norms they form.

Methods

The study measures consensus strength, directional movement, contribution at lexical and conceptual levels, semantic content, and adoption–persistence dynamics in repeated human–AI interactions.

  • Consensus dynamics: Consensus strength measures final group alignment from pairwise cosine similarities among sentence embeddings, with within-individual centering to remove stable expression-style offsets.Final consensus strength is reported as C(T), the value at the last round.
  • Consensus dynamics: Directional movement measures whether participants move toward the opposite group’s initial linguistic space by tracking reductions in cosine distance.A positive Δd_i indicates adaptation toward the opposite group.
  • Contribution metrics: Lexical contribution attributes stable consensus vocabulary to agents or humans using population-normalized early-diffusion counts, producing a frequency-weighted score from 0 to 1.Stable vocabulary consists of words appearing in every final five rounds; the default diffusion window is three rounds.
  • Contribution metrics: Conceptual contribution captures agent influence over semantic territories by clustering expressions into semantic groups, complementing word-level attribution.The Conceptual-Lexical Ratio summarizes the relationship between conceptual and lexical influence.
  • Consensus content: Consensus content is assessed through concreteness, propositional idea density, analogical ratio, holistic ratio, and event framing ratio.These dimensions distinguish tangible grounding, information density, analogy, holistic structure, and dynamic framing.
  • Behavioral dynamics: Behavioral dynamics are decomposed into adoption of a partner’s expression and persistence of one’s own expression across rounds.Adoption is change in similarity to the partner, whereas persistence is consecutive-round self-similarity; prompt engineering manipulates these tendencies.

Experimental Procedure

The experiments used mixed human–AI groups in repeated, randomly paired description games, with agent proportions, traits, and capabilities varied across studies.

  • Participants: 127 human participants were recruited from universities in China through online advertisements and screened for English expression ability and expression diversity.The study received Institutional Review Board approval from Tsinghua University.
  • Experimental conditions: Study A tested 0%, 12.5%, 33.3%, 50%, and 75% agent ratios, while Studies B and C varied agent traits and model capability at fixed 33.3% proportions.Study B used high-adoption and high-persistence conditions; Study C used a smaller base model.
  • Stimuli: Tangram stimuli were selected from KiloGram using a vision-language-model ambiguity index, retaining figures with moderate interpretive ambiguity.The selection criterion was a mean pairwise embedding similarity between 0.65 and 0.70.
  • Procedure: Each session contained 24 participants who completed 40 rounds of randomly changing pairwise description exchanges under homogeneous mixing.Participants independently described the same stimulus and received partner descriptions with similarity feedback.
  • Procedure: Participants were told that partners could be human or AI but were not told the number or proportion of agents.They answered in English using one descriptive sentence and were instructed not to use external language models.
  • Questionnaires: Post-experiment questionnaires collected ratings of final-description agreement, perceived AI leadership, and interaction experiences, with additional items personalized to interaction histories.Questionnaire items sampled early, middle, and late rounds for partner-identity and adoption judgments.

Agent Contribution

Agent contribution is quantified separately at lexical and conceptual levels using early diffusion into stable final consensus vocabulary and semantic clusters.

  • Lexical-level contribution: Lexical contribution uses stable vocabulary words appearing in every final five rounds and population-normalized usage during a three-round early-diffusion window.The resulting per-word score ranges from 0 for entirely human-driven to 1 for entirely agent-driven.
  • Lexical-level contribution: Clex is the frequency-weighted average of per-word agent contribution scores, with weights equal to final-five-round usage counts.The weighting emphasizes words used more frequently in the final consensus.
  • Conceptual-level contribution: Conceptual contribution applies analogous early-diffusion attribution to semantic clusters that persist through all final five rounds.Clusters are formed from scene-graph-based expression similarities using optimized HDBSCAN parameters.

Agent Trait Prompts

Agent behavior was controlled through system prompts that preserved the game rules while adding neutral, high-adoption, or high-persistence tendencies.

  • Neutral condition: The neutral prompt specifies a 40-round semantic-similarity game and instructs agents to maximize accumulated payoff without additional behavioral guidance.The neutral condition serves as the default across experimental conditions.
  • Prompt design: All prompt conditions retain the same base game context, while trait conditions differ through appended behavioral-tendency clauses.T denotes the total number of rounds, and payoff_min and payoff_max define the semantic-similarity payoff range.
  • High-adoption condition: High-adoption prompts instruct agents to observe and imitate partners’ phrasing and structure to increase similarity and payoff.This condition prioritizes progressive alignment with the partner’s way of describing stimuli.
  • High-persistence condition: High-persistence prompts instruct agents to maintain consistent wording, sentence structure, and anchor words across rounds, making only small adjustments.They avoid adopting partner expressions unless those expressions fit their own interpretation.

Content Annotation Using LLM-as-a-Judge

The paper uses an LLM-as-a-judge procedure to score tangram descriptions on independent semantic dimensions, including analogicality, holistic structure, and event framing. Expressions are annotated as continuous 0.00–1.00 JSON scores under a psycholinguistic framework.

  • Annotation procedure: Each unique expression from the final five rounds is annotated once on three independent dimensions.Identical expressions share annotation scores across all instances.
  • Annotation dimensions: Analogicality measures how strongly a description maps the figure onto a recognizable real-world entity, creature, person, or scene.The scale ranges from purely geometric or abstract descriptions to vivid real-world entity mappings.
  • Annotation dimensions: Holistic structure measures whether a description treats the tangram as a unified whole rather than decomposing it into geometric parts.Body parts of an analogical entity are explicitly excluded from geometric decomposition.
  • Annotation dimensions: Event framing measures whether a description presents the figure as engaged in a dynamic action, posture, or scene.Compositional verbs do not increase the score; agentive actions such as sitting, looking, or running do.
  • Output format: The annotator returns only a JSON object containing continuous analogical, holistic, and event scores from 0.00 to 1.00.The procedure asks annotators to use the full range rather than round to convenient values.

Influence Network Analysis

The influence network converts pairwise expression changes into directed edges from influencers to influenced participants, then applies PageRank to the reversed graph. Overall human–agent influence distributions do not significantly differ, although unusually influential agent nodes emerge at higher agent proportions.

  • Network construction: Pairwise directional influence is estimated by comparing a participant’s actual embedding change with the direction toward the partner’s prior expression.A high positive cosine similarity indicates that the participant’s change aligned with movement toward the partner.
  • Network construction: The final directed network retains the stronger direction for each interaction and averages repeated directional edge weights across rounds.Edges point from the more influential participant to the less influential participant, with weights representing net directional influence.
  • Influence scoring: PageRank is computed on the reversed influence graph so higher scores identify participants whose expressions were more frequently adopted by others.The influence vector is defined as the stationary distribution of the row-normalized reversed graph, using α = 0.85 and N = 24.
  • Results: Human and agent influence-score distributions do not differ significantly across conditions, with all Mann–Whitney U tests yielding p > 0.1.This result concerns overall distributions rather than individual high-influence participants.
  • Results: At 50% and 75% agent proportions, individual agents with influence scores substantially above the group mean begin to appear.The authors leave open whether these nodes reflect systematic network topology or stochastic interaction sequences.
Loading 2609.02122v1…