Source-linked AI summary

On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial

Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, Robert West

arXiv:2403.14380v1cs.CY

TL;DR

LLMs raise concerns about tailored persuasion, yet their effectiveness in direct human conversations and the value of personalization remain limitedly documented. This pre-registered study uses randomized, structured online debates comparing human and GPT-4 opponents with and without access to anonymized personal information. Personalized GPT-4 produced the strongest persuasive effect, while the authors note that the forced debate structure may limit generalization to spontaneous online discussions.

  • Problem

    Evidence remains limited on LLM persuasion in direct conversations with humans and on how personalization compares with human performance.

  • Method

    A pre-registered randomized experiment compares human and GPT-4 opponents in short debates with or without access to anonymized demographic information.

  • Results

    81.7% higher odds of increased agreement occurred with personalized GPT-4 versus human opponents (p < 0.01), while non-personalized GPT-4 showed +21.3% with p = 0.31.

  • Takeaways & Limitations

    The results suggest that personalization and AI persuasion are meaningful concerns because GPT-4 can out-persuade humans in online conversations through microtargeting.

  • Takeaways & Limitations

    The predetermined debate structure may diverge from spontaneous online-conversation dynamics, making generalization to social networks and open platforms unclear.

Abstract

from arXiv · show

The development and popularization of large language models (LLMs) have raised concerns that they will be used to create tailor-made, convincing arguments to push false or misleading narratives online. Early work has found that language models can generate content perceived as at least on par and often more persuasive than human-written messages. However, there is still limited knowledge about LLMs' persuasive capabilities in direct conversations with human counterparts and how personalization can improve their performance. In this pre-registered study, we analyze the effect of AI-driven persuasion in a controlled, harmless setting. We create a web-based platform where participants engage in short, multiple-round debates with a live opponent. Each participant is randomly assigned to one of four treatment conditions, corresponding to a two-by-two factorial design: (1) Games are either played between two humans or between a human and an LLM; (2) Personalization might or might not be enabled, granting one of the two players access to basic sociodemographic information about their opponent. We found that participants who debated GPT-4 with access to their personal information had 81.7% (p < 0.01; N=820 unique participants) higher odds of increased agreement with their opponents compared to participants who debated humans. Without personalization, GPT-4 still outperforms humans, but the effect is lower and statistically non-significant (p=0.31). Overall, our results suggest that concerns around personalization are meaningful and have important implications for the governance of social media and the design of new online environments.

1 Introduction

LLMs have shown strong persuasive performance in authored texts, but their effectiveness in direct human conversations and the role of personalization remain insufficiently understood. This pre-registered study addresses that gap through randomized online debates comparing humans and GPT-4 with and without access to opponent information.

  • LLMs can generate persuasive content perceived as at least comparable to, and often more persuasive than, human-written messages.
  • Limited evidence compares LLM and human persuasion directly in conversations or tests how personalization changes LLM performance.
  • The study randomly assigns participants to human or GPT-4 opponents, with personalization enabled or disabled, in short online debates.
  • 81.7% higher odds of increased agreement occurred when participants debated personalized GPT-4 rather than humans (p < 0.01; N = 820).
  • Without personalization, GPT-4 produced a smaller, statistically non-significant effect of +21.3% versus humans (p = 0.31).

2 Related Work

Prior research has studied persuasion psychologically, compared LLM-generated and human-authored messages, examined personalization, and analyzed online debates. However, direct conversational comparisons between LLMs and humans remain limited.

  • Persuasion research has examined psychological determinants of language that drive opinion shifts across several outcomes.
  • Studies comparing human and LLM-authored messages generally find comparable or stronger perceived persuasiveness for modern language models.
  • Personalization research finds mixed evidence about whether LLM-based microtargeting improves persuasion.
  • Debate research has examined autonomous debating systems and persuasion strategies in human dialogues, including effects linked to participants’ backgrounds.

3 Methods

The study uses a multi-stage online experiment that selects broadly understandable, debatable topics, randomly assigns debate conditions and roles, and measures opinion change before and after structured debates. Persuasiveness is analyzed using transformed ordinal agreement outcomes.

  • Topic selection: Candidate propositions were required to be understandable, support reasons for both sides, be broad enough for personal resonance, and divide opinions.
  • Topic selection: Topics requiring advanced knowledge, extensive research, or overly specific evidence were excluded to preserve accessible debate conditions.
  • Topic selection: Annotators rated topics on 1–5 scales for Agreement, Knowledge, and Debatableness, which were used to construct topic-level measures.
  • Topic selection: From 60 candidates, researchers retained 30 topics after filtering unanimous and difficult-to-debate propositions, then grouped them into three Strength clusters.
  • Experimental workflow: Participants completed demographic and political surveys, were randomly assigned to treatments, topics, and PRO or CON roles, and debated synchronously through opening, rebuttal, and conclusion stages.
  • Experimental design: The four treatments crossed human versus GPT-4 opponents with non-personalized versus personalized access to anonymized opponent demographics.
  • Outcome and analysis: Persuasive effect was measured by comparing transformed pre-debate and post-debate agreement, with the post-debate outcome modeled as ordinal.

4 Results

GPT-4 with personalization produced the strongest increase in agreement with opponents, while unpersonalized GPT-4 showed a smaller, non-significant effect. Additional analyses examined agreement distributions, demographics, language, social dimensions, and opinion fluidity.

  • Persuasion outcomes: Personalized Human-Human debates showed a non-significant -17.4% change in persuasiveness relative to Human-Human debates without personalization.The confidence interval was [-46.1%, 26.5%] with p = 0.38.
  • Agreement distributions: Human-Human debates generally concentrated agreement changes in a backfire pattern, whereas personalized Human-AI debates had an average agreement difference of 0.14.The personalized Human-AI condition also showed more probability mass toward the upper antitriangular submatrix.
  • Demographics: Republicans were more likely to be persuaded by opponents, while demographic controls did not significantly alter the treatment effects.Republicans had +60% odds of greater agreement, with p = 0.02.
  • Textual analysis: Personalization produced very similar linguistic-feature distributions across personalized and non-personalized conditions.This similarity was observed for both Human-Human and Human-AI debates.
  • Textual analysis: GPT-4 used substantially more factual knowledge, whereas humans used more similarity, support, trust, and fun-related language.The analysis used social-pragmatic dimensions averaged across sentences in the Opening stage.
  • Opinion fluidity: Topic Knowledge reduced opinion fluidity by 72.9% per one-point increase, while Moderate prior scores increased it by 73.3%.Topic Debatableness increased fluidity by 117.8%, whereas topic clusters had negligible effects.

5 Discussion

The study compares AI- and human-driven persuasion in one-on-one online debates, finding especially strong effects when GPT-4 receives personal information. The authors argue that personalization creates meaningful governance concerns, while noting important limits to generalization and experimental validity.

  • The study used randomized one-on-one debates to compare human and LLM opponents with or without access to participant information.Registered agreements before and after debates measured opinion shifts as an indicator of persuasive power.
  • 81.7% higher odds of increased agreement occurred when participants debated personalized GPT-4 rather than humans.The estimate has a confidence interval of [+26.3%, +161.4%] and p < 0.01.
  • GPT-4 without personalization still produced higher agreement odds than humans, but the +21.3% effect was not statistically significant.The reported p-value was 0.31.
  • The authors conclude that personalization and AI persuasion warrant concern because LLMs out-persuaded humans through microtargeting even with limited personal information.They suggest online platforms should consider measures addressing LLM-driven persuasion.
  • Randomizing debate sides may weaken human arguments when participants do not genuinely hold the assigned position, potentially biasing comparisons with LLMs.A human-only robustness model found the relevant prior-agreement effect was non-significant and of the opposite expected sign.
  • The predetermined debate structure and time limits may constrain how well the findings generalize to spontaneous online discussions.The authors also note that time pressure may especially affect personalized human participants processing opponent information.

Appendix A Debate propositions

The appendix organizes debate propositions into low-, medium-, and high-strength clusters. The listed topics span political, social, educational, economic, health, and technology-related questions.

  • Low-Strength cluster: Low-Strength topics include felon voting rights, statehood, online learning, abortion, the death penalty, fossil fuels, and social media.
  • Medium-Strength cluster: Medium-Strength topics include leadership quotas, military aid, space exploration, taxation, election regulation, the Electoral College, animal research, and tuition-free college.
  • High-Strength cluster: High-Strength topics include school uniforms, artificial intelligence, national service, race-conscious admissions, basic income, unhealthy-food regulation, and surveillance.

Appendix B LLM prompts

The LLM prompt assigns a debate topic and side, structures responses into opening, rebuttal, and closing stages, and optionally supplies anonymized opponent information for personalization.

  • The prompt substitutes the assigned side with “in favor of” or “against” and inserts personalization instructions only when participant information is available.
  • The model is instructed to debate a specified proposition while impersonating the randomly assigned PRO or CON side.
  • Opening, rebuttal, and closing responses are each constrained to one or two concise sentences.
  • At each later stage, the model receives the opponent’s preceding argument and is instructed to address or respond to it.
  • With personalization enabled, the model receives gender, age, race, education, employment status, and political orientation to craft more persuasive arguments.

Appendix C Social dimensions

The appendix describes social dimensions as universal categories of social pragmatics and reports their distribution across Opening-stage sentences using classifier predictions averaged over sentences.

  • Social-dimension scores are computed by averaging classifier predictions across sentences in the Opening stage.The classifier was developed by Choi et al. (2020).
  • The dimensions come from Deri et al. (2018) and are analyzed as categories of social pragmatics.

Appendix D Regression results

Appendix D documents regression specifications for agreement-related outcomes, opinion fluidity, and perceived opponent, including demographic controls, reference categories, and uncertainty estimates.

  • The Brant-Wald test evaluates whether model (6) satisfies the proportional odds assumption.
  • Model (6) is estimated both without demographic covariates and with independently one-hot-encoded survey demographics.
  • The Human-Human debate is the reference category for the reported regression specifications, with demographic baseline characteristics specified for models including survey covariates.
  • Appendix regressions separately model Opinion Fluidity and Perceived Opponent using predictors described in the main-text tables.
  • Regression uncertainty is represented with 95% confidence intervals in the figures, while Table D2 uses Liang-Zeger cluster-robust standard errors.
Loading 2403.14380v1…