Source-linked AI summary
Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments
Lin Chen, Yitong Chen, Yong Li
TL;DR
The paper investigates whether LLMs update beliefs like humans when judging persuasive arguments, a question important for using them as proxies in social simulations. It compares eight models with participant-verified ChangeMyView outcomes and finds slight agreement, systematic cue and strategy differences, and greater resistance under third-person observation. These results identify structural divergence between LLM and human processing of persuasive discourse.
Problem
Whether LLMs revise beliefs in response to persuasive arguments like humans remains poorly understood, despite their increasing use as proxies in social simulations.
Method
The study compares eight LLMs’ binary belief-update judgments with participant-verified persuasion outcomes from matched ChangeMyView replies.
Results
LLMs achieve only slight agreement with humans and diverge in their sensitivity to textual cues, persuasion strategies, proposition types, and perspective framing.
Takeaways & Limitations
LLM judgments should not be treated as faithful general-purpose proxies for human belief updating in persuasive discourse.
Takeaways & Limitations
The CMV user population likely skews toward young, English-speaking, politically engaged Reddit users, limiting generalizability to other demographic contexts.
Abstract
from arXiv · showhide
Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs in response to persuasive arguments, as humans do, remains poorly understood. We conduct a systematic comparison using a naturally occurring online persuasion corpus in which original posters explicitly verify whether a reply changed their view. Our results show that LLMs achieve only slight agreement with humans (Cohen's kappa ranging from 0.079 to 0.178). Content-level analyses show that humans and LLMs agree on the strongest persuasion cues but diverge on finer ones: humans are more swayed by novel content and assertive language, whereas LLMs favor topical similarity and surface-level formatting. At the level of persuasion strategy, LLMs underweight emotional appeals and overweight credibility signals relative to humans, while the type of proposition under debate exerts no measurable effect on the degree of divergence. Furthermore, switching from first-person role-playing to third-person observation shifts all models toward greater resistance to persuasion, with the effect varying across persuasion strategies and textual features. These findings highlight the risk of treating LLM judgments as faithful proxies for human belief updating and point to structural differences in how LLMs and humans process persuasive discourse. Our code is available at https://github.com/tsinghua-fib-lab/LLM-belief-update-cmv.
1 Introduction
The paper asks whether LLMs update beliefs like humans when evaluating persuasive arguments, and identifies the factors associated with human–LLM divergence. Using real-world persuasion data, it finds slight agreement, cue-level differences, and systematic effects of perspective framing.
- Motivation: LLMs are increasingly used to simulate human interactions, making human-like belief updating an important empirical question.The paper frames belief updating as revising or maintaining prior beliefs after encountering evidence or arguments.
- Research questions: The study asks whether LLM judgments differ systematically from humans and whether content, proposition type, persuasion strategy, or perspective framing modulates that divergence.These questions correspond to RQ1–RQ3.
- Approach: The authors compare eight LLMs with participant-verified persuasion outcomes from matched replies in the ChangeMyView corpus.Each pair contains one reply acknowledged as changing the original poster’s view and one that was not.
- Main findings: All models achieve only slight agreement with human labels, while LLMs favor topical similarity and discount novelty, emotional appeals, and some human-relevant textual cues.The divergence is shaped by argument construction and persuasion strategy rather than proposition type.
- Main findings: Third-person observation shifts every model toward greater resistance to persuasion, with effects varying across strategies and textual features.Perspective framing also alters error composition and feature sensitivities.
2 Experimental Design
The experimental design uses ChangeMyView discussions with participant-verified persuasion outcomes and asks eight LLMs to make binary belief-update judgments. Divergence is analyzed through textual features, proposition types, persuasion strategies, and first-person versus third-person perspectives.
- Dataset construction: The study builds its dataset from ChangeMyView, an English-dominant Reddit community where users invite challenges to personal views.Successful persuasion is explicitly acknowledged by the original poster with a delta symbol.
- Dataset construction: Matched reply pairs contain one root reply awarded a delta and one unsuccessful reply, isolating single-argument persuasion outcomes.The filtering yields 2,262 original posts and excludes deltas awarded only after later exchanges.
- Task formulation: Models receive an original post and challenger reply independently and output whether the reply would change the expressed view.The binary output is delta_awarded: true or false.
- Perspective conditions: The task compares first-person role-playing as the original poster with third-person prediction as an external observer.This design tests whether active stance maintenance introduces additional bias.
- Evaluation: Alignment is measured with Cohen’s κ, while True Positive Rate and False Positive Rate characterize directional errors.A high FPR indicates overacceptance of ineffective arguments, whereas a low TPR indicates rejection of effective ones.
- Analytical dimensions: The analysis covers textual features, three proposition types, and three independently annotated persuasion strategies.Proposition types are fact, value, and policy; strategies are logos, pathos, and ethos, with multiple strategies allowed per reply.
- Experimental setup: The model set includes DeepSeek-V3, MiniMax-M2.5, GLM-4.7, GPT-4o-mini, GPT-5.5, Qwen-32B, Qwen-72B, and Gemini-2.5-Flash.The final three names are abbreviations for the corresponding full model names.
3 LLM-Human Divergence in Belief-Update Judgments (RQ1)
Under first-person role-playing, the eight LLMs show only slight agreement with human belief-update labels. Although overall error rates are similar, models differ in whether they tend to reject or accept persuasive arguments.
- Overall alignment: Cohen’s κ ranges from 0.079 for GPT-4o-mini to 0.178 for GPT-5.5, placing every model in the slight-agreement category.The best-performing model captures less than one fifth of the non-chance agreement structure with human labels.
- Error composition: Total anomaly rates span 41.1% to 46.1% across models, but the composition of false negatives and false positives differs substantially.The similar aggregate error rate masks distinct behavioral profiles.
- Error composition: Gemini-2.5-Flash, Qwen-32B, DeepSeek-V3, and GPT-5.5 show receptive profiles with more false positives than false negatives.Gemini-2.5-Flash is the extreme case, with FN = 8.0% and FP = 37.6%.
- Interpretation: The results complicate treating LLM judgments as general-purpose proxies for human belief updating.The divergence is substantial despite evaluating the same persuasion judgments.
4 Content Effects on LLM-Human Divergence (RQ2)
Content affects LLM–human divergence unevenly: humans and LLMs share broad structural cues but differ on finer textual features and persuasion strategies, while proposition type does not meaningfully moderate divergence.
- Textual Features: Humans reward novel reply content, whereas LLMs favor replies that remain topically close to the original post.Reply Dissimilarity has opposite associations: human OR = 1.1476 and LLM OR = 0.9115, both q < 0.001.
- Textual Features: Humans respond to assertive language, while LLMs are more responsive to formatting and stylistic markers.Reply Definitive predicts human persuasion but not LLM persuasion; Reply Formatting survives FDR correction only for LLMs, and LLM judgments also differ on inclusive language and first-person expression.
- Textual Features: Longer, evidence-backed replies are more persuasive for both humans and LLMs, while longer or more definitive original posts are harder to overturn.Reply Length and Reply Link are positive predictors for both groups; OP Length and OP Definitive are negative predictors.
- Proposition Types: The type of proposition under debate does not meaningfully moderate LLM–human divergence.Agreement and error composition remain similar across value, fact, and policy propositions.
- Persuasion Strategies: Combining multiple persuasion strategies increases persuasion for both humans and LLMs, with L+E+P achieving the highest rates: 57.08% for humans and 65.93% for LLMs.Ethos alone has low effectiveness for both groups but consistently boosts effectiveness when added to other strategies.
- Persuasion Strategies: LLMs underweight emotional appeals and overweight credibility signals relative to humans.Pathos persuades humans at 41.67% versus 24.34% for LLMs, while ethos-containing combinations generally produce higher LLM persuasion rates.
5 Perspective-Induced Bias in Belief Update Judgments (RQ3)
Changing from first-person role-playing to third-person observation systematically reshapes LLM–human divergence: models become more resistant, textual sensitivities partially realign with humans, and strategy effects vary.
- Observer perspective and error composition: All eight models show higher false-negative and lower false-positive rates under observer framing than first-person role-playing.The observer condition shifts every model toward rejecting persuasion more often.
- Observer perspective and error composition: The observer condition collapses three first-person error profiles into a resistant profile across all models.Under observation, false negatives substantially exceed false positives.
- Interpretation: First-person role-playing appears to amplify accommodation, whereas third-person observation reduces that pressure and produces a conservative default against persuasion.The paper connects this shift to sycophantic behavior and contrasts it with the human third-person effect.
- Textual features: Observer framing moves most divergent textual-feature coefficients toward human baselines, with the largest correction occurring for OP Length.OP 1stPerson strengthens in the opposite direction, remaining a notable LLM-specific predictor.
- Proposition types and strategies: Proposition type remains unrelated to divergence, while observer framing lowers LLM persuasion rates across strategy combinations with uneven effects.Strategy rankings are largely preserved, and the perspective effect interacts with persuasive content rather than acting as a uniform threshold shift.
- Overall implication: Perspective is therefore an active source of bias in LLM belief-update judgments rather than a neutral evaluation choice.The observer condition suppresses accommodation, partially realigns feature sensitivities, and affects strategies to different degrees.
6 Related Work
Prior work studies LLMs as social simulators, persuaders, and debate participants, but comparatively little work tests whether LLMs update beliefs like humans on individual arguments.
- LLMs as proxies for humans: LLMs are used as proxies in simulated interaction, opinion surveys, online discourse, elections, and collective decision-making, although proxy fidelity remains contested.Prior findings include demographic-level fidelity alongside misalignment for particular groups and topics.
- LLM persuasion and opinion change: Research on LLM persuasion mainly examines generated arguments, audience personalization, or open-ended debate behavior rather than belief-update judgments against human ground truth.Documented tendencies include accuracy bias, social bias, and context-driven shifts in group or debate settings.
- Positioning this study: The closest CMV studies analyze LLMs as persuaders or forecasting tools, whereas this work treats them as persuadees and compares their judgments with human-verified outcomes.The paper diagnoses divergence through textual features, proposition type, and persuasion strategy.
- Perspective and self–other asymmetries: Research on role assignment, perspective-taking, sycophancy, and human third-person effects motivates examining how perspective shapes LLM belief-update judgments.Whether these effects align with or depart from human self–other patterns remains largely unexamined.
7 Conclusion
The study finds a structural mismatch between LLM and human belief updating: divergence is shaped by argument construction and perspective, but not proposition type.
- Conclusion: All tested models achieve only slight agreement with human-verified persuasion outcomes, despite similar overall divergence rates across models.The internal composition of errors differs markedly between models.
- Conclusion: LLMs favor topical overlap and credibility signals while discounting novelty and emotional engagement relative to humans.These differences indicate qualitatively different cues for evaluating persuasive force.
- Conclusion: Third-person observation shifts every model toward greater resistance, with uneven effects across persuasion strategies and textual features.The conclusion identifies perspective as a source of systematic variation in divergence.
- Conclusion: The persistence of mismatch in the strongest tested model leaves whether further scaling resolves it as an open empirical question.The paper also points toward evaluations of multi-turn persuasion trajectories and richer persona information.
Limitations
The study’s conclusions are bounded by the CMV sample and labels, absent persona context, and model-generation changes that may alter specific bias magnitudes.
- Dataset and labels: CMV users likely skew young, English-speaking, and politically engaged, limiting generalizability to other demographic contexts.The corpus is retained because it provides an initial belief, counterargument, and participant-verified persuasion outcome.
- Dataset and labels: CMV labels capture single-root-reply acknowledgment and binary belief change, introducing asymmetric negative-label noise and obscuring degrees of updating.Persuasion occurring through later exchanges may be missed, while partial and complete shifts receive the same positive label.
- Persona context: Default model configurations omit the original poster’s values, background, and prior discussion history, so they are not faithful simulations of specific respondents.The authors present this as a baseline and propose adding richer persona attributes.
- Model evolution: Rapid model development may change specific bias magnitudes in future generations, although the authors expect broader structural patterns to persist.This qualification limits direct projection from the evaluated checkpoints.
Ethical Considerations
The study uses public ChangeMyView data and annotation prompts designed to classify belief updates, proposition types, and persuasion strategies without recruiting human participants. It cautions that LLM judgments should not be treated as faithful proxies for human responses without validation.
- The study uses publicly available ChangeMyView data, without personally identifiable information, and recruits no human participants.
- The authors caution that divergence from human belief updating does not establish defective reasoning, but limits unvalidated use of LLMs as human-response proxies.
- First-person prompts assign models the original poster’s role, while third-person prompts frame them as impartial observers predicting whether the poster’s view changed.
- Annotation prompts define proposition types and persuasion strategies, including logos, pathos, and ethos, while instructing models to classify sensitive texts as research material.
- The strategy schema permits replies to receive multiple labels when they combine modes such as statistics and empathy.
B Judgment Stability across Repeated Runs
Repeated first-person evaluations are generally stable across random seeds, while annotation validation shows substantial or near-perfect agreement for proposition types and persuasion strategies.
- Fleiss’ κ ranges from 0.739 to 0.900 across seven of eight models, with pairwise agreement of at least 86% across repeated runs.MiniMax-M2.5 is the exception, with Fleiss’ κ = 0.438.
- GLM-5.2 and the original annotator achieve Cohen’s κ = 0.751 for proposition type and micro-averaged κ = 0.817 for persuasion strategy.
D Consistency across Models
LLMs agree with one another more than with humans under the first-person condition, although their textual-feature sensitivities are not uniformly aligned.
- Inter-model Cohen’s κ ranges from 0.190 to 0.548 under first-person evaluation, substantially exceeding model–human agreement.
- Most additional textual features show no significant association with persuasion outcomes for either humans or LLMs.
- Type-token ratio, paragraph count, and other features reach significance exclusively for LLMs in separate regressions.
F Regression of Continuous Belief Change Scores
The study tests whether textual features predict a continuous belief-change score derived from model justifications, extending the binary analysis while documenting how perspective framing affects consistency.
- The original belief-change measure is binary, whereas the continuous analysis scores belief change from 0–100 using GPT-5.1 evaluations of model justifications.
- The continuous analysis is an extension necessitated by the binary delta measure and limited access to direct confidence measures for closed-source models.
- The five largest continuous-score coefficients—Reply Length, OP 1stPerson, Reply Link, OP Length, and Reply Dissimilarity—match the directions of the binary logistic regression.
- Cross-perspective consistency ranges from raw agreement = 0.90 and κ = 0.75 for GLM-4.7 to raw agreement = 0.41 and κ = 0.12 for Gemini-2.5-Flash.
- Gemini-2.5-Flash shifts most dramatically from a receptive first-person profile to a resistant observer profile, while GLM-4.7 changes comparatively little.