Source-linked AI summary
Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources
Rana Muhammad Usman, Dominic Williamson
TL;DR
Single-agent benchmarks do not capture recursive social feedback in LLM-agent populations. The paper introduces PV-SST and a frozen, preregistered matched-exposure experiment, finding robust lexical convergence under a peer-ranked feed but no reliable distributed-source stance advantage. Its conclusions are limited to synthetic LLM-agent populations and do not estimate effects on people or production platforms.
Problem
Single-agent benchmarks do not measure recursive platform loops in which LLM agents produce content, react to others, and receive ranked feedback.
Method
PV-SST is an auditable peer-voted testbed paired with a frozen, preregistered matched-exposure experiment across model, topic, and seed blocks.
Results
A peer-ranked feed increases final-post lexical similarity across both model-size panels, while distributed sources show no reliable general stance-moving advantage over one source at fixed adversarial exposure.
Takeaways & Limitations
Feed exposure, ranking, source multiplicity, message diversity, and reach should be tested as separate mechanisms rather than inferred from account count alone.
Takeaways & Limitations
PV-SST models synthetic LLM-agent populations and its feed contrast bundles peer-post exposure with ranking, so it does not estimate human or ranking-only effects.
Abstract
from arXiv · showhide
Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a peer-voted social-platform testbed, and report a separately frozen, preregistered matched-exposure experiment spanning four topics, four unused seeds, four open-weight model families, and three prespecified larger variants. The experiment comprises 448 trials and 112 complete model-by-topic-by-seed blocks. Relative to a topic-only control, a feed of previous-round peer posts ranked by peer-generated likes increases final-round lexical similarity in both the four-family core panel (paired mean difference +0.0082 TF-IDF cosine units, 95% block-bootstrap CI [0.0043, 0.0121], randomization p=0.000105, n=64 blocks) and the three-variant size extension (+0.0109 [0.0069, 0.0151], p=0.000001, n=48). This contrast bundles peer-post exposure with ranking and therefore does not identify a ranking-only effect. Opposite-side survival falls in the core panel (-3.9 percentage points [-6.8, -1.6], p=0.0068) but not conclusively in the larger variants (-1.0 pp [-3.1, 0.4], p=0.50). Holding adversarial impressions fixed, four distributed sources do not reliably move honest-agent stance more than one source. The preregistered distributed-minus-single contrast is positive but inconclusive in the core panel (+0.057 [-0.009, 0.125], p=0.112) and negative in the larger variants (-0.040 [-0.113, 0.035], p=0.332), failing the prespecified cross-model and cross-topic consistency criterion. Thus the robust result is lexical convergence under the tested peer-ranked feed, not general opinion capture or a general coordination advantage. The study evaluates synthetic LLM-agent populations; it does not estimate effects on people or production platforms.
1 Introduction
PV-SST addresses the limits of single-agent benchmarks by testing recursive social feedback in controlled LLM-agent populations. Its confirmatory contributions are an auditable peer-voted testbed and matched-exposure tests of feed effects and source multiplicity.
- Single-agent benchmarks miss recursive platform loops in which agents produce content, react to others, and receive ranked feedback.Existing evaluations mostly measure isolated responses or task completion.
- PV-SST uses persona-conditioned agents, structured votes, platform-mediated feedback, and preserved traces of posts, exposures, stances, and parsing.The confirmatory protocol was frozen before outcomes were inspected and covered 448 trials across seven variants and four model families.
- The matched-exposure coordination test gives one-source and four-source attacks the same adversarial impressions to the same honest population.This design separates source multiplicity from reach and message volume advantages.
- A peer-ranked feed increases final-post lexical similarity across both model-size panels and every tested topic.Minority-view suppression is model-dependent and does not replicate conclusively in larger variants.
- The findings concern the implemented LLM-agent system, not human behavior.The testbed is intended to qualify threat models before human studies.
2 Related Work
Prior work establishes social influence, opinion dynamics, and LLM social simulation, but PV-SST targets a narrower gap: auditable, paired mechanism tests within synthetic agent populations.
- LLM social simulations study opinion formation, diffusion, polarization, and platform-scale interaction, but PV-SST emphasizes peer-voted, fully logged paired mechanism tests.The paper does not claim to be the first social simulator.
- Opinion-dynamics models formalize convergence and influence, while human experiments show popularity signals can alter evaluation and consumption.These traditions motivate the treatments but do not provide effect-size priors for LLM-agent populations.
- The study treats model-topic-seed runs as inferential units rather than generated posts, targeting internal causal contrasts within a synthetic system.It reports model and topic strata and avoids treating thousands of posts as independent samples.
- Figure 1 frames the design around a complete peer-ranked feed contrast and a fixed-exposure one-source versus four-source contrast.The feed comparison estimates the complete feed surface rather than ranking alone.
3 PV-SST Confirmatory Design
PV-SST uses six-round trials with balanced persona-conditioned agents, endogenous peer voting, paired feed and source contrasts, and block-level confirmatory inference. The design holds exposure and key attack features fixed while separating complete-feed effects from distributed-source packages.
- Each trial contains 12 honest agents from a 24-persona pool, balanced across six initial stance levels, and runs for six rounds.Agents receive feedback, emit posts, update stance and confidence, and vote on other posts each round.
- The confirmatory protocol was frozen before outcomes were inspected and spans four core model families, three larger variants, four topics, and four seeds.The core contains 256 trials and the size extension 192 trials.
- The peer-ranked feed shows five previous-round honest posts ranked by peer-generated likes after round 0, unlike the topic-only control.This contrast bundles peer-post exposure with ranking and estimates the complete feed surface, not ranking alone.
- Single-source and distributed-source conditions each expose every honest agent to one adversarial post and four organic posts after round 0.The distributed condition uses four accounts whose posts are rotated across viewers.
- 60 adversarial impressions per trial are fixed across attack conditions, while source multiplicity and cross-viewer message realization vary together.The contrast therefore tests a distributed-source package rather than isolated account count.
- The primary stance contrast is paired distributed_sources - single_source, with PDI requiring positive overall, family-level, and topic-level estimates.The criterion requires positive estimates in at least three of four core families and at least two thirds of completed topics.
- Final-post lexical similarity is the mean off-diagonal cosine similarity among 12 honest agents’ TF-IDF vectors, with vocabulary fitted within each trial endpoint.It is one of the prespecified secondary outcomes.
- The complete paired model-topic-seed block is the inferential unit, with block-bootstrap intervals, sign tests, sign-flip randomization p-values, and released model/topic strata.Individual posts are never treated as inferential units.
4 Confirmatory Results
The peer-ranked feed robustly increases final-round lexical similarity, while evidence for minority suppression and distributed-source influence is panel-dependent or inconclusive. Model-stratified results show meaningful heterogeneity rather than universal effects.
- Pooled contrasts: -3.9 percentage points in the core panel, but -1.0 percentage points in larger variants: opposite-side survival decreases only conclusively in the core panel.The core estimate has CI [-0.0677, -0.0156] and p=0.0068; the larger-variant interval is [-0.0312, +0.0035].
- Pooled contrasts: +0.0573 in the core panel and -0.0399 in larger variants: the distributed-minus-single contrast is uncertain and fails the PDI criterion.The core interval is [-0.0091, +0.1250], while the larger-variant interval is [-0.1128, +0.0347].
- Model and topic heterogeneity: Lexical-similarity effects are positive for three of four core models and all three larger variants, with all four topic means positive in each panel.Gemma 4 E4B is slightly negative, while Qwen 3.5 4B and 9B have the largest positive estimates.
- Model and topic heterogeneity: The core survival effect is concentrated in Qwen 3.5 4B, whereas larger-panel estimates vary across models, supporting model dependence rather than a universal effect.Qwen 3.5 4B has a -12.5 pp estimate; Gemma 4 E4B is exactly null, and larger-panel Gemma 4 12B and Ministral 3 14B are near zero or positive.
- Model and topic heterogeneity: For PDI, only two of four core topic means and one of four size-extension model means are positive, while the larger Qwen variant is significantly negative.These patterns explain why the pooled positive exploratory result cannot be generalized across model variants.
5 Exploratory Origin and Measurement Audit
The confirmatory questions arose from a small exploratory stage, while an exploratory misinformation probe exposed substantial measurement problems. Those probes informed protocol separation and motivated better stance-aware evaluation rather than confirmatory intervention claims.
- Exploratory origin: An 80-trial exploratory stage used two models on one platform-policy topic to screen feedback, coordinated-account, rank-amplification, and misinformation interventions.The stage generated the two confirmatory questions about replicated feed convergence and matched-exposure distributed sources.
- Exploratory origin: The exploratory cells were not confirmatory evidence because they used only two models, two seeds, and unmatched attack mechanisms.Keeping these results separate prevented exploratory choices from becoming retroactively preregistered claims.
- Measurement audit: An all-keyword detector recorded zero matches despite close restatements, while polarity-blind embeddings conflated endorsement with rebuttal.The paper labels these failures paraphrase leakage and stance confounding.
- Measurement audit: Two independent LLM stance judges agreed on only 43.6% of endpoint labels, with Cohen’s kappa=0.258, and disagreed on treatment direction.Consequently, the paper does not rank misinformation interventions and instead motivates human annotation or validated stance-aware outcomes.
6 Discussion
Under the tested PV-SST conditions, peer-ranked feeds reproducibly increase lexical similarity, while equal-exposure distributed sources show no reliable general stance advantage. These findings support decomposing platform risks into separately controlled mechanisms and keeping conclusions within the synthetic-agent setting.
- What the results establish: The peer-ranked feed produces the clearest replicated effect: final posts become more lexically similar under peer-post exposure ranked by endogenous likes.The contrast estimates the complete feed effect, not ranking alone.
- What the results establish: When adversarial impressions are fixed, four distributed sources do not reliably move stance more than one source.The distributed treatment changes source multiplicity and message realization together, so it is not an account-count-only test.
- Design implications: Matched exposure is necessary before comparing coordination mechanisms because source count otherwise remains entangled with reach and message volume.The distributed package is tested under equal adversarial impressions, not as isolated account multiplicity.
- Scope and limitations: The findings concern a synthetic LLM-agent population and do not estimate effects on people or named production platforms.The testbed omits human emotion, long-term relationships, off-platform information, and platform-specific recommendation systems.
- Design implications: Exposure and ranking must be separated in a follow-up unranked-feed arm because the current feed treatment bundles both mechanisms.The present experiment identifies a feed effect rather than a ranking-only effect.
- Scope and limitations: The lexical result is emphasized because it replicates across both model-size panels and topic strata, rather than reflecting one favorable cell.The feed outcomes are prespecified secondary estimands with nominal p-values.
- Scope and limitations: Simulation-to-human transfer remains uncalibrated because the released diagnostic lacks matched human treatment and the planned ChangeMyView replay used synthetic placeholders.This limits interpretation beyond the implemented agent system.
7 Conclusion
The confirmation provides a falsifiable evaluation of a synthetic LLM-agent population across model variants, topics, and seeds. It finds replicated lexical convergence under peer-ranked feeds but no reliable distributed-source stance advantage, motivating mechanism-decomposed evaluation rather than broad human-platform claims.
- Conclusion: A seven-variant, four-topic, four-seed confirmation provides a falsifiable evaluation of this LLM-agent population.The conclusion is limited to the tested synthetic population.
- Conclusion: A peer-ranked feed increases final-post lexical similarity across both model-size panels.This is the result that replicates across the tested panels.
- Conclusion: Distributed sources gain no reliable stance-moving advantage over one source when adversarial impressions are held fixed.The conclusion concerns the tested distributed-source package under matched exposure.
- Methodological lesson: LLM-agent platform risks should be decomposed into feed exposure, ranking, source multiplicity, message diversity, and reach, then tested with paired run-level designs.Generated posts are evidence to audit rather than independent samples for inflating statistical power.
- Methodological lesson: Human-platform claims require human evidence.The study's magnitudes are not predictions for people or production platforms.