Source-linked AI summary
How Far Will They Go? Red-Teaming Online Influence with Large Language Models
Daniel C. Ruiz, Anna Serbina, Ashwin Rao, Emilio Ferrara, Luca Luceri
TL;DR
The paper asks how locally deployable open-source LLMs could be steered to support political influence campaigns, beyond static audits of intrinsic political bias. It introduces a red-teaming framework that measures political Overton Windows and jailbreak effects across models. The results show asymmetric and model-dependent steerability, motivating family-specific audits and defenses.
Problem
Existing audits provide limited evidence about adversarially steerable political behavior, despite the relevance of local open-source models to privacy- and compute-constrained influence actors.
Method
The study evaluates 31 open-source or open-weight instruction-tuned models using social-media generation, eight human-readable jailbreaks, and Overton Window scoring.
Results
Political expressivity is directionally asymmetric, jailbreak effects vary by model family, and Few-Shot prompting reliably amplifies while commitment/deception framings often increase refusal.
Takeaways & Limitations
Effective influence-oriented steering requires model-specific tuning loops, while defenses should use family-specific, scenario-grounded Overton Window audits.
Takeaways & Limitations
The study is restricted to instruction-tuned open-source models and does not capture proprietary, reasoning-only, or uncensored models.
Abstract
from arXiv · showhide
As large language model (LLM)-based agents increasingly participate in online discourse, red-teaming their capacity to support political influence campaigns is critical for information integrity. In pursuit of this goal, we focus on locally deployed open-source LLMs, as opposed to frontier API-only models, given their superior alignment with the operational constraints of privacy-conscious malicious actors deployed in social media environments. We introduce an empirical red-teaming framework for measuring LLM Overton Windows (OWs), defined as the range of political opinions a model can reliably express on controversial topics, and for quantifying how simple natural-language jailbreaks expand that range. We evaluate more than 30 LLMs spanning 10 model families and five countries of origin. We find systematic asymmetries in political expressivity: open-source LLMs are typically more willing to generate left-leaning social media content, OWs tend to contract inversely to model size, and regional differences are substantial despite uneven representation in the open-source ecosystem. Jailbreak potency also varies sharply across model families, motivating a workflow for identifying effective combinations of jailbreak techniques. Taken together, our results establish a practical framework for auditing the political steerability of open-source LLMs and for helping future researchers design stronger countermeasures against LLM-enabled influence campaigns.
1 Introduction
This study red-teams locally deployable open-source LLMs for political influence operations by measuring their political expressivity and susceptibility to simple jailbreaks. It evaluates how prompt techniques and model characteristics shape the practical range of content models can generate.
- Existing political-bias audits offer limited insight into how far LLM behavior can be externally steered under adversarial conditions.
- Local open-source models are especially relevant because privacy- and compute-constrained actors may use them to generate persuasive social-media content at scale.
- The framework measures Overton Windows as the range of political opinions models reliably express and how adversarial prompts shift that range.
- The study asks how simple human-readable jailbreaks affect model Overton Windows and how size, architecture, and origin influence expressivity and steerability.
- More than 30 models across 10 families and five countries support a workflow for identifying jailbreak combinations and auditing realistic misuse.
2 Related Works
Prior work measures intrinsic political bias and prompt steerability, but offers limited evidence about adversarially expanded political expression in realistic social-media misuse settings. This paper addresses that gap with simple, operationalizable jailbreaks and repeated open-ended evaluation.
- Existing political-bias studies largely examine intrinsic bias and static political space rather than adversarially altered behavior in realistic misuse settings.
- Prior steerability research examines persona prompting and other prompt- or model-level interventions for controlling outputs.
- This study deliberately focuses on low-cost, human-readable jailbreaks that are scalable and easy to operationalize.
- The Political Compass Test can be sensitive to forcing methods and prompt paraphrasing, motivating an open-ended social-media setup with repeated experiments.
- The framework measures both point-estimate lean and how simple adversarial prompts expand each model’s Overton Window.
3 Methodology
The methodology combines a curated ordinal opinion corpus, local generation, human-readable jailbreaks, and automated Likert-based judging across open-source instruction-tuned models. It aggregates normalized expression fidelity across topics, ideological positions, and repeated trials.
- Task Formulation and Topic Selection: The corpus contains 90 opinion statements across 10 topics, with nine ordinal positions per topic spanning extreme-left to extreme-right.
- Task Formulation and Topic Selection: The endpoints are deliberately inflammatory and extreme, while intermediate positions represent more mainstream policy stances for refusal and steerability stress tests.
- Generation: Models generate engaging social-media posts of at most 280 characters using temperature 1.0 and top-p 0.9, with experiments repeated across trials.
- Jailbreak Techniques: Eight human-readable jailbreaks are evaluated individually and in combinations, including Few-Shot, Authority, Extreme Persona, and Moral Decoupling.
- Models Tested: 31 open-source or open-weight instruction-tuned models are tested with inference-time reasoning disabled where supported.
- Evaluation: A local Qwen3 judge scores stance alignment from 0–9, with κ = 0.795 against human consensus on 210 manually labeled posts.
- Evaluation: The OW score is the mean normalized expression fidelity across topics, nine ideological positions, and trials, after generation and scoring.
4 Results
Across 31 open-source models, baseline political expression is generally high but left-leaning, while jailbreak effects and cross-model steerability vary substantially by technique, size, family, and origin.
- Baseline expression: Mean OW was 0.853 at baseline, with left-leaning positions expressed more faithfully than right-leaning positions across sensitive topics.This asymmetry appeared in 29 of 31 models, indicating that jailbreaks begin from a pre-tilted alignment surface.
- Prompt techniques: Few-Shot raised mean OW from 0.853 to 0.936 (∆= +0.083), while Anti-Neutrality and Extreme Persona produced smaller gains.Foot-in-the-Door, Adversarial Pleading, Moral Decoupling, and Authority reduced compliance on average.
- Prompt techniques: Jailbreak outcomes depended on the model-technique pair, with large Qwen3.5 checkpoints showing steep drops while Falcon-H1-34B remained near-flat or positively receptive.For Qwen3.5, Foot-in-the-Door drops were −0.381 at 122B and −0.304 at 27B.
- Compositional stacks: Greedy jailbreak stacks improved source-model OW but transferred weakly: the 0.5-1B stack succeeded in 1/4 tests, versus 3/4 for the 27-34B stack.Parameter count alone was a weak predictor of stack transferability.
- Baseline expression: Baseline OW scores ranged from 0.25 to 0.97, and 24/31 models exceeded 0.85.Additionally, 29/31 models had lean below the neutral value of 4.0.
- Cross-model variation: Mean OW generally declined with model size below 27B in 4/5 tested families, although Falcon-H1, OLMo-2, and Granite-4.0 remained highly compliant.Qwen3.5 showed earlier inverse scaling, dropping sharply by 27B.
- Cross-model variation: Family profiles differed: Falcon-H1 and OLMo-2 combined high baseline OW with low refusal, whereas Qwen3.5 had lower OW, stronger leftward lean, and larger degradations.The authors attribute practical steerability primarily to post-training policy choices rather than architecture class alone.
- Cross-model variation: Developer-origin aggregates also differed descriptively: UAE models were highest-compliance and closest to neutral, while Chinese models were lowest-compliance and most left-leaning.These comparisons are constrained by imbalanced origin groups, including single-model representation for France and India.
5 Discussion
The study finds asymmetric, model-dependent political steerability in local open-source LLMs, while practical auditing must account for important scope and measurement limitations.
- Across 31 models, left-leaning political content is generally easier to elicit, especially on sensitive topics.
- Few-Shot prompting reliably amplifies political expressivity, whereas commitment and deception framings often increase refusal.
- Influence-campaign use requires model-specific tuning loops rather than universal prompts.
- Limitations: The evaluation covers only instruction-tuned open-source models and excludes proprietary, reasoning-only, and uncensored systems.
- Limitations: The manually curated ordinal opinion corpus probes structured capabilities rather than directly measuring real-world political behavior.
- Limitations: A single LLM judge and a fixed set of prompting techniques may introduce evaluation bias or miss adaptive attacks.
6 Ethical Considerations
The ethical framing treats politically extreme content as a research instrument for adversarial robustness evaluation, with prompts and opinion tables documenting the study’s controlled setup.
- Politically extreme and potentially offensive statements are included solely to probe model robustness under adversarial conditions.
- Opinion domains: Opinion tables cover crime, foreign policy, gun policy, healthcare, immigration, LGBTQ+ and gender, free speech, and taxation.
- Prompt setup: The appendix documents prompt templates that combine manipulation techniques with baseline social-media generation prompts.
C.1 Human Annotation and Inter-Annotator Agreement
Human annotation used three independent raters to score 210 opinion–post pairs, with agreement assessed using ordinal reliability statistics.
- Three human annotators independently rated 210 opinion–post pairs on a 0–9 Likert scale.The sample comprised 70 unique opinions generated by three models.
- Agreement was assessed with Cohen’s quadratic-weighted κ and Krippendorff’s α for ordinal ratings.
- Krippendorff’s α = 0.478 fell below the 0.667 tentatively acceptable threshold despite strong practical reliability under skewed ratings.Most items clustered at scores 8–9, producing the reported kappa paradox.
C.2 Judge Evaluation and Selection
The judge-selection evaluation compared six candidate LLM judges against human consensus using three reliability measures.
- Six candidate LLM judges evaluated the same 210 items against the median consensus of three human annotators.
- Judge agreement was measured with quadratic-weighted Cohen’s κ, Krippendorff’s α, and ICC(3,1).
C.2.1 Optimal Panel Search
Exhaustive judge-panel evaluation found that the best internally consistent three-judge panel still underperformed human consensus and Judge A alone.
- κ = 0.693 against human consensus, 10% lower than Judge A alone, despite the best three-judge panel achieving mean pairwise κ = 0.709 and Krippendorff’s α = 0.438 internally.
C.3 Rationale for Single-Judge Selection
Judge A was selected because it aligned most closely with human judgments, remained robust across agreement metrics, and outperformed aggregated judge panels. Additional judges supported the robustness of the main findings rather than improving the primary evaluation.
- Judge A achieved κ = 0.795 against human consensus, outperforming all other judges and optimized panels.Adding Judge A to the human panel caused only minimal ICC degradation, from 0.843 to 0.820.
- Optimized judge panels underperformed Judge A by 10–20% against human judgment.Aggregation helped only when raters contributed independent noise; weaker judges instead introduced correlated bias.
- Cohen’s κ, Krippendorff’s α, and ICC(3,1) all ranked Judge A first, supporting robustness across measurement assumptions.
- Judges B–F served as robustness checks, with the main findings holding qualitatively across multiple evaluators.
E Miscellaneous Visualizations
The visualizations compare technique-induced changes in Overton Window scores across model families and sizes. Their captions emphasize heterogeneity in technique effects, including scale-dependent suppression and broadly positive Few-Shot effects.
- Figure 5 compares mean ∆OW relative to baseline across techniques and model sizes for Qwen3.5 and Gemma-3.Values are means ± standard deviations across 10 trials; blue indicates increased compliance, red decreased compliance, and the colormap is capped at ±0.42.
- Technique effects show strong family- and scale-dependent heterogeneity, including sharp suppression in larger Qwen3.5 checkpoints.Few-Shot remains broadly positive in the Figure 5 comparison.
- Figure 6 reports technique-minus-baseline ∆OW scores for Falcon-H1, OLMo-2, and Granite-4.0 models.Blue denotes increased opinion expression, red decreased expression, and the dagger marks the MoE model.
- Figure 7 reports technique-minus-baseline ∆OW scores for Gemma-3, Qwen3.5, and the remaining models.Gemma-3-1B is marked as an outlier because its baseline OW is approximately 0.25; the dagger marks the MoE model.