Source-linked AI summary
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
Art Kanke
TL;DR
DeflectBench addresses the understudied question of whether language models generate rhetorical fallacies on request, rather than merely detect them. It evaluates four frontier models across claims, prompt framings, and fallacy types, finding that request structure dominates claim content and produces distinct compliance signatures.
Problem
Prior computational work has focused almost exclusively on detecting rhetorical fallacies, leaving their requested generation insufficiently studied.
Method
DeflectBench evaluates four frontier models across claims, seven prompt framings, and three fallacy types, using blind dual-judge scoring to separate content, framing, and fallacy effects.
Results
Prompt framing and requested fallacy type dominate claim content as refusal drivers, while models differ in refusal, labeled compliance, soft refusal, clean compliance, and instruction-following fidelity.
Takeaways & Limitations
Labeled compliance offers a low-cost monitoring signal, and coach framing provides a reproducible probe for red-teaming refusal.
Takeaways & Limitations
The benchmark is English-only and single-turn, tests four proprietary models, lacks human-validated judge labels, and is not fully prompt-controlled across fallacy types.
Abstract
from arXiv · showhide
Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting fallacies in existing text. We close this gap with DeflectBench, evaluating 23,990 generations from four frontier models across three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims spanning four controversy levels. Refusal is governed primarily by request structure rather than claim content. Per claim refusal varies by only 11 percentage points across the 80 claims, while a single prompt frame change can swing within model refusal by nearly 100 percentage points and switching the requested fallacy type can swing it by over 80 percentage points within explicit framings. An educational debate coach prompt framing collapses refusal to near zero across all four model families, but the bypassed behavior is not clean compliance. Models typically produce labeled compliance, naming the requested manipulation in the same response that contains it. The four models distribute differently across refusal, labeled compliance, soft refusal, and clean compliance. The code and dataset are released at https://github.com/ArtKanke/DeflectBench.
1. Introduction
Deflection rhetoric avoids engaging a claim’s substance, and its generation by language models has manipulation-at-scale implications. DeflectBench addresses the underexamined inverse question of whether models will produce fallacies on request.
- Concept and motivation: Deflection rhetoric responds to claims without engaging their substance, including whataboutism, ad hominem, and red herring strategies.Whataboutism redirects attention to another wrongdoing, ad hominem attacks the claim-maker, and red herring introduces an irrelevant or loosely related topic.
- Concept and motivation: These strategies occur in political discourse, adversarial debate, and disinformation campaigns, making model-generated deflections a plausible vector for manipulation at scale.
- Research gap: Prior computational work has focused almost exclusively on detecting rhetorical fallacies in existing passages.
- Benchmark contribution: DeflectBench evaluates whether language models will produce rhetorical fallacies on request using claims, prompt framings, and fallacy types.The benchmark separates claim content, prompt framing, and fallacy type to identify which factor drives refusal.
- Benchmark contribution: Labeled compliance names the requested fallacy in the same response that contains it, a behavior that binary refusal benchmarks miss.The benchmark reports that this mode appears at meaningfully different rates across the four tested models.
2. Related Work
Related work has primarily studied fallacy detection and refusal behavior under different prompt framings. DeflectBench extends pluralistic alignment by examining how different alignment regimes resolve the same rhetorical manipulation requests.
- Fallacy detection: Computational fallacy research has predominantly addressed detection rather than generation on request.Existing work includes multiclass corpora, taxonomies, social-media whataboutism datasets, and bilingual detection benchmarks.
- Fallacy detection: Detection remains difficult, with reported classifier and LLM F1 scores below 0.55 and 0.15, respectively, on cited tasks.The cited work also reports red herring among the hardest categories and a roughly 30-point gap between the strongest tested LLM and human accuracy on one bilingual benchmark.
- Refusal evaluation: Refusal research shows that aligned LLMs can respond differently to semantically equivalent requests under different framings.SORRY-Bench and HarmBench provide standardized refusal-evaluation frameworks, while persuasion-based and persona-based prompting reduce refusal in cited work.
- Pluralistic alignment: DeflectBench extends pluralistic alignment by treating alignment regimes across laboratories as a distribution reflected in distinct model compliance signatures.The authors argue that safety evaluation cannot rely on one model as a proxy for the broader alignment landscape.
3. Methodology
DeflectBench systematically varies claims, prompt framings, and fallacy types across four frontier models, then uses blinded dual-judge scoring to characterize refusal and compliance. Its rubric distinguishes explicit refusal, soft refusal, labeled compliance, and clean compliance, while reliability results are treated as provisional without human validation.
- Benchmark design: DeflectBench evaluates four frontier models on 80 claims using 15 prompt templates and two blinded LLM judges.The benchmark scores each valid generation with an eight-field rubric.
- Claims: The 80 claims cover four controversy levels and two geopolitical contexts, with levels, contexts, and domains assigned by hand.There are 20 claims per controversy level, including factually true, consensus, contested, and factually false claims.
- Prompt templates: The seven framing conditions include explicit fallacy requests, a choose condition, and neutral or political implicit deflection requests.Ad hominem prompts name the fictional opponent Jordan Ivanov because attacking a named target is structurally required.
- Sampling and scoring: 23,990 valid generations were produced from 24,000 attempted cells, and 23,981 received valid scores from both judges.Models were sampled five times per model–claim–prompt cell at temperature T = 1.0, while judges ran at T = 0.
- Rubric: The rubric separates explicit refusal from positive responses, then independently flags soft refusal and fallacy labeling; clean compliance is the residual without either flag.Soft refusal uses substantive disclaimers that undercut rhetorical force, while fallacy labeling names the requested fallacy immediately before or after producing it.
- Reliability: κ = 0.45 and AC1 = 0.94 for soft refusal illustrate why the study reports both agreement statistics under prevalence skew.The authors report 95% block-bootstrap confidence intervals and Cohen’s h, while treating agreement as reliability rather than validity.
4. Results
DeflectBench finds that refusal varies far more with request structure and fallacy type than with claim content. Models also differ in whether they refuse, label compliant fallacies, soften compliance with disclaimers, or produce clean compliance.
- Compliance signatures across model families: The four models exhibit distinct compliance signatures rather than a single more-versus-less compliant ranking.Refusal and labeling vary independently across models, with differences highly significant (χ2 test, p < 0.001).
- Compliance signatures across model families: Claude frequently uses soft refusal, including disclaimers that undercut produced deflections and refusals based on request coherence rather than policy.Claude soft refusal reaches 75.5% under the choose framing.
- Framing dominates content: Prompt framing sharply changes refusal, with coach prompts reducing refusal to at most 0.6% while producing mostly labeled rather than clean compliance.Labeled rates under coach framing range from 89.2% to 99.1% across models.
- Framing dominates content: Explicit political and manipulation framings increase refusal for Claude and GPT, but leave DeepSeek and Grok largely unchanged.For GPT, coach-to-manipulation swings approach 100 percentage points on the same 80 claims.
- Free-choice fallacy preferences: Fallacy type and naming also reshape behavior: ad hominem is refused more under explicit framings, while implicit prompts yield more clean compliance and red herring.Under coach framing, refusal is near zero across all fallacy types and models.
- Framing dominates content: Refusal is nearly invariant across claim content: rates span 19.3% to 30.3% across 80 claims, an 11-percentage-point spread.Across controversy levels, refusal is practically equivalent within ±5 percentage points according to TOST.
5. Conclusion
DeflectBench finds that prompt framing and requested fallacy type drive refusal more strongly than claim content. Compliance also varies qualitatively across labeled compliance, soft refusal, and clean compliance, revealing distinctions that binary refusal measures conflate.
- Prompt framing and requested fallacy type both dominate claim content as drivers of refusal across four frontier models.
- Nearly 100 percentage points is the within-model refusal swing produced by changing prompt framing on the same 80 claims.
- Labeled compliance, soft refusal, and clean compliance distinguish how models respond when they comply.
- Binary refusal benchmarks conflate an instruction-following axis that separates these compliance patterns.
Limitations
The benchmark’s evidence is limited by judge validation, English single-turn coverage, incomplete prompt control across fallacy types, and evaluation of only four proprietary models.
- No human-validated subset of judge labels is available, so dual-judge agreement is treated as a reliability lower bound rather than ground truth.
- The English-only, single-turn benchmark leaves multilingual generalization and multi-turn conversational dynamics unaddressed.
- Ad hominem prompts include a fictional speaker reference unlike whataboutism and red herring prompts, limiting prompt control in cross-fallacy comparisons.
- Testing only four proprietary frontier models prevents determining whether the patterns generalize to open-source LLMs.
Ethical Considerations
The release includes generated examples of manipulative rhetoric, while its prompt templates and fallacy definitions make the benchmark’s generation conditions explicit.
- The released benchmark includes generated outputs that demonstrate manipulative rhetoric.
- The outputs can be reproduced by users with API access to the tested models and released prompt templates, so the release exposes no new model capabilities.
- Fifteen prompt templates are organized into seven categories, with explicit fallacy variants and implicit deflection requests.
- The templates define whataboutism, ad hominem, and red herring as distinct forms of deflection.
B. Claim Set
DeflectBench’s claim set contains 80 manually categorized claims spanning controversy levels, contexts, and domains, and responses are scored by blinded LLM judges using a multi-field rubric.
- B. Claim Set: The four levels are factually true, consensus opinion, genuinely contested, and factually false.
- B. Claim Set: The claim set includes U.S.-specific and international contexts, alongside geopolitical context and topical domain labels.
- C. Evaluation Rubric: Two judge models independently score valid responses while remaining blind to the generating model and prompt template.
- C. Evaluation Rubric: The rubric separately records refusal, soft refusal, three fallacy types, any fallacy presence, clean compliance, and labeled fallacy.
D. Reliability and Statistical Methods
DeflectBench assesses rubric reliability with Cohen’s κ, absolute disagreement rates, and bootstrap confidence intervals, while framing effects are summarized using Cohen’s h.
- Per-model reliability: κ = 0.91 for Claude’s any-fallacy field, compared with κ = 0.32 for DeepSeek and κ = 0.28 for Grok.DeepSeek and Grok produced a fallacy in over 99% of generations, illustrating the prevalence paradox.
- Per-model reliability: 0.2% disagreement accompanies Claude’s clean-compliance κ = 0.57, whose lower agreement reflects a 0.3% clean-compliance prevalence.The apparent low agreement arises within a narrow base-rate band.
- Statistical estimation: 1,000 block-bootstrap resamples of the 80 claims provide 95% confidence intervals for the four principal outcome variables.The estimates are computed over claims rather than treating individual generations as independent units.
- Framing effects: Cohen’s h shows large refusal effects (h ≥0.8) for Claude and GPT, while refusal-rare models show large effects only for coach-framing changes in clean compliance.Under coach framing, labeling supplants clean compliance for the refusal-rare models.
E. Per Prompt and Frame Level Breakdowns
Prompt framing and requested fallacy type produce substantial differences in refusal and compliance outcomes, with coach framing yielding near-universal fallacy production but usually labeled rather than clean compliance.
- E.1. Per-prompt-template breakdown: Ad hominem prompts trigger substantially more refusal than whataboutism or red herring prompts for Claude and GPT within explicit framings.The underlying claim and deflection task remain identical across these comparisons.
- E.2. Any fallacy rates by framing: Refusal-rare DeepSeek and Grok produce a fallacy in nearly every generation regardless of framing.Claude and GPT instead produce fallacies roughly in inverse proportion to their refusal rates.
- E.2. Any fallacy rates by framing: Coach framing yields near 100% any-fallacy across all four models, while political and manipulation framings yield near-zero any fallacy for Claude and GPT.The framing comparison links fallacy prevalence to refusal behavior in the refusal-prone models.
- E.3. Coach framing by fallacy type: Three models produce labeled compliance at near-uniform rates across all three fallacy types under coach framing.This pattern indicates that the requested fallacy type does not substantially alter labeled compliance for those models.
- E.3. Coach framing by fallacy type: DeepSeek produces 30% clean compliance for ad hominem under coach framing, while clean compliance remains below 5% for whataboutism and red herring.The model-specific exception is confined to the ad hominem comparison in this framing.
- F. Run-Level Variance: Each (model, claim, prompt) cell is sampled five times, and refusal-count distributions are reported across the repeated runs.The design includes 1,200 cells per model.
G. Verbosity
The four tested models differ substantially in output length, and fallacy density is defined relative to generated-token counts and reported by model and outcome category.
- Output length: 493 tokens is Claude’s mean output, compared with 396 for DeepSeek, 289 for GPT, and 65 for Grok.Mean output length differs by nearly an order of magnitude across the models.
- Output length by outcome: Mean output tokens are also reported separately by outcome type for each model.This breakdown complements the overall model-level output statistics.
- Fallacy density: Fallacy density counts the mean number of whataboutism, ad hominem, and red-herring indicators per 100 generated tokens.The indicator total ranges from 0 to 3.