Source-linked AI summary
Interrupting the Chain: Human Perception of AI-Generated Disinformation Through a Kill Chain Lens
Alexander Loth, Martin Kappes, Marc-Oliver Pahl
TL;DR
Generative AI enables scalable, plausible disinformation while existing defenses act largely after cognitive exploitation. Using human evaluations of AI and human news fragments organized through an adapted kill chain, the paper finds a perception-accuracy gap, frequent human-indistinguishable LLM text, and selective fatigue in fake-news detection. These results identify candidate points for proactive, stage-specific defense, while the framework and several effects remain exploratory or indirectly grounded.
Problem
Generative AI produces culturally nuanced, plausible disinformation at scale, while current defenses operate after cognitive exploitation and lack a proactive intervention framework.
Method
The study evaluates news fragments with 504 participants and organizes origin and veracity perception data through an adapted five-stage cybersecurity kill chain.
Results
The findings show a perception-accuracy gap, frequent human-like LLM outputs, and a 10.2 percentage point decline in fake-news detection under sustained exposure while AI-origin detection remains stable.
Takeaways & Limitations
The results identify candidate intervention points for proactive, stage-specific defenses that target cognitive vulnerabilities rather than relying only on machine-origin detection.
Takeaways & Limitations
The text-only controlled study lacks real-platform social dynamics, uses a primarily European/North American sample, and reports rapidly evolving model-dependent rates as a snapshot.
Abstract
from arXiv · showhide
Generative AI enables customized misinformation at scale, yet defenses remain largely reactive. We present empirical findings from a human-subject study (n=504 participants, n=2,438 judgments) in which users classified news fragments by origin (human vs. machine) and veracity (real vs. fake). We organize results using an adapted cybersecurity kill chain as a taxonomy for intervention, mapping perception data onto stages of a cognitive attack lifecycle. Three key findings emerge: (1) a perception-accuracy gap where heightened suspicion does not improve detection; (2) modern LLMs frequently produce human-indistinguishable text; and (3) an asymmetric cognitive fatigue effect where fake-news detection degrades by 10.2 percentage points under sustained exposure while AI-origin detection remains stable. These findings identify candidate intervention points for proactive defense against AI-driven disinformation.
1 Introduction
Generative AI enables large-scale, culturally nuanced disinformation that can outpace debunking. The paper adapts the cybersecurity kill chain as a taxonomy for identifying proactive intervention points without treating disinformation as a strictly sequential process.
- Problem: Generative AI turns disinformation into systematic cognitive attacks that threaten trust in online information ecosystems.Large language models can produce plausible, culturally nuanced narratives at scale, creating an asymmetry between production and debunking.
- Framework: The adapted five-stage taxonomy comprises Reconnaissance, Weaponization, Delivery, Exploitation, and Post-Exploitation.These stages cover target profiling, AI-content creation, dissemination, cognitive effects, and attribution evasion.
- Framework: The kill chain organizes empirical signals into actionable defense layers and identifies candidate points where defenses can interrupt the attack lifecycle.The framework is presented as a conceptual scaffold rather than a claim that campaigns literally follow discrete sequential phases.
- Evidence scope: Stages 1, 2, and 4 have strong empirical grounding, whereas Stages 3 and 5 rely on indirect evidence.Evidence strength is indicated for each stage in Figure 1.
2 Related Work
Prior disinformation frameworks mainly describe campaign spread, operator tactics, or content analysis, while human AI-text detection has often been studied separately. This paper contributes a compact lifecycle taxonomy anchored in measured human-perception data to locate cognitive vulnerabilities.
- Research gap: Human AI-text detection has been studied in isolation from kill-chain treatments of disinformation.Prior approaches remain grounded in operator behavior or content analysis.
- Prior frameworks: Existing frameworks primarily model operator behavior, including campaign spread, tactics and techniques, or attacks on democratic cognition.Their granularity and intent differ across breakout-scale, DISARM/AMITT-like, and cognitive-attack perspectives.
- Contribution: The paper retains a compact five-stage lifecycle but shifts analysis from attacker actions to where human targets are cognitively vulnerable.Each stage is anchored in measured human-perception data.
- Framework boundary: The kill-chain analogy supports stage-specific, layered intervention but is less suitable as a causal or strictly sequential process model.Disinformation stages can overlap, iterate, and run in parallel, and exploitation may occur without tailored reconnaissance.
- Contribution: The work addresses a gap by mapping large-scale human-perception data onto the attack lifecycle and measuring both veracity and origin judgments using full-text fragments.JudgeGPT adds machine-versus-human origin as a second dimension beyond headline-level veracity testing.
3 Methodology
The study combines multilingual, multi-model news-fragment generation with human evaluations of origin, veracity, and topic familiarity. Accuracy analyses compare participant judgments across demographic, temporal, and model-related factors.
- Stimuli: RogueGPT generated news fragments with seven LLMs across four languages, three styles, and three formats.Human-written legitimate content and verified fake-news material supplemented the generated corpus.
- Human study: JudgeGPT collected evaluations from 504 participants using three 0–1 sliders for origin, veracity, and topic familiarity.Accuracy counted a response as correct when its score was at least 0.5 and matched ground truth.
- Analysis: The analysis used Pearson correlations, independent t-tests, temporal fatigue analysis, and LLM performance comparisons.A causal analysis of susceptibility factors was presented separately.
4 Findings Through the Kill Chain Lens
Across the kill-chain stages, the findings show that suspicion is weakly connected to accurate detection, high-quality AI text can appear human, rapid judgments reduce accuracy, and sustained exposure selectively harms fake-news detection. The mapping identifies cognitive vulnerabilities while distinguishing direct empirical evidence from exploratory or indirect interpretations.
- 4.1 Reconnaissance: Higher suspicion did not translate into better detection: familiarity correlated with suspicion at r=0.20, while accuracy correlations remained negligible at |r|<0.12.Overall accuracy was 58.4% for AI content and 68.1% for fake content.
- 4.2 Weaponization: Modern commercial LLMs frequently produced text perceived as human-written, with HumanMachineScore values of 0.463 for GPT-3.5-turbo and 0.498 for GPT-4o.Smaller open-source models were somewhat more detectable, with scores of 0.519–0.545; scores nearer 0.50 indicate more human-like outputs.
- 4.2 Weaponization: AI-generated legitimate content produced 44.6% origin-detection accuracy, below the 50% chance baseline, whereas AI-generated fake content reached 58.6%.Veracity detection was 65.8% for AI-legitimate content and 75.1% for AI-fake content.
- 4.3 Delivery: Deliberate responses above 34 seconds improved origin accuracy from 47.8% to 56.4% and veracity accuracy from 67.3% to 74.1%.These median-split differences are descriptive and exploratory because responses were nested within participants.
- 4.4 Exploitation: Fake-news detection declined 10.2 percentage points from 71.2% to 60.9% across response bins 1–10 and 21–30, while AI-origin detection stayed near 56%.The trend was descriptive and exploratory, based on aggregated response positions and pending repeated-measures analysis.
- 4.5 Post-Exploitation: Near-midpoint scores for high-quality LLMs suggest inconsistent human-detectable artifacts and motivate stronger empirical grounding for provenance-based defenses.The Post-Exploitation interpretation remains indirectly supported and calls for longitudinal belief-persistence measurements.
5 Discussion and Implications
The discussion frames measured cognitive vulnerabilities as opportunities for proactive, stage-specific defense, while emphasizing that proposed interventions remain testable directions rather than validated countermeasures. It also identifies important boundaries: controlled text-only data, limited geographic representation, model-specific results, and interpretive kill-chain mapping.
- Intervention opportunities: The paper maps empirical perception findings onto concrete intervention opportunities across the adapted kill chain.The framework assigns possible actions to platforms, AI developers, educators, and policymakers while keeping claims modest.
- Intervention opportunities: Platform operators could add resharing friction and pace dense misinformation exposure to support deliberate evaluation and counter fatigue.Suggested measures include interstitials, dwell-time prompts, and rate-limiting dense misinformation feeds.
- Intervention opportunities: AI developers and standards bodies could use provenance credentials and watermarking to backstop synthetic text that humans detect poorly.Domain-level credibility priors such as CRED-1 are presented as a complementary on-device signal for Delivery-stage prebunking.
- Intervention opportunities: Educators could target the perception-accuracy gap with calibration exercises and prebunking that teach when confidence should be distrusted.The proposed cognitive scaffolding focuses on calibration, not only factual instruction.
- Intervention opportunities: Policymakers could require provenance disclosure and support independent evaluation alongside platform- and education-level measures.The paper presents these measures as complementary rather than as independently validated solutions.
- Limitations and future work: The study’s scope is constrained by controlled text-only testing, a primarily European/North American sample, model-specific results, and interpretive stage mapping.Social-platform dynamics, multimodal content, evolving models, and experimental validation of individual stages remain outside the demonstrated scope.
- Limitations and future work: Future work should link stages and data through integrative modeling, cross-language analysis, automated-detector comparisons, multimodal studies, and ecological longitudinal designs.These directions include mixed-effects and multivariate regression accounting for per-participant nesting.
6 Conclusion
The paper reports human-subject evidence organized through an adapted cybersecurity kill chain, showing cognitive vulnerabilities that motivate a shift from reactive fact-checking toward proactive, stage-specific defense. It argues that locating where people become vulnerable is as important as understanding how AI generates misinformation.
- Conclusion: The study analyzed 504 participants and 2,438 judgments using an adapted cybersecurity kill chain.The framework maps human perception onto stages of the cognitive attack lifecycle.
- Conclusion: The perception-accuracy gap, LLM weapon efficacy, and asymmetric cognitive fatigue effect jointly motivate proactive, stage-specific defense.The conclusion contrasts this approach with reactive fact-checking.
- Conclusion: Mapping perception onto the attack lifecycle identifies where—and in whom—to interrupt AI-driven misinformation.The conclusion presents vulnerability location as complementary to understanding AI generation.