Source-linked AI summary

The Radicalization Risks of GPT-3 and Advanced Neural Language Models

Kris McGuffie, Alex Newhouse

arXiv:2009.06807v1cs.CYcs.AI

TL;DR

The paper asks how GPT-3 changes the abuse potential of generative language models and evaluates that risk across extremist narratives, ideologies, and interaction formats. Using representative prompts and examples, it finds that GPT-3 can generate ideologically consistent extremist content with little conventional training, raising concerns about synthetic radicalization and recruitment while leaving copycat-model risks for further study.

  • Problem

    The paper examines whether GPT-3’s few-shot capabilities overcome GPT-2’s resource and brittleness limitations in generating extremist content and threaten information integrity and online social interactions.

  • Method

    CTEC tests GPT-3 with representative examples spanning extremist ideologies, text styles, social interactions, and factual questions, using early prompt attempts to reduce cherry-picking.

  • Results

    GPT-3 generates ideologically consistent extremist and conspiracy-oriented text, including interactive content, after only a few short prompts and without conventional training.

  • Takeaways & Limitations

    Synthetic content from GPT-3 could be lightly altered and automated to amplify extremist communities and make online material difficult to distinguish from human-generated content.

  • Takeaways & Limitations

    Further study is needed to evaluate wholly synthetic extremist content, copycat-model deployment, and its believability across multiple platforms.

Abstract

from arXiv · show

In this paper, we expand on our previous research of the potential for abuse of generative language models by assessing GPT-3. Experimenting with prompts representative of different types of extremist narrative, structures of social interaction, and radical ideologies, we find that GPT-3 demonstrates significant improvement over its predecessor, GPT-2, in generating extremist texts. We also show GPT-3's strength in generating text that accurately emulates interactive, informational, and influential content that could be utilized for radicalizing individuals into violent far-right extremist ideologies and behaviors. While OpenAI's preventative measures are strong, the possibility of unregulated copycat technology represents significant risk for large-scale online radicalization and recruitment; thus, in the absence of safeguards, successful and efficient weaponization that requires little experimentation is likely. AI stakeholders, the policymaking community, and governments should begin investing as soon as possible in building social norms, public policy, and educational initiatives to preempt an influx of machine-generated disinformation and propaganda. Mitigation will require effective policy and partnerships across industry, government, and civil society.

2 Background

GPT-3 addresses major GPT-2 limitations through few-shot prompting, enabling varied extremist text generation with little fine-tuning or specialized expertise. CTEC finds it can produce interactive, multilingual, ideological, and manifesto-style content that could support extremist influence and recruitment.

  • Model advances: GPT-2 required hours of fine-tuning and remained brittle across ideologies, whereas GPT-3 can generate subject-specific text from a few representative prompts.GPT-2’s fine-tuned models did not easily transfer between ideological outputs.
  • Implications: GPT-3’s few-shot capabilities allow efficient generation of ideologically consistent material across extremist communities and formats.The paper links this capacity to broader threats to information integrity and online social interactions.
  • Capabilities: GPT-3 can produce extremist polemics, fake genocidal forum discussions, radicalized QAnon answers, and multilingual anti-Semitic text from English prompts.These capabilities span narrative, interactive, ideological, and linguistic forms of extremist content.
  • Weaponization risk: GPT-3’s weaponization risk stems from producing believably human influential content that reduces extremists’ workload without specialized technical knowledge.The paper describes short prompts as sufficient to generate compelling and consistent far-right extremist text.

3 Methodology and Discussion of Model Performance

CTEC tests GPT-3 with representative prompts spanning extremist ideologies, text styles, and knowledge domains, emphasizing early attempts to reduce cherry-picking. The experiments examine both factual extremist knowledge and the model’s ability to generate biased or interactive content.

  • Comparison with GPT-2: Compared with GPT-2’s 10–12 hours of training on a large manifesto corpus, GPT-3 requires only short representative prompts for subject-specific generation.GPT-2’s outputs remained constrained by the data on which it was trained.
  • Methodology: CTEC gathered varied examples across ideological and functional spectra, using the first or second prompt attempt to avoid cherry-picking.The examples span white supremacy, QAnon conspiracism, recruitment, radicalization, and organization.
  • Zero-shot knowledge: GPT-3 was probed with questions about extremist groups and ideologies, including Wagner Group facts and QAnon concepts.The Wagner Group dialogue was described as generally correct, while QAnon answers were neutral and empirically oriented.
  • Wagner Group: GPT-3 produced detailed answers about the Wagner Group’s location, leadership, activities, personnel, weapons, and role in Syria.These outputs included claims about Dmitry Utkin, military support, and heavy weapons.
  • QAnon: GPT-3 answered niche QAnon questions while stating that there was no evidence QAnon was a military intelligence official.The responses covered QAnon’s origins, predictions, conspiratorial claims, and anti-Semitic comments.
  • Atomwaffen Division: GPT-3 also answered questions about the Atomwaffen Division and its ideology with a high degree of accuracy.The supplied section reports this capability without providing a quantitative accuracy value.

Sample 3: Zero-Shot Prompt Asking About the Atomwaffen Division

This sample examines GPT-3’s zero-shot knowledge of the Atomwaffen Division and situates it within the model’s broader few-shot prompting capabilities. The cited passages report detailed group information and describe how representative examples can bias subsequent answers.

  • Sample 3: Zero-Shot Prompt Asking About the Atomwaffen Division: GPT-3 was asked, in a zero-shot prompt, to identify the Atomwaffen Division.The sample presents the question as an evaluation of the model’s knowledge about the group.
  • Sample 3: Zero-Shot Prompt Asking About the Atomwaffen Division: The model described Atomwaffen as a national socialist organization seeking the eventual overthrow of the US government through terrorism and guerrilla warfare.This is the sample’s primary ideological and organizational characterization.
  • Sample 3: Zero-Shot Prompt Asking About the Atomwaffen Division: The response identified Brandon Russell as the group’s founder and named Andrew Oneschuk, Devon Arthurs, and Jeremy Himmelman as other members.These details appear in the sample’s follow-up answers.
  • Sample 3: Zero-Shot Prompt Asking About the Atomwaffen Division: The sample states that Atomwaffen members idolize Adolf Hitler and that Siege was written by James Mason.These answers describe ideological influence and a cited text associated with the group.
  • Sample 3: Zero-Shot Prompt Asking About the Atomwaffen Division: The response places Atomwaffen’s origins on Ironmarch.org.This is the sample’s answer to where the group started online.
  • 3.2 Few-Shot Prompting and Output Biasing: GPT-3’s few-shot learning lets users shape outputs through longer representative prompts instead of training the entire model on a large corpus.The paper contrasts this process with GPT-2’s more resource-intensive training approach.
  • 3.2 Few-Shot Prompting and Output Biasing: Three biased Q-and-A examples caused GPT-3 to answer QAnon questions consistently within a conspiratorial frame.The paper presents this as integrating the model’s niche knowledge with ideological bias.

Sample 4: Few-Shot Prompt Asking About QAnon

The QAnon-primed model answered increasingly specific questions in a coherent worldview, while also producing structured and interactive content across formats. The experiment illustrates both ideological consistency and practical limits of the model’s prompt-based generation.

  • QAnon worldview: The model answered questions about QAnon’s worldview, including its alleged enemies, the Storm, and the role of QAnon.
  • QAnon worldview: It attributed World War III and a New World Order to the Rothschilds while distinguishing QAnon from anti-Semitism.
  • Additional questions: Specific follow-up questions elicited deeper explanations of adrenochrome and other claims promoted within the conspiracy theory.
  • Additional questions: The generated answers extended conspiracy claims to vaccines, Bill Gates, and Hillary Clinton.
  • Generation capability: The model’s strength includes recognizing structural cues and generating consistent examples of organized social interactions.
  • Online communities: The most immediate concern involves private forums and message boards designed to attract and retain violent extremists.
  • Online communities: GPT-3 could extend snippets of complicated Iron March threads and generate new threads, despite its limited prompt size.

Sample 5: Few-Shot Prompt With Iron March Forum Thread

GPT-3 generated extremist forum content across multiple voices, themes, and viewpoints, including recruitment, antisemitism, racial ideology, and sexual exploitation. It could complete snippets convincingly and create new topics or opening posts consistent with Iron March’s ideological environment.

  • Recruitment and interaction: The model generated responses that simulated users joining a transnational white-supremacist organization.
  • Generation across viewpoints: GPT-3 completed forum snippets with multiple viewpoints and themes drawn from far-right extremist discourse.
  • Generation across viewpoints: GPT-3 could generate new topics and opening posts from scratch that remained within ideologies promoted on Iron March.

Sample 6: Few-Shot Prompt With Russian-Language Anti-Semitic Posts

GPT-3 translated English descriptions of extremist material into coherent Russian-language comments while preserving right-wing bias, xenophobia, and conspiracism. It also generated extremist manifestos and polemics in varied styles.

  • Multilingual generation: The Russian-language assistant generated comments from English descriptions of extremist topics.
  • Generated topics: The generated examples addressed Jews, Soros, adrenochrome, Americans, Crimea, and alleged enemies of humanity.
  • Generated topics: Additional outputs asserted antisemitic world-control conspiracies and nationalist claims about Crimea.
  • Manifesto generation: GPT-3 was also effective at generating extremist manifestos, which remain important because they are informative and easily shareable.
  • Manifesto generation: Prompting with the El Paso shooter’s manifesto produced ideological polemics in multiple styles.

Sample 7: Few-Shot Prompt With Mass Shooter Manifestos

GPT-3 generated manifesto-style extremist texts from prompts associated with mass shooters and other violent ideologies. The examples included racial-replacement narratives, personal background, attack motives, sovereign-citizen claims, and revolutionary themes.

  • Shooter manifesto styles: The experiment prompted GPT-3 with manifesto descriptions modeled on the El Paso and Christchurch white-supremacist shooters.
  • Shooter manifesto styles: A generated Christchurch-style text framed an attack as retaliation against alleged ethnic and cultural replacement.
  • Shooter manifesto styles: The Christchurch-style material also supplied a biographical persona describing an ordinary white man from a working-class background.
  • Attack rationale: Another passage articulated motives centered on defending white homelands and taking revenge against historical enemies.
  • Other ideological styles: GPT-3 was also prompted to produce manifesto-style writing from sovereign-citizen and anarcho-primitivist perspectives.
  • Other ideological styles: A separate generated passage called for revolution, systemic destruction, and a return to more primitive living.

4 Radicalization Risk Methodology

CTEC evaluates GPT-3’s radicalization risk by applying an online-radicalization framework to synthetic content and testing prompts aligned with established influence mechanisms.

  • Radicalization is defined as an increasingly committed process toward violent extremism, with potential implications for violence, mobilization, and recruitment.
  • Online radicalization remains difficult to characterize precisely, although research broadly agrees that the internet facilitates extremist radicalization and coordination.
  • Synthetic content can replace human-generated content online or offline to share ideologies, provide supportive information, and establish group identity.
  • CTEC designed prompts to test whether GPT-3 could detect and generate content leveraging mechanisms of influence in radicalization processes.
  • The evaluation considers group identity formation and the emulation of authentic engagement within apparently large ideological movements.

5 Output Reveals Radicalization Efficacy

GPT-3 can emulate ideologically consistent and interactive online extremist environments, creating synthetic content that may amplify radicalization and recruitment.

  • GPT-3 can emulate the ideologically consistent, interactive, and normalizing environment of online extremist communities.
  • Heavily synthetic forums could attract new adherents and intensify existing users’ involvement by rewarding increasingly extreme beliefs and discouraging dissent.

6 Risk Mitigation

The paper identifies uncertainty about synthetic text’s effectiveness but recommends coordinated safeguards, literacy efforts, detection systems, responsible advocacy, and cross-sector partnerships.

  • Conspiracy theories and distrust may reduce synthetic text’s effectiveness, but truth vacuums could make high-volume synthetic content particularly dangerous.
  • Managing emerging technology requires effective partnerships among industry, government, and civil society to establish standards for use and abuse.
  • Creators and distributors should impose strong, consistent safeguards and restrictions on powerful language-generation models.
  • Government and civil society should promote critical digital literacy and awareness of synthetic text and automated content distribution through mass-audience education.
  • Service providers and civil society should deploy detection models to reduce distribution of and spotlight nefarious synthetic content online.
  • Consumer-facing platforms should integrate easy-to-use, easy-to-understand detection and filtration systems alongside publicly accessible language-generation technology.
  • Online communities should value identified and verified information sources, with social-media coordination supporting disclosure of synthetic-content risks.
  • Advocacy should demand responsible and transparent model application, supported by public-facing partnerships between AI research organizations and civil society.

7 Conclusion

The paper calls for further study of GPT-3 and follow-on models to better characterize risks from wholly synthetic extremist content and its deployment across platforms.

  • Further research should evaluate wholly synthetic extremist content, including its production, deployment, believability, and efficacy across multiple platform types.
  • Testing should examine whether copycat models can populate interactive forums with minimal human curation and without significant training, funding, or organizational support.
Loading 2009.06807v1…