Source-linked AI summary

Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations

Josh A. Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, Katerina Sedova

arXiv:2301.04246v1cs.CY

TL;DR

Generative language models may automate scalable and persuasive propaganda, creating an emerging influence-operations threat. The paper analyzes how these models could change actors, behaviors, and content, and organizes mitigations across the AI-to-target pipeline. It concludes that language models are likely to significantly affect influence operations, that no silver bullet exists, and that coordinated combinations of mitigations are needed.

  • Problem

    Generative language models may enable scalable and persuasive influence operations, creating an emerging threat that requires analysis.

  • Method

    The paper analyzes effects across actors, behaviors, and content and classifies mitigations across model construction, access, dissemination, and belief formation.

  • Results

    Language models are likely to significantly impact the future of influence operations, and no silver bullet exists for minimizing AI-generated disinformation risks.

  • Takeaways & Limitations

    Responses may require multiple mitigation strategies and new coordination among AI providers, social media platforms, governments, and civil society.

  • Takeaways & Limitations

    The paper focuses on language models and covert propaganda rather than other AI models, other information-control behaviors, or specific actors.

Abstract

from arXiv · show

Generative language models have improved drastically, and can now produce realistic text outputs that are difficult to distinguish from human-written content. For malicious actors, these language models bring the promise of automating the creation of convincing and misleading text for use in influence operations. This report assesses how language models might change influence operations in the future, and what steps can be taken to mitigate this threat. We lay out possible changes to the actors, behaviors, and content of online influence operations, and provide a framework for stages of the language model-to-influence operations pipeline that mitigations could target (model construction, model access, content dissemination, and belief formation). While no reasonable mitigation can be expected to fully prevent the threat of AI-enabled influence operations, a combination of multiple mitigations may make an important difference.

1 Introduction

The paper examines how generative language models may reshape influence operations and develops mitigation frameworks while emphasizing uncertainty, scope limits, and the need for combined responses.

  • 1.1 Motivation: The paper asks how language models could affect future influence operations and what mitigation strategies could address those threats.It frames the analysis around emerging capabilities, current limitations, and critical unknowns.
  • 1.2 Potential Effects: Language models may lower propaganda costs, expand the number of actors able to wage campaigns, and give propagandists competitive advantages.These effects follow from models’ potential to automate text production and rival human-written content at low cost.
  • 1.2 Potential Effects: Language models may make influence operations easier to scale, reduce the cost of personalization, enable real-time chatbot tactics, and produce more impactful or less discoverable content.The paper organizes these predicted changes across actors, behavior, and content.
  • 1.4 Outline of the Report: The report classifies mitigations across model construction, model access, content dissemination, and belief formation.It considers a range of strategies rather than a single intervention.
  • 1.3 Scope and Limitations: The study focuses on language models and covert propaganda rather than other AI models, other information-control behaviors, or specific actors.It primarily analyzes technology and trends across possible settings, while acknowledging author-background bias and the need for further research.
  • 1.4 Outline of the Report: No reasonable mitigation is expected to fully prevent AI-enabled influence operations, but multiple strategies may make an important difference and require cross-sector collaboration.The proposed collaborations include social media platforms, AI companies, government agencies, and civil society actors.

2 Orienting to Influence Operations

Influence operations are covert or deceptive efforts to shape target audiences’ opinions, using tactics that can persuade, activate, distract, or alter perceptions. Their effects matter but remain difficult to measure because engagement is an inadequate proxy and causal attribution is often unavailable.

  • 2.1 What Are Influence Operations, and Why Are They Carried Out?: The paper defines influence operations as covert or deceptive efforts to influence a target audience’s opinions.The definition does not depend on whether the message is true or false or on the actor’s identity.
  • 2.1 What Are Influence Operations, and Why Are They Carried Out?: Influence operations may activate believers, persuade audiences, or distract them by competing for scarce attention and diluting unfavorable narratives.Distraction can involve spreading alternative theories or other content that diverts attention.
  • 2.1 What Are Influence Operations, and Why Are They Carried Out?: Common tactics include promoting one’s government or policies, advocating for or against policies, improving allies’ reputations, and destabilizing rivals.These tactics target domestic affairs, foreign relations, or perceptions of political actors.
  • 2.1 What Are Influence Operations, and Why Are They Carried Out?: Operations often conceal information sources through fake personas, news properties, and inauthentic amplification across major, alternative, small-group, and encrypted platforms.Accounts may masquerade as members of targeted communities, making subtle linguistic cues relevant to detection.
  • 2.1 What Are Influence Operations, and Why Are They Carried Out?: Since 2016, Meta and Twitter have removed well over a hundred influence operations, though publicly reported cases likely undercount the total.Influence operations may be domestic or foreign and are designed to remain secret.
  • 2.2 Impacts of Influence Operations: Influence operations can affect beliefs and consequential behavior, but their effectiveness is usually difficult to measure.Engagement metrics such as clicks and shares do not adequately measure social influence, while comparison groups, attribution, detection, and multicausal opinion change remain difficult.

3 Recent Progress in Generative Models

Generative models have advanced rapidly across modalities, producing increasingly realistic outputs while becoming cheaper and more accessible to develop and use. Progress has been driven by larger datasets, improved algorithms, and increased computational power, with fine-tuning offering a lower-cost alternative to training from scratch.

  • Generative models learn patterns from data to produce original digital content, including images, video, audio, and text.
  • Image-generation systems progressed from early text-to-image methods to models capable of elaborate scene construction and composition from detailed language prompts.
  • Training frontier models can cost at least tens of millions of dollars, although GPT-3-level performance was later reached from scratch for less than $500k.
  • Fine-tuning existing foundation models can provide general capabilities for specialized tasks at lower cost than training a model from scratch.
  • Recent progress has been driven by more internet-scale training data, improved neural-network algorithms, and greater computational power for larger models.
  • Moderately capable models are publicly available, while the most capable models remain private or behind monitorable APIs; nation-states and many non-state actors can likely fine-tune public models.

4 Generative Models and Influence Operations

The report connects influence operations with generative models by analyzing how models may transform actors, behaviors, and content. It then considers technical developments and uncertainties, linking expected improvements to future campaign implications.

  • The section applies the actors, behaviors, and content framework to analyze how generative models may transform influence operations.
  • It examines expected machine-learning developments by describing current technology, anticipated improvements, and implications for future influence campaigns.
  • The analysis also identifies critical unknowns that will affect the role generative models can play in influence operations.

4.1 Language Models and the ABCs of Disinformation

Language models could reshape influence operations across actors, behaviors, and content by lowering costs, enabling new tactics, and improving message quality and tailoring.

  • The ABC model distinguishes influence operations by actors, behaviors, and content, revealing how language models may affect each manipulation vector.
  • Actors: Language models could expand the pool of actors able to run campaigns by automating content production, persona creation, and culturally appropriate outputs.Lower costs reduce the resources required to maintain fake personas and streams of relevant content.
  • Behavior: Replacing or augmenting human writers could reduce costs and increase the scalability of mass messaging and unattributable long-form propaganda.
  • Behavior: Generative models could scale cross-platform testing and improve existing influence tactics, potentially increasing campaigns’ overall impact.
  • Behavior: Models could enable real-time, dynamic, and demographic-tailored content, including personalized chatbots that engage targets in one-on-one conversations.The cost-effectiveness of tailoring depends on how well models can use limited demographic information.
  • Content: Language models could improve short- and long-form text quality while producing varied, personalized, and narratively aligned content that is harder for existing detectors to identify.Such content could obscure repeated messaging patterns and support narrative laundering through apparently reputable sources.

4.2 Expected Developments and Critical Unknowns

The report frames future influence-operation impacts around improving model usability, reliability, and efficiency while emphasizing uncertainty about capabilities and deployment conditions.

  • The report presents expected developments as plausible medium-term scenarios rather than explicit forecasts, highlighting critical unknowns that could shape future influence operations.
  • Expected Developments: Usability, reliability, and efficiency are identified as three model features likely to affect deployment in influence operations.
  • Usability and Reliability: Improved usability and reliability could let lower-skilled propagandists use models with less oversight and make existing capabilities cheaper and more efficient.
  • Usability and Reliability: Current models often require skilled operators and monitoring because plausible outputs may contain repetition, incoherence, fabricated facts, or failures to follow complex tasks consistently.
  • Reliability: Lack of current-event awareness limits reliability because models trained once on static corpora may lack context for events after training.
  • Reliability: Reliability may improve through continual retraining, targeted model updates, retrieval from external databases, or task-specific fine-tuning.
  • Efficiency: Efficiency improvements from algorithms, hardware, and inexpensive fine-tuning could reduce the cost of automating influence tactics.
  • Operational Innovation: Future organizations may improve human-machine collaboration through software for overseeing, selecting, correcting, and adapting model outputs.

New Capabilities as a Byproduct of Scaling and Research

Scaling and broader AI research may produce language-model capabilities that were not explicitly targeted. Some of these emergent or side-effect capabilities could become useful to influence operators.

  • Training language models on next-word prediction can yield adjacent capabilities such as summarization, style-specific generation, and other uses.
  • Larger language models already support summarization, translation, basic analogies, and basic conversation, while the emergence of further capabilities remains difficult to predict.
  • Broader AI research may unintentionally produce capabilities relevant to influence operations, because small adjustments can uncover capabilities beyond the original target.

Models Specialized for Influence

Language models could become more useful for influence operations through targeted training, broader capabilities, and combinations with other technologies. Access, tooling, and actors’ willingness to deploy them remain critical uncertainties.

  • Propagandists could intentionally improve models through targeted training, greater generality, or combinations with other technologies.
  • Targeted training could use click-through data or other influence proxies to optimize models for more persuasive text.
  • Systems such as GPT-3 have produced slanted news articles without explicit task-specific training, suggesting future models might generate tailored persuasive texts or sustain long-lived dialogue.
  • Combining modalities could let a single system ingest online images, respond to them, and generate fabricated images and text.
  • Generative models paired with automation software could support more human-like bots, including systems that identify compatible profiles and generate targeted messages.
  • The impact of these capabilities depends on access, unregulated tooling, intent to use, and whether social norms constrain deceptive deployment.

Willingness to Invest in Generative Models

Influence operations may gain access to advanced language models through repurposing, theft, or dedicated development, but large investments depend on resources, feasibility, and expected returns.

  • Propagandists could repurpose or steal state-of-the-art models, or train models specifically for influence operations.
  • Investment in general-purpose models by governments, firms, or wealthy individuals could increase propagandists’ access through legitimate channels or theft.
  • A well-resourced propagandist could fund a dedicated system, requiring extensive computation, bespoke data, and engineering talent.
  • Diminishing returns from advanced capabilities could make large, influence-specific investments less attractive to propagandists.

Greater Accessibility from Unregulated Tooling

Easy-to-use tooling can lower the operational barrier to AI-enabled influence operations. Its proliferation may broaden language-model adoption and enable more automated campaign capabilities.

  • Using language models for propaganda can require operational know-how, including carefully designing inputs or running models on dedicated infrastructure.
  • AI-generated profile pictures are commonplace in influence operations and have also been used for deceptive commercial purposes.
  • Without easy-to-use tooling, influence operations might not have adopted AI-generated profile pictures, or might have used them less extensively.
  • Language-model tools could lower barriers for propagandists lacking machine-learning expertise and support automated chatbots targeting people selected by bad actors.

Norms and Intent-to-use

Norms may constrain whether actors use language models for influence operations, even when the technology is accessible and inexpensive. The report discusses norm development as a possible way to discourage political actors and developers from enabling such use.

  • Norms can constrain state behavior beyond cost-benefit incentives, suggesting they may also influence decisions about language-model-enabled influence operations.The passage gives examples including nuclear weapons, assassinations, and mercenary use.
  • Access to capable models and low deployment costs do not ensure that an actor will build and deploy them for influence operations.The actor must still decide to use the capability, while norms could constrain political actors and encourage developers to inhibit misuse.
  • Creating a norm against using language models for propaganda could involve state coalitions, compliance mechanisms, and advocacy by researchers or ethicists.The report describes both international and substate pathways for norm entrepreneurship.

5.1 A Framework for Evaluating Mitigations

The report organizes language-model influence-operation mitigations around four intervention stages and evaluates them using feasibility, risk, and impact criteria. It concludes that effective mitigation will likely require coordinated combinations rather than a single solution.

  • Intervention framework: Mitigations target four pipeline stages: model construction, model access, content dissemination, and belief formation.These stages correspond to creating capable models, obtaining reliable access, spreading content, and limiting audience influence.
  • Evaluation criteria: The framework evaluates each mitigation by technical feasibility, social feasibility, downside risk, and impact.These criteria distinguish implementability, institutional and political viability, negative externalities, and threat reduction.
  • Scope and limitations: The report does not classify mitigations as worth trying or fine to ignore, leaving stakeholders to weigh their individual advantages and disadvantages.The authors also identify norm development, retaliation, and harms from newly formed beliefs as outside the model.

5.2 Model Design and Construction

Model construction can be targeted through detectable outputs, training-data restrictions, and hardware controls, but each approach faces important technical, legal, coordination, or enforcement limits. The report emphasizes that preventing model construction altogether would be most reliable but is unlikely.

  • Construction choices: Not building large language models would most reliably prevent their misuse, but a complete halt to new model development is extremely unlikely.The section therefore focuses on changing model construction rather than stopping development entirely.
  • Detectable outputs: AI-generated text detection is difficult because larger models produce text harder to distinguish from human writing, while proposed fingerprinting methods remain uncertain.Radioactive-data training and parameter perturbations might improve detectability, but text-based approaches have not been extensively validated.
  • Detectable outputs: Attackers could switch to models without traceability features, so detectable-model strategies require broad coordination among developers and may impose costs rather than eliminate capability.Adversaries able to build their own models may face additional costs without losing access to the capability.
  • Training-data interventions: Radioactive-data strategies raise ethical concerns, may affect only models trained in the same language, and have uncertain requirements for meaningful internet coverage.The report also notes uncertainty about whether the approach would work effectively for text.
  • Training-data interventions: Regulating web scraping could slow model growth, but such measures are out of step with the current United States regulatory environment.The report notes the absence of comprehensive data-privacy laws and relevant legal constraints on scraping publicly available data.
  • Hardware controls: A model 200 times larger than the current largest model could potentially be trained using less than 0.5% of worldwide cloud-computing resources, complicating compute monitoring.Compute is also a general resource that does not reveal whether an organization is training a language model or running another project.
  • Hardware controls: Export controls on semiconductors and related equipment may slow language-model development, but effectiveness depends on enforcement and avoiding stockpiling or re-exports.Restrictions may also incentivize indigenous chip production by affected states.

5.3 Model Access

AI providers can restrict model access through purpose checks, application limits, trusted-user policies, output caps, and input screening. These measures may reduce misuse, but their effectiveness depends on coordination across providers and on the capabilities of publicly released models.

  • Coordination limits: Provider restrictions create a collective-action problem because commercial incentives encourage defection and propagandists can move to less restrictive models.Restrictions are less effective when publicly released models are sufficiently capable for propagandists.
  • Access restrictions: API-based access regimes let providers require stated purposes, restrict indirect user access, limit outputs, screen inputs, or serve only trusted institutions.These options target both direct misuse and applications that expose models to additional users.
  • Access restrictions: Access restrictions can limit the scale of influence operations without necessarily preventing smaller tailored campaigns.Output caps may constrain high-volume activity while leaving lower-volume article generation possible.
  • Coordination limits: Industry norms, standards, or government regulation could support broader adoption of strong access restrictions and impose costs when private models remain more capable than open-source alternatives.The durability and force of emerging provider guidelines remain uncertain.
  • Alternative mechanisms: Alternative access controls include specialized-hardware requirements and model-use licenses, but further research is needed to assess restricted-access mechanisms.These approaches may support both access control and attribution.
  • Research norms: AI research norms have traditionally favored openness, while proposed governance norms include staged release, risk frameworks, misuse reporting, and prepublication safety review.The report does not make specific claims about which research norms are desirable.

5.4 Content Dissemination

The report evaluates mitigations for disseminating AI-generated influence content, emphasizing platform–AI-company collaboration while recognizing technical, institutional, privacy, and evasion limits.

  • AI-generated content matters only if it reaches and influences people, so mitigation must address dissemination as well as generation.
  • Platforms are unlikely to impose blanket bans because AI-generated content has legitimate uses, including parody, comedy, and automated announcements.
  • Detection based only on text statistics and user metadata lacks sufficient confidence for disruptive action and performs poorly on typical social-media posts.
  • Platform–AI-company collaboration could trace stored model outputs to users and identify broader reposting patterns across platforms.
  • Monitoring publicly posted content may be feasible but is difficult on encrypted or noncooperative platforms, requiring substantial bilateral coordination.
  • Hash-based detection is imperfect because propagandists can evade it through small output changes, while close-match methods introduce greater statistical uncertainty.
  • Authentication requirements could disrupt fully automated bot operations but would not prevent humans from copying model outputs into social-media posts.
  • User authentication faces privacy resistance and can be circumvented through burner accounts or inexpensive labor completing personhood checks.

5.5 Belief Formation

The report examines demand-side responses to AI-enabled influence operations, from media literacy to consumer-facing AI tools. These approaches can help users evaluate information, but require updating and careful implementation because defensive systems introduce their own risks.

  • 5.5 Belief Formation: Mitigations that address misinformation supply remain partial when audiences continue seeking information tailored to their beliefs.
  • 5.5.1 Institutions Engage in Media Literacy Campaigns: Media literacy campaigns can improve discernment between real and fake news, but AI systems may avoid the behavioral cues current programs teach people to spot.
  • 5.5.1 Institutions Engage in Media Literacy Campaigns: Media literacy programs therefore require updating as AI-generated influence tactics evolve.
  • 5.5.2 Developers Provide Consumer-Focused AI Tools: Consumer-focused AI tools could help users identify, critically evaluate, and curate information, reducing demand for disinformation.
  • 5.5.2 Developers Provide Consumer-Focused AI Tools: Proposed tools include warning labels, fake-account detection, selective ad blocking, and AI-assisted source vetting, scoring, ranking, and curation.
  • 5.5.2 Developers Provide Consumer-Focused AI Tools: Contextualization engines could connect sources with related high-quality material and identify areas where relevant data is missing.
  • 5.5.2 Developers Provide Consumer-Focused AI Tools: Defensive generative models could explain flaws in tailored arguments, identify manipulated-image artifacts, and help consumers find relevant information.
  • 5.5.2 Developers Provide Consumer-Focused AI Tools: These tools may serve individuals as well as businesses, governments, and organizations seeking greater awareness of influence operations.

6 Conclusions

The report concludes that language models are likely to reshape influence operations, but no single mitigation can address the threat. Effective responses require coordinated institutions, attention to both information supply and demand, and further research.

  • Language models are likely to significantly impact the future of influence operations across actors, behaviors, and content.
  • Wider access could lower propaganda costs, attract additional actors, and enable political campaigns to outsource operations to private firms.
  • Models may enable dynamic responses, automated cross-platform testing, and other unforeseen tactics shaped by evolving defenses.
  • Language models could reduce propaganda costs, increase its scale, and produce persuasive text that is difficult to distinguish from human-generated content.
  • No silver-bullet mitigation is simultaneously technically feasible, institutionally tractable, robust against second-order risks, and highly impactful.
  • Meaningful mitigation requires collaboration among AI developers, social-media companies, policymakers, researchers, and other stakeholders.
  • Supply-side interventions are only partial if demand for misleading information remains unchanged.
Loading 2301.04246v1…