Source-linked AI summary

Red-Teaming for Generative AI: Silver Bullet or Security Theater?

Michael Feffer, Anusha Sinha, Wesley Hanwen Deng, Zachary C. Lipton, Hoda Heidari

arXiv:2401.15897v3cs.CYcs.HCcs.LG

TL;DR

Generative AI's growing risks have made red-teaming prominent, but its meaning, regulatory role, and relationship to cybersecurity practice remain unclear. The paper analyzes industry cases, research literature, and NIST RFI comments, finding substantial variation and arguing that red-teaming is not a universal safety solution. It proposes a question bank to scaffold more precise future evaluations.

  • Problem

    Public definitions and practices of AI red-teaming lack consensus about its scope, role in regulation, evaluation targets, and relationship to broader safety assessment.

  • Method

    The paper analyzes recent industry red-teaming cases, surveys relevant research literature, and examines public comments submitted to a NIST request for information.

  • Results

    AI red-teaming practices diverge across threat models, artifacts, settings, methods, actors, resources, and resulting reporting, disclosure, and mitigation decisions.

  • Takeaways & Limitations

    Red-teaming can be a broad framing for GenAI harm evaluation, but treating it as a panacea risks security theater; the paper offers a question bank for future practice.

  • Takeaways & Limitations

    Red-teaming exercises cover limited vulnerability sets and therefore cannot guarantee safety from all angles.

Abstract

from arXiv · show

In response to rising concerns surrounding the safety, security, and trustworthiness of Generative AI (GenAI) models, practitioners and regulators alike have pointed to AI red-teaming as a key component of their strategies for identifying and mitigating these risks. However, despite AI red-teaming's central role in policy discussions and corporate messaging, significant questions remain about what precisely it means, what role it can play in regulation, and how it relates to conventional red-teaming practices as originally conceived in the field of cybersecurity. In this work, we identify recent cases of red-teaming activities in the AI industry and conduct an extensive survey of relevant research literature to characterize the scope, structure, and criteria for AI red-teaming practices. Our analysis reveals that prior methods and practices of AI red-teaming diverge along several axes, including the purpose of the activity (which is often vague), the artifact under evaluation, the setting in which the activity is conducted (e.g., actors, resources, and methods), and the resulting decisions it informs (e.g., reporting, disclosure, and mitigation). In light of our findings, we argue that while red-teaming may be a valuable big-tent idea for characterizing GenAI harm mitigations, and that industry may effectively apply red-teaming and other strategies behind closed doors to safeguard AI, gestures towards red-teaming (based on public definitions) as a panacea for every possible risk verge on security theater. To move toward a more robust toolbox of evaluations for generative AI, we synthesize our recommendations into a question bank meant to guide and scaffold future AI red-teaming practices.

1 Introduction

Generative AI's expanding use has raised concerns about societal harms, while policymakers and practitioners increasingly position red-teaming as a way to identify and manage related risks. The paper examines what AI red-teaming entails and how its scope, evaluation setting, and outcomes should be characterized.

  • Generative AI adoption has raised concerns about discriminatory, stereotypical, and otherwise harmful outputs, alongside limited transparency and accountability.
  • The US Executive Order defines AI red-teaming as structured, often adversarial testing to identify flaws, vulnerabilities, undesirable behaviors, limitations, and misuse risks.
  • Public definitions leave unresolved which risks red-teaming can address, how activities should be structured, who should participate, and how findings should inform mitigation.
  • The study combines publicly documented industry cases, an extensive literature survey, thematic analysis, and analysis of public comments responding to a NIST request for information.
  • The authors find limited consensus across threat models, evaluation artifacts, settings, actors, resources, methodologies, reporting, disclosure, and mitigation decisions.
  • They argue that red-teaming can provide a broad framing for GenAI evaluation, but treating it as a catch-all response to safety concerns risks security theater.

2 Related Contemporary Work

The paper situates AI red-teaming within its military and cybersecurity history and introduces a question bank for specifying the artifact and conditions of evaluation. The proposed questions address model version, safeguards, lifecycle stage, and release conditions.

  • Red-teaming originated in warfare and religious contexts before the US military formalized the term in the 1960s for adversarial modeling.
  • The paper's question bank asks evaluators to specify the artifact under evaluation, including model version, fine-tuning details, existing guardrails, lifecycle stage, and release conditions.

0. Pre-activity

Related work distinguishes red-teaming from penetration testing and emphasizes that effective evaluation depends on context, knowledge, assumptions, and complementary transparency and mitigation practices. Existing surveys and frameworks further show that AI red-teaming spans external review, UX participation, reliability testing, and broader generative-AI evaluation.

  • Red-teaming models an adversary and maps vulnerabilities from a threat lens, whereas penetration testing involves actively attempting to find system vulnerabilities.
  • Red-teaming is not an audit, and effective exercises require context, knowledge, and assumptions about how the system will be used.
  • Red-teaming is one transparency approach among factsheets, audits, model cards, media literacy, and watermarking rather than a standalone remedy.
  • Prior work includes interviews with LLM attackers, surveys supporting external red teams, proposals for UX participation, NLP reliability frameworks, and layered GenAI evaluation frameworks.

3 Case Studies: AI Red-teaming in practice

The six publicly documented AI red-teaming cases varied widely in goals, team composition, resources, disclosure, and mitigation. These cases also exposed structural tradeoffs: testing rarely covered the full risk surface, and limited reporting made effectiveness difficult to assess.

  • Case selection: The six cases were identified from public reports and news stories, primarily concerning private-sector exercises whose undisclosed methods limit coverage of practice.The selection was not intended to represent the full range of red-teaming activities.
  • Goals and threat models: Red-teaming goals and processes varied from single testing rounds to iterative testing that reprioritized risks for further investigation.Activities also differed in their threat models and areas of focus.
  • Teams and resources: Teams comprised subject-matter experts, crowdsourced participants, or language models, with corresponding differences in time, compute, model access, and task constraints.Crowdsourced efforts were typically time-boxed and API-limited, whereas expert-led efforts were more open-ended.
  • Disclosure and reporting: Only half of the cases publicly shared specific examples of risky behavior, and no standardized reporting procedures governed disclosure.One case released 38,961 red-team attacks, while other findings were limited or withheld under responsible-disclosure concerns.
  • Mitigation and supporting evaluation: Every case identified risky behavior, but none led to a decision not to release the model; mitigations varied and their effectiveness was often difficult to determine.Models had also undergone evaluations beyond red-teaming, but those methods lacked established reporting standards as well.
  • Discussion and limitations: Red-teaming remained ill-structured because open-ended instructions broaden exploration while specific instructions improve relevance, leaving the full risk surface unexamined.Team composition and disclosure introduced additional tradeoffs: experts may be selection-biased, crowdsourced evaluators resource-limited, and incomplete reporting reduces practical utility.

4 A Survey of AI Red-teaming Research

The surveyed AI red-teaming literature spans divergent risk definitions, evaluation methods, adversary assumptions, evaluator groups, and follow-up practices, with no consensus on what red-teaming entails. These variations expose unresolved questions about whose values guide evaluation and how findings should lead to mitigation.

  • Threat models: Dissentive and consentive risks provide contrasting threat models: the former permits contextual disagreement, while the latter concerns harms whose danger is broadly agreed.The survey also identifies work addressing both risk types or neither, where issue definitions must be developed from scratch.
  • Methodologies: Algorithmic search modifies prompts through random or guided perturbations, while targeted attacks deliberately exploit model components or training steps.Automated models may repeatedly attack systems until safeguards are bypassed, and perturbations can also support jailbreak detection.
  • Scope and criteria: The surveyed methods lack an agreed definition of red-teaming, despite being presented as evaluations under that label.The authors characterize these evaluations as necessary but potentially insufficient safety tests.
  • Threat models: Threat models disproportionately target dissentive risk, potentially producing exaggerated safety while diverting attention from consentive risk.The survey notes that some mitigated behaviors may be admissible in particular contexts.
  • Participants and values: Red-teaming practices disagree about adversary resources, evaluator identities, and the values used to judge acceptable outputs.Evaluators include crowdworkers, competition participants, researchers, and casual participants; multilingual and interdisciplinary participation are proposed responses to coverage gaps.
  • Outcomes and mitigation: Public follow-up is generally muted or mixed, with vulnerabilities and affected models often persisting after disclosure.One cited exception involved ChatGPT changes and revised terms of use, but those changes followed the authors’ notification by 90 days.

5 NIST RFI Comment Analysis Summary

Analysis of NIST RFI comments largely supports the survey’s finding that public AI red-teaming lacks clear definitions and rigorous structure. Commenters also distinguish model-level from system-level evaluation and raise broader concerns about GenAI’s development and assessment.

  • Similarities: Industry, academia, and civil society commenters asked NIST to clarify red-teaming and provide resources, guidelines, and best practices.Even firms experienced in GenAI red-teaming requested concrete guidance.
  • Differences: Some commenters recommended evaluating both models and systems, revealing that either level may be described as red-teaming.The authors use this variation to illustrate the need for more concrete evaluation definitions.
  • Scope: The authors note that their work primarily considers model-level evaluations, while RFI comments also express concerns about GenAI beyond evaluation.These comments broaden the policy discussion beyond the survey’s main evaluation focus.

6 Takeaways and Recommendations

The analysis finds that red-teaming is valuable but limited: it is inconsistently scoped, incompletely reported, and often followed by unclear mitigation. The authors therefore recommend treating it as one evaluation approach among others and offer a question bank to structure future exercises.

  • Red-teaming exercises covered limited vulnerability sets and cannot guarantee safety across all risks.Different approaches may identify harmful text responses or phishing vulnerabilities, while some issues, including algorithmic monoculture and environmental impacts, lie beyond red-teaming alone.
  • Red-teaming should be combined with other evaluation paradigms, while organizations should support participation from technical, user-facing, and legal roles.
  • Red-teaming is currently an unstructured procedure with undefined scope, motivating publicly available guidelines.
  • Reporting lacks unified protocols, and some studies omit findings or the resources required for evaluation.
  • Mitigation and alignment follow-ups are often vague, risking an assurance that red-teaming occurred without specifying problems discovered or fixed.Common solutions such as RLHF did not represent the full range of possible responses; monitoring, prediction modification, and refusing deployment were rarely mentioned.
  • The proposed question bank is a starting point for considering red-teaming's benefits, limitations, and design choices before, during, and after evaluation.The authors explicitly present it as an initial contribution rather than finalized guidelines.

A Research Survey and Case Study Details

The appendix materials document how the surveyed literature was classified and provide supporting research artifacts. They include detailed classifications by content type and evaluation approach, alongside additional project notes and analyses.

  • The appendix provides further details on the research papers and case studies examined in the work.It also includes access to a Google Sheets project containing notes and thematic analyses.
  • The survey classification includes a dimension labeled type of risk investigated.
  • Table 4 classifies surveyed papers by the type of content produced and the approach used for evaluation.

B Extended NIST RFI Comment Analysis

The extended NIST RFI analysis examines who commented, which issues they addressed, and how they framed red-teaming. Comments call for clearer definitions, model-and-system coverage, shared resources, and guidance that balances cybersecurity practices with AI-specific constraints.

  • The RFI analysis examined submitted comments to understand respondents, issue areas, and argument construction.The RFI solicited feedback on AI red-teaming, content watermarking, and AI development standards.
  • Comments came from individuals, civil society, industry, academic groups, and government agencies, with recurring attention to watermarking, copyright infringement, and red-teaming.Individual comments often focused on data rights and transparency, while industry submissions more often discussed products, regulation, or internal evaluation approaches.
  • Many respondents described red-teaming as vaguely defined and called for clear, standardized, or globally aligned definitions.Google characterized the term as a catch-all, while Hugging Face distinguished natural-language red-teaming prompts from adversarial machine-learning prompts.
  • Respondents emphasized that future guidance should distinguish red-teaming AI models from red-teaming complete AI systems.System red-teaming includes the model alongside data infrastructure, user interfaces, and other components.
  • Comments requested that NIST disseminate case studies, illustrations, and examples of red-teaming practices.Meta specifically proposed collecting materials that exemplify best practices or the state of the art.
  • Industry and civil society diverged over external red-teaming, with companies stressing feasibility and selective disclosure while civil society urged external participation throughout the lifecycle.Civil society proposals included prompt-engineering or white-hat-jailbreaking expertise, domain experts, and red-teamers representative of expected users; cybersecurity-informed comments also called for remediation time before reporting.
Loading 2401.15897v3…