Source-linked AI summary
Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data
Dmitry Nikolaev, Ashley A. Mattheis
TL;DR
The study addresses limited evidence about extremist speech in LLM training corpora, develops a conservative semi-automated detection pipeline, and estimates that such content occurs in Dolma at roughly 1/2000 documents.
Problem
Limited research has examined the composition of LLM training data, including whether it contains unfiltered, uncontextualised extremist speech.
Method
The paper combines large-scale automated classification and filtering with manual expert analysis to conservatively identify extremist texts in Dolma.
Results
Approximately 100 of 200,000 sampled Dolma documents contained extremist narratives or hate speech, yielding a conservative prevalence estimate of 1/2000.
Takeaways & Limitations
Dolma contains extremist content of several types, while borderline material and contextual dislocation complicate decisions about training-data inclusion.
Takeaways & Limitations
False negatives from documents rejected at the pipeline’s first stage mean the study provides only a lower-bound prevalence estimate.
Abstract
from arXiv · showhide
Despite a strong interest on the part of the research community in the topic of trustworthy and safe AI, the composition of the text corpora that large language models (LLMs) encounter in pre- and post-training has not yet drawn much attention. In this work, we address the question of whether LLMs are exposed to unfiltered, uncontextualised extremist speech. Using several definitions of extremist speech, stemming from official documents and research literature, and an extraction pipeline combining automated text processing with expert verification, we provide a lower bound on the prevalence of extremist documents in Dolma, an open training corpus underpinning the OLMo series of models. We show that Dolma is likely to include hundreds of thousands of documents containing extremist content and hate speech of several types, including direct calls for violence, and discuss the implications of this for data curation and model pre-training.
1 Introduction
The study addresses the underexamined curation of LLM training data by estimating the prevalence of unfiltered extremist speech in Dolma. It develops a conservative, expert-verified pipeline that yields a reliable lower bound while examining how estimates vary across three definitions of extremism.
- Motivation: Trustworthy-AI research emphasizes architectures and post-training, while data curation and filtering receive comparatively less attention.
- Motivation: Even small amounts of extremist training content may pose risks because LLMs can memorize texts verbatim, while automated identification remains difficult under data scarcity.
- Contribution: The study estimates extremist-content prevalence in Dolma by combining large-scale automated classification and filtering with manual expert analysis.
- Method and results: The conservative identification pipeline provides a reliable lower bound on prevalence, with outcomes varying slightly across the three tested extremism definitions.
2 Extremism definitions
The paper uses multiple definitions of extremism, centered on ideological opposition to rights, democracy, or pluralism and on hostile action against out-groups. These definitions vary in emphasis, from democratic-order threats to in-group survival and concrete extremist traits.
- UK Government definition: The UK-adapted definition treats extremism as promoting an ideology based on violence, hatred, or intolerance that targets rights, democracy, or a permissive environment for such outcomes.It focuses on opposition to the democratic order and prevailing rights and freedoms.
- Berger definition: The Berger-adapted definition links extremism to the belief that in-group success or survival requires hostile action against an out-group.Hostility may include verbal attacks, diminishment, discrimination, violence, or genocide.
- Schmid definition: The Schmid-adapted definition describes extremism as political expression unacceptable to mainstream politics and characterized by anti-democratic, intolerant, authoritarian, and uncompromising traits.It also identifies preferences for violence over persuasion, uniformity over diversity, unity over pluralism, and orders over dialogue.
3 Data and methods
The study samples 200,000 documents from Dolma’s approximately 3-trillion-token corpus and applies a multi-stage pipeline combining language-model extraction, filtering, and expert validation. Because false negatives cannot be reliably screened, prevalence estimates are treated as lower bounds.
- Sampling: 200,000 documents were randomly sampled from Dolma’s approximately 3 trillion English-language tokens using reservoir sampling over one pass.The sample was used because holistic analysis of the full corpus was infeasible.
- Automated extraction: Several mid-size language models classified sampled documents as extremist speech or otherwise under hardware-constrained parameter and quantisation settings.The comparison included 4-bit Llama 3.3 70B and Qwen 2.5 72B alongside smaller models.
- Union check: A union prompt checked whether models identified texts satisfying any extremist-speech definition and whether results overlapped reasonably with single-definition extractions.The authors illustrate this with a target of at least 7 overlapping texts when the UK definition yields 10 and the union yields 100.
- Automated filtering: False negatives could not be screened convincingly, so the estimated prevalence is explicitly treated as a lower bound on extremist material in training data.Consistency checks and rechecking extracted documents address false positives but cannot establish how many extremist documents were discarded.
- Expert evaluation: Ten texts per definition were sampled from Gemma 4 extractions validated by GPT 4o mini and assessed by a domain expert in extremist speech.The authors note that the small sample was difficult to analyse because texts were long, partly incoherent, and used hateful and coded language.
4 Prevalence estimates
Prevalence estimates vary substantially by extraction model and extremist-speech definition. Qwen 3.5 27B offers the best balance of consistency and precision, while Gemma 4 is more conservative and Llama 3 70B is unreliable.
- Model comparison: Llama 3 70B is unreliable: it extracted more Berger-definition texts than union texts, produced poor union checks, and had most results disqualified by GPT 4o mini.Its apparent recall is undermined by an extremely high false-positive rate.
- Model comparison: Qwen 2.5 returns considerably more false positives than Qwen 3.5, whereas Qwen 3.5 is more conservative and better aligned with GPT 4o mini estimates.Despite model sizes of 72B vs. 27B parameters, the Qwen models show similar overall patterns; Qwen 2.5 performs well on union checks.
- Model comparison: 20% fewer texts returned on average across definitions compared to Qwen 3.5, Gemma 4 trades recall for the lowest GPT 4o mini rejection percentage.Gemma’s union discrepancy averages 11%, compared with 3% for Qwen 3.5.
- Model comparison: Qwen 3.5 27B provides the best balance between internal consistency and precision, provided larger-model or human screening follows extraction.Its recall is not noticeably lower on average than Llama 3 70B.
- Definitions: Berger’s definition produces high-variance estimates, making it more suitable for exploratory work than practical screening applications.Schmid’s definition returns more texts than the official UK definition, but filtering by GPT 4o mini narrows the difference.
5 In-depth analysis
Manual close reading found that definitions strongly shaped which texts the models classified, while at most about half of returned texts were extremist. The analysis also revealed substantial hate speech, borderline content, contextual ambiguity, and an almost complete absence of Jihadist extremist material.
- Cross-model findings: None of the models returned Jihadist extremist narratives, and only one news story about Jihadist terrorists appeared across the full sample.The analysis also identified many borderline texts whose extremist significance may be obscured when corpus extraction removes original context.
- Overall sample: 10 (33.34%) of 30 reviewed texts were extremist, including nine right-wing narratives, two explicitly calling for violence, and one left-wing text.A further 6 (20%) texts were hate speech without common extremist narratives.
- UK Government definition: The UK Government definition produced a relatively good extremist hit rate, but half the sampled texts were non-extremist; four general right-wing texts and one violent text dominated its extremist content.Three of the five non-extremist texts nevertheless involved potential harms or indeterminacy.
- Berger definition: The Berger model returned the broadest coding range but relatively few extremist texts, emphasizing hate speech and including both right-wing and left-wing extremism.Its right-wing extremist text called for violence, while its left-wing text presented a detailed radicalization narrative; it also identified mainstreaming and context-dependent material.
- Schmid definition: The Schmid model was most narrowly focused on extremist and hate-speech content, but three international or regional political texts were indeterminate without context.Those cases could indicate stronger detection of explicit extremism if contextualized, although they appeared related to political polarization as coded.
- Overall sample: At best, only about half of model-returned texts were extremist under close reading, and each model also returned non-extremist texts.The classifications varied according to the definition used, with problematic, polarized, violent, and indeterminate material among the non-extremist outputs.
6 Source-domain check
The source-domain check found that metadata statistics offer only limited value for pre-screening harmful documents, although KiwiFarms’ five documents suggest some domain filtering may be warranted.
- 6 Source-domain check: Domain metadata had limited value for pre-screening harmful documents because most evidently extremist resources contributed only one document each.The analysis extracted domains for XR/LW(-V) and HS documents and counted domain representation in the 200k sample.
- 6 Source-domain check: 5 documents from KiwiFarms in the analysed sample suggest that at least some domain filtering may be warranted.Other documents came from comments on different message boards and informational resources.
7 Conclusion
The study proposes and validates a conservative semi-automated pipeline for identifying extremist texts in LLM training data and uses its findings to motivate improved data curation, stakeholder discussion, and technical responses. It also identifies reverse-engineering models’ operational definitions of extremism as a future research direction.
- Contribution: The study’s primary contribution is a semi-automated pipeline for identifying extremist texts in LLM training data and an overview intended to spur further pre-training-data analysis and curation.The authors characterize the extraction procedure as rather conservative.
- Implications: Potential material harms to individuals and society make addressing extremist and hate-speech content in training data urgent, beyond ensuring compliant AI-system functioning.The passage also refers to reported instances of self-harm and public violence by AI chatbot users.
- Implications: The findings create urgency around whether violent, politically polarised, extremist, and hate-speech texts are appropriate training data for LLM-based interactive technologies.The authors present the work as a foundation for multistakeholder discussions while emphasizing continued technical responses.
- Future work: Future work should reverse-engineer the operational definitions of extremism that models apply in practice.Qwen 2.5 appears to produce many false positives while applying the union definition robustly, revealing opaque criteria that do not perfectly align with the researchers’ definitions.
Limitations
The study likely underestimates extremist-document prevalence because its pipeline can produce false negatives, and its low observed prevalence creates high sampling variance. The authors nevertheless find this variance limitation non-crucial at this stage because the sample is i.i.d. and results remain significant under tenfold overcounting.
- Limitations: False negatives mean the method provides only a lower-bound estimate of extremist-document prevalence in the training data.Documents rejected by the first analytic stage were difficult to recheck.
- Limitations: 1/2000 extremist-document prevalence makes the estimate vulnerable to high sampling variance.The initial document sample is i.i.d.
- Limitations: Tenfold overcounting would still leave the results significant, so sampling variance is not considered crucial at this analysis stage.This assessment depends on the initial sample being i.i.d.
Appendix A Prompts
Appendix A provides the prompts used to classify documents for extremist speech. The prompts instruct an expert political-language evaluator to apply multiple definitions using only document content and return a binary Yes/No judgment.
- By-definition prompts: Definition 1 characterizes extremism as promoting or advancing violence-, hatred-, or intolerance-based ideology aimed at undermining rights, democracy, or conditions enabling those outcomes.The definition lists three aims: negating rights, replacing liberal democracy, or intentionally creating a permissive environment for others to do so.
- By-definition prompts: The by-definition prompt asks an expert in political language to check whether a document satisfies a specified definition of extremist speech.The evaluator is framed as an expert and the task is embedded in Python code.
- By-definition prompts: The by-definition evaluator must use only the document content, ignore its source and author, and answer with one word: Yes or No.This constrains classification to the text itself and produces a binary output.
- Union prompt: The union prompt asks an expert in political language to check whether a document satisfies any of the listed extremist-speech definitions.It combines the definitions into a single screening question.
- Union prompt: The union prompt includes Definition 1, which links extremist ideology to violence, hatred, or intolerance and aims involving rights, liberal democracy, or permissive conditions for such aims.This reproduces the rights- and democracy-focused definition used in the by-definition prompt.
- Union prompt: Definition 2 defines extremism as treating an in-group’s success or survival as inseparable from hostile action against an out-group, ranging from verbal attacks to violence.The listed hostile acts include verbal attacks, diminishment, discriminatory behavior, and violence.
- Union prompt: Definition 3 describes extremism in democratic societies as political expression outside moderate mainstream acceptance, often on the far left or far right, with groups tending to be anti-constitutional or anti-democratic.The passage also characterizes extremist groups and parties as anti-pluralistic and fanatical.
- Union prompt: The union evaluator likewise uses only document content, ignores source and author, and returns one word: Yes or No.The same binary, document-only decision format is applied to the union of definitions.
Appendix B Source-domain statistics for extremist/hate-speech documents
Table 4 reports source-domain prevalence among documents identified as containing extremist narratives or hate speech in the analysed 200k-document sample.
- Source-domain prevalence: Table 4 reports the prevalence of source domains in the analysed 200k-document sample.The sample concerns documents examined for extremist narratives or hate speech.
- Source-domain prevalence: The analysed sample contains documents identified through close reading as containing extremist narratives or hate speech.Close reading was used to identify the relevant documents.
- Source-domain prevalence: The reported domains are those that provided documents identified as containing extremist narratives or hate speech.The table links identified documents to their providing domains.
Appendix C Manually analysed documents · Appendix C.1 The UK definition
The manually analysed Dolma documents contain varied extremist and hateful material, including threats, calls for violence, dehumanising abuse, racial and antisemitic conspiracy claims, and anti-government or revolutionary rhetoric. The examples span both explicit advocacy and material framed as commentary, propaganda, or political grievance.
- Appendix C.1 The UK definition: The manually analysed material is heterogeneous and sometimes embedded within longer, mixed-topic pages rather than appearing as isolated extremist documents.Several passages shift between political commentary, historical discussion, news items, personal anecdotes, and abusive or threatening statements.
- Appendix C.1 The UK definition: The sample includes direct threats and violent retaliation against perceived wrongdoers, including calls to expose police officers and wishes for sexualised assault against a convicted child molester.One passage urges posting officers’ contact information and addresses, while another endorses placing a convicted offender with violent gang members.
- Appendix C.1 The UK definition: The corpus also contains extremist propaganda about terrorism, including warnings about Islamic attacks and claims that a terrorist event was manipulated to rally politically awakened Americans.The examples present terrorism through both promotional fear messaging and conspiratorial political interpretation.
- Appendix C.1 The UK definition: Religious extremist speech includes demands that a Christian convert be tried or executed, alongside an explicit threat to resort to violence against the government.The passage reports clerics and followers demanding the convert’s return and death, then threatening violence if the government did not comply.
- Appendix C.1 The UK definition: Several documents promote racial or ethnic hatred through slurs, assertions of White racial superiority, and hostility toward Black, Jewish, and Somali populations.The material portrays demographic change as an attack on White civilization and uses racialised abuse against multiple groups.
- Appendix C.1 The UK definition: Political extremist rhetoric threatens mass upheaval, predicts that plutocrats will not survive an uprising, and frames revolution or world war as consequences of ideological conflict.The passages describe a minor spark triggering violence and accuse political opponents of destabilising society.
- Appendix C.1 The UK definition: Other examples combine historical grievance, nationalist claims, and support for armed mobilisation, including discussion of colonial violence, separatism, and civilians carrying AR-15 rifles.These passages connect political arguments about Sri Lanka and Britain with accounts of violence and imagery of an armed irregular force.
- Appendix C.1 The UK definition: Antisemitic conspiracy narratives portray Jews as controlling media, politics, immigration, and public opinion, while urging White people to organise for civil war.The examples combine conspiracy claims about Jewish influence with explicit collective mobilisation against Jews and perceived racial outsiders.
Appendix C.2 Berger’s definition
The passages illustrate Berger’s definition through racist dehumanization, anti-Muslim and anti-immigrant threat framing, advocacy of preemptive violence, and extremist propaganda celebrating martyrdom. They also include scapegoating and broad political or conflict-related accusations.
- A racist passage claims that Black people have no place in modern society and portrays their presence as causing urban decline.
- Other excerpts attribute political or military failures to liberals, critics, governments, or opposing factions and contain bitter accusations about armed conflicts.The material includes claims that media and domestic critics are being blamed for wartime losses, alongside accusations exchanged by opposing sides in Sudan.
- Several passages frame Muslims, Islamists, immigrants, or ethnic groups as collective security threats requiring profiling, internment, exclusion, or preemptive violence.One passage calls killing Islamist infiltrators overseas and apprehending them domestically necessary, while another defends threat profiling and internment analogies.
- An extremist nasheed calls on “soldiers of Allah” to crush enemies and celebrates martyrdom, prompting an investigation for promoting terrorism.
Appendix C.3 Schmid’s definition
The examples under Schmid’s definition span white-nationalist and antisemitic rhetoric, threats and advocacy of violence, racial hostility, and political or religious extremism. They include both direct extremist statements and reported or contested extremist content.
- White-nationalist and antisemitic rhetoric: White-nationalist passages frame Jews and Christianity through racialized hostility and present radical views as ideas that non-nationalists might gradually accept.The passages describe Jews as a central adversarial concern, praise white-nationalist framing, and discuss radicalization through preliminary exposure.
- Political and religious extremism: Other passages portray civil conflict and revolution as impending remedies, condemn religious and political institutions, and describe attacks as vengeance against an alleged infidel.These examples combine treason and civil-war rhetoric, denunciations of the Vatican, and justification of an attack on Governor Samuel Ortom.
- Threats and violence: Violence-related examples include a planned Columbine-style massacre, stated willingness to die for a cause, and threats to rip a political opponent’s tongue out.The passages describe terrorism charges involving guns and explosives and a public threat against anyone speaking against Nawaz Sharif.
- Racial hostility: A separate passage expresses hostility toward Chinese migrants, blaming a small minority for disorder and economic harm in other countries.The passage generalizes about Chinese people moving abroad and characterizes their conduct as uncivil.
- Contested extremist framing: One discussion rejects the assumption that people are supposed to be hated and distinguishes a nuanced position from the BNP’s perspective.The passage explicitly challenges a generalized imperative to hate while acknowledging oppression and violent tendencies among some individuals.