Source-linked AI summary
The Effect of Extremist Violence on Hateful Speech Online
Alexandra Olteanu, Carlos Castillo, Jeremy Boy, Kush R. Varshney
TL;DR
The paper asks how extremist attacks affect the volume and type of online hate speech. It analyzes Twitter and Reddit using categorized hate-speech collections and counterfactual time-series methods. Extremist violence tends to increase hate speech, particularly messages promoting violence, while some events also increase counter-hate speech.
Problem
The paper investigates how external extremist events affect the prevalence and types of online hate and counter-hate speech targeting Muslims and Islam.
Method
The study analyzes 150M+ messages from Twitter and Reddit across 13 attacks, categorizing speech and comparing observed post-event time series with counterfactual series.
Results
Extremist violence increased hate speech targeting Muslims and Arabs, especially high-severity speech, while Islamist attacks also increased counter-hate speech.
Takeaways & Limitations
The findings provide a blueprint for monitoring hate speech around extremist attacks and suggest boosting the visibility of counter-hate speech.
Takeaways & Limitations
The study is limited to English-language data and 13 events in Western countries, with possible collection and annotation biases.
Abstract
from arXiv · showhide
User-generated content online is shaped by many factors, including endogenous elements such as platform affordances and norms, as well as exogenous elements, in particular significant events. These impact what users say, how they say it, and when they say it. In this paper, we focus on quantifying the impact of violent events on various types of hate speech, from offensive and derogatory to intimidation and explicit calls for violence. We anchor this study in a series of attacks involving Arabs and Muslims as perpetrators or victims, occurring in Western countries, that have been covered extensively by news media. These attacks have fueled intense policy debates around immigration in various fora, including online media, which have been marred by racist prejudice and hateful speech. The focus of our research is to model the effect of the attacks on the volume and type of hateful speech on two social media platforms, Twitter and Reddit. Among other findings, we observe that extremist violence tends to lead to an increase in online hate speech, particularly on messages directly advocating violence. Our research has implications for the way in which hate speech online is monitored and suggests ways in which it could be fought.
1 Introduction
The paper examines how extremist attacks affect hate and counter-hate speech on Twitter and Reddit. It builds a large, multidimensional dataset and uses causal time-series analysis to compare post-event speech with counterfactual behavior.
- The study asks how external events affect the prevalence and type of hate and counter-hate speech targeting Muslims and Islam.
- Extremist violence tends to increase messages directly advocating violence, supporting concern about feedback between offline violence and online hate speech.
- The analysis covers 19 months of Twitter and Reddit messages collected through an expanded lexicon of hate-related terms.
- The framework categorizes messages by target, speaker stance, hateful severity, and content framing before analyzing selected extremist attacks.
- The study selects 13 extremist attacks involving Arabs and Muslims as perpetrators or victims in Western countries.
- Causal impact is estimated by comparing post-event observed time series with counterfactual predictions of behavior absent the event.
2 Background & Prior Work
Prior work frames online hate speech as harmful to individuals and communities while documenting challenges in defining, detecting, and moderating it. The paper also treats counter-speech as an important complement to suppression-focused approaches.
- Definitions of hate speech vary across jurisdictions, so the study combines legal and scholarly definitions covering intolerance, animosity, and disparagement based on group characteristics.
- Counter-speech is studied alongside hate speech because international guidance generally prefers it to suppression where free-speech protections create tension.
- Research on Muslims identifies visible Muslim identity as especially vulnerable to hostility, including online and offline intimidation, abuse, and threats of violence.
- Social platforms have expanded avenues for harassment, motivating efforts to detect, understand, quantify, filter, and moderate harmful speech.
- Existing online-hate research commonly categorizes messages by target group, basis for hate, and severity, but distinguishing hate from offensive speech remains difficult.
3 Data Collection
The paper constructs high-recall hate-speech collections from Twitter and Reddit using iterative keyword expansion and external lexicons. It supplements social-media data with news time series for event analysis.
- The collection targets hate and counter-hate speech about Muslims and Islam after Islamist terrorist and Islamophobic attacks in Western countries.
- Reddit posts are collected through the Search API, with comments retained for matching posts; GDELT news data provide exogenous variables for time-series modeling.
- The researchers bootstrap keyword queries, iteratively add terms discovered in retrieved messages, and favor recall in annotation and expansion.
- The initial term process produced 91 terms, including expressions such as “ban islam” and “kill all muslims.”
- Twitter data combine the Decahose 10% public-stream sample with expanded and HateBase-derived queries, yielding over 1TB of raw data.
4 Characterizing Hate Speech
The paper characterizes post-event messages along stance, target, severity, and framing. These dimensions distinguish supportive and hostile positions, targeted groups, violent or nonviolent severity, and the ways messages diagnose problems or propose solutions.
- Exploration of Post-Event Messages: The study examines messages after the Manchester Arena bombing, Portland train attack, and Quebec City mosque shooting.
- Framing: After the Manchester attack, recurring frames blamed Muslim exclusion or immigration for terrorism, while other messages offered counter-narratives.
- Stance: Stance categories include favorable support, unfavorable attack, commentary on negative actions or speech, and neutral, factual, or unclear content.
- Target: Targets include Muslims and Islam, religious groups, Arabs or Middle-Easterners, ethnic or foreign-descent groups, immigrants or refugees, and other vulnerable groups.
- Severity: Severity distinguishes messages that promote violence, intimidate, or offend or discriminate against a target.
- Framing: Framing identifies messages that diagnose causes, suggest solutions, perform both functions, or perform neither.
5 Methodological Framework
The framework treats extremist violence as interventions in daily social-media time series and estimates effects by comparing observed post-event behavior with synthetic counterfactual controls. It uses relative lift or drop over seven post-event days, with transformations and a one-week window addressing low-volume series and short-lived effects.
- Extremist attacks are modeled as interventions, with causal effects estimated from differences between treated and synthetic control time series.The control represents behavior had the event not occurred and is modeled from correlated series not affected by the event.
- The control series combines 11 pre-event weeks, analogous periods one and 23 weeks earlier, and external Twitter, Reddit, and news signals.The state-space model predicts the counterfactual from these observed sources using maximum likelihood estimation.
- Relative effect compares summed treated-control differences with summed control values across the seven days after each event.Terms are ranked for each event by this relative lift or drop measure.
- A large constant C is added before control synthesis to prevent negative synthetic controls and skewed estimates in low-volume series.The transformation is intended to preserve the impact's shape and amplitude; C is arbitrarily set to 1000.
- The one-week post-event observation window is intended to capture the bulk of effects, while the dataset spans 19 months.The cited prior evidence found related messages after a UK terror attack propagated for less than two days.
6 Experimental Results
The study estimates how 13 extremist attacks affected hate and counter-hate speech across Twitter and Reddit. Islamist terrorist attacks were associated with increased anti-Muslim hate, especially severe speech, while Islamophobic attacks produced less consistent responses.
- Event selection: The analysis selects 13 extremist attacks involving Arabs and Muslims as perpetrators or victims during a 19-month observation period.Events were selected from Western countries and covered by international news media.
- Data and annotation: The researchers annotate speech by target, stance, severity, and framing, using crowdsourced labels for 1890 unique terms and message-level framing annotations.Terms receive at least three annotations, with additional annotations when consensus is not reached.
- Impact estimation: Effects are identified when the 90% confidence interval for the observed-counterfactual difference excludes zero.The counterfactual forecast becomes less certain farther into the post-event period.
- Hate speech effects: +3.0 on Twitter and +2.9 on Reddit are the mean relative effects for hate speech targeting Muslims after Islamist terrorist attacks.High-severity speech targeting Muslims and Arabs increased more strongly: +10.1 on Twitter and +6.2 on Reddit, with one exception for the 2016 Istanbul Airport attack.
- Event differences: Islamophobic attacks did not produce a consistent cross-platform pattern: high-severity terms increased after Finsbury Park but decreased after the Olathe Kansas shooting.The aggregate distribution in Figure 4 supports this lack of consistency.
- Counter-hate and framing: +1.8 on Twitter and +2.9 on Reddit are the mean relative effects for counter-hate speech after Islamist terrorist attacks, but no corresponding increase appears after Islamophobic attacks.Speech framed as solutions was more prevalent among highly impacted terms on Twitter than Reddit, especially calls to ban or deport immigrants, Muslims, or Arabs.
7 Conclusions
The paper argues that measuring hate speech after extremist violence requires causal inference alongside automated collection and manual annotation. Its evidence can inform monitoring, but interpretation is constrained by collection bias, retrospective deletion, and limited geographic and linguistic scope.
- Conclusions: Causal inference, automated processes, and manual annotations are combined to measure hate speech at scale while managing a finite annotation budget.The methodology is designed to balance large-scale analysis with the limits of human annotation.
- Conclusions: Observed hate-speech time series are compared with counterfactual series across two platforms and two classes of extremist events.The counterfactual comparison estimates how speech might have evolved had the events not occurred.
- Conclusions: The findings provide a blueprint for monitoring online hate speech, especially the relationship between online calls for violence and deadly extremist attacks.The paper also reports increases in counter-hate speech and suggests platforms could boost its visibility.
- Conclusions: The analysis is limited by potential bias from the bootstrap term list and ambiguous query-level annotations.These collection choices can affect which messages enter the dataset.
- Conclusions: The study covers English-language data from only 13 events in Western countries, limiting direct support for other regions, languages, and event types.The authors also describe the framing analysis as an initial effort requiring deeper analysis.
- Conclusions: Because harmful content may be deleted before retrospective analysis, increases in some speech types are more reliable than observed decreases.Platform moderation can produce incomplete data collections.