Source-linked AI summary

A Quantitative Approach to Understanding Online Antisemitism

Savvas Zannettou, Joel Finkelstein, Barry Bradlyn, Jeremy Blackburn

arXiv:1809.01644v2cs.CYcs.SI

TL;DR

Online antisemitism is difficult to measure at Web scale, despite its importance and the limitations of traditional qualitative approaches. The paper develops a large-scale quantitative framework using fringe-community data and finds increasing antisemitic language, discoverable rhetorical patterns, and cross-community meme propagation. It also identifies scope limits from conservative measurement, selected keywords, and focus on two evolving communities.

  • Problem

    The Web's scale and speed limit traditional qualitative approaches to understanding the growth and spread of online antisemitism, despite its societal importance.

  • Method

    The paper analyzes over 100M posts from /pol/ and Gab using quantitative temporal analysis, word2vec embeddings, meme analysis, and influence modeling.

  • Results

    The study finds increasing antisemitism and racially charged language, discovers distinct antisemitic-language facets, and identifies /pol/ and The Donald as the most influential and efficient communities, respectively, for spreading Happy Merchant.

  • Takeaways & Limitations

    The results provide a data-driven quantitative framework for understanding online antisemitism and augmenting qualitative anti-hate efforts.

  • Takeaways & Limitations

    Measurements are generally lower bounds, keyword growth is focused on two keywords, and treating rapidly evolving Gab as stable may underestimate its influence.

Abstract

from arXiv · show

A new wave of growing antisemitism, driven by fringe Web communities, is an increasingly worrying presence in the socio-political realm. The ubiquitous and global nature of the Web has provided tools used by these groups to spread their ideology to the rest of the Internet. Although the study of antisemitism and hate is not new, the scale and rate of change of online data has impacted the efficacy of traditional approaches to measure and understand these troubling trends. In this paper, we present a large-scale, quantitative study of online antisemitism. We collect hundreds of million posts and images from alt-right Web communities like 4chan's Politically Incorrect board (/pol/) and Gab. Using scientifically grounded methods, we quantify the escalation and spread of antisemitic memes and rhetoric across the Web. We find the frequency of antisemitic content greatly increases (in some cases more than doubling) after major political events such as the 2016 US Presidential Election and the "Unite the Right" rally in Charlottesville. We extract semantic embeddings from our corpus of posts and demonstrate how automated techniques can discover and categorize the use of antisemitic terminology. We additionally examine the prevalence and spread of the antisemitic "Happy Merchant" meme, and in particular how these fringe communities influence its propagation to more mainstream communities like Twitter and Reddit. Taken together, our results provide a data-driven, quantitative framework for understanding online antisemitism. Our methods serve as a framework to augment current qualitative efforts by anti-hate groups, providing new insights into the growth and spread of hate online.

1 Introduction

Online antisemitism is increasingly consequential because fringe Web communities can spread racist ideology at the scale and speed of social media. The paper addresses limited large-scale evidence by proposing a transparent quantitative framework and asking how antisemitism rises, evolves, and spreads online.

  • Motivation: Online communities enable alt-right groups to congregate, organize, and disseminate weaponized information through humor and memes.The Web's scale and speed support the spread of racist ideology beyond fringe communities.
  • Motivation: Recent increases in hate crimes, fascist and white-power groups, and white-nationalist propaganda parallel the rise of online racist rhetoric.The paper presents this correspondence as motivating closer study of online hate and real-world events.
  • Research gap: Antisemitism is central to alt-right ideology and historically associated with authoritarian ideologies, yet large-scale evidence about online antisemitism remains limited.Existing organizations have generated valuable insights, but primarily qualitative methods are constrained by Web scale.
  • Approach: The paper presents an open, scientifically rigorous, transparent, and generalizable framework for quantitative analysis of online antisemitism.The framework analyzes over 100M posts from 4chan's /pol/ and Gab using word2vec models to discover new antisemitic terms.
  • Research questions: The study asks whether online antisemitism has risen, how emerging antisemitic language can be categorized automatically, and how fringe communities influence the wider Web.These questions organize the paper's quantitative investigation of online antisemitism.

2 Related Work

Prior work studies hate speech prevalence, detection, online behavior, and antisemitism through diverse computational and qualitative methods. This paper differs by quantitatively examining the dissemination of antisemitic content across fringe Web communities, including racial slurs and memes.

  • Hate speech on Web communities: Prior studies measure hate speech across communities such as /pol/, Gab, Twitter, Whisper, Reddit, Voat, and Web forums.These studies examine prevalence, targets, community differences, and cross-community transfer.
  • Hate speech detection: Hate-speech detection research uses decision lists, SVMs, Naive Bayes, lexicons, crowdsourcing, deep learning, and semi-supervised methods.The reviewed work addresses explicit and implicit hate, multiple target categories, context, and annotation costs.
  • Hate speech detection: Some detection approaches improve performance by modeling new or misspelled words, community differences, or multiple forms of hateful content.The cited studies include language-model error features, community-driven transfer, blended labels, and deep-learning comparisons.
  • Antisemitism: Antisemitism research has used surveys, focus groups, and studies of beliefs, predictors, perceived consequences, and effects on Jewish communities.Examples address motives, support-seeking, authoritarianism, conspiracy beliefs, and education about discrimination.
  • Novelty: In contrast, this study performs a large-scale quantitative analysis of antisemitic dissemination in /pol/ and Gab, focusing on racial slurs and the Happy Merchant meme.Its emphasis is propagation rather than only detection, perception, or isolated case studies.

3 Datasets

The study constructs two large-scale datasets from 4chan's Politically Incorrect board (/pol/) and Gab to measure online antisemitism. /pol/ is an influential, high-hate anonymous imageboard, while Gab is a lightly moderated network welcoming banned users.

  • Dataset construction: The researchers collect two large-scale datasets from /pol/ and Gab to study antisemitism on the Web.The datasets contain posts and images from both communities.
  • /pol/: /pol/ is an anonymous 4chan imageboard where users create image-based threads and replies, and it exhibits high racism and hate speech.The paper selects /pol/ because it is also described as influential on the Web.
  • Gab: Gab is a social network founded in August 2016 that welcomes banned users and uses mild moderation, excluding illegal pornography, terrorist promotion, and doxing.The paper obtains Gab posts and images using methodology from prior work.
  • Ethical considerations: The analysis uses publicly available data collected from /pol/ and Gab.The ethical-considerations passage identifies the data source as public postings.

4 Results

The results quantify rising ethnic and antisemitic language in /pol/ and Gab, relate temporal changes to real-world events, and use semantic and image-based analyses to characterize terminology and meme propagation.

  • Trends in ethnic terms: Ethnic terms generally increased even after normalization, with “kike” rising most on both communities, followed by “jew” on /pol/ and “nigger” on Gab.By the dataset endpoints, “jew” appeared in 4.0% of /pol/ daily posts and 3.1% of Gab posts.
  • Event-related changes: Changepoint analysis identified temporal changes in “jew” and “white” usage on /pol/ near several political and Middle Eastern events.The PELT algorithm fit timeseries models with changing means and variances while penalizing additional changepoints; real-world events were then inspected manually.
  • Semantic analysis: Word2vec associations linked “jew” to antisemitic symbols, slurs, and derogatory terminology across both communities, while graph clusters separated moral-corruption, geopolitical-conspiracy, identity, lore, and theological themes.The triple-parentheses symbol was most similar to “jew” on /pol/ (cos θ = 0.80) and seventh on Gab (cos θ = 0.69).
  • Semantic analysis: The same semantic procedure found ethnic-nationalist, racial-purity, race-mixing, identity, and political-correctness themes around “white.”“huwhite” was highly similar to “white” on /pol/ (cos θ = 0.78) and Gab (cos θ = 0.70).
  • Meme propagation: Image tracking used perceptual hashing and clustering to compare antisemitic meme propagation across /pol/, Gab, Twitter, and Reddit, with Happy Merchant counts treated as lower bounds.The pipeline labels only clusters unambiguously identified as Happy Merchant, making the reported counts conservative.

5 Discussion

The paper argues that the Web’s scale and speed require large-scale quantitative approaches to understand rapidly proliferating online antisemitism. It reports evidence of increasing antisemitic language and meme propagation while identifying methodological and scope limitations.

  • The Web has enabled antisemitism and hate to grow and proliferate rapidly, motivating new techniques for understanding and combating this behavior.
  • The study analyzes over 100M posts from /pol/ and Gab to quantify antisemitic language, its contexts, and the propagation of the Happy Merchant meme.
  • The authors find increasing antisemitism and racially charged language that correlates in large part with real-world political events, while word2vec reveals distinct linguistic facets including slurs and conspiracy theories.
  • The Happy Merchant analysis identifies /pol/ and Reddit’s The Donald as the most influential and efficient communities, respectively, for spreading the meme across the Web.
  • The results should be treated as lower bounds because the meme pipeline is conservative, language analysis centers on two keywords, and the study focuses primarily on two fringe communities.
  • The authors recommend combining qualitative expertise with open, data-driven analysis and greater scientific effort to measure and combat online antisemitism.
Loading 1809.01644v2…