Source-linked AI summary

Spread of hate speech in online social media

Binny Mathew, Ritam Dutt, Pawan Goyal, Animesh Mukherjee

arXiv:1812.01693v1cs.SI

TL;DR

Online hate speech creates an urgent need to understand how hateful posts spread. This first study analyzes diffusion on Gab using repost and belief networks, finding that hateful posts spread farther, faster, and wider while hateful users are densely connected and generate nearly one-fifth of Gab’s content despite comprising 0.3% of users.

  • Problem

    The study addresses the need for better understanding of how hateful posts spread through online social media amid increasing hate crimes.

  • Method

    The paper conducts the first study of hateful-post diffusion on Gab, identifying hateful and non-hateful users with repost networks and DeGroot’s model before analyzing their diffusion and account characteristics.

  • Results

    Hateful users’ posts spread farther, faster, and wider; 0.3% of users generated 18.65% of Gab’s posts, and hateful users formed a dense network.

  • Takeaways & Limitations

    The findings broaden understanding of hate-speech diffusion by showing that hateful users are disproportionately influential and cohesive within Gab.

  • Takeaways & Limitations

    The model uses a random sample of 200 accounts per class to keep monetary costs manageable.

Abstract

from arXiv · show

The present online social media platform is afflicted with several issues, with hate speech being on the predominant forefront. The prevalence of online hate speech has fueled horrific real-world hate-crime such as the mass-genocide of Rohingya Muslims, communal violence in Colombo and the recent massacre in the Pittsburgh synagogue. Consequently, It is imperative to understand the diffusion of such hateful content in an online setting. We conduct the first study that analyses the flow and dynamics of posts generated by hateful and non-hateful users on Gab (gab.com) over a massive dataset of 341K users and 21M posts. Our observations confirms that hateful content diffuse farther, wider and faster and have a greater outreach than those of non-hateful users. A deeper inspection into the profiles and network of hateful and non-hateful users reveals that the former are more influential, popular and cohesive. Thus, our research explores the interesting facets of diffusion dynamics of hateful users and broadens our understanding of hate speech in the online world.

1 Introduction

This paper studies how hateful content diffuses online, using Gab as a setting with fewer restrictions on potentially hateful posts. It finds that hateful users and their posts achieve greater reach and faster, farther, wider diffusion than non-hateful users.

  • Research motivation and contribution: Online hate speech has been linked to severe real-world harms, including communal violence and hate crimes, motivating closer study of its spread.The paper cites allegations involving Facebook and anti-Muslim mob violence in Sri Lanka, alongside increasing hate crimes.
  • Research motivation and contribution: The study examines hate-speech diffusion on Gab, where users can post potentially hateful content without fear of repercussions.The authors describe this as the first study of hate diffusion dynamics in online social media.
  • Key findings: 21M posts from 341K users over 20 months show that hateful users’ posts spread faster, farther, and wider than normal users’ posts.The dataset spans October 2016 to June 2018.
  • Key findings: 0.3% of users generated 18.65% of Gab’s posts, indicating disproportionate content production by hateful users.The paper characterizes this group as densely connected and influential.

2 Dataset Description

The authors crawl Gab to assemble a large user-post-network dataset, identify hateful users through keywords and diffusion, and classify users by propagated belief. They validate the resulting hateful and non-hateful account sets with human judgments.

  • Data collection: Gab’s API was crawled using a snowball methodology to collect user details, posts, followers, and followings.The resulting dataset contains information gathered by expanding from a popular user through follower and following relationships.
  • Hateful-user identification: The method initializes 1,863 keyword-identified hateful users with belief 1, assigns others belief 0, and diffuses beliefs through a normalized repost network using DeGroot’s model.The diffusion process is intended to identify users with high hate potential through homophily despite lacking explicit hate keywords.
  • Hateful-user identification: After five diffusion iterations, users are stratified by belief, with high-belief users labeled hateful and low-belief users non-hateful when they have at least five posts.The final sets contain 1,055 hateful users and 62,827 non-hateful users.
  • Validation: Human evaluation found 86.9% and 93.2% of sampled hate accounts judged hateful, versus 92.2% and 99.4% of non-hateful accounts judged non-hateful.The corresponding Cohen’s κ scores were 0.69 and 0.87.

3 Diffusion dynamics of posts

The study models repost cascades on Gab to compare how posts from hateful and non-hateful users diffuse. Hateful-user posts generally reach larger audiences, spread more broadly and deeply, and propagate faster, with differences varying across community settings.

  • Model description: Cascades trace repost paths from a root user through the network, with LRIF modeling diffusion over follower-following connections.The analysis uses DAGs generated by the Least Recent Influencer Model because exact influence paths cannot be observed.
  • Characteristic cascade parameters: Size counts reachable users, breadth measures the largest number of nodes at one depth, depth is the longest root-to-node path, and structural virality averages pairwise DAG distances.Average depth is the average path length of nodes reachable from the root.
  • Characteristic differences in cascades: Hateful-user posts generate larger mean cascade size and breadth, indicating larger audiences and wider spread, although non-hateful cascades can have the larger maximum size.The size difference is especially pronounced during the initial stages of the distribution.
  • Characteristic differences in cascades: Hateful-user cascades have significantly greater depth, average depth, and structural virality, indicating deeper diffusion and more viral cascades.These properties remain consistently larger for hateful users across their distributions.
  • Attachments and community perspective: Posts with attachments and posts in groups or topics show stronger diffusion characteristics, while hateful and non-hateful posts do not differ significantly within groups.Groups are more exclusive than topics, and topic membership is larger, producing larger characteristic cascade values.
  • Early adopters and temporal evolution: Hateful users are early propagators of hateful cascades, and hateful cascades reach given diffusion values faster initially; hateful users are therefore described as more proactive and cohesive.The early-propagator pattern reflects strong homophily, while cascades deeper than four levels are extremely rare: 0.0065% for KH and 0.0057% for NH.

4 RELATED WORK

Prior work examined diffusion across several online platforms and content types, but had not studied hate diffusion; this paper addresses that gap using Gab, where hateful posts spread farther, wider, and faster initially.

  • Research had studied diffusion in fake news, LinkedIn, retweet cascades, rumors, Facebook, and Tumblr, but not the diffusion of hate in online social media.
  • Earlier Gab research found the platform centered on news, world events, and politics, with 2.4 times more hate speech than Twitter and hate disseminated by users exploiting weak moderation.
  • Most hate-speech research focused on detection across platforms including Twitter, Facebook, Yahoo! Finance and News, and Whisper.
  • Figure 5 shows hateful users’ posts spreading farther, wider, and deeper more quickly during the initial stages of diffusion.

5 Discussion

Hateful users differed significantly from non-hateful users in account and network characteristics, including activity, influence, density, reciprocity, and popularity. The study’s account-based labeling assumes that most reposts from hateful accounts are hateful, so some reposts may not be hateful.

  • All measured account characteristics differed significantly between hateful and non-hateful users, with p-value<0.001 for every comparison.Measures included time-normalized posts, followers, and followings, plus post-normalized likes, dislikes, replies, and reposts.
  • 0.3% of users generated 18.65% of all posts, indicating significantly greater influence among hateful users.Hateful users generated 3.95M posts, while non-hateful users generated 11.22M posts, or 52.94% of all posts.
  • The hateful-user network was approximately 20 times denser than the non-hateful-user network and had higher reciprocity, 58.3% versus 55.9%.
  • Non-hateful users were 6.8 times more likely to follow hateful users than the reverse, consistent with greater popularity among hateful users.
  • The account-based analysis assumes hateful accounts generate most reposts, so some reposts attributed to them may not themselves be hateful.The authors state that this assumption means the analysis does not capture the full picture of hateful-post cascades.

6 Conclusion and Future work

The study compares hateful and non-hateful users’ post diffusion on Gab, finding that hateful posts spread farther, faster, and wider despite hateful users comprising only 0.3% of users.

  • Hateful users comprise 0.3% of Gab users yet generate almost 1/5th of its content.The study also finds these users are densely connected with each other.
  • Posts by hateful users tend to spread farther, faster, and wider than posts by non-hateful users.
  • The analysis identifies image-heavy hateful-user posts as a direction for future hateful image and video classification.
  • Future research could examine diffusion characteristics of individual hateful posts rather than only hateful accounts.
Loading 1812.01693v1…