Source-linked AI summary
Peer to Peer Hate: Hate Speech Instigators and Their Targets
Mai ElSherief, Shirin Nilizadeh, Dana Nguyen, Giovanni Vigna, Elizabeth Belding
TL;DR
Online hate speech research has lacked comparative evidence about the actors who instigate and receive hate. This paper curates a high-precision Twitter dataset and compares instigators and targets with general users, finding that hate participation and visibility are correlated while both groups show distinctive personality traits. The study’s scope is limited by keyword-based filtering and its focus on explicit hate speech.
Problem
Little is known about online hate speech instigators and targets, and accurate detection lacks a commonly accepted benchmark corpus.
Method
The authors curate a high-precision dataset through semi-automated, keyword-based filtering and compare instigators, targets, and general Twitter users across profiles, visibility, activity, and personality.
Results
Hate instigators target more visible users, and participating in hate commentary is associated with higher visibility; instigators and targets also show distinctive personality characteristics.
Takeaways & Limitations
The findings provide comparative actor-level information that may support hate speech classification, detection, and mitigation.
Takeaways & Limitations
The analysis focuses on explicit hate speech and keyword-based methods, so it does not capture a complete representation of hate speech on Twitter.
Abstract
from arXiv · showhide
While social media has become an empowering agent to individual voices and freedom of expression, it also facilitates anti-social behaviors including online harassment, cyberbullying, and hate speech. In this paper, we present the first comparative study of hate speech instigators and target users on Twitter. Through a multi-step classification process, we curate a comprehensive hate speech dataset capturing various types of hate. We study the distinctive characteristics of hate instigators and targets in terms of their profile self-presentation, activities, and online visibility. We find that hate instigators target more popular and high profile Twitter users, and that participating in hate speech can result in greater online visibility. We conduct a personality analysis of hate instigators and targets and show that both groups have eccentric personality facets that differ from the general Twitter population. Our results advance the state of the art of understanding online hate speech engagement.
Introduction
The paper addresses limited knowledge about online hate speech actors by comparing instigators, targets, and general Twitter users. It contributes a curated dataset and analyses differences in profiles, visibility, activity, and personality.
- 60% of Internet users had witnessed offensive name calling, while 25% saw physical threats and 24% saw sustained harassment.
- The study focuses on hate speech that denigrates people because of innate and protected characteristics.
- The paper presents the first comparative study of hate instigators and targets, using 27,330 hate tweets and 25,278 instigator and 22,287 target accounts.
- The authors compare hate instigators, targets, and general users across profile self-presentation, Twitter visibility, and personality traits.
- The dataset includes a 51-term Hatebase lexicon spanning eight hate classes and a semi-automated classification method for directed explicit hate speech.
- Targets often have older accounts, both groups are more active than general users, and targets include 60% more verified accounts than instigators.
Related Work
Prior research has examined abusive language detection, hate speech triggers, trolling, and targeted abuse, but this paper adds a comparative analysis of instigators and targets.
- Earlier work studied abusive messages, cyberbullying, personal insults, and offensive language across platforms including Twitter and YouTube.
- Hate speech characterization research has linked online hate to trigger events such as terrorist attacks, crime, and news.
- Cheng et al. found that prior negative mood and discussion context can double baseline engagement in trolling behavior.
- Unlike studies focused on sentiment or generalized abuse, this work analyzes both instigators and targets and compares them with general Twitter users.
Preliminaries
The paper studies directed explicit hate speech between Twitter accounts and defines the roles of hate tweets, instigators, and targets.
- Directed abuse targets a specific individual or entity, whereas generalized abuse addresses a group sharing a characteristic.
- A hate tweet is an explicit directed tweet containing one or more hate speech terms used against a Twitter account holder.
- A hate instigator is a Twitter account that posts one or more hate tweets.
- A hate target is an account explicitly mentioned in a hate tweet, and role labels are not mutually exclusive.
Data and Methods
The paper builds a high-precision directed hate-speech dataset through multi-stage filtering, compares hate instigators, targets, and general users, and validates the resulting labels with human annotation.
- Motivation: The study addresses the difficulty of hate-speech detection and role labeling, including scarce benchmark data, definitional disagreement, and incomplete conversation threads.These challenges complicate identifying instigators, targets, and bystanders from social-networking data.
- Dataset construction: The authors define hate speech as direct and serious attacks on protected categories and collect explicit directed attacks between Twitter accounts.The final filtering retains tweets mentioning another account and containing second-person pronouns.
- Dataset construction: The multi-step pipeline combines Hatebase keyphrases, Perspective API toxicity and attack-on-commenter scores, and account-directedness filters.The selected thresholds are 0.8 for toxicity and 0.5 for attack on commenter.
- Datasets: The resulting dataset contains 27,330 hate tweets, while the general comparison dataset samples 60,000 users from Twitter accounts outside the hate-speech dataset.The general sample is drawn from the same 18-month collection window and is designed to give users equal selection probability regardless of activity.
- Dataset composition: Gender, disability, and sexual-orientation classes contribute 52%, 32%, and 14% of instigators despite using only 4, 2, and 2 keywords, respectively.The authors interpret these contributions as evidence of the prevalence of the corresponding hate keywords on Twitter.
- Human validation: 97.8% of sampled tweets were labeled hate speech and 94.3% were labeled attacks on the mentioned account by majority vote.Agreement was 92.8% for hate speech and 82.6% for direct attack, based on annotations from at least three independent annotators per tweet.
Analysis
Hate targets are generally older, more active, more visible, and more established than instigators and general users, while instigators are less likely to be verified. Personality distributions show that instigators and targets resemble each other more than they resemble general Twitter users, with shared differences across several Big Five facets.
- Profile self-presentation: Targets have older accounts than instigators and general users, while instigators have slightly younger accounts than general users.Mean account ages are 4.40 years for targets, 3.67 for instigators, and 3.73 for general users.
- Profile self-presentation: Targets disclose more profile information and have longer descriptions than instigators and general users.Targets are more likely to add images, URLs, locations, and timezones; mean description lengths are 63, 53, and 45 characters for targets, instigators, and general users, respectively.
- Activity and visibility: Targets are more established and visible, including 12% verified accounts, while instigators are less likely to be verified than general users.Targets have more friends, tweets, followers, and retweets, and are listed more often; their visibility differences are strongest for followers, lists, and retweets.
- Activity and visibility: 14.64, 6.92, and 57.94 are the target-to-other-user incidence-rate ratios for followers, lists, and retweet counts, respectively.These differences remain significant when targets are compared with instigators alone, and targets are more visible across quartiles.
- Activity and visibility: Instigators and targets are both positively associated with visibility, and participating in hate speech is related to greater popularity and visibility.The reported associations hold after controlling for activity, profile self-presentation, gender, and other mentioned variables.
- Personality traits: Instigators and targets have more similar personality distributions to each other than to general users, including lower Agreeableness, Conscientiousness, Emotionality, and Adventurousness.They also show higher Imagination and lower several Extraversion facets; the Hellinger distance between instigators and targets is always less than or equal to their distances from general users across the reported traits.
Discussion and Conclusion
The study links hate-speech participation and actor characteristics to visibility, personality, and potential mitigation strategies. It also emphasizes that its high-precision dataset captures only explicit hate speech and is not a complete representation of Twitter hate speech.
- Hate mitigation and counter speech: Personality analyses could inform more effective counter-speech bots by tailoring messages to users’ personality facets.The paper connects personality-tailored messaging with counter-speech design and reports that 50% of hate instigators and targets score above 0.53 for Openness to change.
- Profile-based data collection: Personality scores could also be incorporated as features in profile-based hate-speech data collection.This extends existing collection methods that classify accounts using offensive-term usage.
- Critique of methodology and limitations: The methodology focuses on explicit hate speech and keyword-based filtering, which can miss hateful speech and cannot provide a complete representation of Twitter hate speech.The authors prioritize a high-precision dataset of instigator and target accounts rather than complete coverage.
- Discussion and Conclusion: The analysis found that hate instigators target more visible users and that hate-commentary participation is associated with higher visibility.The study also identifies distinctive personality characteristics among instigators and targets, including anger, depression, and immoderation.