Source-linked AI summary

Directions in Abusive Language Training Data: Garbage In, Garbage Out

Bertie Vidgen, Leon Derczynski

arXiv:2004.01670v3cs.CL

TL;DR

Online-abuse detection depends on training datasets, yet their quality, coverage, and focus have received insufficient systematic attention. This paper reviews 63 datasets and dataset-sharing practices, finding substantial problems in representativeness, context, modality, annotation documentation, and data integrity. It synthesizes these findings into evidence-based recommendations and introduces an open website for cataloguing abusive-language data.

  • Problem

    Training datasets are fundamental to abusive-content detection, but their quality, coverage, focus, and sharing have not been examined systematically enough.

  • Method

    The paper systematically reviews publicly available abusive-content datasets, analyzing their content, tasks, annotations, and sharing practices.

  • Results

    The 63 reviewed datasets are often unrepresentative of abuse in the wild, mostly post-level and text-only, and frequently lack detailed annotation guidelines.

  • Takeaways & Limitations

    Dataset creation and evaluation should account for purpose, class distribution, contextual and multimodal coverage, annotation practice, and responsible sharing.

  • Takeaways & Limitations

    Dataset integrity is constrained by content loss and degradation when researchers must rehydrate datasets from annotations and IDs.

Abstract

from arXiv · show

Data-driven analysis and detection of abusive online content covers many different tasks, phenomena, contexts, and methodologies. This paper systematically reviews abusive language dataset creation and content in conjunction with an open website for cataloguing abusive language data. This collection of knowledge leads to a synthesis providing evidence-based recommendations for practitioners working with this complex and highly diverse data.

Background

Abusive-content detection has expanded alongside recognition of online harms and policy pressure, but its systems depend on training datasets that remain underexamined and flawed. Prior reviews identify concerns involving data scarcity, bias, annotation quality, and variation in task definitions.

  • Background: Research on online abuse has grown sharply as awareness of Internet communication, social harms, and regulatory developments has increased.The reviewed background connects this expansion to initiatives including the EU Code of Conduct, the UK Online Harms white paper, and national regulations.
  • Background: Online abuse detection supports social research, content moderation, and interventions for vulnerable users, but must process enormous volumes of online content.Human moderation remains important but can impose health and economic costs, motivating combined computational and human approaches.
  • Background: Nearly all abusive-content detection systems rely on training datasets that teach systems what counts as abuse.The paper identifies training data as a crucial point where social-scientific insights are needed.
  • Background: Existing reviews discuss abusive-content detection broadly, but none examine training datasets with sufficient breadth or depth.This gap is notable because datasets are fundamental to detection systems and several existing datasets have documented flaws.
  • Background: Prior work identifies small, linguistically narrow datasets, inaccessible metadata, systematic target biases, annotation uncertainty, and substantial variation in dataset quality and class balance.These concerns span data sparsity, dataset bias, degradation, annotation, and variety.

Objectives

The study systematically examines the quality, coverage, focus, sharing, and creation of abusive-content training datasets. It combines critical dataset analysis with open-science infrastructure and recommendations for dataset practice.

  • Objectives: The paper provides an in-depth and critical analysis of available training datasets for abusive online-content detection.This is the first stated research aim.
  • Objectives: It identifies best practices for creating abusive-content training datasets.The recommendations are intended to be evidence-based and grounded in the review findings.
  • Objectives: It maps ways to address limited dataset sharing and the lack of open science in online-abuse research.The aim concerns improving how datasets can be shared and used by researchers.
  • Objectives: It introduces hatespeechdata.com as a mechanism for enabling more dataset sharing.The website is presented as the paper’s open-science contribution.
  • Objectives: The selection criteria, protocol, and methodology follow the PRISMA framework.The process is documented as following published PRISMA guidance.

Sources of content

The review identifies abusive-content training datasets across three academic sources using related keyword searches. The sources span multidisciplinary research, computational linguistics, and open-access preprints.

  • Sources of content: The researchers identified datasets from Scopus, the ACL Anthology, and arXiv to cover a range of academic venues.These sources represent cross-disciplinary research, computational linguistics and NLP, and open-access preprints, respectively.
  • Sources of content: Scopus was used as a cross-disciplinary database indexing journal papers and conference proceedings with quality-control processes.The paper describes Scopus as comparable in coverage to Web of Science and Google Scholar.
  • Sources of content: The ACL Anthology supplied research in computational linguistics and NLP, including workshops on abusive language online.It includes proceedings from the first three workshops on abusive language online.
  • Sources of content: ArXiv provided open-access preprints that do not undergo full peer review.The source includes early drafts, pre-accepted research, and accepted research after official publication.
  • Sources of content: The researchers searched all three databases with related abuse and online-content keywords, modifying the Scopus query with wildcards.Additional terms were excluded because they returned many irrelevant publications.

Study selection

The study manually screened publications for publicly available annotated datasets and analyzed their content, taxonomies, motivations, and annotation practices. It highlights difficult trade-offs between taxonomic detail, dataset compatibility, representativeness, and annotation quality.

  • Study selection: Two researchers screened publications for publicly available abusive-content training datasets, followed by third-researcher re-review and attempted dataset collection.The process corrected identification errors and excluded cases where datasets were not actually described or obtainable.
  • Study selection: The analysis critically examined dataset content rather than conducting a meta-analysis or applying publication-bias weighting.The authors state that common statistical tests were inappropriate because the total number of datasets was very low.
  • Study selection: Dataset analysis partly adopted the data-statements framework to document how NLP datasets were created and how decisions were made.The framework was combined with other schema and process-oriented work for analyzing NLP artefacts.
  • Study selection: Dataset motivations included reducing harm, removing illegal content, improving online conversations, and reducing the burden on human moderators.These motivations connect dataset design to social problems, platform governance, community health, and moderator welfare.
  • Detection tasks: Abuse datasets vary across phenomena, taxonomies, targets, and intended uses, making their categories difficult to organize and datasets difficult to combine.The review extends the person-versus-group distinction to four primary categories plus a Mixed category.
  • Detection tasks: Dirty-word detection can mistake intensifiers or colloquialisms for abuse, while incivility judgments may require inferring speaker intent and social norms.These conditions make some forms of abuse difficult to annotate or detect reliably.
  • Detection tasks: More detailed taxonomies require enough examples per category and consistent annotation, while overly broad categories can create within-class variation.Dataset creators therefore face a trade-off between conceptual distinctions and reliable, usable model training.

The content of training datasets

The 63 datasets vary substantially in what they contain, how data was sampled, and which languages and platforms they represent. These choices create constraints on linguistic coverage, contextual understanding, representativeness, and generalisability.

  • The level of content: 62 of 63 datasets are annotated at the post level, none at the thread level, and 62 contain text only, limiting conversational and multimodal coverage.Only three datasets indicate that annotators saw entire conversation threads, and only one dataset contains images.
  • Language: English appears in 25 datasets, while languages such as Farsi and Russian are absent and several European languages are poorly represented.The review identifies considerable unevenness in the linguistic and cultural focus of abusive-language classification.
  • Source of data: Twitter accounts for 39 of 63 datasets, while many major platforms are sparsely represented or absent, limiting cross-platform robustness.Platform affordances, user demographics, and norms shape content style, users, and abuse patterns, making cross-domain application difficult.
  • Dataset size: Dataset size ranges from 469 posts to 17 million, but larger datasets are not necessarily better when sampling, categories, or annotation are poor.Smaller datasets can lack linguistic variation and increase overfitting, while no established guideline defines the necessary size.
  • Class distribution and sampling: 36.7% of dataset content is abusive on average, ranging from 1% to 100%, whereas abuse prevalence in the wild is likely below 1%.This makes most datasets and their evaluations radically unrepresentative of likely deployment settings.
  • Class distribution and sampling: Sampling strategy shapes class distribution and dataset integrity, with purposive keywords, time windows, or communities introducing bias and limiting generalisability.Oversampled keywords can produce racial bias, while the overrepresentation of far-right communities may cause classifiers to learn political-discourse styles rather than core hate-speech features.
  • Identity of the content creators: Content-creator identities are fully described in only four datasets, and concentrated authorship can make user-level metadata artificially predictive of abuse.In the Waseem and Hovy dataset, 70% of sexist tweets came from two creators and 99% of racist tweets from one.

Annotation of training datasets

The review finds substantial variation in how abusive-language datasets are annotated, with limited reporting about annotators and annotation guidelines. Expert annotation is common, but ambiguity in abuse concepts and annotator backgrounds complicates consistency and interpretation.

  • Annotation process: Five annotation processes appear across the 63 datasets: experts or trained workers, crowdsourcing, professional moderators, mixed approaches, and synthetic data creation.Experts or trained workers account for 28 datasets (44%), crowdsourcing for 21 (33%), professional moderators for 3 (5%), mixed approaches for 7 (11%), and synthetic creation for 4 (6%).
  • Annotation process: 28 datasets (44%) use experts or trained workers, while 21 (33%) use crowdsourcing, trading annotation quality for scalability and lower cost.The review describes expert annotation as time-intensive and generally higher quality, whereas crowdsourcing uses many non-experts and trades quality for quantity.
  • Identity of the annotators: A homogeneous annotator pool may miss covert slang and coded abuse because no individual is likely to understand all such meanings.The review connects annotator diversity with the ability to identify varied and deliberately obfuscated forms of abuse.
  • Identity of the annotators: Annotator identities are undocumented in 27 datasets, minimally described in 24, and detailed in only 12, including just four English-language datasets.The review stresses that demographic information, expertise, and personal experiences can shape annotation and should be reported more systematically.
  • Guidelines for annotation: Only 16 datasets (25%) provide detailed annotation guidelines; 27 (43%) provide none and 20 (32%) provide only highly summarized guidance.The paper argues that fuller annotation information would support understanding dataset contents and improving or extending datasets.
  • Ambiguous abuse phenomena: Irony, calumniation, intent, and truthfulness create difficult annotation problems because their interpretation can be fundamentally indeterminate or hard to infer from text.The paper notes that guidelines or explicit positions can improve consistency, while truth and falsity remain notably underexamined.

Challenges and opportunities of achieving open science

Open sharing could improve collaboration, reproducibility, and research access, but abusive-language data raises acute ethical, legal, privacy, integrity, and security challenges. The paper considers synthetic data, donations, platform-backed sharing, and data trusts as possible responses.

  • Benefits of open science: Dataset sharing could increase collaboration, reproducibility, and researchers’ ability to identify limitations and new directions.The paper also frames access as a fairness and power issue because less-well-funded researchers and organizations could contribute more substantially.
  • Challenges: Abusive-language sharing raises privacy, consent, ethical, legal, and platform Terms of Service concerns, especially when users are identifiable.Researchers often rely on implicit consent from public or semi-public posting, while some highly cited datasets remain unavailable.
  • Challenges: Dataset integrity is threatened when content must be rehydrated from platform IDs, as demonstrated by substantial degradation after tweets became unavailable.The paper identifies integrity maintenance as one of two central requirements for future dataset sharing.
  • Challenges: Secure permissioning must enable legitimate research while respecting platform rules and restricting malicious access.The paper uses the history of Jihadology to illustrate the difficulty of providing access to sensitive research materials.
  • Potential solutions: Synthetic datasets avoid platform Terms of Service restrictions but may introduce bias, lack sustainability, and remain ethically sensitive.Their non-authenticity can limit representativeness, although carefully created synthetic data may broaden coverage of abuse types.
  • Potential solutions: Platform-backed sharing could provide original content and moderation metadata while removing Terms of Service limitations, but requires stronger academia–industry interfaces.The paper proposes secure platform/API access or waiving Terms of Service restrictions for approved research.
  • Potential solutions: The paper presents a permissioned data trust as a promising framework for combining sharing solutions while controlling access and reducing researchers’ discovery burden.Different permission levels could address varying privacy, ethical, commercial, or research sensitivities.

A new repository of training datasets: Hatespeechdata.com

The authors launched hatespeechdata.com as an interim repository for abusive-language datasets and related documentation while larger sharing infrastructure is developed.

  • Repository scope: Hatespeechdata.com catalogs the datasets analyzed in the review, partial data statements, and previously published abusive keyword dictionaries.The site includes only information and data already made publicly available by the original dataset creators and is intended to be updated.
  • Interim solution: The website provides an interim sharing mechanism because a dedicated data trust and API would require substantial resources and field-wide researcher buy-in.The repository is therefore a practical step toward greater dataset sharing rather than a replacement for the proposed infrastructure.

Best practices for training dataset creation

The paper derives dataset-creation best practices from existing efforts and organizes them around four stages of the training-data process.

  • Best-practice framework: The review identifies best practices at four points: task formation, dataset creation, annotation, and documentation.These stages provide the paper’s organizing framework for recommendations on creating abusive-language training datasets.

Task formation: Defining the task addressed by the dataset

Dataset creation should be problem driven, beginning with a clear motivation and a well-defined, specific task. The task should guide taxonomy design and account for abusive language’s complexity and terminological disagreement.

  • Dataset creators should define a clear motivation that addresses a specific abusive-language task.
  • The task should directly inform taxonomy design through engagement with relevant social scientific theory.
  • Explicit task definition is especially important because abusive language is complex and terminology remains contested.

Creating datasets for abusive language annotation

Dataset construction requires deliberate choices about languages, sources, sampling, annotator expertise and diversity, worker support, and iterative guidelines. These choices should address abusive language’s rarity, variation, context dependence, and potential annotation bias.

  • Dataset creators should choose annotated languages, data sources, sampling procedures, target size, and class distribution deliberately.
  • Because abuse is rare, creators should compare sampling options and consider combining data across times, locations, users, and platforms.
  • Annotator pools should combine relevant skills and experience with diversity, avoiding overreliance on narrow groups such as university students or men.
  • Annotators should receive emotional and practical support and appropriate compensation because annotating abuse can be harmful and triggering.
  • Guidelines should use clear examples and edge cases, address difficult phenomena such as irony and intent, and be developed iteratively.

Documenting methods, data, and annotation

Reliable abusive-language systems depend on documented datasets, appropriate data selection, context-sensitive validation, and deployment-aware evaluation. The review finds that many surveyed datasets provide insufficient methodological information, limiting interpretation and appropriate use.

  • Well-documented datasets can support new analyses and preserve annotator-level codings for alternative labeling strategies.
  • Most of the 63 surveyed datasets had limited methodological descriptions, making dataset biases, limitations, and appropriate system uses difficult to assess.
  • Dataset documentation should report the full creation process, including crowdsourcing parameters and annotation-interface design.
  • Dataset selection should match the target task, text genre, expected class balance, and breadth of abusive phenomena.
  • Validation data is particularly valuable for varied abusive language and should include live, in-situ data when matching a specific application.
  • Model evaluation and selection should consider the relative costs of false negatives and false positives in the intended deployment setting.

Best practice summary

The paper recommends purpose-driven, diverse, theoretically grounded, carefully annotated, and fully documented datasets. These practices aim to improve dataset quality and interoperability while balancing standardization with research innovation and freedom.

  • Dataset design should address a clear research purpose and anticipate implications for privacy, protection from harm, and freedom of speech.
  • Creators should seek diverse data sources, account for sampling bias, and size datasets according to sparsity and positive-class needs.
  • Taxonomies should use meaningful, theoretically sound categories and include complex phenomena such as irony, sarcasm, and adversarial content.
  • Guidelines should be developed iteratively with trained, contextually aware annotators, while addressing category boundaries and annotation bias.
  • A Data Statement should document every research step, including challenges, biases, limitations, and details useful to future researchers.
  • The review examines 63 publicly available datasets and uses its evidence-driven analysis to recommend better availability, usefulness, open-science infrastructure, and dataset sharing.
Loading 2004.01670v3…