Source-linked AI summary
"HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
Lingyao Li, Lizhou Fan, Shubham Atreja, Libby Hemphill
TL;DR
Harmful social-media content is difficult and costly to annotate manually, creating a need for scalable detection methods. This study compares ChatGPT with MTurker annotations across HOT classification experiments and varies prompts and requested outputs. ChatGPT reaches approximately 80% accuracy, is more consistent on non-HOT than HOT comments, treats hateful and offensive content as subsets of toxic content, and is affected by prompt choice.
Problem
Manual annotation of harmful content is costly and exposes annotators to offensive material, while ChatGPT’s ability to reliably distinguish HOT concepts remains insufficiently understood.
Method
The study compares ChatGPT classifications with MTurker annotations for Hateful, Offensive, and Toxic content using five prompts across four experiments.
Results
Approximately 80% accuracy was achieved against MTurker annotations, with greater consistency for non-HOT than HOT comments; ChatGPT also grouped hateful and offensive content within toxic content, and prompt choice affected performance.
Takeaways & Limitations
ChatGPT shows potential for scalable HOT annotation, but its reliability, concept discrimination, and performance depend on the prompting approach.
Takeaways & Limitations
MTurker annotations may not represent accurate ground truth, so the study establishes agreement with MTurkers rather than accuracy against an independent standard.
Abstract
from arXiv · showhide
Harmful content is pervasive on social media, poisoning online communities and negatively impacting participation. A common approach to address this issue is to develop detection models that rely on human annotations. However, the tasks required to build such models expose annotators to harmful and offensive content and may require significant time and cost to complete. Generative AI models have the potential to understand and detect harmful content. To investigate this potential, we used ChatGPT and compared its performance with MTurker annotations for three frequently discussed concepts related to harmful content: Hateful, Offensive, and Toxic (HOT). We designed five prompts to interact with ChatGPT and conducted four experiments eliciting HOT classifications. Our results show that ChatGPT can achieve an accuracy of approximately 80% when compared to MTurker annotations. Specifically, the model displays a more consistent classification for non-HOT comments than HOT comments compared to human annotations. Our findings also suggest that ChatGPT classifications align with provided HOT definitions, but ChatGPT classifies "hateful" and "offensive" as subsets of "toxic." Moreover, the choice of prompts used to interact with ChatGPT impacts its performance. Based on these in-sights, our study provides several meaningful implications for employing ChatGPT to detect HOT content, particularly regarding the reliability and consistency of its performance, its understand-ing and reasoning of the HOT concept, and the impact of prompts on its performance. Overall, our study provides guidance about the potential of using generative AI models to moderate large volumes of user-generated content on social media.
1. Introduction
Harmful social-media content motivates safer, scalable detection, but human annotation is costly and exposes annotators to offensive material. This study examines ChatGPT’s reliability, HOT reasoning, and sensitivity to prompting compared with MTurker annotations.
- Harmful online behavior can create hostile environments, reduce participation, and harm individuals, making detection and moderation important.
- Human annotation exposes workers to harmful content and requires substantial time and financial resources, motivating alternative detection approaches.
- Generative AI may detect subtle toxicity, including sarcasm or irony, because it can identify patterns and produce human-like responses.
- ChatGPT’s HOT classifications are evaluated against MTurker annotations across reliability, reasoning, and prompt effects.The study asks how ChatGPT compares with MTurkers, understands distinctions among HOT concepts, and responds to different prompts.
- The study aims to support harm-free annotation and scalable moderation workflows for large volumes of social-media content.
2. Backgrounds
Background research frames HOT detection as important because harmful language damages users and communities, while definitions and annotation practices remain varied and costly. Generative AI, particularly LLMs, offers a promising but still insufficiently understood alternative for identifying and explaining HOT content.
- Harmful language can harm individuals and communities, while automated detection may reduce moderation delays and support timely intervention.
- Different HOT definitions matter because they shape training-data quality and downstream content-moderation models.
- Prior HOT detection progressed from lexicons and text-vectorization methods toward machine-learning and natural-language-processing models.
- Human annotation underlies supervised HOT models but can expose workers to psychological harm, reduce accuracy, and impose substantial costs.
- LLMs use large-scale pretraining and fine-tuning to process and generate token sequences through transformer-based architectures.
- Although LLMs may match human annotators and explain implicit HOT content, their reliability, consistency, and reasoning remain insufficiently understood.
3. Data and methods
The study compares ChatGPT with MTurker annotations for hateful, offensive, and toxic content using a shared dataset, standardized definitions, and multiple prompts. It evaluates reliability, consistency, HOT comprehension, and prompt effects through a framework of experiments.
- Research framework: The framework compares ChatGPT classifications with MTurker annotations across experiments addressing reliability, consistency, HOT reasoning, and prompt effects.Experiments 1–3 address reliability and consistency, Experiment 4 examines reasoning, and multiple prompts assess prompt effects.
- Dataset and annotations: The dataset contains comments from Reddit, Twitter, and YouTube, with purposive sampling used because HOT comments are relatively rare.The dataset focuses primarily on popular political news stories.
- Dataset and annotations: Five independent MTurkers determine each final HOT label by majority vote, with at least three True annotations required for a HOT classification.The study avoids treating the five annotations as a probability because their representativeness and reliability are uncertain.
- Dataset and annotations: MTurker annotations include 2,381 non-HOT and 263 HOT comments among 3,481 comments, with substantial overlap among HOT concepts.For example, 622 of 803 toxic comments were also classified as offensive; ChatGPT is considered accurate when matching the MTurker majority label.
- HOT definitions: The study provides common definitions of hateful, offensive, and toxic content to both MTurkers and ChatGPT to standardize concept interpretation.The supplied definitions characterize hateful content as targeting or insulting a group and offensive content as hurtful, derogatory, or obscene.
- ChatGPT and prompt design: The study uses gpt-3.5-turbo and five prompts that vary output format, explanations, and instructions for binary or probabilistic HOT classification.Prompts 2 and 3 request binary or probability outputs without explanations, while Prompts 4 and 5 request explanations alongside those outputs.
4. Results
Across four experiments, ChatGPT’s HOT classifications were generally consistent but depended on the prompt and category. The model aligned hateful and offensive content largely within toxic content, while showing stronger performance on non-HOT comments than HOT comments.
- 4.1. Results of Experiment 1 – direct comparison with MTurkers: ChatGPT achieved better F1-scores for non-HOT than HOT comments across hateful, offensive, and toxic categories, especially for non-hateful comments.It showed higher agreement than MTurkers for offensive and hateful categories but lower agreement for toxic comments.
- 4.3. Results of Experiment 3 – consistency: ChatGPT generated consistent HOT annotations across binary and probability prompts and temperatures, with agreements above 90%.Agreement was slightly lower for hateful comments in two prompt–temperature combinations.
- 4.4. Result of Experiment 4 – annotation reasoning: ChatGPT classified hateful and offensive comments largely as toxic, reflecting a lower threshold for toxicity than for hateful content.Among 3,470 comments, 491 were toxic and offensive but not hateful, while 433 were toxic without being hateful or offensive.
5. Discussion
The study finds that ChatGPT can provide reliable HOT annotations compared with MTurkers, but agreement varies by HOT category and prompt design. Its outputs align with supplied definitions while revealing limitations in reasoning, probability interpretation, and ground-truth validity.
- Reliability and consistency: Approximately 80% accuracy and over 90% response consistency show that ChatGPT can reliably annotate HOT content compared with MTurkers.Agreement is stronger for non-HOT than HOT comments.
- Reliability and consistency: ChatGPT agrees more consistently with MTurkers on non-HOT comments than HOT comments, with especially substantial disagreement for hateful classifications.The study reports higher F1-scores for non-HOT comments and less agreement for HOT comments.
- Reasoning about HOT concepts: ChatGPT tends to treat hateful and offensive comments as subsets of toxic content and often repeats the supplied HOT definitions when explaining classifications.The authors interpret this pattern as a lower threshold for toxic labels and note that reasoning may reflect prompt conformity rather than generalization.
- Prompt effects: Prompt design affects performance: explicit definitions and requested output formats matter, while explanations may increase HOT classifications without improving agreement with human annotators.Prompts requesting explanations did not show clear F1-score or accuracy improvements over simpler prompts.
- Practical implications: Probability outputs require caution because ChatGPT produces few uncertain scores and intermediate probabilities may not reliably represent HOT intensity.The study specifically questions the interpretability of probabilities in the uncertain range.
- Limitations and future work: The reported ChatGPT–MTurker agreement is not equivalent to validated HOT accuracy because MTurker labels may not represent an appropriate ground truth.The authors recommend expert annotations or additional datasets for future evaluation.
6. Conclusions
The study evaluates ChatGPT for annotating hateful, offensive, and toxic comments against MTurker labels. ChatGPT reaches approximately 80% accuracy and over 90% response consistency, but agreement is weaker for HOT comments and probability outputs, reasoning, and prompts remain consequential.
- Conclusions: Approximately 80% accuracy and over 90% repeated responses indicate ChatGPT can annotate HOT comments reliably relative to MTurkers.Agreement is higher for non-HOT than HOT comments.
- Conclusions: ChatGPT agrees less with MTurkers on HOT comments, especially hateful comments, and produces relatively few intermediate probability values.The conclusion also reports conformity to supplied definitions and uncertainty about reasoning generalization.
- Conclusions: Different prompts affect performance, while explanations may yield more conservative outputs without necessarily improving agreement with human annotators.The study presents ChatGPT as potentially useful for quickly and cheaply annotating large content samples.
Appendix A.
Appendix A lists the five prompt designs used to elicit HOT classifications from ChatGPT. The prompts vary output type, explanation requirements, and whether the task mirrors MTurker binary annotation.
- Prompt designs: The five prompts vary binary versus probability outputs and whether ChatGPT must explain its classification.Prompts 2 and 3 request binary or probability outputs without explanations, whereas Prompts 4 and 5 request explanations.
- Prompt designs: Prompt 1 asks ChatGPT to provide binary HOT judgments in the same format used for MTurkers.The prompt uses the question format “Do you think this comment is hateful? (1) Yes, (2) No.”
- HOT definitions: The prompts define hateful content as targeted-group hatred or insult, offensive content as hurtful, derogatory, or obscene language, and toxic content as rude, disrespectful, or unreasonable language likely to drive readers away.These definitions are embedded in the binary and probability prompt templates.
- Output constraints: Binary prompts explicitly restrict ChatGPT to yes-or-no outputs, while probability prompts request scores from 0 to 1.Some probability prompts additionally require an explanation after the score.
Appendix B.
Appendix B describes an apple-to-apple comparison between ChatGPT probability outputs and MTurker annotation scores. ChatGPT probabilities are mapped to six score intervals corresponding to MTurker vote proportions.
- Comparison procedure: ChatGPT probability outputs are transformed into six intervals so they can be compared directly with MTurker annotation scores.The comparison uses the probability output from Prompt 2.
- Score mapping: Probabilities from 0.0–0.1 map to MTurker score 0.0, while 0.9–1.0 maps to score 1.0.Intermediate intervals map to scores 0.2, 0.4, 0.6, and 0.8.
- Evaluation metrics: The appendix table reports support, precision, recall, F1-score, and accuracy for the hateful category.The displayed table header identifies the evaluation metrics and category.
Appendix C.
Table C illustrates how ChatGPT applies HOT categories across comments, including cases classified as toxic without being hateful or offensive. The examples also show that ChatGPT may distinguish group-targeted hatred from insults, derogatory remarks, and broader toxicity.
- Table C presents examples of ChatGPT outputs for comments classified as toxic but not hateful or offensive, toxic and offensive but not hateful, and offensive but not hateful or toxic.
- ChatGPT distinguishes insults directed at individuals from hateful comments targeting specific groups, classifying accusations and derogatory remarks about individuals as non-hateful.
- Some critical, sarcastic, or opinion-based comments are judged non-offensive and non-hateful when they lack explicit derogatory language or group targeting.
- Comments containing disrespectful, unreasonable, or inflammatory language are described as toxic because they may provoke negative reactions or discourage participation.
- ChatGPT labels derogatory or hurtful remarks about individuals as offensive, including comments about physical appearance, intelligence, or personal conduct.