Source-linked AI summary

Disturbed YouTube for Kids: Characterizing and Detecting Inappropriate Videos Targeting Young Children

Kostantinos Papadamou, Antonis Papasavva, Savvas Zannettou, Jeremy Blackburn, Nicolas Kourtellis, Ilias Leontiadis, Gianluca Stringhini, Michael Sirivianos

arXiv:1901.07046v3cs.SIcs.CY

TL;DR

Inappropriate videos targeting toddlers remain difficult to identify on YouTube and can be surfaced through recommendations despite existing mitigation efforts. The paper characterizes this content at scale and develops a binary classifier, finding 84.3% accuracy and a 3.5% chance of encountering an inappropriate video within ten recommendation hops from a benign starting point.

  • Problem

    YouTube contains inappropriate videos targeting toddlers, while recommendation similarity and limited detection make the problem difficult to control.

  • Method

    The authors manually review and categorize toddler-oriented videos, then train a deep learning classifier and use it for large-scale content and recommendation-walk analysis.

  • Results

    84.3% accuracy: the binary classifier outperforms several baselines, while simulations find a 3.5% chance of encountering an inappropriate video within ten recommendation hops.

  • Takeaways & Limitations

    Toddlers following YouTube recommendations from benign videos can encounter inappropriate content, and YouTube’s current mitigations remove only a minority of reviewed disturbing videos.

  • Takeaways & Limitations

    The study analyzes YouTube rather than YouTube Kids, and its collected videos are not representative of the entirety of YouTube.

Abstract

from arXiv · show

A large number of the most-subscribed YouTube channels target children of a very young age. Hundreds of toddler-oriented channels on YouTube feature inoffensive, well-produced, and educational videos. Unfortunately, inappropriate content that targets this demographic is also common. YouTube's algorithmic recommendation system regrettably suggests inappropriate content because some of it mimics or is derived from otherwise appropriate content. Considering the risk for early childhood development, and an increasing trend in toddler's consumption of YouTube media, this is a worrisome problem. In this work, we build a classifier able to discern inappropriate content that targets toddlers on YouTube with 84.3% accuracy, and leverage it to perform a first-of-its-kind, large-scale, quantitative characterization that reveals some of the risks of YouTube media consumption by young children. Our analysis reveals that YouTube is still plagued by such disturbing videos and its currently deployed counter-measures are ineffective in terms of detecting them in a timely manner. Alarmingly, using our classifier we show that young children are not only able, but likely to encounter disturbing videos when they randomly browse the platform starting from benign videos.

1 Introduction

YouTube hosts substantial toddler-oriented content, but inappropriate videos can mimic benign material, evade timely moderation, and appear during recommendation browsing. The paper characterizes this problem at scale and develops a classifier that detects inappropriate toddler-targeted videos with 84.3% accuracy.

  • Mitigation context: YouTube Kids and user-report-based manual review have not prevented disturbing videos from appearing, while manual inspection does not scale easily to YouTube’s volume.The paper identifies difficulty in detecting such videos as a reason they remain available, including in YouTube Kids.
  • Study scope: The study collects and manually reviews toddler-oriented, random, and popular videos, categorizing them as suitable, disturbing, restricted, or irrelevant.It defines toddlers as children aged 1 to 5 years and presents the first study focused on toddler-oriented disturbing content on YouTube.
  • Problem characterization: Disturbing videos often imitate benign toddler-oriented content, using innocent thumbnails and familiar characters to attract young viewers.The videos may contain mild violence or sexual connotations and can accumulate substantial view counts.
  • Classification: 84.3% accuracy: the binary classifier outperforms several baselines at distinguishing inappropriate from appropriate toddler-oriented videos.The authors collapse four labels into two categories for this analysis because the finer distinctions are difficult to classify reliably.
  • Large-scale analysis: 3.5% of simulated recommendation walks reached an inappropriate video within ten hops from a toddler-appropriate starting result.The starting video was among the top ten results for a toddler-appropriate keyword search such as Peppa Pig.
  • Moderation: Only 20.5% of manually reviewed disturbing videos and 2.5% of restricted videos had been removed by YouTube.These results indicate that current mitigation struggles to keep pace with the problem.

2 Methodology

The study combines broad YouTube crawling, manual annotation, and metadata analysis to characterize toddler-oriented disturbing videos. It examines how titles, thumbnails, categories, engagement statistics, and platform enforcement relate to distinguishing disturbing from suitable content.

  • Data Collection: 12,097 seed videos expanded to 844K recommended videos across three recommendation hops, providing broad coverage of toddler-oriented and related YouTube content.The collection used multiple crawling strategies and gathered titles, descriptions, thumbnails, tags, and video statistics.
  • Dataset Scope: The dataset covers multiple YouTube content subsets and categories, including Elsagate-related, other child-related, random, and popular videos.The collection design aimed to broaden coverage beyond Elsagate-related material and support classifier generalization across video types.
  • Manual Annotation Process: 4,797 videos received manual ground-truth labels after three annotators inspected their content and metadata using toddler-specific suitability categories.The labeling framework separates suitable, disturbing, restricted, and irrelevant videos, while majority agreement produced the final labels.
  • Ground Truth Dataset Analysis: 82.6% of videos containing “spiderman” and 80.4% containing “mous” in their titles were disturbing, showing that seemingly innocent cartoon terms can obscure content differences.Other high disturbing proportions included “peppa” at 78.6%, “superhero” at 76.7%, “pig” at 76.4%, “frozen” at 63.5%, and “elsa” at 62.5%.
  • Ground Truth Dataset Analysis: 47.4% of videos with spoofed thumbnails, 60.0% with violent thumbnails, and 34.8% with racy thumbnails were disturbing in the Elsagate-related subset.Adult and medical thumbnail content was more commonly associated with restricted videos, while thumbnail patterns differed for other child-related videos.
  • Ground Truth Dataset Analysis: None of the available metadata clearly identified disturbing videos, and enforcement remained limited: only 20.5% were removed and 6.9% of remaining disturbing videos were age-restricted.The authors report that this pattern was not explained by the videos being too recently uploaded for detection.

3 Detection of Disturbing Videos

The paper develops a multimodal deep learning classifier for toddler-oriented disturbing videos and evaluates it against baseline models for multi-class and binary classification.

  • 3 Detection of Disturbing Videos: The ground-truth dataset contains 4,797 videos, and the model is trained and tested using five-fold stratified cross-validation with oversampling for class imbalance.The model processes the four collected feature types during training and evaluation.
  • 3 Detection of Disturbing Videos: The classifier combines title, tags, thumbnail, and statistics/style branches before a fully connected network produces the final classification.Title and tags use recurrent text-processing branches, thumbnails use transfer learning, and statistics/style features use a dense network.
  • 3 Detection of Disturbing Videos: Thumbnail features are more important for classification performance than the other input feature types in the evaluated feature combinations.This comparison is reported for the proposed model across combinations of the four input types.
  • 3 Detection of Disturbing Videos: The binary classifier outperforms the strongest baseline, CNN-DDNN, by 12.3% in accuracy, 13.3% in precision, 16.6% in recall, 13.9% in F1 score, and 11.0% in AUC.The comparison covers all reported performance metrics.

4 Analysis

Using the binary classifier and recommendation-graph analyses, the paper measures inappropriate-video prevalence and simulates toddlers following YouTube recommendations from child-related searches.

  • 4.1 Recommendation Graph Analysis: 1.1% of Elsagate-related videos and 0.4% of other child-related videos are classified as inappropriate, indicating exposure risk at YouTube’s scale.The Elsagate-related subset contains 231K appropriate and 2.5K inappropriate videos, while the other child-related subset contains 99.5% appropriate videos.
  • 4.1 Recommendation Graph Analysis: A toddler randomly following one of the top ten recommendations from an Elsagate-related benign video has a 0.6% probability of reaching a disturbing or restricted video.The recommendation graph contains videos as nodes and recommendation links as directed edges.
  • 4.2 How likely is it for a toddler to come across inappropriate videos?: The analysis uses a snowball-sampled dataset that may not generalize to YouTube’s billions of videos and many-hop recommendation graph.The authors therefore conduct live random walks to assess the problem beyond the original dataset.
  • 4.2 How likely is it for a toddler to come across inappropriate videos?: The study performs 100 ten-hop random walks per seed keyword, selecting one video from the top ten search results and then one recommendation at each step.Each visited video is classified with the binary classifier.
  • 4.2 How likely is it for a toddler to come across inappropriate videos?: 3.5% of random walks seeded with sanitized Elsagate-related keywords encounter at least one inappropriate video, compared with 1.3% for other child-related keywords.Most inappropriate videos are encountered at the first hop, and the percentage decreases at later hops.
  • 4.2 How likely is it for a toddler to come across inappropriate videos?: 84.6% of the 338 detected inappropriate videos are disturbing, but the binary setup can misclassify some restricted or irrelevant videos because of category proximity.The authors explain that collapsing the four classes corrects some multi-class errors while retaining potential ambiguity.

5 Related Work

Prior work examined child-directed YouTube content and other malicious activity, whereas this paper focuses specifically on disturbing videos that explicitly target toddlers.

  • 5 Related Work: Existing studies addressed inappropriate content for children, parental controls, recommendation systems, spam, hate, extremism, and clickbait on YouTube.These lines of work cover related safety and platform-abuse problems but not the paper’s specific target.
  • 5 Related Work: The paper claims to be the first study to characterize and detect inappropriate videos explicitly targeting toddlers.It distinguishes disturbing toddler-targeted videos from broader inappropriate content and malicious activity.
  • 5 Related Work: The authors combine manual annotation, deep learning classification, and large-scale prevalence and recommendation analysis to study this problem.The classifier reaches 84.3% accuracy in the reported overall evaluation.

6 Conclusion and Discussion

The paper concludes that disturbing toddler-targeted videos remain a meaningful YouTube risk, while noting that its evidence and classifier are constrained by sampling, platform coverage, and training-data size.

  • 6 Conclusion and Discussion: The study reports 84.3% classifier accuracy, 1.05% inappropriate Elsagate-related videos, and a 3.5% chance of recommendation exposure within ten recommendations.These results combine automated detection with large-scale content and browsing analyses.
  • 6 Conclusion and Discussion: The authors identify crowd-sourced, uncurated content and engagement-oriented recommendation systems as a more pressing concern than screen time alone within their findings.They connect this concern to the risks of algorithmically amplified inappropriate content.
  • 6 Conclusion and Discussion: The dataset is not representative of all YouTube videos, although the authors argue that its seed keywords cover a wide range of child-related content.The study includes Elsagate-related, other child-related, random, and popular videos.
  • 6 Conclusion and Discussion: The study analyzes YouTube rather than YouTube Kids because YouTube does not provide an open API for collecting videos appearing in YouTube Kids.The authors also state that classifier performance is highly affected by the small training size.
Loading 1901.07046v3…