Source-linked AI summary
ChildSafeAds Shared Task 2026: Commercial Content in Child-Facing YouTube Videos
Thales Bertaglia, Catalina Goanta, Gerasimos Spanakis, Gunes Acar
TL;DR
ChildSafeAds addresses inconsistent disclosure and limited structured evidence about commercial content in child-facing YouTube videos. It introduces a SponsorBlock-based benchmark spanning three classification tasks and cumulative evidence levels, with model agreement ranging from 94.4% for ST1 labels to 44.4% for full ST3 flag sets.
Problem
Child-facing YouTube advertising is difficult for young audiences to recognise, while disclosure remains inconsistent and commercial-content evidence is limited.
Method
The paper constructs a 3,360-instance SponsorBlock-based benchmark with three classification tasks and cumulative evidence levels from transcripts to linked promotional pages.
Results
Model agreement was 94.4% for ST1 labels, 66.3% for exact ST2 label sets, and 44.4% for complete ST3 flag sets.
Takeaways & Limitations
The benchmark enables comparison of classification performance across evidence-access levels and their associated data-collection costs.
Takeaways & Limitations
The dataset requires a SponsorBlock mark and a crawlable promotional page, introducing selection bias and excluding non-commercial examples.
Abstract
from arXiv · showhide
ChildSafeAds is a shared task on commercial content in YouTube videos likely to reach children and teenagers. It contains 3,360 videos from 939 channels. Each instance begins with a segment submitted to SponsorBlock, an open-source crowdsourced browser extension whose users mark sponsor segments so that others can skip them. We pair the segment with its available transcript, video and channel information, and a sales or service page linked from the video description. Systems determine what kind of offer is being promoted (ST1), assign product categories (ST2), and identify legal risk flags (ST3). The evidence is divided into four cumulative access levels, from the transcript to the linked page, so results can be compared against the cost of collecting the data. 45.5\% of videos in our data failed to properly use the in-platform ad disclosure method (the ``Includes paid promotion'' label). GPT-5.4 produced the labels after the expert organiser team reviewed samples and iterated on the taxonomy, prompts and model choices. GPT-5.6-luna independently labelled the development set. This report describes the task, data and evaluation. An updated version will add participating systems and shared-task results.
1 Introduction
ChildSafeAds addresses the difficulty of detecting and classifying creator-delivered commercial content that children and teenagers may not recognise, using SponsorBlock candidates rather than platform disclosures alone. It defines three linked classification tasks and evaluates systems across cumulative evidence levels from transcripts to linked promotional pages.
- Motivation: 45.5% of videos failed to properly use YouTube’s “Includes paid promotion” label, so monitoring systems cannot rely on platform disclosure alone.SponsorBlock provides an independent crowdsourced signal for identifying candidate sponsor segments.
- Task design: ChildSafeAds asks systems to identify what is promoted, assign product categories, and flag aspects requiring closer legal review.These form three linked classification tasks.
- Task design: Systems may use transcripts alone or add video metadata, channel information, and linked promotional pages, with submissions recording their evidence level.This enables accuracy comparisons across the cost of collecting progressively richer evidence.
- Report scope: The report introduces a 3,360-instance dataset, its splits, three tasks, legal taxonomy, and evaluation setup.An updated version will describe submitted systems and final results.
2 Legal Context
The task’s legal labels are grounded in EU consumer and audiovisual advertising law, using the CRD for offer descriptions and the UCPD, AVMSD and DSA for legal-risk flags. ST3 covers disclosure, misleading content or exhortation, and product-related risks, while flags support review rather than establish infringement.
- Legal foundations: ST1 describes what the consumer receives using distinctions from the Consumer Rights Directive (CRD).
- Legal foundations: ST3 draws mainly on the Unfair Commercial Practices Directive (UCPD), Audiovisual Media Services Directive (AVMSD) and Digital Services Act (DSA).
- ST3 flag groups: ST3 flags cover disclosure recognition, misleading claims or omissions, direct exhortation, and age-restricted, prohibited, or unhealthy-food products.Disclosure flags assess whether audiences can recognise commercial content; content flags apply a narrow test to pressure on children to make purchases happen.
- Interpretation and limitations: Flags indicate cases for further review and are not findings of infringement.The released data maps each flag to relevant provisions and explains the connection through citations and notes rather than reproducing the law verbatim.
3 Related Work
Prior work studies disclosure, product type, child and adolescent advertising, sponsored-segment detection, and legal analysis, but does not benchmark joint classification from the same child-facing videos. CHILDSAFEADS builds on these strands by connecting video, platform, and promoted-product evidence to legally grounded labels.
- Disclosure and child-facing advertising: Prior studies examine influencer-advertising disclosure, children’s recognition of commercial persuasion, economic exploitation, advertising literacy, and food marketing.These studies establish why disclosure and product type matter for child-facing content.
- Benchmark gap: Existing work does not provide a benchmark for classifying disclosure and product type from the same videos.
- Computational detection: Computational studies detect sponsored YouTube segments, use model explanations for human labelling, and assess influencer-marketing compliance.CHILDSAFEADS instead uses SponsorBlock’s crowd-submitted segment boundaries for discovery and classifies the resulting cases.
- Legal analysis: Related legal-NLP work benchmarks legal-text analysis and connects legal rules to data-driven analysis, while LegalLens links detected violations to supporting legal provisions.CHILDSAFEADS connects evidence from a video, the platform, and the promoted product to legally grounded labels.
4 What is ChildSafeAds?
ChildSafeAds classifies likely commercial segments in child-facing YouTube videos using shared instances and cumulative evidence, rather than detecting commercial content. Its three tasks identify the offer, product category, and potential legal-risk flags.
- Dataset: Each instance pairs a crowd-submitted SponsorBlock sponsorship interval with its transcript, video and channel information, and one likely sales or service page.Commerciality and child-facing status are dataset-construction assumptions, not system predictions.
- Dataset: The benchmark evaluates classification after a likely commercial segment has been found and contains no non-commercial examples.Teams may enter any subset of the three tasks, which share the same instances and evidence.
- ST1: Offer type: ST1 assigns one of five labels describing what the consumer receives through the promoted transaction: physical goods, digital content or services, physical services, no identifiable offer, or other.
- ST2: Product category: ST2 assigns one or more of twelve product categories, including apps, hardware and electronics, food, fashion, health, education, financial products, gambling, toys, creator communities, and other products.The categories also include gambling-adjacent mechanics.
- ST3: Legal risk: ST3 assigns potential legal-review flags for misleading claims, inadequate or undisclosed advertising, direct exhortation, restricted or prohibited products, and HFSS foods.The standalone labels are no flag and insufficient context; systems predict flags, not severity or applicable legal provisions.
5 Dataset Construction
The dataset combines crowdsourced sponsorship intervals with transcript, video, channel, and linked promotional-page evidence for videos from screened child-facing channels. Its channel-disjoint, cumulative-access design supports controlled evaluation, while substantial label imbalance and collection constraints shape interpretation of results.
- Candidate selection: SponsorBlock sponsorship intervals with more upvotes than downvotes formed the initial pool, retaining the highest-voted eligible interval per video.These submissions were used to identify likely commercial segments, not to make legal determinations.
- Candidate selection: Channels passed keyword, public-information, and GPT-5.4 screening designed to remove clearly adult-oriented channels; systems receive this child-facing designation rather than predicting it.The released definition of “child-facing” is operational: a channel is child-facing if it passes the screening process.
- Evaluation design: The v1.0 split is channel-disjoint, preventing recurring channel presentation styles and disclosure habits from crossing boundaries, although brands and destination pages may recur.This setup evaluates generalisation to new channels rather than new advertisers.
- Released evidence: 3,360 videos from 939 channels include marked-segment transcripts, video metadata, channel information, and a usable promotional page; requiring such a page removed 1,116 of 4,985 candidates (22.4%).GPT-5.4 matched pages for 2,092 instances, while 1,208 used the first usable page as a fallback; 60 instances lack transcripts.
- Evaluation design: The four cumulative access levels progress from marked transcript to video metadata, channel information, and crawled promotional page, enabling comparison by data-collection cost.Macro-F1 weights labels equally, but product-risk flags each have fewer than 60 training examples; the rarest labels are ST1 other (2), ST2 gambling (12), and ST3 insufficient context (15).
6 Annotation and Reliability
The study used GPT-5.4 as its primary judge, with an iterative taxonomy and prompt refinement process guided by organiser review and legal expertise. GPT-5.6-luna independently labelled the development set to assess cross-model stability, not human inter-annotator agreement.
- Annotation procedure: Expert review iteratively revised the taxonomy, prompts, and model choices, narrowing direct exhortation and adding a new annotation pass.The organiser team included a legal expert, but sample review was not systematic double annotation.
- Annotation procedure: GPT-5.4 judged labels from task-specific evidence, including transcripts, descriptions, extracted brands, linked-page text, paid-promotion fields, and legal-provision summaries.The revised direct-exhortation pass used only the brand and transcript.
- Cross-model stability: 94.4% of cases received the same ST1 label from GPT-5.4 and GPT-5.6-luna.GPT-5.6-luna independently labelled the 504 development instances for this cross-model comparison.
- Cross-model stability: 66.3% of cases received exactly the same ST2 label set, with average Jaccard similarity of 0.74.ST2 is multi-label, so set overlap complements exact label-set agreement.
7 Evaluation Setup
The evaluation ran on CodaBench across three tasks, with submissions scored per task and tracked for coverage and ST3 risk families. The final export included 21 participating teams, while the supplied majority-class baseline scored 0.093 overall on development data.
- Competition and submissions: 21 teams submitted to the final export, excluding the organiser baseline, and teams could enter any subset of the three tasks.Each submission required one machine-readable prediction record per instance; missing predictions received zero credit.
- Competition and submissions: The development phase ran from 20 July to 10 August 2026, followed by evaluation from 11 to 18 August, closing at 00:00 UTC on 19 August.
- Scoring: Evaluation reported prediction coverage and a family-level ST3 score grouping the six flags into disclosure, content and product risks.Prediction coverage is the share of instances for which a system returned a prediction.
- Scoring: The supplied majority-class baseline scored 0.093 on the development set, with 0.151 for ST1, 0.042 for ST2 and 0.085 for ST3.
- Additional reporting: Teams reported their highest data-access level, method, estimated cost and legal-material retrieval, but legal grounding was not scored because the format could not reliably evaluate retrieved provisions.System reports instead describe legal-material use and the contribution of each additional data source.
8 Discussion, Limitations, and Ethics
The benchmark makes evidence-access cost part of evaluation, but its SponsorBlock- and crawl-based sampling, predominantly English text-based coverage, and limited legal context constrain interpretation. Ethical safeguards restrict use to research and human review rather than profiling or automatic enforcement.
- Discussion: Access levels encode a practical cost trade-off: transcripts are cheap to process, whereas promotional pages require finding, resolving and crawling.The shared-task results will test whether the additional evidence improves performance enough to justify that cost.
- Limitations: The dataset represents only videos with SponsorBlock sponsorship marks and crawlable promotional pages, introducing selection bias and excluding non-commercial examples for realistic discovery-rate estimation.Channel screening estimates likely teenage reach from public metadata rather than measuring actual audience.
- Limitations: 1,208 instances lacked alignment between the selected page and marked segment, while 60 instances had no transcript.Pages may also change after collection, adding uncertainty to the paired evidence.
- Limitations: The benchmark is predominantly English and text-based, omitting video frames, visual disclosure timing, non-verbal audio and transaction details relevant to legal assessment.These omissions are especially consequential for ST3, where exact cross-model agreement is 44.4%.
- Limitations: GPT-5.4 produced final labels after expert review and sample validation shaped the task definitions and prompts, with systems evaluated against the released taxonomy.No human inter-annotator agreement figure is available.
- Ethics: The release excludes viewer records and measured demographics, and its agreement prohibits redistribution, re-identification, contact, harassment and publications singling out creators.Worked examples redact identifiers, and the dataset is intended for research and human review, not profiling or automatic enforcement.