Source-linked AI summary

MigrationNarrate: A Dataset for Detection of Migration Narratives in YouTube Videos

Fatima Haouari, Carolina Scarton, Kalina Bontcheva

arXiv:2608.20984v1cs.CVcs.CLcs.CY

TL;DR

MigrationNarrate addresses the lack of dedicated annotated resources for migration narratives and the limited study of narratives in videos. It constructs and annotates a UK YouTube transcript dataset using a two-level taxonomy, then benchmarks encoder models and open- and closed-source LLMs. Fine-tuned open-source models can outperform closed-source models at the super-narrative level, whereas closed-source models remain stronger for fine-grained narrative detection.

  • Problem

    Migration narratives and narratives in videos remain understudied, partly because dedicated publicly annotated datasets are lacking.

  • Method

    The paper constructs MigrationNarrate from UK YouTube transcripts, applies a 12-super-narrative and 53-narrative taxonomy, and evaluates encoder models alongside open- and closed-source LLMs.

  • Results

    Fine-tuned open-source LLMs can outperform closed-source LLMs for super-narrative classification, while closed-source LLMs can still surpass them for fine-grained narrative detection.

  • Takeaways & Limitations

    The dataset provides a foundation for migration narrative detection in videos and supports future research through benchmarking and error analysis.

  • Takeaways & Limitations

    Manual channel and search-phrase selection may introduce selection bias and limit dataset representativeness.

Abstract

from arXiv · show

Narratives are central to how social communication is framed, making their detection critical for understanding and analysing public discourse. Prior work has explored narrative detection and extraction across diverse domains; however, migration narratives remain significantly understudied, primarily due to the absence of dedicated annotated datasets. Furthermore, public communication has recently shifted towards video-centric platforms, where narratives are conveyed through multimodal signals and consumed at scale. Despite this shift, narratives in videos remain largely unexplored. To bridge these gaps, we introduce MigrationNarrate, the first multimodal dataset for detection of migration narratives in the UK, consisting of 1,115 YouTube video transcripts annotated using a two-level taxonomy of 12 migration super-narratives and 53 narrative labels. This paper details the dataset design, collection, and annotations; together with benchmark results using a combination of pre-trained encoder models and both open- and closed-source Large Language Models. Finally, a thorough error analysis offers insights for future work.

1 Introduction

MigrationNarrate addresses the underrepresentation of migration narratives and video-based narrative detection by introducing an annotated UK YouTube transcript dataset.

  • Motivation: Video platforms shape information exposure and public discourse at scale, while potentially amplifying misleading or biased content.The supplied passage reports over 2.83 billion YouTube users worldwide as of March 2026 and more than 1.8 million hours of daily uploads.
  • Research gap: Prior narrative detection research has predominantly studied textual data from social media platforms and news articles.
  • Research gap: Video-transcript narrative detection remains largely understudied despite videos differing from written text in structure and delivery.
  • Research gap: Migration narratives have not previously been studied in this research context, although they can frame migration through humanitarian, solidarity, exclusion, policy, or security perspectives.
  • Contribution: MigrationNarrate provides 1,115 annotated YouTube video transcripts using 12 migration super-narratives and 53 narrative labels.The dataset targets UK-related migration narratives and is accompanied by model benchmarks and error analysis.

2 Related Work

Prior research has examined migration narratives and narrative detection across topics, platforms, countries, and media contexts, but no publicly annotated migration-narrative dataset was available.

  • Migration narratives: Migration narratives have been studied conceptually and empirically in political and media contexts, including their roles in policymaking and public discourse.
  • Migration narratives: Existing studies examine narrative distribution across media platforms and countries, with variation linked to context, event type, and media genre.
  • Migration narratives: Research has also explored factors associated with narrative influence, success, circulation, and framing across time and political parties.
  • Research gap: Prior migration-narrative studies did not release a publicly annotated dataset for the task.
  • Narrative detection: Narrative detection studies cover climate change, the Ukraine war, COVID-19, and elections, primarily using social media and news data.

3 The MigrationNarrate Dataset

MigrationNarrate combines a JRC-based taxonomy, targeted YouTube collection, automated filtering, model-assisted preannotation, and human adjudication to produce a multi-source annotated dataset.

  • Taxonomy: The adopted JRC taxonomy contains 12 super-narratives and 53 narratives covering humanitarian, solidarity, exclusion, policy, and security framings.
  • Taxonomy: GPT-5.1 generated concise definitions for individual narratives from super-narrative definitions, which authors manually reviewed and validated.
  • Collection: Videos were collected through YouTube channels, playlists, and search-based retrieval spanning institutional, political, media, commentator, activist, and independent sources.
  • Filtering and sampling: A multi-step filtering pipeline used metadata constraints, Whisper-based transcripts, embedding similarity, and GPT-5.1 preannotation to reduce annotation costs and select balanced samples.
  • Annotation: Two annotators independently labeled each batch, with an adjudicator resolving disagreements; Cohen’s kappa was 0.45 for super-narratives and 0.30 for narratives before resolution.
  • Dataset statistics: The final dataset contains 1,115 human-annotated videos from 684 YouTube channels, alongside a larger filtered corpus supporting weak, semi-supervised, and future expanded annotation.
  • Dataset statistics: The most prevalent super-narrative has 190 videos, while the most frequent individual narrative has 65, and the class distribution is imbalanced.

4 Experiments and Evaluation

The experiments benchmark encoder models and open- and closed-source LLMs for hierarchical migration-narrative classification, comparing prompting with supervised adaptation. Results show that fine-tuning generally improves LLM performance, while model advantages differ between super-narrative and fine-grained classification.

  • Experimental setup: The benchmark compares encoder models with LLMs, prompting with supervised fine-tuning, and open-source with closed-source systems.RoBERTa is fine-tuned for super-narrative and narrative classification; LLMs are evaluated under zero-shot, few-shot, and, for open-source models, fine-tuning setups.
  • Experimental setup: Hierarchical prompting first identifies a super-narrative and then its corresponding narrative, mirroring the annotation procedure and reducing confusion among labels.Zero-shot and few-shot prompts include label definitions, while few-shot prompts additionally provide one example per class.
  • Experimental setup: The evaluation uses stratified 70% training, 10% development, and 20% testing splits, with four extremely low-frequency narratives excluded from training.The corresponding split contains 774 training, 117 development, and 224 test videos; the rare classes remain evaluation-only.
  • Results and discussion: For super-narrative classification, fine-tuned Llama and Gemma surpassed RoBERTa with Macro-F1 scores of 0.525 and 0.502, respectively, while GPT-5.4 achieved 0.480 versus RoBERTa’s 0.446.Fine-tuned Qwen remained below the encoder baseline, showing that adaptation did not improve every open-source model equally.
  • Results and discussion: Few-shot prompting usually improved over zero-shot prompting, but fine-tuning all open-source LLMs produced substantially larger gains than prompting alone.Llama was the exception, with a slight decline after examples were introduced, suggesting sensitivity to prompt formulation or example selection.
  • Results and discussion: For fine-grained narrative classification, closed-source GPT-4o and GPT-5.4 remained strongest, while fine-tuned open-source models consistently outperformed RoBERTa but did not surpass the closed-source systems.GPT-4o and GPT-5.4 reached Macro-F1 scores of 0.350 and 0.324, compared with 0.304 for Llama and 0.282 for Gemma.

5 Error Analysis

The error analysis identifies five recurring error types, showing that models struggle with surface–narrative misalignment, fine-grained distinctions, overlapping interpretations, counter-narratives, and weak narrative signals.

  • Models often prioritise salient lexical cues over the dominant narrative framing in surface–narrative misalignment cases.Explicit references to record migrant arrivals led models toward migration-volume or border-strain labels, although the dominant framing concerned a failing policy.
  • Fine-grained label confusion occurs when closely related narratives differ through subtle framing distinctions.Models confused financial-burden narratives with claims that immigrants receive better benefits when both themes appeared in the same video.
  • Multi-narrative overlap reflects ambiguity because multiple interpretations can be plausible for one video.
  • Models struggle to detect counter-narratives because argumentative or conversational contexts require capturing stance and negation.
  • Weak narrative signals lead models to over-predict labels even when the gold label is None.In one example, both models assigned a false-hope narrative despite the absence of a migration narrative label.
  • The findings point to surface-cue reliance and implicit framing as priorities for improving narrative detection, alongside clearer label boundaries and multi-label annotation.

6 Conclusion

The paper presents MigrationNarrate as a dataset and benchmark for migration-narrative detection in video transcripts, combining model evaluations with error analysis. Its results indicate different strengths across model types and narrative-granularity levels.

  • MigrationNarrate is presented as the first dataset for detecting migration narratives in video transcripts.The dataset is intended to fill a research gap and provide a foundation for migration-narrative detection tools.
  • The benchmarks combine pre-trained encoders with open- and closed-source large language models.The paper also includes a thorough error analysis to inform future research directions.
  • Fine-tuned open-source LLMs can outperform closed-source LLMs in zero-shot and few-shot super-narrative detection, while closed-source LLMs can surpass them for fine-grained narratives.

Limitations

The dataset’s representativeness and label reliability are constrained by manual content selection and difficult single-label annotation, while the analysis uses transcripts without other video modalities.

  • Manual channel, playlist, and search-phrase selection may bias the dataset toward expected narrative frames and selected viewpoints.The curation limits representativeness by filtering content through taxonomy-driven queries and selected sources.
  • Annotation by two annotators, with adjudication, produced relatively low agreement because narratives overlap and annotators assign one dominant label.These conditions create ambiguity and potential label noise.
  • The approach uses textual transcripts and excludes visual, auditory, and other multimodal signals.Consequently, the analysis may capture only a partial view of the underlying narratives.

Ethics Statement

The study reports ethics approval, informed consent and compensation procedures, platform-compliant data handling, and privacy safeguards for sensitive public content.

  • The research received approval from the university Research Ethics Committee.
  • Annotators were informed about potential risks, reviewed consent materials, and could opt out at any point.They received an hourly rate equivalent to Graduate Teaching Assistant work under university guidelines.
  • The dataset follows YouTube’s Terms of Service by releasing video identifiers and annotations rather than audio, video, or transcripts.The dataset is intended for research use and will be released under a CC-BY-NC-SA 4.0 license.
  • The study restricts itself to public content without private or sensitive personal information and frames annotations as categories rather than judgments about speakers.The authors acknowledge potential reputational harm from sensitive narratives.

A Narrative Definitions

The paper defines migration narratives as concise, non-overlapping claims organized under broader super-narratives. Definitions cover cultural, security, demographic, economic, welfare, housing, and border-related framings.

  • Definition format: Narrative definitions are formatted as super-narrative/narrative pairs to identify each label’s place in the taxonomy.The definitions were initially generated with GPT-5.1 and then manually reviewed.
  • Definition guidelines: The guidelines require definitions to fit logically under their super-narrative while remaining conceptually distinct and nonjudgmental.They describe what each narrative claims without expressing personal, ethical, or political judgments.
  • Cultural framings: Cultural narratives frame immigration as threatening European identity, social cohesion, or integration.Examples include claims that immigration weakens established customs or that some immigrants cannot meaningfully integrate.
  • Security framings: Security narratives portray immigration as increasing risks involving individual safety, national security, crime, sexual harm, terrorism, or disease.These definitions describe claims about vulnerability, criminality, extremist violence, and public-health threats.
  • Economic and welfare framings: Economic and welfare narratives claim that immigration strains public resources, housing, employment, welfare, healthcare, or border management.They include claims about excessive arrivals, insufficient controls, job competition, welfare abuse, housing pressure, and preferential treatment for natives.
  • Economic and welfare framings: Other definitions portray asylum seekers as economic migrants and promote prioritizing native-born citizens in access to resources and support.These labels frame asylum claims as financially motivated and advocate a natives-first allocation of public support.

B Videos Collection Source

Videos were collected through curated YouTube channels and playlists, supplemented by general and narrative-oriented search phrases. The workflow also used semantic filtering and two-step LLM classification.

  • Curated sources: The collection included YouTube channels and playlists selected for migration-related content.The supplied passages identify a channel-and-playlist source table but do not provide its full entries.
  • Search phrases: Search-based retrieval used general phrases such as immigration, asylum, refugees, Channel crossings, small boats, and border control in the UK.These queries targeted broad UK migration-related content.
  • Search phrases: Additional queries addressed illegal immigration, people smuggling, deportation schemes, and UK immigration policy and law.These phrases broadened retrieval toward specific policy and enforcement topics.
  • Filtering and classification: The LLM classifier first selected a migration super-narrative and then classified the transcript into the most relevant narrative within it.Each step returned None when no applicable label was identified.

G MigrationNarrate Statistics

MigrationNarrate is a human-annotated collection of UK migration-related YouTube videos, with dataset statistics and example records reported in the appendices.

  • Dataset examples: Examples from the MigrationNarrate dataset are presented in Table 9.

I Hyperparameters Fine-tuning

The experiments fine-tuned RoBERTa-Large and LLMs with specified training, optimization, and regularization settings, while the appendices document data-collection and annotation materials.

  • Encoder fine-tuning: RoBERTa-Large was fine-tuned for five epochs with sequence length 512 and batch size 16.Learning rates and dropout values were selected from predefined grids.
  • LLM fine-tuning: LLMs were trained with QLoRA adapters applied to all linear layers using rank r = 16 and α = 32.LLM training used three epochs, batch size 1, and searched learning rates and dropout values.
  • Model selection: The paper reports hyperparameter settings selected according to development-set Macro-F1 score.These settings are listed in Table 8.
  • Annotation materials: The annotation interface presented transcripts with direct video links and revealed narratives after selecting a super-narrative.Annotators could hover over labels to view definitions.
  • Data-collection materials: The appendices include search phrases covering UK migration, borders, crime, jobs, welfare, housing, political messaging, and cultural identity.These phrases span both migration-related topics and narrative framings.
Loading 2608.20984v1…