Source-linked AI summary
MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation
Nadav Borenstein, Greta Warren, Desmond Elliott, Isabelle Augenstein
TL;DR
Multimodal misinformation is difficult to understand and annotate at scale because existing taxonomies are incomplete and current VLMs have limitations. The paper constructs a multilingual real-world dataset, develops and operationalizes a comprehensive taxonomy with a VLM pipeline, and validates it across two datasets. The resulting analysis identifies mechanism variation across topics and languages, including especially prevalent AI-generated images in technology and science and news images used for vaccination-related credibility.
Problem
Existing multimodal misinformation taxonomies lack comprehensive grounding in real-world data, while annotation remains difficult and resource-intensive at scale.
Method
The paper builds a seven-language X dataset, develops a theory- and qualitatively grounded taxonomy, and applies it through an automated VLM annotation pipeline validated with human annotators.
Results
The approach shows good agreement with expert annotators across Community Notes and AMMEBA and reveals topic-linked differences in multimodal misinformation mechanisms.
Takeaways & Limitations
The findings provide a foundation for more targeted analysis, detection, and mitigation of multimodal misinformation.
Takeaways & Limitations
The taxonomy excludes some prior axes, and the datasets do not support robust analysis of every narrative because many have too few samples.
Abstract
from arXiv · showhide
Multimodal misinformation on social media is highly prevalent, potent, and harmful, yet difficult to detect and counter, and still poorly understood compared to its text-only counterpart. Research on the properties and deceptive strategies of multimodal misinformation is hindered by a lack of taxonomies grounded in real-world contexts and by the limitations of current multimodal machine learning models, which prevent the automation of annotation and analysis at scale. We address these shortcomings in three steps. First, we collect a large-scale, high-quality dataset of real-world misinformation instances from Twitter/X in seven languages. Second, we develop a novel, comprehensive taxonomy of multimodal misinformation grounded in an in-depth qualitative analysis of the data and prior theoretical work. Finally, we operationalise the taxonomy through an automated multi-step annotation pipeline using a Vision-Language Model (VLM), and perform human-validation. Our novel approach leads to previously undocumented insights about how social media users combine images with text to spread misinformation in the wild, e.g., that AI-generated content is particularly prevalent in technology and science, while vaccination misinformation disproportionately utilises images from news outlets to assert credibility. Our method and findings provide guidance for targeted approaches for detecting multimodal misinformation, and suggest that mitigation efforts should be developed and applied strategically rather than uniformly.
1 Introduction
Multimodal misinformation is consequential yet difficult to classify and analyze at scale. The study addresses these gaps with a comprehensive taxonomy, a VLM annotation pipeline, and validation across real-world datasets.
- Multimodal misinformation can be more convincing and emotionally salient than text alone, increasing concern across predominantly multimodal social platforms.
- Existing taxonomies are disparate, ambiguous, and unable to classify many real-world multimodal misinformation cases.
- Reliable large-scale analysis remains difficult because extensive taxonomies challenge crowdsourced annotators and cases often require domain-specific expertise.
- The study develops a comprehensive taxonomy with five additional orthogonal axes, grounded in prior theory and qualitative analysis of multilingual Community Notes data.
- A VLM-based pipeline applies the taxonomy at scale, while evaluation on Community Notes and AMMEBA shows good agreement with expert human annotators.
- AI-generated content is particularly prevalent in technology and science, whereas vaccination misinformation disproportionately uses news-outlet images to assert credibility.
2 Background
Prior research offers fragmented multimodal misinformation taxonomies and task-specific datasets, while VLM applications remain constrained by multimodal reasoning and instruction-following limitations.
- Few studies explicitly define taxonomies of multimodal misinformation mechanisms, and existing schemes vary in categories and scope.
- The AMMEBA taxonomy organizes misinformation across orthogonal axes including image type, media type, and manipulation mechanism.
- Benchmark datasets commonly target specific tasks such as claim detection, out-of-context prediction, image manipulation detection, or deepfake detection.
- VLMs have been applied to several multimodal misinformation tasks, including claim detection, out-of-context identification, manipulation detection, and justification generation.
- Limitations in multimodal reasoning and instruction following have constrained the utility of VLMs for these tasks.
3 Datasets
The study builds a multilingual, image-containing misinformation dataset from X Community Notes and supplements it with an AMMEBA sample for external generalization assessment.
- The dataset development process uses real instances of online misinformation to guide taxonomy development and enable empirical validation at scale.
- X Community Notes was selected for its quality, topical and linguistic diversity, contemporary coverage, and consistent username–claim–image–note format.
- The adapted Community Notes collection retains helpful notes in English, Spanish, Portuguese, Japanese, French, German, or Hebrew and keeps posts containing images.
- The resulting dataset contains 24,596 samples before additional examples are collected from the official Community Notes release.
- Recovered posts from 25k sampled instances are filtered to retain image-containing posts, producing a final merged dataset.
- A sample from AMMEBA is used to assess generalization beyond Community Notes, restricted to cases where the image actively contributes to the misinformation.
4 Taxonomy Creation
The taxonomy unifies fragmented prior schemes through iterative alignment, refinement, and empirical annotation, resulting in six orthogonal axes and 25 mechanisms. It adds categories for previously uncovered cases while omitting axes that cannot be reliably annotated from content alone or fall outside the study’s scope.
- Existing multimodal misinformation taxonomies differ in naming, coverage, granularity, and hierarchical structure, making them poorly aligned.
- The authors unify prior taxonomies by aligning exact matches, consolidating near-matches, and organising remaining categories hierarchically.
- Taxonomy development iteratively used dataset annotation to refine vague definitions and merge theoretically distinct categories that were indistinguishable in practice.
- The final taxonomy contains six axes and 25 mechanisms, including newly added categories for scientific errors, conspiracies, exaggeration, and textual claims embedded in images.
- The orthogonal axes cover classification, image type, emotion, rhetorical role, topic and message, and multimodal mechanism, enabling narrative–mechanism analysis beyond frequency counts.
- Intent and recipient axes are omitted because intent cannot be reliably annotated from content alone and recipient falls outside the study’s scope.
5 VLM Annotations
The authors automate taxonomy annotation with a multi-stage VLM pipeline that processes specific axes and evaluates predictions against human annotations. The pipeline achieves generally good agreement, while performance remains lower on ambiguous axes and varies substantially across mechanism labels.
- VLM automation enables rapid, low-cost annotation of large datasets where expert annotation is slow and crowd annotation is difficult.
- The pipeline predicts taxonomy axes in stages, using image-only input for image type and post-based inputs for emotion, classification, topic, rhetorical role, and mechanism.
- The mechanism stage first determines whether the image actively contributes to deception, assigning Decorative when it does not, before predicting mechanism categories.
- Three pipeline runs showed that subjective and ambiguous axes had lower inter-run agreement, broadly matching the human–human agreement pattern.
- 79.3% of non-distractor mechanism predictions were judged correct, with Manipulated image–removal and Unreliable source–satire above 90% and two other sub-mechanisms below 60%.
- On AMMEBA’s seven aligned categories, F1 scores ranged from 0.61 to 0.83, except for Time mismatch at 0.26.
6 Analysis
The analysis examines how deceptive mechanisms vary across sub-mechanisms, topics, languages, and datasets, revealing recurring and context-specific patterns. Fine-grained narrative analysis further links particular image types and emotions to misinformation narratives.
- General Results: Slanted is the most common Community Notes mechanism, exceeding 7,000 samples, while Unreliable source, Other mechanism, and Deny authenticity total fewer than 800.Fake image occurs as frequently as Manipulated image.
- General Results: AI-generated and forged images occur at comparable rates, and both exceed every Image manipulation subtype except textual manipulation.This finer-grained comparison comes from the sub-mechanism distribution across the datasets.
- Language Differences: Decorative is over-represented in Portuguese, Slanted is disproportionately used in Japanese, and Japanese Slanted posts often involve pseudoscientific earthquake-prediction claims.Portuguese patterns are associated with Sports, Celebrities, and Entertainment, while Japanese patterns are attributed to scientific errors and conspiracies.
- Topic and Mechanism Interactions: Mismatch, especially Place, Time, and Event mismatch, co-occurs strongly with Conflict, while Textual claim is over-represented in Law, Economics, and Health.The analysis uses Fisher’s exact tests with Benjamini–Hochberg FDR correction and omits cells with p ≥0.05.
- Cross-Dataset Comparison: Mismatch–Conflict and the under-representation of Celebrities and Textual claim persist across Community Notes and AMMEBA, whereas Fake image diverges between datasets.AMMEBA is dominated by Mismatch and contains considerably less Decorative content; AI-generated images are also rarer there.
- Narrative Case Studies: News screenshot image types are particularly over-represented in the athletic-rivalry narrative, suggesting that journalistic authority is appropriated as a deception strategy.The analysis also finds reliance on negative emotions such as Anger and Fear in anti-immigration narratives.
7 Conclusion
The paper introduces a unified, real-world-grounded taxonomy of multimodal misinformation mechanisms and operationalises it through an automated VLM annotation pipeline. Validation across Community Notes and AMMEBA supports its use for large-scale analysis, which reveals variation in deceptive mechanisms and informs targeted detection and mitigation.
- Conclusion: The unified taxonomy consolidates prior frameworks, adds novel categories grounded in real-world data, and is operationalised through an automated VLM annotation pipeline.The taxonomy is validated across the Community Notes and AMMEBA datasets.
- Conclusion: Applying the taxonomy at scale reveals that AI-generated images are especially prevalent in technology and science, while vaccination misinformation disproportionately uses news-organisation screenshots to assert credibility.These findings provide a foundation for more targeted analysis, detection, and mitigation of multimodal misinformation.
Limitations
The authors identify limitations involving taxonomy coverage, dataset size, model selection, multilingual processing, label structure, and data quality.
- Several prior taxonomy axes were excluded because they were ambiguous, difficult to annotate reliably, dependent on external sources, or outside the study’s scope.The authors identify incorporating these axes as an open opportunity.
- The Community Notes dataset cannot support robust analysis of every narrative because many narratives have too few samples.The AMMEBA data used is also only a subsample.
- Computational constraints and avoidance of proprietary systems limited model selection, and Qwen3.5 27B does not match state-of-the-art VLM performance.The authors state that a larger or task-fine-tuned model would likely yield better predictions.
- The multilingual approach relies on Qwen3.5’s built-in support rather than explicitly modeling multilinguality.The authors suggest translation pipelines or language-specific VLMs may perform better.
- The Multimodal Mechanism axis uses single-label classification even though some misinformation instances exhibit multiple mechanism types.The authors propose multi-label assignment as a natural direction for future work.
- Manual inspection found some low-quality instances in both datasets, motivating better filtering independent of existing quality annotations.
B.3 Analysis
The analysis compares vaccination-related misinformation with the overall distribution and uses automated prompts to classify several multimodal properties.
- The vaccination analysis reports divergence from the overall distribution across all misinformation instances.
- The annotation prompts classify image type, evoked emotions, misinformation status, topic and message, and rhetorical role.The topic-and-message prompt also extracts abstracted message versions for high-level analysis.
D Reproducibility
The reproducibility materials describe dataset construction, narrative clustering, annotation agreement, and language-model summarisation steps.
- Dataset construction: The AMMEBA data are constructed by joining released files on image_id and fact_check_url, filtering disqualified or incomplete rows, and retrieving images and fact-check metadata.The pipeline uses Google’s Fact Check Tools API and downloads full fact-check articles from publishers’ websites.
- Narrative analysis: Messages are embedded as 768-dimensional vectors with a multilingual sentence-transformer, optionally using abstract paraphrases before dimensionality reduction and clustering.
- Validation: Table 3 reports agreement across annotation settings by step, including matches, label counts, and skip conditions.Classification, Mechanism, and Sub-mechanism are single-label tasks, so Jaccard and at-least-one-match metrics are omitted.
- Narrative analysis: Posts associated with multiple topics are expanded into independent post-topic observations, then KMeans creates size-proportional sub-clusters of approximately 50 members.
- Narrative analysis: For each sub-cluster, an LLM receives the 20 messages nearest the centroid and produces a short narrative summary cached in a structured JSON mapping.Researchers manually scan these summaries to identify narratives of interest.
E Full Taxonomy
The full taxonomy materials combine dataset and analysis figures with multilabel image categories and mechanism labels for multimodal misinformation.
- Dataset and cross-cutting analyses: The section includes figures showing dataset timing, mechanism-topic co-occurrences, topic-language distributions, and language-related mechanism patterns.
- Narrative analyses: Additional figures and tables examine athletic rivalry, vaccination, immigration, and conflict-related misinformation, including mechanisms, image types, emotions, and examples.
- Image type: The image taxonomy supports multilabel classification across photographs, screenshots, documents, visualisations, digital art, memes, annotated images, captioned images, collages, and complex composites.A meme that is also a social-media screenshot can receive both labels.
- Image type: The taxonomy distinguishes additional image categories including news screenshots, official documents, other text images, data visualisations, graphic designs, and composite formats.The instructions require assigning all applicable categories and selecting composition labels when embedded photographs appear within larger designs.
- Annotation instructions: The image-classification instructions distinguish simple photographs from larger compositions and require multilabel assignment when multiple categories apply.CAPTIONED_IMAGE is reserved for a single image paired with one text element, whereas additional elements indicate COMPLEX_COMPOSITE.
- Multimodal mechanism: Mechanism labels include decorative, fake image, manipulated image, unreliable source, mismatch, and slanted representation, with finer-grained subcategories.Mismatch covers identity, time, place, event, and other discrepancies; slanted covers misleading interpretation such as exaggeration.