Source-linked AI summary
RCMN: Understanding Misleadingness in Influential Public Discourse
Peiling Yi
TL;DR
Research on misleadingness has focused less on how framing, omission, context, and communication shape readers’ interpretations beyond claim-level factuality. This paper introduces RCMN, an evidence-grounded reader-centric taxonomy, dataset, and benchmark, finding that interpretations are often recoverable from lightweight representations while precise misleading mechanisms require richer grounding.
Problem
Existing research has less comprehensively addressed how misleadingness arises and shapes reader interpretations beyond individual claim veracity.
Method
RCMN operationalises misleadingness across five dimensions and constructs an evidence-grounded dataset and benchmark for influential public discourse.
Results
Likely reader interpretations and broad affective and communicative cues can often be recovered from lightweight claim-and-context representations, whereas precise misleading mechanisms are considerably less reliable.
Takeaways & Limitations
Lightweight representations may support scalable initial misleadingness analysis, while deeper mechanism identification remains dependent on richer contextual, evidential, or multimodal information.
Takeaways & Limitations
The dataset contains few non-misleading controls, limiting evaluation of whether models distinguish misleading communication from ordinary evidence-compatible communication.
Abstract
from arXiv · showhide
Influential public discourse shapes public beliefs and can also mislead, not only through what is stated, but also through how information is framed, omitted, contextualised, and communicated. Yet less research has focused on how such misleadingness arises and shapes the interpretations formed by readers. To address this gap, we introduce Reader-Centric Misleadingness Understanding (RCMN), a framework that operationalises misleadingness through five dimensions: misleading mechanism, likely reader interpretation, evidence-warranted interpretation, emotional arousal, and communicative intent. Based on this framework, we construct an evidence-grounded dataset of influential public discourse. Empirical findings show that misleadingness is diverse and extends well beyond fabrication, with unsupported inference, exaggeration, and omission among the prevalent mechanisms, and is frequently associated with heightened emotional arousal and distortive communicative intent. Moreover, we investigate whether lightweight claim-and-context representations retain sufficient cues for understanding reader-centric misleadingness without access to richer contextual, evidential, and multimodal information. Evaluation across five recent generative foundation models shows that reader-level interpretations can often be recovered from such limited representations, whereas identifying how misleadingness is produced remains considerably more challenging. These findings highlight the potential of lightweight representations for scalable misleadingness analysis, while reliable understanding of misleading mechanisms continues to require richer contextual and evidential grounding.
1 Introduction
Influential public discourse shapes interpretations of public issues and can mislead through framing, omission, contextualisation, and communicative intent rather than factual inaccuracy alone. RCMN addresses this gap with a reader-centric, evidence-grounded framework, dataset, and benchmark.
- Influential public discourse shapes how audiences understand political, social, economic, and other public-interest issues.
- Factually accurate information can still mislead when selective presentation or omitted context encourages an unsupported policy interpretation.
- Existing computational approaches predominantly verify individual claims, overlooking misleading effects produced by framing, combination, contextualisation, and presentation.
- RCMN characterises misleadingness through misleading mechanisms, likely reader interpretations, Evidence-Warranted Interpretation, emotional arousal, and communicative intent.
- The framework constructs an evidence-grounded dataset from publicly circulated messages linked to fact-checking attention, public figures, organisations, and major political or social events.
- Findings show misleadingness extends beyond fabrication, is often associated with heightened emotion and distortive intent, and is harder to identify by mechanism than by reader interpretation.
2 Related work
Related work progresses from claim-level factual verification toward contextual, multimodal, intent-aware, and reader-centric misleadingness analysis. However, existing datasets and methods remain partial or fragmented representations of reader-level misleadingness.
- Datasets and Task Evolution: Early benchmarks operationalised misinformation primarily through veracity or factuality prediction, including fine-grained labels and evidence-based claim verification.
- Datasets and Task Evolution: Out-of-context misinformation datasets assess authentic media reused with mismatched context, including miscaptioned image–text pairs and reconstructed original context.
- Datasets and Task Evolution: Recent datasets incorporate claimant intent, image contextualisation, location verification, verdict prediction, and omission-sensitive comparisons between preview- and article-supported interpretations.
- Datasets and Task Evolution: These approaches still capture only part of the broader phenomenon of reader-level misleadingness.
- Methods: Reader-centric research models interactions among content, context, communicative intent, emotion, and reader interpretation, drawing on human-centred and emotion-aware approaches.
- Methods: Existing methodological strands remain fragmented rather than unified within a single reader-centric misleadingness framework.
3 RCMN taxonomy
RCMN defines misleadingness through divergence between a message’s encouraged interpretation and the interpretation warranted by evidence and context. Its taxonomy operationalises this divergence across mechanisms, interpretations, emotional arousal, and communicative intent.
- Misleadingness is the extent to which a message’s encouraged interpretation diverges from the interpretation warranted by relevant evidence and context, regardless of literal truth.
- The taxonomy includes mechanisms such as fabrication, miscontextualisation, omission, misattribution, exaggeration, unsupported inference, and not misleading.
- Unsupported inference covers causal, evaluative, predictive, or generalised conclusions insufficiently supported even when underlying statements are factually accurate.
- Interpretations: Likely Reader Interpretation captures the inference a message encourages, while Evidence-Warranted Interpretation captures the conclusion justified by fact-checking and contextual evidence.
- Emotional arousal: Emotional arousal measures the intensity of the response content is likely to evoke, ranging from low and moderate to high alerting or emotionally charged presentation.
- Communicative intent: Communicative intent distinguishes informative, persuasive, and distortive goals, with distortive communication steering readers toward inadequately warranted interpretations.
4 RCMN Datasets
RCMN constructs an evidence-grounded, reader-centric dataset from fact-checking records by recovering original communications, context, provenance, and evidence before annotation. The resulting dataset separates factual veracity from misleadingness and supports analysis of mechanisms, communicative signals, and reader interpretations.
- 4.1 Data Source: RCMN uses fact-checking records as candidate entry points for reconstructing potentially misleading public communication.The source is used for candidate identification and evidence recovery rather than as the target-label source.
- 4.1 Data Source: The source records span 1970–2025 and contain 260,863 entries across 57 fields, but exhibit substantial heterogeneity, multilinguality, and sparsity.English accounts for 7% of records, while the data include 1,007 fact-checking organisations and 24,869 claim-rating labels.
- 4.2 Data construct: Fact-check articles serve as recovery gateways for reconstructing source messages and retrieving contextual and evidential information.Original posts, articles, images, videos, advertisements, or archived copies are recovered when available, with reconstruction details recorded when they are not.
- 4.3 Annotation: RCMN annotation follows evidence grounding, veracity–misleadingness separation, multi-level annotation, and auditability principles.Annotations and explanations must be supported by explicit evidence and remain open to human verification and adjudication rather than deriving directly from fact-check verdicts.
- 4.3 Annotation: 93.6% of 2,216 instances had sufficient evidence, and 95.5% contained direct two-sided evidence for comparing annotations with evidence-supported context.Repeated review covered 34.9% of instances with a second check and 14.6% with a third check.
- 4.4 Empirical Findings: The dataset shows diverse misleading mechanisms and strong associations between misleadingness, emotional arousal, and distortive communicative intent.Misleading cases diverge more from evidence-warranted interpretations, while high arousal occurs in 58.3% of mechanism-identified misleading messages versus 12.7% of non-misleading messages.
5 Benchmarks
The benchmark tests whether lightweight claim-and-context representations preserve cues for reader-centric misleadingness, without the richer evidence and multimodal context used for annotation. Models often recover interpretations and affective or communicative cues, but identifying the precise misleading mechanism remains harder.
- 5.1 Task Definition: The benchmark estimates reader-centric annotations from limited claim-and-context inputs rather than the complete evidence used to establish gold labels.The target labels include misleading mechanism, likely reader interpretation, emotional arousal, and communicative intent.
- 5.4 Results: 84–97% of generated interpretations are fully semantically equivalent despite modest ROUGE-L scores of 0.228–0.387.Claude Fable 5 combines the lowest ROUGE-L, 0.228, with the highest full-equivalence rate, 97%, and semantic adequacy of 0.98.
- 5.4 Results: GPT-5.6 Sol achieves Macro-F1 = 0.643 for emotional arousal and Macro-F1 = 0.607 for communicative intent.Its strongest class-level results include high arousal F1 = 0.871, persuasive intent F1 = 0.966, and distortive intent F1 = 0.959.
- 5.4 Results: Claude Fable 5 achieves the strongest misleading-mechanism classification, with Macro-F1 = 0.520.The not misleading class remains difficult for most models, with F1 scores ranging from 0.031 to 0.184, while Claude Fable 5 reaches 0.393.
- 5.4 Results: Unsupported-inference recognition is intermediate: models can detect stronger-than-supported conclusions, but reliability depends on knowing what the evidence warrants.This makes mechanism recoverability dependent on contextual and evidential grounding.
- 5.5 Discussion: Claim-and-context information preserves substantial but incomplete cues, while precise mechanism identification is less reliable when misleadingness depends on omission, displaced context, or external evidence.Lightweight representations therefore support preliminary analysis but do not replace full contextual and multimodal reasoning.
6 Limitations
The study’s limitations concern dataset coverage, interpretive subjectivity, and incomplete recovery of original multimodal communications. These constraints affect how broadly benchmark results and annotations should be interpreted.
- 6 Limitations: Only a small proportion of instances are labelled not misleading because the dataset is biased toward disputed or potentially misleading claims.This imbalance limits evaluation of distinguishing misleading communication from ordinary, evidence-compatible communication and may contribute to low not misleading F1 scores.
- 6 Limitations: Reader-centric labels remain partly interpretive despite AI-assisted annotation, human verification, and adjudication.The labels are evidence-grounded reference annotations rather than deterministic representations of every possible reader response.
- 6 Limitations: The original post or other media is not recoverable for every instance, so some communications are reconstructed from fact-check articles and provenance information.The original source is used as supplementary reference when available.
7 Conclusion
RCMN introduces a reader-centric framework, evidence-grounded dataset, and benchmark for analysing misleadingness beyond claim-level factuality. Its findings distinguish observable communicative cues from deeper evidence-dependent mechanisms, motivating adaptive analysis that retrieves richer information when needed.
- RCMN provides a framework, dataset, and benchmark for understanding misleadingness in influential public discourse beyond claim-level factuality.
- Table 5 evaluates per-class and Macro-F1 performance across reader-centric misleadingness dimensions on 2216 shared instances.
- Table 6 evaluates generated likely reader interpretations using ROUGE-L and semantic-equivalence measures.
- The framework captures how messages become misleading, the interpretations they encourage relative to evidence, and affective and communicative signals shaping those interpretations.
- Lightweight contextual signals may support initial assessment, while richer contextual, evidential, or multimodal information can be selectively retrieved for deeper verification.
8 Ethics Statement
RCMN is constructed from publicly circulated discourse and professionally reviewed source material using AI assistance followed by human verification and adjudication. The study limits its scope to research use and defined communicative signals, without inferring private attributes or intentions beyond that framework.
- RCMN uses publicly circulated discourse and professionally reviewed source material, with AI-assisted evidence recovery and initial annotation followed by human verification and adjudication.
- The dataset may contain sensitive or politically salient public communication.
- The study focuses on research use and does not infer private attributes or intentions beyond the communicative signals defined in the annotation framework.