Source-linked AI summary
Definitional Sensitivity in Media Bias Detection: A Multi-Definition Dataset and Benchmark
Martin Wessel, Timo Spinde, Jürgen Pfeffer, Gianluca Demartini
TL;DR
Media-bias datasets often use varying or unstated definitions, leaving open whether similarly named categories measure the same construct. This paper tests definitional sensitivity with 354 participants and four LLMs, finding that conceptual framing shifts ratings while construct-preserving elaboration does not, with stronger sensitivity among LLMs.
Problem
Media-bias datasets rely on varying or unstated definitions, making it unclear whether similarly named categories measure the same construct or different phenomena.
Method
A preregistered between-subjects 3 × 2 experiment compares three conceptual definitions and two elaboration lengths across four bias categories, with parallel evaluation by four LLMs.
Results
Conceptual specification consistently affects human and LLM ratings, whereas construct-preserving detail does not; LLMs amplify definitional sensitivity by approximately 1.3–4× relative to humans.
Takeaways & Limitations
Definitions should be treated as experimental interventions in media-bias annotation and prompt-based measurement, supported by the released MUDD dataset.
Takeaways & Limitations
The findings concern four bias categories and U.S. English-speaking annotators, so generalization to other bias types, populations, languages, and larger models remains open.
Abstract
from arXiv · showhide
Media bias detection relies on definitions and examples that specify what counts as bias, yet these specifications often vary across datasets or remain implicit, even when given the same name. Such variation makes it unclear whether models trained for the same bias category learn the same construct or different phenomena, a problem largely overlooked in prior work. We examine how definition choice affects bias annotation in a between-subjects experiment with 354 participants and a parallel evaluation with four LLMs. Participants and models rate six news articles across four bias categories using definitions that vary in conceptual framing and elaboration. Across 8,496 human and 28,800 LLM ratings, we find that the conceptual target of a definition drives annotation divergence, while construct-preserving elaboration does not: conceptual framing significantly shifts annotations for humans and does so even more strongly for LLMs. We discuss implications for construct specification in annotation protocols and prompt-based measurement, and consider how definitional sensitivity may propagate to downstream classification beyond media bias. We also release MUDD, the Multi-Definition Bias Detection Dataset.
1 Introduction
Media bias datasets depend on definitions that connect the abstract construct to annotation labels, but those definitions vary and can produce divergent judgments on identical articles. This paper experimentally tests that effect across humans and LLMs.
- Research gap: Most media-bias datasets use varying or unstated definitions, leaving unclear whether similarly named categories measure the same construct.A review of 115 datasets documents this definitional inconsistency.
- Research gap: Three standard concepts of linguistic bias produce significantly different mean ratings for the same articles.The concepts concern biased word choice, grammatical agency manipulation, and demeaning labels, evaluated on 300 articles.
- Research gap: Definitional choices may cause datasets labeled as measuring the same construct to capture different phenomena, weakening cross-study comparability and classifier validity.Prior work documents variation across datasets but does not isolate its effect on annotation behavior.
- Approach: The study uses a preregistered between-subjects 3 × 2 concept-by-length experiment with 354 participants rating six of 300 articles across four bias categories.Participants receive one of six randomly assigned definitions per category; four LLMs perform a parallel evaluation.
- Findings: Conceptual specification, rather than construct-preserving elaboration, drives annotation divergence across all four categories, while tested LLMs amplify sensitivity by approximately 1.3–4× relative to humans.The study releases 8,496 human annotations, 28,800 LLM annotations, 24 definitions, and experimental materials as MUDD.
2 Related Work
Prior work documents inconsistent definitions and guideline effects in media-bias and related annotation datasets, but has not directly tested how conceptual definitional variation changes media-bias judgments.
- Definitions and taxonomies: A review of 115 MBIB datasets finds inconsistent and often unstated media-bias definitions, while taxonomies organize their differing conceptualizations.Prior work documents definitional divergence but does not test its causal effect on annotation.
- Annotation guidelines: Narrowly defined bias types yield more focused datasets, whereas clearer definitions may improve quality at the cost of generalizability.Persistent low inter-annotator agreement has also been reported even with clear guidelines.
- Hypotheses: Existing hypotheses predict that conceptual definitions and definition length can change bias ratings and confidence, with concept effects exceeding political-orientation and gender-attitude effects.These hypotheses motivate separating conceptual content from elaboration.
- LLM annotation: Research on LLM annotation reports that reliability varies substantially with task framing and prompt design, motivating direct comparison of human and model sensitivity.LLMs are increasingly considered scalable annotation alternatives, but framing remains consequential.
3 Methodology
The methodology isolates conceptual definition effects from wording length and stimulus variation through a preregistered factorial annotation experiment spanning four bias categories, diverse articles, human participants, and LLMs.
- Research design: The preregistered design varies three theoretically separable concepts and two definition lengths across linguistic, gender, political, and text-level context bias.The study aims to isolate conceptual content from length, register, and example count.
- Research design: Participants receive one of six definitions per category, with short and long versions differing through construct-preserving elaboration.Participants do not mix definition lengths within the task.
- Stimuli and sampling: The stimulus set contains 300 English-language news articles sampled across four topics and stratified by predicted bias label and outlet political leaning.The articles come from 39 outlets spanning left, center, and right positions.
- Participants: 354 U.S.-based English-speaking participants were recruited after power analysis indicated N = 342 for the most demanding interaction test.Prolific prescreening approximated U.S. Census distributions for political orientation, gender, and age.
- Procedure: Participants rate each article on a 6-point bias scale and a 7-point confidence scale after reading six articles sequentially.Definitions remain accessible during annotation, and an attention check permits one retry.
- Coverage: Counterbalanced assignment ensures all 300 articles are annotated under all six definitions across four categories, yielding 7,200 article–category–definition combinations.The design uses 27 covering sets, each applied once per length group.
- LLM evaluation: Four LLMs replicate the human task to compare definitional sensitivity across frontier commercial and same-size open-weight models.The tested models span provider and cost tiers.
4 Results
Across human and LLM ratings, conceptual definition choice produced systematic annotation differences across bias categories, whereas construct-preserving elaboration had limited effects. LLMs generally amplified these concept effects, while within-definition agreement remained modest.
- Concept effects: Concept effects were significant for all four bias categories, strongest for linguistic bias and smaller for political and context bias.For linguistic bias, η2 = .044; gender η2 = .011; context η2 = .004; political η2 = .003.
- Length effects: Definition length had negligible effects overall, with a significant effect only for linguistic bias and a smaller effect than the concept effect.For linguistic bias, long definitions averaged 3.50 versus 3.24 for short definitions; η2 was .006 for length versus .043 for concept.
- Robustness checks: Annotator characteristics and definition clarity contributed little: political orientation and gender attitudes stayed within ±0.10 scale points, while clarity ratings were comparable.All 24 definitions received clarity means of approximately 3.84–4.07.
- Within-definition reliability: Within-definition agreement was modest, with Krippendorff’s α = 0.39 and 63% of rating pairs agreeing within one scale point.Agreement rose to α = 0.47 among combinations with at least three raters, and longer definitions yielded only slightly higher agreement (α = 0.40 versus 0.35).
- LLM comparison: All four LLMs showed larger concept effects than humans across almost all categories, with amplification of roughly 1.3–4× and architecture-dependent magnitudes.Category ordering was broadly preserved, but model rankings varied by category; Llama’s near-zero gender effect reflected floor compression.
- LLM rating behavior: LLMs rated below humans on average, and their near-constant confidence scores did not signal uncertainty when ratings shifted across definitions.LLM means were 2.09–2.45 versus 3.11 for humans; model confidence SDs were 0.69–0.94 versus 1.56 for humans.
- Article-level sensitivity: Article-level rating ranges reached 5.0 points, with mid-scale articles more definition-sensitive than extremes for three human and LLM bias categories.The exception was context bias, where the mid-scale versus extreme comparison was not significant.
5 Discussion
Across human and LLM annotations, conceptual framing—not construct-preserving elaboration—drives divergent ratings, with consequences for construct validity, model evaluation, and annotation practice.
- Across all four bias categories, conceptual framing significantly shifts ratings, whereas construct-preserving elaboration does not consistently do so.Length reaches significance only for linguistic bias, with a much smaller effect size.
- The authors interpret rapid schema formation as a mechanism: annotators absorb added elaboration into an initial conceptual schema rather than refining it.Re-click rates suggest definitional dependence is greatest where the construct is least intuitive.
- 5.1 Same Label, Different Constructs: Different concepts within one bias category likely produce systematically different datasets, even when guidelines share the same category label.The divergence reflects different measured phenomena rather than merely different levels of specificity.
- Definitional choices matter most for borderline articles, while clear-cut cases are comparatively robust across definitions.For three of four categories, mid-scale articles show the largest cross-definition rating shifts.
- Confidence tracks rating extremity (r = .42–.45), not inter-annotator convergence, so it indexes decisiveness rather than reliability.This distinction matters for aggregation schemes that use confidence as a reliability proxy.
- 5.4 Practical Recommendations: For subjective NLP constructs, concept selection appears more consequential than surface phrasing, shifting attention toward construct specification.The authors recommend publishing exact definitions, testing multiple concepts, and reporting definition-sensitive variance.
- 5.5 Future Work: The study does not establish whether these effects generalize beyond the four tested bias categories or to broader model and linguistic settings.The authors call for replication across other tasks, larger models, and multilingual settings.
6 Conclusion
The paper shows that conceptual content, rather than added detail, drives definitional sensitivity in human and LLM bias ratings, and releases MUDD to study this effect.
- Across 354 humans and four LLMs, conceptual specification drives annotation divergence more than construct-preserving detail.LLMs amplify this sensitivity in architecture-dependent ways.
- MUDD contains 37,296 annotations across 24 definitions for studying definitions as experimental interventions.
Limitations
The study’s limitations concern the breadth of bias categories, participant and model coverage, prompt scaffolds, and self-reported covariates. These boundaries leave open whether the findings generalize across bias types, populations, model capabilities, and prompting conditions.
- Scope of bias categories: The experiment covers four bias categories, leaving generalization to racial, reporting-level, and framing bias open for future work.The studied categories are linguistic, gender, political, and text-level context bias.
- Participant and model coverage: Findings are based on U.S.-based, English-speaking Prolific crowdworkers, so transfer across languages, cultures, and expertise is constrained.The passage notes that bias perception varies with language, culture, and expertise, and that bias measurement may not transfer cleanly across languages.
- Participant and model coverage: Larger open- or closed-weight models remain untested, leaving model-capability interactions with framing and prompt effects unresolved.The supplied passage specifically identifies 70B-class open-weight models and larger closed-weight models as untested.
- Prompting and covariates: The study holds prompt scaffolds constant, so other scaffolds such as chain-of-thought or role prompting could compress or amplify concept effects.Keeping prompt form identical preserves comparability across conditions but does not test scaffold-dependent changes.
- Prompting and covariates: Political orientation, gender attitudes, news consumption, and attention covariates are subject to social-desirability bias.This limitation affects the interpretation of the measured covariates rather than the definition manipulation itself.
Ethical Considerations
The study reports participant consent and voluntary withdrawal procedures, alongside measures for ratings, confidence, reading behavior, attention, and experimental conditions.
- The study received ethics approval, and participants gave explicit informed consent after reviewing study purpose, data collection, duration, and compensation.Participation was voluntary, with withdrawal permitted at any time without explanation.
- Bias ratings used a 6-point Likert scale, while confidence used a 7-point scale.
- The between-subjects manipulation varied three conceptual levels and two definition lengths.
- The study measured political orientation, gender attitudes, definition clarity, reading attention, reading time, re-click frequency, and attention-check performance.
B LLM Robustness Check
A temperature-0.7 GPT-4o rerun shows that conceptual definition effects remain significant and exceed within-cell sampling variation across all four bias categories.
- Concept effects remain significant at T=0.7 across linguistic, gender, political, and context bias.All reported p-values are <.001, with η2 values of .230, .205, .067, and .041, respectively.
- Between-concept rating differences exceed within-cell standard deviations by 1.3–4.1× across the four categories.The ratios are 3.7× for linguistic, 3.5× for gender, 4.1× for political, and 1.3× for context bias.
- The robustness rerun uses a stratified 60-article × 24-definition subset at T=0.7 with five samples per cell.
C.1 Definition Derivation Procedure
The study selects four bias categories from a broader literature-based taxonomy, covering both mechanism-based and topic-based forms of media bias.
- The four categories are linguistic, gender, political, and text-level context bias.
- Mechanism-based categories: Linguistic and context bias represent complementary textual levels, from local linguistic choices to global informational structure.
- Topic-based categories: Gender and political bias are selected as prominent topic-based categories in media bias research.
C.2 Selection of Concepts
The concepts are designed to isolate conceptual content, reflect genuine theoretical distinctions, and vary independently, while acknowledging that they sample only part of definitional space.
- Design goals: Definitions are matched on surface features so rating differences can be attributed to conceptual content rather than length, register, or example count.Each concept has short and long versions, with construct-preserving elaboration varied while the concept remains fixed.
- Design goals: Each concept is grounded in an existing research tradition rather than being an invented variant.
- Design goals: The three concepts within each category target analytically distinguishable mechanisms that can vary independently within a single article.This supports interpreting between-definition rating differences as effects of the definition rather than inherent article properties.
- Scope: The concepts sample a larger definitional space rather than claiming taxonomic completeness, and claims concern between-subconcept differences rather than absolute rating levels.
- Definition construction: The study uses separated mechanism phrasings and examples to reduce dependence on any single contestable wording choice.Small wording changes could still shift absolute effect sizes, although the authors state they are unlikely to erase the divergence pattern.
- Validation: All 24 definitions undergo author review and a Prolific pilot testing conceptual accuracy, parity, register consistency, and comprehensibility.
- Selected concepts: The selected concepts span linguistic, gender, political, and context bias, with three separable concepts chosen for each category.
C.5 Full Definition Texts
The appendix presents the complete definition texts used in the experiment for all four bias categories, organized by category-specific tables.
- Each bias category contains three conceptual sub-definitions, C1–C3, each presented in short and long versions.Short versions contain one to two sentences and one example; long versions provide extended explanations and three examples.
- Table 4 contains the complete definition texts for linguistic bias.
- Table 5 contains the complete definition texts for gender bias.
- Table 6 contains the complete definition texts for political bias.
- Table 7 contains the complete definition texts for text-level context bias.