Source-linked AI summary
Framing War Across Languages: Power, Agency, and Sentiment in Wikipedia's Multilingual War Narratives
Jiarui Xia, Diego Gomez-Zara
TL;DR
Wikipedia’s neutral war narratives may still reflect the perspectives of distinct language communities, but large-scale evidence on relational portrayals across languages is limited. This study analyzes multilingual war narratives using connotation frames and finds divergence when communities are directly involved, but largely convergent narratives otherwise.
Problem
Prior studies have not systematically examined how combatants are relationally portrayed across many wars and languages, or how portrayals depend on communities’ relationship to conflict.
Method
The study analyzes 2,542 battles from 158 post-1900 wars across 20 Wikipedia language editions using connotation frames measuring power, agency, and sentiment.
Results
Wikipedia narratives diverge systematically when language communities directly participate in conflicts, but largely converge when they are not directly involved.
Takeaways & Limitations
War-framing asymmetries vary across languages and reflect self-referential identification rather than broader alliance-based identification.
Takeaways & Limitations
Translating all articles into English may obscure subtle semantic distinctions and culturally specific connotations, while sentiment scores showed lower cross-language annotator agreement.
Abstract
from arXiv · showhide
While Wikipedia promotes a neutral point of view on historical conflicts, its language editions are written by editors from distinct linguistic and cultural communities. In this study, we analyze 158 wars since 1900 to examine how the descriptions of combatants vary across 20 Wikipedia language editions. Using connotation frames---which assess power, agency, and sentiment toward an entity---we examine how each language portrays the parties involved in the conflict. We find systematic differences when language editions describe wars involving their own communities, although the direction of these asymmetries varies across languages. However, when language editions describe conflicts that do not involve their own linguistic communities, their narrative structures exhibit high cross-linguistic similarity. These findings show how linguistic communities influence war narratives on Wikipedia, revealing that shared historical accounts remain shaped by the perspectives of the language communities that produce them.
Introduction
This study examines how 20 Wikipedia language editions relationally portray combatants in 158 post-1900 wars, despite shared Neutral Point of View norms. It finds that direct involvement produces systematic but directionally varying differences, whereas uninvolved languages largely converge in their war narratives.
- Background: Wikipedia consists of independent language-edition communities embedded in distinct cultural, linguistic, and geopolitical contexts despite its Neutral Point of View policy.These communities can narrate the same conflict differently, including in responsibility assignment and actor portrayal.
- Research gap: Existing research documents cross-language differences and ingroup bias but leaves relational portrayals of combatants and their dependence on language-community involvement unresolved.Prior work often relies on coverage measures, coarse sentiment analysis, or small-scale qualitative studies.
- Approach: 2,542 battles from 158 post-1900 wars across 20 Wikipedia language editions are analyzed using connotation frames derived from subject-verb-object relations.The framework measures combatant portrayals in terms of power, agency, and sentiment.
- Findings: Systematic differences emerge in portrayals of combatants, with some languages depicting self-aligned entities as more powerful and agentic and others as less so.The direction of asymmetry varies across language editions.
- Findings and contributions: When describing wars that do not directly involve them, language editions largely converge in narrative structure, identifying direct involvement as a source of divergence.The study extends framing analysis to Wikipedia at scale and compares relational combatant portrayals beyond coverage metrics.
Related Work
Prior research shows that war memories and Wikipedia narratives are shaped by political, cultural, linguistic, and governance contexts rather than simply recording conflicts neutrally. Existing studies identify ingroup bias and cross-linguistic variation, but have not systematically analyzed relational portrayals of combatants across many wars and languages.
- Collective memory and mediation: War memories reflect political agendas, national identity, and cultural traditions, while media reconstruct the past by selectively elevating actors within moral frameworks.These processes help explain why collective remembrance is not a fixed record of historical events.
- Cross-linguistic variation: Wikipedia’s more than 300 language editions represent shared topics across cultural contexts but exhibit systematic cross-linguistic differences in content and emphasis.Prior examples include differing emphases on personal details and national identity across English and Polish articles.
- Conflict framing: Conflict-focused studies find ingroup bias, with language editions portraying their own groups as less immoral and less responsible than opposing groups.Qualitative studies have examined conflicts including Srebrenica, the Portuguese Colonial War, and the Second Sino-Japanese War.
- Governance and editorial outcomes: Governance structures also shape editorial outcomes, as concentrated administrative power enabled nationalist revisionism in Croatian Wikipedia but not in the more distributed Serbian edition.The comparison links institutional differences to divergent distortions between linguistically similar editions.
- Methodological foundations: Connotation frames capture implied sentiment, power, and agency in predicate–argument structures, extending beyond coverage measures and small-scale human-coded analyses.Prior applications include gender bias in film scripts, multilingual sentiment lexicons, and other social and cultural contexts.
- Research gap: Despite evidence of cross-linguistic variation, prior work has not systematically examined relational portrayals of combatants across a large set of wars and languages.This gap motivates studying how power, agency, and sentiment are attributed through the relational structure of text.
Data Collection
The study assembled a multilingual Wikipedia dataset of wars and battles using MediaWiki APIs, expanding from wars after 1900 to related battle articles. Articles were filtered for valid content and sufficient editorial activity before analysis.
- War selection: 788 wars remained after collecting post-1900 wars from English Wikipedia and removing entries with invalid article or belligerent links.The initial collection used Wikipedia’s “List of wars by date” category and retained major-war metadata.
- Article expansion: 1,558 additional battle and conflict articles were identified by expanding links from the 158 war pages across available language editions.Linked pages were retained when infoboxes listed belligerents and a valid conflict timeframe, then filtered for known, non-conflicting languages.
- Battle collection: 2,581 battles formed the final dataset after deduplicating and aggregating link-based and category-based collections.World War I and World War II battles were identified through curated list articles instead of English Wikipedia categories.
- Editorial activity: Articles were retained only if they had at least 10 edits, a threshold chosen to exclude only the least-edited articles while preserving lower-resourced language coverage.The 10th percentile of edit counts was 11, and a 20-edit threshold would disproportionately reduce Persian and Romanian coverage.
Methods
The study uses a six-stage pipeline to compare multilingual war narratives through entity-level power, agency, and sentiment frames. It links entities to national alignments, tests self–enemy and allied–opponent differences, and examines cultural, editorial, and outcome-related correlates.
- Data and pipeline: The six-stage pipeline collects articles, translates them, extracts SVO triples, computes connotation frames, links entities to national alignments, and compares portrayals across languages.The scripts and materials for reproducing the analysis are available on OSF.
- Connotation frames: The adapted RIVETER pipeline retains each SVO triple, its connotation scores, and its source sentence for event-level analysis of power, agency, and sentiment.Power captures implied dominance or control, agency captures intentionality and capacity to influence, and sentiment captures positive or negative affect.
- Entity alignment: Entities are linked to Wikidata metadata to infer country and geopolitical affiliations, including citizenship information for people when available.These alignments support categorization by self versus enemy and allied versus opposing combatants.
- Narrative analysis: For RQ2, BERTopic models compare topics in sentences where self-aligned entities receive higher versus lower power or agency scores after removing entity-dominated topics.The analysis maps aligned SVO structures back to their original sentences within each language.
- Statistical analysis: For RQ3, exploratory Spearman correlations and nine OLS models relate framing gaps to cultural values, editorial behaviors, and battle outcomes, with Benjamini–Hochberg-adjusted significance levels.The correlations use language-level averages and are interpreted cautiously because the sample contains n = 20 language editions.
Results
Results show that Wikipedia war narratives differ most in self-referential contexts, with language-specific asymmetries in power, agency, sentiment, and framing. By contrast, portrayals of allied or uninvolved actors show no significant asymmetries or broadly high cross-linguistic similarity.
- Self-aligned versus enemy-aligned portrayals: Languages diverged in self-aligned portrayals: English showed higher agency, Japanese and Hebrew higher power and agency, Russian differences across all dimensions, while Arabic, Ukrainian, and Vietnamese showed lower power or agency.These results compare self-aligned with enemy-aligned entities after multiple-comparison correction.
- Self-aligned versus enemy-aligned portrayals: No language showed significant allied-versus-enemy differences on any dimension after correction, indicating that asymmetries primarily emerged in self-referential contexts.Ukrainian and Hebrew were excluded because of insufficient sample sizes.
- Linguistic framing patterns: Russian used award in 3% of positive self-aligned descriptions, compared with 0.02% in Japanese and 0.01% in Hebrew, with a significant asymmetry only in Russian.The verb appeared more often in positive descriptions of self-aligned than enemy-aligned entities in Russian.
- Linguistic framing patterns: Arabic, Ukrainian, and Vietnamese emphasized resistance when self-aligned entities scored higher but failure or inability when scores were lower, whereas stronger-framing languages foregrounded ingroup losses.Chinese and Spanish showed no clear or consistent narrative pattern across conditions.
- Correlates of framing differences: Win rate correlated positively with power and agency gaps (ρ = .95 and .92), while loss rate correlated negatively (ρ = −.94 and −.91); Long-Term Orientation correlated moderately positively.All reported associations were significant after adjustment.
- Cross-linguistic similarity: Non-involved-war narratives had uniformly high similarity scores from 0.7 to 0.9, with no clear clusters, although Arabic and Dutch were central while Persian, Czech, Vietnamese, and Finnish were peripheral.Peripheral languages branched off earlier in the hierarchical dendrogram.
Discussion
Wikipedia’s war narratives diverge when language communities are directly involved, but largely converge for conflicts outside their communities. These patterns reflect locally situated perspectives interacting with shared editorial norms, while the study’s descriptive findings remain subject to translation, sampling, geographic, and pipeline limitations.
- Core findings: Directly involved language communities show pronounced narrative differences, whereas narratives largely converge when languages describe battles in which they are not directly involved.Similarity between non-involved languages ranges from 0.7 to 0.9.
- Core findings: Self-aligned portrayals vary by language: English, Japanese, Hebrew, and Russian emphasize greater power or agency, while Arabic, Ukrainian, and Vietnamese emphasize less.No language showed significant allied-versus-enemy differences, indicating that the asymmetries were self-referential rather than alliance-based.
- Explanatory factors: Languages with higher victory proportions and long-term orientation tended to exhibit larger self-enemy power and agency gaps, although the 20-language correlation analysis requires caution.Article-level regression models also corroborated the role of battle outcomes.
- Editorial implications: Higher editorial concentration is associated with suppressed agentic portrayals of opposing entities, linking unequal editorial influence to conflict narratives.The finding quantitatively complements prior qualitative work on governance capture in specific language editions.
- Editorial implications: Non-involved language editions may provide comparative references for editors revising wars involving their own communities and could support editorial assistance tools.The recommendation follows their observed high cross-linguistic similarity.
- Limitations: The study is descriptive rather than causal and is limited by English translation, unequal battle counts, a small language sample, geographic heterogeneity, and incomplete whole-corpus pipeline evaluation.Validation found comparable error propagation across languages for agency and power, but not for sentiment.
Conclusion
Wikipedia’s war narratives diverge systematically across language editions for conflicts involving their communities but largely converge for conflicts that do not. Framing asymmetries vary across languages and appear for self-referential identification, not broader alliance structures.
- Conclusion: Wikipedia war narratives diverge systematically across language editions when communities describe conflicts in which they are directly involved.The study finds largely convergent narratives when communities describe conflicts in which they are not directly involved.
- Conclusion: The direction and mechanisms of framing asymmetries vary across languages, extending prior findings on ingroup bias.The conclusion presents this variation as an extension of earlier ingroup-bias findings.
- Conclusion: Framing asymmetries occur for self-referential identification but not for broader alliance structures.This distinction separates community-based identification from alliance-based relationships.
Paper Checklist to be included in your paper · A Language Assignment for Belligerent Entities
The paper checklist documents ethical, methodological, theoretical, reproducibility, and reporting practices, while belligerent entities receive official-language assignments through Wikidata and Wikipedia country affiliations. The study emphasizes descriptive, non-causal interpretation, limitations, responsible release, and transparency about potential misuse.
- Paper Checklist to be included in your paper: The paper presents a descriptive, comparative analysis and does not overclaim causal relationships.The abstract and introduction are described as accurately reflecting the study’s contributions and scope.
- Paper Checklist to be included in your paper: Connotation frames, the RIVETER pipeline, entity linking, and statistical testing are justified as appropriate for relational portrayals of entities.The methodological justification includes procedures for entity linking and statistical testing.
- Paper Checklist to be included in your paper: The study reports limitations including translation artifacts, sample-size variation, a small language sample, lack of human evaluation, and geographic-distribution differences.The analysis is explicitly characterized as descriptive rather than causal, and language-specific nuances may be obscured by translation.
- Paper Checklist to be included in your paper: The authors discuss potential misuse, including delegitimizing Wikipedia or inflaming intergroup tensions, and frame transparency about collective-memory construction as the research aim.They also connect the findings to Wikipedia’s governance model, neutrality norms, and future social-science research.
- Paper Checklist to be included in your paper: Publicly available Wikipedia data, documented processing procedures, released code, and supporting materials support reproducibility without collecting or releasing personally identifiable information.The paper states that Wikipedia content is publicly available and that no new dataset is being curated or released.
- Paper Checklist to be included in your paper: The paper acknowledges that editor composition, source availability, and cultural-memory traditions may explain differences that cannot be fully disentangled.It also notes that connotation frames were developed primarily from English-language sources and may miss language-specific connotations.
- Paper Checklist to be included in your paper: Annotators were paid $40 each for approximately 70 minutes, with 12 annotators across four languages and total compensation of $480.The study received institutional IRB approval, assessed participant risk as minimal, and used pseudonymized annotation data.
- A Language Assignment for Belligerent Entities: Belligerent entities’ official languages are assigned by linking English Wikipedia entities to Wikidata QIDs and using entity type and country affiliation.Countries and states use official languages from English Wikipedia infoboxes, while organizations and groups inherit languages from their affiliated sovereign country.
B Implementation Details … D.2 Task Instructions
The study used BERTopic with pre-trained sentence embeddings and default dimensionality-reduction settings, then validated translated connotation annotations through a structured multilingual human study. The study also specified an OLS framing-outcome model and implemented comprehension, consistency, and translation-alignment checks.
- B Implementation Details: BERTopic version 0.15.0 performed topic extraction and clustering, largely using its default configuration.The framework’s default settings were retained unless otherwise specified.
- B Implementation Details: Document embeddings used the pre-trained allmpnet-base-v2 SentenceTransformer model without additional training or task-specific fine-tuning.The model was selected for its sentence-level semantic-similarity performance.
- B Implementation Details: Dimensionality reduction used BERTopic’s default UMAP configuration, while HDBSCAN clustering set the minimum cluster size to 20 and left other hyperparameters at default values.HDBSCAN was chosen as a noise-robust density-based method for discovering topics of varying sizes.
- C OLS Formulation: The OLS outcome yij represented one of nine battle-level framing measures covering self-, enemy-, or gap-based power, agency, and sentiment scores.The formulation additionally included revert rate, edit concentration, log(unique editorsij), and log(talk lengthij) predictors.
- D Translation Human Validation Study: The translation validation study assessed data quality and examined how translation could affect connotation scores.Because the connotation lexicon is English-based, non-Latin-script languages were considered especially susceptible to translation-induced distortion.
- D.1 Language Choice: Chinese, Arabic, and Japanese were selected as primary validation languages because they represent distinct language families and NLP challenges.The stated challenges were Arabic’s morphological complexity, Chinese’s lack of inflectional morphology, and Japanese’s agglutinative structure.
- D.2 Task Instructions: For each language, 50 SVO triples were randomly sampled, and annotators completed an IRB-approved Qualtrics study recruited through Prolific without collecting personally identifiable information.Participants viewed an introduction and consent form before beginning the task.
- D.2 Task Instructions: Annotators received 10 examples, had to pass a five-question quiz with at least 80% accuracy, and were monitored with a repeated triple used as an attention check.Each triple included source-language context, an English translation, article metadata, and extracted English Subject, Verb, and Object fields; native-speaker verification checked source–translation alignment before release, except in English.
D.3 Task Settings · D.4 Annotator Quality Control · D.5 Inter-Annotator Agreement
The study used language- and geography-specific annotator eligibility rules, quantified disagreement across language-specific judgment sets, and evaluated agreement with question-appropriate Krippendorff’s α. The agreement analysis also accounts for restricted-range effects that can depress α despite high observed agreement.
- D.3 Task Settings: Japanese annotators had to reside in Japan, Chinese annotators in China, and English annotators had to be native English speakers without geographic restrictions.All annotators also needed English proficiency because the SVO triples and connotation labels were presented in English.
- D.3 Task Settings: All annotators were required to be proficient in English because annotation materials were presented in English.The materials included extracted SVO triples and connotation score labels.
- D.4 Annotator Quality Control: Disagreement rate was defined as the proportion of instances where an annotator’s response differed from both other annotators.This reliability measure followed Park et al. (2021).
- D.4 Annotator Quality Control: 350 judgments per annotator were used for Chinese, Arabic, and Japanese, versus 300 for English.The totals came from 50 items × 7 questions for Chinese, Arabic, and Japanese, and 50 items × 6 questions for English.
- D.4 Annotator Quality Control: The annotation questions and response options were documented for the translation human validation study.These materials are presented in Table D.1.
- D.5 Inter-Annotator Agreement: Krippendorff’s α used nominal agreement for Q1–Q2 and ordinal agreement for Q3–Q7.The choice reflected binary categorical responses for Q1–Q2 and ordered scales for Q3–Q7.
- D.5 Inter-Annotator Agreement: Low α for Q1–Q2 can result from restricted response ranges even when percent agreement is generally high.A skew toward one response increases expected chance agreement, depressing α regardless of true annotator reliability.
D.6 Error Propagation Analysis
The analysis stratifies agreement by translation quality and SVO accuracy, finding robust Power and Agency correspondence but substantially weaker Sentiment consistency. Translation deviations have little effect on Power and Agency agreement, whereas Sentiment remains limited by subjectivity and cultural specificity.
- Method: Items were stratified by Perfect, Minor, or Major translation quality and Accurate or Inaccurate SVO accuracy, with Wrong translations excluded because none occurred.Subject and object were treated separately, and annotator judgments were mapped to (−1, 0, +1) values.
- Power and Agency: 78.3%/80.4%, 80.0%/80.0%, and 70.0%/65.4% Power/Agency agreement occurred for Perfect-translation subject-accurate triples in Arabic, Chinese, and Japanese, respectively.Arabic and Chinese consistently agreed more with lexicon scores than Japanese across conditions.
- Power and Agency: Minor translation deviations did not substantially change Power or Agency agreement relative to Perfect translation across languages.This stability held across the analyzed conditions.
- Sentiment: Sentiment consistency was substantially lower than Power and Agency across all conditions and languages, likely because sentiment attribution is subjective and culturally specific.The English-based RIVETER lexicon may not capture affective presuppositions in war narratives, and sentiment dimensions also showed low inter-annotator agreement.
D.7 Robustness Check: Alternative Translation Service
The RQ1 analysis was replicated with the Wikimedia Foundation’s MinT translation service instead of Google Translate. Results were consistent with the main analysis, supporting robustness across translation pipelines.
- Alternative Translation Service: MinT is an open-source, Wikipedia-native machine translation system used independently of the Google Translate API.The independent comparison helps rule out artifacts of a single translation engine.
- Alternative Translation Service: The replication yielded results consistent with the main analysis, supporting the validity of the findings across translation pipelines.Tables D.7 and D.8 report the consistent replication results.
E Mixed-Effects Regression Robustness Check
A robustness analysis re-estimated the framing regressions as linear mixed-effects models with language-specific random intercepts, while retaining the main analysis outcomes and fixed effects. The mixed-effects results were substantively similar to OLS, indicating robustness to modeling language-level clustering.
- Model specification: The study re-estimated self-, enemy-, and gap-based framing scores across power, agency, and sentiment using linear mixed-effects models.Language was specified as a random intercept, with battle outcomes and the same editorial behavior covariates included as fixed effects.
- Model specification: The model represents each framing outcome as a function of covariates, a language-specific random intercept, and residual error.Here, y_ij denotes the framing outcome for battle i in language edition j, while u_j captures language-specific variation.
- Estimation checks: Continuous predictors were standardized, count variables were log-transformed, and residual and multicollinearity diagnostics were inspected.Assumption checks included residual normality, skewness, kurtosis, Q–Q plots, heteroscedasticity, and multicollinearity.
- Robustness result: The mixed-effects models produced substantively similar results to the main-text OLS models.This similarity suggests that the observed framing patterns are robust to alternative specifications accounting for language-level clustering.