Source-linked AI summary
A Social Media Analysis of Discourse on the Israel--Palestine Conflict on Telegram
Michail Zafeiropoulos, Despoina Antonakaki, Sotiris Ioannidis
TL;DR
Research on conflict discourse has not systematically compared pro-Israel and pro-Palestine Telegram communities at scale. This study analyzes longitudinal Telegram data with complementary sentiment, stance, and framing methods, finding that the communities use death- and victim-related vocabulary in different emotional registers. Fine-tuned BERTweet performs best for stance detection, while the discourse findings distinguish predominantly neutral, report-style pro-Israel messaging from more negative pro-Palestine messaging.
Problem
Large-scale longitudinal comparisons of pro-Israel and pro-Palestine Telegram communities using complementary NLP methods remain limited.
Method
The study analyzes 87,617 messages from 16 Telegram channels using sentiment analysis, three stance detection methods, and framing analysis, with stance methods evaluated on 736 manually annotated messages.
Results
Fine-tuned BERTweet achieved 72.1% accuracy and 0.721 macro F1, while the communities used death- and victim-related vocabulary in contrasting emotional registers.
Takeaways & Limitations
Reading stance, sentiment, and framing together reveals pro-Israel channels as predominantly neutral and report-style, while pro-Palestine channels are markedly more negative.
Takeaways & Limitations
The findings are scoped to politically aligned channels and do not extend to general news channels, multilingual content, or communities between the two poles.
Abstract
from arXiv · showhide
Social media has become a central arena in which armed conflicts are contested, yet the pro-Israel and pro-Palestine communities on Telegram, whose broadcast architecture yields an unusually direct record of deliberate political communication, have not been systematically compared at scale. This study presents a multi-method computational analysis of 87,617 messages from sixteen Telegram channels, eight pro-Israel and eight pro-Palestine, spanning May 2021 to June 2026 and covering multiple conflict escalations. It combines sentiment analysis, three stance detection methods drawn from distinct paradigms (keyword matching, zero-shot DeBERTa via natural language inference, and a fine-tuned BERTweet model), and a framing analysis, all evaluated against 736 manually annotated messages. The fine-tuned model performed best (72.1% accuracy, 0.721 macro F1 under 5-fold cross-validation), outperforming both label-free baselines by 8 to 11 points; the baselines stalled in the low-to-mid 60s, indicating a hard ceiling for stance detection not adapted to in-domain language. The central finding emerges only when sentiment, stance, and framing are read together: the two communities deploy the same death- and victim-related vocabulary in opposite emotional registers, pro-Israel channels predominantly neutral and report-style, pro-Palestine channels markedly more negative, consistent with writing from the distinct discourse positions of acting party and affected party.
1 Introduction
Telegram is an important venue for conflict-related political communication, but pro-Israel and pro-Palestine communities have not been systematically compared there at scale. This study addresses that gap through a longitudinal, multi-method analysis of opposing channels.
- Research gap: Existing research usually treats Telegram as one platform among several or focuses on Twitter/X and individual conflict episodes.These studies are predominantly qualitative or descriptive and rarely examine Telegram as a dedicated object of study.
- Study scope: The study presents a large-scale comparison of pro-Israel and pro-Palestine Telegram channels across multiple conflict phases.The corpus contains 87,617 messages from 16 public channels collected between May 2021 and June 2026.
- Data contribution: The study contributes a curated stance-annotated Telegram dataset for evaluating supervised and zero-shot methods on conflict-related social media content.This enables systematic comparison of methods adapted to politically committed Telegram discourse.
- Method comparison: Three stance methods are compared: keyword matching, zero-shot DeBERTa inference, and a fine-tuned BERTweet classifier.The fine-tuned model achieves 72.1% accuracy and a 0.721 macro F1-score under five-fold cross-validation.
- Analytical approach: The analysis combines sentiment analysis, stance detection, and narrative framing to examine how conflict discourse changes over time.The framework links messaging patterns to major conflict events and military escalations.
2 Theoretical and Computational Background
The section frames social media as a central infrastructure for political communication in conflict and introduces computational tools for analyzing emotional tone, stance, and discourse. It emphasizes Telegram’s direct broadcast architecture, the target-dependent nature of stance, and the difficulty of applying standard methods to reportorial, implicitly framed conflict messages.
- Social Media and Political Communication: Social media embeds political communication in commercial architectures that shape how user-generated content circulates and is contested.The section contrasts claims of empowerment with the non-neutral infrastructures through which political discourse is produced and exchanged.
- Telegram as a Political Communication Platform: Telegram channels broadcast administrator-selected posts chronologically to unlimited subscribers without intermediary social connections or engagement-based ranking.This architecture makes channel content an unusually direct record of deliberate political communication.
- Sentiment Analysis: Sentiment analysis classifies subjective orientation as positive, negative, or neutral, but emotional valence does not necessarily indicate political stance.A message may express negative sentiment while supporting an actor or describe an adversary’s defeat neutrally.
- Stance Detection: Stance detection relates text to a target, requiring inference from vocabulary, actor references, and framing rather than explicit statements alone.In this study, messages receive pro-Palestine, pro-Israel, or neutral labels toward the Israel–Palestine conflict, with 736 hand-labeled messages used for evaluation.
- Transformer-Based Methods: The study uses zero-shot DeBERTa through natural-language-inference capabilities and fine-tuned transformer models adapted to social-media language.DeBERTa’s disentangled attention and enhanced mask decoder support its NLI performance, while fine-tuning updates model parameters on labeled data.
- Political Stance Analysis and Dataset Challenges: Conflict-related Telegram stance is difficult because reportorial language, multi-actor references, and unstable neutral boundaries embed political positioning indirectly.This instability anticipates weak neutral-class performance across the three detection methods.
3 Related Work
Prior research establishes Telegram as important conflict infrastructure and documents extensive social-media contestation over Israel–Palestine, while relying mainly on cross-platform or Twitter-focused analyses. The present study addresses the resulting gap through a dedicated longitudinal comparison of opposing Telegram communities using complementary computational and framing methods.
- Review Scope: Related work is organized around Telegram, Israel–Palestine social-media studies, sentiment analysis, stance detection, and narrative analysis.The review compares earlier studies’ approaches, data, and findings before stating the gap addressed by the present analysis.
- Telegram as Conflict Infrastructure: Telegram research has treated the platform as moderation-resistant infrastructure and as a venue for politically motivated communication.Earlier work documented migration from Twitter to Telegram and examined the platform’s distinctive communication architecture.
- Israel–Palestine Social-Media Research: Earlier Israel–Palestine social-media studies examined large Twitter corpora, activist and witness discourse, partisan volume, and humanitarian framing.These studies predate the dedicated Telegram comparison undertaken here.
- Stance Detection: Across stance-detection studies, fine-tuned in-domain transformers generally outperform general-purpose transformers and lexicon- or feature-based methods.A cited benchmark reports a fine-tuned BERTweet model at macro-F1 74.8%, narrowly ahead of a DeBERTa-based system at 73.9%.
- Narrative and Framing Analysis: Narrative and framing research explains how conflict actors construct meaning and identity across system, national, issue, and community levels.This tradition motivates aggregate rather than purely message-level interpretation of recurring framing patterns.
- Research Gap: Existing Telegram studies of Israel–Palestine have mainly used cross-platform designs or focused on the period after October 2023.The present work is described as the first dedicated longitudinal NLP comparison of ideologically opposed Telegram communities using multiple complementary methods.
4 Dataset
The study analyzes publicly accessible Telegram broadcasts from ideologically aligned pro-Palestine and pro-Israel channels, covering multiple conflict-related contexts. After balanced collection and preprocessing, the corpus was reduced to 87,617 messages, alongside a manually labeled evaluation sample.
- 4 Dataset: The dataset included broader regional conflict content because axis-of-resistance channels presented Gaza alongside related military operations as one unified conflict.This scope became especially pronounced after the April 2024 Damascus consulate strike and subsequent Iranian missile exchanges.
- 4 Dataset: Sixteen manually selected Telegram channels comprised eight pro-Palestine and eight pro-Israel sources.Selection emphasized subscriber count, posting frequency, and clear ideological alignment.
- 4 Dataset: Approximately 47,000 messages were collected from each stance group to avoid structural sampling imbalance.The crawler retrieved posts in reverse chronological order, beginning with the most recent messages.
- 4 Dataset: The final records covered May 2021 to June 2026 and included message text, timestamps, channel identifiers, post identifiers, and stance-group labels.Most messages were published after October 2023 because collection began with the newest available posts.
- 4 Dataset: 5,712 duplicate messages were removed, reducing the initial corpus from 94,290 to 88,578 messages.Normalization lowercased text, removed URLs and punctuation, and retained the first occurrence of exact duplicates.
- 4 Dataset: Manual annotation produced 736 messages labeled pro-Palestine, pro-Israel, or neutral for model training and evaluation.Because messages came from politically aligned channels, genuinely neutral content was comparatively uncommon.
5 Methodology
The methodology combines sentiment analysis, three stance-detection paradigms, and keyword-based framing analysis. Stance methods were evaluated against manually annotated messages using accuracy and macro F1.
- 5 Methodology: The pipeline applied sentiment analysis, keyword matching, zero-shot classification, fine-tuned stance detection, and framing analysis to the full dataset.Framing examined recurring themes including civilian casualties, military operations, and legal accountability.
- 5.2 Sentiment Analysis: The sentiment model was Twitter-RoBERTa, selected because Telegram posts resemble short, informal, emotionally charged tweet-style text.It was fine-tuned on the TweetEval sentiment benchmark.
- 5.3 Keyword Stance Detection: Keyword stance detection scored messages against manually constructed pro-Palestine and pro-Israel lexicons.Word-boundary matching and longest-match precedence reduced substring false positives and resolved multi-word terms to one side.
- 5.4 Zero-Shot Stance Detection: Zero-shot stance detection used DeBERTa-v3-large natural language inference to assign labels without task-specific training examples.The selected checkpoint was fine-tuned for natural language inference, the basis of the zero-shot approach.
- 5.5 Fine-Tuned Stance Detection: Fine-tuned stance detection adapted BERTweet, a social-media language model pretrained on approximately 850 million English tweets.The model was trained on the manually labeled data using 5-fold stratified cross-validation and a three-class output layer.
- 5.6 Framing Analysis: Framing classified death, victim, military, and legal language with keyword categories that could overlap within a message.The resulting assignments were analyzed alongside sentiment and stance over time.
- 5.7 Evaluation: Accuracy and macro F1 evaluated all three stance methods against 736 manually annotated messages.Macro F1 gave equal importance to the pro-Palestine, pro-Israel, and neutral classes despite unequal class counts.
6.1 Sentiment Analysis
Sentiment across the Telegram corpus was predominantly neutral, with negative messages forming the main non-neutral category and positive messages remaining uncommon.
- 6.1 Sentiment Analysis: 64% of the 87,617 messages were neutral, compared with 32% negative and approximately 3.5% positive.The distribution was calculated across the full dataset using the Twitter-RoBERTa sentiment model.
6.2 Evaluation of Stance Detection Methods
The fine-tuned BERTweet model performed best on manually annotated Telegram messages, while keyword and zero-shot methods commonly defaulted to neutral when stance signals were implicit.
- Overall comparison: 72.1% accuracy and 0.721 macro F1 were achieved by fine-tuned BERTweet under 5-fold cross-validation.Evaluation used the same 736 manually annotated messages for all three methods.
- Keyword-based method: 60.6% accuracy and 0.608 macro F1 were achieved by the keyword method, whose dominant error was assigning stance-bearing messages to neutral.The method correctly classified 156 pro-Palestine, 134 pro-Israel, and 156 neutral messages.
- Overall comparison: 11.3 points and 7.8 points were the fine-tuned model’s macro F1 advantages over keyword and zero-shot methods, respectively.The keyword and zero-shot methods reached macro F1 scores of 0.608 and 0.643.
- Zero-shot DeBERTa: 43.9% pro-Palestine recall was achieved by zero-shot DeBERTa, whose errors mainly involved neutral rather than cross-stance predictions.Pro-Israel and neutral recall were 70.4% and 80.9%, respectively.
- Fine-tuned BERTweet: 75.1% pro-Palestine recall and 73.1% pro-Israel recall were achieved by the fine-tuned model, compared with 67.8% neutral recall.The neutral class remained the weakest of the three classes.
- Fine-tuned BERTweet: Fine-tuning adapted BERTweet to the dataset’s political vocabulary, framing patterns, and contextual signals, recovering stance that the other methods often left neutral.The model used examples from the same Telegram channels and conflict context as evaluation.
6.3 Inter-Method Agreement
The three stance methods disagreed on roughly 40% of the full corpus, with the greatest disagreement involving pro-Palestine labels assigned by the keyword method.
- Overall agreement: 61.3% agreement occurred between keywords and zero-shot, 60.2% between keywords and fine-tuned, and 62.7% between zero-shot and fine-tuned.Agreement was calculated across all 87,617 messages.
- Pro-Palestine disagreement: 41.9% of keyword-assigned pro-Palestine messages were reassigned to neutral by zero-shot, while the two methods agreed on 37.4%.The fine-tuned model agreed with keywords on pro-Palestine labels in 54.6% of cases.
6.4 Stance Distribution Across the Full Dataset
The methods produced substantially different stance distributions across the full dataset, especially in their proportions of neutral and explicit political labels.
- Neutral and explicit stance shares: 41.5% of messages were neutral under the fine-tuned model, versus 55.9% under keywords and 58.1% under zero-shot.The fine-tuned model assigned 27.9% pro-Palestine and 30.6% pro-Israel labels.
- Explicit stance shares: 30.6% pro-Israel and 27.9% pro-Palestine labels gave the fine-tuned model the most balanced distribution across explicit stance classes.Keywords assigned 15.8% pro-Israel, while zero-shot assigned 14.5% pro-Palestine.
6.5 Sentiment by Stance-All Methods
Across all three methods, neutral messages were least negative, while explicitly pro-Palestine or pro-Israel messages were more negative; pro-Palestine messages were consistently more negative than pro-Israel messages.
- Cross-method pattern: Explicitly pro-Palestine and pro-Israel messages carried more negative sentiment than neutral-stance messages under every method.Neutral messages had the highest neutral-sentiment share and the lowest negative-sentiment share.
- Cross-method pattern: Approximately 54% of zero-shot pro-Palestine messages were negative versus 42% neutral, making it the only stance with more negative than neutral sentiment in that aggregation.Positive sentiment never exceeded roughly 5% across categories.
- Death and victim framing: 62% negative sentiment appeared in death-context pro-Palestine messages versus 45% for pro-Israel, while victim framing showed 58% versus 35%.The emotional-register gap was larger in these framing subsets than across all messages in each stance group.
6.6 Framing Analysis
Framing frequencies alone suggest pro-Israel channels use death and victim categories more often, but sentiment reveals opposite emotional registers: predominantly neutral pro-Israel reporting versus markedly negative pro-Palestine discourse.
- Framing rates: 51.8% of pro-Israel messages and 41.1% of pro-Palestine messages used military framing.Rates are computed within fine-tuned stance groups and exclude stance-neutral messages.
- Framing rates: 27.8% of pro-Israel messages used death-context framing versus 17.7% of pro-Palestine messages.The same direction held for victim framing, at 19.8% versus 11.9%.
- Sentiment within framing: 54% of pro-Israel death-context posts were neutral versus 45% negative, while approximately 62% of pro-Palestine posts were negative.The emotional contrast reverses the frequency pattern observed for death-context framing.
- Sentiment within framing: Approximately 58% of pro-Palestine victim-framing messages were negative compared with approximately 35% of pro-Israel messages.Thus, overlapping death- and victim-related vocabulary appears in different emotional registers.
6.7 Event Analysis
The event analysis compares sentiment and stance around five major milestones using ±14-day windows, but sparse coverage and overlapping aftermaths constrain interpretation of some events.
- Method and scope: Five major milestones were analyzed using ±14-day windows because they had sufficient message volume for day-level comparison.Earlier events contain fewer messages because collection proceeded in reverse chronological order.
- Method and scope: Event comparisons were descriptive, with formal significance testing confined to corpus-level framing comparisons.Per-event significance tests were not reported because many events, groups, and windows were compared.
- Israel strikes Iranian air defense: The October 25, 2024 strike produced low day-0 volume and no pronounced event-date neutral spike across groups.Its window overlapped the October 1 escalation aftermath, confounding pre/post sentiment comparisons.
- Israel strikes Iranian air defense: Pro-Israel sentiment rose from 35% to 47% in the first week before reverting to 34%, while pro-Palestine negativity remained 42%, 42%, and 45%.The isolated shift was not interpreted because the contaminated baseline could not distinguish residual movement from an event response.
- Israel and US attack Iran: The February 28, 2026 attack produced the dataset’s largest single-day volume increase, with 553 neutral, 194 pro-Israel, and 136 pro-Palestine messages.The neutral surge reflected descriptive reporting rather than a neutral community, while pro-Palestine messages showed stronger explicit stance expression.
- Cross-event comparison: Around major escalations, negativity sometimes rose jointly but the October 1 one-sided strike produced mirror-image movement between the communities.Around February 28, pro-Israel negativity rose immediately while pro-Palestine negativity rose sharply by the second week; around June 13 both shifted upward slightly.
7 Discussion
The discussion finds that domain-adapted stance detection outperforms label-free methods and that framing-sensitive sentiment distinguishes the communities’ discourse registers, while dataset and measurement limits bound the conclusions.
- 7.1 Model Performance and Method Selection: 72.1% accuracy and 0.721 macro F1 were achieved by fine-tuned BERTweet under 5-fold stratified cross-validation.The model outperformed zero-shot DeBERTa and keyword matching, which reached 64.5% and 60.6% accuracy respectively.
- 7.1 Model Performance and Method Selection: Fine-tuning on 736 domain-specific labeled examples outperformed both lexicon-based and zero-shot approaches for politically polarized Telegram content.The cross-validation estimate was 0.721 ± 0.027 across folds.
- 7.2 Interpretation of Key Findings: Across all three methods, explicitly political messages carried substantially more negative sentiment than messages classified as neutral.The pattern held regardless of which method assigned the stance label.
- 7.2 Interpretation of Key Findings: Pro-Israel messages used death-context and victim framing more often, but pro-Palestine messages within those categories were more negative.Death-context rates were 27.8% versus 17.7%, while victim-framing rates were 19.8% versus 11.9%; sentiment was approximately 62% negative for pro-Palestine death-context posts versus 45% for pro-Israel.
- 7.3 Limitations: The corpus covers only 16 politically aligned channels, with earlier periods underrepresented by reverse-chronological collection.The authors state that stronger longitudinal claims would require a larger and more temporally balanced dataset.
- 7.3 Limitations: The study’s conclusions do not extend to general news channels, multilingual content, or communities between the two political poles.The channel selection was intended to compare opposing communities rather than map the full conflict-discourse environment.
- 7.4 Broader Context: The channels function as closed information environments because subscribers self-select into streams consistently framed from one political position.This context links the observed discourse to echo-chamber research on reinforcing content exposure.
8 Conclusion
The study finds that interpreting Telegram discourse requires reading sentiment, stance, and framing together: both communities use death- and victim-related vocabulary, but in different emotional and narrative registers.
- 8 Conclusion: 72.1% accuracy and 0.721 macro F1 made fine-tuned BERTweet the strongest stance detector, outperforming keyword and zero-shot DeBERTa methods by roughly 8 to 11 points.The estimate came from 5-fold cross-validation over 736 manually labeled messages.
- 8 Conclusion: Pro-Israel channels used death and victim vocabulary more often but with predominantly neutral sentiment, whereas pro-Palestine channels expressed substantially more negative sentiment around the same topics.The interpretation combines framing-vocabulary rates with sentiment register rather than relying on separately validated frame analysis.
- 8 Conclusion: The contrasting registers are consistent with pro-Israel channels writing more often as reporting actors and pro-Palestine channels writing more often as affected parties.This distinction becomes visible when stance, sentiment, and framing are examined together rather than in isolation.