Source-linked AI summary

Analyzing Polarization in Social Media: Method and Application to Tweets on 21 Mass Shootings

Dorottya Demszky, Nikhil Garg, Rob Voigt, James Zou, Matthew Gentzkow, Jesse Shapiro, Dan Jurafsky

arXiv:1904.01596v2cs.CL

TL;DR

The paper asks how political polarization is expressed linguistically on social media beyond topic differences. It develops a framework for topic choice, framing, affect, and illocutionary force, and applies it to 4.4M tweets about 21 mass shootings. The study finds substantial polarization, with partisan framing differences playing the central role described in the supplied abstract.

  • Problem

    Prior work identifies polarized messages and topics but lacks a broad account of the ways polarization can be instantiated linguistically.

  • Method

    The paper combines lexical measures for four linguistic dimensions with embedding-based tweet clustering to identify salient, cross-event topics.

  • Results

    The discussion of mass shootings is highly polarized, with leave-out partisanship ranging from .517 to .547 across events and increasing after events during the first 10 days.

  • Takeaways & Limitations

    Polarization in these discussions is primarily associated with partisan framing differences rather than topic choice, while Republicans emphasize shooters and event facts and Democrats emphasize victims and policy change.

  • Takeaways & Limitations

    The studied events do not fully disentangle shooter race from other factors because school and worship-place shootings overwhelmingly involve white perpetrators.

Abstract

from arXiv · show

We provide an NLP framework to uncover four linguistic dimensions of political polarization in social media: topic choice, framing, affect and illocutionary force. We quantify these aspects with existing lexical methods, and propose clustering of tweet embeddings as a means to identify salient topics for analysis across events; human evaluations show that our approach generates more cohesive topics than traditional LDA-based models. We apply our methods to study 4.4M tweets on 21 mass shootings. We provide evidence that the discussion of these events is highly polarized politically and that this polarization is primarily driven by partisan differences in framing rather than topic choice. We identify framing devices, such as grounding and the contrasting use of the terms "terrorist" and "crazy", that contribute to polarization. Results pertaining to topic choice, affect and illocutionary force suggest that Republicans focus more on the shooter and event-specific facts (news) while Democrats focus more on the victims and call for policy changes. Our work contributes to a deeper understanding of the way group divisions manifest in language and to computational methods for studying them.

1 Introduction

The paper addresses the limited understanding of how polarization is linguistically instantiated on social media. It introduces a framework spanning topic choice, framing, affect, and illocutionary force, applied to tweets about mass shootings.

  • Prior NLP studies identify polarization in message sharing and topics but do not broadly capture its linguistic manifestations.
  • The framework analyzes four dimensions of linguistic polarization: topic choice, framing, affect, and illocutionary force.
  • The study examines more than 4.4M tweets about 21 mass shooting events, analyzing polarization within and across events.
  • The social-media focus extends prior framing research centered largely on news media and politicians.
  • 1.2 The Role of the Shooter’s Race: The paper considers shooter race as a potential factor but cannot fully disentangle race from other characteristics because some event types overwhelmingly involve white perpetrators.

2 Data: Tweets on Mass Shootings

The dataset comprises tweets collected for 21 mass shootings between 2015 and 2018, filtered for event relevance and sufficient volume. User partisanship is inferred from followed political accounts and validated against state voting patterns.

  • The authors compile mass shootings between 2015 and 2018 from the Gun Violence Archive and retrieve tweets from an archived Twitter firehose.
  • Tweets cover the two weeks after each event and require both a location keyword and a shooting-related lemma; retweets and deactivated users are removed.
  • The analysis retains 21 events with more than 10,000 tweets remaining after filtering.
  • Users are labeled Democrat or Republican according to whether they follow more politicians from one party than the other.
  • 51–72% of users per event receive partisan labels, and inferred state-level Republican shares correlate with two-party vote shares.

3 Quantifying Overall Polarization

The study measures partisan language using a leave-out estimator of phrase partisanship and finds substantial polarization across mass-shooting discussions. Polarization generally increases after events, while the temporal pattern is not explained by users who tweet on multiple days.

  • 3.1 Methods: Partisanship is the expected posterior probability of correctly inferring a tweeter’s party from one randomly drawn token, with .5 indicating no partisan difference in token usage.
  • 3.1 Methods: The estimator combines between-group differences in token posteriors with within-group similarity between each user and their party.
  • 3.1 Methods: The leave-out estimator excludes the speaker and tokens used by fewer than two speakers when computing empirical posterior probabilities.
  • 3.2 Results and Discussion: .517 to .547: leave-out partisanship across events indicates that discussions of each mass shooting are highly polarized.
  • 3.2 Results and Discussion: The Fort Lauderdale event is excluded from leave-out experiments because only its first post-event day is available, making it incomparable.
  • 3.2 Results and Discussion: slope = .002, p < 0.05: post-event polarization increases across events during the first 10 days after shootings.
  • 3.2 Results and Discussion: ∼10% of users tweeted on multiple days but contributed ∼28% of tweets; removing them preserves the temporal patterns with the same statistical significance.

4 Topics and Framing

The paper develops embedding-based topics that are comparable across mass-shooting events, then separates polarization arising from topic choice from polarization arising within topics. Across events, within-topic framing is more polarized than topic choice, with partisan topic preferences also differing systematically.

  • Topic assignment: Embedding-based clustering is designed to identify salient, comparable topics across events whose discourse is tied to event-specific details.The approach is compared with MALLET and Biterm Topic Model baselines trained on sampled tweets.
  • Topic assignment: The embedding-based model outperforms LDA-based methods on both word-intrusion and tweet-intrusion evaluations, especially tweet intrusion.The authors use k = 8 for subsequent analysis because it slightly outperforms other k values in tweet intrusion; model-level differences across k are not significant.
  • Measuring polarization: Within-topic partisanship measures how users discuss each topic, whereas between-topic partisanship measures party inference from topic assignments alone.Within-topic values are computed by reapplying the leave-out estimator to tweets in each topic and averaging by topic prevalence; between-topic values replace tweets with assigned topics.
  • Results: Within-topic polarization is higher than between-topic polarization for most events, increases over time, and supports distinguishing topic choice from topic-level framing.Between-topic polarization remains stable while within-topic polarization increases.
  • Results: Shooter identity and ideology (.55) and laws and policy (.54) are the most polarized topics on average, while news (.51), victims and location (.52), solidarity (.52), and remembrance (.52) are less polarized.For Las Vegas, solidarity has the lowest and shooter identity and ideology the highest polarization; most topics increase over time, with news showing the steepest increase.
  • Results: Across events, Republicans more often discuss investigation, news, and shooter identity and ideology, while Democrats more often discuss laws and policy and solidarity.The reported topic preferences suggest Republican-focused topics relate more to the shooter, whereas Democratic-focused topics relate more closely to victims.

5 Specific Framing Devices

The paper examines partisan framing through token choice and contextual grounding, finding that shooter race shapes how Democrats and Republicans describe mass shootings. Contrasting uses of “terrorist” and “crazy,” along with partisan historical references, are central framing devices.

  • 5.1 Methods: The analysis estimates token partisanship using event-level log odds ratios comparing Democratic and Republican vocabularies.Within-event and within-topic z-scores are used for token comparisons, while cross-event comparisons retain the original signs because verbosity affects magnitude ranges.
  • 5.2 Results: “Terrorist” is more Democratic for white shooters but more Republican for shooters of color, while “crazy” shows the reverse, weaker pattern.These terms exhibit differential partisan patterns across events grouped by the shooter’s race.
  • 5.2 Results: The complete partisan reversal associated with shooter race suggests that race strongly shapes how each party frames the shooter’s identity and mental health.The authors connect this pattern to prior work on binary racial conceptualization in the United States and call for further exploration.
  • 5.2 Results: Contextual grounding is also partisan: Democrats most often invoke Sandy Hook, Republicans 9/11, and Democrats more often reference school, worship, and white-shooter events.The contextual events used after shootings therefore differ by party as well as by the characteristics of the focal event.

6 Affect

The paper measures partisan affect with a domain-adapted emotion lexicon and finds systematic differences in emotional expression across parties and shooter-race contexts. Democrats more often express positive sentiment, sadness, and trust, while Republicans more often express fear and disgust, especially for shooters of color.

  • 6.1 Methods: The study adapts the NRC Emotion Lexicon to mass-shooting discourse using representative stems, embeddings, and label propagation.For each emotion category, the authors retain 30 stems closest to selected domain-relevant representatives in GloVe space.
  • 6.1 Methods: Partisan affect is measured by aggregating emotion-stem frequencies for each event and party, then calculating each category’s partisan log odds ratio.The analysis covers positive and negative valence plus disgust, fear, trust, anger, and sadness.
  • 6.2 Results: Positive sentiment, sadness, and trust are more Democratic across events, whereas fear and disgust are more Republican, particularly when the shooter is a person of color.Anger, trust, and negative sentiment are described as similarly likely across parties in the reported results.
  • 6.2 Results: The fear and disgust findings align with prior research associating higher fear and disgust sensitivity with conservative ideology.The paper presents this as agreement with existing literature rather than as a causal explanation of the observed Twitter patterns.

7 Modality and Illocutionary Force

The paper uses necessity modals to study illocutionary acts in post-shooting tweets, especially calls for action and expressions of mental state. Modals are disproportionately associated with Democrats and the laws-and-policy topic, supporting a focus on policy-change appeals.

  • 7.1 Methods: Modality concerns necessity and possibility, and the paper uses modal language to study calls for action, blame, emotion, and factual assertions.The analysis treats modal use as a lexical signal of the kinds of acts users perform through tweets.
  • 7.1 Methods: The study analyzes should, must, have to, and need to using partisan log odds ratios and annotations of 200 modal-containing tweets.Annotations distinguish calls for change or action from expressions of the user’s mental state.
  • 7.2 Results: Approximately 78% of the 200 modal uses express calls for change or action, while approximately 40% express the user’s mental state.The examples and annotations support the hypothesis that these modals are primarily used for calls to action.
  • 7.2 Results: Modals are over-represented in the laws-and-policy topic, suggesting that calls for policy change, especially gun control, dominate calls for action.The topic representation is computed from modal and topic tweet frequencies.
  • 7.2 Results: All four modals are more likely to be used by Democrats: have to mean −.39, must mean −.3, should mean −.18, and need to mean −.18.Each reported party difference is statistically significant at the level given in the passage; negative values indicate Democratic predominance under the paper’s convention.
  • 7.2 Results: The partisan pattern for should have is similar to should, but unlike should, it does not vary significantly with shooter race or presidential administration.The passage reports mean log odds of −.22 for should have and nonsignificant race and administration differences.

8 Conclusion

The paper presents a multi-faceted framework for studying polarization in social-media language and applies it to mass-shooting discussions. The results show substantial political polarization across linguistic dimensions, including topic, framing, affect, and illocutionary force.

  • Reactions to mass-shooting events are highly polarized politically, as shown by leave-out estimates of phrase partisanship.
  • The tweet-clustering approach produces cohesive topic representations that are robust to differences in event vocabularies and tweet counts.
  • Republicans preferentially discuss the shooter’s identity and ideology, investigation, and news, whereas Democrats preferentially discuss solidarity and policy-related topics.
  • Republicans express more fear and disgust, while Democrats express more sadness and positive sentiment, make calls for action, and assign blame.
  • The framework examines polarization through topic choice, framing, affect, and illocutionary force.
  • The measures provide convergent evidence of complex ideological division in public life.

A Data

The dataset combines event-specific Twitter collections, partisan user labels, and checks for data quality and foreign-account presence. Events are represented through location keywords, while Washington, DC is excluded from a sanity check because its Twitter population differs from its voting population.

  • Partisan tweets are distributed across events, with the distribution summarized in Figure 12.
  • The study identifies mass-shooting tweets using location-specific keywords for 21 events.The supplied event lists include terms such as “Orlando” and “pulse nightclub,” and “Pittsburgh” and “tree of life.”
  • The data properties and Russian-account breakdown are reported in Tables 3 and 4, while topic-word summaries for MALLET and BTM appear in Tables 5 and 6.
  • DC is excluded from the sanity check because it is not an official state and its Twitter user population is expected to differ from its voting population.Figure 13 reports DC values separately.
  • The analysis finds no substantial presence of Russian accounts after preprocessing, although the banned-account list may underestimate foreign influence.Orlando had one such account with 115 tweets, while four accounts in Vegas produced 70 tweets in total.

C.2 Topic Tweets

The topic examples span event-specific news, shooter identity and ideology, victims and solidarity, and policy or security responses. They also include partisan interpretations involving race, religion, mental health, and responsibility.

  • News-oriented topic tweets report shootings, fatalities, suspects, arrests, and other developing event facts.Examples include breaking reports from Annapolis, suspect updates from Kalamazoo, and reports from San Francisco and Colorado Springs.
  • Shooter-identity and ideology tweets discuss perpetrators’ backgrounds, motives, religion, race, and alleged extremism.Examples reference the Capital Gazette shooter, the Waffle House gunman, and alleged radicalization or hate crimes.
  • Several examples frame responsibility differently by contrasting mental illness, race, religion, security failures, and gun policy.Tweets describe shooters as “lone wolves,” invoke racialized or religious comparisons, and dispute whether shootings are primarily security or gun-law failures.
  • Policy and security tweets call for gun-control measures, school security, government action, or responses to mental-health concerns.Examples include calls for stronger gun-control laws, hardened school access, and action on mental health, security, and guns.
  • Solidarity tweets express condolences, prayers, memorial observances, and support for victims and affected communities.Examples include messages about Chattanooga, Pittsburgh, San Bernardino, Parkland, Roseburg, and Orlando.

D Topic Model Evaluation

The evaluation compares an embedding-based topic model with MALLET and BTM using crowdsourced word- and tweet-intrusion tasks. The tasks test whether words and tweets assigned to topics are cohesive and distinguishable.

  • The study compares MALLET, BTM, and an embedding-based model through crowdsourced topic-model evaluation tasks.
  • The word-intrusion task presents five words close to a topic and one word distant from it but close to another topic.Workers select the odd word out; 2,850 experimental items were created.
  • The tweet-intrusion task evaluates whether tweets assigned to one topic are more uniquely close to that topic than to a second topic.Workers select the odd tweet from sets containing three close tweets and one contrasting tweet.
  • The evaluation uses lexical seed sets covering positive, negative, sadness, disgust, anger, fear, and trust-related affect.

F.1 Most Partisan Phrases Overall

Partisan phrase usage differs across mass-shooting events, with Republicans often emphasizing shooters, attacks, security, and event facts, while Democrats more often emphasize victims, violence, and gun policy.

  • The analysis lists the 20 most partisan unigrams and bigrams occurring at least 100 times for each event, using log-odds z-scores to identify partisan language.Absolute z-scores greater than 2 are treated as significantly partisan.
  • Republican-associated phrases frequently reference shooters, attacks, policing, terrorism, Islam, Obama, or event-specific details.Examples include “gun free zone” in Chattanooga, shooter identities and investigation terms in Las Vegas, and terrorism-related terms in Orlando and San Bernardino.
  • Democrat-associated phrases frequently reference guns, violence, victims, families, grief, and gun-sensitivity or gun-control policy.Orlando prominently features victim- and solidarity-related terms, while Las Vegas prominently features gun-control and gun-violence terms.
  • The partisan lexicons vary substantially by event, reflecting event-specific language alongside recurring differences in attention to shooters, victims, violence, and policy.For example, Republican phrases in Fresno emphasize “Allahu akbar,” terrorism, and Islam, whereas Democrat phrases emphasize victims, families, and police.

G.2 Results

Pronoun and modal analyses connect partisan language with different orientations toward solidarity, policy, action, emotion, and third-person event descriptions. Democrats more often use personalized and collective-action language, while Republicans more often focus on third parties and use modals epistemically or in idiomatic constructions.

  • Pronouns: Democrats use first- and second-person pronouns more often across events, while “SheHe” is used similarly by both parties.Mean partisan log odds are I: −.26, We: −.26, You: −.13, They: −.06, and SheHe: .05.
  • Pronouns: Pronoun differences vary with shooter race: Democrats use “SheHe” and “You” more for white shooters, whereas Republicans use them more for shooters of color.The reported differences are significant for SheHe at p < 0.01 and You at p < 0.05.
  • Pronouns and topics: “I” is concentrated in solidarity and other topics, while “We” is overrepresented in laws & policy, linking Democratic language with personal feeling and collective action.These patterns are interpreted alongside topic preferences and the modal analysis.
  • Pronouns and topics: “SheHe” is most frequent in investigation, shooter’s identity & ideology, and victims & location, topics more likely to be discussed by Republicans.The result suggests Republican tweets more often focus on a third person, such as the shooter.
  • Modal collocations: Democrats more often use necessity modals for collective or proactive action and emotional expression, whereas Republicans more often use them epistemically or idiomatically.Democratic examples include “we need to act” and “something needs to be done”; Republican examples include “it must have been” and “I have to say.”

I Results: Additional Plots

The additional plots visualize partisan differences in necessity-modal usage and topic polarization over time, including topic prevalence across the Orlando timeline.

  • Figure 16 plots the log odds ratio of necessity modals.
  • Figure 17 plots Orlando topic polarization over time using leave-out phrase partisanship, with bars showing each topic’s proportion of the data at each time.
Loading 1904.01596v2…