Source-linked AI summary

Virality Prediction and Community Structure in Social Networks

Lilian Weng, Filippo Menczer, Yong-Yeol Ahn

arXiv:1306.0158v2cs.SIcs.CYphysics.data-anphysics.soc-ph

TL;DR

The paper asks whether memes generally spread as complex contagions and whether early diffusion patterns can predict future popularity. It quantifies meme concentration across communities and uses those early features for prediction. Most memes resemble complex contagions, while viral memes spread like simple contagions and community structure provides predictive knowledge about virality.

  • Problem

    The paper examines whether memes generally spread as complex contagions and whether community structure can help predict which memes become viral.

  • Method

    The paper measures early meme concentration across communities and applies community-based prediction to distinguish viral from non-viral memes.

  • Results

    Most memes behave like complex contagions, whereas viral memes permeate many communities like simple contagions.

  • Takeaways & Limitations

    Community structure can be translated into predictive knowledge about what information will spread virally.

  • Takeaways & Limitations

    The authors report competing financial interests involving a social media analytics company co-founded by Y.-Y.A., who is also a major shareholder.

Abstract

from arXiv · show

How does network structure affect diffusion? Recent studies suggest that the answer depends on the type of contagion. Complex contagions, unlike infectious diseases (simple contagions), are affected by social reinforcement and homophily. Hence, the spread within highly clustered communities is enhanced, while diffusion across communities is hampered. A common hypothesis is that memes and behaviors are complex contagions. We show that, while most memes indeed behave like complex contagions, a few viral memes spread across many communities, like diseases. We demonstrate that the future popularity of a meme can be predicted by quantifying its early spreading pattern in terms of community concentration. The more communities a meme permeates, the more viral it is. We present a practical method to translate data about community structure into predictive knowledge about what information will spread widely. This connection may lead to significant advances in computational social science, social media analytics, and marketing applications.

Results

Most memes show community concentration and social reinforcement consistent with complex contagions, but viral memes permeate communities like simple contagions. Early community-based diffusion features predict meme virality better than community-blind or random baselines.

  • Communities and Communication Volume: People communicate more within their communities, and intra-community links carry more messages than inter-community links.The findings are statistically significant and robust across community-detection methods.
  • Meme Concentration in Communities: Non-viral memes exhibit concentration comparable to or stronger than reinforcement and homophily baselines, consistent with complex contagion.The baselines model structural trapping, social reinforcement, and homophily.
  • Meme Concentration in Communities: Viral memes match the simple-cascade concentration pattern, indicating that community structure traps them less and they permeate many communities.Their adoption pattern differs from that of most non-viral memes.
  • Strength of Social Reinforcement: Viral memes require as little social reinforcement as the simple-cascade model, whereas non-viral memes require as many exposures as the complex-contagion baselines.Reinforcement is measured from adopter exposures during the first 50 tweets.
  • Prediction: A random-forest model using early popularity and community features predicts virality with higher precision and recall than random guessing and community-blind prediction.The features include infected communities, usage and adoption entropy, and within-community interaction fractions, computed from the first 50 tweets.
  • Prediction: At θU = 90, the method is about seven times as precise as random guess and over three times as precise as community-blind prediction.Recall is over 350% better than random guess and over 200% better than community-blind prediction, with similar results across community-detection methods.

Discussion

The findings show that community structure plays an important role in meme spreading and can be translated into predictions of viral diffusion. This approach does not use message content and is relevant to social media and marketing applications.

  • Community structure has an important role in the spreading of memes, though the role of weak ties between communities is also recognized.
  • The study provides a direct approach for translating community-structure data into predictive knowledge about which information will spread virally.
  • The method does not exploit message content and can be applied to socio-technical networks using a small sample of data.
  • The results can be relevant for online marketing and other social media applications.
  • Further analysis of network community structure may support characterization and forecasting of social behavior.

Methods

The study analyzes sampled Twitter data with community structure identified from reciprocal-following networks, tests robustness across alternative network constructions and detection methods, and predicts meme popularity using community-concentration features.

  • The dataset comprises 121,807,378 tweets from 14,599,240 users containing at least one of 10,393,465 hashtags.
  • Researchers constructed an undirected, unweighted network from reciprocal-following relationships among 595,460 randomly selected users.This conservative network definition excludes link direction and weights.
  • Community structure was identified with Infomap, while link clustering produced similar results in robustness analyses.The network remained unweighted for community identification.
  • The analysis focused on new memes with fewer than 20 tweets during the preceding month, with hashtag-filtering sensitivity tests reported separately.
  • Prediction used a random forest with 500 decision trees and 10-fold cross-validation, while diffusion simulations replicated 10% API sampling across 100 × 10 samples.Each tree used 4 random features independently, and simulation outputs were averaged across the repeated samples.

Additional Information

The additional information includes figures describing community concentration, contrasting meme diffusion patterns, and prediction performance, alongside disclosures about competing financial interests.

  • Y.-Y.A. disclosed co-founding and being a major shareholder of Janys Analytics, while the other authors declared no competing financial interests.
  • Figure 1 illustrates structural trapping, social reinforcement, and the role of clustering in producing multiple exposures and cascades.
  • Figure 2 measures meme concentration through community edge weight and user community focus using retweets or mentions.Box plots show the central 50% and 95% whisker ranges, with medians and means marked.
  • Figure 3 relates usage and adoption dominance and entropy to meme popularity, averaging ratios across hashtags within popularity bins.Popularity is defined by tweet count T or adopter count U, with standard errors shown.
  • Figure 4 contrasts viral and non-viral meme evolution by representing communities as nodes sized by tweet production and colored by first-use timing.
  • Figure 5 evaluates viral-meme prediction at percentile thresholds θ = 70, 80, 90 using community-concentration features from the initial n = 50 tweets.The caption states that performance is robust across networks and community-detection methods.
Loading 1306.0158v2…