Source-linked AI summary
Influence and Passivity in Social Media
Daniel M. Romero, Wojciech Galuba, Sitaram Asur, Bernardo A. Huberman
TL;DR
Social-media influence depends not only on attracting attention but also on overcoming audiences’ passivity in forwarding content. The paper introduces an influence–passivity algorithm based on information forwarding activity. Its influence score predicts URL clicks, outperforms several conventional measures, and shows that popularity and influence need not coincide.
Problem
The paper addresses how to identify influential users when popularity, propagation, and audience passivity jointly affect how much attention content receives.
Method
The paper develops an efficient HITS-like algorithm that jointly scores users’ influence and passivity from network structure and information diffusion.
Results
The influence model outperforms PageRank, H-index, follower count, and retweet count, and predicts the upper bound on URL clicks.
Takeaways & Limitations
High popularity does not necessarily imply high influence, so influence assessment should account for audience passivity and propagation behavior.
Takeaways & Limitations
The evaluation uses fixed 300-hour windows, although influence changes over time and timestamp-aware modeling remains necessary.
Abstract
from arXiv · showhide
The ever-increasing amount of information flowing through Social Media forces the members of these networks to compete for attention and influence by relying on other people to spread their message. A large study of information propagation within Twitter reveals that the majority of users act as passive information consumers and do not forward the content to the network. Therefore, in order for individuals to become influential they must not only obtain attention and thus be popular, but also overcome user passivity. We propose an algorithm that determines the influence and passivity of users based on their information forwarding activity. An evaluation performed with a 2.5 million user dataset shows that our influence measure is a good predictor of URL clicks, outperforming several other measures that do not explicitly take user passivity into account. We also explicitly demonstrate that high popularity does not necessarily imply high influence and vice-versa.
1. INTRODUCTION
Social media content competes for scarce attention, making popularity and actual propagation central to understanding influence. The paper models influence together with user passivity, distinguishing reach from the difficulty of getting audiences to forward content.
- Content competes for scarce user attention as social media enables massive-scale generation and consumption.
- Popularity reflects received attention through followers, whereas influence reflects the actual propagation of a user’s content.
- User passivity creates a barrier to propagation that influential users must overcome.
- The proposed model combines network structure and diffusion behavior to score influence while accounting for audience passivity.
- The model outperforms PageRank, H-index, follower count, and retweet count, predicts URL-click upper bounds, and finds popularity and influence weakly correlated.
2. RELATED WORK
Prior research examined network structure, audience affinity, influential users, and Twitter propagation. These studies motivate distinguishing influence from activity and separating homophily from influence as explanations for information spreading.
- Earlier work studied how scale-free networks and member affinity affect information propagation.
- Research on key influentials focused on users responsible for broad information dissemination in social networks.
- Twitter studies examined sparse interaction structure, word-of-mouth advertising, and prediction of which users would tweet particular URLs.
- Related work found that influential bloggers were not necessarily the most active and distinguished homophily from influence as drivers of propagation.
3. TWITTER
Twitter is a large directed microblogging network where users follow accounts, post short updates, and forward posts through retweets. The study samples URL-containing activity over a 300-hour period and constructs a user graph from observed URL publishers.
- Twitter had more than 105 million users in April 2010 and represents follow relationships as a directed social network.
- Tweets are short status updates, often containing personal information, news, or links to external content.
- Retweets forward another user’s post and help propagate posts and links through the Twitter community.
- The dataset contains a continuous 300-hour stream of 22 million URL-containing tweets, estimated at one-fifteenth of Twitter activity.
- The user graph includes observed URL publishers but excludes users without URL tweets and users with private streams.
- Each followed-user list was fetched only once, so graph changes during the observation period are not captured.
4. THE IP ALGORITHM
The IP algorithm models influence and passivity from users’ forwarding relationships, treating passivity as resistance to influence. It iteratively computes both scores on a weighted retweet graph, with influence depending on audience size and audience behavior.
- An average Twitter user retweets only one in 318 URLs, providing evidence of substantial user passivity.
- Passivity is measured using user retweeting rate and audience retweeting rate.
- The model treats influence as depending on both the quantity and quality of the audience a user influences.
- The IP algorithm assigns every user relative influence and passivity scores from pairwise influence information in a weighted directed graph.
- Influence depends on audience size, audience dedication, and passivity, while passivity depends on exposed influence and rejected influence.
- The algorithm iteratively computes influence and passivity simultaneously in a manner similar to HITS.
- The influence graph contains users who tweeted at least three URLs, with arcs representing observed retweets and weights based on retweet frequencies.
5. EVALUATION
The evaluation compares IP-influence with established user measures and tests whether influence predicts URL attention. IP-influence incorporates network structure and passivity, converges in tens of iterations, and better predicts URL popularity while remaining distinct from follower popularity.
- 5.1 Computations: The Influence-Passivity algorithm computes influence and passivity on a weighted graph of approximately 450k nodes and 1 million arcs, converging in tens of iterations.PageRank is computed on the reversed weighted graph for comparison.
- 5.2 Influence as a correlate of attention: Follower count, past retweets, PageRank, and H-index are relatively weak or poor predictors of the maximum clicks that URLs can receive.The IP algorithm differs from PageRank by accounting for the passivity of the people a user influences.
- 5.2 Influence as a correlate of attention: Influencing users who are difficult to influence is a better indicator of URL popularity than PageRank.
- 5.2 Influence as a correlate of attention: IP-influence predicts the maximum number of clicks a URL can receive, with 99.9% certainty bounding URLs from high-average-IP users below 100,000 clicks and low-average-IP users below 100 clicks.Clicks are evaluated using 3.2M Bit.ly URLs and the 99.9th percentile of clicks at each attribute value.
- 5.2 Influence as a correlate of attention: IP-influence is not well correlated with follower count, showing that many followers do not necessarily imply power to influence users to click a URL.
6. IP ALGORITHM ADAPTABILITY
The IP algorithm can use alternative influence graphs when explicit retweet signals are unavailable. Co-mention and retweet graphs produce different influence rankings, with retweet-based scores predicting URL traffic slightly better while co-mentions surface different candidates.
- 6. IP ALGORITHM ADAPTABILITY: The algorithm accepts any influence graph, including a co-mention graph based on users who follow one another and mention URLs previously mentioned by those they follow.The co-mention edge weight uses the counts of URLs uniquely mentioned by one user and shared between the two users.
- 6. IP ALGORITHM ADAPTABILITY: Co-mention edges are less explicit than retweet edges and may connect users who do not influence one another, whereas retweet graphs may omit influence when reposts fail to credit the source.
- 6. IP ALGORITHM ADAPTABILITY: Retweet-based influence scores predict the maximum number of URL clicks better than co-mention-based scores, although co-mention scores remain predictive.
- 6. IP ALGORITHM ADAPTABILITY: The most IP-influential users with at least 10 posted URLs are dominated by news services, whose links are frequently forwarded by other users.
- 6. IP ALGORITHM ADAPTABILITY: Influence scores from co-mention and retweet graphs do not correlate well, so explicit and implicit influence signals can substantially change IP-algorithm outcomes.Retweet signals yield slightly better URL-traffic predictions, while co-mentions may identify a different set of influential users.
7. CASE STUDIES
The case studies show that IP-influence distinguishes users who spread content from users who merely attract attention, and identifies both highly passive accounts and influential users with few followers.
- Most influential: IP-influence rankings identify news services that post many links forwarded by other users.The ranking is constrained to users posting 10 URLs.
- Most passive: Highly passive users follow many people but retweet only a very small percentage of consumed information.Robots, suspended accounts, and extremely frequent posters appear among the most passive users.
- Most passive: The IP algorithm can automatically identify robot accounts, including aggregators and spammers, through their high passivity scores.Robots attend to many tweets but retweet only selected content, producing low forwarding percentages.
- Least influential with many followers: Users with many followers can have relatively low influence because their messages are consumed but not forwarded.These users attract substantial attention without spreading their messages far through the network.
- Most influential with few followers: Users with few followers can nevertheless achieve high influence through successful retweeting contests, high-quality drawings, or locally relevant content.The measure surfaces content that popularity rankings based on follower counts might miss.
8. DISCUSSION
The discussion presents IP-influence as useful for predicting attention, ranking content, and filtering likely spam, while noting unresolved questions about graph scale and changing influence over time.
- Influence as predictor of attention: IP-influence accurately predicts the upper bound on total URL clicks and can use activity-derived weighted graphs even without retweets or likes.The graph’s arc weights represent one user’s influence over another.
- Topic-based and group-based influence: Whether IP-influence remains equally accurate at different graph scales is an open question.The algorithm can be run on topic- or group-specific subgraphs, but scale-dependent accuracy remains unresolved.
- Content ranking: IP-influence can rank content likely to receive attention based on which users mention it early.The ranking can be computed for particular topics or user groups.
- Content filtering: High passivity is associated primarily with robots or spammers, motivating content filtering that limits feeds to influential users.The proposed extension aims to reduce spam in Twitter feeds.
- Influence dynamics: The study computes influence over a fixed 300-hour window, leaving open how influence spikes compare with sustained influence over time.A time-aware modification incorporating tweet timestamps may be needed.
9. CONCLUSION
The study finds that popularity is only weakly correlated with influence because information propagation requires active forwarding rather than passive reading. Its influence measure is not specific to Twitter and can identify users with broader reach regardless of popularity.
- The correlation between popularity and influence is weaker than expected.
- Information propagation requires users to forward content actively rather than read it passively.
- The influence measure is applicable beyond Twitter to other social networks.
- Influential individuals can have greater reach than others regardless of their popularity.