Source-linked AI summary
Quantifying Information Overload in Social Media and its Impact on Social Contagions
Manuel Gomez Rodriguez, Krishna Gummadi, Bernhard Schoelkopf
TL;DR
Information overload arises when social-media users receive information faster than they can process it, raising questions about its effects on dissemination. The paper models Twitter users as queueing systems and uses timestamped activity to estimate processing behavior and limits. It finds bounded production but unbounded receipt, a roughly 30 tweets/hour overload threshold, and reduced forwarding and cascade propagation under heavier information inflow.
Problem
Social media exposes users to information at rates that can exceed their cognitive processing capacity, motivating quantitative study of overload and dissemination.
Method
The paper models Twitter users as LIFO information-processing queues and reverse engineers their behavior from timestamps of received and forwarded tweets.
Results
The study finds bounded information production but unbounded receipt, with retweet probability falling sharply above ∼30 tweets/hour.
Takeaways & Limitations
Higher information inflow reduces forwarding effectiveness and can make large social cascades rarer.
Takeaways & Limitations
The analysis assumes stationary information-flow rates and a static social graph.
Abstract
from arXiv · showhide
Information overload has become an ubiquitous problem in modern society. Social media users and microbloggers receive an endless flow of information, often at a rate far higher than their cognitive abilities to process the information. In this paper, we conduct a large scale quantitative study of information overload and evaluate its impact on information dissemination in the Twitter social media site. We model social media users as information processing systems that queue incoming information according to some policies, process information from the queue at some unknown rates and decide to forward some of the incoming information to other users. We show how timestamped data about tweets received and forwarded by users can be used to uncover key properties of their queueing policies and estimate their information processing rates and limits. Such an understanding of users' information processing behaviors allows us to infer whether and to what extent users suffer from information overload. Our analysis provides empirical evidence of information processing limits for social media users and the prevalence of information overloading. The most active and popular social media users are often the ones that are overloaded. Moreover, we find that the rate at which users receive information impacts their processing behavior, including how they prioritize information from different sources, how much information they process, and how quickly they process information. Finally, the susceptibility of a social media user to social contagions depends crucially on the rate at which she receives information. An exposure to a piece of information, be it an idea, a convention or a product, is much less effective for users that receive information at higher rates, meaning they need more exposures to adopt a particular contagion.
Introduction
The paper frames information overload as a growing social-media problem and studies it by modeling Twitter users as information-processing systems. Timestamped tweets and retweets reveal users’ queueing behavior, processing limits, and overload.
- Social media sharply increases information exposure, and surveys report that two thirds of Twitter users feel they receive too many posts.
- The study reverse engineers users’ information-processing behavior from timestamps of received and forwarded tweets.
- Users are modeled as LIFO queues that process incoming information at unknown rates and forward some processed items.
- The analysis estimates queueing policies and processing limits while assessing how behavior varies with incoming information rates.
- The study finds bounded information production, unbounded information receipt, and a roughly 30 tweets/hour threshold associated with sharply lower retweet probability.
Related Work
Prior work examines how social ties and competing contagions shape information exchange. This paper instead emphasizes users’ processing limits and the effects of simultaneous background traffic on propagation.
- Related studies analyze attention across contacts, propagation from highly connected users, and finite communication capacity.
- This paper differs by identifying a sharp processing threshold that can reveal overloaded users, unlike the absence of a sharp threshold for distributing attention across ties.
- Prior contagion models study pairwise interactions or lack experimental validation, whereas this work examines many simultaneously propagating contagions with background traffic.
Methodology
The methodology treats Twitter users and feeds as information-processing systems and reconstructs queue behavior from observational timing and network data. The dataset covers public Twitter activity and follow relationships from 2009.
- The study infers processing behavior from timestamps of information receipt and forwarding without directly observing reading or processing.
- Users are modeled with persistent LIFO queues, where feeds contain tweets from followed users and processed items may be retweeted.
- Twitter’s 2009 inverse-chronological feed ordering permits reconstruction of a user’s queue at any given time.
- The framework estimates queue sizes, processing delays, and source-prioritization behavior from generated and forwarded information times.
- The Twitter data include 52 million user profiles, 1.9 billion directed follow links, and 1.7 billion public tweets.
Information Processing Limits
The analysis finds limits on information production but not receipt, with overload emerging when incoming information exceeds users’ processing capacity. Higher inflow reduces forwarding and can weaken information dissemination.
- Limits on information generation & forwarding: ∼40 tweets/day marks a sharp fall-off in total information produced by a user, providing evidence of a daily production limit.
- Limits on information generation & forwarding: Retweet counts follow a power-law distribution, unlike total tweet production, suggesting distinct policies for publishing and forwarding information.
- Origins of information overload: 50% of users receive fewer than 50 tweets/day, while 10% receive more than 500 tweets/day, with no apparent limit on tweet inflow.
- Origins of information overload: Tweet inflow rises linearly with followee count, indicating that oversubscription contributes to receiving more information than users can process.
- Evidence of information overload: Above ∼30 tweets/hour, retweet probability falls as βr ∝λ^-0.65, identifying a regime associated with information overload.
- Evidence of information overload: Higher background traffic can strongly reduce information dissemination and make large social cascades rarer.
User Processing Behaviors
The paper analyzes how Twitter users process incoming information through queues, focusing on queue position, delays, and source prioritization as in-flow increases. It finds that overload shifts processing away from recently received tweets, changes queueing delays, and encourages selective attention to influential sources.
- Queue position vs. processing probability: Users process tweets asynchronously from LIFO feeds, and the analysis asks how queue position, processing delays, and source priorities vary with incoming information.The study estimates these behaviors from timestamped tweet activity and the social graph.
- Queue position vs. processing probability: At 30−1000 tweets/hour, retweets shift down the queue, but tweets beyond the first 100 slots are retweeted at more than an order of magnitude lower probability than those in the first 10.The results suggest bounded effective queues: once information moves beyond the top few positions, processing probability drops sharply.
- Queue position vs. processing probability: At approximately 30 tweets/hour, average and median retweet queue positions rise sharply, coinciding with the threshold at which users begin experiencing information overload.When overloaded, retweets are no longer drawn from queues and may be found through other mechanisms.
- Queueing delays for forwarded information: Higher tweet in-flow shifts queueing-delay peaks earlier: users receiving 5−10 tweets/hour peak at 5 minutes, versus less than 2 minutes at 100−200 tweets/hour.The delay distributions share a family well fit by a convolution of observation-time and reaction-time lognormal distributions, while their mean, variance, and peak vary with in-flow.
- Queueing delays for nonforwarded information: Using Little’s Theorem, the analysis derives a lower bound on non-retweeted reading time from unread-queue size, in-flow rates, and retweeted-item delay.Because ∆nr and ∆* exceed ∆r, the paper concludes that retweeted tweets are those users happen to read and decide on earlier.
- Priority queueing to select sources: As in-flow rises, the retweet source set grows sub-linearly, indicating diminishing returns and more selective processing of influential users during overload.The paper describes overloaded users as prioritizing influential sources at the expense of remaining users.
- Priority queueing to select sources: Exposure curves show that lower-in-flow users have larger adoption-probability increases for hashtags, conventions, and URL-shortening services.Users receiving information at higher rates are less susceptible to each exposure across these contagion types.
Impact on Social Contagions
The paper examines how background information traffic affects adoption probabilities and cascade size and duration across social contagions on Twitter and synthetic networks.
- Exposure curves: Users with smaller in-flow rates show much larger increases in hashtag-adoption probability after exposure.Maximum adoption probabilities differ by one order of magnitude between users with small and large in-flows.
- Exposure curves: Exposure to the “RT” convention is much less effective for users with larger in-flows.
- Exposure curves: Hashtags, retweet conventions, and bit.ly adoption are evaluated with exposure curves grouped by users’ in-flow rates.
- Cascade sizes: Larger background traffic produces shorter cascades in the extended independent cascade model.
- Cascade sizes: For µ = 1, more than 12% of cascades reach at least 3 nodes, compared with 0.1% for µ = 100.
- Cascade duration: When µ > 10, longer cascades emerge even as quicker-extinguishing cascades become more frequent with increasing average out-flow rate.The continuous-time model links prolonged lifetimes to overloaded users delaying diffusion by searching profiles or sorting incoming tweets.
Conclusions
The study quantitatively characterizes information processing and overload in social media, showing that information-receipt rates strongly affect users’ susceptibility to social contagions. Its scope is limited by stationary-flow and static-network assumptions and by exclusive reliance on Twitter data.
- Contributions: The study characterizes how frequently, from how many sources, and how quickly social media users forward information.
- Contributions: Users’ information-processing limits and susceptibility to social contagion depend dramatically on their information-receipt rates.
- Limitations and future work: The analysis assumes stationary information-flow rates and a static social graph, while relying exclusively on Twitter data.The authors propose extending the work to dynamic flows, time-varying networks, and other platforms.