Source-linked AI summary
Is the Sample Good Enough? Comparing Data from Twitter's Streaming API with Twitter's Firehose
Fred Morstatter, Jürgen Pfeffer, Huan Liu, Kathleen M. Carley
TL;DR
Researchers lack clear evidence about whether Twitter’s sampled Streaming API represents overall activity adequately. This paper compares identically parameterized Streaming and Firehose datasets across statistical, topical, network, and geographic analyses, finding that Streaming API reliability depends strongly on coverage and analysis type. The results provide evidence to help researchers and institutions interpret Streaming API data.
Problem
The paper asks whether Twitter’s sampled Streaming API is a sufficient representation of overall Twitter activity despite limited documentation about its sampling.
Method
The study collects matched Streaming API and Firehose data and compares them through statistical, topic, network, and geographic analyses.
Results
Streaming API results depend strongly on coverage and analysis type, with topical accuracy highest at greater coverage and only 50–60% of top 100 network key-players identified on average from one day.
Takeaways & Limitations
The findings offer evidence to help researchers, business analysts, and government institutions better ground conclusions based on Twitter data.
Takeaways & Limitations
The authors hope to test whether the methodology generalizes to Twitter data from other domains, indicating that this scope remains unresolved.
Abstract
from arXiv · showhide
Twitter is a social media giant famous for the exchange of short, 140-character messages called "tweets". In the scientific community, the microblogging site is known for openness in sharing its data. It provides a glance into its millions of users and billions of tweets through a "Streaming API" which provides a sample of all tweets matching some parameters preset by the API user. The API service has been used by many researchers, companies, and governmental institutions that want to extract knowledge in accordance with a diverse array of questions pertaining to social media. The essential drawback of the Twitter API is the lack of documentation concerning what and how much data users get. This leads researchers to question whether the sampled data is a valid representation of the overall activity on Twitter. In this work we embark on answering this question by comparing data collected using Twitter's sampled API service with data collected using the full, albeit costly, Firehose stream that includes every single published tweet. We compare both datasets using common statistical metrics as well as metrics that allow us to compare topics, networks, and locations of tweets. The results of our work will help researchers and practitioners understand the implications of using the Streaming API.
Introduction
Twitter offers unusually open access to social-media data through a sampled Streaming API, but the sample’s validity for representing overall activity is unclear. This paper compares Streaming API data with the comprehensive Firehose across statistical, topical, network, and geographic analyses.
- Introduction: Twitter data is widely sought by computer and social scientists because the platform captures rapid communication and user behavior.The paper situates this interest in Twitter’s large user and tweet volumes and its role in major events.
- Introduction: Twitter’s Streaming API returns at most a 1% sample once matching tweets exceed 1% of all tweets, using keywords, geographic boxes, or user IDs.The API’s sampling method is undocumented.
- Introduction: Researchers must choose between the freely available but limited Streaming API and the costly, resource-intensive Firehose containing all public tweets.The Firehose requires substantial cost and storage, network, and server resources.
- Introduction: The study compares the two datasets using classic statistics, topic extraction, network measures, and geographic distributions of geolocated tweets.These analyses target how sampling affects common measures performed on Twitter data.
Related Work
Prior work uses Twitter samples for statistical, topical, network, and geolocation research, while sampling theory offers stronger guarantees for random data than for networks. These applications motivate direct comparison of sampled and complete Twitter datasets.
- Related Work: Researchers have used Twitter’s Streaming API for topic modeling, network analysis, and statistical analysis of content.The paper describes these examples as evidence of widespread reliance on the API.
- Related Work: Classical sampling theory supports approximating population statistics when a randomly selected subsample is sufficiently large, but network sampling is more complicated.The discussion invokes the law of large numbers and the Glivenko–Cantelli theorem for statistical sampling.
- Related Work: Prior Twitter studies examine hashtag counts, disaster-related topics, geographically relevant topics, and the geographic properties of users’ tweets.These studies span prediction, topic discovery, and geolocation applications.
- Related Work: The paper compares datasets across facets commonly used in the literature, including hashtags, topics, networks, and geolocation.The supplied figure caption identifies daily tweet-count comparison as one data-analysis view.
The Data
The study collected matched Streaming API and Firehose tweets about Syria and found that Streaming coverage varied substantially across days. The data collection used keywords, geographic boundaries, and users, with daily counts visualized for both sources.
- The Data: 528,592 Streaming API tweets and 1,280,344 Firehose tweets were collected from December 14, 2011, through January 10, 2012.Both sources used exactly the same keywords, geographic bounding boxes, and users.
- The Data: 43.5% is the average daily coverage of the Streaming API relative to the Firehose during the collection period.The paper defines coverage as Streaming API data divided by Firehose data.
- The Data: Coverage rates vary widely across days, with higher absolute or relative Twitter activity associated with lower coverage and fewer collected tweets.Figure 3 displays the daily coverage distribution with whiskers marking extreme values.
- The Data: The collection parameters target Syria and include keywords, geographic bounding boxes, and users.Table 1 specifies the collection parameters and identifies southwest and northeast boundary-box coordinates.
Statistical Measures
The paper compares Streaming API and Firehose data through hashtag rankings and LDA topic distributions, using correlation and divergence measures across coverage levels. Results examine whether sampled data reproduces patterns in the full dataset and how topic similarity changes with coverage.
- Top Hashtag Analysis: Kendall’s τβ compares the ordering of top hashtags between Streaming and Firehose datasets, accounting for concordant, discordant, and tied pairs.The statistic ranges from -1 for perfect negative correlation to 1 for perfect positive correlation.
- Top Hashtag Analysis: Figure 4 evaluates τβ for n from 10 to 1000 across five days representing minimum, quartile, median, and maximum Streaming coverage.The reported results show mixed behavior at small values of n.
- Topic Comparison: LDA represents each discovered topic as a word-probability distribution, enabling comparison between Streaming topics, Firehose topics, and random Firehose subsamples.The analysis treats the number of topics and Dirichlet-prior parameters as LDA inputs, while focusing on similarity of resulting topics.
- Topic Comparison: Topics are matched by maximum-weight bipartite matching using Jaccard similarity between topic word sets, then compared with Jensen-Shannon divergence.The matching step addresses the lack of an implicit ordering among LDA topics.
- Topic Comparison: Lower coverage produces higher topic divergence, while higher coverage produces lower divergence between Streaming and Firehose topics.The authors interpret this pattern as showing that decreased Streaming coverage causes variance in discovered topics.
- Comparison with Random Samples: Random Firehose samples can yield topics close to Streaming topics at higher coverage levels, but random data produces significantly better topics than Streaming API data under a 3-sigma threshold.The comparison uses average Jensen-Shannon scores and z-scores across 100 random-data runs.
Network Measures
The study compares Twitter retweet networks at node and network levels, finding that sparse sampled data limits daily accuracy while longer observation periods improve key-player identification.
- Network construction: Retweet networks are directed User × User graphs whose nodes are tweeting or retweeted users, analyzed at both network and node levels.The analysis ignores repeated edge weights and self-loops, and uses ORA for network metrics.
- Node-Level Measures: Sparse Twitter networks make sampled network metrics less accurate when coverage is smaller, and daily rankings are unlikely to be fully correct with missing data.Prior work indicates network measures are more stable in denser networks.
- Node-Level Measures: ∼50% of key-players can be identified correctly for a single day, with accuracy increasing over longer observation periods.Potential Reach metrics are also quite stable on some days in aggregated data.
- Network-Level Measures: Network-level comparisons focus on the largest connected component and calculate clustering, degree, and centralization-related measures.Smaller disconnected components contain less than 1% of nodes, motivating their exclusion.
- Network-Level Measures: Node and link coverage generally resembles tweet coverage, but Streaming centralization values vary more across days because coverage fluctuates.The Firehose contains a larger proportion of peripheral nodes, reflected by fewer nodes with non-zero in-degree.
Geographic Measures
The geographic analysis examines how Streaming API sampling affects geotagged-tweet distributions, finding high overall coverage but strong sensitivity to the collection boundary box.
- Geographic Measures: The study compares the geographic distribution of geolocated tweets between Streaming API and Firehose data.Geolocation is treated as a facet for evaluating the effects of Streaming API sampling.
- Geographic Measures: 16,739 Streaming and 18,579 Firehose tweets were geotagged, representing 3.17% and 1.45% of their respective datasets.Despite different total tweet counts, the Streaming data covered 90.10% of geotagged tweets.
- Geographic Measures: Excluding tweets from the collection boundary box reduced Streaming coverage of geotagged tweets to 39.19%.More than 90% of geotagged tweets from both sources were excluded, producing a more even Asia–North America representation.
Conclusion and Future Work
The Streaming API’s usefulness depends strongly on coverage and analysis type: it shows bias in several comparisons, but performs well for some large-scale or geographically bounded analyses. The study provides evidence to help researchers judge when sampled Twitter data can support their conclusions.
- The Streaming API’s coverage decreases as the number of tweets matching the requested parameters increases.More specific parameter sets involving users, bounding boxes, and keywords may mitigate this reduction.
- Top hashtags are estimated well for large n but can be misleading when n is small, while topic similarity improves with greater API coverage.The API performed worse than randomly sampled Firehose data, especially at low coverage.
- One day of Streaming API data identifies 50–60% of the top 100 retweet-network key-players on average.Aggregating data across multiple days can substantially increase accuracy, and network centralization correlates with API coverage.
- The Streaming API almost captures all geotagged tweets when geographic boundary boxes are used, although geotagged tweets comprise only about 1% of tweets overall.After removing tweets collected this way, the datasets diverge more, but retain similar continental distributions.
- Overall, Streaming API results depend on coverage and analysis type, so researchers should account for these nuances when grounding findings in Twitter data.The authors present initial evidence that may help estimate API coverage and plan to develop methods to compensate for sampling bias.