Source-linked AI summary
The Pushshift Reddit Dataset
Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, Jeremy Blackburn
TL;DR
Researchers need large-scale social-media datasets, but collecting and managing Reddit data remains technically demanding and time-consuming. The paper presents Pushshift’s Reddit dataset with historical data, monthly dumps, a searchable API, and community tools; it reports a large dataset already used in over 100 peer-reviewed publications.
Problem
Collecting and systematically analyzing Reddit’s large-scale data requires substantial engineering and computational resources, creating barriers for research.
Method
The paper releases Pushshift’s Reddit dataset through monthly dumps, a searchable API, and additional community resources for querying, visualization, and discussion.
Results
The dataset contains 651,778,198 submissions and 5,601,331,385 comments from June 2005 through April 2019, and had been used in over 100 peer-reviewed publications by late 2019.
Takeaways & Limitations
Pushshift provides an accessible, large-scale Reddit data resource for data-intensive research and reduces time spent on collection, cleaning, and storage.
Abstract
from arXiv · showhide
Social media data has become crucial to the advancement of scientific understanding. However, even though it has become ubiquitous, just collecting large-scale social media data involves a high degree of engineering skill set and computational resources. In fact, research is often times gated by data engineering problems that must be overcome before analysis can proceed. This has resulted recognition of datasets as meaningful research contributions in and of themselves. Reddit, the so called "front page of the Internet," in particular has been the subject of numerous scientific studies. Although Reddit is relatively open to data acquisition compared to social media platforms like Facebook and Twitter, the technical barriers to acquisition still remain. Thus, Reddit's millions of subreddits, hundreds of millions of users, and hundreds of billions of comments are at the same time relatively accessible, but time consuming to collect and analyze systematically. In this paper, we present the Pushshift Reddit dataset. Pushshift is a social media data collection, analysis, and archiving platform that since 2015 has collected Reddit data and made it available to researchers. Pushshift's Reddit dataset is updated in real-time, and includes historical data back to Reddit's inception. In addition to monthly dumps, Pushshift provides computational tools to aid in searching, aggregating, and performing exploratory analysis on the entirety of the dataset. The Pushshift Reddit dataset makes it possible for social media researchers to reduce time spent in the data collection, cleaning, and storage phases of their projects.
1 Introduction
Researchers face shrinking and stratified access to social-media data, alongside persistent inefficiencies and reproducibility concerns. The paper addresses these constraints by releasing Pushshift’s Reddit dataset with dumps, an API, and collaborative tools.
- Open social-media data access has narrowed through resource deprecation, stratified access, and increased fear of prosecution around terms-of-service violations.These changes have curtailed researchers’ ability to collect timely data, share tools, instruct students, and reproduce findings.
- Earlier API-driven research also faced inefficient collection, unequal access, ethical concerns, weak data sharing, and limited validation or generalization.Open platforms such as Reddit nevertheless continued to offer access to social data.
- Pushshift releases monthly dumps containing 651M submissions and 5.6B comments posted on Reddit between 2005 and 2019.The release is presented as part of the goal of providing open APIs and data dumps to researchers.
- The release includes a researcher API and Slackbot in addition to monthly dumps, supporting access to and interaction with the collected Reddit data.The API supports querying the whole dataset without downloading the dumps, while the Slackbot supports real-time visualization and discussion.
- These resources reduce time spent on data collection, cleaning, and storage before research teams begin analyzing the data.The API also reduces the need for substantial storage capacity by enabling queries over the whole dataset.
2 Pushshift
Pushshift combines technical and social infrastructure to collect, store, index, and disseminate Reddit data. Its platform supports scalable ingestion, querying, aggregation, visualizations, and researcher communities.
- 2 Pushshift: Pushshift provides technical infrastructure for collecting, storing, cataloging, indexing, and disseminating social-media data, alongside organizational processes for governing and discussing it.The platform therefore includes both software and hardware components and a social infrastructure for responsible data use.
- 2 Pushshift: Its core subsystems are an ingest engine, PostgreSQL database, Elasticsearch document store cluster, and API for dynamic access and aggregation.The database supports advanced querying and metadata storage, while Elasticsearch supports indexing and aggregation.
- 2.1 Data collection process: The ingest engine orchestrates heterogeneous data-collection programs through job scheduling and common storage APIs.Individual programs may interact with web APIs, scrape HTML pages, or use available data streams, without requiring a particular programming language.
- 2.1 Data collection process: Elasticsearch scales horizontally through clustering and maintains redundancy with multiple index replicas, while supporting large-scale storage and analysis.Pushshift’s API exports much of Elasticsearch’s search and aggregation functionality and serves 500M requests per month.
- 2.1 Data collection process: Pushshift’s Reddit and Slack communities support announcements, questions, bug reports, feature feedback, data-science discussion, and visualization.The Reddit community has more than 2,100 subscribers, while the Slack team has nearly 300 registered users and more than 260,000 messages across 53 channels.
- 2.1 Data collection process: A Slack chatbot analyzes and visualizes Pushshift data from channel queries, returning time-series plots and summary statistics within a few seconds.The chatbot can also be shared with other non-Pushshift workspaces.
3 Description of the Pushshift Reddit Dataset
The Pushshift Reddit dataset provides large-scale Reddit submissions and comments in monthly JSON files, with metadata described for each record type. It is also designed for public access and interoperability through monthly dumps, an API, and JSON formatting.
- Dataset coverage: 651,778,198 submissions and 5,601,331,385 comments from 2,888,885 subreddits span June 2005 through April 2019.Comments exceeded 1M per day after August 2013 and reached 5M per day by April 2019; submissions exceeded 500K per day consistently near the dataset’s end.
- File structure: The dataset is organized into separate monthly newline-delimited JSON files for submissions and comments.Each line contains one submission or comment represented as a JSON object.
- Record structure: Submission and comment records include documented keys and values describing their data fields.The paper provides separate descriptions for submission and comment JSON objects.
- FAIR principles: The dataset aligns with FAIR principles through public monthly dumps, a persistent sample DOI, API and Slackbot access, and JSON interoperability.The entire dataset could not be uploaded to Zenodo because of its 100GB limit, while the full collection is several terabytes.
4 Dataset Use Cases
Pushshift data supports research across diverse areas, including online governance, extremism, disinformation, web science, big data science, health informatics, and intelligent systems. By late 2019, over 100 peer-reviewed publications had used the dataset.
- Over 100 peer-reviewed publications had used Pushshift data by late 2019, spanning topics such as toxicity, personality, virality, and governance.
- Online community governance: Pushshift data supports studies of volunteer-led online community governance and responses to social movements, fringe identities, hate speech, and harassment campaigns.
- Online extremism: Researchers use Pushshift data to study online extremism despite difficulties obtaining timely data from rapidly changing fringe and mainstream spaces.
- Online disinformation: Pushshift data has been used to investigate disinformation and social-media trustworthiness, an area where limited data access constrains research.
- Pushshift enables research on technology diffusion, interface effects, online communities, network science, cloud computing, very large databases, health, and intelligent systems.
5 Related Work
Pushshift belongs to a broader ecosystem of large-scale, researcher-oriented data services and dataset releases. Related platforms provide combinations of APIs, data dumps, query tools, visualization, and analysis infrastructure for studying online information and society.
- Researchers have proposed alternatives to cloud-hosted storage buckets that are better tailored to research needs.
- Media Cloud tracks hundreds of millions of news stories and offers aggregated counts and topical data through a free, semi-public API.
- GDELT provides a database and knowledge graph for filtering and visualizing global news events, related topics, and imagery.
- Stack Exchange combines hosted data dumps, an activity API, and a Data Explorer for SQL queries against a regularly updated database.
- Wikimedia provides data dumps, robust APIs, interactive services, and Jupyter Notebook access to replication databases for revisions and content.
- Other dataset papers have released large collections from platforms such as Gab, extending research access to social-media data beyond major services.
6 Discussion & Conclusion
The paper presents Pushshift as a large-scale Reddit dataset with historical coverage, monthly dumps, a searchable API, and community resources. Its existing use across disciplines supports its continuing value as a research resource.
- Pushshift includes hundreds of millions of submissions and billions of comments from 2005 until the present.
- The dataset is available through monthly dumps, a searchable API, and additional tools and community resources.
- Having been used in over 100 papers across numerous disciplines, Pushshift is positioned as an ongoing resource for the research community.