Source-linked AI summary

How the Scientific Community Reacts to Newly Submitted Preprints: Article Downloads, Twitter Mentions, and Citations

Xin Shuai, Alberto Pepe, Johan Bollen

arXiv:1202.2461v3cs.SIcs.DLphysics.soc-ph

TL;DR

The paper asks how scientific preprints receive online attention and whether Twitter activity relates to downloads and early citations. Using a cohort of arXiv submissions, it compares response timing and tests correlations among these measures. Twitter responses are faster and more short-lived than downloads, while mention volume correlates with downloads and early citations.

  • Problem

    The study addresses limited understanding of the temporal relations among readership, Twitter mentions, and subsequent citations for scientific preprints.

  • Method

    The authors analyze arXiv downloads, Twitter mentions, and early citations, comparing their timing and testing regression and correlation relationships.

  • Results

    Twitter mentions have shorter delays and narrower time spans than arXiv downloads, and mention volume is statistically correlated with both downloads and early citations.

  • Takeaways & Limitations

    Social-media attention may provide information about scholarly usage and early citation patterns, supporting broader assessment of scholarly impact.

  • Takeaways & Limitations

    Citation data cover only early citations, and the observational methodology confounds factors, so the results do not establish that tweeting increases citation rates.

Abstract

from arXiv · show

We analyze the online response to the preprint publication of a cohort of 4,606 scientific articles submitted to the preprint database arXiv.org between October 2010 and May 2011. We study three forms of responses to these preprints: downloads on the arXiv.org site, mentions on the social media site Twitter, and early citations in the scholarly record. We perform two analyses. First, we analyze the delay and time span of article downloads and Twitter mentions following submission, to understand the temporal configuration of these reactions and whether one precedes or follows the other. Second, we run regression and correlation tests to investigate the relationship between Twitter mentions, arXiv downloads and article citations. We find that Twitter mentions and arXiv downloads of scholarly articles follow two distinct temporal patterns of activity, with Twitter mentions having shorter delays and narrower time spans than arXiv downloads. We also find that the volume of Twitter mentions is statistically correlated with arXiv downloads and early citations just months after the publication of a preprint, with a possible bias that favors highly mentioned articles.

1 Introduction

The paper examines how online scholarly communication, especially social media, relates to readership and citation behavior. It analyzes the temporal relations among Twitter mentions, article downloads, and subsequent citations for scientific preprints.

  • Social media may influence scholarly citation behavior as research increasingly becomes an online process.
  • Prior research examined scientists’ Twitter use and found that article mentions can predict future citations.
  • The study extends this work by examining temporal relations among readership, Twitter mentions, and subsequent citations.
  • The authors compare Twitter mentions, arXiv downloads, and early citations in both magnitude and time.
  • Download and social-media responses follow distinct temporal patterns, while social-media mentions correlate significantly with downloads and citation counts.

2 Data and study overview

The study follows 4,606 arXiv submissions using download, Twitter, and early-citation data, then characterizes response timing and relationships among these measures. It defines delay and span to compare how quickly responses peak and how long they persist.

  • 2.1 Data collection: 4,606 articles submitted between October 4, 2010 and May 2, 2011 form the study cohort.Weekly downloads and daily Twitter mentions were recorded through May 9, 2011; citations were collected from Google Scholar on September 30, 2011.
  • 2.1 Data collection: 2,904,816 downloads were recorded for the 4,606 articles in weekly arXiv logs.
  • 2.1 Data collection: Twitter data scanned 1,959,654,862 tweets, identifying mentions of 4,415 articles, approximately 95% of the cohort.Tweets contained explicit or shortened links to arXiv papers; the collection used a randomly sampled 10% Gardenhose feed.
  • 2.1 Data collection: Early citations were manually collected for the 70 most Twitter-mentioned articles, which together had 431 citations by September 30, 2011.The most cited article had 62 citations, while most articles received hardly any citations.
  • 2.1 Data collection: The study focuses on immediate responses and therefore treats citations accumulated over 5 months to 1 year as early citations rather than total potential citations.
  • 2.2 Definitions: delay and time span.: Delay measures the time from submission to the peak response, whereas span measures the interval between the first and last download or Twitter mention.These measures are defined for both arXiv downloads and Twitter mentions.

3 Results and discussion

The analyses characterize distinct temporal patterns for arXiv downloads and Twitter mentions, then examine how downloads, mentions, and early citations relate. Twitter activity is faster and more ephemeral, while Twitter mentions show stronger citation associations than downloads in the analyzed sample.

  • Domain-level descriptive statistics: The corpus shows broad subject-domain coverage, with Physics—especially Astrophysics, High Energy Physics, and Mathematics—prominent among downloaded and Twitter-mentioned papers.Download totals increase over time, partly because earlier papers have had longer to accumulate reads, whereas total Twitter mentions decrease.
  • Domain-level descriptive statistics: Twitter mentions are strongly skewed toward low frequencies, with few papers receiving relatively many mentions across the five most frequent subject domains.The analysis warns that the 10% Twitter Gardenhose sample may underestimate absolute mention counts by a factor of 10.
  • Temporal analysis of delay and span: All articles took more than four days to reach peak arXiv downloads, and most remained downloaded beyond 100 days.The download-delay distribution indicates that nearly all articles required at least five days to reach peak downloads.
  • Temporal analysis of delay and span: Twitter mentions peak rapidly after submission and usually disappear quickly, whereas arXiv downloads persist much longer.Nearly 80% of articles reached peak Twitter mentions one day after submission, while most were downloaded for over 100 days.
  • Regression between article downloads, Twitter mentions, and citations: The citation regression models include Twitter mentions, arXiv downloads, and elapsed publication time as predictors of early citations.The citation analysis is restricted to the 70 most Twitter-mentioned articles, with publication time included because articles had unequal opportunities to accumulate citations.
  • Regression between article downloads, Twitter mentions, and citations: Twitter mentions are more strongly associated with early citations than arXiv downloads, while publication period is also a non-negligible predictor.ArXiv downloads do not show a statistically significant relationship to early citations when the other predictors are accounted for.

4 Discussion

Online scholarly impact signals arise from overlapping communities and show distinct temporal and correlational patterns. The observed associations may reflect social-media influence, manuscript quality or appeal, or other confounded factors rather than a demonstrated causal pathway.

  • 70 most-mentioned articles show statistically significant correlations among Twitter mentions, arXiv downloads and citations, despite strongly skewed distributions.
  • ArXiv downloads and Twitter mentions should not be treated as purely scientific impact or public chatter because their user communities overlap.
  • Twitter mentions and arXiv downloads follow distinct temporal patterns, with mentions showing shorter delays and narrower activity spans.
  • The observed correlations could reflect causal effects of social-media exposure or shared manuscript quality and popular appeal.
  • The methodology confounds distinct or overlapping factors, so the results do not establish that social-media coverage fully determines impact or increases citations by itself.

5 Materials

The materials describe how Twitter mentions of arXiv papers were identified by resolving direct, shortened and indirect links. The process expanded shortened URLs before organizing mentions into four categories.

  • Tweets were filtered for URLs that directly or indirectly linked to arXiv papers.
  • The classification includes direct arXiv links, expanded shortened arXiv links, and direct or shortened links to pages containing arXiv links.
  • Shortened URLs were expanded through APIs for 16 popular shortening services, resolving 98,377,880 URLs.
Loading 1202.2461v3…