Source-linked AI summary

A Systematic Identification and Analysis of Scientists on Twitter

Qing Ke, Yong-Yeol Ahn, Cassidy R. Sugimoto

arXiv:1608.06229v2cs.DLcs.SIphysics.soc-ph

TL;DR

Prior arguments about scientists’ activity on social media lacked empirical grounding, while bibliographic reliance introduced bias. The paper presents a systematic, large-scale approach to identifying scientists across disciplines and examines their demographics and connections, finding disciplinary and gender-representation differences.

  • Problem

    Existing arguments lacked empirical grounding, and reliance on bibliographic databases bound studies to traditional citation indicators and introduced bias.

  • Method

    The paper presents a large-scale, systematic study that improves scientist identification by integrating how scientists self-identified with how they were identified through other evidence.

  • Results

    Scientists were over-represented in social disciplines and under-represented among mathematical and physical scientists, while disciplinary patterns also appeared in network centralities.

  • Takeaways & Limitations

    The improved identification approach supports analysis of scientists across disciplines and provides a basis for studying their representation and interconnections.

  • Takeaways & Limitations

    The gender estimate is rough, and the method does not establish the extent to which women’s more equal representation is due to the identification procedure.

Abstract

from arXiv · show

Metrics derived from Twitter and other social media---often referred to as altmetrics---are increasingly used to estimate the broader social impacts of scholarship. Such efforts, however, may produce highly misleading results, as the entities that participate in conversations about science on these platforms are largely unknown. For instance, if altmetric activities are generated mainly by scientists, does it really capture broader social impacts of science? Here we present a systematic approach to identifying and analyzing scientists on Twitter. Our method can identify scientists across many disciplines, without relying on external bibliographic data, and be easily adapted to identify other stakeholder groups in science. We investigate the demographics, sharing behaviors, and interconnectivity of the identified scientists. We find that Twitter has been employed by scholars across the disciplinary spectrum, with an over-representation of social and computer and information scientists; under-representation of mathematical, physical, and life scientists; and a better representation of women compared to scholarly publishing. Analysis of the sharing of URLs reveals a distinct imprint of scholarly sites, yet only a small fraction of shared URLs are science-related. We find an assortative mixing with respect to disciplines in the networks between scientists, suggesting the maintenance of disciplinary walls in social media. Our work contributes to the literature both methodologically and conceptually---we provide new methods for disambiguating and identifying particular actors on social media and describing the behaviors of scientists, thus providing foundational information for the construction and use of indicators on the basis of social media metrics.

Introduction

The paper addresses whether altmetrics reflect broader social impact by empirically identifying who participates in science-related Twitter activity. It introduces a large-scale, bibliographic-data-independent approach for identifying scientists across disciplines and other stakeholder groups.

  • Altmetrics may misrepresent broader impact because the entities generating social-media activity are largely unknown.
  • Validating broader-impact claims requires identifying scientists and non-scientists participating in Twitter activity.
  • Earlier identification studies were narrow, small-scale, or dependent on purposive samples and external bibliographic databases.
  • The study presents a large-scale, systematic approach that identifies scientists across many disciplines without external bibliographic databases.
  • The method can identify other stakeholder groups, occupations, and entities, supporting research on scholarly communication and altmetric impact.

Background

Prior work examined both the production of scholarly attention on Twitter and the relationship between altmetrics and traditional citations. It found limited coverage, weak or variable citation correlations, and sampling constraints in studies of scientists.

  • Twitter mentions cover a limited share of research papers, with coverage varying across disciplines and favoring medical and social-science papers.One cited study found Twitter mentions for 9.4% of papers indexed by PubMed and Web of Science.
  • The correlation between altmetrics and citations is generally negligible but varies substantially by discipline, journal, and time window.A meta-analysis reported r = 0.003.
  • Survey and small-scale studies show that scientists use Twitter for scholarly sharing and community interaction, with usage differing across disciplines.
  • Many scientist studies manually searched for researchers selected outside Twitter, producing small and popularity-biased samples.
  • More systematic approaches still relied on external bibliographic data and were confined to a single discipline.

Identifying Scientists

The study identifies scientists by combining occupational titles with Twitter-list-based discovery and profile validation. It expands a seed set through list memberships, yielding a large candidate pool and a refined final dataset while acknowledging coverage limitations.

  • Scientist occupations are defined from SOC and Wikipedia sources, producing a lexicon of 322 titles and related title variants.
  • The procedure infers identity from how multiple Twitter users classify an account in list names and descriptions.
  • The method combines 8,545 seed users with breadth-first snowball sampling through Twitter lists to discover additional users.
  • The approach excludes some users through name-format filtering and is blind to scientists who are not listed.
  • 110,708 users appeared in 4,920 lists containing scientist titles, and profile-description filtering produced 45,867 final users.

Analyzing Scientists

The study systematically identifies scientists on Twitter and analyzes their demographics, URL-sharing behavior, and network connections. Scientists span diverse disciplines, but representation and interaction patterns are uneven across disciplines and genders.

  • Identification: The method assigns disciplines to 30,793 users and demonstrates coverage across diverse scientific fields.Discipline assignment uses scientist titles extracted from profiles and list names, prioritizing profile information when available.
  • Demographics: Social and computer-and-information scientists are over-represented on Twitter, whereas mathematical, life, and physical scientists are under-represented.The comparison uses Twitter counts against U.S. science-workforce data, but the authors describe the estimate as rough.
  • Demographics: Women comprise 38.6% of identified scientists, and the gender ratio is less skewed on Twitter than among U.S. scientific authorships.Gender was identified for 71.9% of the sample, comprising 12,732 females and 20,232 males.
  • URL sharing: Scientists share discipline-specific scientific domains, but scientific links remain a minor activity for most scientists on Twitter.Nature.com is popular across fields, while arxiv.org and aps.org lead among physicists and acm.org is popular among computer scientists; biological scientists share more scientific links than other groups.
  • Network connections: Scientist networks are assortative by discipline but not by gender, with follower, retweet, and mention assortativity coefficients of 0.548, 0.492, and 0.537, respectively.Gender assortativity coefficients are 0.054 for following and 0.074 and 0.086 for retweeting and mentioning.

Discussion

The study broadens scientist identification on Twitter beyond external bibliographic data and finds substantial disciplinary, gender, sharing, and network structure among identified users. These patterns both inform altmetric interpretation and expose important sampling limitations.

  • Methodological contribution: The method identifies scientists across disciplines without relying on external bibliographic databases, broadening earlier approaches and supporting analysis beyond paper-centric data.It combines list- and bio-based classifications, integrating self-identification with community identification while favoring precision over recall.
  • Limitations: The sample favors precision over recall because list reliance, language filtering, private lists, unknown curation, and profile filtering leave many scientists unidentified.Twitter lists may also skew toward elite or high-profile science communicators, while excluding users whose names lack spaces biases sampling toward English-speaking users.
  • Disciplinary representation: Social scientists are overrepresented on Twitter, whereas mathematicians are particularly underrepresented; scholars from many disciplines are nevertheless represented.Historians, physicists, political scientists, computer scientists, biologists, economists, sociologists, psychologists, and nutritionists appear among identified users.
  • Gender representation: 38.6% of identified users with inferable gender were female and 61.4% were male, a more equal representation than in scholarly publishing.The authors suggest that scientists on Twitter may be more gender-balanced than the population of publishing scientists.
  • Sharing behavior: Scientists share both mainstream and scholarly domains, but science-related URLs form only a small fraction of their tweets, indicating heterogeneous content.Nature, Science, and arXiv are prominent scholarly destinations alongside Instagram, Facebook, YouTube, and general news sites.
  • Network structure: Disciplinary assortativity means scholars tend to follow others in their own disciplines, while mathematicians, computer scientists, and historians show relative network isolation.Because disciplinary walls persist, social-media metrics may not provide unfettered access to scholarship; altmetric practices should therefore normalize for field differences.

Conclusion

The paper develops a systematic method for identifying scientists on Twitter and analyzes their demographics, sharing behavior, and network interconnectivity. It finds disciplinary and gender imbalances, limited science-related URL sharing, and assortative mixing across disciplines.

  • The method systematically identifies scientists on Twitter through profile and Twitter-list information.
  • Social and computer and information scientists are over-represented, while mathematical and physical scientists are under-represented; women are better represented than in scholarly publishing.
  • Only a small portion of shared URLs are science-related, despite a distinct scholarly imprint in scientists’ sharing behavior.
  • Follower, retweet, and mention networks show assortative mixing by discipline among scientists.
  • Future work should examine machine learning, network-based identification improvements, gender-representation explanations, and alignment with bibliometric communities.

S1 Table. Scientist occupations from 2010 Standard Occupational

The table uses the 2010 Standard Occupational Classification released by the US Department of Labor.

  • The occupation classification is based on the 2010 Standard Occupational Classification released by the US Department of Labor.

S3 Table. Top scientist titles from Twitter list names.

The passage identifies incoming degree and strength measures and PageRank as scientist-ranking measures.

  • Scientist rankings use in-degree, in-strength, and PageRank measures.

S1 Text

The study identifies scientists from Twitter using list-based attributes, stringent seed selection, and snowball sampling, then examines demographics and communities. It finds disciplinary organization and connected clusters among scientists.

  • Identification pipeline: The pipeline scans a Gardenhose dataset, filters users by list coverage, and obtains attributes for 2,436,889 users.
  • Identification pipeline: It selects 8,545 high-precision seed users whose attributes include “science” and a scientist title, then applies snowball sampling.
  • Demographic analysis: Academic rank is inferred from profile keywords for students, postdocs, and professors, with the first matching category selected.
  • Community structure: The analysis uses Infomap to identify 343 follower-network communities with more than 10 nodes and labels them from profile-description words.
  • Community structure: Scientists organize by discipline and tend to follow others within their own scientific communities.
  • Community structure: Ecologists and biologists, astronomers and physicists, and political scientists, economists, and sociologists form tightly connected community groupings.
Loading 1608.06229v2…