Source-linked AI summary
A Glimpse of the First Eight Months of the COVID-19 Literature on Microsoft Academic Graph: Themes, Citation Contexts, and Uncertainties
Chaomei Chen
TL;DR
The rapidly expanding and uncertain COVID-19 literature is difficult to navigate, while available datasets may lack citation information and other metadata needed for scientometric analysis. The paper introduces a flexible method combining structural, temporal, thematic, citation-context, and uncertainty analyses, and demonstrates it on a MAG-based COVID-19 dataset. The method supports visual exploration at multiple levels and enables comparison of bibliographic databases, although citation contexts covered only about 10% of records at the time.
Problem
Rapid COVID-19 literature growth and uncertainty create a navigation challenge, while some datasets lack cited references, abstracts, or adequate coverage for scientometric analysis.
Method
The paper combines network and temporal analysis, thematic clustering, citation-context analysis, and linguistic uncertainty measures using Microsoft Academic Graph data.
Results
The method provides a flexible, extensible visual-analytic workflow and enables comparison of bibliographic databases for rapidly growing research domains.
Takeaways & Limitations
Researchers can construct tailored datasets and study literature through citation contexts and network-centric visual analytics at multiple levels of granularity.
Takeaways & Limitations
Only about 10% of the dataset’s records had citation contexts at the time of writing, although those records appeared concentrated in the network core.
Abstract
from arXiv · showhide
As scientists worldwide search for answers to the overwhelmingly unknown behind the deadly pandemic, the literature concerning COVID-19 has been growing exponentially. Keeping abreast of the body of literature at such a rapidly advancing pace poses significant challenges not only to active researchers but also to the society as a whole. Although numerous data resources have been made openly available, the analytic and synthetic process that is essential in effectively navigating through the vast amount of information with heightened levels of uncertainty remains a significant bottleneck. We introduce a generic method that facilitates the data collection and sense-making process when dealing with a rapidly growing landscape of a research domain such as COVID-19 at multiple levels of granularity. The method integrates the analysis of structural and temporal patterns in scholarly publications with the delineation of thematic concentrations and the types of uncertainties that may offer additional insights into the complexity of the unknown. We demonstrate the application of the method in a study of the COVID-19 literature.
Introduction
The rapidly expanding COVID-19 literature creates urgent challenges for researchers and society, while existing datasets can be inadequate for scientometric analysis. Citation-based methods and bibliographic databases provide essential foundations for studying the research landscape.
- COVID-19 presents major scientific and societal challenges because its origins and transmission routes remain largely unknown and effective vaccines were not yet available.
- Rapid growth of COVID-19 research makes it difficult for researchers and society to keep abreast of the literature.
- CORD-19 and similar datasets may lack cited references, abstracts, or adequate data coverage and quality for scientometric research.
- Citation-based scientometric indicators, including the h-index and g-index, depend on references in scholarly publications.
CiteSpace
CiteSpace analyzes scholarly knowledge domains through networks, clusters, and temporal patterns. Its framework also measures linguistic uncertainty in citation contexts to help interpret research developments and unresolved issues.
- CiteSpace constructs networks of entities and relations, combining structural and temporal patterns to identify significant developments in a knowledge domain.
- Its workflow decomposes networks into clusters of frequently co-cited references and examines their meanings and interrelationships.
- CiteSpace uses modularity and silhouette scores to evaluate cluster decomposition, with values closer to 1.0 indicating clearer configurations.
- Structural holes are represented by references with high betweenness centrality, highlighting potentially boundary-spanning positions in the network.
- The uncertainty framework measures epistemic uncertainty, hedging, and transitional uncertainty through linguistic cues and information entropy.
Microsoft Academic Services
Microsoft Academic Services provides access to the Microsoft Academic Graph and enriches publication metadata with references and citation contexts. The study uses these capabilities to build a citation-aware COVID-19 literature dataset for visual analysis.
- Microsoft Academic Services allows users to access the Microsoft Academic Graph through an API or maintain their own copy.
- MAG records can be retrieved by MAG IDs, DOIs, and fields of study, and citing articles can be found for sets of references.
- Citation contexts give Microsoft Academic Services a distinctive capability for enriching bibliographic records and integrating citation analysis with visual analytics.
- The COVID-19 dataset was constructed by selecting records matching specified coronavirus-related fields of study.
- As of September 5, 2020, the dataset contained 79,476 records, including 7,693 with corresponding citation contexts.
Data
The study builds a self-constructed MAG dataset and focuses its analysis on COVID-19 publications from the first eight months of 2020. Comparisons across databases show substantial differences in retrieved literature volume.
- The analysis focuses on the subset of COVID-19 publications from the first eight months of 2020.
- The MAG query used coronavirus-related field-of-study matches to construct the demonstrative dataset.
- The CORD-19 dataset contained 130,000 articles and served as the volume reference for comparing bibliographic databases.
- Web of Science represented about 23% of the CORD-19 volume, while MAG represented over 62%.
- Full-text searches on Dimensions and Lens represented over 75% of the CORD-19 volume.
Overview
The study uses citation-network analysis to overview the COVID-19 literature and augments conventional co-citation analysis with citation contexts that allow references to be weighted more selectively.
- The approach represents references as nodes and their co-occurrence strengths as links, allowing clusters of frequently co-cited references to emerge.
- Citation contexts enhance conventional Document Co-Citation Analysis by retaining frequently cited references and discarding below-average references.Users may control this filtering option.
- Figure 1 overviews 1,330 top-cited articles within 77,897 COVID-19 articles.
Concept Trees
Concept trees expose the thematic structure of citation clusters and individual references, while interactive citation-context inspection surfaces specific concerns and reported values.
- The concept tree for Cluster #5, pregnant women, identifies vertical transmission as the primary concern and lets users inspect its citation contexts.
- Concept trees can represent the structure of a cluster, a reference, or a phrase.
- For a reference with the highest epistemic uncertainty score, incubation period and mean incubation period emerge as major citation themes.
- Interactive inspection repeatedly surfaces a mean incubation period of 5.2 days from citation contexts.
Uncertainties
The method measures linguistic uncertainty in citation contexts and supports interactive exploration of uncertain themes and concepts in the COVID-19 literature.
- Uncertainty framework: Citation-context uncertainty is measured using uncertainty cue words, with semantically equivalent cues learned from hand-picked seed words.The framework distinguishes epistemic uncertainty, hedging, and transitional uncertainty; epistemic cues target unknown, incomplete, controversial, and contradictory aspects.
- Uncertainty framework: Table 3 combines epistemic uncertainty cues with rhetorical words in conclusions to reveal additional insights into research challenges.
- Citation contexts: Figure 4 uses citation contexts for Li, Guan et al. (2020) to visualize epistemic uncertainty and highlight cue words such as unknown, uncertainty, and controversial.The orange-bar length represents the epistemic uncertainty score for each context.
- Concept exploration: Concept trees organize concepts co-occurring with a reference or phrase, allowing interactive inspection of themes and their underlying citation contexts.For vaccine or vaccination, the tree represents the hierarchy of co-occurring concepts and supports interactive exploration.
- Concept exploration: The method illustrates how uncertainty analysis and concept trees can support sense-making across rapidly growing COVID-19 literature.
Structural Variation Analysis (SVA)
Structural Variation Analysis identifies newly published articles that connect distinct thematic clusters and may therefore have transformative potential. Its visualizations also expose how epistemic uncertainty and cross-cluster citation patterns are distributed across the literature.
- SVA purpose and design: SVA is built on Structural Hole Theory and uses structural variation metrics including modularity change and centrality divergence.
- SVA purpose and design: SVA identifies new articles that may have profound future impacts by detecting novel links between distinct literature clusters.Transformative-potential scores distinguish novel inter-cluster links from incremental within-cluster links.
- Network visualization: The smaller demonstration network preserves similar clusters, with spike protein as the largest cluster and transmission as the second largest.
- Network visualization: Most epistemic uncertainty is concentrated in clusters 0, 2, and 5, while cluster 6 has very low overall uncertainty.Cluster 0 contains the most nodes with significant epistemic uncertainty; Li, Guan et al. (2020) is among the three most uncertain nodes.
- Transformative-potential examples: SVA identified a newly published article with high transformative potential by centrality divergence and a high ranking in modularity change.Its cited references span different clusters, producing a diverse structural footprint.
- Illustrative citation contexts: Citation evidence presents conflicting remdesivir findings, with compassionate-use improvement contrasted against a randomized trial finding no significant clinical benefit.The contrast suggests that additional efficacy studies were needed to demonstrate remdesivir’s benefit.
- Illustrative citation contexts: Cytokine storm syndrome is discussed as a possible contributor to critical disease and death, while cytokine and biomarker changes distinguish critically ill patients.Higher inflammatory-marker levels were reported in ICU patients, and biochemical measurements were linked to disease severity and mortality.
- Transformative-potential examples: The dataset’s citation-context coverage enabled analysis of promising areas such as vaccine development and innovative connections between thematic islands.Nearly 15% of records contained vaccine or vaccination, and SVA results were described as consistent with citation-based ranking.
Discussion
The examples support using MAG to study rapidly expanding literature at monthly timescales and to compare bibliographic databases. The main boundary is limited citation-context coverage, although available contexts appear concentrated in the network core.
- Contributions and applications: MAG integration supports scientometric analysis of rapidly growing literature over monthly timescales and enables end users to construct datasets at their preferred pace.
- Contributions and applications: MAG enables comparison across bibliographic databases, including differences in COVID-19 record coverage between Web of Science, CORD-19, MAG, Lens, and Dimensions.The cited comparison reports 29,858 Web of Science records, 130,000 CORD-19 records, and roughly 70,000–90,000 records for MAG, Lens, and Dimensions.
- Limitations: About 10% of records had citation contexts, but those records appeared concentrated in the core of the larger network while missing records tended toward the periphery.The authors describe this distribution as assuring and state that the available contexts still revealed specific answers such as consensus on mean incubation period.
Conclusion
The paper introduces a flexible visual analytic method for scientometric studies that addresses shortcomings in existing literature-analysis practices. It supports customizable datasets, citation-context access, and integrated analysis across levels of detail.
- The method addresses significant shortcomings in conducting scientometric studies of the literature.
- Researchers can tailor dataset breadth and depth to their requirements and apply existing tools at their own pace.
- Easy access to citation contexts contributes valuable information to visual analytic workflows.
- Combining citation contexts with network-centric analysis may support integrated study of scientific literature at multiple levels of detail.