Source-linked AI summary

Visualizing a Field of Research: A Methodology of Systematic Scientometric Reviews

Chaomei Chen, Min Song

arXiv:1906.04800v1cs.DL

TL;DR

Systematic scientometric reviews need bibliographic datasets that are comprehensive without depending entirely on what reviewers already know. This paper introduces cascading citation expansion, unifies global and local science mapping as points on a continuous expansion spectrum, and demonstrates the approach with five literature-based-discovery datasets. The comparisons show that full-text search can miss recently emerged topics, while combining query search with citation expansion is likely to provide more balanced domain coverage.

  • Problem

    Query-based search can require complex queries and remains limited by currently known vocabulary, risking omission of significant or unfamiliar relevant literature.

  • Method

    The paper introduces cascading citation expansion and applies it across three use scenarios to construct and compare five literature-based-discovery datasets.

  • Results

    Network-overlay comparisons of five datasets revealed topics that a commonly used full-text search strategy could have missed, particularly recently emerged topics.

  • Takeaways & Limitations

    Combining query-based search with cascading citation expansion is likely to provide more balanced research-domain coverage and reduce the risk of missing unfamiliar topics.

  • Takeaways & Limitations

    The paper leaves open how seed-article choice affects expansion stability and how many expansion generations are optimal.

Abstract

from arXiv · show

Systematic scientometric reviews, empowered by scientometric and visual analytic techniques, offer opportunities to improve the timeliness, accessibility, and reproducibility of conventional systematic reviews. While increasingly accessible science mapping tools enable end users to visualize the structure and dynamics of a research field, a common bottleneck in the current practice is the construction of a collection of scholarly publications as the input of the subsequent scientometric analysis and visualization. End users often have to face a dilemma in the preparation process: the more they know about a knowledge domain, the easier it is for them to find the relevant data to meet their needs adequately; the little they know, the harder the problem is. What can we do to avoid missing something valuable but beyond our initial description? In this article, we introduce a flexible and generic methodology, cascading citation expansion, to increase the quality of constructing a bibliographic dataset for systematic reviews. Furthermore, the methodology simplifies the conceptualization of globalism and localism in science mapping and unifies them on a consistent and continuous spectrum. We demonstrate an application of the methodology to the research of literature-based discovery and compare five datasets constructed based on three use scenarios, namely a conventional keyword-based search (one dataset), an expansion process starting with a groundbreaking article of the knowledge domain (two datasets), and an expansion process starting with a recently published review article by a prominent expert in the domain (two datasets). The unique coverage of each of the datasets is inspected through network visualization overlays with reference to other datasets in a broad and integrated context.

Introduction

Systematic reviews synthesize and assess research to clarify a field’s state of the art, while science mapping tools help visualize its structure and dynamics. Scientometric approaches also shift some review work from domain experts toward computational analysis of scholarly texts.

  • Systematic reviews: Systematic reviews synthesize original research, assess consensus or disagreement, and identify challenges and future directions.They can orient newcomers, update experienced researchers, and help other stakeholders understand a scientific field.
  • Science mapping: Science mapping tools take bibliographic records as input to visualize the structure and dynamics of a research field.
  • Computational scientometrics: Computational techniques can extract meaningful information from text documents and synthesize thematic patterns, reducing reliance on direct domain-expert involvement.
  • Literature-based discovery: Literature-based discovery seeks to reveal and use undiscovered public knowledge in existing scholarly literature.The study focuses on LBD rather than the broader literature-related discovery concept, which also includes literature-assisted discovery.

Mapping the Scientific Landscape

Science mapping ranges from broad global representations to focused local maps, while expansionism conceptualizes these states along a continuous spectrum. Local mapping is especially suited to problem-driven domains but faces the challenge of constructing representative literature collections.

  • Mapping approaches: Science mapping approaches are characterized as global, local, or hybrid according to their intended scope and applications.
  • Globalism: Global maps seek holistic coverage across scientific disciplines, commonly using journals and journal clusters to represent disciplinary structure.They provide a relatively stable organizing framework for understanding scientific knowledge.
  • Globalism: Global maps can reveal structural insights unavailable at smaller scales and support detection of potentially transformative links when both link contexts are represented.
  • Localism: Local maps focus on problem-driven knowledge domains whose boundaries follow relevance to central research questions rather than disciplinary or institutional boundaries.Their practical challenge is constructing a representative literature body from a much larger database.
  • Search and relevance: Semantic techniques and external resources can augment query terms with synonyms or related concepts to establish article relevance to a topic.
  • Expansionism: Expansionism unifies globalism and localism as states on a continuous spectrum formed by progressively adding publications to the represented knowledge.The process begins at local-map coverage and can extend toward a global endpoint encompassing all published articles.
  • Expansionism: Cascading citation expansion demonstrates one pathway between local and global approaches and can begin from different types of seed articles.

Cascading Citation Expansion

Cascading citation expansion constructs bibliographic datasets by extending an initial query or seed set through forward and backward citation chains. In the literature-based discovery demonstration, citation-expanded datasets provided broader coverage and stronger network-clustering properties than the full-text search dataset.

  • Rationale: Citation-based searching reduces reliance on specifying potentially relevant topics in the initial query, helping limit missed topics.A robust strategy combines a query-based initial set with citation-chain expansion.
  • Demonstration: The study compares five literature-based discovery datasets across three scenarios: full-text search, expansion from a groundbreaking article, and expansion from a recent review.S3 and S5 begin with Swanson’s 1986 article, while NF and NB begin with Smalheiser’s review.
  • Method: The expansion process can start from a singleton or multiple-article set and use forward citation expansion, backward citation expansion, or both.The paper describes m-generation backward and n-generation forward expansion as a generic formulation.
  • Dataset distributions: The dataset distributions differ substantially: S3 concentrates mostly between 2006 and 2018, whereas S5 increases steadily over time and NB reaches back to 1936.S3 uses a minimum inclusion threshold of 10 citations, while S5 uses 20 to reduce processing time.
  • Dataset coverage: 46,756 unique articles span the five datasets, with S5 representing 93.47% of the combined set, NB 5.21%, and F 3.80%.Overlap is measured as the ratio of articles in common to the set union.
  • Network properties: S3 has the highest silhouette score at 0.80, while NB has the highest modularity at 0.95 and S5 the second highest at 0.91.The full-text search network has modularity 0.82 but the lowest silhouette value, 0.09; the authors identify S3 and S5 as potential systematic-review bases.

Literature-Based Discovery

This section visualizes literature-based discovery from multiple perspectives using five datasets, beginning with full-text search results and then cascading citation expansions.

  • The thematic landscape of literature-based discovery is examined across five datasets using full-text search and cascading citation expansions.

Full Text Search

The full-text search uses two field-specific phrases and retrieves a large dataset for visualizing literature-based discovery. However, query-based strategies remain limited by the terminology specified in advance.

  • The Dimensions query combines ‘literature-based discovery’ with ‘undiscovered public knowledge,’ including terminology from Swanson’s 1986 publications.
  • 1,777 records were retrieved, including 431 matched to PubMed abstracts for later exploration.
  • The full-text search supports a thematic overview of literature-based discovery through document co-citation analysis.
  • Query-based search can miss significant literature because its coverage depends on known concepts and chosen vocabularies.

Full Text Search vs. Forward Citation Expansion

The comparison shows that full-text search and forward citation expansion cover different parts of the literature-based discovery landscape. Cascading citation expansion makes these additions and omissions explicit for verifying a field map’s scope.

  • S5 is a five-generation forward citation expansion from Swanson’s fish-oil and Raynaud’s-syndrome article, while the full-text network is overlaid for comparison.
  • The full-text search overlaps with S5 in one region but leaves a large upper-left area uncovered.
  • The shared region contains three Swanson articles, while additional clusters concern migraine, magnesium, migraine aura, visual cortex excitability, and total blood magnesium.
  • The extra clusters cannot be classified uniformly as gains for expansion or losses for full-text search because their value depends on the reviewer’s purpose.
  • Cascading citation expansion makes potential omissions and additions explicit, providing an additional way to verify a field map’s scope.

The Structure of the Thematic Landscape

A combined network enables direct comparison of the five datasets and reveals both shared core areas and strategy-specific coverage. Forward, backward, and full-text approaches reach substantially different parts of the thematic landscape.

  • The combined base map identifies clusters uniquely covered by each dataset and areas covered by multiple datasets.
  • NB covers the earliest upper-left clusters most comprehensively, whereas S5 provides the best coverage around the upper-right clusters.
  • The full-text network appears to be a superset of S3 but a subset of S5, placing individual datasets in a broader context.
  • EF_n(S) denotes n-step forward expansion, while EB_m(S) denotes m-step backward expansion; NB combines one forward and one backward step from Smalheiser’s 2017 review.
  • Cluster #2 biomedical literature is consistently shared, while different approaches diverge in coverage outward from this core area.
  • Full-text search misses much of one region, S5 substantially covers it, and backward expansions cover different branches of the same landscape.
  • S5 fully covers clusters #32 deep learning and #30 big data, which full-text search F and backward expansion NB miss categorically.

Major Themes

The combined network is organized around major themes including drug discovery, medical literature, and biomedical literature. Cluster exploration and concept trees expose their internal structures and associated subthemes.

  • Major Themes: CiteSpace cluster exploration represents co-cited reference clusters as underlying research themes and divides top-level clusters into second-level sub-clusters.A concept tree generated from phrases in relevant articles examines the main branches within a cluster.
  • Major Themes: The largest cluster, #0, is labeled drug discovery and includes sub-clusters on protein interaction and computational drug discovery.Its concept tree highlights branches on drug repositioning and drug repurposing.
  • Major Themes: The second-largest top-level cluster, #1, is labeled medical literature and includes LBD systems, text mining, spatial relationship, and network-based retrieval model themes.LBD is identified as an abbreviation of literature-based discovery.
  • Major Themes: The third-largest top-level cluster, #2, is biomedical literature, with sub-clusters on literature-based discovery, entity recognition, big data, and the genomic era.This thematic area is consistently covered by all individual network overlays.

Special Themes

Overlay comparisons identify clusters shared across datasets and reveal specialized structures within fish oil, Raynaud’s syndrome, and deep learning themes. These examples connect local sub-clusters to broader literature-based discovery structures.

  • Special Themes: The top-five second-level clusters of top-level cluster #15 provide a further decomposition of that higher-level theme.
  • Special Themes: Fish oil and Raynaud’s syndrome appear in the title of Swanson’s groundbreaking article and correspond to clusters #10 and #45.The fish oil cluster contains tightly coupled sub-clusters on fish oil, blood viscosity, and chemical structure.
  • Special Themes: Only overlay S5 fully covers cluster #32 deep learning, illustrating topics that full-text search alone could miss.Its largest second-level clusters are deep learning and convolutional neural network, and its concept tree branches into deep learning approaches and drug design.

Discussions and Conclusions

The paper proposes cascading citation expansion to improve systematic scientometric reviews and uses multiple network overlays to examine dataset coverage. It recommends combining search strategies while recognizing that seed and stopping choices affect results.

  • Discussions and Conclusions: Cascading citation expansion is proposed to improve the quality of systematic scientometric reviews.The approach can begin with either a pioneering article or a recent review article.
  • Discussions and Conclusions: Combining query-based search with cascading citation expansion is likely to provide more balanced domain coverage and reduce reliance on topics specified in initial queries.The strategy may uncover emerging topics connected to the main literature through chains of weak links.
  • Discussions and Conclusions: Multiple cascading expansions with different seed articles are recommended, while multi-generation expansions reduce the risk of missing unfamiliar or unknown topics.
  • Discussions and Conclusions: Overlapping networks indicate core topics, and decomposing top-level clusters into second-level clusters can provide additional insights.The recommendations include multiple levels of network and cluster analysis.
  • Discussions and Conclusions: Different starting and ending points can produce different results, with network complexity and threshold selections affecting reproducibility.The paper raises questions about seed validity, expansion stability, and the optimal number of generations.

Notes

The paper identifies access points for reproducing the analysis and obtaining the datasets.

  • Notes: CiteSpace is available through SourceForge, and the five datasets are available through Google Drive.
Loading 1906.04800v1…