Source-linked AI summary

Text mining and visualization using VOSviewer

Nees Jan van Eck, Ludo Waltman

arXiv:1109.2058v1cs.DL

TL;DR

Researchers need ways to create and interpret bibliometric maps from large text collections. This report presents VOSviewer’s text mining functionality, which builds co-occurrence-based term maps and demonstrates applications in scientific corpora. The examples reveal field structure, institutional focus, thesis organization, and topic-level citation-impact differences.

  • Problem

    Analyzing large amounts of text data requires a way to create, visualize, and explore bibliometric maps of science.

  • Method

    VOSviewer identifies relevant noun phrases and creates two-dimensional term maps from English-language document corpora using co-occurrence-based relatedness.

  • Results

    The applications show term maps revealing scientific-field structure, Leiden University’s research focus, thesis chapter organization, and differences in citation impact among journal topics.

  • Takeaways & Limitations

    VOSviewer’s text mining functionality supports analysis and visualization of large amounts of scientific text data and can also be applied to nonscientific texts.

Abstract

from arXiv · show

VOSviewer is a computer program for creating, visualizing, and exploring bibliometric maps of science. In this report, the new text mining functionality of VOSviewer is presented. A number of examples are given of applications in which VOSviewer is used for analyzing large amounts of text data.

1. Introduction

VOSviewer is a freely available program for creating, visualizing, and exploring bibliometric maps of science. The report presents its new text mining functionality for analyzing large amounts of text data.

  • VOSviewer analyzes bibliometric network data, including citation, collaboration, and term co-occurrence relations.

2. Text mining functionality

The new functionality creates two-dimensional term maps from English-language document corpora, positioning terms by co-occurrence-based relatedness. It identifies and selects relevant noun phrases using linguistic processing and distributional differences.

  • Term maps position terms so smaller distances indicate stronger relatedness based on document co-occurrences.
  • VOSviewer supports term maps from publications, patents, and newspaper articles, but only for English-language documents.
  • The workflow identifies noun phrases through part-of-speech tagging, linguistic filtering, and singularization of plural phrases.
  • Noun-phrase relevance increases with the Kullback-Leibler distance between its co-occurrence distribution and the overall distribution.

3. Applications

Three applications show term maps revealing field structure, institutional research activity, thesis organization, and topic-level citation impact across scientific texts.

  • A map of about 10,000 LIS publications identified three roughly equal, well-separated subfields: bibliometrics/scientometrics, library science, and information science/information retrieval.The bibliometrics and library science subfields appeared slightly more connected to each other than either was to information science.
  • Coloring the LIS map by Leiden University activity showed a strong focus on the bibliometrics subfield.
  • In a PhD-thesis map, clusters of 218 relevant terms corresponded reasonably well with the thesis’s different chapters.Examples included clusters for automatic term identification and VOSviewer.
  • The JASIST map used 468 terms to compare average citation impact across topics and showed large differences between them.It indicated strong separation between bibliometric/scientometric and information science/information retrieval topics.

4. Conclusion

The report presents VOSviewer’s new text mining functionality through applications analyzing large amounts of text data. Although the examples focus on scientific texts, the functionality also applies to nonscientific texts.

  • The report presents VOSviewer’s new text mining functionality through examples analyzing large amounts of text data.
  • The functionality can also be applied to nonscientific texts such as newspaper articles.
Loading 1109.2058v1…