Source-linked AI summary
Text mining and visualization using VOSviewer
Nees Jan van Eck, Ludo Waltman
TL;DR
Researchers need ways to create and interpret bibliometric maps from large text collections. This report presents VOSviewer’s text mining functionality, which builds co-occurrence-based term maps and demonstrates applications in scientific corpora. The examples reveal field structure, institutional focus, thesis organization, and topic-level citation-impact differences.
Problem
Analyzing large amounts of text data requires a way to create, visualize, and explore bibliometric maps of science.
Method
VOSviewer identifies relevant noun phrases and creates two-dimensional term maps from English-language document corpora using co-occurrence-based relatedness.
Results
The applications show term maps revealing scientific-field structure, Leiden University’s research focus, thesis chapter organization, and differences in citation impact among journal topics.
Takeaways & Limitations
VOSviewer’s text mining functionality supports analysis and visualization of large amounts of scientific text data and can also be applied to nonscientific texts.
Abstract
from arXiv · showhide
VOSviewer is a computer program for creating, visualizing, and exploring bibliometric maps of science. In this report, the new text mining functionality of VOSviewer is presented. A number of examples are given of applications in which VOSviewer is used for analyzing large amounts of text data.
1. Introduction
VOSviewer is a freely available program for creating, visualizing, and exploring bibliometric maps of science. The report presents its new text mining functionality for analyzing large amounts of text data.
- VOSviewer analyzes bibliometric network data, including citation, collaboration, and term co-occurrence relations.
2. Text mining functionality
The new functionality creates two-dimensional term maps from English-language document corpora, positioning terms by co-occurrence-based relatedness. It identifies and selects relevant noun phrases using linguistic processing and distributional differences.
- Term maps position terms so smaller distances indicate stronger relatedness based on document co-occurrences.
- VOSviewer supports term maps from publications, patents, and newspaper articles, but only for English-language documents.
- The workflow identifies noun phrases through part-of-speech tagging, linguistic filtering, and singularization of plural phrases.
- Noun-phrase relevance increases with the Kullback-Leibler distance between its co-occurrence distribution and the overall distribution.
3. Applications
Three applications show term maps revealing field structure, institutional research activity, thesis organization, and topic-level citation impact across scientific texts.
- A map of about 10,000 LIS publications identified three roughly equal, well-separated subfields: bibliometrics/scientometrics, library science, and information science/information retrieval.The bibliometrics and library science subfields appeared slightly more connected to each other than either was to information science.
- Coloring the LIS map by Leiden University activity showed a strong focus on the bibliometrics subfield.
- In a PhD-thesis map, clusters of 218 relevant terms corresponded reasonably well with the thesis’s different chapters.Examples included clusters for automatic term identification and VOSviewer.
- The JASIST map used 468 terms to compare average citation impact across topics and showed large differences between them.It indicated strong separation between bibliometric/scientometric and information science/information retrieval topics.
4. Conclusion
The report presents VOSviewer’s new text mining functionality through applications analyzing large amounts of text data. Although the examples focus on scientific texts, the functionality also applies to nonscientific texts.
- The report presents VOSviewer’s new text mining functionality through examples analyzing large amounts of text data.
- The functionality can also be applied to nonscientific texts such as newspaper articles.