Source-linked AI summary

A unified approach to mapping and clustering of bibliometric networks

Ludo Waltman, Nees Jan van Eck, Ed C. M. Noyons

arXiv:1006.1032v1cs.DLphysics.data-anphysics.soc-ph

TL;DR

Bibliometric studies often combine mapping and clustering even though the techniques may use different principles and assumptions. This paper derives VOS mapping and weighted, parameterized modularity-based clustering from one principle, then illustrates the approach on frequently cited information-science publications from 1999–2008. The resulting combined analysis provides an overview of the field and supports maps at different levels of detail.

  • Problem

    Mapping and clustering are often combined in bibliometric analyses despite typically relying on different ideas and assumptions.

  • Method

    The paper derives VOS mapping and a weighted, parameterized variant of modularity-based clustering from a shared underlying principle.

  • Results

    The approach produces a combined map and clustering that reveals the structure of information science, including ISR and informetrics subfields and 25 publication clusters.

  • Takeaways & Limitations

    A unified approach can support multiple maps of the same domain at different levels of detail, including detailed expert-validated maps and general cluster-level maps.

  • Takeaways & Limitations

    Mapping usually cannot display relations in more than two dimensions, whereas clustering avoids dimensional restrictions but provides a coarser binary picture.

Abstract

from arXiv · show

In the analysis of bibliometric networks, researchers often use mapping and clustering techniques in a combined fashion. Typically, however, mapping and clustering techniques that are used together rely on very different ideas and assumptions. We propose a unified approach to mapping and clustering of bibliometric networks. We show that the VOS mapping technique and a weighted and parameterized variant of modularity-based clustering can both be derived from the same underlying principle. We illustrate our proposed approach by producing a combined mapping and clustering of the most frequently cited publications that appeared in the field of information science in the period 1999-2008.

1. Introduction

Bibliometric mapping and clustering are commonly combined to reveal network structure, but they typically rely on different ideas and assumptions. The paper motivates a unified approach to make such analyses more transparent and consistent.

  • Motivation: Mapping and clustering techniques are used to study the structure of networks of documents, keywords, authors, and journals.They address questions about topics, relationships among fields, and domain development over time.
  • Existing practice: Researchers often combine a map showing individual network nodes with a clustering displayed over the map.Clusters may be marked as areas or represented by coloring nodes according to cluster membership.
  • Existing practice: The most commonly used combination is multidimensional scaling for mapping and hierarchical clustering.Other alternatives include Kamada–Kawai, pathfinder networks, VxOrd, VOS, and factor analysis.
  • Research gap: Although mapping and clustering share the objective of revealing network structure, they have generally been developed separately.Consequently, techniques used together may be based on different ideas and assumptions.
  • Paper focus: The paper proposes a unified approach to mapping and clustering and illustrates it with frequently cited information-science publications from 1999–2008.The paper also relates the approach to earlier physics literature before summarizing its conclusions.

2. Mapping and clustering: A unified approach

The proposed approach represents mapping and clustering through a shared objective based on distances and association strengths. It yields VOS mapping and a weighted, resolution-parameterized modularity-based clustering formulation.

  • Common formulation: The approach minimizes a common objective over node representations, using association strength to reward proximity between strongly related nodes.For mapping, each node receives a vector location; for clustering, each node receives a positive integer denoting its cluster.
  • Common formulation: High-association nodes are pulled together while low-association nodes are pushed apart by attractive and repulsive forces.The repulsive force does not depend on association strength, producing the stated separation effect.
  • Clustering: The clustering case is equivalent to maximizing a weighted, parameterized modularity function.The resolution parameter γ controls cluster granularity: larger values produce more clusters.
  • Mapping: The mapping case is equivalent to the VOS mapping technique, which is closely related to multidimensional scaling.This establishes the mapping side of the unified formulation.
  • Clustering: Setting γ and the weights w_ij to 1 reduces the clustering function to Newman and Girvan’s modularity function.The proposed method is therefore a weighted variant of modularity-based clustering with an added resolution parameter.

3. Related work

The unified formulation is connected to parameterized mapping and modularity work from physics. It differs from related approaches by directly linking established mapping and clustering techniques through one formulation.

  • Relation to earlier work: The approach resembles Noack’s parameterized objective function, which connects force-directed mapping techniques with Newman–Girvan modularity.Noack’s framework includes techniques such as Fruchterman–Reingold.
  • Relation to earlier work: Unlike Noack’s result, this paper directly relates well-known mapping techniques to modularity-based clustering rather than treating their objective functions as special cases of one function.The supplied passage identifies this as the first of three differences from Noack’s result.
  • Relation to earlier work: Setting the weights w_ij equal to 1 makes the proposed clustering technique essentially the generalized modularity function of Reichardt and Bornholdt.This connects the clustering formulation to parameterized modularity work in the physics literature.

4. Illustration of the proposed approach

The unified approach was applied to 1,242 frequently cited information-science publications from 1999–2008, producing a mapped and clustered overview with 25 clusters. The map distinguishes major subfields and finer research areas, including informetrics clusters summarized in Table 1.

  • Map structure: The resulting map separates information seeking and retrieval on the left from informetrics on the right.Within information seeking and retrieval, hard system-oriented research occupies a small upper-left area, while soft user-oriented research occupies a larger middle and lower-left area.
  • Cluster structure: The clustering produced 25 clusters, averaging 49.7 publications with a standard deviation of 31.5 per cluster.The smallest cluster contains two publications, while the largest contains 123 publications and concerns citation analysis and related bibliometric and scientometric topics.
  • Cluster structure: Eight of the 25 clusters cover the informetrics subfield, and their contents are summarized in Table 1.The table also lists the four authors with the largest number of publications in each cluster as important authors.
  • Interpretation: The authors describe the application as yielding an accurate and detailed picture of information science’s structure.The combined map and clustering can be inspected using VOSviewer, and the clustering is also available in a spreadsheet file.

5. Conclusions

The paper argues that mapping and clustering should be unified because they are complementary but traditionally rely on different ideas and assumptions. It proposes combining VOS mapping with weighted, parameterized modularity-based clustering and identifies uses across map detail levels.

  • Complementarity and limitations: Mapping provides a detailed network picture but is usually restricted to two dimensions, making higher-dimensional relations invisible.This limitation concerns the practical dimensionality of maps rather than the underlying network.
  • Complementarity and limitations: Clustering avoids dimensional restrictions but uses binary rather than continuous dimensions, producing a coarser picture of network structure.The paper discusses clustering techniques in which each node is assigned to exactly one cluster, not overlapping-cluster methods.
  • Rationale: A unified approach aligns mapping and clustering around similar ideas and assumptions, helping avoid inconsistencies between their results.The motivation follows from their complementary roles and frequent combined use in bibliometric analysis.
  • Contribution: The proposal unifies VOS mapping with a weighted and parameterized variant of modularity-based clustering.It also connects a longstanding statistical research stream with a more recent physics research stream.
  • Applications: The unified approach is especially useful when multiple maps of the same domain are needed at different levels of detail.The paper gives science-policy mapping as an example, where expert and managerial audiences may need maps showing individual nodes and clusters respectively.
  • Availability and scope: Algorithms implementing the unified approach are incorporated into VOSviewer, with open-source MATLAB algorithms also available.The paper’s stated scope excludes clustering methods that allow nodes to belong to multiple clusters.

Appendix A

The appendix proves the equivalence between the clustering formulation based on minimizing the mapping-related objective and maximizing the weighted modularity formulation. It does so by applying a constant transformation and substituting the relevant definitions.

  • Equivalence proof: The appendix states that minimizing equation (3) is equivalent to maximizing equation (6) with weights w_ij given by equation (7).This establishes the link between the clustering objective and the weighted modularity formulation.
  • Definitions: The indicator δ(x_i, x_j) equals 1 when x_i = x_j and 0 otherwise.This indicator encodes whether two nodes belong to the same cluster.
  • Transformation: Equation (9) is obtained from equation (8) by multiplying by a constant and adding a constant.Because the multiplicative constant is always negative, minimizing equation (8) is equivalent to maximizing equation (9).
  • Transformation: Substituting equation (8) into equation (9) yields equation (10), which can then be rewritten as equation (6) using weights w_ij from equation (7).The appendix concludes that this completes the proof of equivalence.

Appendix B

The appendix illustrates how weighting changes the proposed clustering technique’s assignment of an ambiguous node relative to modularity-based clustering. The proposed assignment is preferred because it reflects relative rather than absolute connectivity.

  • The proposed clustering technique and modularity-based clustering both identify three clusters, but disagree on node 31’s assignment.Both methods group nodes 1–10, 11–20, and 21–30 into separate clusters, while assigning node 31 differently.
  • The proposed technique assigns node 31 with nodes 1–10, whereas modularity-based clustering assigns it with nodes 11–20.This disagreement results from the different weighting schemes used by the two methods.
  • The proposed technique gives less weight to nodes with larger total link counts and consequently gives relatively more weight to nodes 1–10.Nodes 11–20 have almost an order of magnitude more total links than nodes 1–10.
  • Node 31 has 2.5 times more links with nodes 11–20 than with nodes 1–10, despite the much larger total link count of nodes 11–20.From a relative point of view, node 31 is more strongly connected to nodes 1–10.
  • The authors consider assigning node 31 to nodes 1–10 preferable because that assignment reflects its stronger relative association with that group.They state that, in this particular example, the proposed clustering results are preferable to those of modularity-based clustering.
Loading 1006.1032v1…