Source-linked AI summary

The semantic mapping of words and co-words in contexts

Loet Leydesdorff, Kasper Welbers

arXiv:1011.5209v2cs.CLstat.AP

TL;DR

Semantic mapping addresses how meaning and semantic dynamics can be measured in discourse. The paper introduces computer-assisted statistical approaches, illustrates them with an empirical example, and summarizes analyst choices and software options.

  • Problem

    Measuring the dynamics of meaning remains a challenge despite semantic approaches such as latent semantic analysis, semantic web research, and semantic mapping.

  • Method

    The paper provides an introduction, empirical example, software pointers, and guidance on selecting words and analyzing word-document relations using statistical and computer-assisted techniques.

  • Results

    The empirical search produced 195 documents containing 59 words, with “Factor” and “Impact” identified as central in the semantic map.

  • Takeaways & Limitations

    Semantic mapping makes visualization of semantic relations more accessible and supports informed choices by analysts and scholars using relevant software.

  • Takeaways & Limitations

    The analysis must account for skewed word distributions, and the described software is available in standard packages such as SPSS.

Abstract

from arXiv · show

Meaning can be generated when information is related at a systemic level. Such a system can be an observer, but also a discourse, for example, operationalized as a set of documents. The measurement of semantics as similarity in patterns (correlations) and latent variables (factor analysis) has been enhanced by computer techniques and the use of statistics; for example, in "Latent Semantic Analysis". This communication provides an introduction, an example, pointers to relevant software, and summarizes the choices that can be made by the analyst. Visualization ("semantic mapping") is thus made more accessible.

Abstract

The paper concerns semantic mapping of words, documents, and latent meaning.

  • Semantic mapping relates words and documents to represent latent meaning.

Introduction

The paper introduces co-word and latent semantic approaches for studying semantic relations in document sets. It frames documents as cases and words as variables, while providing guidance, software, and analyst choices.

  • Co-word maps were developed as an alternative for studying semantic relations in scientific and technology literatures.
  • Latent Semantic Analysis further developed these mapping techniques using word-document matrices.
  • Factor-analytic techniques cluster words according to their distributions across documents.
  • Singular value decomposition combines alternative representations but is not readily available in standard software packages such as SPSS.
  • The communication provides an overview, software pointers, and arguments for implementation choices while keeping applications versatile.

The word-document matrix

The word-document matrix represents word occurrences across documents and supports multiple criteria for selecting words and comparing their distributions. The paper presents frequency-based, tf-idf, chi-square, and observed/expected measures for this purpose.

  • The word-document matrix records word occurrences, with documents as cases and words selected as variables for analysis.
  • Documents are units of analysis whose size and aggregation level can vary, so researchers must first select relevant units.
  • Tf-idf increases with a term’s frequency but decreases when that term occurs in more documents.
  • Chi-square compares observed cell values with expectations calculated from matrix margins, and its cell values can be aggregated by word.
  • Four word-selection criteria are available: frequency, tf-idf, column contribution to chi-square, and observed/expected margin totals.
  • Case studies found observed/expected margin totals most convenient, although all four measures are available.

The analysis

The analysis transforms word-document data into relational or vector-space representations for semantic mapping. Factor analysis, SVD, cosine normalization, and related choices affect interpretation, statistical properties, and the resulting maps.

  • The asymmetrical word-document matrix can be transformed into a symmetrical co-occurrence matrix for network analysis.
  • Factor analysis, multidimensional scaling, and SVD analyze latent dimensions of the word-document matrix.
  • Factor analysis and SVD use Pearson correlations among variable distributions, changing similarity from relations to correlations.
  • Because word distributions are skewed, using Pearson correlation is debatable; cosine normalization avoids the mean but sacrifices orthogonal rotation and statistical testing.
  • Factor analysis is preferable on the available two-mode word-document matrix, and cosine-normalized patterns or factor analysis can approximate semantic maps without producing precisely similar representations.
  • Semantic maps provide a systemic perspective in which meaning is represented through semantic structures in document sets.
  • The resulting structure is induced from data rather than imposed by an a priori scheme, potentially reducing the indexer effect.
  • Automated content analysis extends measurement from Likert scales to semantic maps of intrinsic meaning in document sets.

An empirical example

The empirical example applies factor analysis and co-word mapping to 195 documents containing “impact factor,” showing how alternative inputs and visualizations reveal different semantic structures. The maps identify central terms, factor groupings, and connections while also exposing interpretive limits.

  • Dataset and setup: The analysis used 195 documents containing 59 words occurring more than twice, after stop-word correction.The search targeted titles containing “impact factor” published in 2008 or 2009.
  • Factor analysis: Figure 1 used cosine-normalized word occurrences and suppressed factor loadings and cosine values below 0.1.The rotated factor matrix colored nodes according to their factor associations.
  • Alternative input: Using observed/expected ratios instead of word frequencies changed the factor structure: “Impact” and “Factor” no longer loaded positively on any factor and showed interfactorial complexity.The alternative input was the cellwise observed/expected matrix rather than the word-document frequency matrix.
  • Alternative visualizations: Figure 3 incorporated negative loadings by visualizing the rotated component matrix directly as an asymmetrical 2-mode matrix.Factor loadings correspond to Pearson correlations between variables and latent dimensions, allowing both to be projected into one vector space.
  • Co-word mapping: The co-word map kept “Factor” and “Impact” central because they co-occurred in all items, while critic-associated words formed a distinct visible grouping.“Journal” and “Indicator” appeared near the search terms, and the Factor 2 grouping remained separate near the center.
  • Interpretive limits: The star-shaped center-periphery structure dominated the network graph and tended to obscure semantic relations in the vector space.Visual inspection also became increasingly difficult for factors explaining less common variance.

Conclusions and summary

The paper distinguishes semantics as a property of language from meaning as use and treats meaning as generated by relating information at the systems level. It concludes that modeling the dynamics of meaning requires further elaboration.

  • Open problem: The measurement of the dynamics of meaning remains in its infancy, so modeling those dynamics requires further elaboration.The paper frames this as an unresolved extension of semantic analysis rather than as a completed method.
  • Systems-level meaning: Meaning is generated when different bits of information are related at the systems level and positioned in a vector space.The paper presents a discourse, such as a set of documents, as one possible system in which such relations can be studied.
  • Conceptual distinction: Semantics is treated as a property of language, whereas meaning is often defined in terms of use at the level of agency.This distinction separates language-level semantic analysis from pragmatic accounts of meaning.
  • Conceptual distinction: Semantic measurement has shifted toward the intrinsic meaning of textual elements in discourses and texts at an objective, supra-individual level.Pragmatic meaning can instead be measured through Likert scales and respondent reports.
Loading 1011.5209v2…