Source-linked AI summary
Distributional Semantics and Linguistic Theory
Gemma Boleda
TL;DR
The review addresses distributional semantics’ limited impact in theoretical linguistics despite its computational success. It synthesizes methods and findings across semantic change, polysemy and composition, and grammar-semantics interfaces, concluding that distributional structure captures several graded semantic phenomena while retaining important data and context limitations.
Problem
Distributional semantics has been useful in computational linguistics and cognitive science, but its impact in theoretical linguistics has remained limited.
Method
The survey critically synthesizes computational distributional-semantic research relevant to semantic change, polysemy and composition, and syntax- and morphology-semantics interfaces.
Results
Distributional methods capture semantic relations, graded semantic phenomena, compositional effects, and aspects of semantic change, with composition exploiting conceptual and referential meaning.
Takeaways & Limitations
Distributional semantics offers theoretical linguistics scalable empirical tools for exploring phenomena, identifying instances, testing hypotheses, and discovering theoretically relevant trends.
Takeaways & Limitations
Distributional methods require large amounts of data and remain challenged by bias, context dependence, and highly context-dependent interpretations of larger constituents.
Abstract
from arXiv · showhide
Distributional semantics provides multi-dimensional, graded, empirically induced word representations that successfully capture many aspects of meaning in natural languages, as shown in a large body of work in computational linguistics; yet, its impact in theoretical linguistics has so far been limited. This review provides a critical discussion of the literature on distributional semantics, with an emphasis on methods and results that are of relevance for theoretical linguistics, in three areas: semantic change, polysemy and composition, and the grammar-semantics interface (specifically, the interface of semantics with syntax and with derivational morphology). The review aims at fostering greater cross-fertilization of theoretical and computational approaches to language, as a means to advance our collective knowledge of how it works.
1. INTRODUCTION
Distributional semantics learns semantic representations from natural-language data and models meaning through geometric relations among vectors. This survey reviews its relevance to theoretical linguistics and seeks greater integration between computational and theoretical approaches.
- The survey critically reviews distributional-semantic methods and results relevant to semantic change, polysemy and composition, and the grammar-semantics interface.
- Distributional semantics represents units as vectors learned from large text collections, with contexts determining positions in a multidimensional semantic space.
- Vector similarity models semantic relatedness: words with similar contexts receive similar vectors and occupy nearby positions.
- Automatic induction scales to large vocabularies, languages, and domains when sufficient linguistic data are available.
- Distributional models support rich semantic distinctions through hundreds of dimensions, including nuanced differences among near-synonyms.
- Continuous values make distributional similarity graded, corresponding to graded semantic relations such as degrees of near-synonymy.
2. SEMANTIC CHANGE
Distributional approaches model semantic change by comparing word representations across time and inspecting shifts in similarity and nearest neighbors. They can detect, locate, characterize, and test hypotheses about change, but remain constrained by data scarcity, language coverage, and confounding contextual effects.
- Distributional semantic change methods build word representations for different time points and compare them to detect change and track its temporal evolution.
- Cosine similarity for gay relative to its 1900 representation fell from around 0.75 in the late 1970s to around 0.3 in 2000.
- Nearest-neighbor inspection traces semantic shifts, including gay from cheerful-related neighbors toward homosexual and broadcast from concrete to abstract use.
- Distributional methods can detect narrowing and broadening, with contexts separating over time for broadening cases and converging for narrowing.
- Across four datasets comprising tens of thousands of words, synonyms stayed closer in space than control pairs across the 20th century, supporting parallel change over differentiation.
- Semantic-change research can detect, temporally locate, and characterize shifts, and test competing theories, although identifying shift types remains under-researched.
- Data scarcity makes diachronic models vulnerable to spurious effects and contributes to concentration on English and the Google Book Ngrams corpus.
- Research has focused mainly on lexical change, while newer neural models create possibilities for studying grammaticalization and function words.
3. POLYSEMY AND COMPOSITION
Distributional semantics models polysemy through single or sense-specific representations and composes meanings into larger constituents. Evidence shows composition captures graded semantic effects, while context-dependent interpretation, function words, and larger constituents remain challenging.
- 3.1. Single representation, polysemy via composition: Single word vectors abstract over all contexts and encompass the word’s attested senses, while alternative approaches assign different vectors to different senses.Both strategies parallel longstanding linguistic treatments of polysemy, although distributional semantics does not resolve when distinct senses require separate vocabulary items.
- 3.1. Single representation, polysemy via composition: Compositional distributional methods build phrase representations from their parts, with matching semantic dimensions reinforcing shared properties and shifting ambiguous words toward contextually relevant uses.In cut cost, composition shifts cut from its predominantly physical sense toward its abstract sense associated with reducing costs.
- 3.1. Single representation, polysemy via composition: An average cosine similarity of 0.6 was obtained by the best composition method between synthetic and corpus-based phrase vectors in Boleda et al. (2013).Synthetic phrases are evaluated by comparing their vectors with corpus-extracted vectors; closer vectors indicate better composition.
- 3.1. Single representation, polysemy via composition: Vector addition is surprisingly good and often outperforms more sophisticated composition methods, suggesting that distributional power lies substantially in lexical representations.Distributional models also distinguish acceptable from unacceptable unattested adjective-noun phrases, including semantically anomalous combinations.
- 3.3. Discussion: Composition models content words in short phrases but struggle with larger constituents, function words, and highly context-dependent interpretations requiring referential or speaker meaning.They model general uses such as red box for a box red in color, but not context-specific uses such as a brown box containing red objects.
- 3.2. Different representations, polysemy via word senses: Distributional representations alleviate the problem of encoding information shared across senses by representing similarities and differences across dimensions in a graded fashion.They do not improve the decision about when senses are different enough to warrant separate vocabulary items.
4. GRAMMAR-SEMANTICS INTERFACE
Distributional methods connect linguistic form and meaning by modeling graded selectional preferences, morphological composition, and semantic variation. The reviewed work supplies empirical data and suggests theoretically relevant patterns, while performance remains variable across words and derivational processes.
- 4.1. Syntax-semantics interface: Selectional restrictions can be modeled as graded similarities between candidate arguments and verb-specific argument prototypes.Erk et al. compare candidate–prototype similarity with human plausibility ratings for agent and patient roles.
- 4.1. Syntax-semantics interface: Erk et al.’s model correlated with human ratings at 0.33 and 0.47 Spearman correlation across two English datasets.Both correlations were statistically significant at p < 0.001.
- 4.2. Morphology-semantics interface: Derivational morphology requires representations that encode both grammatical and semantic properties of stems and affixes, including polysemy and gradedness.The suffix -er illustrates agentive and instrumental interpretations and combines with verbs under morphosyntactic and semantic constraints.
- 4.2. Morphology-semantics interface: Compositional distributional methods capture derivational phenomena such as affix polysemy and context-sensitive stem-sense selection.Synthetic vectors distinguish agentive carver from instrumental broiler and select the journalism sense in columnist.
- 4.3. Discussion: The evidence remains limited by substantial performance variation across individual words and derivational patterns.Word frequency and lexicalization can reduce the quality of corpus-based representations; one study found mean similarity of 0.47 for derived and base forms versus 0.56 for the best compositional measure.
- 4.2. Morphology-semantics interface: Vector addition performs surprisingly well for derivation, inflection, and compound compositionality, although modifier information can matter more than head information.For inflection, Mikolov et al. report 40% average accuracy over a vocabulary of 82,000 words; in compound compositionality, the modifier is a much better predictor than the head.
- 4.3. Discussion: The reviewed research contributes data for theoretical work, including participant ratings, typicality and relatedness measures, and information about derived words.These resources can support material selection and empirical investigation in theoretical linguistics.
- 4.3. Discussion: Distributional models reveal theoretically relevant patterns, including argument-structure effects and interactions between derivational valence effects and concreteness.Derivations affecting argument structure are reported as harder to model, while diminutives and prefixation show different valence effects for concrete and abstract words.
5. CONCLUSION AND OUTLOOK
The conclusion identifies multiple ways distributional semantics can inform linguistic theory while emphasizing empirical biases, data requirements, and scope limitations. It also points toward closer integration with neural-network research.
- 5. CONCLUSION AND OUTLOOK: Distributional research contributes to linguistic theory through exploration, phenomenon identification, hypothesis testing, and discovery of theoretically relevant trends.The fourth role is described as hardest and requiring collaboration between computational and theoretical linguists.
- 5. CONCLUSION AND OUTLOOK: Distributional models mirror their input data, yielding radically empirical representations but also exposing them to biases in the underlying data.The survey presents this duality as both an advantage and a challenge of data-driven methods.
- 5. CONCLUSION AND OUTLOOK: Distributional methods require large amounts of data to learn reasonable representations.
- 5. CONCLUSION AND OUTLOOK: The survey focuses on simple vector methods and omits machine-learning techniques commonly used with semantic spaces.The author directs readers to the survey’s references for more information about the omitted methods.
- 5. CONCLUSION AND OUTLOOK: The discussion also excludes neural-network research not specifically targeted at building semantic spaces.Neural networks are described as inducing representations while learning to perform tasks such as translation.
- 5. CONCLUSION AND OUTLOOK: The complexity and success of distributional models have increased interest in understanding what aspects of language they capture, alongside arguments for integrating neural networks into linguistic research.The author endorses this integration.
DISCLOSURE STATEMENT
The author reports no affiliations, memberships, funding, or financial holdings that might be perceived as affecting the review’s objectivity.
- DISCLOSURE STATEMENT: The author reports no affiliations that might be perceived as affecting the review’s objectivity.
- DISCLOSURE STATEMENT: The author reports no memberships that might be perceived as affecting the review’s objectivity.
- DISCLOSURE STATEMENT: The author reports no funding or financial holdings that might be perceived as affecting the review’s objectivity.