Source-linked AI summary

The Changing Geometry of Grammar: Dimensionality and Neighborhood Reorganization across Transformer Layers

Samuele Vallisa, Federico Ravenda, Claudio Palominos, Rui He, Andrea Raballo, Antonietta Mira, Philipp Homan, Wolfram Hinzen

arXiv:2608.25166v1cs.CL

TL;DR

The paper asks whether grammatical role shapes the local geometry of transformer representations. It measures PoS-conditioned intrinsic dimension and neighborhood reorganization across layers in encoders and decoders, finding architecture-dependent trajectories that recover grammatical role from geometry alone.

  • Problem

    The paper examines whether tokens’ grammatical roles produce systematic geometric differences, addressing limited token-level evidence about how lexical and relational content traverse transformer layers.

  • Method

    The study compares PoS-conditioned intrinsic dimension and Information Imbalance trajectories across four transformer models and uses geometric features for grammatical-role classification.

  • Results

    Closed-class items expand earlier and collapse sooner than open-class items; encoder and decoder trajectories differ, while geometry-only classification reaches F1 = 0.540 versus shuffled baselines of 0.389 and 0.167.

  • Takeaways & Limitations

    Local geometric trajectories encode grammatical role independently of lexical identity and connect semantic informativeness with identifiable neighborhood reorganizations.

  • Takeaways & Limitations

    The analysis is restricted to English and four models, so persistence across typologically diverse languages, larger scales, and broader model families remains open.

Abstract

from arXiv · show

Transformer representations describe trajectories through high-dimensional vector spaces, which are shaped dynamically as tokens incorporate relational context across layers. Such data tend to concentrate on lower-dimensional sub-manifolds, a form of compression quantified by the Intrinsic Dimensionality (ID), the minimum number of independent variables needed to represent them without significant information loss. In this work, we ask whether the grammatical role of tokens, as marked by their part-of-speech (PoS) tag, shapes the local geometry of this manifold. To this end: (1) We investigate the layer-wise evolution of ID, finding that closed-class items expand earlier and collapse sooner than open-class ones; (2) We show its expansion and contraction to be explained by changes in the neighborhood structure, and hence in the relations between words within a sentence; (3) We compare encoders (ModernBERT, bigbird-roberta-large) and decoders (gemma-2-2B, Llama-3.2-3B), finding that the two families evolve differently across layers, consistently with how each integrates context;(4) We show that geometric features alone recover a token's grammatical role, and use them to interpret how the semantic content of each PoS evolves across layers in a downstream classification task.

1 Introduction & Related Works

The paper studies whether grammatical class shapes the geometry of transformer token representations. It focuses on distinct geometric trajectories for lexical content words and relational function words across network layers.

  • Transformer activations occupy lower-dimensional manifolds despite residing in high-dimensional spaces.Intrinsic Dimension relates this geometric compression to the effective degrees of freedom of representations.
  • Open-class and closed-class items differ in whether their content is lexical or relational.Function words contribute relationally and are near/vacuous in isolation, whereas content words denote lexically.
  • The study conditions geometric measures on parts of speech to test whether function and content words trace distinct trajectories.This brings representation-geometry analysis to the token level and asks whether lexical class systematically shapes those trajectories.
  • The paper links ID and neighborhood reorganization to context integration and tests whether geometric trajectories predict grammatical role.Its contributions include PoS-conditional geometric measures, architecture comparisons, and classification from geometry without lexical identity.

2 Methods

The method measures PoS-conditional manifold dimension and neighborhood change within documents, then compares their layer-wise behavior across encoder and decoder architectures. It also relates these trajectories to grammatical categories and their correlations.

  • Research Questions: The study asks whether lexical class, neighborhood reorganization, and context direction shape token geometry across layers.The research questions distinguish open versus closed classes, contraction versus expansion during dependency resolution, and bidirectional versus causal attention.
  • Data and Models: The analysis uses 539 Pile-10k documents of 1’000–1’500 words and examines ModernBERT, bigbird-roberta-large, gemma, and Llama.Documents average 1’218 word representations per sentence, supporting stable intrinsic-dimension estimates.
  • Data and Models: Eight frequent non-punctuation PoS categories cover 83% of occurrences, spanning open classes and principal closed classes.The categories include NOUN, VERB, ADJ, ADV, PROPN, DET, PRON, and ADP.
  • Geometric Descriptors: PoS-conditional intrinsic dimension estimates effective degrees of freedom using the nearest-neighbor ABIDE estimator.Low cID indicates constrained contextual variation, while high cID indicates variation across many dimensions.
  • Geometric Descriptors: Information Imbalance measures how well one layer’s neighborhood structure predicts another’s, complementing cID’s compression measure.Low values indicate preserved neighborhoods; high values indicate independent neighborhood structure, with separate input- and final-layer comparisons.
  • Geometric Descriptors: Figure 1 compares PoS-conditional neighborhood change and ID across layers for two encoders and two decoders.Its panels report input-anchored and final-layer-anchored II, cID trajectories, and Pearson/Spearman correlations.

3 Results

Across models, PoS-conditioned geometry reveals architecture-dependent neighborhood reorganization, and geometric trajectories alone predict both lexical class and fine-grained PoS.

  • Layer-wise geometry: Function words and content words follow distinct layer-wise cID and neighborhood trajectories across the evaluated transformer models.ModernBERT shows a middle-layer function-word plateau before content-word peaks; bigbird-roberta exhibits a separate early trajectory for DET, PRON, and PROPN.
  • Discriminative capacity: The results support a link between layer-wise geometric evolution and grammatical role, including differences in how contextual identity develops across encoder and decoder layers.The reported analyses combine cID, cII, and logistic regression over geometric trajectories.
  • Neighborhood reorganization: Encoder neighborhood structure remains tied to the input for function words, whereas decoder input-level predictivity weakens in intermediate layers.The paper relates this contrast to bidirectional versus causal context integration and to which layer provides the relevant contextual role.
  • Discriminative capacity: Geometric features alone classify function versus content words with F1 = 0.878 and specific PoS categories with F1 = 0.540.These scores exceed shuffled baselines of 0.389 and 0.167, respectively, without lexical identity.

4 Conclusions

PoS-conditioned ID and II reveal that grammatical status shapes token geometry, with architecture-dependent trajectories that remain predictive of grammatical role independently of lexical identity.

  • 4 Conclusions: Closed-class items expand into higher-dimensional neighborhoods earlier and collapse sooner than open-class items.These trajectories track the resolution of the elements with which the tokens combine.
  • 4 Conclusions: The timing of geometric change depends on whether attention accesses related elements from the input layer or only after further processing.The conclusion connects this architecture dependence to encoder and decoder context integration.
  • 4 Conclusions: Geometric trajectories alone recover grammatical role, indicating that grammatical status is encoded in local manifold geometry independently of lexical identity.This conclusion follows from the reported geometric-feature classification results.

5 Limitations

The analysis is limited by its English-only setting and its evaluation of only two encoders and two decoders.

  • Language scope: The findings are restricted to English, whose grammatical properties may not represent languages with different typological profiles.The paper specifically identifies head directionality and morphological marking as factors that could alter geometric trajectories.
  • Model coverage: The study examines two encoders and two decoders, so its models cannot represent the full space of transformer language models.Architecture, pretraining data, model scale, and training objective may influence representation geometry.
  • Model coverage: Whether the reported phenomena persist at larger scales and across a broader range of models remains open.The four evaluated models show consistent findings, but broader verification is still required.

A Derivation of the Geometric Descriptors

The paper estimates PoS-conditioned intrinsic dimension from local neighbor counts in concentric balls, using a likelihood whose density dependence cancels under local homogeneity. The procedure selects neighborhood sizes globally, then averages estimates within each PoS group with uncertainty controls.

  • Intrinsic dimension is the minimum number of variables needed to describe high-dimensional data without significant information loss.
  • ABIDE estimates dimension from neighbor counts inside two concentric balls with radii rB and rA = τrB.The estimator models how many neighbors fall within each ball around every point.
  • Under local density homogeneity, the volume ratio τ^d determines the probability that a neighbor inside the larger ball also lies in the smaller ball.Because local density cancels, the per-point factor depends on local geometry rather than the point’s density regime.
  • The restricted likelihood over tokens in a PoS group provides a conditional estimator of its intrinsic dimension.Its separability across points makes restriction to a token subset G legitimate.
  • Neighbor counts include tokens of every PoS, while the final average is restricted to the selected PoS group.The full iterative procedure determines optimal per-point neighborhood sizes globally before group-level averaging.

A.2 POS-Conditional Information Imbalance Across Layers

POS-conditional information imbalance tracks how token neighborhoods relate across layers, while cID measures describe category-specific geometric trajectories. The method averages over tokens of each category but compares neighbors drawn from the full sentence, enabling detection of cross-category reattachment.

  • POS-conditional information imbalance adapts a sequence-level measure to token representations within one document and conditions it on PoS.
  • The nearest neighbor and rank quantities are computed over all tokens in the sentence, irrespective of part of speech.PoS conditioning enters only through the averaging set S_p, not through the candidate-neighbor search.
  • Restricting only the averaging set allows the measure to detect cross-category reattachment, such as a determiner’s neighbor shifting to its head noun.
  • Δp → 2/N indicates preserved neighborhoods, whereas Δp → 1 indicates statistically independent neighborhoods; asymmetry carries the relevant signal.An approximately zero A→B imbalance means A predicts B for that category, rather than the reverse.
  • For each layer and PoS category, the method computes imbalance toward the first and last layers to track neighborhood preservation and convergence.A crossover between the profiles localizes when tokens detach from input-level identity and approach their final contextual role.

B.1 Semantic Absorption in Function Words

Function words become nearly as informative about document topics as content words, with this semantic absorption coinciding with localized neighborhood reorganization and compressed local geometry.

  • A downstream multiclass task tracks the emergence of semantic content in categories that carry little or none lexically.
  • Content words are highly discriminative from early layers, while function-word accuracy rises sharply around layer 19 and nearly converges with content words.The task uses PubMed abstracts labeled across 15 medical specialties, where technical nouns provide a natural upper bound and function words a natural floor.
  • In ModernBERT, determiner neighborhood detachment coincides with the shrinking accuracy gap, whereas in bigbird-roberta-large cII profiles for all function categories track that gap.
  • Layer-mean cID correlates significantly negatively with classification accuracy, strongest for DET, PRON, and ADP.The result links semantic informativeness with compressed local manifolds and interprets both as descriptions of the same dependency-resolution event.

B.2 Normalized cID trajectories

The study rescales PoS-conditioned ID trajectories to make cross-category and cross-model shapes comparable while preserving the locations of peaks and troughs.

  • Each PoS category’s layer-wise cID estimates is affine-rescaled to the unit interval because absolute cID scales differ across categories and models.The bounds use the confidence band, ensuring the full interval lies within [0, 1].
  • The rescaling preserves relative peak and trough positions across layers while enabling direct comparison of trajectories with different baseline dimensionality.Figure 4 presents the normalized PoS-conditioned ID trajectories for the four language models.

C Point-wise Geometric Features

Point-wise ID and information-imbalance trajectories convert group-level geometric analyses into token-level features for grammatical-role classification, with results reported across models and PROPN settings.

  • Group-level cID and cII averages are replaced by one geometric feature vector per token for classification experiments.The move is motivated by systematic PoS variation in intrinsic dimension, which makes a single global estimate insufficient.
  • The local ID estimator models distances to k nearest neighbors within each token’s context representation space.This produces a layer-wise value reflecting the geometric complexity surrounding the specific token.
  • Per-token information imbalance retains each token’s exact neighborhood-rank changes to the first and last layers, forming ID and II trajectory features.Tables 3 and 4 report per-model and combined-model results, while Tables 5 and 6 include PROPN.
Loading 2608.25166v1…