Source-linked AI summary

When Can We Work in Embedding Space? What Text Embeddings Preserve

Simon Freyaldenhoven

arXiv:2608.31059v1econ.EMcs.CLstat.ML

TL;DR

The paper asks when replacing text with a low-dimensional embedding preserves information needed for empirical analysis. Under a latent-topic model, it studies clustering and embedding-based controls, finding that clusters reflect topic mixtures and that control validity depends on whether those mixtures capture confounding. In 363 U.S. metropolitan areas, LLM-generated-description clusters yield interpretable economic types and separate local employment dynamics more sharply than curated industry and demographic covariates.

  • Problem

    Text embeddings may discard information needed for empirical analysis, making it unclear when replacing high-dimensional text with a low-dimensional representation is valid.

  • Method

    The paper analyzes embeddings under a generative model in which documents are mixtures of K latent topics, focusing on clustering and conditioning on embeddings as controls.

  • Results

    In 363 U.S. metropolitan areas, embedding-based clusters of LLM-generated descriptions produce interpretable economic types and separate local employment dynamics more sharply than curated industry and demographic covariates.

  • Takeaways & Limitations

    Embedding clusters correspond to documents with similar topic mixtures, while embedding-based control is valid when the topic mixture captures the confounding.

  • Takeaways & Limitations

    Exact recovery of a latent partition requires stronger conditions, and the application does not satisfy the paper’s separation condition, so it claims no latent-partition recovery there.

Abstract

from arXiv · show

When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can trade text for its low-dimensional embedding, and lose little in doing so. I make that assumption precise under a generative model in which documents are mixtures of latent topics. I study two uses---clustering units in embedding space and controlling for high-dimensional text. A cluster of embeddings is a set of documents with similar topic mixtures; controlling for the embedding is equivalent to controlling for the topic mixture, so validity reduces to whether that mixture captures the confounding. In an application to 363 U.S. metropolitan areas, embedding-based clusters of LLM-generated economic descriptions recover interpretable economic archetypes and separate local employment dynamics more sharply than clustering on model residuals, or on a curated set of industry and demographic covariates.

1 Introduction

The paper asks when low-dimensional text embeddings preserve information needed for empirical analysis and makes that question precise under a latent-topic model. It studies embeddings for clustering and as controls, linking their validity to whether topic mixtures capture the relevant confounding.

  • Motivation: Text embeddings reduce extremely high-dimensional documents to relatively low-dimensional vectors, but this reduction may discard information needed for analysis.The paper frames the central justification as an assumption that little is lost when text is represented in embedding space.
  • Approach: The paper makes embedding validity precise using a generative model in which documents are mixtures of K latent topics.This topic-model setup is presented as interpretable and familiar in economics, while also connecting to latent-variable views of language models.
  • Theory: Under the topic model, word co-occurrence structure has rank K −1, and embeddings matching that structure inherit the geometry of topic loadings.Words with similar topic loadings receive similar embeddings, while standard methods instead factorize a logarithmic transform of co-occurrence ratios.
  • Document embeddings: Averaging word embeddings produces an invertible linear image of each document’s topic mixture, connecting document-level embedding distances to latent composition.This result supports interpreting clusters as groups of documents with similar topic mixtures rather than merely nearby vectors.
  • Applications: For clustering, embedding-space groups represent similar latent mixtures, while exact recovery requires concentration and stronger separation conditions.The clustering use case is motivated by balancing pooled models’ implausible homogeneity against separate city models’ noisy estimates.
  • Applications: For control use, adjusting for an embedding is equivalent to adjusting for the topic mixture, so validity requires that the mixture captures the confounding.This translates the high-level sufficiency assumption into an explicit substantive condition rather than an algorithm-only property.

2 Theoretical Results

Under a latent-topic model, document embeddings preserve topic-mixture geometry, making clustering interpretable as grouping similar mixtures and reducing embedding-based adjustment to a condition on confounding. These results require sufficient embedding dimension and identification, while exact latent-partition recovery needs stronger separation conditions.

  • Generative model: Documents are generated by independent word draws from mixtures of K latent topics, with word probabilities Π = BΘ.The model represents each document’s word distribution through a word-topic matrix B and topic-document matrix Θ.
  • Word embeddings: Words with proportional topic loadings receive identical embeddings, while valid factorizations recover the marginal-normalized topic-loading geometry.The geometry is invariant across valid factorizations and can be recovered directly with a rank-(K −1) factorization.
  • Document embeddings: Expected document embeddings are linear in topic mixtures and lie in the simplex spanned by topic centroids, with an invertible encoding under the stated assumptions.Embedding distances can therefore be expressed as weighted differences in latent topic mixtures.
  • Clustering documents: Embedding clusters represent similarity in topic mixtures rather than necessarily similarity in dominant topics, and the number of clusters L need not equal K.This interpretation relies on the document-embedding metric and does not require concentration around a finite set of mixture-types.
  • Clustering documents: Exact latent-partition recovery requires concentration and separation conditions; in the application, those conditions are not satisfied, so the paper claims no recovery of a latent partition.The weaker result is that stable clusters correspond to similar topic mixtures.
  • Embeddings as controls: Controlling for document embeddings is equivalent to controlling for topic mixtures when the embedding map is injective, so validity depends on whether topic mixtures capture confounding.More generally, the embedding dimension must satisfy r ≥s and the embedding geometry must recover the confounding-relevant summary rather than merely provide a low-dimensional projection.

3 Numerical Illustration

Simulations compare SVD-β and SGNS embeddings under two- and three-topic designs. Both recover interpretable topic structure, while document embeddings reflect mixtures across the topic simplex and SGNS captures most relevant spectral mass.

  • Design 1: Two Topics: Both SVD-β and SGNS recover similar qualitative word geometry in the two-topic design, separating fruit and vegetable words into distinct clusters.Their geometries are nearly identical up to rotation and scaling, and both clouds are one-dimensional as theory predicts.
  • Design 1: Two Topics: Anchor words lie at topic extremes, while mixed words fall nearer cluster boundaries according to their relative topic loadings.This pattern holds exactly for SVD-β and approximately for SGNS.
  • Design 2: Three Topics: With three topics, both embeddings clearly separate non-citrus fruits, citrus fruits, and vegetables despite topics 1 and 2 lacking anchor words.Kale, the anchor for vegetables, lies on the convex-hull edge; nonanchor words are pulled toward mixtures of topic centroids.
  • Document Embeddings: Averaged document embeddings are predicted to lie in the simplex spanned by topic centroids, with similar topic mixtures producing similar embeddings.The Design 2 documents are projected into two dimensions and colored by dominant topic.
  • SGNS Approximation: πK−1 equals 0.88 for Design 1 and 0.91 for Design 2, so the leading K −1 directions capture most of log R’s spectral mass.The authors interpret this as approximate support for Proposition 2 describing trained SGNS embeddings, not only its target.

4 Application: Clustering Metropolitan Economies

The paper applies five embedding-based clustering methods to 363 U.S. metropolitan areas, finding interpretable economic archetypes and sharper separation of employment dynamics than non-text benchmarks.

  • Clustering Metropolitan Economies: 363 U.S. CBSAs are assigned to five clusters using 500-word LLM-generated economic descriptions and k-means in embedding space.The five embeddings include SVD-β, corpus-trained SGNS, corpus-trained CBOW, pretrained Word2Vec, and OpenAI’s text-embedding-3-large.
  • Clustering Metropolitan Economies: Embedding-based clusters group CBSAs whose descriptions imply similar topic mixtures, with each cluster centroid corresponding to a well-defined mixture.This gives the empirical pipeline a direct interpretation under the paper’s topic model.
  • Clustering Metropolitan Economies: The clusters identify recognizable archetypes including energy dependence, university and government anchoring, deindustrialization, and diversified growth.These qualitative groupings remain similar across different choices of L.
  • Local Employment Dynamics: SVD-β, corpus-trained SGNS, and LLM embeddings tend to perform best, while pretrained Word2Vec averaging tends to perform worst among text-based methods.A closed-form factorization computed from the corpus matches a commercial transformer embedding on this task.
  • Local Employment Dynamics: The exercise is entirely in-sample, so it does not establish that text-based peer groups are preferred for forecasting this panel.The textual descriptions also span the outcome window, unlike the effectively predetermined 1970 covariates.

5 Conclusion

The conclusion argues that embeddings preserve useful topic-mixture structure when their geometry matches the text and the analysis-relevant latent summary. In the metropolitan application, LLM-generated descriptions produce interpretable types and sharper employment-dynamics separation than curated covariates.

  • Conclusion: The paper asks what replacing text with an embedding preserves under a latent-topic generative model.It studies the question for clustering and controlling for high-dimensional text.
  • Conclusion: Document-embedding proximity corresponds to proximity in topic mixture, so clusters represent similar mixtures and embedding adjustment equals topic-mixture adjustment.This conclusion follows from the paper’s document-level theoretical results.
  • Conclusion: Embedding validity depends on two alignments: recovering the text’s low-dimensional structure and recovering the structure required by the analysis.Thus, dimension alone is not sufficient; the confounding-relevant topic mixture must also be captured.
  • Conclusion: In 363 U.S. metropolitan areas, LLM-generated descriptions yield interpretable economic types that separate local employment dynamics more sharply than curated industry and demographic covariates.SVD-β and corpus-trained SGNS perform similarly, while pretrained averaging generally trails.
  • Conclusion: Transformer embeddings remain outside the paper’s theoretical framework, motivating work on their topic structure and on dynamic topics over time.These are identified as extensions rather than resolved results.
  • Conclusion: Generating text with a language model and inspecting downstream embedding steps preserves model representational power while keeping the pipeline inspectable and reproducible.The conclusion presents this division of labor as a reasonable way to incorporate large models into empirical work.

Model

The model represents embedding algorithms as weighted low-rank fits whose population targets are determined by corpus responses and loss generators. This framework identifies the target of each algorithm and characterizes skip-gram’s full-softmax solution.

  • The algorithms as weighted low-rank fits: Each algorithm scores word pairs with an inner product and fits that score using a corpus response, convex generator, and positive pair weights.The response is fixed by the data, while the score is the rank-constrained parameter.
  • The algorithms as weighted low-rank fits: Bregman divergences measure the discrepancy between each corpus response and fitted value, with squared-error, Kullback–Leibler, and Bernoulli log-loss cases represented.The divergence is nonnegative but not symmetric and is not a metric.
  • The algorithms as weighted low-rank fits: The link function ϕ′ maps real-valued inner-product scores to the response scale, defining the fitted value used in the loss.Strict convexity makes ϕ′ invertible.
  • The algorithms as weighted low-rank fits: When the rank constraint is nonbinding, the population optimum equals the algorithm’s target entrywise and is independent of positive weights.The target is determined by the response and generator alone.
  • Skip-gram: The population skip-gram objective replaces corpus-position averages within a context window with population co-occurrence probabilities.Its score matrix collects the word-pair inner products.
  • Skip-gram: Full-softmax skip-gram targets log R plus word- and context-specific additive terms, with the context-side term reflecting softmax’s column-shift invariance.The unconstrained maximizers satisfy X∗ = log R + log q 1⊤ − 1g⊤.

A.3 SGNS

The paper formalizes how standard embedding objectives relate to topic-model co-occurrence structure and explains what document embeddings preserve. It then applies these embeddings to clustering, showing robust separation of employment dynamics across cluster counts.

  • Computationally difficult full-softmax skip-gram motivates negative sampling as the practical approximation analyzed next.
  • SGNS has unconstrained target X∗ = log R − log ν 1_V1_V^⊤ under the topic-model assumptions.
  • The weighted-Bernoulli formulation identifies the response and pair weights determining SGNS’s population target.This places SGNS in the same target-based framework as the other embedding algorithms.
  • GloVe instead targets log M entrywise after learned biases absorb marginal word frequencies, leaving the centered ratio structure log R.
  • For multi-word contexts, CBOW lacks an analogous V × V target matrix, so the topic-model factorization theory does not cover it.
  • Across cluster counts, every method remains above the permutation floor, while a text method leads at every tested L and residual k-means trails nearly everywhere.The ordering among corpus-based text methods is not stable, but text-based groupings outperform the benchmarks regardless of cluster count.

B.6 Sensitivity to the Embedding Dimension

The application is insensitive to embedding dimension for SVD-β and corpus SGNS across the tested range, whereas corpus CBOW exhibits unstable clustering outcomes. This supports the primary dimension choice without requiring knowledge of the latent topic count.

  • Any r ≥ K − 1 recovers the same population geometry, so the embedding dimension only needs to be sufficiently large.
  • SVD-β and corpus SGNS remain flat from r = 10 to r = 200, supporting r = 50 without identifying K.Their between-group ratios and cluster balance remain comparable throughout the sweep.
  • Corpus CBOW’s between-group ratio changes by a factor of two because k-means selects among near-equivalent optima.

B.7 Sensitivity to the Autoregressive Order

The autoregressive order affects estimated cross-city impulse-response dispersion, with p = 2 providing the most favorable low-dimensional comparison under the simulation benchmark. The sensitivity analysis also highlights dependence on shock assumptions, context windows, and covariate timing.

  • On the simulation benchmark, estimation error accounts for 74% of dispersion at p = 2 and at least 94% at every other order.The share reaches 100% at p = 5 and exceeds 100% from p = 6.
  • The benchmark independently redraws innovations across units, although employment shocks may be positively correlated, potentially biasing all depicted ratios upward.
  • Seven of eight approaches peak at p = 2, while all remain above the permutation floor across autoregressive orders.
  • The paper treats p = 2 as a parsimonious projection for cross-sectional comparison, not as a correctly specified dynamic model.Longer models use scarce per-unit degrees of freedom and can spread heterogeneity across noisy coefficients.
  • The SVD-β ratio falls from 21.8× at the full-document window to 12.1× at J = 5, while corpus CBOW and SGNS show no clear trend.The narrowing-window result reflects noisier estimated co-occurrence structure for SVD-β.
  • Dating the observables later raises the ratio from 15.5× at 1970 to 18.6× at 2000, approximately a 20% look-ahead gain.The paper notes that the text descriptions are also concurrent with the outcome window.

Cluster Descriptions by Method

Across methods, clustering metropolitan areas produces interpretable economic archetypes, though the groupings and balance differ by method. Text-based clusters reveal themes including amenity growth, anchor institutions, deindustrialization, energy dependence, and diversified economies.

  • Text-based methods identify recognizable archetypes including amenity migration, university and hospital anchors, manufacturing transition, and severe deindustrialization.
  • Pretrained Google News Vectors (PW): Pretrained Google News vectors yield deindustrialized manufacturing, energy and resource extraction, agricultural, and eds-and-meds regional-hub groupings.
  • Residual k-means (RK): Residual k-means forms less-balanced clusters, including a singleton and a largest group containing 47% of CBSAs, while still revealing interpretable economic patterns.
  • Closed-Form SVD-β at r = 50 (SB): SVD-β produces diversified, energy-dependent, government-anchored, manufacturing-heavy, and military-outlier groups.The groups contain 113, 17, 170, 62, and 1 CBSAs, respectively.
  • The alignment between text-based themes and employment-based labels provides qualitative external validation for the economic groupings.

C.1 Large Corpus

The large-corpus appendix compares embedding behavior at D = 1,000 and Nd = 10,000 with the main-text designs. SVD-β displays the predicted low-dimensional geometry, while SGNS uses a smaller feasible training sample.

  • At K = 2, SVD-β’s embedding cloud collapses onto a line when the corpus grows to D = 1,000 and Nd = 10,000.The larger design increases ordered word–context pairs from 8.6 × 10^7 to 1.0 × 10^11.
  • SGNS uses J = 5 at the larger scale because full-document context is infeasible, consuming roughly 6.0×10^7 pairs per epoch.
  • The effective SGNS training sample is about 30% smaller in pairs than in the main-text counterpart.
  • At K = 3, both SVD-β and SGNS separate the three topics in principal-direction projections.CBOW’s raw coordinate basis is arbitrary, so principal directions provide the comparable visualization.
  • With α = 1, document embeddings fill the triangle spanned by the three topic centroids, as Proposition 3 predicts.

C.2 Sparse Topic Mixtures: Dirichlet(α = 0.01)

Under Dirichlet(α = 0.01), topic mixtures concentrate near simplex vertices rather than spreading across the simplex. Consequently, document embeddings concentrate near topic centroids and dominant-topic clusters become cleanly separated.

  • Dirichlet(α = 0.01) makes most documents nearly pure-topic, with mixtures close to a single topic-simplex vertex.
  • Document embeddings therefore concentrate at topic centroids rather than filling the centroid simplex.
  • At α = 0.01, Pr(θd,k∗ > 0.5) ≈ 1 for both K = 2 and K = 3.
  • For K = 2, documents cluster at the two endpoints of the line segment [c1, c2] rather than filling it.
  • For K = 3, both embeddings cleanly recover the well-separated dominant-topic partition.Documents concentrate at the simplex vertices, contrasting with the α = 1 design.

C.3 Interior Archetypes (L < K)

When documents concentrate around interior archetype mixtures, embedding-space clusters reflect those archetypes rather than dominant-topic vertices. This regime allows fewer clusters than latent topics and mirrors the empirical application’s use of five clusters.

  • The empirical application uses L = 5 clusters whose centroids are interior mixtures of an unknown number of latent topics.
  • A two-archetype mixture is generated with K = 3 topics, with each document drawn from a tight Dirichlet distribution around its assigned archetype.
  • Both archetypes lie inside the topic simplex rather than at its vertices, producing the L = 2 < K = 3 regime.
  • The simplex becomes much flatter because the topic-mixture covariance is nearly rank one, despite the unchanged word-distribution matrix.
  • Cluster identity is governed by archetype assignment, not dominant-topic vertex, so L need not equal K.
  • UMAP preserves cluster structure at both application and larger-corpus scales.

C.6 Rank of the SGNS Target under a Sparser Prior

Under the sparser Dirichlet(α = 0.1) prior, the rank-(K −1) approximation captures less spectral mass and loses its invariance to vocabulary size for moderate K. The larger diagnostic departure from independence also places most cases outside the expansion used for the α = 1 analysis.

  • Experiment: The α = 0.1 results average eight population draws of B and topic mixtures, with maximum cell-level standard deviation 0.016 and no corpus sampling.The negative-sampling shift contributes one additional eigenvalue of size V log ν, which is excluded.
  • Sparse-prior results: At α = 0.1, πK−1 ranges from 0.68 to 0.89, compared with 0.91 to 0.99 at α = 1.The rank-(K −1) approximation therefore misses between a tenth and a third of the spectral mass of log R.
  • Sparse-prior results: Invariance to V breaks for moderately sized K under the sparse prior, suggesting that the rank constraint becomes costlier as vocabulary grows.
  • Departure from independence: The diagnostic ϱ ranges from 0.84 to 6.5 and exceeds one in 25 of 30 cells, whereas it stays below 0.6 throughout the α = 1 grid.
  • Departure from independence: Because Remark 3 expands log R = ββ⊤+ O(ϱ2) only under ϱ < 1, that expansion does not cover most entries in the sparse-prior table.
Loading 2608.31059v1…