Source-linked AI summary

The Geometry of Culture: Analyzing Meaning through Word Embeddings

Austin C. Kozlowski, Matt Taddy, James A. Evans

arXiv:1803.09288v1cs.CL

TL;DR

Existing cultural text analysis is difficult to replicate, so the paper uses word embeddings to model cultural meaning and finds that they reveal persistent shifts and subtle differences in gender and class associations.

  • Problem

    Cultural analysis often relies on techniques that are difficult to replicate and on analysts’ intuition and finesse.

  • Method

    The paper uses word embeddings that position words with similar textual contexts nearby while remaining naïve about what words signify.

  • Results

    The analysis traces persistent shifts in the relation between gender and class in American culture and subtle differences in the meanings of cultural markers across the 20th century.

  • Takeaways & Limitations

    Word embeddings support investigating changing relations among cultural markers across historical periods and national contexts.

  • Takeaways & Limitations

    Embedding associations can reflect textual patterns without always reflecting material or historical realities.

Abstract

from arXiv · show

We demonstrate the utility of a new methodological tool, neural-network word embedding models, for large-scale text analysis, revealing how these models produce richer insights into cultural associations and categories than possible with prior methods. Word embeddings represent semantic relations between words as geometric relationships between vectors in a high-dimensional space, operationalizing a relational model of meaning consistent with contemporary theories of identity and culture. We show that dimensions induced by word differences (e.g. man - woman, rich - poor, black - white, liberal - conservative) in these vector spaces closely correspond to dimensions of cultural meaning, and the projection of words onto these dimensions reflects widely shared cultural connotations when compared to surveyed responses and labeled historical data. We pilot a method for testing the stability of these associations, then demonstrate applications of word embeddings for macro-cultural investigation with a longitudinal analysis of the coevolution of gender and class associations in the United States over the 20th century and a comparative analysis of historic distinctions between markers of gender and class in the U.S. and Britain. We argue that the success of these high-dimensional models motivates a move towards "high-dimensional theorizing" of meanings, identities and cultural processes.

1 University of Chicago, Department of Sociology · 2 University of Chicago, Booth School of Business · 3 Amazon

The paper introduces neural-network word embeddings as a high-dimensional method for analyzing cultural meaning, showing that vector-space dimensions capture culturally salient categories and associations. It demonstrates applications to contemporary validation, historical change, and cross-national comparison.

  • 3 Amazon: Word embeddings address the challenge of representing rich, complex corpora in forms that are intuitively understandable, analytically useful, and theoretically relevant.The paper also emphasizes the need to evaluate statistical significance and discipline interpretation when turning text into data.
  • 3 Amazon: Word embedding models capture more complex semantic relations than prior computational text-analysis methods, enabling analysis of cultural categories and associations.Each word is represented as a vector in a shared high-dimensional space.
  • 3 Amazon: Across these applications, the findings demonstrate the broad utility of word embedding models for cultural analysis and motivate further development in sociology of culture.The proposed approach is intended to recover subtle and complex cultural associations from large collections of text.
  • 3 Amazon: Vector-space dimensions correspond to meaningful cultural dimensions such as race, class, and gender, while word positions capture relations within those categories.Words sharing contexts are positioned nearby, whereas words inhabiting different contexts are farther apart.
  • 3 Amazon: Projecting occupations onto a gender dimension places traditionally feminine occupations such as “nurse” and “nanny” opposite traditionally masculine occupations such as “engineer” and “lawyer.”Local contexts associated with gendered pronouns and terms nudge occupation vectors toward opposite poles.
  • 3 Amazon: The models can analyze cultural differences and change by comparing embeddings trained on different social groups, cultures, or historical periods.These comparisons examine differences in meanings and categories across contexts and shifts in cultural associations over time.
  • 3 Amazon: The analyses compare embedding-derived semantic relations with survey measures to assess ecological validity for widespread race, class, and gender associations.The paper also analyzes twentieth-century United States texts and texts from distinct cultures.
  • 3 Amazon: The longitudinal analysis identifies slow but persistent changes in the interrelationship of gender and class categories in twentieth-century United States texts.A cross-national analysis compares the United States and Great Britain to identify differences in class and gender markers.

FORMAL TEXT ANALYSIS IN THE STUDY OF CULTURE

Text is central to cultural analysis, but dominant qualitative methods are difficult to reproduce, have limited inter-coder reliability, and cannot efficiently analyze very large corpora. Formal methods such as semantic networks and topic models improve systematic analysis yet remain poorly suited to representing multifarious associations, cultural valences, and relations among cultural categories.

  • Qualitative approaches: Qualitative text analysis is constrained by weak reproducibility, low inter-coder reliability for complex themes, analyst intuition, and the pace of human reading.These limitations make close reading and qualitative coding poorly suited to very large corpora or entire sociocultural domains.
  • Formal methods: Semantic network analysis represents words as nodes linked by textual co-occurrences and uses network structure to study word relationships and conceptual organization.It can examine central words and bridges between semantic or cultural holes.
  • Formal methods: Topic modeling learns sparse word distributions for inductively discovered topics, enabling analysis of polysemy and heteroglossia across documents.Polysemy is traced through words appearing in multiple topics, while heteroglossia is traced through mixtures of topics across documents.
  • Limitations of formal methods: Existing formal methods struggle to represent multifarious word associations and cultural valences or answer how objects are positioned on dimensions such as masculine/feminine, good/bad, and high/low-status.They also do not readily capture relationships between cultural categories, such as how a culture’s good/bad distinction relates to its masculine/feminine distinction.
  • Formal methods: Large semantic networks become difficult to interpret because topological metrics fail to distinguish conceptual distance, while dense weighted networks produce an analytically unwieldy “hairball.”Topic modeling instead distorts optimal geometric distances by forcing sparse, human-readable representations.

WORD EMBEDDING MODELS AND COMPLEX SEMANTIC RELATIONSHIPS

Word embeddings represent words as vectors whose geometric relationships encode semantic structure from shared contexts. Their high-dimensional geometry supports semantic neighborhoods, cultural associations, and analogy-solving through vector arithmetic.

  • Contextual Geometry: Word embeddings place words with similar contexts near one another and disconnected-context words farther apart, allowing semantic information to be read from surrounding vectors.Words need not co-occur: shared contextual associations can position terms such as “doctor” and “lawyer” nearby.
  • Vector Geometry: Proximity is typically measured by cosine angle between normalized vectors, reflecting distances along a hypersphere rather than Euclidean straight-line distance.Normalization places vectors on the surface of a hypersphere, where angular proximity corresponds to surface distance.
  • Word2vec Method: Word2vec learns word vectors with a shallow neural network that predicts words from context, using either continuous bag-of-words or skip-gram architectures.CBOW predicts a center word from surrounding words, whereas skip-gram predicts context given a word.
  • Semantic Dimensions: High-dimensional geometry enables direct identification of semantic dimensions onto which other words meaningfully project, beyond relations between nearby words.The models’ analogy-solving ability derives from both their number of dimensions and their underlying geometry.

CULTURAL DIMENSIONS OF WORD EMBEDDINGS

Word embeddings provide a geometric method for identifying cultural dimensions such as gender, class, and race, positioning words along them to analyze shared associations and cultural relations. The approach extends bias detection by interpreting embedding dimensions as meaningful cultural categories across contexts and time.

  • Method: The method constructs cultural dimensions by averaging antonym-pair differences, then projects normalized word vectors onto them to compare words’ cultural connotations.Gender dimensions can use pairs such as man–woman, he–she, and male–female; class and race use pairs such as rich–poor and black–white.
  • Illustration: Sports projections illustrate the method: softball and tennis are feminine, soccer is nearly unassociated, and hockey, boxing, and baseball are masculine.These positions are obtained by projecting sports names onto the gender dimension.
  • Cultural associations: Embedding positions correlate with both unconscious associations measured by implicit tests and widely shared conscious associations measured by surveys.The authors argue that the method detects shared cultural meanings, not only hidden biases.
  • Cultural associations: Examples show that “doctor” is more white than black and “scientist” more masculine than feminine in the embedding associations.The paper treats such patterns as cultural associations to interrogate rather than simply distortions to remove.
  • Contribution: The framework interprets embedding dimensions as meaningful cultural categories and uses them to illuminate complex relations within, across, and over time.This extends prior work focused primarily on bias and establishes word embeddings as tools for sociological and cultural inquiry.

WORD EMBEDDINGS AND CULTURAL THEORY

Word embeddings model meaning relationally, positioning words by contextual usage and relative location in a high-dimensional semantic space. This framework supports cultural analysis of binary dimensions, intersectional categories, varied cultural capitals, and actor-specific perspectives.

  • WORD EMBEDDINGS AND CULTURAL THEORY: Word embeddings derive meaning from contextual usage and relative position among words, rather than intrinsic referents, operationalizing a fundamentally relational theory of meaning.Their representations align with practice-oriented and structuralist accounts in which meaning emerges through relations within broader systems.
  • WORD EMBEDDINGS AND CULTURAL THEORY: Constructing dimensions from antonym pairs such as man−woman and rich−poor preserves the structuralist insight that culture is organized through binary dimensions.The method remains free of many structuralist assumptions while retaining binary cultural classifications.
  • WORD EMBEDDINGS AND CULTURAL THEORY: Unlike low-dimensional correspondence-analysis spaces, embeddings preserve semantically rich, high-dimensional cultural relations, enabling theories of multiple distinct cultural capitals and status dimensions.Their use of large corpora positions words in high-dimensional spaces without requiring reduction for interpretation.
  • WORD EMBEDDINGS AND CULTURAL THEORY: Embeddings make intersectional analysis tractable by locating objects simultaneously across categories such as race, class, and gender and comparing associations at their intersections.Comparing words high on masculinity and whiteness with those high on masculinity and blackness can reveal how masculinity differs across racial lines.
  • WORD EMBEDDINGS AND CULTURAL THEORY: The framework models intersectional identity as a high-dimensional tensor of hundreds or thousands of cultural associations rather than a low-dimensional matrix.This follows from the models’ empirical success in representing cultural dimensions and their ability to represent cross-cutting identities.
  • WORD EMBEDDINGS AND CULTURAL THEORY: Because embeddings represent cultural spaces from a given vantage point, they can illuminate how distinct social actors perceive cultural fields and enable formal perspective comparisons.This extends their similarity to Bourdieusian social spaces while emphasizing that the represented relations are perspective-specific.

DATA & METHODS

The study combines varied data sources to evaluate word embeddings for cultural analysis. It validates embedding-based associations against survey responses, then uses Google Ngrams for historical and cross-national analyses.

  • The study draws on multiple, varied data sources to demonstrate word embeddings’ utility for cultural analysis.
  • A survey measures associations between everyday objects and cultural categories of race, class, and gender for ecological validation.
  • Survey associations are compared with associations calculated by embedding models trained on contemporary text sources.
  • The contemporary models include three publicly available pretrained models and a fourth trained on the Google Ngrams corpus, which supports historical and cross-national analyses.

Survey of Cultural Associations

The study surveyed 398 U.S.-based Mechanical Turk respondents to compare reported cultural associations with those represented in word embeddings. Respondents rated 59 items across gender, race, and class dimensions, producing weighted mean association scores from 0 to 100.

  • Sample: Post-stratification weights adjusted responses by race, education, and sex, while unweighted analyses produced substantively similar findings.The reported results use the weighted models.
  • Survey design: Respondents rated 59 items on gender, race, and class scales spanning occupations, foods, clothing, vehicles, music genres, sports, and first names.The topical diversity was intended to test cultural associations across different subjects.
  • Measurement: For each item, weighted mean responses yielded a 0-to-100 cultural-association rating on each of the gender, class, and race dimensions.The resulting means served as estimates of general cultural association.

Word embedding data

The study uses three public word embeddings plus a Google Ngram model to represent contemporary and historical cultural associations. It validates robustness across corpora and applies decade- and country-specific models to study cultural change and comparison.

  • Historical change: For twentieth-century U.S. change, the authors train ten independent word-embedding models on Google Ngram 5-grams, one for each decade from 1900–1909 through 1990–1999.Comparing these models traces macro-cultural change over the 100-year period.
  • Cross-cultural comparison: For cross-cultural comparison, the authors train separate U.S. and U.K. models on Google Ngram texts published between 1890 and 1910.The twenty-year window is intended to provide enough text to capture subtle but meaningful cultural traces, enabling comparison of gender and class associations.
  • Limitations: Although Google Ngram’s yearly corpus composition may not represent total literary output, embedding positions depend on shared contexts rather than word frequency.The authors argue this makes corpus-composition concerns less problematic than in analyses based on raw frequency counts.
  • Validation: Gender, class, and racial associations are similarly captured across books, news, Wikipedia, and webpages, suggesting these associations are shared across textual genres.The analyses also find that gender-name classification accuracy remains high across all ten decades.

RESULTS

Word embeddings recover cultural dimensions that align strongly with survey associations and classify culturally patterned words accurately. Their dimensions also reveal intersectional and ideological associations while showing that cultural meaning remains distributed across many dimensions.

  • Dimensionality of cultural meaning: Gender, race, and class explained only 0.57%, 0.48%, and 0.47% of total variance, while the top principal component explained 3.65%.The results indicate that the embedding space cannot be reduced to a very low-dimensional model without losing much cultural information.
  • Survey validation of models: Class correlations remained substantial but weaker than gender, ranging between 0.40 and 0.60, while racial correlations ranged from 0.17 to 0.70 across models.Google News performed best for racial associations at 0.70, whereas Google N-grams performed relatively poorly at 0.17.
  • Classification performance: Correct classification exceeded 80% in most substantive domains and often 90%, reaching 100% for first names but only 51.8% for clothing items.Embedding performance was stronger where cultural associations were more pronounced.
  • Applications to cultural categories: Projecting words across multiple dimensions revealed intersectional associations, including jazz as both African American and high-class, while hip hop and rap tended working class.The method also depicted class-stratified ideological markers, such as Subaru and Prius for upper-class liberalism and golf and steak for high-class conservatism.

Analyzing Historical Change in Cultural Dimensions

Decade-specific word embeddings reconstruct twentieth-century changes in gender and class associations, showing both occupation-specific trajectories and a broad shift in how class relates to gendered and cultural dimensions.

  • Historical corpus and method: Word embeddings trained independently on U.S. Google Ngrams by decade trace gender and class associations from 1900 through 1999.Words were projected onto gender and class dimensions in models trained with word2vec for each ten-year period.
  • Historical corpus and method: Because Google Ngrams recovered gender and class associations better than racial associations, the historical analysis was limited to gender and class.Changing racial terminology would also require decade-specific construction of the race dimension.
  • Occupation trajectories: Nurse became less strongly feminine, engineer less masculine, and journalist shifted from masculine to increasingly feminine across the twentieth century.Most associations remained stable between adjacent decades, while journalist changed more dramatically and gradually.
  • Occupation trajectories: Class trajectories generally diverged from gender trajectories: journalist became more upper-class while more feminine, nursing rose in class as it became less feminine, and engineer remained middle-spectrum.Journalism’s mid-century movement toward upper-class association aligns with its historical professionalization.
  • Relations among cultural dimensions: Class and gender were strongly associated in early twentieth-century U.S. embeddings, with femininity linked to higher class, but became orthogonal in the century’s final decades.The early association remained stable through the first half of the century before diminishing rapidly.
  • Relations among cultural dimensions: Class shifted from associations with beauty and grace toward education and employment, increasingly identifying upper-class people as educated and employed.Early upper-class associations included beautiful, graceful, and fine, while lower-class associations included ugly, awkward, and coarse.

Cross-National Comparison: United States and Great Britain

Word embeddings reveal similar broad gender and class conceptions in the United States and Great Britain around 1900, but markedly different national markers and associations. These differences include stronger British links between class and social rank, sharper American links between class and race, and stronger British feminization of colonized regions.

  • Cross-National Comparison: United States and Great Britain: Around 1900, the United States and Great Britain shared similar gender and class conceptions, while the specific markers carrying those connotations differed markedly.Four of the five dimensions closest to gender in U.S. text also ranked among Britain’s top five; both associated masculinity with ruggedness, loudness, toughness, and boldness, and femininity with delicateness, softness, tenderness, and timidity.
  • Cross-National Comparison: United States and Great Britain: British class associations emphasized rank and nobility, while American associations emphasized cotton, wilderness, and tenement, reflecting distinct national class markers.“Rank” and “nobility” were more upper-class in Britain, “colonies” were upper-class there but not in the United States, “cotton” and “wilderness” were more affluent in U.S. text, and “tenement” was sharply lower-class only in the United States.
  • Cross-National Comparison: United States and Great Britain: Gender associations varied by national context: Western powers were masculine in both countries, while Africa, Asia, and India were more feminine in Britain than in America.England and the U.S.A. were each masculine in the other country’s text, Europe was masculine in both, and Eastern colonized regions were especially feminized in British text.
  • Cross-National Comparison: United States and Great Britain: The stronger British feminization of Eastern locales is consistent with Britain’s greater nineteenth-century colonial dominance, whereas the United States lacked equally strong associations for regions it did not dominate.China was feminine in both countries, but the British text showed stronger feminization of Asia, Africa, and India than the American text.

DISCUSSION:

Word embeddings provide a high-dimensional method for analyzing cultural categories and associations while preserving complex semantic relations. The paper’s applications show that these models capture changing and cross-cultural meanings, support systematic interpretation, and motivate high-dimensional theories of culture.

  • Method: Word embeddings represent cultural meanings as relations among vectors, preserving semantic richness while enabling systematic analysis of cultural dimensions and associations.The paper emphasizes word2vec and high-dimensional vector spaces as a compact representation of large text collections.
  • Applications: Analyses reveal persistent shifts in the relationship between gender and class in American culture across decades and subtle differences in cultural markers between the United States and Britain.These applications demonstrate both longitudinal analysis within one culture and comparative analysis across cultures.
  • Theoretical implications: High-dimensional embeddings align with and extend structural and field theories of culture by revealing a stable geometry that amplifies competing cultural framings.The authors present this geometry as capturing intersecting cultural dimensions and enabling coordinated and spontaneous social action.
  • Interpretation and limitations: Systematic comparison of words and dimensions helps constrain selective interpretation of unsupervised embeddings, although meaningful analysis requires relatively large text corpora.The approach supports complete or systematically sampled interpretations, while subtle associations generally depend on extensive textual data.
  • Findings: The models capture complex cultural associations and dimensions from large bodies of text at a level previous computational methods could not approach.The authors argue that contemporary embeddings preserve more semantic information and more complex relations than earlier text-analysis methods.

APPENDIX A: SURVEY OF CULTURAL ASSOCIATIONS · APPENDIX B: SUPPLEMENTARY FIGURES

Appendix A describes a Mechanical Turk survey and its word list for comparing human cultural evaluations with word-embedding results. Appendix B provides supplementary projections of political figures and popular names onto embedding dimensions, including a two-decade lag for name analyses.

  • APPENDIX A: SURVEY OF CULTURAL ASSOCIATIONS: The survey’s word list covered occupations, clothing, sports, music genres, vehicles, food, and first names.Examples included banker, blouse, baseball, bluegrass, bicycle, beer, and Aaliyah.
  • APPENDIX A: SURVEY OF CULTURAL ASSOCIATIONS: Words marked with daggers were excluded from 2000–2012 Google Ngrams analyses because they lacked sufficient frequency for the embedding model.The marked names include Aaliyah, Hiphop, and Shanice.
  • APPENDIX A: SURVEY OF CULTURAL ASSOCIATIONS: The cultural-associations survey supplied human-rated evaluations for comparison with word-embedding models and was administered through Amazon Mechanical Turk.The survey was listed as a 15-minute “Sociological Survey” and used paid online workers.
  • APPENDIX A: SURVEY OF CULTURAL ASSOCIATIONS: Mechanical Turk results used post-stratification weighting, while additional analyses reportedly found that weighting did not substantively alter the results.The appendix also presents demographic comparisons between the Mechanical Turk and Census CPS samples.
  • APPENDIX B: SUPPLEMENTARY FIGURES: Supplementary Figure B1 projects U.S. presidents, vice presidents, and presidential candidates onto the liberal-conservative dimension of the Google News embedding.The figure concerns political figures’ positions in the embedding space.
  • APPENDIX B: SUPPLEMENTARY FIGURES: Supplementary Figure B2 projects popular boys’ and girls’ names onto the gender dimension of decade-specific Google Ngram embeddings.The analysis uses the ten most popular newborn names for each sex in each decade, retrieved from Social Security records.
  • APPENDIX B: SUPPLEMENTARY FIGURES: The name analysis imposed a two-decade lag between popularity data and the text used to train the embedding.For example, names popular in the 1880s were projected using embeddings trained on 1900s text because many names were too infrequent in contemporaneous corpus data.
Loading 1803.09288v1…