Source-linked AI summary
Diffusion of Lexical Change in Social Media
Jacob Eisenstein, Brendan O'Connor, Noah A. Smith, Eric P. Xing
TL;DR
The paper asks how lexical change spreads through online communication and whether it tends toward standardization or fragmentation. It analyzes word-use dynamics with a latent vector autoregressive model and tests demographic and geographic predictors of inferred influence networks. The results support geographic proximity and population effects while showing that racial demographic similarity is especially associated with shared linguistic influence.
Problem
Existing methods struggle to characterize online lexical change because word counts and city sizes are sparse, and prior work has limited quantitative evidence on competing geographic and demographic influences.
Method
The authors use a latent vector autoregressive model with Bayesian inference to recover diffusion networks from word counts and analyze them with post hoc logistic regression.
Results
Racial demographic similarity is the strongest predictor of transmitted linguistic influence, while geographic distance, population, and other demographic factors also contribute.
Takeaways & Limitations
Online language evolution reproduces existing social and geographic divisions rather than converging on a single unified dialect.
Takeaways & Limitations
The analysis uses metropolitan areas and word frequencies, leaving within-metropolitan diversity and linguistic levels beyond word frequencies for future work.
Abstract
from arXiv · showhide
Computer-mediated communication is driving fundamental changes in the nature of written language. We investigate these changes by statistical analysis of a dataset comprising 107 million Twitter messages (authored by 2.7 million unique user accounts). Using a latent vector autoregressive model to aggregate across thousands of words, we identify high-level patterns in diffusion of linguistic change over the United States. Our model is robust to unpredictable changes in Twitter's sampling rate, and provides a probabilistic characterization of the relationship of macro-scale linguistic influence to a set of demographic and geographic predictors. The results of this analysis offer support for prior arguments that focus on geographical proximity and population size. However, demographic similarity -- especially with regard to race -- plays an even more central role, as cities with similar racial demographics are far more likely to share linguistic influence. Rather than moving towards a single unified "netspeak" dialect, language evolution in computer-mediated communication reproduces existing fault lines in spoken American English.
Introduction
Online written language is changing rapidly, but its diffusion may reflect existing social and geographic divisions rather than a single uniform dialect. This paper develops a statistical approach to infer linguistic influence networks and explain them using demographic and geographic predictors.
- Social media has produced new written forms, including emoticons, abbreviations, phonetic spellings, and neologisms, often treated as one uniform dialect.
- The paper asks which communities influence one another online and whether written language is moving toward standardization or fragmentation.
- The authors infer linguistic diffusion networks from raw word counts and use logistic regression to relate transmission pathways to demographic and geographic predictors.
- Tracking linguistic change is difficult because it involves temporal bursts and lulls, while word counts and city sizes are long-tailed and sparse.
- Existing models emphasize interaction density, geographic distance, and population size, while cultural and racial differences can shape both diffusion and resistance to language change.
- The study supports roles for population and geography but identifies strong racial homophily in city-to-city linguistic influence.
Materials and methods
The study analyzes geolocated Twitter messages with a latent dynamical model designed to separate lexical activation from sampling and city-size effects. It then infers linguistic-diffusion networks and evaluates their geographic and demographic structure.
- Data: 107 million Twitter messages from more than 2.7 million accounts were aggregated into 165 weekly bins and assigned to 200 U.S. metropolitan areas.The corpus covers public Twitter data from 2009–2012 and includes GPS coordinates.
- Lexical data: 2,603 English words were selected from frequent terms whose usage changed significantly over time, excluding names, hashtags, and foreign-language terms.The initial 100,000 terms were narrowed to 4,854 dynamically changing terms before manual refinement.
- Latent dynamics: The latent vector autoregressive model represents each word’s underlying activation in each metropolitan area and week, using a binomial observation model with a logistic link.Latent variables are estimated from observed user counts rather than treating raw frequencies as the direct dynamical process.
- Model motivation: The model addresses unequal city sizes, changing Twitter sampling rates, changing user verbosity, short-lived global events, and many zero word counts.Raw counts require normalization, while direct frequency models remain sensitive to sampling-rate changes and other nuisance effects.
- Diffusion network: 510 dynamics coefficients survived the FDR < 0.05 threshold, defining high-probability pathways of linguistic influence.A more stringent FDR < 0.001 threshold produced a sparser network with dense regional connections and relatively few cross-country links.
- Geographic and demographic correlates: Model-linked metropolitan areas had more short-distance connections than expected by chance, and many linked cities also showed greater demographic similarity.The comparison used networks of randomly selected metropolitan-area pairs as a baseline.
Results
Linguistic influence reflects both demographic similarity and asymmetric population effects, while geographic distance remains prominent. The logistic regression was additionally evaluated through cross-validation.
- The absolute difference in the proportion of African Americans is the strongest predictor of linguistic influence between metropolitan areas.More demographically similar cities are more likely to transmit linguistic influence; Hispanic proportion, urbanization, and median income differences are also strong predictors.
- Larger cities are more likely to transmit linguistic influence to smaller cities.In B0.05, New York, Los Angeles, and Chicago have 40 outgoing connections and 15 incoming connections combined.
- Wealthier and younger cities are significantly more likely to lead than to follow in linguistic influence.The authors distinguish wealthy cities from wealthy individuals when relating this pattern to earlier sociolinguistic findings.
- The logistic regression was assessed by holding out 10% of city-pair instances and fitting the model on the remaining 90%.Predicted links were compared with directed pairs present in B(k).
Discussion
The analysis finds that demographic similarity, especially racial similarity, is strongly associated with lexical influence between metropolitan areas. It also identifies methodological scope boundaries and positions large-scale social-media analysis as complementary to traditional sociolinguistics.
- Demographically similar areas are significantly more likely to transmit language change, especially when they share racial demographics.
- The study identifies homophily between geographical communities as an important factor in the observable diffusion of lexical change.
- Metropolitan areas provide robust analytical units, but they can conceal linguistic diversity within those areas.
- A single first-order dynamics matrix across all words simplifies estimation but may miss region-specific influence for different word types.
- The analysis examines word frequencies, while future work could address structural changes such as phonetic processes.
- Large-scale social-media analysis complements traditional sociolinguistics by aggregating linguistic decisions from millions of individuals.
Appendix S2. Term examples. Examples for each term considered in our analysis.
The appendix materials describe data-processing procedures, term annotations, and preprocessing software associated with the analysis.
- Data-processing procedures include Twitter acquisition, geocoding, content filtering, word filtering, and text processing.
- Table S1 provides term annotations in a tab-separated file.
- Preprocessing software is provided as source code for data preprocessing.
Figures
The figures visualize word-frequency changes, the statistical modeling workflow, inferred linguistic-influence networks, distance comparisons, and demographic predictors of links.
- Figure 1: Figure 1 maps changing frequencies for six words across cities, with circle presence marking usage by at least 0.1% of users and area indicating probability.The six words are ion, - -, ctfu, af, ikr, and ard.
- Figure 2: Figure 2 presents the statistical modeling procedure, repeating the outlined computation across sequential Monte Carlo samples.The dotted outline identifies the repeated portion of the procedure.
- Figure 3: Figure 3 compares empirical term frequencies with Monte Carlo smoothed estimates and shows smoothed estimates of η.Empirical frequencies are circles, smoothed frequencies are dotted lines, and η appears in the right panel.
- Figure 4: Figure 4 displays significant influence coefficients among the 40 most populous MSAs, yielding 254 links at FDR < 0.001.Blue edges are bidirectional, whereas orange links are unidirectional.
- Figure 5: Figure 5 contrasts inferred networks for all 200 cities with negative networks sampled from Q, preserving empirical sender and receiver marginals.Blue lines indicate bidirectional edges and orange lines indicate unidirectional edges.
- Figure 6: Figure 6 compares distance distributions for connected city pairs in inferred networks and Q-sampled negative networks.The top histograms show model-inferred networks and the bottom histograms show negative networks.
- Figure 7: Figure 7 plots standardized logistic-regression coefficients for predicting city-pair links with 95% confidence intervals, standard errors, means, and standard deviations.The coefficients characterize predictors of links between MSA pairs.
Tables
The tables define the study’s metropolitan-area statistics and model notation, then summarize linked-pair differences and link-prediction accuracy.
- Table 1: Table 1 reports means and standard deviations for demographic attributes across the 200 metropolitan statistical areas studied.The table summarizes statistics for the MSAs included in the analysis.
- Tables 2: The notation tables define observed word counts, posting counts, empirical probabilities, latent activations, autoregressive parameters, Monte Carlo weights, networks, and Q.Q is a random network distribution matching empirical sender and receiver marginal frequencies.
- Table 3: Table 3 summarizes mean differences and standard errors between linked and sampled non-linked city pairs.These comparisons quantify how linked pairs differ from the sampled reference pairs.
- Table 4: Table 4 reports average accuracy for predicting MSA-pair links and Monte Carlo standard errors from K = 100 simulation samples.Feature groups are defined in Table 3, and “population” denotes “raw diff log population.”