Source-linked AI summary

When the Feature Pool Goes Algorithmic: Extending Mufwene's Ecology of Language Evolution to LLM-Mediated Exposure

Kunmei Han

arXiv:2608.21088v1cs.CL

TL;DR

LLMs return model-generated language to human interaction, raising questions about variant competition when deployed systems reshape speakers’ exposure. This paper treats LLMs as distributional mediators upstream of human selection, finding evidence for lexical uptake and selected pathway links but not population-level convergence.

  • Problem

    The paper asks how variant competition changes when widely deployed LLMs reshape the distribution of linguistic forms speakers encounter.

  • Method

    The paper extends Mufwene’s feature-pool ecology by treating LLMs as distributional mediators that algorithmically reweight speaker-accessible alternatives before human selection.

  • Results

    Evidence supports model-specific lexical weighting, subsequent human uptake, and selected pathway links, but does not establish population-level convergence.

  • Takeaways & Limitations

    LLMs alter the ecology of human selection without relocating language evolution’s locus from people to machines, creating testable predictions about uptake, versions, convergence, and reversal.

  • Takeaways & Limitations

    The evidence does not establish population-level convergence, and exposure does not determine subsequent speaker selection.

Abstract

from arXiv · show

Mufwene's ecological model locates language evolution in competition among variants contributed by individual idiolects and in speakers' selection from linguistic material made available through interaction. Large language models (LLMs) complicate this architecture without requiring the locus of selection to move away from human speakers. This article argues that LLMs are best treated as distributional mediators: they aggregate language produced across human populations, transform its distribution through training and post-training, and redistribute model-specific outputs at scale. I call the resulting ecological process algorithmic reweighting of the speaker-accessible distribution: model mediation can alter the relative frequencies with which competing variants reach human selectors. Emerging evidence on model-specific linguistic profiles and lexical uptake is consistent with parts of this pathway, but does not establish inevitable convergence. Human social evaluation remains decisive: model-associated forms may diffuse and become conventionalized, become socially recognizable as 'AI-like' and subsequently avoided, or fail to diffuse in the first place. The proposal extends Mufwene's feature-pool ecology one step upstream of speaker selection and yields testable predictions about uptake, model-version effects, convergence, and social reversal.

1 The ecological problem posed by LLM-mediated exposure

LLMs create feedback loops in which human-produced language is transformed and redistributed back into human communication, potentially affecting which variants become available for selection. The article extends Mufwene’s ecology upstream of human selection by treating model-specific distributional weighting as a source of testable effects on uptake, version-related change, and social reversal.

  • The ecological problem posed by LLM-mediated exposure: LLMs are trained on human language, but their outputs can re-enter human communication and subsequent human production.This creates a feedback loop in which model-generated language becomes part of the material available for later interaction.
  • The ecological problem posed by LLM-mediated exposure: Models display recurrent lexical, grammatical, and rhetorical preferences rather than neutrally reproducing human linguistic distributions.Related findings also connect model-side preferences with human uptake, while written-corpus evidence cannot by itself distinguish AI assistance from human internalization.
  • The ecological problem posed by LLM-mediated exposure: The proposal preserves human selection while extending Mufwene’s ecology to predict variant uptake, lagged changes after model-version shifts, and reversal when concentrated forms become socially recognizable.These predictions connect model behavior and human-AI interaction to how variants become available for selection and accumulate into population-level change.
  • The ecological problem posed by LLM-mediated exposure: LLMs add an upstream layer by aggregating human linguistic material, transforming output distributions through training and post-training, and shaping sampling through generation procedures.Repeated redistribution to users can change the relative frequencies with which competing variants reach human selectors.

2 Mufwene’s theoretical baseline: feature pools, access, and speaker selection

Mufwene’s ecology treats communal language as a population-level construct emerging from partially overlapping idiolects, with speakers accessing and selecting among socially structured variants. Introducing LLMs does not move selection away from human speakers: model generation probabilities remain distinct from speaker-level selection.

  • Mufwene’s theoretical baseline: Communal language is a population-level construct extrapolated from the partially overlapping idiolects of individuals sharing a communicative code.The model uses a population-genetic analogy: communal language resembles a species-level construct, while idiolects are individual-level realizations.
  • Mufwene’s theoretical baseline: Speakers contribute competing phonological, morphological, lexical, and grammatical variants to a feature pool, but individuals acquire options through particular interactions rather than wholesale.Language acquisition is characterized as recreation, and interaction-specific experience shapes the options available to learners.
  • Mufwene’s theoretical baseline: Access to linguistic material is structured by acquisition settings, interaction partners, and communication networks, which together form part of the ecology of language.Contact can place variants associated with different dialects or languages in the same pool, while their population-level effects remain interactionally conditioned.
  • Mufwene’s theoretical baseline: Competition describes unequally weighted coexistence among variants, whereas selection occurs through the development and use of individual idiolects.Variants may be favored and reproduced more often, remain marginal, or fail to spread; these outcomes arise from speaker-level selection rather than variant agency.
  • Mufwene’s theoretical baseline: LLM generation probabilities are computational weightings, not the speaker-level selection through which human idiolects form and communal language changes.The extension retains the commitment that the individual idiolect remains the locus at which selection operates.

3 What is an LLM in this ecology?

LLMs resemble human speakers functionally because prior linguistic input conditions later production, but their ecological structure differs fundamentally. They are best analyzed as distributional mediators that aggregate, transform, and redistribute population-derived linguistic variation at scale.

  • Human idiolects arise from locally limited, socially situated interaction, whereas LLM distributions are population-derived, algorithmically transformed, and redistributable at scale.Mufwene’s account emphasizes exposure to only a subset of communal production, while model development aggregates large collections of human-produced text.
  • Post-training and decoding shape LLM behavior and output diversity, producing model-specific lexical, grammatical, and rhetorical distributions unlike matched human production.Corpus comparisons identify systematic differences between resulting model outputs and matched human production.
  • LLMs and human speakers are functionally similar because prior linguistic input conditions subsequent production, but an LLM is not simply another idiolect.Both can affect the linguistic material encountered by human speakers; the difference is ontological and ecological rather than functional.
  • LLM-mediated exposure combines individualized one-to-one outputs with concentration at the provider level, presenting many users with samples from a shared model distribution.This network structure differs from ordinary human exposure, which typically draws on many distinct authors and speakers.

4 Extending the feature-pool model one step upstream

The proposed extension places LLM mediation one step upstream of human speaker selection by allowing models to restructure the distribution of variants speakers encounter. Preliminary evidence supports model-specific divergence and lexical uptake, but not population-level convergence.

  • 4 Extending the feature-pool model one step upstream: The extension asks how speakers’ accessible distribution of linguistic variants may be structured before socially situated human selection occurs.Baseline exposure varies with interlocutors, networks, demographics, communicative histories, and unequal encounter frequencies.
  • 4 Extending the feature-pool model one step upstream: Algorithmic reweighting describes model mediation altering the relative probabilities with which competing variants reach human linguistic environments.Human-produced language enters model development, training and post-training transform behavior, and model outputs circulate through direct or indirect communication.
  • 4 Extending the feature-pool model one step upstream: The theoretical claim is that P(fi | Es, M) may differ systematically from P(fi | Es), not that model-mediated input is necessarily narrower, more frequent, or more influential.M denotes the model-mediated component of exposure, while Es denotes the speaker’s broader ecology of exposure.
  • 4 Extending the feature-pool model one step upstream: Reinhart et al. (2025) show that model-generated language differs systematically from matched human production across lexical, grammatical, and rhetorical distributions.This supports the pathway’s first necessary condition: model mediation can transform human linguistic distributions.
  • 4 Extending the feature-pool model one step upstream: Yakura et al. (2026) find that ChatGPT-preferred words increased in spontaneous human speech and that AI-supplied variants were reused after interaction ended.These findings support model-side distributional divergence and lexical uptake, but not population-level convergence, which requires reduced diversity among competing alternatives.

5 Human selection after model-mediated exposure: convergence, enregisterment, and reversal

Model-mediated exposure can increase diffusion without producing convergence because human speakers remain socially situated selectors. As model-associated forms acquire AI-related indexical meanings, social evaluation may stabilize, redirect, or reverse their uptake.

  • Convergence and selection: Algorithmic reweighting may increase a model-preferred variant’s frequency without convergence, which requires competing alternatives to become less evenly distributed in human production.Diffusion and convergence are therefore distinct outcomes.
  • Convergence and selection: Concentrated model exposure constrains alternatives entering speaker-level competition but does not determine selection, because speakers remain socially situated selectors.Linguistic forms can acquire new social meanings as they diffuse.
  • Enregisterment and reversal: A model-preferred form may become socially recognizable as AI-like, acquiring indexical meanings that prompt either maintenance or avoidance.The proposed trajectory is model concentration -> increased exposure -> human uptake -> social salience -> AI indexicality -> maintenance or avoidance.
  • Enregisterment and reversal: Model reweighting and human social evaluation interact: increased exposure can be followed by stabilization, redirection, or reduced reproduction without positing a separate evolutionary mechanism.The process remains human selection operating on variants whose social meanings may have changed through model-mediated diffusion.
  • Outcome trajectories: Three outcomes follow: diffusion and conventionalization, boom and bust through costly AI indexicality, or ecological resistance from norms, identities, goals, and entrenched alternatives.Recent vocabulary trajectories suggest reversal, but reduced model supply and changed human evaluation remain distinct mechanisms.

6 Predictions and a research agenda

The research agenda derives four falsifiable predictions linking model-side distributions, human exposure, and later human production. It calls for designs measuring both sides of the relation while accounting for ecological boundary conditions and distinguishing linguistic from cognitive convergence.

  • Predictions: Four predictions test distinct links: preference uptake, version tracking, concentration-convergence, and salience-reversal.They respectively address later human change, temporal effects of model changes, variation, and human social evaluation.
  • Measurement: Testing requires controlled estimates of model-side distributions and longitudinal human data modeling time, register, and ideally speaker or author identity.Generations should record model version, prompt, and settings; analyses should cover lexical, constructional, syntactic, discourse, and information-structuring alternatives.
  • Causal designs: Complementary designs include parallel human-LLM corpora, longitudinal corpora, model-version shifts, and controlled exposure experiments testing persistence after exposure ends.These designs provide evidence on model-specific distributions, later population change, quasi-experimental variation, and persistence.
  • Boundary conditions: Uptake should be analyzed under ecological boundary conditions because diffusion can vary across communities, registers, platforms, and languages.Demography, interaction patterns, existing alternatives, and social evaluation can alter population-level outcomes.
  • Evidential limits: Linguistic convergence alone cannot establish cognitive or ideological convergence, which requires independent measures of conceptual organization, judgment, memory, or reasoning.This boundary prevents a distributional account of linguistic change from becoming an unsupported theory of cognitive homogenization.

7 Conclusion

The conclusion retains Mufwene’s feature-pool model while extending it upstream to account for LLM-mediated changes in the distribution of alternatives available before human selection. LLMs act as distributional mediators, but human social evaluation remains decisive in determining diffusion, conventionalization, avoidance, or reversal.

  • Mufwene’s core commitments remain useful: language is grounded in idiolect populations, variants are unequally weighted ecologically, and speaker-level selection drives change.
  • LLMs are distributional mediators that aggregate, algorithmically transform, and redistribute language, altering the speaker-accessible distribution before human selection.
  • Model-side reweighting can favor diffusion without determining outcomes, because speakers may reproduce, conventionalize, reinterpret, resist, or avoid encountered variants.
  • Model-associated concentration may support convergence in some conditions, but may instead make forms socially recognizable as AI-associated and trigger reversal.
  • LLMs alter the ecology of human selection rather than relocating language evolution to machines, creating an empirical setting for testing access, weighting, diffusion, and social selection.
Loading 2608.21088v1…