Source-linked AI summary
Asymmetric Discourse Homogenization and Shared Language Technology: Evidence from Reddit
Fengming Liu
TL;DR
The paper asks how shared language technology reshapes political discourse across communities with different pre-existing norms. Using six million Reddit comments and multiple identification strategies, it finds that conservative communities homogenized around late 2022 while progressive communities showed no comparable change.
Problem
The paper examines how shared language technology reshapes political discourse across communities with different pre-existing discursive norms.
Method
The study analyzes six million comments from two cross-partisan political subreddits using multiple causal designs and temporal specifications.
Results
Conservative communities experienced a within-group similarity shift around late 2022, while progressive communities showed no comparable change.
Takeaways & Limitations
The findings establish a reproducible asymmetry in community-level discursive convergence across ideological groups.
Takeaways & Limitations
The cumulative AI exposure index is a single time series, and the data cannot cleanly separate ecological convergence from concurrent secular change.
Abstract
from arXiv · showhide
I document an ideologically asymmetric break in the pre-existing diversification trend of political discourse, emerging around late 2022, using 6 million Reddit comments from two cross-partisan forums, 2019-2025. Conservative users experienced an interruption of their prior diversification trajectory; progressive users showed no comparable change. The asymmetry is consistent across estimation strategies (ITS, DiD, RDiT, propensity-score matching) and temporal aggregations. A daily-frequency permutation test over 2,377 candidate cutoff dates shows the ChatGPT threshold produces an unremarkable estimate (49.8th percentile): the shift builds gradually instead of breaking at a single date. A continuous cumulative LLM index, tracking AI exposure across seven model releases, remains significant under a quadratic trend specification that eliminates the binary estimate. A stayer analysis narrows the mechanism: the homogenization effect disappears when the sample is restricted to authors active throughout the study period, and the stayer confidence interval excludes within-author effects even a tenth the size of the full-sample estimate. The mechanism is most parsimoniously ecological (community-level discursive convergence) rather than individual-level AI adoption, though the data cannot cleanly separate this account from concurrent secular change.
1 Introduction
Using 6 million Reddit comments from 2019–2025, the paper finds asymmetric discourse homogenization: conservative communities’ similarity increased around late 2022, while progressive communities showed no comparable change. Multiple analyses indicate gradual, community-level accumulation rather than a single ChatGPT-timed break or within-author effect.
- Timing and robustness: 49.8th percentile for the right-wing shift and 49.9th for the right-left asymmetry show that the ChatGPT threshold is not unusually large among 2,377 candidate cutoffs.The shift is detectable at many possible dates, indicating gradual build-up rather than a dated break.
- Timing and robustness: The asymmetry holds across ITS, DiD, and RDiT estimation strategies and monthly, weekly, and daily aggregation levels.Propensity-score matching yields a significant differential of d = −1.11, p = 0.004, while RDiT detects the right-wing discontinuity with p < 0.001 in all specifications.
- Mechanism: +0.0026, p = 0.003 for the continuous cumulative LLM index remains significant under a quadratic trend, whereas the binary post dummy is not (p = 0.667).The index tracks accumulated AI exposure across seven major model releases, supporting a cumulative process rather than a dated break.
- Mechanism: γ = −0.0001, 95% CI [−0.00016, +0.00003], p = 0.194: the homogenization effect disappears among authors active before and after ChatGPT.The stayer confidence interval excludes within-author effects even a tenth the size of the full-sample γ = +0.0081 estimate, favoring a community-level mechanism.
2 Theoretical Framework
The framework asks whether shared language technology compresses discourse unequally across ideological communities and whether community cohesion shapes resistance to this pressure. It distinguishes individual-level AI adoption from ecological, community-level narrowing, while treating cohesion predictions as exploratory because the data do not support the simple version.
- Motivation: Shared language technology may compress discourse unequally across ideological communities, motivating an ecological question about communities’ capacity to resist homogenizing pressure.Group-polarization theory emphasizes amplification among like-minded individuals, while AI-writing research links language models to reduced collective diversity.
- Community cohesion: The simple hypothesis that low-cohesion communities are more permeable is unsupported: embedding-space cohesion does not predict which right-wing communities homogenize.The cohesion–change correlation is partly a regression-to-the-mean artifact, so the discussion remains exploratory rather than a set of confirmed predictions.
- Competing mechanisms: The individual-level mechanism predicts aggregate change through many independent within-author shifts caused by specific users adopting AI-assisted writing.This pathway predicts within-author change among users active both before and after the threshold.
- Competing mechanisms: The ecological mechanism predicts community-level narrowing of shared discursive space affecting participants regardless of their personal relationship to AI, without within-author change.The design tests this divergence in Section 4.6.
3 Data and Methods
The study analyzes a filtered, ideologically classified Reddit panel spanning 84 months and 2,557 days, using DeBERTa embeddings to measure within-group semantic similarity. Identification combines interrupted time series with complementary comparative, matching, temporal-discontinuity, cumulative-index, and stayer analyses.
- Data and sample: 6,056,593 comments from r/AskALiberal and r/AskConservatives span January 2019 through December 2025 after filtering deleted and bot content.The analytic sample retains comments ≥20 characters with identifiable political flairs and covers 84 months and 2,557 unique days.
- Ideology classification: Commenters are classified as Left, Right, or Center using keyword sets covering 136 political flairs, with 99.2% accuracy; Center flairs are excluded from Left-Right analyses.The classifier matched self-declared ideology in all but 4 cases in a random 500-flair subsample.
- Semantic measurement: 768-dimensional DeBERTa v3 embeddings yield mean pairwise cosine similarity for each group-period cell, with higher values indicating more semantically homogeneous discourse.An accumulation formulation reduces memory complexity from O(n2) to O(d), enabling daily computation across thousands of group-day cells.
4 Results
Results show an ideologically asymmetric interruption of diversification: right-wing discourse shifts toward homogenization after late 2022, while left-wing discourse shows no comparable disruption. The pattern is robust to alternative specifications and is more consistent with gradual cumulative exposure than a single dated shock.
- Robustness: Left-wing discourse shows no comparable disruption, while right-wing discontinuities remain significant across monthly, weekly, bi-weekly, and daily aggregations.The right-wing discontinuity has p < 0.001 in every aggregation; left-wing p-values are 0.982, 0.469, 0.469, and 0.454, respectively.
- Permutation test: 49.8th percentile: the ChatGPT cutoff is unremarkable among 2,377 candidate dates, ruling out a ChatGPT-specific event-driven break.A total of 1,194 randomly chosen dates produce larger right-wing estimates, and the shift builds across many possible dates.
- Functional-form sensitivity: β = +0.0026 (p = 0.003): the cumulative LLM index remains significant under a quadratic trend that eliminates the binary treatment indicator.It is the only AI-related predictor surviving that quadratic specification and adds explanatory power beyond the trend alone, F(1, 81) = 57.7.
- Mechanism: The evidence supports community-level discursive convergence: the measure tracks community variation rather than compositional shifts, consistent with an ecological homogenization mechanism.The community-level account predicts progressive shaping of the public discursive space rather than a break tied to one specific release.
5 Discussion
The discussion argues that right-wing discourse homogenization accumulated gradually through community-level processes rather than a single release or individual AI adoption. The effect is robust across specifications but remains subject to multiple-comparison, identification, measurement, and generalizability limitations.
- Robustness and timing: Right-wing communities homogenized more than left-wing communities around late 2022 across estimation strategies, aggregation levels, and matching specifications.The magnitude is specification-sensitive.
- Robustness and timing: p = 0.003 for the cumulative index under a quadratic trend, whereas the binary post dummy yields p = 0.667.This supports accumulation over an instantaneous break and away from individual adoption tied to one release date.
- Robustness and timing: 2,377 candidate cutoff dates place the ChatGPT threshold at the 49.8th percentile for the right-wing shift and 49.9th percentile for the asymmetry.No single date produces an unusually large estimate; the shift builds over time.
- Mechanism: +0.0081 ITS level shift represents a 0.87% increase in within-group similarity relative to the pre-treatment mean, approximately 0.8 standard deviations.The ecological mechanism explains the stayer null without requiring personal AI use.
- Mechanism: β = +0.0006 predicts next-month homogenization from community-level AI-influenced content exposure, stronger after ChatGPT at βambient×post = +0.0003.By contrast, author AI-likeness predicts diversification at β = −0.0007.
- Mechanism: Right-wing entropy dispersion contracted more sharply, with d = −2.49 versus −1.57, while the entropy variance ratio converged from 1.03 to 1.00.This fits a low-variance signal compressing the more dispersed register.
- Limitations: The core asymmetry survives Bonferroni correction: the two right-wing ITS level shifts remain significant at p < 0.001 and p = 0.002.The threshold is α = 0.05/4 = 0.0125, although dozens of hypotheses were tested without formal multiple-comparison correction.
- Limitations: The analysis cannot fully separate concurrent treatments, AI accumulation from time passing, compositional from ecological channels, or broader population effects.The DeBERTa effect is statistically robust but small on the 0.90–0.95 similarity scale, and Reddit users are not representative of the general population.
6 Conclusion
The paper finds a reproducible late-2022 asymmetry in discourse homogenization: conservative communities became more similar, while progressive communities showed no comparable change. The shift builds gradually and supports an ecological interpretation, but remains sensitive to modeling choices and cannot be cleanly separated from secular change.
- Core finding: Late-2022 similarity increased within conservative communities, while progressive communities showed no comparable change.The asymmetry was robust across estimation strategies, aggregation levels, and matching methods.
- Timing and specification: 2,377 candidate cutoff dates produced no unusually large estimate, indicating that the shift builds gradually rather than at one date.The magnitude and significance remain sensitive to trend functional form, matching strategy, and binary-versus-continuous parametrization.
- Interpretation and limits: Community-level discursive convergence best fits the evidence, but distinguishing it from concurrent secular change requires argument-level or network-level data.The conclusion is therefore a bounded finding about what large-scale cross-partisan discourse data can establish regarding shared language technologies.
Appendix
The appendix reports pre/post semantic-similarity comparisons and difference-in-differences estimates, while showing that a sign contradiction arises from parallel-trends specification. Group-specific slopes reveal a significant pre-period gap for AskCons. and a marginal gap for AskALib.
- Appendix: The appendix presents pre/post semantic-similarity comparisons by group and difference-in-differences estimates.Pre denotes before December 2022, and Post denotes December 2022 onward.
- Appendix: The sign contradiction in the DiD reconciliation is attributed to the parallel-trends specification.The pooled DiD forces parallel trends where they fail.
- Appendix: Group-specific slope interactions yield p = 0.010 for AskCons. and p = 0.005 for AskALib.The specification uses t + Right × t slopes, HC1 standard errors, and NW-HAC with 6 lags.
- Appendix: −0.00038 per month is the AskCons. pre-period slope gap, with p < 0.001.The gap is measured as Right minus Left.
- Appendix: −0.00008 is the AskALib. pre-period slope gap, with p = 0.054.The gap is measured as Right minus Left.