Source-linked AI summary

CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Target

Qianwen Gao, Zichang Su, Yiwen Hou, Arlen Kumar, Leanid Palkhouski

arXiv:2608.30466v1cs.AIcs.IR

TL;DR

Population-level effects of repeated optimization for LLM visibility remain poorly understood. The paper introduces CHASE, a controlled simulation that iterates ranking, feature discrimination, rewriting, and evaluation, and finds that quality–ranking alignment weakens across six domains over 20 rounds. The resulting ecosystem dynamics are domain-dependent, with controls associating the divergence with ranking-derived incentives rather than rewriting alone.

  • Problem

    The paper addresses limited evidence about how repeated GEO adaptation changes content ecosystems and ranking incentives beyond single-round optimization.

  • Method

    CHASE simulates repeated content adaptation through RANK, DISCRIMINATE, REWRITE, and EVALUATE stages using an LLM ranking signal.

  • Results

    Quality–ranking alignment decreases across all six domains over the 20-round simulation, while ecosystem dynamics vary substantially by domain.

  • Takeaways & Limitations

    Repeated optimization against a fixed LLM ranker can reshape document features and the relationship between ranking success and content quality.

  • Takeaways & Limitations

    CHASE isolates a fixed ranker and a finite 20-round horizon, so its trajectories do not establish behavior under adaptive systems or long-run equilibrium.

Abstract

from arXiv · show

Generative Engine Optimization (GEO) is increasingly used to improve content visibility in LLM-based retrieval systems, yet its population-level effects under repeated optimization remain poorly understood. We introduce Content Homogenization under rAnking Signal Exploitation (CHASE), a controlled simulation framework for studying how content ecosystems are reshaped when creators repeatedly adapt documents to an LLM ranking signal. We use ranking as a proxy for source visibility and validate this abstraction against citations in grounded generated responses, obtaining a rank-citation AUC of 0.853 $\pm$ 0.093 across six domains. CHASE then iterates ranking, feature discrimination, rewriting, and evaluation over 20 rounds across different domains. Quality-ranking alignment decreases in all six domains: from R0 to R20, the change in Spearman's rho ranges from -0.107 to -0.018, with a mean change of -0.068, which means documents closer to the ranking feature profile become less aligned with independently judged document quality over the simulation horizon. A random-target control has shown that it is associated with adaptation toward ranking-derived incentives rather than iterative rewriting alone. The resulting ecosystem dynamics are strongly domain-dependent. Together, these findings show how repeated optimization against a fixed LLM ranking signal can reshape both content populations and the incentives faced by content creators.

1 Introduction

The paper asks how repeated creator adaptation to LLM ranking success reshapes content ecosystems, extending GEO beyond single-round edits. CHASE simulates this feedback loop and finds weakening quality–ranking alignment across six domains over 20 rounds.

  • Motivation: GEO research has largely studied single-round edits, whereas CHASE examines repeated, population-level adaptation to ranking-associated features.The paper asks how competing creators’ repeated adoption of successful strategies changes document distributions and ranking incentives.
  • Framework: CHASE iterates RANK, DISCRIMINATE, REWRITE, and EVALUATE stages to simulate ranking-driven content adaptation.An interpretable discriminator identifies features associated with ranking success, and a stochastic subset of documents is rewritten toward the resulting target profile.
  • Findings: Across six domains and 20 rounds, proximity to ranking-derived feature profiles becomes less aligned with independently judged document quality.The evaluation uses separate model families for ranking, rewriting, and quality evaluation, plus five random seeds.
  • Controls: Random-target and no-rewrite controls indicate that the observed quality–ranking divergence is associated with ranking-driven adaptation rather than rewriting alone.The study also analyzes how ranking-predictive features evolve, including directional associations with evidentiary features.

2 Related work

Related work connects CHASE to proxy optimization failures, generative-engine visibility interventions, and competitive adaptation in search. These strands motivate studying repeated strategic document adaptation under LLM ranking signals.

  • Goodhart effects in model optimization: Goodhart-type failures describe weakened alignment between an optimized proxy and the underlying objective.Related machine-learning work includes reward hacking and degradation under continued optimization of learned reward models.
  • Generative Engine Optimization: GEO studies content interventions that affect visibility in generated responses and recommendation outcomes.Examples include statistics, quotations, citations, and strategic product-description modifications.
  • Competitive search and strategic adaptation: Information retrieval research has examined adversarial search manipulation and publisher adaptation to ranking incentives.Recent work extends this setting to LLM-based publishers and ranking competitions.

3 The CHASE Framework

CHASE simulates how documents repeatedly adapt to a fixed LLM ranking signal through ranking, feature discrimination, rewriting, and evaluation. The framework uses controlled assumptions and evaluates evolving document populations, quality–ranking alignment, ecosystem properties, and grounded responses.

  • 3.1 Scope and Problem Formulation: CHASE models document-pool transitions through four stages: RANK, DISCRIMINATE, REWRITE, and EVALUATE.The framework isolates content-side adaptation to ranking incentives rather than simulating the full retrieval and response-generation pipeline.
  • 3.1 Scope and Problem Formulation: CHASE assumes myopic, non-strategic creators and a fixed ranker, making it a controlled stress test rather than an equilibrium model.The discriminator also provides an idealized, cleaner inference signal than creators may observe in practice.
  • 3.2 Stage 1: Rank: The ranker orders associated documents for each query, with five shuffled presentation orders aggregated by mean rank.Domain prompts frame the model as a recommendation engine for Retail, Video Games, and Books, or an information-retrieval engine for Web, News, and Debate.
  • 3.3 Stage 2: Discriminate: The discriminator labels the top 10% as winners, fits L2-regularized logistic regression, and selects five features with the largest absolute coefficients.The selected features and their winning-document means form the inferred ranking signal passed to rewriting.
  • 3.4 Stage 3: Rewrite: Only a stochastic subset of non-winning documents is rewritten each round, with expected participation probability 2/7 ≈0.29.Participation is independently resampled across documents and rounds, and rewrites are instructed to preserve factual claims and avoid invented information.
  • 3.5 Stage 4: Evaluate: Evaluation combines document-level ecosystem metrics, independently judged document quality, quality–ranking alignment, grounded-response measures, and human validation.Discriminator AUC, homogeneity, and ranking stability are computed every round; quality and response-level metrics are evaluated at rounds 0, 5, 10, 15, and 20.

4 Results

Across six domains, repeated ranking-derived adaptation reduced quality–ranking alignment over 20 rounds, while controls and audits indicate the pattern was not caused by rewriting alone or accumulating corruption. The magnitude and form of ecosystem change varied by domain, with ranking-predictive features becoming more separable and domain-specific regimes emerging.

  • Ranking validation: 0.853 ± 0.093 rank–citation AUC across six domains supports ranking as a useful proxy for source visibility in CHASE.AUC remained positive in every domain and was stable from 0.849 at R0 to 0.863 at R20.
  • Quality–ranking divergence: −0.018 to −0.107 was the R0–R20 decline in quality–ranking alignment across the six domains.Alignment remained positive at R20 in every domain, but documents closer to ranking-associated feature profiles became less strongly associated with independently judged quality.
  • Quality–ranking divergence: Mean document quality did not necessarily decline monotonically, so the divergence reflects a less informative ranking-derived optimization direction rather than universal quality degradation.Mean quality was nearly unchanged in several domains and increased in Retail.
  • Ecosystem dynamics: Discriminator AUC increased in every domain, whereas homogeneity and presentation-order stability were comparatively stable except for a substantial Web stability decline.These metrics motivate treating CHASE as domain-dependent dynamics rather than one universal trajectory.
  • Controls: No-rewrite controls left homogeneity unchanged in Retail (.732 → .732) and Debate (.633 → .633), indicating population changes required active rewriting.Ranking and feature extraction continued while documents remained frozen.
  • Controls: Random-target controls produced smaller or mixed ∆ρ changes than CHASE, associating divergence with adaptation toward ranking-derived targets rather than arbitrary iterative rewriting.For Retail, Debate, and News, ∆ρ was +.001, +.014, and −.036 under random targets versus −.047, −.065, and −.104 under CHASE.
  • Ecosystem dynamics: Retail, Debate, and News exemplified structural convergence, signal instability, and feature dominance, respectively, within the observed 20-round horizon.The regimes are descriptive patterns, not distinct failure classes or long-run endpoints.
  • Robustness and validation: 93.0% of 3,472 audited rewrites passed all integrity checks, and the pass rate did not deteriorate over time.Detected fabrication occurred in 3.4% of rewrites, citation removal in 2.3%, and quote removal in 2.1%.

5 Discussion

CHASE indicates that repeated adaptation to a fixed ranking system can weaken the relationship between ranking success and document quality without necessarily reducing quality itself. The resulting ecosystem responses are domain-dependent, and the framework motivates evaluating ranking incentives under repeated optimization.

  • Core finding: Repeated adaptation toward ranking-derived targets made ranking success a less reliable indicator of independently judged document quality across domains.Quality remained relatively stable in several domains, so the result concerns alignment rather than necessary quality decline.
  • Domain dependence: Structural convergence, signal instability, and feature dominance describe distinct domain-dependent responses observed within CHASE.Their emergence depends on the content domain and the ranking criteria available for creators to exploit.
  • Implications: Ranking systems in adaptive environments should be assessed for the incentives they create when repeatedly optimized against, including feature-level preferences that may become undesirable when amplified.CHASE provides a controlled stress test for whether ranking signals remain aligned with desired content properties as populations adapt.

6 Limitations

The study isolates repeated adaptation to a fixed ranking signal rather than modeling a complete generative-search system. Its conclusions may also depend on the selected model configurations, fixed feature space, finite horizon, and primarily automated quality measurement.

  • Scope and setting: CHASE omits retrieval, response generation, personalization, and user feedback, while creators receive clean signals and act myopically and independently.Real systems may involve noisy or delayed feedback and strategic reasoning about competitors.
  • Model and feature assumptions: Each pipeline role uses a single model configuration, so the dynamics may depend on the particular models and prompts used.The discriminator also operates over a fixed 25-dimensional feature space.
  • Model and feature assumptions: Unmeasured properties may influence both ranking and quality, so declining alignment may partly reflect feature insufficiency rather than optimization pressure alone.This limits a direct attribution of the observed divergence to optimization pressure.
  • Temporal and system limits: The fixed ranker isolates content-side distribution shift but excludes ranking-model adaptation through retraining, user feedback, or policy changes.The experiments also end after 20 rounds, so trajectories do not establish convergence or long-run equilibrium behavior.
  • Quality measurement: Document quality is primarily measured by an independent LLM judge and validated on a sampled subset with human annotation.Human validation cannot establish the quality or verifiability of every document in the ecosystem.

7 Conclusion

CHASE shows that repeatedly adapting content to a fixed LLM ranking signal can weaken ranking–quality alignment and reshape document features and content populations, with outcomes varying by domain.

  • Repeated optimization weakens the relationship between ranking success and independently judged content quality.The ranker remains fixed while document features, populations, and ranking–quality relationships change.
  • Ecosystem dynamics vary substantially across domains rather than following a single uniform trajectory.
  • Rank–citation analysis supports ranking as a useful source-visibility proxy within CHASE.
  • Control experiments distinguish ranking-driven adaptation from arbitrary rewriting effects.
  • The conclusions are limited to the simulated 20-round horizon and do not represent the full dynamics of deployed generative-search systems.

Ethics statement

The paper frames CHASE as a controlled simulation for studying ranking-driven content adaptation while limiting claims about deployed generative-search systems. Its methodology uses multiple LLM components and fixed, explicitly specified simulation procedures.

  • CHASE studies content adaptation to a fixed LLM ranking signal without modeling the full deployed generative-search pipeline.The framework is intended to identify possible incentives and simulated effects, not establish that deployed systems produce them.
  • The experiment uses separate model families for ranking, rewriting, quality evaluation, and rewrite-integrity auditing.The supplied implementation passage names Gemini 3.1 Flash-Lite, GPT-5.4-mini, Claude Haiku 4.5, and Claude Sonnet 4.6 for these roles.
  • Documents are represented with 25 features spanning structural, evidentiary, and semantic categories.

C Rank–Citation Validation

CHASE validates ranking as a source-visibility proxy by comparing document rankings with citations in grounded responses, then uses controls to separate ranking-target adaptation from rewriting alone.

  • 0.853 ± 0.093 rank–citation AUC supports ranking as a source-visibility proxy across six domains.The association remains similar from 0.849 at R0 to 0.863 at R20 and is positive in every domain.
  • Ranking should not be treated as equivalent to a complete generative-search pipeline.The broader pipeline may also include retrieval, reranking, generation, and personalization.
  • No-rewrite controls keep homogeneity unchanged in Retail (.732 → .732) and Debate (.633 → .633) from R0 to R20.This indicates that changes in document-level population statistics require active rewriting rather than repeated ranking alone.
  • Quality–ranking alignment changes by +.001, +.014, and −.036 under random targets versus −.047, −.065, and −.104 under CHASE in Retail, Debate, and News.The contrast associates larger declines with adaptation toward ranking-derived targets rather than arbitrary iterative rewriting.
  • Discriminator AUC changes smoothly across winner thresholds {5%, 10%, 20%} and feature counts J ∈ {3, 5, 7, 10}.Feature-selection overlap between adjacent winner thresholds is 39–42%, versus approximately 11% under random selection.
  • A constant participation probability of 0.30 produces qualitatively similar Debate trajectories to the default Beta(2, 5) mechanism.

E.1 Round-by-Round Ecosystem Metrics

The appendix documents round-by-round evaluation, feature evolution, prompts, and rewrite-integrity and human-evaluation procedures underlying the CHASE ecosystem analysis.

  • The canonical cross-family CHASE experiment reports round-by-round ecosystem metrics as means across five independent seeds.
  • The discriminator is refit on the evolving document population, so selected features can change while the underlying ranker remains fixed.Coefficient signs indicate conditional associations with top-ranked status, not causal feature effects.
  • Ranking prompts cover recommendation and question-answering domains, with domain-specific instructions for products and documents.
  • Non-winner documents are rewritten toward discriminator-derived feature targets while prompts instruct preservation of factual claims and avoidance of invented information.
  • 93.0% of 3,472 audited rewrites pass all integrity checks, with fabrication in 3.4%, citation removal in 2.3%, and quotation removal in 2.1%.The aggregate pass rate rises from 90.3% at the earliest audited round to 96.4% at R15.
  • Human evaluation assesses factual accuracy, completeness, usefulness, and verifiability on documents sampled across domains and evaluation rounds.

H.4 Results

Human validation used 60 documents assessed by three annotators, with substantial but imperfect agreement and moderate agreement with the independent LLM judge on shared quality dimensions.

  • 60 documents were evaluated by 3 annotators in the human-validation sample.
  • Ordinal Krippendorff’s α = 0.72 indicates substantial but not perfect inter-annotator agreement.
  • Spearman’s ρ = 0.58, 95% CI [0.49, 0.66], indicates moderate agreement between human ratings and the independent LLM judge.
Loading 2608.30466v1…