Source-linked AI summary

Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)

Liwei Jiang, Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yulia Tsvetkov, Maarten Sap, Alon Albalak, Yejin Choi

arXiv:2510.22954v1cs.CL

TL;DR

Open-ended LM diversity lacks scalable evaluation grounded in realistic user interactions and pluralistic human preferences. The paper introduces INFINITY-CHAT, a large dataset and taxonomy paired with large-scale studies of model homogeneity and evaluator calibration. It finds pronounced repetition within models, homogeneity across models, and miscalibration when human preferences diverge despite comparable response quality.

  • Problem

    Scalable evaluation of LM diversity remains limited beyond narrow tasks and often overlooks the multiple valid answers and divergent preferences inherent in open-ended interactions.

  • Method

    The paper constructs INFINITY-CHAT from real-world open-ended queries, develops a 6-category and 17-subcategory taxonomy, studies 70+ LMs, and collects dense human preference annotations.

  • Results

    The study finds pronounced intra-model repetition, inter-model homogeneity, and evaluator miscalibration on similarly good responses eliciting divergent annotator preferences.

  • Takeaways & Limitations

    INFINITY-CHAT provides a foundation for diagnosing and benchmarking Artificial Hivemind effects and for studying more pluralistic LM alignment.

  • Takeaways & Limitations

    The study does not provide a full causal analysis of whether repetition arises from pretraining data, alignment, memorization, contamination, or generalization.

Abstract

from arXiv · show

Language models (LMs) often struggle to generate diverse, human-like creative content, raising concerns about the long-term homogenization of human thought through repeated exposure to similar outputs. Yet scalable methods for evaluating LM output diversity remain limited, especially beyond narrow tasks such as random number or name generation, or beyond repeated sampling from a single model. We introduce Infinity-Chat, a large-scale dataset of 26K diverse, real-world, open-ended user queries that admit a wide range of plausible answers with no single ground truth. We introduce the first comprehensive taxonomy for characterizing the full spectrum of open-ended prompts posed to LMs, comprising 6 top-level categories (e.g., brainstorm & ideation) that further breaks down to 17 subcategories. Using Infinity-Chat, we present a large-scale study of mode collapse in LMs, revealing a pronounced Artificial Hivemind effect in open-ended generation of LMs, characterized by (1) intra-model repetition, where a single model consistently generates similar responses, and more so (2) inter-model homogeneity, where different models produce strikingly similar outputs. Infinity-Chat also includes 31,250 human annotations, across absolute ratings and pairwise preferences, with 25 independent human annotations per example. This enables studying collective and individual-specific human preferences in response to open-ended queries. Our findings show that LMs, reward models, and LM judges are less well calibrated to human ratings on model generations that elicit differing idiosyncratic annotator preferences, despite maintaining comparable overall quality. Overall, INFINITY-CHAT presents the first large-scale resource for systematically studying real-world open-ended queries to LMs, revealing critical insights to guide future research for mitigating long-term AI safety risks posed by the Artificial Hivemind.

1 Introduction

The paper addresses limited scalable evaluation of diversity in real-world open-ended LM use by introducing INFINITY-CHAT and studying homogeneity within and across models. It finds an Artificial Hivemind effect and shows that current evaluators can miss divergent human preferences among similarly rated responses.

  • Motivation: Existing diversity evaluations often use stylized, narrow tasks and do not capture the open-endedness of real-world user interactions.The paper highlights limitations in scalable diversity evaluation beyond repeated sampling from one model or synthetic tasks.
  • Contribution: INFINITY-CHAT contains 26K real-world open-ended queries with no single correct response, organized into 6 top-level categories and 17 subcategories.The queries were mined from naturally occurring chatbot-user interactions and span diverse prompt types.
  • Findings: Across 70+ open and closed source LMs, the study finds intra-model repetition and stronger inter-model homogeneity in open-ended generation.Different models often converge on similar ideas with only minor phrasing differences.
  • Human preferences: The dataset includes 31,250 human annotations covering absolute ratings and pairwise preferences, with 25 independent annotations per example.These annotations support analysis of collective and individual-specific preferences.
  • Human preferences: LMs, reward models, and LM judges are often miscalibrated on responses that elicit divergent annotator preferences despite comparable overall quality.The finding challenges evaluation pipelines that assume a single consensus notion of quality.
  • Implications: INFINITY-CHAT provides a framework for diagnosing Artificial Hivemind effects and guiding safer, more expressive, and resourceful language models.The framework combines real-world queries, query taxonomy, and dense human annotations.

2 INFINITY-CHAT: Real-World Open-Ended Queries with Diverse Responses

INFINITY-CHAT is built from in-the-wild user queries to characterize the diverse landscape of open-ended LM interactions. Its taxonomy reveals both dominant and underexplored query types, while the resource supports systematic study of response repetition and diverse outputs.

  • Dataset construction: INFINITY-CHAT is constructed from filtered and refined WildChat inputs, yielding 26,070 open-ended and 8,817 closed-ended queries.The source queries were English, non-toxic, single-turn GPT-4 interactions between 15 and 200 characters.
  • Query taxonomy: The taxonomy combines manual label refinement with GPT-4o annotation to produce 6 high-level categories and 17 fine-grained subcategories.The process begins with approximately 100 mined queries and iteratively builds a hierarchical structure.
  • Query taxonomy: Creative Content Generation accounts for 58.0% of queries, while Brainstorming & Ideation accounts for 15.2%.Other prominent types include Alternative Writing Genres at 38.5% and Concept Explanation at 23.6%.
  • Query taxonomy: The mining process identifies 314 novel open-ended categories beyond the predefined taxonomy.Prominent keywords include Cultural, Analysis, Ethical, Historical, Media, and Humor.
  • Resource value: INFINITY-CHAT provides a resource for evaluating whether LMs can generate varied, appropriate outputs in naturally occurring open-ended settings.The taxonomy is intended to support research on diverse responses and pluralistic alignment.

3 Artificial Hivemind: Intra- and Inter-Model Homogeneity in LMs

Across open-ended queries, LMs exhibit both strong repetition within individual models and substantial homogeneity across different models. This Artificial Hivemind appears in semantic similarity, phrase overlap, and convergence on recurring ideas.

  • Intra-model repetition: Using 50 responses per query, 79% of cases had same-model average similarity above 0.8 despite high-stochasticity decoding.The sampling used top-p = 0.9 and t = 1.0.
  • Intra-model repetition: 61.2% of response pairs exceeded 0.8 similarity under min-p sampling, although the method reduced extreme repetition above 0.9.Additionally, 81% of pairs exceeded 0.7 similarity.
  • Inter-model homogeneity: Cross-model response similarity ranged from 71% to 82%, including 0.82 for DeepSeek-V3 and qwen-max-2025-01-25.DeepSeek-V3 and gpt-4o-2024-11-20 reached 0.81 similarity.
  • Inter-model homogeneity: Models sometimes share extended verbatim phrases or identical responses, demonstrating instance-level cross-model overlap on open-ended queries.Examples include matching motto outputs from qwen-max-2025-01-25 and qwen-plus-2025-01-25.
  • Inter-model homogeneity: For “Write a metaphor about time,” 50 responses from each of 25 models formed only two semantic clusters centered on “time is a river” and “time is a weaver.”This illustrates convergence at the level of abstract concepts, not only surface wording.
  • Inter-model homogeneity: The top-N most similar responses are evaluated by counting their unique source models, with higher counts indicating stronger cross-model similarity.This measurement compares how many models contribute to the most similar outputs for each query.

4 How Do LMs, Reward Models, and LM Judges Handle Alternative Responses to Open-Ended Queries?

The study tests whether model-based ratings reflect pluralistic human judgments when alternative open-ended responses are similarly good or elicit disagreement. Alignment with human ratings drops substantially in both settings, especially for disagreement.

  • Human preference distributions: 31,250 human annotations cover absolute quality ratings and pairwise preferences, with 25 independent annotators per query–response example.The dataset contains 18,750 absolute-rating labels and 12,500 pairwise-preference labels.
  • Human preference distributions: Human preference distributions vary substantially across alternative responses, with some pairwise examples receiving near-uniform support across options.This reflects disagreement despite multiple responses being valid for open-ended queries.
  • Alignment with human judgments: The evaluation compares perplexity-based LM scores, standardized reward-model outputs, and LM-judge ratings against human annotations.LM judges use overall-quality and HHH rubrics.
  • Subset construction: Similar-quality subsets are constructed by removing Tukey-fence outliers while varying k from 0.5 to 3.0.The resulting subsets are compared with the full set using correlations between model and human ratings.
  • Alignment with human judgments: Model–human rating correlations drop significantly on similar-quality subsets in both absolute and pairwise preference setups.The result remains consistent across alternative subset-selection methods.
  • Alignment with human judgments: Correlations with human ratings also drop substantially for high-annotator-disagreement examples across both rating setups.The findings motivate more nuanced modeling of idiosyncratic human disagreement.

5 Related Work

Prior work identifies diversity collapse, creativity measurement, and pluralistic alignment as related challenges, while emphasizing the difficulty of representing varied human values and open-ended generation.

  • Diversity collapse: Diversity collapse has been linked to synthetic-data training, LM alignment, and insufficient diversity in training data.Its potential consequences include reduced creativity.
  • Creativity measurement: Creativity and divergent thinking in LMs are often assessed through psychometric tasks such as divergent association, alternate uses, and Torrance tests.Human evaluation and LLM-as-a-judge are also used.
  • Pluralistic alignment: Pluralistic alignment addresses the challenge that AI systems may otherwise represent values monolithically rather than serving varied population demands.This work situates disagreement and diverse preferences within that alignment concern.

6 Conclusion

INFINITY-CHAT is presented as a resource for evaluating LM diversity in naturally occurring open-ended settings. Its taxonomy and dense human annotations support diagnosing and benchmarking the Artificial Hivemind effect.

  • Conclusion: INFINITY-CHAT evaluates LM diversity in naturally occurring open-ended settings and identifies both intra-model repetition and inter-model homogeneity.The resource is designed to diagnose mode collapse across current LMs.
  • Conclusion: By combining diverse prompt categories with dense human preference annotations, INFINITY-CHAT provides a foundation for diagnosing, benchmarking, and mitigating generative mode collapse.The authors aim to encourage genuine output diversity and guard against homogenization of human expression.

A.1 Limitations

The paper identifies scope, measurement, and evaluation boundaries that constrain how broadly its findings should be interpreted. It also outlines future work on multilingual coverage, causal mechanisms, and diversity-aware development.

  • 26K queries provide only a snapshot of possible open-ended queries and may miss creative divergence across contexts.
  • English prompts derived from WildChat may underrepresent linguistic, cultural, and regional diversity, limiting generalizability beyond English contexts.
  • Semantic similarity of text embeddings may not fully capture the multidimensional spectrum of creative variation in generated responses.
  • Future work targets multilingual and multicultural extensions, foundation-model analysis, diversity-aware training, decoding evaluation, and integration into red-teaming and curriculum design.
  • The analysis focuses primarily on diversity under fixed decoding configurations and does not study quality as a coequal evaluation dimension.
  • The study reveals Artificial Hivemind patterns but does not establish whether shared training data, memorization, generalization, or other factors cause homogenization.

B INFINITY-CHAT: A Dataset of In-The-Wild Open-Ended User Queries

INFINITY-CHAT is constructed by filtering real-world WildChat queries, classifying them along semantic and response-type dimensions, and retaining open-ended prompts for diversity-focused analysis.

  • The dataset mines diverse real-world user conversations from WildChat to identify open-ended and closed-ended queries.
  • 37,426 English, non-toxic, GPT-4-directed queries of 15–200 characters are extracted as candidates for analysis.
  • GPT-4o classifies queries by meaningful information seeking, greeting or model inquiry, and whether responses should be single or multiple.
  • Automatic filtering yields 26,070 open-ended queries and 8,817 closed-ended queries in INFINITY-CHAT.
  • The taxonomy annotation process uses GPT-4o-2024-11-20 to classify open-ended queries into categories and subcategories.

B.3 Human Validation of the Open-endedness of Queries in INFINITY-CHAT

Human validation confirms that most sampled INFINITY-CHAT queries are open-ended, while annotator judgments about the breadth of possible answers vary substantially.

  • 89% of queries were judged open-ended by majority vote, while 100% qualified under an inclusive one-annotator criterion.
  • Annotators estimated answer breadth using fewer than 3, 3 to 10, 10 to 20, or more than 20 reasonable alternatives.
  • 81.27% of queries were judged to allow at least 3 alternatives, and 34.66% to permit over 20.
  • Individual annotator judgments about open-endedness varied considerably across prompts.
  • Using a coarse rule, prompts were high-open-endedness if any annotator predicted more than 20 responses and low-open-endedness otherwise.
  • Across 42 models, average sentence similarity was 0.800 for high-open-endedness prompts and 0.837 for low-open-endedness prompts.

C.1 Evaluation Setups

The evaluation reuses model outputs generated from 100 representative INFINITY-CHAT prompts under a unified protocol across open- and closed-source models.

  • The study uses 100 representative open-ended prompts from INFINITY-CHAT100 as the seed set for intra- and inter-model analyses.
  • The same model outputs are reused across both intra-model and inter-model analyses under a unified generation protocol.
  • Generations come from HuggingFace models, provider APIs, or TogetherAI for models exceeding local GPU capacity.
  • All models follow the same decoding configurations regardless of generation method.
  • Extended intra-model repetition and inter-model homogeneity results are reported in Tables 6–11.

C.4 Examining How Paraphrased Queries Influence Response Homogeneity

The study finds that paraphrasing queries does not substantially reduce response homogeneity: models produce highly similar outputs for original and semantically equivalent prompts. Examples show similarity ranging from shared concepts to repeated wording across models.

  • The evaluation measures within-prompt similarity among responses to one prompt and cross-paraphrase similarity between responses to original and paraphrased prompts.
  • 0.821 within-prompt similarity versus 0.781 cross-paraphrase similarity shows high response consistency across original and paraphrased queries.The difference is only 0.04 across all 42 models.
  • Responses often share high-level concepts across paraphrases, such as framing time with the metaphor “time is a river.”
  • Different models also exhibit surface-level homogeneity by reusing words such as “profound” in opening sentences.

D.3 Similar-Quality Responses

The paper tests whether calibration findings for comparable-quality responses depend on how those subsets are selected. Results remain consistent across alternative selection procedures, while the authors note that no single gold-standard approach exists.

  • No single gold-standard method exists for selecting comparable-quality response subsets in the dataset.Tukey’s fences are one possible method rather than a uniquely justified standard.
  • Alternative similar-quality subset methods produce findings consistent with the reported results.The alternatives include optimized sliding, centroid-based, distance-based, Tukey, median expansion, and gap-based selection.
  • The alternative procedures differ in how they identify tightly grouped subsets, including minimizing range, variance, pairwise distance, or local gaps.
  • Table 19 reports Spearman correlations between human scores and perplexity, reward-model, and LM-judge scores across similar-quality subsets.

2. Limitations

The paper reports a limitations discussion and documents reproducibility, experimental details, statistical analyses, compute resources, ethics, and broader-impact considerations.

  • The authors report a thorough limitations discussion in §Appendix A.
  • No theoretical results are included, so theory assumptions and proof completeness are not applicable.
  • The paper states that its main experimental results are reproducible, with code release and necessary details in Appendices B, C, and D.
  • The authors state that experimental settings, compute and annotation resources, and statistical significance analyses are documented in the appendices.
  • The authors confirm compliance with the NeurIPS Code of Ethics and report discussing broader impact in §Appendix A.
Loading 2510.22954v1…