Source-linked AI summary
We're Different, We're the Same: Creative Homogeneity Across LLMs
Emily Wenger, Yoed Kenett
TL;DR
Prior studies showed that individual LLMs can narrow creative outputs, leaving unclear whether this effect is model-specific or general to LLMs. The paper compares human and broad multi-LLM responses on standardized creativity tests and finds substantially lower population-level variability among LLMs, even after controls. It concludes that the evaluated LLMs produce a narrower and more homogeneous range of creative responses.
Problem
Prior studies examined single LLMs, leaving unclear whether creative-output homogeneity is model-specific or occurs across LLMs generally.
Method
The study elicits human and LLM responses on standardized creativity tests and compares their population-level variability while controlling for response structure and model similarities.
Results
LLMs exhibit much lower population-level output variability than humans and return a narrower range of creative responses, even after controlling for model similarities and structural differences.
Takeaways & Limitations
The findings suggest that using LLMs as creative partners generally may drive users toward a limited set of similar creative outputs.
Takeaways & Limitations
The study demonstrates LLM homogeneity only for certain creativity tests and does not prove that LLMs generally produce homogeneous outputs in all settings.
Abstract
from arXiv · showhide
Numerous powerful large language models (LLMs) are now available for use as writing support tools, idea generators, and beyond. Although these LLMs are marketed as helpful creative assistants, several works have shown that using an LLM as a creative partner results in a narrower set of creative outputs. However, these studies only consider the effects of interacting with a single LLM, begging the question of whether such narrowed creativity stems from using a particular LLM -- which arguably has a limited range of outputs -- or from using LLMs in general as creative assistants. To study this question, we elicit creative responses from humans and a broad set of LLMs using standardized creativity tests and compare the population-level diversity of responses. We find that LLM responses are much more similar to other LLM responses than human responses are to each other, even after controlling for response structure and other key variables. This finding of significant homogeneity in creative outputs across the LLMs we evaluate adds a new dimension to the ongoing conversation about creativity and LLMs. If today's LLMs behave similarly, using them as a creative partners -- regardless of the model used -- may drive all users towards a limited set of "creative" outputs.
1 INTRODUCTION
Prior research found that using individual LLMs as creative partners can homogenize outputs, but it remained unclear whether this effect was model-specific or characteristic of LLMs generally. This work compares human and multi-LLM responses on standardized creativity tests and finds substantially greater cross-LLM homogeneity than human similarity.
- The study tests cross-model creative-output convergence by comparing population-level variability in human and LLM responses across standardized creativity tasks.
- LLMs match or outperform humans on standard tests of individual creativity, but this performance coexists with narrower variation across LLM outputs.
- LLM responses to creative prompts are much more similar to each other than human responses, even after controlling for model-family overlap and response structure.
- Encouraging higher creativity through system-prompt changes slightly increases LLM creativity and inter-LLM variability, but human responses remain more variable.
- The findings suggest that relying on today’s popular models as creative partners could collectively narrow users’ creative outputs regardless of the model selected.
2 RELATED WORK
Prior work documented creative homogenization for particular LLMs and studied model similarity, but rarely examined whether creative outputs converge across models. This paper addresses that gap by comparing creative-response diversity across many LLMs with human diversity.
- Earlier studies found that LLM-assisted creative outputs become more similar across writing, research ideas, surveys, creative ideation, and art.
- Existing work typically focused on specific models, leaving open whether observed homogeneity stems from one model or from LLM use generally.
- Prior studies considered narrow output distributions or algorithmic monocultures, but did not examine similarity across models.
- Research on model representations reports feature similarity and feature universality across LLMs, while downstream consequences of that similarity remain limited.
- This paper compares creative-response diversity across many LLMs with human diversity using standard creativity tests.
3 METHODOLOGY
The study compares the population-level variability of human and LLM responses to structured creativity tests, while addressing concerns about applying human creativity measures to LLMs. It uses multiple tests, models, and participants to evaluate response diversity while controlling relevant sources of similarity.
- 3.1 How do we elicit creative responses from LLMs?: The authors evaluate response diversity rather than inherent LLM creativity, treating variability across responses as an empirical property.They justify using creativity tests because the goal is to compare response distributions, not psychological creativity.
- 3.1 How do we elicit creative responses from LLMs?: Structured creativity tests are chosen to separate content similarity from shared response structure, such as passive voice or gerund use.Short stories could confound these factors, whereas AUT, FF, and DAT elicit more directly comparable outputs.
- 3.1 How do we elicit creative responses from LLMs?: The study compares human and LLM response variability using the Alternative Uses Test, Forward Flow, and Divergent Association Test.These tests provide structured ways to elicit comparable creative responses across populations.
- 3.3 Test subjects: The analysis administers AUT, FF, and DAT to human participants and a broad set of publicly accessible LLMs, including models from distinct families.The study also examines within-family behavior and varies the system prompt in additional experiments.
4 KEY RESULTS
Across standardized creativity tests, LLMs achieve roughly similar individual originality to humans but produce substantially more homogeneous responses. Their responses show lower semantic variability, stronger clustering, and greater lexical overlap than human responses.
- LLMs slightly outperform humans on AUT and DAT originality, while humans slightly outperform LLMs on FF.Overall, the groups show roughly equal measured originality, reducing individual creativity as a confound in variability comparisons.
- Across all tests, LLMs have significantly lower mean population-level variability than humans.Variability is measured using cosine distances between embedded responses.
- LLM responses cluster more tightly in embedded feature space than human responses.TSNE visualization and k-means clustering of AUT sentence embeddings provide visual evidence of lower LLM response variability.
- LLM responses contain many more words in common than human responses across all three creativity tests.Lexical overlap partially explains their higher semantic similarity because overlapping words map responses to similar feature vectors.
5 ADDITIONAL ANALYSIS
Additional analyses test whether LLM homogeneity reflects response structure, model family, prompting, or sampling differences. Homogeneity persists after structural controls, across prompt variants, and despite creative system prompts, while same-family models are only slightly less diverse.
- 5.1 Controlling for AUT Response Structure: Prompt version 3 most closely matches human AUT response lengths and is therefore used in the main experiments.It produces roughly the same proportion of single-word answers as humans while reducing two-word answers.
- 5.1 Controlling for AUT Response Structure: LLMs retain much lower response variability than humans after controlling for AUT response structure.The difference persists across three prompt versions and a single-word-response comparison.
- 5.1 Controlling for AUT Response Structure: LLM variability increases slightly from prompt version 1 to version 3, but response substance rather than structure remains the primary difference.Low variability also persists on FF and DAT, where response structure does not matter.
- Models from the same family show slightly lower response diversity than models from different families, but the difference is not statistically significant.The reported family comparison is based partly on visual inspection of the embedding distributions.
- Creative system prompts slightly increase individual LLM creativity but do not substantially improve response variability.Across prompts, LLM variability remains much lower than human variability.
- The online human study broadly matches prior results, supporting its use as a baseline for the cross-population analysis.Relative to prior studies, the sample has different individual scores but equal or greater population-level variability on AUT and FF.
6 DISCUSSION
Across standardized creativity tests, LLMs produced less population-level variability than humans, even after controlling for model similarities and response structure. The authors interpret this cross-LLM homogeneity as evidence that LLM use may homogenize creative outputs, while noting important scope limitations.
- LLMs exhibited much lower population-level output variability than humans, even after controlling for model similarities and structural differences.
- LLMs performed well on divergent-thinking tests, but their responses covered a narrower range than human responses.
- The findings extend prior work on single-model effects by showing that homogeneity persists across the evaluated LLMs.
- If LLMs respond similarly to creative requests, users may converge toward a limited set of creative outputs.
- This convergence may constrain the divergent creativity associated with recognized artistic geniuses.
- The evidence is limited to certain creativity tests and does not establish that LLMs generally produce homogeneous creative outputs.
- The study measures originality through semantic similarity, leaving flexibility, fluency, and elaboration for future investigation.
7 ETHICAL CONSIDERATIONS
The study reports ethical safeguards for its human-participant survey and describes the associated risks as minimal.
- The user study received IRB approval, and participants signed a clearly written consent form before completing the survey.
- Participant data were anonymized and stored on secure servers to protect privacy.
- The authors characterize other ethical risks as minimal because the experiments used no sensitive data and elicited only benign model responses.
A DIVERGENT THINKING TEST WORDING
The appendix records slightly different prompt wording for humans and LLMs while standardizing response formatting for analysis.
- The study reports the exact wording used for creativity tests administered to humans and LLMs.
- Human and LLM prompts differed slightly because LLMs were instructed to use a processing-friendly output format.
- Without formatting instructions, LLMs often explained their word choices, which muddied the data.
A.1 AUT prompts.
The appendix specifies the Alternative Uses and Forward Flow prompts, including their starting words, response limits, and formatting constraints for humans and LLMs.
- A.1 AUT prompts.: The Alternative Uses Test used five start words in the original experiments and ten in the expanded LLM evaluation.
- A.1 AUT prompts.: Human Alternative Uses prompts asked participants to list up to 10 creative uses for an object.
- A.1 AUT prompts.: LLM Alternative Uses prompts likewise allowed up to 10 uses but required semicolon-separated words or phrases without additional text.
- A.2 Forward Flow prompts.: Forward Flow used candle, table, bear, snow, and toaster as starting words.
- A.2 Forward Flow prompts.: Human Forward Flow prompts requested the next single word in each mental association sequence and prohibited proper nouns.
- A.2 Forward Flow prompts.: LLM Forward Flow prompts imposed the same single-word and proper-noun constraints while requiring at least 22 words.
- A.2 Forward Flow prompts.: The LLM Forward Flow response had to begin with candle and contain only a comma-separated word list.
A.3 DAT Prompts.
The DAT asks participants to provide 10 mutually different English nouns under constraints on word form, vocabulary, originality, and time. The LLM prompt closely mirrors the human instructions.
- A.3 DAT Prompts: The DAT requires 10 English nouns that are as different from one another as possible in meaning and use.Responses must contain only single words, exclude proper nouns and specialized vocabulary, and be generated independently.
- A.3 DAT Prompts: Both human and LLM prompts impose the same core constraints, including a four-minute limit and returning only the word list.The LLM version reproduces the noun, English-language, non-specialized-vocabulary, and independent-generation requirements.
- A.3 DAT Prompts: Figure 10 shows LLM responses to the DAT clustering more tightly in feature space than human responses.
B ORIGINALITY SCORES FOR AUT, FF, AND DAT
The paper computes originality for AUT, FF, and DAT populations using embedding-based semantic distances, with task-specific scoring procedures and TF-IDF weighting for AUT.
- Scoring framework: The scoring framework represents words and responses with embeddings, computes cosine similarity, and indexes scores by test and population.Here, t denotes AUT, FF, or DAT, and P denotes a human or LLM population.
- AUT scoring: AUT originality is a TF-IDF-weighted sum of semantic distances between the prompt and each response word.Common words receive lower weights, while unusual words receive higher weights in the final score.
- FF scoring: FF originality measures each thought’s average distance from all preceding thoughts in the sequence.Population-level FF scores are then defined for collections of n-word sequences.
- DAT scoring: DAT originality averages semantic distances between all pairs of words in each response.
C TSNE OF FF AND DAT
TSNE visualizations of DAT and FF sentence embeddings show that LLM responses cluster more closely than human responses, corresponding to lower population-level originality measurements.
- TSNE of FF and DAT: Figure 10 uses TSNE of sentence embeddings to visualize DAT and FF response distributions.
- TSNE of FF and DAT: LLM responses cluster closer in feature space than human responses for DAT and FF, producing lower population-level originality measurements.Figure 10 visualizes the DAT and FF embedding distributions, extending the clustering trend observed for AUT.
- TSNE of FF and DAT: Because DAT has no varying start words, the visualization includes all LLM and human responses rather than separate start-word clusters.