Source-linked AI summary

Homogenization Effects of Large Language Models on Human Creative Ideation

Barrett R. Anderson, Jash Hemant Shah, Max Kreminski

arXiv:2402.01536v2cs.HCcs.AI

TL;DR

The paper asks whether LLM-based creativity support tools support genuinely creative outputs while potentially reducing diversity across users. Using a 36-participant comparative study of ChatGPT and a non-AI CST, it finds greater group-level homogenization with ChatGPT alongside greater idea quantity and detail, but lower perceived responsibility.

  • Problem

    The paper asks whether LLM-based creativity support tools support genuinely creative outputs while potentially reducing diversity across users.

  • Method

    A 36-participant comparative user study tested ChatGPT against an alternative CST across divergent ideation tasks and examined semantic similarity, creativity measures, and user experience.

  • Results

    ChatGPT assistance produced less semantically diverse ideas at the group level, while each participant’s ideas were similarly diverse across conditions.

  • Takeaways & Limitations

    LLM-based CSTs may support rapid enumeration of relatively obvious possibilities, but design and model interventions may be needed to mitigate homogenization.

  • Takeaways & Limitations

    The lab study used fixed short ideation periods and researcher-supplied prompts, which may not reflect organic divergent ideation in the wild.

Abstract

from arXiv · show

Large language models (LLMs) are now being used in a wide variety of contexts, including as creativity support tools (CSTs) intended to help their users come up with new ideas. But do LLMs actually support user creativity? We hypothesized that the use of an LLM as a CST might make the LLM's users feel more creative, and even broaden the range of ideas suggested by each individual user, but also homogenize the ideas suggested by different users. We conducted a 36-participant comparative user study and found, in accordance with the homogenization hypothesis, that different users tended to produce less semantically distinct ideas with ChatGPT than with an alternative CST. Additionally, ChatGPT users generated a greater number of more detailed ideas, but felt less responsible for the ideas they generated. We discuss potential implications of these findings for users, designers, and developers of LLM-based CSTs.

1 INTRODUCTION

This paper examines whether LLM-based creativity support tools assist creative ideation while potentially homogenizing ideas across users. A 36-participant comparative study investigates group- and individual-level semantic similarity, perceived responsibility, and other creativity facets.

  • Creative success depends partly on producing ideas that are new, surprising, and valuable.
  • Researchers have raised concerns that centralized AI systems may decrease diversity in creative outputs.
  • The study compares ChatGPT with an alternative non-AI creativity support tool across four divergent ideation tasks and 1,271 participant-generated ideas.
  • The research questions address group-level and individual-level semantic similarity, users’ responsibility for ideas, and fluency, flexibility, and elaboration.
  • Uniquely among recent studies, the paper compares LLM homogenization with an alternative CST, separates group- and individual-level effects, and extends beyond writing.
  • The results suggest that LLM homogenization arises from similar ideas being provided to different users rather than increased individual fixation.

2 BACKGROUND AND RELATED WORK

Prior research provides evidence that AI assistance can homogenize creative outputs, but leaves open how these effects compare with alternative CSTs and whether they occur mainly across users or within individuals. The paper situates its evaluation in broader CST assessment practices and creativity dimensions.

  • Prior studies of homogenization: Earlier studies found homogenization from instruction-tuned LLM assistance in argumentative writing and from GPT-4 assistance in fictional narrative writing.
  • Prior studies of homogenization: Only two previous direct studies examined LLM-driven creative homogenization, both comparing LLM-assisted users with unassisted users in writing tasks.
  • Prior studies of homogenization: Existing evidence left unclear how severe homogenization is, whether it occurs outside writing, and whether it is primarily individual-level or group-level.
  • Why LLMs may homogenize ideas: Algorithmic monoculture may homogenize creative outcomes when many people use one consistent AI system instead of varied humans or systems.
  • Why LLMs may homogenize ideas: LLMs may also encourage fixation by presenting complete-seeming ideas early, potentially constraining later solution variation.
  • Evaluating creativity and CSTs: CST evaluation includes process measures such as user reports and actions, and product measures such as quantity, quality, and artifact characteristics.
  • Evaluating creativity and CSTs: The study combines process and product data, focusing on homogenization through semantic similarity of participant-generated ideas.

3 METHODS

The researchers conducted a within-subjects comparison of ChatGPT and the Oblique Strategies deck for divergent ideation. Participants completed timed tasks, generated ideas with both tools, and provided process and experience measures.

  • The within-subjects experiment compared ChatGPT with the Oblique Strategies deck as two creativity support tools.
  • The study recruited 36 participants, with three later excluded from all analyses.
  • ChatGPT versions 3.5 released on May 3 and August 3, 2023 were used in the study.
  • The Oblique Strategies deck supplied creative prompts as the non-AI control condition.
  • Participants completed Product Improvement and Improbable Consequences tasks, generating as many ideas as possible within fixed sessions.
  • After each tool, participants completed the Creativity Support Index and rated responsibility for their output or attribution to the tool.

4 RESULTS

ChatGPT-assisted ideas were more homogenized across participants than ideas produced with the alternative CST, while individual participants’ idea diversity did not differ observably. The analysis used process data and embedding-based semantic similarity, with an explicit validation caveat.

  • Evaluation approach: Creative outcomes were evaluated using participant idea lists and cosine similarity between sentence embeddings and average embeddings.
  • Evaluation caveat: The sentence embeddings had not previously been validated for creativity assessment, so the researchers conducted a small human-judgment validation experiment.
  • Group-level homogenization: ChatGPT ideas were less divergent from the task-wide average than OS ideas, indicating greater group-level homogenization.ChatGPT: M = .24, SD = .07; OS: M = .28, SD = .08; t(32) = 2.154, p = 0.038, d = .47, 95% CI [.00,.07].
  • Individual-level homogenization: ChatGPT and OS ideas did not differ observably in divergence from each participant’s own average ideas.ChatGPT: M = .65, SD = .07; OS: M = .66, SD = .08; t(32) = .944, p = 0.352, d = .12, 95% CI [-.04,.01].

4.2 Sense of Responsibility

ChatGPT users assigned less responsibility to themselves and more to the CST for their ideas than OS users. Individual-level semantic homogeneity did not differ observably between tools.

  • 4.2 Sense of Responsibility: 48.17% of responsibility was assigned to participants themselves with ChatGPT, versus 63.63% with OS.The difference was statistically significant, t(32) = 3.21, p = 0.003, d = .67.

4.3 Other Facets of Creativity

ChatGPT increased idea fluency, category coverage, and elaboration relative to OS, while the tested uniqueness measures showed no observed advantage for either tool.

  • 4.3.1 Fluency: Simple Idea Count: ChatGPT users generated 8.39 ideas versus 7.32 with OS, an approximately 15% increase.The comparison was significant, t(32) = 2.10, p = 0.044, d = .32.
  • 4.3.2 Flexibility: Human-Coded Categories: ChatGPT users covered 8.58 idea categories versus 6.77 with OS, about 27% more categories.The comparison was significant, t(32) = 3.50, p = 0.001, d = .54.
  • 4.3.3 Elaboration: Stoplisted Word Count: ChatGPT ideas had a stoplisted word count of 8.25 versus 6.46 with OS, indicating greater elaboration by this measure.Stoplisted word count was used as a computational correlate of human-judged elaboration.
  • 4.3.4 Originality: Uniqueness: ChatGPT did not produce more unique ideas than OS under the primary or alternative uniqueness measures.The primary comparison was M = .74 versus M = .97, and the 5% threshold comparison was M = .87 versus M = .90.

4.4 Retrospective Reflections

Participants generally found ChatGPT easier and faster to use, but less rewarding and engaging than OS. Overall Creativity Support Index ratings did not differ between the tools.

  • 4.4.1 Interview: 27.78% of participants described ChatGPT as easier to use but less rewarding, the most common interview theme.The theme included 10 participants.
  • 4.4.1 Interview: 25.00% of participants expressed positive sentiment about the speed and accuracy of ChatGPT responses.This was the second most common interview theme, reported by 9 participants.
  • 4.4.1 Interview: 22.22% of participants found ChatGPT less engaging, while some reported repetitive or overly specific responses.Participants reported repetitive responses at 8.33% and responses becoming too specific too quickly at 5.56%.
  • 4.4.2 Creativity Support Index: CSI ratings were 78.03% for ChatGPT and 73.98% for OS, with no observed overall or subscale differences.The overall comparison was not significant, t(35) = 1.028, p = 0.312, d = .24.

4.5 Creative Process

Participants commonly began with their own ideas, copied the task prompt, and iterated when using ChatGPT. More prompting related to more ideas and higher weighted uniqueness, but process choices showed no observed outcome differences in the reported comparisons.

  • 4.5.1 Initial Ideation: 63.89% of participants began ChatGPT sessions by entering their own ideas, similar to their pattern with OS.Participants initially supplied a mean of 5.65 ideas before interacting with ChatGPT.
  • 4.5.2 Prompting: 86.11% of participants directly copy/pasted the creative ideation prompt, and 72.22% iterated on their prompts.Iteration included adding context, asking specific variations, or requesting regenerated or additional responses.
  • 4.5.3 Use of LLM Output: 41.67% of participants directly copied some ChatGPT output, while most of these participants selected, edited, or combined copied material.Only one participant copied the entire output without curation or editing.
  • 4.5.4 Process Impact on Outcome: The number of ChatGPT prompts correlated positively with idea count and weighted uniqueness, but not with average idea homogeneity.The correlations were r(32)=.43, p=.008; r(32)=.46, p=.004; and r(32)=.23, p=.185, respectively.
  • 4.5.4 Process Impact on Outcome: Starting with personal ideas or using creative prompting strategies produced no observed differences in idea count, weighted uniqueness, or homogeneity.The reported comparisons found no significant outcome differences for either process choice.

5 DISCUSSION

ChatGPT increased idea quantity without increasing individual-level diversity, while reducing semantic diversity across users relative to the alternative CST. The discussion attributes this group-level homogenization partly to similar model suggestions and the low inferential distance of finished-looking outputs.

  • ChatGPT-assisted ideas were significantly less semantically diverse at the group level than ideas produced with the non-AI CST.
  • Individual participants produced similarly diverse idea sets with ChatGPT and the non-AI CST despite generating more ideas with ChatGPT.
  • The findings suggest homogenization stems from ChatGPT giving different users similar ideas rather than increasing individual-level fixation.
  • The authors associate homogenization partly with the low inferential distance between ChatGPT outputs and apparently complete ideas.
  • ChatGPT users felt less responsible for their ideas, and some described the model as doing the heavy lifting during ideation.
  • Potential mitigations include oblique outputs, cliché alerts, user information about typical outputs, and movement among diverse CSTs.
  • More sophisticated prompting alone may not reliably elicit diverse responses, and prompt-level randomness requires thorough testing.

6 LIMITATIONS

The study’s controlled laboratory setting and use of general-purpose ChatGPT constrain how broadly its findings apply to real-world divergent ideation and other CST deployments.

  • The laboratory setting, fixed short response times, and researcher-supplied prompts may not reflect organic divergent ideation scenarios.
  • Because the study used general-purpose ChatGPT rather than a divergent-ideation-specific CST, alternative models or interfaces may avoid homogenizing ideas.

7 CONCLUSION

The study finds stronger group-level homogenization with ChatGPT than with a plausible alternative CST, alongside higher fluency, flexibility, and elaboration. It concludes that current general-purpose instruction-tuned LLMs can support rapid ideation but are not well suited to developing truly original ideas.

  • ChatGPT produced stronger group-level homogenization than at least some plausible alternative CSTs in human-in-the-loop divergent ideation.
  • ChatGPT users showed greater fluency, flexibility, and elaboration, suggesting usefulness for rapidly enumerating relatively obvious possibilities.
  • Current general-purpose instruction-tuned LLMs are not well suited to helping users develop truly original ideas.
  • The authors propose wider use of homogenization analysis and interventions at CST design and model-development levels.

A VALIDATING SENTENCE EMBEDDINGS FOR HOMOGENIZATION ANALYSIS

The paper validates a transformer-based sentence embedding approach for measuring semantic similarity in creativity research, while noting imperfect agreement with human judgments and replication constraints from proprietary alternatives.

  • The primary analysis uses all-MiniLM-L6-v2 to compute semantic similarity between participant ideas expressed as short strings.
  • Semantic-similarity creativity research assigns numeric similarity scores to creative artifacts relative to a fixed reference point.
  • Transformer-based sentence embeddings are preferred because they incorporate sentence structure, but had not been specifically validated for creativity research.
  • The validation experiment compared embedding-assigned categories with human coder categories across four creativity tasks and several models.
  • all-MiniLM-L6-v2 agreed with human categorizations more than half the time across all four tasks and outperformed GloVe and the random baseline.
  • all-MiniLM-L6-v2 also consistently outperformed all-mpnet-base-v2 by a small margin.
  • OpenAI embedding models were excluded because their proprietary, costly, non-local, and closed-source nature limits research replicability.
  • Benchmark evidence indicated text-embedding-ada-002 was not clearly much better or worse than several SentenceTransformers models across tasks.
Loading 2402.01536v2…