Source-linked AI summary
AI and Its Impact on Creativity and Diversity: An Empirical Study of LLM-Generated Product Ideas
Christian Terwiesch, Lennart Meincke, Karan Girotra, Ethan Mollick, Gideon Nave, Karl T. Ulrich
TL;DR
The paper examines fundamental trade-offs in using AI when quantity and variance in ideas are desirable. It provides a detailed empirical assessment of LLMs such as GPT-4, finding strong quality performance while identifying reduced exploration of the solution landscape as an important concern.
Problem
Using AI for idea generation involves fundamental trade-offs when quantity and variance in outcomes are desirable, because the solution landscape may be less fully explored.
Method
The paper provides a detailed empirical assessment of large language models, such as GPT-4, for AI-assisted product innovation.
Results
GPT-4-generated ideas are 7x more likely to appear among the top 10% than human ideas, while GPT-4 also generates better ideas on average.
Takeaways & Limitations
AI can produce especially capable top-tier ideas, but its brainstorming can leave the underlying solution landscape less fully explored.
Takeaways & Limitations
The paper documents a critical limitation of AI-powered brainstorming in its setting: a loss of idea diversity.
Abstract
from arXiv · showhide
This research examines how well large language models, or LLMs, generate new product ideas for college students priced under $50. Across a series of studies, we identify key strengths and weaknesses of using LLMs for product innovation. Our first study shows that LLM-generated product ideas have higher average quality than human ideas, based on purchase intent, and are 7 times more likely to rank in the top 10%. Our second study shows that this AI-induced creativity boost is not explained by the LLM's more persuasive pitching skills. Our third and fourth studies identify a weakness of using LLMs for brainstorming: AI-generated ideas are less novel at the idea level and less diverse at the set level. In our fifth study, we analyze prior LLM-based creativity studies and find consistently lower idea diversity across all of them, demonstrating the generalizability of these findings. Our sixth and seventh studies investigate techniques to mitigate this diversity loss. We compare LLMs from different vendors and versions and find that more recent models generate more diverse ideas, though they still fall short of human-level diversity. We also demonstrate techniques that increase idea diversity almost to the level of human idea generation: pooling ideas across vendors; prompt engineering, including Chain-of-Thought prompting and injecting heterogeneous personas or constraints; and creative agents that broadly explore the solution landscape to restore diversity. Finally, in our eighth study, we show that exploiting the near-zero marginal cost of AI idea generation by scaling the number of ideas steadily improves coverage of the idea space, approaching human-level coverage. We conclude by presenting actionable recommendations for innovation managers who want to identify better new product ideas with the help of LLMs.
1. Introduction
This research compares human and LLM-generated product ideas to assess both innovation performance and the trade-off between idea quality and diversity. It finds stronger average and top-tier quality for AI ideas, but reduced diversity and incomplete exploration that can be mitigated through model, prompting, agent, pooling, and scaling strategies.
- Idea quality: LLM-generated product ideas have higher average quality than human-generated ideas when evaluated using customer purchase intent.
- Idea quality: AI’s superior idea performance cannot be attributed to pitching or communication skills, because AI-rephrased human ideas did not meaningfully change purchase intent.
- Idea diversity: Reduced diversity makes the underlying solution landscape less likely to be fully explored, despite AI’s stronger top-tier idea performance.
- Idea diversity: AI-generated ideas are less diverse than human-generated ideas, producing sets whose ideas are more similar to one another.
- Diversification strategies: More recent models, heterogeneous prompts, Chain-of-Thought prompting, cross-vendor pooling, creative agents, and larger idea pools can increase diversity or coverage toward human levels.
2. Study 1: The Impact of AI on Idea Quality
Study 1 finds that GPT-4-generated product ideas receive higher purchase-intent ratings than human ideas and are especially overrepresented among the highest-quality ideas. This advantage is not explained by writing style, while AI ideas are slightly less novel and few-shot examples may further reduce novelty.
- Study 1a: GPT-4 ideas had higher average purchase intent than human ideas: 46.4% with zero-shot and 49.3% with few-shot prompting versus 40.4%.The differences versus human ideas were statistically significant, while the zero-shot and few-shot AI pools did not differ significantly.
- Study 1c: Human ideas had higher average novelty than GPT-4 ideas: 40.6% versus 36.7% for zero-shot and 36.1% for few-shot GPT-4.The authors report that AI ideas were only slightly less novel and did not simply reproduce existing ideas.
- Study 1a: 7x more GPT-4 ideas appeared among the top 10% of ideas than human ideas.Of 40 top-decile ideas, 5 came from humans, 15 from zero-shot GPT-4, and 20 from few-shot GPT-4.
- Study 1 discussion: Few-shot prompting improved average and best-idea quality but appeared to reduce novelty, suggesting that examples may anchor the model.The reported novelty decrease was more pronounced with few-shot prompting.
- Study 1b: The AI quality advantage was not attributable to persuasive writing style: rephrasing human ideas with AI did not significantly change purchase intent.The estimated rephrasing effect was 0.009, with a 95% confidence interval of [-0.005, 0.023].
3. Study 2: The Impact of AI on Idea Diversity
Study 2 examines whether LLM-generated idea portfolios cover as much of the solution space as human portfolios. Across visualizations, similarity analyses, and multiple datasets, AI ideas are less diverse, concentrating in commercially plausible but less novel regions while humans explore more unusual regions.
- Conceptualization of idea diversity: Idea diversity is defined as semantic distance between ideas in a set, measured here through embedding vectors and cosine similarity.Ideas are represented as positions in a high-dimensional idea space; closer positions and more similar vector directions indicate greater similarity.
- Study 2a: College products: 20 of 35 embedding-space squares contained human ideas, compared with 16 for AI ideas.The visualization indicates broader human coverage of the projected idea space.
- Study 2a: College products: GPT-4 zero-shot ideas were more similar than human ideas (B = 0.193; 95% CI [0.182, 0.204]), while few-shot ideas were less diverse than both human and zero-shot ideas.Few-shot ideas differed from human ideas by B = 0.207 and from zero-shot ideas by B = 0.014.
- Study 2b: Meta-analysis: Pooling the five datasets, LLM-generated ideas were less diverse than human ideas (B = 0.090; 95% CI [0.038, 0.142]; t(4.01) = 3.38, p = .028).The cross-dataset estimate supports a persistent diversity gap between LLM and human idea generation.
- Interpretation: LLMs can produce stronger average and top-tier ideas while concentrating in commercially plausible regions and potentially overlooking less-explored areas containing breakthrough innovations.The paper characterizes this as a tension between superior performance within explored regions and narrower exploration of the solution landscape.
- Interpretation: The reduced-diversity pattern appeared across creative tasks, AI models, diversity methods, and prompting strategies, suggesting it may be a general LLM property.Effect sizes varied with task complexity, domain specificity, and the creative challenge; the effect was especially pronounced in complex, open-ended tasks.
4. Study 3: Increasing Diversity
Study 3 evaluates ways to increase the diversity of LLM-generated ideas through model selection, pooling, prompting, creative agents, and scaling. Newer models and targeted interventions improve diversity, while pooling across models and human-guided exploration approach human-level performance, though some regions remain difficult to reach.
- Models, Parameters and Pooling: Newer LLMs generally generated more diverse ideas than the older GPT-4 model, although performance varied across models and release dates.All but one tested model significantly exceeded the old GPT-4 model, but no individual model approached human diversity.
- Models, Parameters and Pooling: Pooling ideas across models performed nearly as well as humans, offering a strategy to reduce individual model limitations.The composite approach achieved B = -0.010 with a 95% CI of [-0.022, 0.002].
- Prompting and Creative Agents: Chain-of-Thought prompting, heterogeneous personas, and functional constraints increased GPT-4o idea diversity, whereas adding sonnets produced much smaller effects.Independent sessions generating one idea sharply decreased diversity, showing that the intervention's structure matters.
- Prompting and Creative Agents: The Bold Ideas agent reduced diversity by repeatedly refining ideas toward common features, while human-directed problem exploration achieved the second-strongest performance.The refinement process converged on characteristics such as scents, resembling local search trapped near a limited solution peak.
- Prompting and Creative Agents: More eccentric personas and constraints pushed ideas farther from human ideas but left them heavily clustered in underexplored regions.Eccentricity did not monotonically increase diversity, indicating that displacement from human ideas and dispersion among AI ideas are distinct outcomes.
- Scaling the Number of Ideas: Increasing the number of AI ideas steadily improved coverage of the human idea space, with mean nearest-neighbor distance falling from 0.546 at N = 100 to 0.478 at N = 1,000.The decline showed no sign of plateauing within the tested range, although the hardest-to-match human ideas remained difficult for LLMs to reach.
- Implications: The practical diversity gap depends on how managers use AI, with large inexpensive idea pools followed by curation offering a viable innovation-tournament strategy.The authors recommend involving humans in ideation and evaluation because targeted human guidance can substantially improve AI-generated diversity.
5. Conclusion
The conclusion finds that LLMs generate higher-quality product ideas but less diverse idea sets, creating a quality–diversity trade-off. It identifies model pooling, prompt and agent design, human direction, and large-scale generation as practical ways to improve coverage.
- Findings: LLM-generated ideas score higher in average purchase intent than human-generated ideas and are seven times more likely to reach the top decile.The performance reflects idea quality rather than superior pitching or communication skills.
- Findings: AI-generated ideas exhibit greater semantic similarity and converge on a narrower region of the solution space than human ideas.This diversity loss can leave conceptual regions underexplored and reduce the likelihood of identifying radical innovations.
- Findings: Reanalysis of four prior LLM creativity studies confirms reduced diversity as a robust empirical pattern.The finding suggests that lower diversity may be a general property of LLM-assisted ideation rather than an isolated result.
- Mitigating diversity loss: More recent models, cross-vendor idea pooling, prompt engineering, heterogeneous personas or constraints, and creative agents can increase idea diversity.Creative agents that explore distinct regions can restore diversity to near-human levels in this setting.
- Mitigating diversity loss: LLMs are most effective when humans direct them toward specific open needs, combining human insight about unmet needs with AI-generated solutions.The conclusion proposes this as a division of labor in AI-augmented innovation.
- Scaling generation: Increasing the number of AI ideas steadily improves coverage of the human idea space, with no diminishing returns observed at 1,000 ideas.Some regions of the human idea space remain systematically harder to reach.
- Scaling generation: The study recommends generating ideas at scale and selecting among them as a complementary, low-effort strategy for innovation managers.This approach exploits AI’s near-zero marginal generation cost rather than optimizing only one model or prompt.
- Limitations: The main analysis uses a single college-market product-idea tournament, limiting generalizability across innovation challenges and settings.The authors call for replication across other tournaments, product-development settings, and innovation tasks.
A. Prompting
The prompting procedure uses a novice-oriented GPT-4 setup with controlled randomness, repeated batches, and prompts intended to preserve context and encourage distinct ideas. It also includes zero-shot and few-shot variants, while context limits require compression for larger idea pools.
- A. Prompting: GPT-4 was prompted with the same product-idea task given to students, targeting physical goods for U.S. college students priced below about USD 50.The prompt requested numbered ideas with names and separate paragraphs.
- A. Prompting: The study used minimal prompt engineering to represent a novice-user scenario rather than an extensively optimized workflow.The authors acknowledge that many strategies could potentially improve LLM performance.
- A. Prompting: The zero-shot condition supplied contextual information and requested ten ideas at a time, each described in 40–80 words.The prompt included an explicit request for distinct ideas in successive batches.
- A. Prompting: The model used gpt-4-0314 with temperature 0.7, balancing coherence and creativity through controlled output randomness.Lower temperature produces more deterministic output, whereas higher temperature increases variability.
- A. Prompting: Because GPT-4’s context window held roughly 80 ideas, earlier ideas were compressed into summaries before later batches were generated.The summaries were supplied back to the model so it could avoid repetition while remaining within token limits.
- A. Prompting: The few-shot condition appended six highly rated student ideas to provide examples of well-received product concepts.Six examples were used because of the context-window limitations at the time.
- B. Idea Rephrasing Prompt: A rephrasing prompt instructed GPT-4 to rewrite product descriptions in a similar style and length without adding unspecified details.Examples included earplugs, multifunctional furniture, snack boxes, study booths, and multitools.
C. Idea Descriptions
The idea descriptions illustrate products aimed at common college-student needs, including ergonomic study support, portable organization, hydration, and multifunctional dorm-room use.
- Study and organization: BookBuddy is a lightweight, foldable holder that supports textbooks, tablets, and laptops at an ergonomic viewing angle.Adjustable arms and a non-slip base provide stability across settings.
- Study and organization: An Adjustable Laptop Riser elevates laptops to improve viewing angle, posture, and eye comfort during extended use.Its durable construction, non-slip surface, and foldable design support portability.
- Mobility and hydration: Hydration Harness is an adjustable strap that attaches large water bottles to backpacks while keeping students’ hands free.The concept addresses backpack side pouches that are too small for some bottles, especially for gym-goers and athletes.
D. USE Details
The study represents each idea as a fixed-length semantic embedding and measures pairwise similarity to quantify how closely ideas relate in meaning.
- Representation: The Universal Sentence Encoder maps variable-length idea text to fixed-length 512-dimensional embedding vectors.The pretrained model is intended to capture semantic relationships without task-specific fine-tuning.
- Representation: Semantically related texts map to nearby embedding-space points, while semantically distinct texts map farther apart.This geometric representation provides a standardized basis for comparing ideas.
- Similarity measure: Pairwise cosine similarity between idea embeddings serves as the primary measure of idea similarity.The measure follows prior creativity and ideation research cited by the study.
E. Rephrasing Details
The study rephrased human and GPT-4 ideas into a standardized style to test whether stylistic variation explained the observed diversity gap. The gap persisted after standardization, although removing stylistic differences reduced it modestly.
- Standardization procedure: The authors standardized 400 ideas—200 human, 100 GPT-4 zero-shot, and 100 GPT-4 few-shot—using Claude Sonnet 4.5.The procedure preserved content while imposing consistent third-person, declarative prose and similar lengths.
- Measures: Embedding-based diversity was measured with within-group mean pairwise cosine similarity, while fidelity measured similarity between each idea's original and standardized embeddings.Higher fidelity indicates that standardization changed style without altering semantic content.
- Zero-shot results: B = 0.154, 95% CI [0.141, 0.168], t(298) = 21.82, p < .001 after standardization, retaining 80% of the original zero-shot diversity gap.Controlling for word count reduced the effect only marginally to B = 0.149, 95% CI [0.136, 0.161].
- Few-shot results: B = 0.186, 95% CI [0.172, 0.201], t(298) = 25.69, p < .001 after standardization, retaining 90% of the original few-shot gap.With word count controlled, the effect remained B = 0.177, 95% CI [0.164, 0.190].
- Interpretation: The residual diversity gap was interpreted as reflecting topical differences after accounting for stylistic dissimilarity.The authors note that human stylistic heterogeneity may have contributed 10–20% of the original gap.
G. Entropy Prompting Approaches
The study tests whether prompt-induced entropy can move LLM product ideas beyond their usual narrow region of idea space. It varies personas, functional constraints, and unrelated textual stimuli while holding the model and core prompt structure constant.
- Motivation: The study asks whether narrow LLM ideation is fundamental or an artifact of prompting, motivating interventions that target different regions of idea space.The proposed interventions vary who the model is, which constraints it must respect, or what stimulus it has recently attended to.
- Experimental design: Three entropy-injection variants sampled stochastic elements from fixed pools while holding GPT-4o, temperature 1.0, and the original prompt structure constant.The variants were persona prompts, functional constraints, and Shakespearean sonnets as content-irrelevant stimuli.
- Persona prompting: Persona prompts assigned varied student identities intended to shift the model's representation of the target user and the experiences used during ideation.The persona pool covered majors, extracurriculars, life circumstances, and backgrounds.
- Constraint prompting: Functional constraints required each proposed product to satisfy an added condition, such as folding flat, using no electricity, or serving two functions.The constraint pool contained 20 requirements, each used once per session, to rule out easy modal solutions.
- Stimulus prompting: Shakespearean sonnets served as content-irrelevant entropy injections by exposing the model to fresh tokens before product ideation.The sonnets had no semantic connection to product design but could steer subsequent generation through the residual stream.
H. Persona and Constraint Variations
Persona and constraint eccentricity did not produce a simple monotonic increase in diversity. Although some intermediate conditions approached the human baseline, every tested level remained significantly less diverse than human ideas.
- Variation design: Five-level eccentricity gradients ranged from homogeneous to outlandish personas and from mild to surreal constraints.The persona scale moved through wealthy professionals, diverse students, unusual identities, and outlandish fictional or non-human figures.
- Overall results: All five levels of each strategy were significantly less diverse than the human baseline, with p < .001 except the extreme-constraint tier at p < .01.All reported p-values were uncorrected.
- Personas: Persona diversity was lowest for homogeneous prompts, but greater eccentricity did not keep increasing diversity.The diverse-student set exceeded the most eccentric persona set significantly, while its advantage over the second-most eccentric set was not significant.
- Constraints: Constraint diversity peaked at the extreme-yet-physically-possible level rather than at the demanding or surreal levels.This condition was closest to the human baseline, but remained significantly less diverse; surreal constraints reversed the gain.
I.1. Supplemental Figures
Supplemental analyses describe the figures, compare AI and human ideas on purchase intent and novelty, and examine coverage of the human idea space as the LLM sample grows. They also note model- and measurement-related limitations.
- Caveats: The supplemental analyses caution that diversity comparisons may reflect model-specific creativity differences, prompt sensitivity, and limitations of embedding-based similarity.The study used random samples of 100 ideas from larger pools, and embedding models may conflate stylistic with topical variation.
- Best-of-both-worlds ideas: Humans and AI produced 48 high-purchase-intent/high-novelty ideas each, so the main difference was not in that quadrant.The contrast instead concerned the distribution of ideas across the other purchase-intent and novelty combinations.
- Quality–novelty trade-off: AI ideas were more common in the high-purchase-intent/low-novelty quadrant, while human ideas were more common in the low-purchase-intent/high-novelty quadrant.The counts were 74 AI versus 30 human ideas in the former and 69 human versus 34 AI ideas in the latter.
- Idea-space coverage: For coverage analysis, each human idea was compared with its nearest neighbor in increasingly large GPT-4o pools from N = 100 to 1,000.Lower nearest-neighbor distance indicates better coverage of the human idea space.
I.2. Supplemental Tables
The supplemental material documents the paper’s research questions, study designs, analyses, and evidence on AI-generated idea quality, novelty, and diversity. It also reports comparisons across models, prompting conditions, and idea-generation strategies.
- Supplemental evidence: The supplementary tables include purchase-intent, novelty, similarity, regression, model, and temperature analyses supporting the paper’s broader comparisons.The listed tables cover human baselines, GPT-4 conditions, prior studies, different models, and different temperatures.
- Research questions: The studies examine AI brainstorming’s effects on idea quality, pitching explanations, novelty, and diversity.The research also asks whether vendor choice, model configuration, prompting strategies, creative agents, and scaling can improve AI idea sets.
- Prior research: Prior work reports mixed evidence, with LLMs sometimes matching or exceeding human creativity while also producing less diverse outputs.The cited literature includes findings on divergent thinking, strategic viability, novelty, persuasion, and diversity loss.
- Methods: The supplemental analyses use purchase-intent ratings, regression models, cosine similarity, minimum spanning trees, and multiple diversity metrics.Several analyses compare human and GPT-generated ideas, while others examine rephrasing, models, temperatures, personas, and solution multiplicity.
- Diversity interventions: Multiple solutions and multiple problems show large reported differences from single-solution or single-problem conditions, including d = 2.35, d = 2.71, and d = 4.403.The comparisons include expert and random persona conditions and are evaluated with similarity-based diversity measures.