Source-linked AI summary
Artificial muses: Generative Artificial Intelligence Chatbots Have Risen to Human-Level Creativity
Jennifer Haase, Paul H. P. Hanel
TL;DR
The paper tests whether humans remain more creative than generative AI chatbots by comparing their ideas and having humans and AI assess them. The chatbots produced ideas judged as original as human ideas, supporting their use as assistants while leaving key parts of the creative process to humans.
Problem
The paper tests the widespread belief that creativity is one of the areas where humans remain better than AI.
Method
The study compares human-generated ideas with ideas from six GAI chatbots, assessed independently by humans and a specifically trained AI.
Results
GAI chatbot ideas were judged as original as human-generated ideas, while chatbots generated more ideas and differed in how they produced them.
Takeaways & Limitations
GAI chatbots can support creative work by generating and expanding ideas, but humans remain responsible for understanding problems, selecting solutions, and implementing them.
Takeaways & Limitations
The study did not measure usefulness, and the validity of its originality measure, the AUT, remains under debate.
Abstract
from arXiv · showhide
A widespread view is that Artificial Intelligence cannot be creative. We tested this assumption by comparing human-generated ideas with those generated by six Generative Artificial Intelligence (GAI) chatbots: $alpa.\!ai$, $Copy.\!ai$, ChatGPT (versions 3 and 4), $Studio.\!ai$, and YouChat. Humans and a specifically trained AI independently assessed the quality and quantity of ideas. We found no qualitative difference between AI and human-generated creativity, although there are differences in how ideas are generated. Interestingly, 9.4 percent of humans were more creative than the most creative GAI, GPT-4. Our findings suggest that GAIs are valuable assistants in the creative process. Continued research and development of GAI in creative tasks is crucial to fully understand this technology's potential benefits and drawbacks in shaping the future of creativity. Finally, we discuss the question of whether GAIs are capable of being truly creative.
1. Main
The paper tests whether humans remain more creative than generative AI chatbots, focusing on everyday creative idea generation and the meaning of machine creativity. It frames GAI as capable of producing original ideas but not the full human creative process.
- 1. Main: The study compares human creativity with six GAI chatbots to test whether creativity remains a distinct human advantage.The comparison uses both human and AI judges.
- 1. Main: GAI is increasingly used for complex tasks, raising interest in whether it can support and enhance human creativity.The paper also notes debate over whether AI creates genuinely new ideas or recombines existing knowledge.
- 1. Main: The paper situates GAI within broader efforts to generate novel insights and support human work across domains such as healthcare, education, and entertainment.It also reviews examples including AI-generated music, game strategies, and tools for artists and designers.
- 1. Main: Creativity is operationalized as producing something perceptibly new and useful, a standard that does not require machines to replicate human experiences or emotions.Under this pragmatic view, creative output matters more than whether the system possesses human-like attitudes or behaviors.
- 1. Main: The research focuses on everyday creativity, while humans remain responsible for understanding the problem, judging fit, and implementing selected ideas.GAI therefore contributes primarily to idea generation rather than the full holistic creative process.
2. Results
The study compares human and GAI responses on the Alternative Uses Test, measuring originality and fluency across multiple everyday-object prompts. It uses a widely applied creativity test to assess both idea quality and quantity.
- 2. Results: The Alternative Uses Test was administered to 100 humans and five GAIs using five everyday objects as prompts.Participants generated multiple original uses for pants, a ball, a tire, a fork, and a toothbrush.
- 2. Results: Originality and fluency were assessed as measures of creativity quality and quantity.The Alternative Uses Test is described as one of the most frequently used creativity tests with good predictive validity.
Originality
Across the originality analyses, human and GAI-generated ideas showed no overall mean difference, although GPT-4 performed best among the tested GAIs. A minority of humans still exceeded GPT-4.
- Originality: rs = .78 - .94, ps < .0001: human and AI originality ratings showed very high agreement across the five prompts.Human-rater intraclass correlations ranged from .85 to .94, indicating strong agreement among human raters as well.
- Originality: B = -0.21, SE = 0.15, p = .218 for human ratings and B = -0.18, SE = 0.13, p = .241 for AI ratings: neither model found a mean originality difference between humans and GAIs.The results were mostly replicated in between-subject t-tests.
- Originality: 32.8 humans were more original than the most original GAI across all prompts.This comparison counts human participants exceeding the best GAI chatbot.
- Originality: None of the five GAIs was more original than the other four across all five prompts.The chatbot rankings varied by prompt rather than producing one consistently dominant system.
- Originality: GPT-4 outperformed the other five GAIs except on the ball prompt, where it ranked second.Its responses were evaluated only by the AI because human raters might have recognized them as nonhuman.
- Originality: 9.4 humans were more creative than GPT-4 across all prompts on average.The prompt-level counts were 2 for pants, 29 for ball, 0 for tire, 3 for fork, and 13 for tooth.
Fluency
GAI chatbots generated substantially more ideas than humans when prompted repeatedly, while fluency and originality were largely unrelated.
- 2-3 times more ideas were generated by GAI chatbots than by humans.Most chatbots were prompted multiple times.
- Fluency and originality were mostly unrelated, with correlations ranging from -.28 to .26.
3. Discussion
The discussion presents GAI chatbots as capable of human-level everyday ideation while remaining dependent on humans to define problems, interpret ideas, and implement solutions. It also identifies methodological and ethical boundaries that constrain how broadly these findings should be applied.
- Creativity assessment: GAI chatbot ideas were judged as original as human-generated ideas on a standardized broad-associative creativity measure.The output was judged by both humans and AI and was indistinguishable from human output.
- Creativity assessment: GAI chatbots generated more ideas than humans on the same simple prompts, and their ideas were on average similarly original.The authors caution against emphasizing fluency because assessment styles and the sheer number of ideas were not fully comparable.
- Scope of creativity: The study found that chatbots can compete with human ideation skills for everyday creativity, but complex solutions require domain knowledge, experience, emotion, cultural background, and abstract thinking.ChatGPT performed well on complex knowledge-intensive tasks but showed limitations in emotional responses, evaluations, and more complex tasks.
- Human role: GAI can generate ideas from specific input but cannot create the need to ideate; human motivation and problem understanding must initiate the creative task.
- Human role: Because GAI responds only to a prompt, humans currently define problems and judge whether generated ideas fit them.AI was considered sufficient for assessing AUT output in the particular context studied.
- GAI as assistant: GAI can identify seemingly new connections and generate ideas, while humans must make sense of them, embed them in physical reality, and determine the achievement level.ChatGPT-4 showed the best originality results, followed by Copy.ai, ChatGPT-3, and YouChat at similarly high levels.
- Limitations: The experimental design may have underestimated both human and GAI creativity because participants lacked intrinsic motivation and the chatbots were not tested with smart, tailored prompts.The study also did not test profession-specific prompting or answer reshaping, which the authors suggest could improve relevance and quality.
- Limitations: The AUT's validity remains debated, possible training-data overlap could affect chatbot responses, usefulness was not measured, and fluency assessment was not very meaningful.
Participants
The study recruited 100 working native English speakers and evaluated five initially selected GAI chatbots, later adding GPT-4 responses.
- Participants: 100 participants were recruited through Prolific Academic; all were native English speakers from the USA with full- or part-time work.The sample included 50 women and 50 men, with Mage = 41.00 and SD = 12.25.
- GAI chatbots: Five chatbots—Alpa.ai, Copy.ai, ChatGPT version 3, Studio, and YouChat—were initially selected for free usability and similar functions.The selection was intended to ensure comparability.
- GAI chatbots: Copy.ai responses were generated with its “freestyle” template because its chatbot function appeared limited.The template’s “more like this” feature was used three times to generate additional ideas.
- GAI chatbots: ChatGPT was described as a language model based on the GPT-3 database that generates human-like responses to natural-language inputs.GPT-4 responses were added after the initial data collection and analyses.
- GAI chatbots: AI21 Studio was used through its Playground interface, which the authors considered closest to a chatbot tool.Studio is based on AI21 Labs’ Jurassic-1 large language model.
- GAI chatbots: YouChat was included as a messaging platform and AI-powered search assistant created by You.com.The platform supports questions, explanations, recommendations, translation, and summarization.
Materials
Participants completed the Alternative Use Test across five everyday objects, generating as many ideas as possible within a fixed time.
- Materials: Participants completed the Alternative Use Test five times for a ball, fork, pants, tire, and tooth.These objects are commonly used in creativity tests and were considered reliably assessable by the AI rater.
- Materials: Human participants had three minutes per object to write down as many ideas as possible.
Procedure
Responses were rated blindly by human evaluators using the CAT method, while fluency was calculated from the assessed number of ideas.
- Procedure: Data was collected in early February 2023, before GPT-4’s release on 14 March.GPT-4 was therefore not included among the five chatbots rated in the initial collection.
- Procedure: Six human raters evaluated human responses and responses from five GAI chatbots while blind to response origin.Prompt order and idea lists were randomized throughout the rating process.
- Procedure: Human raters followed the CAT method and used the full 1–5 originality scale.The study also assessed originality scores.
- Procedure: Fluency scores were calculated as the sum of ideas from each participant and chatbot.Rater differences in coding non-relevant answers as no-answer caused slight variation in the sums.