Source-linked AI summary
Art or Artifice? Large Language Models and the False Promise of Creativity
Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, Chien-Sheng Wu
TL;DR
Objective evaluation of creativity in LLM-generated writing remains difficult. The paper adapts TTCT with expert judgments into TTCW, a 14-test product evaluation for short stories, and compares human and LLM-generated stories while testing LLM assessors. LLM-generated stories pass three to ten times fewer TTCW tests than professional stories, and the evaluated LLM assessors do not positively correlate with experts.
Problem
Objectively evaluating the creativity of writing is challenging despite LLMs’ demonstrated writing capabilities across genres.
Method
The paper adapts TTCT and the Consensual Assessment Technique into 14 binary TTCW tests covering fluency, flexibility, originality, and elaboration, then evaluates 48 stories with 10 experts and tests LLM assessors.
Results
LLM-generated stories pass three to ten times fewer TTCW tests than professional stories, and the three evaluated LLM assessors achieve correlations with experts close to zero.
Takeaways & Limitations
TTCW provides an expert-grounded framework for assessing creativity in fictional short stories, while current LLMs do not reproduce expert TTCW assessments reliably.
Takeaways & Limitations
TTCW was designed for short fiction, and its adequacy for scripts, novels, slogans, poetry, and other creative-writing forms remains to be empirically verified.
Abstract
from arXiv · showhide
Researchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by the Torrance Test of Creative Thinking (TTCT), which measures creativity as a process, we use the Consensual Assessment Technique [3] and propose the Torrance Test of Creative Writing (TTCW) to evaluate creativity as a product. TTCW consists of 14 binary tests organized into the original dimensions of Fluency, Flexibility, Originality, and Elaboration. We recruit 10 creative writers and implement a human assessment of 48 stories written either by professional authors or LLMs using TTCW. Our analysis shows that LLM-generated stories pass 3-10X less TTCW tests than stories written by professionals. In addition, we explore the use of LLMs as assessors to automate the TTCW evaluation, revealing that none of the LLMs positively correlate with the expert assessments.
1 INTRODUCTION
The paper introduces TTCW, an expert-grounded protocol for evaluating creativity in short stories as a product, and applies it to human- and LLM-written stories and LLM assessment. Expert-written stories substantially outperform LLM-generated stories, while LLM assessors show near-zero agreement with experts.
- Protocol and motivation: TTCW adapts TTCT’s process-oriented creativity framework to product evaluation of short stories using the Consensual Assessment Technique.The protocol centers on fluency, flexibility, originality, and elaboration.
- Evaluation design: The benchmark contains 48 short stories: 12 written by professionals and 36 generated by ChatGPT, GPT4, and Claude 1.3.The stories average 1400 words, and 10 experts administer the 14 TTCW tests with three evaluations per story.
- Evaluation design: Fleiss Kappa 0.41 across the 14 TTCW tests and Pearson correlation 0.69 for aggregated tests indicate moderate item-level and strong aggregate expert agreement.The findings are based on more than 2,000 expert-administered tests.
- Main findings: 84.7% of tests were passed on average by expert-written stories, compared with 9% for ChatGPT-generated stories and up to 30% for Claude-generated stories.LLM-generated stories were three to ten times less likely than expert-written stories to pass individual TTCW tests.
- Main findings: GPT4 was more likely to pass Originality tests, whereas Claude V1.3 was more likely to pass Fluency, Flexibility, and Elaboration tests.The TTCW dimensions reveal specialization differences among LLMs beyond the overall creativity gap.
- Main findings: Three evaluated LLMs achieved correlations with expert assessments close to zero when administering TTCW tests.The study expanded each test into a detailed prompt and compared LLM judgments with collected expert judgments.
2 RELATED WORK
Prior work frames creativity and creative-writing evaluation through divergent thinking, rubrics, structural complexity, narrative analysis, and human assessment. Research also questions whether non-experts can reliably distinguish or evaluate machine-generated writing.
- Creativity measurement: Divergent thinking remains a common indicator of creativity, alongside consensual, peer or teacher, and self-assessment methods.The paper grounds its evaluation approach in divergent-thinking research and TTCT.
- Creative-writing evaluation: Writing rubrics assess prominent characteristics relevant to a specific discourse type, including fiction-specific narrative and stylistic elements.Prior fiction rubrics cover elements such as narrative voice, characterization, setting, mood, dialogue, plot, and mechanics.
- Creative-writing evaluation: The SOLO taxonomy evaluates creative writing by structural complexity, ranging from incoherent and linear forms to integrated and metaphoric forms.The taxonomy organizes products into prestructural, unistructural, multistructural, relational, and extended-abstract categories.
- Evaluator expertise: Crowd workers may be unsuitable for evaluating AI-generated writing, while non-experts can struggle to discriminate model-generated text from human references.The related work motivates careful consideration of who assesses creative-writing quality.
- Evaluator expertise: Studies report that untrained evaluators can distinguish GPT3-generated from human-authored text only at random-chance levels across stories, news articles, and recipes.This supports examining evaluator expertise when assessing generated writing.
3 DESIGN CONSIDERATIONS
The paper designs TTCW as an artifact-centered creativity protocol for short fiction, adapting TTCT’s four dimensions into additive binary tests. This design addresses limits of process-based evaluation when creative processes are unobservable or difficult to assess objectively.
- TTCW design: 14 TTCW tests evaluate short-story creativity across fluency, flexibility, originality, and elaboration, adapting TTCT from process to product assessment.The tests were developed with domain experts and are intended to assess fictional artifacts.
- Artifact-centric testing: Artifact-centric testing evaluates the final writing rather than the cognitive process that produced it.The paper uses “artifact” for the product of creative writing.
- Artifact-centric testing: Process-based evaluation can be weakened by unobserved thoughts and activities, difficulty separating process from artifact, and inaccessible processes in preexisting works or LLMs.These constraints motivate evaluating the artifact itself.
- Binary testing: Each TTCW test uses a binary Yes/No question, with open-ended rationales supporting the assessment.The binary format is applied to multiple questions within each Torrance dimension.
- Additive testing: The final assessment is the number of tests passed, because individual tests are independent and no single pass or failure constitutes a complete creativity judgment.Aggregate scores are intended to provide a fuller picture of an artifact’s creativity.
4 FORMATIVE STUDY: FORMULATING THE TORRANCE TESTS FOR CREATIVE WRITING
The formative study converts expert judgments about fiction into 14 actionable TTCW tests across four Torrance dimensions. Experts proposed and refined measures that address narrative craft, including pacing, scene–exposition balance, perspective, emotion, literary devices, coherence, and character development.
- 4 FORMATIVE STUDY: FORMULATING THE TORRANCE TESTS FOR CREATIVE WRITING: Experts were selected for formal creative-writing training, traditional publication, or university-level fiction instruction to align recruitment with the Consensual Assessment Technique.The study excluded self-published authors.
- 4 FORMATIVE STUDY: FORMULATING THE TORRANCE TESTS FOR CREATIVE WRITING: The formative study briefed experts on fiction-focused creativity metrics and Torrance dimensions, then collected their proposed measures through a web application.The study was organized in three parts, beginning with a briefing and followed by online measure submission.
- 4.1 From Measures to Actionable Tests: Experts’ proposed measures showed semantic convergence, including similar formulations for originality through formal novelty and elaboration through three-dimensional characters.These overlaps supported consolidation into shared test categories.
- 4.1 From Measures to Actionable Tests: Three authors inductively grouped expert measures into shared low- and high-level categories before constructing the TTCW framework.The grouping process reduced category overlap through repeated discussion.
- 4.1 From Measures to Actionable Tests: 14 distinct test groups were formed: five for Fluency and three each for Flexibility, Originality, and Elaboration.A representative measure was selected for each group, and GPT4 converted some measures into Yes/No questions.
- 4.2 The Torrance Test for Creative Writing: Narrative Pacing tests whether time compression or stretching creates appropriate, balanced storytelling speed and rhythm.The test operationalizes experts’ advice to control how quickly a story unfolds.
- 4.2 The Torrance Test for Creative Writing: Scene vs Exposition tests whether a story balances real-time dramatization with summarized information such as backstory and setting details.The balance supports pacing, reader engagement, and information delivery.
- 4.2 The Torrance Test for Creative Writing: Other TTCW tests assess sophisticated literary devices, narrative coherence, diverse and convincing perspectives, emotional balance, and appropriately complex character development.Together these measures extend the framework beyond surface novelty to multiple aspects of fictional craft.
5 TTCW IMPLEMENTATION WITH EXPERTS AS ASSESSORS
The study evaluates 48 stories with TTCW assessments by creative-writing experts, comparing professional New Yorker stories with outputs from three LLMs. Human-written stories pass substantially more tests, while aggregate scores show stronger agreement than individual-test judgments.
- Study design: Stories were generated from one-sentence plots of New Yorker originals so evaluation could focus on writing form rather than plot-line creativity.The plots were automatically summarized by GPT-4 and verified by humans; models were prompted to produce similarly long stories.
- Comparative results: 84.7% was the overall pass rate for New Yorker stories, equivalent to 11.9 of 14 TTCW tests on average.No individual test was passed by 100% of the high-quality human stories, supporting use of the tests as an additive set.
- Comparative results: GPT3.5 stories passed less than 10% of TTCW, while GPT4 and Claude v1.3 stories were closer to 30.0%.Across models, Fluency had the highest pass rate; Claude led Fluency, Flexibility, and Elaboration, while GPT4 led Originality.
- Reproducibility: Individual-test agreement averaged Fleiss Kappa 0.41, whereas agreement on aggregate numbers of tests passed reached Pearson correlation 0.69.Experts agreed more strongly on the overall count than on whether particular TTCW tests were passed.
- Comparative results: New Yorker stories were ranked most preferred 89% of the time, while GPT3.5 stories ranked least preferred roughly two-thirds of the time.Claude was almost twice as likely as GPT4 to rank second behind the human story and was preferred in three of four non-human wins.
6 TTCW IMPLEMENTATION WITH LLMS AS ASSESSORS
The study examines whether LLMs can automate TTCW assessments of creative stories, using the same 14 tests across 48 stories and comparing model judgments with expert annotations. None of the evaluated LLMs positively correlated with expert assessments, limiting their reliability as automated assessors.
- Assessment setup: LLMs were prompted to answer the 14 TTCW tests for all 48 stories using the same data selection and three models as the human evaluation.The models were GPT3.5, GPT4, and Claude.
- Assessment setup: Chain-of-thought prompting was used to elicit step-by-step reasoning before each LLM verdict.The resulting explanations were often procedural and longer than expert explanations.
- Results: None of the LLMs produced assessments that correlated positively with expert assessments, with average correlations close to zero.GPT4 exceeded 0.2 on only two of fourteen tests, which still did not qualify as moderate agreement.
- Results: Few-shot prompts did not produce significant correlation gains on the TTCW tests examined.The main prompts were zero-shot and contained no example labels or expert justifications.
- Implications: Reliable automated TTCW assessments could support iterative story editing, including Self-Refine-style revision until drafts pass more tests.The authors release the TTCW benchmark with binary judgments and expert justifications to support future evaluation research.
7 DISCUSSION
The discussion considers how experts distinguish AI-written stories and how TTCW can support targeted writing feedback. It also emphasizes that TTCW is not universal because its design reflects a narrow expert base and particular literary assumptions.
- Expert judgments: Experts were asked to classify stories as expert-written, amateur-written, or AI-generated, although no amateur stories were included.The three-category task was intended to make dataset tracing less likely and provide a more granular prediction scale.
- Expert judgments: The rubric evaluates writing quality rather than penalizing AI-generated text, because fluent and human-like AI writing is difficult to detect directly.Experts nevertheless reported recurring AI writing tics, including predictable paragraph structures and unusual sentence constructions.
- LLMs as research tools: LLMs were also used as research tools for tasks such as generating an initial expanded expert measure and clustering expert explanations.The paper reflects on both the utility and limitations of this use.
- Implications: TTCW tests could provide structured, targeted feedback for planning and reviewing phases of writing.The motivation is to support specificity rather than generic feedback.
- Scope: TTCW is not a universal benchmark because it relies on a narrow expert base and may reproduce Western highbrow-literary biases.The authors identify possible disadvantages for singular viewpoints, experimental forms, non-cathartic endings, and traditional formal structures.
8 LIMITATIONS AND FUTURE WORK
The paper identifies limitations involving generation settings, overlap among TTCW tests, prompt and model dependence, generalization beyond short fiction, and the subjectivity of creativity judgments.
- Generation settings: Default temperature T=1.0 evaluates models in a non-optimized setting, and alternative parameters might produce stories passing more TTCW tests.The authors call for further study of generation parameters as evaluation costs decrease.
- TTCW coverage: TTCW tests may be correlated because several address shared story elements, but the study does not measure overlap between test pairs.Future work should refine the tests and consider additional measures.
- Prompts and models: Results remain dependent on prompt engineering and on closed-source models whose generation quality can change over time.The authors open-source the prompts but note that there is no upper bound on optimizing a task-specific prompt.
- Generalization: TTCW was designed for short fiction, and its adequacy for scripts, novels, slogans, poetry, and other creative forms remains unverified.Different forms may require different fine-grained evaluation metrics.
- Evaluation objectivity: Creativity judgments are shaped by personal tastes, expectations, hindsight, and unclear boundaries between expert and amateur writers.These factors complicate claims of objective evaluation.
9 CONCLUSION
The paper adapts TTCT into TTCW for evaluating creativity in short fiction and compares expert-authored with LLM-generated stories. It finds a substantial holistic gap between seasoned writers and LLMs while identifying dimension-specific differences among models.
- Contribution: TTCW adapts TTCT from process-based creativity evaluation to product-based evaluation of short fictional stories.Its development and validation involved creative-writing experts.
- Findings: Expert- and LLM-authored stories provide a comparative evaluation of creative writing performance across the TTCW dimensions.The tests also identify areas where LLM-generated stories perform weakest.
- Findings: Different LLMs show strengths in different creativity dimensions, but remain far behind human expertise when assessed holistically.The conclusion distinguishes dimension-specific proficiency from overall creative performance.
A.1 NewYorker Data for evaluation
The evaluation data include 12 expert-written New Yorker stories, summarized into single-sentence plots and used as an upper bound for creativity.
- Figure 5 reports the distribution of story word counts in the test set.
- 12 stories from The New Yorker form the expert-written evaluation set.They were summarized into single-sentence plots.
- The stories were selected from highly established literary experts as an upper bound for creativity.The set spans complex themes.
A.2 Expert Perception on the TTCW tests
Experts largely regarded the TTCW rubric as thorough and effective, and said it helped structure their evaluation of storytelling.
- Almost every expert agreed that the TTCW rubric was thorough and effective.
- Experts said the rubric helped them consider different aspects of storytelling more systematically.
A.3 Common themes in TTCW of Originality and Elaboration
This appendix documents the TTCW evaluation materials, expert annotations, and the planned exploration of less costly alternatives to expert assessment.
- Originality and Elaboration: Table 13 collects common themes and issues in expert explanations for Originality and Elaboration tests.
- Assessment explanations: Table 14 compares LLM-generated explanations with expert explanations for binary TTCW assessments.
- Future evaluation: Future work may test whether non-experts correlate with experts, potentially enabling more cost-effective TTCW evaluation.
- Annotation design: The study relies on experts for annotation to maximize experimental validity and assess agreement when evaluating stories with TTCW.