Source-linked AI summary

The Limits of Automatic Evaluation of Creativity in Large Language Models

Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi

arXiv:2608.23705v1cs.CLcs.AIcs.CY

TL;DR

Creativity evaluation for LLM-generated text is difficult because creativity is multidimensional, subjective, and difficult to reduce to standardized metrics. The paper compares human and LLM-generated stories using human ratings, automated metrics, and LLM-as-a-Judge evaluations, finding substantial misalignment with human judgments. LLM judges favor machine-generated text and fail to reliably capture deeper semantic dimensions.

  • Problem

    The paper addresses whether current automatic evaluation schemes appropriately assess human and artificial creativity and correlate with human judgments across multiple dimensions.

  • Method

    The study conducts a multifaceted comparison of human- and LLM-generated texts using human ratings, LLM-as-a-Judge ratings, and commonly used quantitative metrics.

  • Results

    LLM judges show a strong preference for machine-generated text and fail to reliably capture deeper semantic dimensions, while automated scores fail to correlate with human judgments.

  • Takeaways & Limitations

    Automatic evaluation methods, including LLM-as-a-Judge scores, are unreliable proxies for human evaluation of creative story writing.

  • Takeaways & Limitations

    The analysis was restricted in scope, and the study did not isolate differences in writing and evaluative capabilities.

Abstract

from arXiv · show

Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity, yet evaluating creativity in LLM-generated content remains a significant challenge. Here, we investigate whether current automatic evaluation methods can reliably capture human judgments of creativity. We collect human evaluations of human- and AI-generated short stories from the WritingPrompts dataset across 11 dimensions of creativity, and compare these judgments with automated objective metrics and LLM-as-a-Judge evaluations. Our experiments reveal substantial misalignment between automatic evaluations and human assessments. In particular, LLM-based judges exhibit a systematic preference for AI-generated stories, consistently favoring their stylistic characteristics over the unpredictability and other qualities of human-authored texts. Furthermore, correlation analyses show that widely used automatic metrics exhibit near-zero alignment with human judgments across both human- and AI-generated stories, suggesting that they fail to capture important dimensions of creativity. These findings highlight fundamental limitations in current approaches to the automatic evaluation of creative text and underscore the difficulty of reducing the multidimensional and subjective nature of creativity to computational metrics.

Introduction

Creativity evaluation for LLM-generated text remains difficult because creativity is task-dependent, subjective, and hard to capture with standardized metrics. The paper compares human and automated assessments of human- and LLM-generated stories to examine this problem.

  • Evaluation challenge: Creativity depends on the task and observer, making standardized ground-truth-based evaluation metrics particularly difficult to establish.Observers’ prior knowledge, skills, and preferences influence judgments within a domain.
  • Evaluation challenge: The assumption that divergent-thinking abilities correlate with creativity remains unproven for both humans and AI.
  • Existing approaches: Human evaluation captures subjective value, novelty, and surprise but is difficult to collect, validate, replicate, and standardize.Its outcomes also depend on evaluators’ skill levels.
  • Existing approaches: Quantitative metrics are easy to compute but operate at a lower level of abstraction, while LLM-as-a-Judge methods assume LLMs possess evaluative capabilities.
  • Study approach: The study compares 100 human-authored and 100 LLM-generated stories using human ratings, LLM-as-a-Judge ratings, and five quantitative metrics.The assessments cover creativity and 10 related concepts, followed by aggregate-score and correlation analyses.
  • Findings: None of the automated evaluation methods reliably proxies human evaluation, with most metrics showing negative or no correlation with human scores.Elaboration is the only dimension reported to show a positive, albeit weak, correlation with LLM judgments.
  • Findings: LLM judges favor AI-generated stories and appear to prioritize surface-level properties over the broader semantic dimensions considered by humans.

Results

Across automated metrics and LLM-as-a-Judge evaluations, results show weak alignment with human assessments of creativity. LLM judges additionally exhibit systematic preference for LLM-generated stories and capture some structural properties better than subjective qualities.

  • Automatic metrics: Most automatic metric correlations with human creativity judgments fall within [−0.2, 0.2], indicating very weak or negligible relationships.The authors describe this as a substantial disconnect between automated text evaluation and human assessment of narrative creativity.
  • Automatic metrics: The explicitly creativity-oriented Creativity Index fails to exceed ρ = 0.11 across the eleven subjective dimensions.The finding underscores the difficulty of reducing creativity to rigid mathematical and data-driven metrics.
  • Automatic metrics: Negative associations for lexical and semantic variation suggest that diversity alone may not increase perceived creativity without sufficient narrative or thematic coherence.The paper proposes that creatively successful stories require balance between semantic diversity and thematic coherence.
  • LLM-as-a-Judge: Human evaluators found human- and LLM-written stories largely indistinguishable in quality, with significant advantages for LLM stories only in Authenticity and Elaboration.This contrasts with the LLM judge, whose scores differed significantly across all eleven dimensions.
  • LLM-as-a-Judge: The LLM judge assigned perfect scores to LLM-generated texts across several dimensions while rating human texts lower and more variably.The affected dimensions included Effectiveness, Elaboration, Fluency, Originality, Value, and Creativity, producing a pronounced ceiling effect.
  • LLM-as-a-Judge: Surprise has near-zero agreement between LLM and human judgments (τb = 0.01, p = 0.867), indicating misalignment in assessing narrative unpredictability.Four dimensions—Effectiveness, Novelty, Surprise, and Usefulness—have nonsignificant correlations.
  • LLM-as-a-Judge: LLM judgments show at most modest agreement with human evaluations, with Elaboration reaching the highest correlation (τb = 0.31, p < 0.001).Only Elaboration exceeds τb = 0.30 among the eleven dimensions.
  • Automatic metrics: Perplexity shows negative correlations with all eleven subjective dimensions, strongest for Effectiveness (ρ = −0.23) and Elaboration (ρ = −0.21).The authors interpret this as a slight tendency for humans to rate lower-perplexity texts more highly.

Discussion

The study finds substantial misalignment between human assessments of creative narratives and both automated metrics and LLM judges. These systems struggle to capture creativity’s subjective, multidimensional qualities and may favor machine-like stylistic regularity over human expression.

  • Core findings: The comprehensive evaluation reveals a substantial disconnect between quantitative metrics and human-perceived creativity.The study examines relationships between machine and human assessments across multiple dimensions.
  • LLM-as-a-Judge: LLM judges strongly preferred machine-generated narratives and their scores failed to correlate with human judgments regardless of writer origin.Their evaluations were consequently misaligned with human assessments.
  • Implications: Current quantitative frameworks remain fundamentally limited in capturing creativity’s subjective and multifaceted nature.The paper cautions against replacing human aesthetic judgment with automated metrics and recommends human-in-the-loop strategies.
  • Interpretation: Human creativity scores correlated with novelty, value, surprise, originality, and effectiveness, whereas LLM scores correlated primarily with surface-level syntactic properties such as elaboration.The mismatch became especially apparent for semantic properties including usefulness, novelty, and surprise.
  • Systemic risks: Using LLMs to generate training data, evaluate outputs, and fine-tune later generations without human supervision may create a self-reinforcing feedback loop.Such systems may increasingly favor algorithmic peers and homogenize writing while reducing diversity.
  • Automatic metrics: Automatically computed metrics showed near-zero correlations with all 11 subjective dimensions rated by human judges.Token-based metrics therefore miss semantic and emotional qualities valued by human readers.

Methods

The study compares human and machine-generated short stories using blinded human and LLM judgments alongside statistical measures. It balances 100 stories from each source, pairing them with WritingPrompts premises and generating machine stories from five models.

  • Experimental design: The pipeline triangulates statistical probability measures with human cognitive judgments to assess automated creativity evaluation.
  • Dataset: The dataset contains 200 stories: 100 human-generated and 100 machine-generated texts.Human stories were randomly sampled from WritingPrompts, while machine stories were generated from a separate set of prompts.
  • Dataset: Each story is paired with a unique WritingPrompts prompt serving as its creative premise.
  • Machine-generated stories: Five state-of-the-art models generated 20 stories each, yielding 100 machine-generated stories.The models were GPT-5.2, DeepSeek-V3.2, Mistral Large 3, Claude Sonnet 4.5, and Gemini 3 Pro.
  • Evaluation protocol: All evaluations were conducted blindly without identifiers distinguishing human-authored from AI-generated stories.Source information was retained for later analysis but withheld during evaluation to support fair comparisons.

Automatically Computed Evaluation Metrics

The study evaluates creativity through five automatic metrics spanning lexical, phrase-level, syntactic, and semantic properties. These include originality, predictability, syntactic repetition, and embedding-based diversity measures.

  • Metric set: Five automatic metrics assess creativity across granularities from phrase-level lexical choices to high-level narrative structures.The metrics are Creativity Index, Perplexity, Syntactic Templates, Expectation-Adjusted Distinct n-grams, and Sentence-BERT Embedding Diversity.
  • Lexical originality: Creativity Index measures how much text cannot be attributed to existing web content through L-uniqueness scores.Higher L-uniqueness indicates greater novelty at the corresponding n-gram scale.
  • Predictability: Perplexity quantifies language-model uncertainty and serves as a proxy for statistical conventionality.Lower values indicate more predictable text, whereas higher values suggest greater unpredictability.
  • Syntactic structure: Syntactic-template metrics quantify structural repetition using POS-tag compression, template coverage, and template density.CR-POS, Template Rate, and Templates-Per-Token characterize redundancy and repeated grammatical patterns.
  • Semantic diversity: Sentence-BERT Embedding Diversity measures semantic variation by computing the complement of average sentence similarity.

Subjective Metrics

Subjective creativity evaluation is decomposed into 11 dimensions rated on a five-point scale by humans and an LLM judge. The evaluations use blind protocols and isolated metric-specific inference calls to reduce cross-metric contamination.

  • Dimensions and rubric: The rubric covers authenticity, effectiveness, elaboration, fluency, flexibility, novelty, originality, surprise, usefulness, value, and overall creativity.Each dimension is rated from 1 = Poor to 5 = Excellent.
  • Evaluation protocol: Human and LLM-as-a-Judge evaluations were conducted blindly without authorship labels.
  • LLM judge: Llama 3.3 70B Instruct evaluated texts as an objective literary critic according to the 11 dimension definitions.The open-weight model was selected for a reproducible and transparent evaluation pipeline.
  • Isolated evaluation: Batched scoring produced high inter-metric correlations that obscured differences between fluency and originality.The study therefore used independent inference calls for each text-metric pair.
  • Human evaluation: Human evaluators included experts, students, and laypeople with varied AI experience to represent both technical and public perceptions.

Data Availability

The study’s generated and collected data are publicly available, including the research repository and the original WritingPrompts source dataset.

  • Public resources: All generated and collected data are publicly available in the study’s GitHub repository.
  • Source dataset: Original prompts and human-written short stories were taken from the publicly available WritingPrompts dataset.

Code Availability

The paper makes its complete source code and evaluation scripts publicly available to support reproducibility and further research.

  • The complete source code is publicly available.
  • Public release is intended to facilitate reproducibility and encourage further research.
  • The evaluation scripts are publicly available.

Appendix A Extended Correlation Matrices

Appendix A reports provenance-stratified correlation matrices comparing automatic metrics with subjective creativity dimensions for LLM-generated and human-authored stories. No machine-based metric shows a strong correlation with human judgments, although the patterns differ between provenance groups.

  • Overall comparison: No machine-based metric shows a strong correlation with human judgments across the provenance-stratified analyses.
  • LLM-generated stories: For LLM-generated stories, only SBERT-Div exceeds 0.2 in absolute correlation with human scores.
  • LLM-generated stories: SBERT-Div has a weak but consistent negative correlation with human scores for authenticity and effectiveness in LLM-generated stories.
  • Human-authored stories: For human-authored stories, Creativity Index and SBERT-Div show weak or no correlation with human scores, while TR shows no correlation.
  • Human-authored stories: CR-POS correlates positively with all dimensions for human-authored stories, peaking at elaboration and creativity.
  • Human-authored stories: Perplexity, TPT, and EAD show negative correlations with almost all dimensions in human-authored stories.

Appendix B Human Survey Statistics

Appendix B describes the human-evaluation sample and its distribution across the dataset. The evaluator pool combines varied LLM familiarity and backgrounds to provide a balanced mix of expert and general-user perspectives.

  • Evaluator background: The evaluator pool had heterogeneous LLM familiarity and included participants with varied educational and linguistic backgrounds.
  • Survey sample: 441 individual evaluations were collected from 115 unique participants.
  • Survey sample: Texts received an average of 2.21 independent human ratings each.
  • Evaluator background: Nearly 65% of evaluators held a university degree.
  • Evaluator background: The distribution was designed to reflect a balanced mix of expert and general-user perspectives.

Appendix C LLM Generation Prompt

Appendix C specifies the prompt and generation constraints used to produce the AI-authored stories. The setup standardizes format and length while introducing variation across generations and models.

  • Prompt design: Length variation was introduced to prevent nearly identical output lengths across generations.
  • Prompt design: The creative-writing task instructs models to write narrative text based on a supplied story premise.
  • Output constraints: Outputs were required to use Markdown formatting and standard spacing between paragraphs.
  • Output constraints: Stories were constrained to 200–500 words, with a simulated random length within that range.

Appendix D LLM-as-a-Judge System Instructions

The appendix specifies an isolated LLM-as-a-Judge evaluation procedure in which a fixed system prompt and dynamic metric prompt instruct the model to assess text objectively and return structured JSON. It also defines scoring guidance, ambiguity handling, and strict output-format requirements.

  • The evaluation uses a fixed system prompt together with a dynamic user prompt containing the metric definition and input.
  • The judge is instructed to evaluate each text objectively according only to the specified metric definition.
  • Each response must be a single valid JSON object containing a metric field with score, justification, and excerpt keys.
  • The template requests an excerpt of no more than 20 words supporting the evaluation and prohibits chain-of-thought or explanations outside the JSON.
  • If the text is ambiguous or too short to judge, the evaluator should assign 3 and note insufficient evidence.
  • Scores range from 5 for clear, strong evidence to 1 for no evidence or direct counter-evidence, with 3 indicating ambiguous or mixed evidence.
Loading 2608.23705v1…