Source-linked AI summary
CLIN: an Objective Framework for Evaluating Creativity in Short Persian Literary Text
Mohammad Reza Modarres, Armin Tourajmehr, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar
TL;DR
The paper addresses uncertainty about how reliably LLMs judge multidimensional creativity in short Persian literary text. It systematically evaluates LLM judges and introduces CLIN, which uses separate interpretable proxies for structured dimensions. Alignment varies by dimension, while CLIN matches or exceeds the strongest zero-shot LLM judge in the setting at lower cost.
Problem
Substantial disagreement remains about which signals reflect human creativity judgments and when automated evaluations can be trusted.
Method
The study evaluates LLM creativity judgments across multiple models, strategies, coordination mechanisms, and prompt formulations, then tests CLIN's separate proxies for three TTCT-derived dimensions.
Results
LLM–human alignment is stronger for Originality, Fluency, and Elaboration than for Emotion and Attractiveness, while CLIN achieves comparable or better alignment than the strongest zero-shot LLM judge at substantially lower cost.
Takeaways & Limitations
Structured creativity dimensions can be evaluated with simple, interpretable proxies that provide a lower-cost alternative to generative judging in this setting.
Takeaways & Limitations
The study is restricted to short Persian literary text, and its proxies approximate rather than fully capture subjective aspects such as emotional depth and aesthetic attractiveness.
Abstract
from arXiv · showhide
Evaluating creativity in large language model (LLM) outputs remains challenging because creativity is multidimensional and human-centered. We examine how reliably LLMs evaluate short literary text in Persian, a low-resource language, across multiple evaluation strategies and prompt formulations. We find that LLM-human agreement varies substantially across dimensions: alignment is stronger for structured TTCT-derived properties such as Originality, Fluency, and Elaboration, but considerably weaker for more subjective dimensions, particularly Emotion and Attractiveness. Judgments are also sensitive to prompt formulation, while few-shot prompting, ensembling, and multi-agent debate provide no consistent improvement. Motivated by this dimension-dependent behavior, we investigate whether structured creativity dimensions can instead be approximated using simple, interpretable proxy metrics. We introduce CLIN, which evaluates three TTCT-derived dimensions separately using topic-aware novelty for Originality, contextual lexical clustering for Fluency, and lexical diversity for Elaboration. These proxies achieve human alignment comparable to or better than the strongest zero-shot LLM judge in our setting while requiring substantially lower evaluation cost.
1 Introduction
This study examines the reliability of LLM creativity judgments for short Persian literary text and introduces CLIN as a lower-cost alternative for structured dimensions. It finds dimension-dependent human alignment and no consistent benefit from more elaborate judging procedures.
- Substantial disagreement remains over which signals reflect human creativity judgments and when automated evaluations can be trusted.
- LLM–human alignment is stronger for structured dimensions such as Originality, Fluency, and Elaboration than for subjective dimensions including Emotion and Attractiveness.
- More elaborate judging procedures provide no consistent improvement, and LLM-based evaluations remain sensitive to prompt formulation.
- CLIN separately evaluates Originality, Fluency, and Elaboration using novelty, contextual lexical clustering, and lexical diversity proxies.Its proxies achieve human alignment comparable to or better than the strongest zero-shot LLM judge at substantially lower evaluation cost.
- The study systematically evaluates LLM creativity judgments across models, evaluation strategies, coordination mechanisms, and prompt formulations in short Persian literary text.Settings include zero-shot evaluation, reference-based comparison, few-shot prompting, ensemble voting, multi-agent debate, and prompt variation.
2 Related Work
Prior work treats creativity as multidimensional and explores both LLM-based judges and automated metrics. CLIN differs by applying separate, transparent proxies to predefined human-rated dimensions of individual literary products.
- TTCT distinguishes Originality, Fluency, Flexibility, and Elaboration, supporting evaluation of creativity as complementary dimensions rather than a single property.
- Prior studies report variable LLM–human agreement, with reference-based evaluation and specialized evaluators improving alignment in some settings.
- Automated creativity measures include writing constraints, semantic associations, perplexity, corpus overlap, semantic entropy, and multi-agent methods.
- CLIN evaluates individual literary products along predefined human-rated dimensions using separate transparent proxies for Originality, Fluency, and Elaboration.The proxies use global and topic-relative novelty, contextual lexical clustering, and lexical diversity, respectively.
3 Creativity Evaluation Test
The paper combines structured TTCT-derived dimensions with selected subjective dimensions to evaluate creativity in short literary text. It distinguishes interpretable components from holistic human judgments rather than treating overall creativity as their sum.
- The study adopts Originality, Fluency, and Elaboration as three TTCT-derived dimensions for evaluating creativity.Flexibility is excluded because prior studies found it highly correlated with Fluency and redundant in practice.
- Originality measures novelty and non-cliché content, Fluency captures diversity of ideas conveying the topic, and Elaboration evaluates textual depth.
- Because short literary texts make some narrative dimensions difficult to assess reliably, the study incorporates Creativity and Attractiveness as compatible subjective dimensions.Creativity captures overall inventiveness, while Attractiveness reflects aesthetic appeal and reader engagement.
- Emotion is added to assess the intensity and vividness of affective content, motivated by concerns about modeling emotional depth.
- TTCT-derived dimensions decompose creativity into interpretable components, whereas the overall Creativity score represents an integrated holistic human judgment.The overall score is not defined as a simple aggregation of the other dimensions.
4 Evaluating LLM as Judge
The study evaluates LLM judgments of short Persian literary texts against human ratings across dimensions, text sources, and evaluation strategies. Agreement is stronger for structured dimensions than subjective ones, while alternative prompting and aggregation strategies provide no consistent improvement.
- Dataset: The dataset contains 200 short Persian literary texts, split evenly between human-authored and GPT-3.5-turbo-generated texts across five themes.Five human annotators rated each text on Originality, Fluency, Elaboration, Creativity, Attractiveness, and Emotion using a 3-point scale.
- Human-rated dimensions: Originality correlates most strongly with Creativity (ρ = 0.58) and Attractiveness (ρ = 0.56), while Creativity and Attractiveness correlate at ρ = 0.57.Fluency and Elaboration show a moderate association (ρ = 0.48), and Creativity is more weakly associated with Fluency, Elaboration, and Emotion.
- Single-judge evaluation: Claude 3.7 Sonnet is the strongest overall zero-shot judge across Originality, Fluency, and Elaboration, although the best model varies by dimension and text source.DeepSeek-V3 performs best on originality, LLaMA-4 on creativity, and Claude 3.7 Sonnet on fluency and elaboration for human-authored texts; Claude has the strongest average alignment for model-generated texts.
- Single-judge evaluation: LLM–human agreement is consistently weak for Attractiveness and Emotion, with most emotion correlations failing to reach statistical significance.Agreement decreases considerably for overall Creativity and Emotion on model-generated texts, while Attractiveness remains essentially unchanged.
- Alternative evaluation strategies: Reference-based evaluation yields less reliable results than single-judge evaluation because outputs are highly sensitive to the chosen reference.No single reference consistently improves performance across dimensions or instances.
- Alternative evaluation strategies: Few-shot prompting, ensemble voting, and multi-agent debate provide no consistent improvement over zero-shot or single-judge evaluation.Majority voting has comparable alignment at increased computational cost, while debate performs similarly to majority voting with negligible differences; few-shot prompting can reduce alignment.
- Prompt sensitivity: LLM evaluations differ substantially when questions are presented jointly or paraphrased, especially for Fluency and Elaboration.Even the relative ordering assigned by an evaluator can depend on prompt formulation despite unchanged evaluation criteria.
5 CLIN Framework
CLIN evaluates Originality, Fluency, and Elaboration separately with transparent lexical, semantic, and statistical proxies. These proxies significantly correlate with human ratings and match or exceed the strongest zero-shot LLM judge at lower cost.
- Framework: CLIN evaluates Originality through global and topic-relative novelty, Fluency through contextual lexical clustering, and Elaboration through lexical diversity.Each dimension is evaluated separately rather than collapsed into a single creativity score.
- Originality: Global originality measures rarity in language use, while local originality measures novelty within a constrained topic-specific distribution.The two components are combined to capture globally rare expressions and context-dependent novelty.
- Fluency: DBSCAN clusters contextual token embeddings to estimate semantically distinct lexical ideas, with Fluency defined as the number of non-noise clusters.DBSCAN allows the number of semantic groups to vary across texts and identifies isolated points as noise.
- Elaboration: Elaboration is approximated by counting unique normalized content-bearing tokens after stopword removal.This simplified unigram measure serves as a lexical-diversity proxy.
- Validation: All three proxies show positive, statistically significant correlations with their corresponding human-rated dimensions at p < 0.05.The correlations indicate that the proxies preserve the ranking structure induced by human evaluations.
- Validation: CLIN significantly outperforms Claude 3.7 Sonnet on elaboration, while differences on originality and fluency are not statistically significant.Thus, CLIN performs comparably on originality and fluency and better on elaboration against the strongest overall zero-shot judge.
6 Conclusion and Future Work
The paper finds that LLM–human agreement for Persian literary creativity is dimension dependent, with structured dimensions aligning better than subjective ones. CLIN provides a lower-cost alternative for the structured dimensions, while broader validation remains necessary.
- Conclusion: LLM–human agreement is substantially stronger for Originality, Fluency, and Elaboration than for Emotion and Attractiveness.The evaluation covered multiple models, prompting strategies, and coordination mechanisms.
- Conclusion: More elaborate judging procedures do not consistently improve performance, and judgments remain sensitive to prompt formulation.
- Conclusion: CLIN’s separate interpretable proxies achieve human alignment comparable to or better than the strongest zero-shot LLM judge at substantially lower evaluation cost.
- Future Work: The results support dimension-specific creativity evaluation using simple, transparent measurements for some human-rated properties.
- Limitations: The study is limited to short Persian literary text, and CLIN does not fully capture subjective aspects such as emotional depth and aesthetic attractiveness.Broader models, languages, longer texts, and other creative tasks require further evaluation.
A.1 Rubric Questions
The supplied passage points to Table A.1 for the complete list and descriptions of creativity dimensions.
- Rubric: Table A.1 lists all creativity dimensions and their descriptions.
A.2 Annotator Agreement
The paper aggregates annotator ratings and assesses their reliability with ICC(2,k), using k = 5 annotators.
- Annotator Agreement: Annotator scores are averaged to obtain a more stable estimate of overall human judgment.The task is inherently subjective, so individual annotators may differ in judgments and assigned scores.
- Annotator Agreement: Human inter-rater reliability is measured with the two-way random-effects, average-measure ICC(2,5).Here, k = 5 corresponds to the five annotators.
A.3 Zero-Shot Prompt Format
The zero-shot prompt asks evaluators to score Persian sentences numerically across six creativity-related dimensions, using explicit three-level rubrics. These dimensions operationalize originality, fluency, elaboration, emotion, attractiveness, and overall creativity through distinct textual criteria.
- Evaluators assign numeric scores only to Persian sentences in a zero-shot setting.The prompt requires scores of 1, 2, or 3 without explanations.
- Originality measures avoidance of clichés and unexpected word combinations, with higher scores indicating increasingly novel combinations.The rubric ranges from recognizable clichés to genuinely surprising predicate or attribute combinations.
- Fluency measures the number of distinct concepts or clauses integrated smoothly into one sentence.Scores increase from one simple idea to three or more coherent informational elements.
- Elaboration measures specific, sensory, or concrete details that ground abstract ideas.Higher scores require vivid and potentially multisensory detail rather than generic or abstract wording.
- Emotion measures whether the sentence evokes a specific discrete emotion, ranging from none to strong and unmistakable impact.The highest rubric level highlights visceral imagery, rhythm, or word choice that creates an emotional response.
- Attractiveness measures internal harmony among word choice, rhythm, imagery, and tone.Highly attractive sentences are described as harmonious, purposeful, and pleasing, potentially using prosody or fitting metaphors.
- Overall creativity is defined as an emergent sense that the sentence is novel, surprising, and valuable beyond the sum of its parts.The highest level requires interacting elements that produce meaningful, vivid, insightful, and memorable results.
A.4 Multi-Agent Debate Template
The appendix details evaluation procedures, perplexity- and corruption-based metrics, reference-based comparisons, and the paired bootstrap used to compare CLIN with Claude 3.7 Sonnet. Controlled shuffling produced monotonic quality degradation, while CLIN matched Claude on originality and fluency and significantly outperformed it on elaboration.
- A.4 Multi-Agent Debate Template: The multi-agent debate begins with independent sentence evaluations, then exposes prior responses and asks models to revise or confirm their scores.Final disagreements about updating are resolved by majority vote.
- A.5 Normalized Perplexity: Perplexity is computed from per-token negative log-likelihood under a pretrained language model and represents sequence predictability.The probability assigned to each token is conditioned on its preceding context.
- A.5 Normalized Perplexity: Raw perplexity is clipped between predefined bounds and normalized to [0, 1].The experiments use PPLmin = 1.0 and PPLmax = 1000.0; lower scores correspond to conventional sentences and higher scores to less likely ones.
- A.6 Quality Metric under Controlled Corruption: The quality metric compares a sentence’s perplexity with the average perplexity of minimally corrupted variants created by token deletion and local shuffling.A larger gap indicates better-formed text, whereas incoherent or degenerate sequences yield smaller gaps.
- A.6 Quality Metric under Controlled Corruption: A controlled experiment randomly shuffles 0.0 to 0.5 of each sentence’s tokens while preserving the underlying token distribution.Quality scores are averaged at each corruption level across the dataset.
- A.6 Quality Metric under Controlled Corruption: Quality scores decrease monotonically as token perturbation increases, supporting the metric’s sensitivity to disruptions in coherence and well-formedness.This behavior is reported in Table A.2 as evidence for its use as a textual-quality proxy.
- A.7 Reference-based Evaluation: Experimental Details and Results: Reference-based evaluation compares candidate and same-topic reference sentences pairwise, converts human annotations into preferences, and correlates model decisions with those preferences.References are randomly sampled, self-comparisons are excluded, and results are averaged over three runs and both human-authored and model-generated texts.
- A.8 Paired Bootstrap Comparison with the Best Zero-shot LLM Judge: Paired item-level bootstrap compares CLIN and Claude 3.7 Sonnet through differences in Spearman correlation with human judgments.The difference is defined as ∆ρ = ρCLIN − ρClaude, with positive values favoring CLIN; CLIN significantly outperforms Claude on elaboration, while originality and fluency differences are not significant.