Source-linked AI summary

A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing

Carlos Gómez-Rodríguez, Paul Williams

arXiv:2310.08433v1cs.CLcs.CY

TL;DR

Creative-writing evaluation lacks broad evidence comparing current LLMs as standalone systems, especially under human assessment. The paper compares 12 instruction-aligned models with human writers on a zero-shot, open-ended combat prompt designed to reduce training-data reuse. State-of-the-art commercial models perform at a very competent level, while humans retain an advantage in originality and models divide sharply on humor.

  • Problem

    A direct evaluation comparing current LLMs as standalone creative-writing systems was lacking, while creative writing remained absent from major benchmark suites because human assessment is costly.

  • Method

    The study uses human evaluation of stories from 12 instruction-aligned LLMs and human writers, generated zero-shot from a purpose-designed prompt intended to limit training-data reuse.

  • Results

    State-of-the-art commercial LLMs perform at a very competent level and outperform the sampled human writers in most rubric categories, while humans lead in originality and models split sharply on humor.

  • Takeaways & Limitations

    The results suggest that leading commercial models are not distinguishably worse than reasonably trained humans on this creative-writing task, although performance varies by creative criterion.

  • Takeaways & Limitations

    Creative-writing ratings are subjective, the volunteer raters assessed only subsets of the 65 stories, and the five human writers are not representative of human creative-writing ability as a whole.

Abstract

from arXiv · show

We evaluate a range of recent LLMs on English creative writing, a challenging and complex task that requires imagination, coherence, and style. We use a difficult, open-ended scenario chosen to avoid training data reuse: an epic narration of a single combat between Ignatius J. Reilly, the protagonist of the Pulitzer Prize-winning novel A Confederacy of Dunces (1980), and a pterodactyl, a prehistoric flying reptile. We ask several LLMs and humans to write such a story and conduct a human evalution involving various criteria such as fluency, coherence, originality, humor, and style. Our results show that some state-of-the-art commercial LLMs match or slightly outperform our writers in most dimensions; whereas open-source LLMs lag behind. Humans retain an edge in creativity, while humor shows a binary divide between LLMs that can handle it comparably to humans and those that fail at it. We discuss the implications and limitations of our study and suggest directions for future research.

1 Introduction

LLMs have advanced across language tasks and can produce high-quality creative writing, but a direct standalone comparison of current models on creative writing was missing. The paper addresses this gap with a zero-shot, open-ended combat scenario designed to reduce training-data reuse.

  • Creative writing requires imagination, coherence, fluency, originality, and other complex skills that make evaluation challenging.
  • A direct evaluation comparing current LLMs as standalone creative-writing systems was missing despite substantial experimentation with LLM-generated stories.
  • The study compares 12 recent instruction-aligned LLMs with human writers using a task-adapted creative-writing rubric.
  • The zero-shot prompt asks for an epic narration of combat between Ignatius J. Reilly and a pterodactyl, with the scenario designed to limit reuse of training material.

2 Related work

Creative-writing evaluation remains underrepresented in LLM benchmarks because it demands broad literary and linguistic abilities and is costly to assess with human judgment. Prior studies typically examined one model, used task-specific prompting, or relied on existing stories that may overlap with training data.

  • LLMs in creative writing: Earlier story-generation models suffered from long-range incoherence and were used with specialized fine-tuning or external knowledge and planning systems.
  • Evaluation gap: Major benchmark suites lacked creative-writing tasks, reflecting the cost of human evaluation compared with easily automated metrics.
  • Prior comparisons: The closest prior study used prompt-based learning, evaluated one LLM, and drew on pre-existing story datasets rather than comparing multiple models in a zero-shot setting.
  • Prior comparisons: Other recent work likewise focused on a single model or limited distinguishability tests, while a concurrent study evaluated only three LLMs despite using human evaluation in a zero-shot setting.
  • Creative writing evaluation: Creative writing combines craft expertise, cultural and literary competence, fluency, coherence, metaphorical understanding, innovation, originality, and imagination.

3 Materials and Methods

The study evaluates instruction-following LLMs on a fresh, open-ended creative-writing prompt and scores their stories with a holistic human rubric. The design emphasizes model comparison, creative latitude, and reduced risk of training-data reuse.

  • 3.1 Task: The task requires an epic single combat between Ignatius J. Reilly and a pterodactyl, written in John Kennedy Toole’s style, from a fresh prompt without previous context.
  • 3.1 Task: The scenario is deliberately unconventional and open-ended, leaving setting, weapons, outcome, and other story elements unspecified to preserve creative latitude.
  • 3.1 Task: The task challenges models to capture a scarce literary character and authorial style while minimizing opportunities to reuse existing stories.
  • 3.2 Models: The model corpus contains 12 instruction-following or conversational systems available by April 20, 2023, using the largest practical version of each distinct model.
  • 3.3 Evaluation rubric: The evaluation rubric measures creative writing holistically across craft-based criteria, with scores from 1-10 grouped as Emerging, Competent, or Sophisticated.

4 Results

The evaluation finds that leading commercial LLMs can match or exceed humans on many creative-writing dimensions, while humans retain advantages in originality and humor. Performance varies sharply by rubric category, with humor especially divisive and epic narration favoring several LLMs.

  • Agreement: Inter-rater reliability was moderate: linearly weighted Cohen’s kappa was 0.48 with a 95% CI of [0.43, 0.54], and overall-score correlation was 0.58.The authors interpret these values as reasonable consistency given the subjectivity of creative-writing evaluation.
  • General overview: GPT-4 produced the highest overall-rated stories and led 8 of 10 individual rubric categories, while humans led originality and Claude led dark humor.GPT-4 also showed unusually low standard deviations relative to humans and other LLMs.
  • General overview: Commercial models generally outperformed open-source systems: Koala scored 60.0 overall versus 80.2 for GPT-4.The strongest LLMs were generally better across categories, although some models showed category-specific strengths.
  • Humor: Humor was the hardest rubric item, averaging 3.4, while structural elements were easiest at 7.3 across all stories.Incorporating John Kennedy Toole’s style was the second-hardest category, averaging 4.7.
  • Humor: Humor ratings formed a binary divide: Claude, Bing, GPT-4, and humans scored between 6 and 6.5, whereas the remaining models scored 3.4 or less.A significance test confirmed that the lower-performing group was worse than both humans and the higher-performing models; the authors suggest humor may emerge in larger LLMs.
  • Creativity: Humans outperformed all LLMs in creativity and originality, although the three strongest humor models were not significantly less original than humans.The authors conclude that LLMs can produce creative stories, while humans retain an edge in creativity.
  • Epicness: Both ChatGPT versions significantly outperformed humans in epicness, while six additional models had higher average ratings without significant differences.OpenAssistant and GPT4All also outperformed humans and Bing in epic narration despite ranking in the overall-score bottom half.

5 Discussion

The study presents a broad zero-shot comparison of 12 instruction-aligned LLMs and human writers using a human-evaluated, task-adapted creative-writing rubric. Its results differ from related work partly because the studies use different human-story sources, writing constraints, lengths, prompt scopes, and protections against training-data reuse.

  • Study scope and design: The evaluation compares 12 instruction-aligned LLMs with human writers in a zero-shot creative-writing task using a 10-item rubric adapted to the scenario.The prompt concerns combat between Ignatius J. Reilly and a pterodactyl, and the study is designed to avoid training-data reuse.
  • Methodological differences: The comparison differs from related work because humans and LLMs receive the same prompt, whereas another study uses unconstrained published human stories and LLM adaptations of their plots.The related study uses New Yorker stories by highly successful authors, while this study uses Creative Writing students.
  • Methodological differences: A single prompt enables a rubric tailored to humor and Toole style, multiple alternative stories per LLM, within-model distributions, and statistical testing.The narrower setting trades breadth of story prompts for more detailed analysis.
  • Methodological differences: The study prevents training-data reuse, unlike evaluations based on existing stories published online that may appear in models’ training data.This distinction is presented as a central methodological difference between the studies.
  • Interpreting divergent results: Another study reports LLM stories clearly behind human-authored stories, which the authors hypothesize results mainly from its higher-bar comparison and plot-adaptation setup.The authors note that genre, target length, and other methodological factors may also benefit humans or LLMs in nonobvious ways.

6 Conclusion

The study finds that leading commercial LLMs perform creative writing at a competent level, often matching or exceeding the evaluated humans across rubric categories. Humans remain strongest in originality, while commercial models lead overall and humor separates models that achieve human-like ratings from those that fail.

  • Main findings: GPT-4 and Claude achieve high scores and outperform the evaluated human writers in most rubric categories.The authors caution that the sample is too small for categorical claims of superhuman storytelling and that five human writers may not represent human ability generally.
  • Main findings: Commercial LLMs achieve the best results, while open-source models clearly lag behind in the study.This conclusion is reported as the main model-family pattern.
  • Dimension-level findings: Humans retain the lead in originality, whereas LLMs tend to excel in technical aspects such as readability and structure.The conclusion summarizes complementary strengths rather than a uniform advantage for either group.
  • Dimension-level findings: Humor is especially difficult: most LLMs fail, but the best three models achieve human-like ratings.The study describes a binary divide between models that handle humor and those that do not.
  • Future work: Future work should examine other literary genres, non-English languages, and whether prompt engineering or fine-tuning improves generated-story quality.These are identified as directions for extending the study.

Limitations

The study’s comparisons are constrained by commercial-model reproducibility, subjective and limited rating data, nonrepresentative human writers, and a narrow English-language genre scope.

  • Commercial LLMs and reproducibility: Commercial models may be difficult to reproduce because their products can change without notice and prior versions are unavailable.The authors report versions and access dates where possible and publish generated outputs, but future prompting and generation may not remain reproducible.
  • Limitations of the analysis: Creative-writing ratings are subjective, and volunteer raters assessed only subsets of the 65 stories, limiting the sample size.The authors provide sample sizes, standard deviations, and inter-rater agreement to help assess variability.
  • Limitations of the analysis: The human-writer sample is not representative of human creative-writing ability, so the study cannot establish whether LLMs are better, equal, or worse than humans at creative writing generally.The evaluation is focused on a specific genre and uses human writers only as a reference point.
  • Scope: Results from this English-language, specific-genre evaluation do not necessarily generalize to other genres or languages.The authors fixed these variables to conduct a detailed evaluation of many LLMs within available resources.

Ethics Statement

The evaluation used volunteer participants, fixed generation settings, and a ten-item rubric covering general creative-writing abilities and task-specific proficiency. Ratings included cohesion, narrative and structural control, plot logic, creativity, style, genre, combat description, character accuracy, and dark humor.

  • Participants: All raters and writers were volunteers, and the study kept the time demand correspondingly low.The evaluation therefore relied on participants who opted into the writing or rating tasks.
  • Contamination control: The study used model access dates and generation dates to assess possible training-set contamination, with pre-2023-10-09 cutoffs likely posing minimal risk.The paper was first publicly disclosed online on 2023-10-09, while human authors, raters, and reviewers had earlier access.
  • Generation settings: Commercial models were run through their presented web interfaces, while open-source models used the default LMSYS web-interface parameters, including temperature 0.7.The study did not tune model hyperparameters; Bing Chat was run in Creative mode.
  • Rubric: The rubric awards 10 points on each of 10 criteria, totaling 100 points, with criteria 1–5 measuring general creative-writing capacities and criteria 6–10 measuring task-specific proficiency.The criteria are designed to assess writing craft while avoiding formulaic, rule-based writing.

E Sample stories

The paper presents sample stories selected by rating, alongside rubric box plots covering cohesion, narrative and structural elements, plot logic, creativity, style, genre, combat, character accuracy, and dark humor. Because different stories were assigned to different raters, selecting stories by rating is necessarily noisy.

  • Sample selection: The sample includes the three top-rated stories, the best human-written story, the median-ranked story, and the worst-rated story.The best human-written story ranked fourth overall.
  • Task-specific criteria: The figures also compare epic genre, combat description, character accuracy, and dark humor across the human and LLM-generated stories.These are task-specific criteria in the evaluation rubric.
  • Selection caveat: Rating-based story selection is noisy because different stories were assigned to different raters.The methodology was designed to compare models fairly, not to support precise comparisons between individual stories.

This story was generated by GPT-4. The ratings

The paper includes a GPT-4-generated story depicting Ignatius J. Reilly’s combat with a pterodactyl, followed by ratings identified as belonging to the corpus’s best overall-rated story.

  • Opening: The story opens with Ignatius J. Reilly confronting a pterodactyl that emerges from a portal in New Orleans.The narration frames the encounter as an anachronistic conflict between a twentieth-century man and a prehistoric creature.
  • Confrontation: Ignatius responds to the creature with sarcastic rhetoric and declares that he will defeat it.His dialogue presents the battle as another challenge to his supposedly indomitable will.
  • Combat: Ignatius stuns the pterodactyl with a shopping cart and avoids its subsequent talon attack.The combat narration emphasizes improbable physical agility and improvised weapons.
  • Victory: Ignatius ultimately kills the pterodactyl with an umbrella and claims victory before the assembled crowd.He presents the encounter as a legendary triumph while retaining the story’s comic, self-aggrandizing voice.
  • Conclusion: The story closes by casting the battle as a lasting legend and having Ignatius leave with a discarded hot dog.Its concluding image combines epic framing with an incongruous detail about his sustenance.
  • Ratings: Table 4 reports the ratings for the corpus’s best overall-rated story, produced by ChatGPT with GPT-4.The table identifies the rated story as the GPT-4 output described in this section.

This story was generated by Bing Chat. The ratings

Ignatius encounters a pterodactyl in Audubon Park and attempts to defeat it with indignation, rhetoric, and improvised weapons before being carried away.

  • Ignatius encounters a pterodactyl that has escaped from a natural-history museum while walking through Audubon Park.
  • He first tries to hide beneath his hunting cap, then confronts the creature as a refined scholar defending civilized society.
  • Ignatius attacks the pterodactyl with a rolled newspaper, stabbing it in the eye and briefly creating an opportunity to escape.
  • The wounded pterodactyl recovers, catches Ignatius by his coat tails, and lifts him from the ground.
  • Ignatius protests loudly, but no one rescues him before the pterodactyl carries him to its skyscraper nest.

This story was generated by Claude. The ratings

A pterodactyl descends over New Orleans and confronts Ignatius, whose umbrella blows and insults drive the creature away.

  • A pterodactyl descends through a roiling gray sky, casting a shadow over New Orleans streets.
  • Ignatius walks beneath it absorbed in his Valencia and fantasies, while the creature reacts indignantly to his presence.
  • He dismisses the creature as prehistoric nonsense and threatens it with his umbrella.
  • Ignatius bats the pterodactyl into a lamppost, then strikes its head and neck while punctuating each blow with insults.
  • After being thrashed, the pterodactyl flees, allowing Ignatius to resume his walk through New Orleans.

E.4 Best-rated human story (and tied for fourth overall best-rated story)

The human story combines domestic absurdity, wordplay, correspondence, and a surreal basement duel between Ignatius and Terry-dactyl.

  • Opening domestic scene: The narrative begins with Ignatius’s failed “Nomad Grey” joke and his mother Irene’s attempts to understand and humor him.
  • Challenge and correspondence: Terry-dactyl challenges Ignatius to a duel, while Ignatius accepts only after insulting the challenger and refusing to remove his hat.
  • Basement duel: Ignatius relocates the contest to his basement and exchanges escalating puns with Terry, including jokes about hands, flapping, and points.
  • Basement duel: The duel becomes physical when Terry cuts off Ignatius’s arm and Ignatius responds with paint and bodily gases that incapacitate both combatants.
  • Aftermath: After the explosion, Ignatius records the incident as wit, while Irene finally understands the “Nomad Grey” joke.
  • Park encounter: In another combat sequence, Ignatius uses a valve, a hot dog, and medievalist rhetoric while deciding whether to flee or fight the attacking pterodactyl.
Loading 2310.08433v1…