Source-linked AI summary

On the Creativity of Large Language Models

Giorgio Franceschelli, Mirco Musolesi

arXiv:2304.00008v5cs.AIcs.CLcs.CY

TL;DR

The paper asks whether LLMs that produce impressive creative writing should be considered genuinely creative. It evaluates them through creativity theories, especially Boden’s value, novelty, and surprise criteria, and examines machine-creativity problems and societal impacts. It concludes that LLMs achieve value and weaker forms of novelty and surprise, but their autoregressive nature appears to prevent transformational creativity while creating substantial legal and societal implications.

  • Problem

    The paper examines whether LLMs that generate impressive creative writing can genuinely be considered creative.

  • Method

    The paper analyzes LLM creativity through classic theories, focusing on Boden’s value, novelty, and surprise criteria, and examines easy and hard machine-creativity problems and societal impacts.

  • Results

    LLMs can produce value and weak forms of novelty and surprise, but their autoregressive nature appears to prevent transformational creativity.

  • Takeaways & Limitations

    LLMs have considerable legal and societal implications, including difficult questions for creative professions and a legal framework not fully suited to generative AI.

  • Takeaways & Limitations

    Current LLMs are immutable after training and cannot directly adapt to changes in the domain without relying on in-context prompts.

Abstract

from arXiv · show

Large Language Models (LLMs) are revolutionizing several areas of Artificial Intelligence. One of the most remarkable applications is creative writing, e.g., poetry or storytelling: the generated outputs are often of astonishing quality. However, a natural question arises: can LLMs be really considered creative? In this article, we first analyze the development of LLMs under the lens of creativity theories, investigating the key open questions and challenges. In particular, we focus our discussion on the dimensions of value, novelty, and surprise as proposed by Margaret Boden in her work. Then, we consider different classic perspectives, namely product, process, press, and person. We discuss a set of ``easy'' and ``hard'' problems in machine creativity, presenting them in relation to LLMs. Finally, we examine the societal impact of these technologies with a particular focus on the creative industries, analyzing the opportunities offered, the challenges arising from them, and the potential associated risks, from both legal and ethical points of view.

1 Introduction

LLMs have become prominent tools for creative writing, producing remarkable poems and stories while raising the unresolved question of whether such systems are genuinely creative. The paper addresses this question through established creativity theories and positions its analysis as an early theoretical and philosophical investigation of LLM creativity.

  • LLMs are widely used for poetry and storytelling, and their outputs are often considered remarkable despite uncertainty about whether the systems are truly creative.The question is framed against Ada Lovelace’s account of machine originality.
  • The paper analyzes LLM creativity using Boden’s three criteria alongside cognitive-science and philosophical theories.Its stated focus is on the dimensions by which LLMs should be evaluated.
  • The paper presents itself as one of the first theoretical and philosophical investigations of LLM creativity.
  • The paper traces LLM development, examines creativity theories, and discusses practical implications for the arts, creative industries, design, and scientific and philosophical inquiry.It concludes by outlining open challenges and a research agenda.

2 A Creative Journey from Ada Lovelace to Foundation Models

Machine creativity developed from early rule-based systems for poems and stories through neural language models and transformers trained on increasingly large datasets. Modern LLMs can generate specialized creative content through in-context learning, although selecting effective prompts and demonstrations remains challenging.

  • Early machine-creativity systems generated poems, stories, paintings, and coherent characters using techniques including planning, case-based reasoning, and evolutionary strategies.
  • Neural language models enabled scalable text generation by learning probabilistic patterns in corpora and sampling new characters, words, or syllables.Examples include recurrent networks, LSTMs, and GRUs.
  • Transformers became central because earlier models scaled poorly to long sequences and often failed to capture the entire context.
  • Large modern language models use vast parameter counts, larger training corpora, and zero-shot or few-shot learning to produce specialized poems and stories from task descriptions and examples.
  • Finding suitable inputs and high-quality demonstrations remains challenging when using in-context learning for creative tasks.
  • RLHF-based ChatGPT and GPT-4 helped lead to a broader generation of related models, including Gemini, Llama, and Mixtral.

3 Large Language Models and Boden’s Three Criteria

Under Boden’s framework, LLMs can produce valuable artifacts and some forms of novelty and surprise, but their autoregressive training makes transformational creativity difficult. Their outputs may appear creative to observers even when they are not novel or surprising relative to the wider domain.

  • Boden’s three criteria: Boden defines creativity through three criteria: ideas or artifacts must be new, surprising, and valuable.
  • Novelty: LLMs may generate novel outputs because stochastic sampling and varied prompts produce texts absent from training data, but their imitation-based operation limits claims of computational novelty.
  • Surprise: Autoregressive LLMs are unlikely to generate highly surprising products because they follow existing data distributions, restricting them mainly to combinatorial or exploratory creativity.
  • Observer-relative creativity: LLM outputs can appear creative to observers through B-novelty and observer-relative surprise, even when they lack creativity under theory-relative standards.
  • Boden’s three criteria: LLMs produce valuable artifacts, while achieving psychological or historical novelty and genuine surprise is more difficult.The conclusion distinguishes value from stronger forms of novelty and surprise.
  • Boden’s three criteria: Transformational creativity likely requires alternative learning architectures because current probabilistic solutions are intrinsically limited in expressivity.

4 Easy and Hard Problems in Machine Creativity

The paper distinguishes easier technical problems from the harder problem of intentionality and self-awareness in machine creativity, arguing that current LLMs can simulate some creative capacities but lack core properties of a creative person and adaptive process.

  • Current LLMs may generate creative products, but producing such outputs does not make them intrinsically creative because creativity concerns both what is achieved and how it is achieved.The paper links intrinsic creativity to intentionality, flair, and the capacities underlying creative production.
  • Process: LLMs lack intrinsic motivation, intention to write, a self-feedback loop for validating outputs, and direct optimization for value, novelty, or surprise.
  • Easy and hard problems: The hard problem of machine creativity is intentionality and self-awareness: current LLMs are causal rather than intentional agents and cannot self-evaluate their own responses.Ranking or assigning quality scores can recognize limitations after generation, but does not establish creative self-awareness.
  • Press: Creative systems must account for press, because products are shaped by their social and historical environment rather than being explainable through product and process alone.
  • Press: Current LLMs are frozen after training and cannot adapt through repeated domain iterations, while in-context learning can simulate adaptation but longer contexts may degrade performance.
  • Person: The person perspective remains largely unavailable to LLMs: despite sometimes promising psychological-test performance, current systems are described as non-conscious and lacking an authentic self.

5 Practical Implications

LLMs create practical opportunities for creative work while raising unresolved legal, reliability, employment, and copyright concerns. Their near-term role is primarily collaborative, with humans retaining responsibility for prompting, curation, validation, and production.

  • Legal implications: LLM adoption raises copyright questions because current law does not clearly accommodate non-human authorship, while protection may depend on a human’s original contribution.The paper identifies prompting as one possible human contribution and discusses the user as a potential rights holder.
  • Legal implications: Training-data memorization can reproduce protected works, potentially infringing reproduction or adaptation rights.The paper suggests creative-oriented training could mitigate this risk and facilitate fair-use applications.
  • Societal risks: LLMs may displace some professional writing work, but unreliable information, inherited biases, and prompt dependence constrain their replacement of human writers.The quality of outputs can require substantial human skill and time for prompting and validation.
  • Opportunities: LLMs can transform creative roles by helping humans validate news, develop ideas, adapt texts to audiences, and generate research hypotheses.The paper characterizes their overall impact as positive while retaining humans in prompting, curation, and pre- and post-production.
  • Opportunities: LLMs can support human-AI co-creativity through targeted story generation, brainstorming, plot ideas, metaphors, and story plans.The paper also notes that these systems sometimes fail to perform such tasks at a human-like level.

6 Conclusion

The paper concludes that LLMs exhibit value and weak forms of novelty and surprise but do not yet achieve transformational creativity across product, process, and social dimensions. It also finds substantial legal and societal implications, alongside opportunities for human-AI cooperation and future technical development.

  • Conclusion: LLMs show value and weak novelty and surprise, but their autoregressive nature appears to prevent transformational creativity.The conclusion distinguishes this product-level assessment from broader requirements involving motivation, perception, self-awareness, and social adaptation.
  • Conclusion: Current legal frameworks are not fully suited to generative AI, while the effects on creative professions and the arts are difficult to forecast but considerable.The paper presents human-AI cooperation as an important opportunity within this uncertain impact.
  • Conclusion: Fine-tuning and continual learning may diversify outputs and support deployment across contexts, but these techniques would only simulate certain aspects of creativity.Whether such simulation constitutes non-human creativity remains a human judgment.
Loading 2304.00008v5…