Source-linked AI summary
The Creativity of Text-to-Image Generation
Jonas Oppenlaender
TL;DR
The paper asks whether text-to-image generation is creative when anyone can produce high-quality digital art through prompts, and whether product-centered evaluation captures the human creativity involved. Using an online-ecosystem perspective grounded in Rhodes’ four P model, it argues that creativity lies in practitioners’ interactions, creative practices, and communities, while identifying information asymmetries that complicate evaluation.
Problem
The paper addresses whether text-to-image art is creative and argues that product-centered evaluation may not capture the human creativity involved.
Method
The paper combines an online ethnography with analysis of prompt engineering, curation, communities, and the four perspectives of Rhodes’ creativity model.
Results
The paper argues that human creativity in text-to-image art lies in interaction with systems, iterative prompting, curation, and online communities rather than solely in generated images.
Takeaways & Limitations
Evaluating text-to-image creativity requires considering the creator’s process and environment alongside the digital image.
Takeaways & Limitations
Creativity assessment is complicated by information asymmetries about the systems, prompts, and processes used to produce opaque text-to-image artworks.
Abstract
from arXiv · showhide
Text-guided synthesis of images has made a giant leap towards becoming a mainstream phenomenon. With text-to-image generation systems, anybody can create digital images and artworks. This provokes the question of whether text-to-image generation is creative. This paper expounds on the nature of human creativity involved in text-to-image art (so-called "AI art") with a specific focus on the practice of prompt engineering. The paper argues that the current product-centered view of creativity falls short in the context of text-to-image generation. A case exemplifying this shortcoming is provided and the importance of online communities for the creative ecosystem of text-to-image art is highlighted. The paper provides a high-level summary of this online ecosystem drawing on Rhodes' conceptual four P model of creativity. Challenges for evaluating the creativity of text-to-image generation and opportunities for research on text-to-image generation in the field of Human-Computer Interaction (HCI) are discussed.
1 INTRODUCTION
Text-to-image systems make high-quality digital art accessible through natural-language prompts, raising questions about creativity. The paper argues that creativity involves practitioners’ interactions, practices, and communities, not only generated products.
- Text-to-image systems can synthesize aesthetically high-quality images from natural-language prompts, while an online ecosystem of tools and resources has emerged.
- Prompt engineering lets practitioners create images with little or no understanding of the underlying technologies.The practice is also called prompt programming, prompt design, or prompting.
- The product-centered definition of creativity evaluates observable artifacts as original and effective, but may miss creativity in text-to-image practices.
- Rhodes’ four P framework examines creativity through product, person, process, and press, with the paper emphasizing the latter three perspectives.
- Prompt engineering, image-level curation, portfolio-level curation, and community-driven resources are presented as creative practices shaping text-to-image art.
2 RELATED WORK
Related work presents creativity as a complex, socially situated concept commonly assessed through products that are original and effective. Rhodes’ four P model broadens this view by including person, process, environment, and product.
- Text-to-image generation became more accessible through CLIP-related progress, open-source systems, and free or easy-to-use Colab notebooks.
- Creativity is difficult to define, and HCI uses multiple definitions that contribute to conceptual fuzziness in the literature.
- Csikszentmihalyi’s systems model situates creativity among the individual, a domain, and a field that provides social confirmation.
- Small-c creativity concerns personally valuable or meaningful creations, whereas Big-C creativity requires a major contribution to society.
- Rhodes’ four P model defines person, process, press, and product as complementary perspectives on creativity and its resulting artifacts.
- The standard product-centered definition requires an artifact to be both novel and effective, appropriate, useful, or valuable.
3 METHOD
The paper uses an online ethnography and the researcher’s experience with text-to-image systems to examine the emerging creative ecosystem. It draws on Rhodes’ model to develop a broad overview of practitioners, practices, resources, and communities.
- The study is grounded in an online ethnography of the text-to-image art community conducted between October 2021 and March 2022.
- The researcher experimented with VQGAN–CLIP, CLIP-guided diffusion, and other systems available through Google Colab.
- The researcher learned prompt writing through Twitter and Midjourney, where members shared images alongside their prompts.Midjourney’s chat-based community made the relationship between prompts and generated images visible.
- The research uses Rhodes’ multiperspective model to develop a broad overview of the creative ecosystem surrounding text-to-image synthesis and text-based generative art.
4 FEEDING RANDOM SNIPPETS OF TEXT TO THE MACHINE
High-quality images can result from random text or verbatim lyrics without substantial creative input from the practitioner. This demonstrates why the generated image alone is an imperfect proxy for human creativity.
- The product-centered definition is presented as insufficient for fully assessing human creativity in text-to-image artwork.
- State-of-the-art systems can potentially turn arbitrary textual inputs into high-fidelity images.
- Two illustrative scenarios use encyclopedia text or verbatim music lyrics as inputs to a text-to-image system.
- When lyrics are merely transmitted after being correctly understood, the practitioner’s role may involve interpretation but not necessarily creative transformation.
- Little or no creativity may be involved beyond basic literacy when detailed textual inputs produce sophisticated images, paralleling the Chinese Room thought experiment.
- Practitioners can produce surprising, interesting, and high-quality artworks from lyrics, random text, single characters, words, quotes, and emojis.
- Generated images may be high-quality even when the interaction is skill-free apart from basic literacy, making the image an imperfect proxy for the person’s creativity.
5 THE HUMAN CREATIVITY OF TEXT-TO-IMAGE GENERATION
Text-to-image creativity arises through human interaction with generative systems and the social practice of prompt engineering. Online communities help users find inspiration and learn this craft.
- Text-to-image creativity arises from text-based interaction between human users and generative systems.
- Online communities shape how users find inspiration and learn prompt engineering.
5.1 Product: The Digital Image
Text-to-image systems can produce aesthetically high-quality images, but the resulting image is an imperfect measure of human creativity. The paper therefore shifts attention from product to the wider creative process.
- Text-to-image systems can synthesize images of high aesthetic quality, especially with prompt modifiers.
- From a product-based perspective, generated images appear creative, but they may result from computational rather than human creativity.
- Because current artworks remain imperfect in depicting anatomy and faces, the image product is an imperfect measure of human creativity.
- The paper proposes assessing person, process, and press alongside the artifact in Rhodes’ model.
5.2 Person: The Practitioner
The practitioner’s creativity depends on learned prompt-engineering skills and knowledge of the system’s training data, latent space, modifiers, and configuration choices.
- Practitioners vary widely, with some producing innovative prompts while others struggle with long, specific prompts.
- Some practitioners produce beautiful images with relatively minimalistic prompts.
- Prompt engineering combines knowledge of training data and latent space with experience using prompt modifiers.
- Producing high-fidelity images also requires choices about aspect ratio, training data, and configuration parameters.
- Prompt engineering is learned because effective keywords and modifiers are not immediately apparent.
5.3 Process: Iterative Prompt Engineering and Image Curation
Text-to-image creativity extends beyond generation into iterative exploration and curation. Practitioners develop images through successive prompts, select among outputs, and assemble portfolios of their strongest works.
- 5.3 Process: Iterative Prompt Engineering and Image Curation: The creative process may span multiple AI systems and applications, from initial generation through enhancement and graphics editing.
- 5.3 Process: Iterative Prompt Engineering and Image Curation: Iteration and curation are the two creative-process components treated as native to text-to-image generation.
- 5.3 Process: Iterative Prompt Engineering and Image Curation: Practitioners probe the model’s latent space with prompts, linking ideas across successive iterations.
- 5.3 Process: Iterative Prompt Engineering and Image Curation: Several iterations may be needed to reach a subjectively satisfactory result, with chance and unanticipated directions affecting the process.
- 5.3.1 Iterative prompt engineering.: Images emerge progressively because each generation step can use the previous output as its input.
- 5.3.1 Iterative prompt engineering.: Practitioners may abort development when results become unsatisfactory, as GAN-based systems can degenerate to sub-optimal solutions.
- 5.3.1 Iterative prompt engineering.: Image-level curation involves selecting one image from generated sets, and the best image need not be the final generation.
- 5.3.2 Image-level and portfolio-level curation.: Portfolio curation involves retaining the strongest works and discarding or withholding images that do not meet the practitioner’s standard.
5.4 Press: The Emerging Text-to-Image Ecosystem
Text-to-image art is supported by an online ecosystem of communities, tools, resources, and practitioners. These communities help practitioners learn prompt engineering and develop creative practices such as image-level and portfolio-level curation.
- Communities and ecosystem: Online communities, social platforms, dedicated services, tools, and resources form an ecosystem around text-to-image art.The ecosystem includes general platforms such as Twitter and Reddit, dedicated communities, and web-based services.
- Community roles: Community roles include innovators, porters, conservators, service and resource providers, and practitioners.These roles are fluid; for example, innovators may also make technologies available through Colab notebooks.
- Learning and practice: Online communities help practitioners learn prompt engineering through experimentation, shared prompts, and engagement with prior work.This support can help novices overcome the learning curve associated with prompt engineering.
- Creative practices: Image-level curation involves selecting an intermediate generated image as the most interesting and artistic result in a batch.In the Figure 4 example, the practitioner aborted generation after 175 steps and selected step 50.
- Learning and practice: The distributed, sequential presentation of community messages can make it difficult to pursue specific learning goals.Dedicated resources increasingly support learning goals such as understanding prompt modifiers.
- Communities and ecosystem: The combination of people, technology, services, tools, and resources makes text-to-image generation more accessible and supports the community’s growth.The ecosystem draws in non-technically minded practitioners and is described as healthy for the community to thrive.
6 DISCUSSION
The paper argues that product-centered creativity assessments miss human creativity in text-to-image art because the image alone reveals little about the system, prompts, and process. It therefore calls for assessment informed by all three aspects.
- Creativity and evaluation: Product-centered creativity assessment is insufficient because text-to-image art involves human creativity throughout the creative process.The paper contrasts evaluating only the artifact with examining the wider process.
- Challenges for evaluation: Text-to-image art is opaque, creating information asymmetries between creators and viewers about the system, prompts, and process.These three aspects are identified as challenges for evaluating creativity.
- Challenges for evaluation: The image alone reveals little about the generative system or its configuration parameters, which can distinguish purposeful mastery from default use.Some systems have dozens of individually adjustable parameters.
- Challenges for evaluation: Unshared prompts and omitted prompt modifiers can prevent viewers from understanding how an image was generated.Prompt modifiers determine aspects including the style and quality of the generated image.
- Challenges for evaluation: Because systems may accept textual prompts, image prompts, or both, the inputs used to generate an artwork are often difficult to determine.Images can guide scene composition, distort existing images, or provide targets for optimization.
- Challenges for evaluation: The image alone cannot distinguish a near skill-free or copied generation from an arduous iterative process of prompt engineering.The paper describes prompt engineering as a focused, iterative interaction with the AI system.
- Creativity and evaluation: Assessing the full extent of human creativity in text-to-image art requires information about the system, prompts, and process.The paper treats all three aspects as necessary for fuller assessment.
6.2 Opportunities for Future Research on Text-to-Image Generation
The paper identifies research opportunities spanning technology, humans, and community, including improved language understanding, co-creative interfaces, and study of the text-to-image art community.
- Technology: Advances in natural language understanding could improve how text-to-image systems interpret practitioners’ prompts.The paper notes that current textual understanding is imperfect and learned concepts may not correspond to concepts in the visual world.
- Technology: Prompt engineering for text-to-image synthesis can draw on practices and resources developed for large-scale language models such as GPT-3.The paper specifically mentions input templates and emerging community resources as relevant precedents.
- Technology: Web-based interfaces for text-to-image systems offer opportunities to design novel co-creative systems and creativity support tools.The paper points to dedicated interfaces such as MindsEye as examples of efforts to better support practitioners.
- Community: The text-to-image art community provides a setting for studying a creative practice that is emerging among practitioners with varied technical expertise.Practitioners include novices, amateurs, semi-professionals, skilled artists, and professionals.
- Humans: Generative AI may democratize art and creative production while changing the role of the creator and society’s relationship with images and creative work.The paper frames this possibility as analogous to photography’s expansion of creative opportunity, while preserving its speculative qualification.
- Humans: Interaction with AI through natural language may shape future digital society, including how people use language, communicate, and work.The paper describes this as a possible bidirectional effect of prompt engineering on society.
7 CONCLUSION
The conclusion argues that evaluating creativity in text-to-image art requires attention beyond the digital image to the creator’s process, environment, interactions, and communities.
- Conclusion: Product-centered creativity measures do not fully assess the human creativity involved in text-to-image-generated art.The paper supports this conclusion with a scenario where the product-centered definition fails to capture the full extent of human creativity.
- Conclusion: Rhodes’ four P model provides a framework for examining creativity across the creator’s process and environment as well as the outcome.The paper uses the model to expand its account of human creativity in text-to-image generation.
- Conclusion: Image-level and portfolio-level curation are practices through which text-to-image creators express creativity.The conclusion treats curation as part of the creative process rather than only evaluating generated images.
- Conclusion: Human creativity in text-to-image art lies in interactions with generation systems and in online communities that help practitioners learn prompt engineering.The paper identifies the growing practitioner community and its craft as an area for HCI research.