Source-linked AI summary
What is in a Text-to-Image Prompt: The Potential of Stable Diffusion in Visual Arts Education
Nassim Dehouche, Kullathida Dehouche
TL;DR
The paper asks how Text-to-Image AI can be understood and used in visual arts education while addressing concerns about artistic ownership. Using 72,980 Stable Diffusion prompts, it formalizes prompt structure and connects it to teaching art history, aesthetics, and technique. It concludes that the systems can be valuable didactic tools with educator guidance, but their use raises unresolved intellectual-property and economic concerns.
Problem
The paper addresses the limited attention given to using generative AI in visual arts education and the related concern that these systems may reproduce protected artistic work.
Method
The authors analyze 72,980 Stable Diffusion prompts to formalize their structure and propose a procedural classification for educational use.
Results
The analysis identifies recurring prompt categories and supports a structured procedural framework connecting prompt elements to art-history, aesthetics, and technique instruction.
Takeaways & Limitations
With educator guidance and curation, Stable Diffusion can serve as a didactic tool for transmitting technical concepts and exploring artistic genres, movements, and aesthetics.
Takeaways & Limitations
Prompt outputs remain partly random and highly sensitive to wording, while coherent results depend on the interplay among prompt elements and randomness.
Abstract
from arXiv · showhide
Text-to-Image artificial intelligence (AI) recently saw a major breakthrough with the release of Dall-E and its open-source counterpart, Stable Diffusion. These programs allow anyone to create original visual art pieces by simply providing descriptions in natural language (prompts). Using a sample of 72,980 Stable Diffusion prompts, we propose a formalization of this new medium of art creation and assess its potential for teaching the history of art, aesthetics, and technique. Our findings indicate that text-to-Image AI has the potential to revolutionize the way art is taught, offering new, cost-effective possibilities for experimentation and expression. However, it also raises important questions about the ownership of artistic works. As more and more art is created using these programs, it will be crucial to establish new legal and economic models to protect the rights of artists.
1. Introduction
AI systems can generate original text and images from natural-language prompts, prompting debate about AI-generated art and creating an underexplored opportunity for visual arts education. The paper examines their use for teaching art history, aesthetics, and technique.
- AI systems can generate relevant and original text and images from simple natural-language prompts, with some outputs recognized in traditional art contests.
- The paper explores how generative AI models could support visual arts education, especially teaching art history, aesthetics, and technique.
- The paper addresses a gap in attention to the educational potential of generative AI despite ongoing debate over whether AI-generated outputs qualify as art.
- The paper proposes examining these models as a compressed version of centuries of human artistic creations.
2. A Brief History of AI-generated Art
AI-generated art developed from early text-generation experiments into deep-learning systems capable of producing convincing images from natural-language descriptions. The emergence of open-source Stable Diffusion expanded access while raising concerns about bias, protected works, and misuse.
- Early systems such as ELIZA demonstrated that software could generate original text responses from human input, although they were not strictly artistic systems.
- Deep learning advanced AI-generated art from early experimental outputs toward increasingly realistic images that attracted attention from the art world and public.
- GPT-3’s general-purpose text generation preceded DALL-E, which could generate convincing images from text descriptions.
- Stable Diffusion emerged as an open-source Text-to-Image model with performance comparable to DALL-E and permissions for commercial and non-commercial use.
- Training on indiscriminate internet data can reproduce social biases and stereotypes, while the use of protected works raises legal concerns.
3. Stable Diffusion
Stable Diffusion is a publicly available latent-diffusion model that generates images from text prompts through CLIP-based semantic selection and denoising. Its parameters support reproducibility and post-processing, while the examples demonstrate both stylistic control and concerns about reproducing living artists’ styles.
- 3. Stable Diffusion: Stable Diffusion is a 2022 text-to-image model whose public code and weights allow operation on most consumer hardware.
- 3. Stable Diffusion: CLIP maps a text prompt into a joint text-image space, after which latent diffusion denoises a semantically related rough image into the final output.
- 3. Stable Diffusion: A fixed seed preserves some generated-image characteristics across prompts, enabling reproducible comparisons such as the paired Figure 1 outputs.
- 3. Stable Diffusion: Figures 1 and 2 use prompts naming Brandon Stanton and Magali Villeneuve to generate photographic and illustrative outputs in their respective styles.
- 3. Stable Diffusion: Inpainting alters masked image regions with prompted content, while outpainting extends images beyond their original dimensions.
- 3. Stable Diffusion: The demonstrated ability to reproduce contemporary artists’ styles is identified as controversial and examined further in the paper.
4. Data and Methods
The study analyzes how Stable Diffusion prompts are structured and what semantic content they contain. It examines 72,980 prompts through tokenization, topic extraction, and classification, including named-entity analysis of artist, brand, and collective references.
- 4. Data and Methods: The dataset contains 72,980 Stable Diffusion prompts collected from Lexica, where curated outputs are paired with their generating prompts.
- 4. Data and Methods: The analysis begins by tokenizing each prompt into linguistic units such as words, phrases, symbols, and other meaningful elements using BERT Tokenizer.
- 4. Data and Methods: GPT-3 is used to extract the main topics or themes represented in the prompts, which describe images in detail.
- 4. Data and Methods: Tokens are classified into one or more identified linguistic topics using GPT-3, while BERT named-entity recognition identifies artist, brand, and collective names.
5. Results and Discussion
Analysis of 72,980 Stable Diffusion prompts identified recurring semantic elements that align with traditional photography concepts and support a procedural prompt framework. The findings also highlight prompt coherence, artist-style references, and unresolved questions about compensation and rights.
- 5.1. Formalizing Stable Diffusion Prompts: Primary prompt categories were identified from the 72,980-prompt sample, while secondary topics provided extensions or additional details.Table 1 presents primary elements, and Table 2 lists less frequent secondary elements.
- 5.1. Formalizing Stable Diffusion Prompts: The identified prompt elements align with traditional photography concepts and can be procedurally classified into a framework for Stable Diffusion prompts.The framework organizes prompt content around semantic elements used in image creation.
- 5.2. Proposed Procedural Classification: Mise-en-scène captures displayed objects, settings, and actors, whereas dispositif captures how the image is created through photographic or digital techniques.Examples include post-apocalyptic settings and lighting for mise-en-scène, and close-up framing, black-and-white treatment, aperture, resolution, and sharpness for dispositif.
- 5.2. Proposed Procedural Classification: Cultural-object elements describe an image’s artifact, purpose, medium, genre, artistic references, meaning, and reception.The examples contrast references to Annie Leibovitz and Michelangelo with descriptors such as religious and award-winning.
- 5.2. Proposed Procedural Classification: Prompt elements are not independent or exclusive, and coherent combinations generally produce better outcomes while generation retains a degree of randomness.The proposed classification therefore treats mastering Text-to-Image as understanding interactions among elements and guiding denoising.
- 5.3. The Need for New Economic Models for Visual Arts: Greg Rutkowski appeared in 41.06% of prompts, while ArtStation appeared in 63.35%, underscoring the prominence of contemporary artists and digital-art platforms as references.The authors link this prevalence to detailed ArtStation labels useful for training and propose compensation models based on artist-name frequency, while noting that implicit uses remain unaccounted for.
6. Conclusions
The paper frames Stable Diffusion as a structured but imperfect medium for connecting prompt-based creation with established art education concepts. With educator guidance, it may support technical teaching and inexpensive experimentation, while requiring ethical and legal clarity.
- 6. Conclusions: Stable Diffusion can connect a new medium of art creation to established art education concepts.The paper acknowledges that outputs are partly random and highly sensitive to wording.
- 6. Conclusions: With proper guidance and curation, Stable Diffusion can support teaching technical concepts, artistic genres, movements, and aesthetics.
- 6. Conclusions: Constant-seed variations of mise-en-scène and dispositif offer a fast, cheap method for experimentation before costly studio time.
- 6. Conclusions: Ethical and legal clarity is necessary for integrating Stable Diffusion and similar software into the art world.