Source-linked AI summary
A Taxonomy of Prompt Modifiers for Text-To-Image Generation
Jonas Oppenlaender
TL;DR
Prompt engineering is difficult to understand because practitioners use diverse, often unintuitive prompt modifiers. This paper uses a three-month online ethnography to develop a taxonomy of six modifier types and structure research on the practice.
Problem
No previous study had investigated different types of prompt modifiers, limiting understanding of prompt engineering in text-to-image digital art.
Method
A three-month online ethnography analyzed how practitioners applied prompt modifiers in prompt writing within the text-to-image art community.
Results
The paper proposes six prompt-modifier types: subject terms, image prompts, style modifiers, quality boosters, repeating terms, and magic terms.
Takeaways & Limitations
The taxonomy provides a foundation for structured investigations and clarifies how prompts are designed for text-to-image generation and AI-generated art.
Takeaways & Limitations
Early text-to-image systems struggled to reliably reproduce specific subjects when an artist’s distinctive style lacked descriptive artwork titles.
Abstract
from arXiv · showhide
Text-to-image generation has seen an explosion of interest since 2021. Today, beautiful and intriguing digital images and artworks can be synthesized from textual inputs ("prompts") with deep generative models. Online communities around text-to-image generation and AI generated art have quickly emerged. This paper identifies six types of prompt modifiers used by practitioners in the online community based on a 3-month ethnographic study. The novel taxonomy of prompt modifiers provides researchers a conceptual starting point for investigating the practice of text-to-image generation, but may also help practitioners of AI generated art improve their images. We further outline how prompt modifiers are applied in the practice of "prompt engineering." We discuss research opportunities of this novel creative practice in the field of Human-Computer Interaction (HCI). The paper concludes with a discussion of broader implications of prompt engineering from the perspective of Human-AI Interaction (HAI) in future applications beyond the use case of text-to-image generation and AI generated art.
1 INTRODUCTION
The paper examines prompt engineering as a non-intuitive creative practice in text-to-image generation, where prompts must follow certain formats to produce desired image styles. Based on a three-month online ethnography, it introduces a taxonomy of prompt modifiers to advance understanding of this practice in HCI.
- Background: Text-to-image systems generate digital images from short descriptive texts, but effective prompts require particular formats to produce desired styles.The practice has become popular in academia and among AI-art practitioners.
- Problem: Prompt engineering has a steep learning curve because modifiers are often unintuitive, image prompts cannot be inferred reliably, and artists frequently withhold complete prompts.Practitioners learn the skill through extensive experimentation and trial.
- Contribution: The paper addresses a gap in prior research by developing a taxonomy of prompt modifiers used by text-to-image practitioners.It focuses specifically on digital art generated with text-to-image systems and draws on an ethnographic study of prompt engineering practices.
- Contribution: A three-month online ethnography analyzes how practitioners apply prompt modifiers in prompt writing to improve understanding of prompt engineering within HCI.The study aims to enhance theoretical understanding of how people write and use prompts and prompt modifiers in human interactions with artificial intelligence.
2 BACKGROUND
The background traces text-to-image generation to multimodal image-text models, particularly CLIP, and frames prompt engineering as the creative practice of controlling generated images. It also describes prompt modifiers, community-developed templates, and iterative experimentation as central to refining outputs.
- Evolution of Text-to-Image Generation: Multimodal models trained on large web-scraped image-text datasets drove rapid growth in deep-learning image synthesis, initially spurred by OpenAI’s CLIP.CLIP is an unsupervised contrastive language-vision model designed for zero-shot image classification.
- Prompt Engineering: Practitioners control image style and quality by adding keywords and key phrases to textual inputs, a creative practice examined through an HCI lens.The paper emphasizes that generating desired images requires more than choosing the subject’s words.
- Prompt Engineering: Prompt engineering is the practice of writing textual inputs and composing sentences to achieve a desired visual style in synthesized images.The practice is also called prompt design, prompt programming, or prompting, and extends beyond text-to-image generation.
- Online Community Practices: Online AI-art communities have developed prompt templates and resources that structure inputs for text-to-image systems.One example organizes prompts as medium, subject, artist(s), details, and image-repository support.
- Prompt Modifiers and Iteration: Prompt engineering is iterative: practitioners run a prompt, observe the result, and adapt it using experimentation, experience, and community resources.Prompt modifiers are keywords or phrases added to direct the resulting image in particular directions.
3 METHOD
The study used autoethnography and online ethnography to investigate prompt engineering and text-to-image art generation from a hands-on, human-centered perspective. It combined iterative experimentation, Twitter community engagement, literature review, and inductive taxonomy development.
- 3.1 Autoethnography: Prompt engineering was studied as an acquired skill learned through iterative experimentation and community-provided resources, including guides, reports, and shared social-media prompts.The authors characterize this learning process as akin to “brute-force trial and error.”
- 3.1 Autoethnography: The author conducted a 3-month autoethnographic study between October and December 2021, experimenting with text-to-image synthesis through Google Colaboratory notebooks.The study provided a practitioner’s perspective through learning from self-use.
- 3.2 Online ethnography: The online ethnography treated the author as a participant-as-observer who engaged with Twitter discussions, posted generated images, and followed trending text-to-image art hashtags.Hashtags included #vqganclip, #VQGAN, #clipguideddiffusion, #digitalart, #AIArt, and #generativeart.
- 3.2 Online ethnography: The research also reviewed literature on text-to-image generation and prompt engineering, relying substantially on gray literature because scholarly HCI research on AI-generated art was scarce.Liu and Chilton’s design guidelines were identified as an exception.
- Taxonomy development: The taxonomy was developed inductively and iteratively by compiling, grouping, and continually reinterpreting candidate prompt modifiers as novel prompt instances appeared.The ethnographic research was documented visually and textually to support ongoing understanding and verification of the taxonomy.
- Research perspective: The study adopted a human-centered rather than technical lens, focusing on users’ text-based interactions with text-to-image systems and novel creative practices.The author’s background was in Computer Science, HCI, and Social Computing, and the author was not an artist.
4 TAXONOMY OF PROMPT MODIFIERS
The taxonomy identifies six types of prompt modifiers used in text-to-image art: subject terms, image prompts, style modifiers, quality boosters, repeating terms, and magic terms. These modifiers control subjects, visual targets, style, quality, associations, and surprise through varied prompt forms.
- Taxonomy: The taxonomy distinguishes six modifier types: subject terms, image prompts, style modifiers, quality boosters, repeating terms, and magic terms.The categories reflect practitioners’ understanding of prompt modifiers in the text-to-image art community.
- Subject terms: Subject terms specify the desired subject and are essential for controlling generation, although descriptive training data can reduce their control over outcomes.Early systems such as VQGAN–CLIP struggled to reproduce specific subjects in works by Zdzisław Beksiński, whose artworks lacked titles.
- Visual targets and style: Image prompts provide visual targets for subject and style, while style modifiers consistently evoke artistic styles, periods, schools, materials, or media.Image prompts are typically supplied as one or more URLs, whereas style modifiers include artist attributions and terms such as “oil painting.”
- Quality and associations: Quality boosters increase aesthetic qualities and detail, while verbosity may improve overall quality at the expense of subject controllability; repeating terms strengthen semantic associations.Repeating terms can use different phrasing or synonyms to more reliably activate latent-space regions associated with subjects.
- Unpredictability and form: Magic terms introduce randomness, unpredictability, and surprise, and modifiers can appear as hashtags, attribution phrases, or complex composite statements.The example “control the soul” was added to encourage more magical, wizard-like imagery.
5 PROMPT ENGINEERING IN PRACTICE
Prompt engineering for text-image generation is presented as an iterative process that begins with subject terms and progressively adds modifiers, solidifiers, and weights. These prompt elements can control subjects, styles, quality, and exclusions, with overlapping effects among modifier types.
- Iterative prompt design: Controlled image generation typically begins by denoting the subject with one or more terms, while other prompt components remain optional.Practitioners can generate artworks from a single term such as “car,” although subject terms are fundamental to controlled generation.
- Iterative prompt design: Style modifiers and quality boosters are added to alter image style or improve quality, but their effects can overlap.The modifier “by Greg Rutkowski” illustrates how the distinction between these categories is sometimes unclear.
- Iterative prompt design: Solidifiers reinforce and stabilize styles by repeating terms, most often subject terms, while remaining applicable to subjects, styles, and quality boosters.Image prompts can carry both subject and style information because of their visual nature.
- Iterative prompt design: Weights can be assigned to all six modifier types, including negative weights that exclude unwanted subjects or styles from generation.For example, “heart:-1” can prevent VQGAN–CLIP from generating heart-shaped objects associated with “love.”
- Iterative prompt design: The overall workflow adds modifiers and solidifiers iteratively or through learned experience, then applies weights to exclude or mix subjects and styles.Subject terms are usually written first because they are most important for controlled image generation.
6 DISCUSSION
The accessibility of text-to-image generation, supported by an ecosystem of technologies and resources, has driven an explosion of AI-generated artworks shared online. Prompt modifiers are central to prompt engineering, and the six-type taxonomy offers an initial framework for examining this practice.
- Emerging creative practice: The accessibility of text-to-image generation and its supporting technologies and resources have fueled an explosion of AI-generated artworks shared online.The discussion characterizes prompt engineering as an emerging creative practice and artistic medium.
- Emerging creative practice: Prompt modifiers are key to the emerging creative practice of prompt engineering, which the paper approaches through a taxonomy of six modifier types.The taxonomy is presented as an initial contribution for understanding prompt engineering.
- Access and scale: 15 million members were using Midjourney at the time of writing, while Stable Diffusion could run on cloud or local hardware.These examples illustrate the growing availability of text-to-image systems through free or relatively inexpensive means.
- Access and scale: 80% of technology products and services were estimated to be built by people who were not technology professionals by 2024.Gartner made this estimate in 2021, in the context of broader implications for productivity and creativity.
6.1 Broader Implications for Human-AI Interaction
Prompt engineering has implications beyond text-to-image synthesis and AI-generated art for human interaction with deep learning models and AI generally. As foundation-scale models power diverse applications, studying prompt design is timely for understanding future creative work and Human-AI Interaction.
- Prompt engineering is relevant to human interaction with deep learning models and artificial intelligence beyond text-to-image synthesis and AI-generated art.
- Deep learning may disrupt and transform creative-economy sectors as generative systems synthesize increasingly complex outcomes, including text-to-video generation.
- AI-based systems may transform not only computer interaction and online work but also work content and human agency.
- Text-to-image art is one application of prompt engineering, whose broader HAI implications include relationships with increasingly opaque models.
- Research on prompt design is timely because foundation-scale models enable diverse applications, including tool-augmented language models that use prompts to access external tools.
6.2 Opportunities for Research on Prompt Engineering in HCI
The paper identifies HCI research opportunities in prompt engineering spanning community dynamics, emerging creative workflows, bias in AI-driven systems, computational aesthetics, and human-AI alignment. It highlights social learning, practitioner workflows, encoded and imposed biases, and prompts as expressions of aesthetic intent.
- 6.2 Opportunities for Research on Prompt Engineering in HCI: HCI research could investigate the social components of text-to-image generation, including how users anticipate descriptions and reactions associated with images on the Web.Because systems are trained on Web-scraped images and text, detailed description alone may not achieve optimal results.
- 6.2.1 Social aspects of prompt engineering.: Dedicated Discord-based communities provide social features for studying profiles, shared creations, prompts, collaborative iteration, daily themes, and long-running discussions.Midjourney profiles connect successful creations with their prompts, while group jams support developing members’ works.
- 6.2.1 Social aspects of prompt engineering.: Community learning is a research opportunity, including how practitioners receive and seek inspiration within text-to-image art communities.The paper proposes detailed ethnographic investigation using its taxonomy as a conceptual starting point or framework.
- 6.2.1 Social aspects of prompt engineering.: Prompt engineering can extend into complex creative workflows that combine multiple text-to-image systems with photo editing.Practitioners may generate initial images for inspiration in one system, continue in another, and finalize them in a photo editor.
- 6.2.2 Human-AI co-creation.: Research should examine bias encoded in text-to-image systems, including CLIP bias and Western-oriented training data reflected in images prompted with “princess.”OpenAI announced reduced bias in DALL-E 2, but the passage indicates this involved a potential cost without specifying it further.
- 6.2.3 Bias in image generation systems.: Responsible deployment and content policies can impose organizational values, create user frustration, and raise concerns about restricting human aesthetic values and experience.DALL-E 2 users may receive notices for terms related to war or sexual content, with possible account closure after repeated warnings.
- 6.2.3 Bias in image generation systems.: Prompts offer a resource for computational-aesthetics research because they encapsulate a person’s stated intent, although that intent is likely only partially explicit.Text-to-image systems increasingly consider human aesthetics when attempting to produce better images.
7 CONCLUSION
The paper proposes a six-type taxonomy of prompt modifiers that structures understanding and investigation of prompt engineering for text-to-image generation and AI-generated art. It also identifies prompt modifiers as tools for improving image quality and controlling image creation, while pointing to ethical, societal, social, and co-creative directions for future research.
- Contribution: The paper proposes six prompt-modifier types: subject terms, image prompts, style modifiers, quality boosters, repeating terms, and magic terms.The taxonomy provides a foundation for structured investigations into prompt engineering for text-to-image generation and AI-generated art.
- Prompt engineering practice: Subject terms support controlled image creation, while practitioners use modifiers to improve quality and exert greater control over the creation process.Modifiers can alter image style or enhance quality, and their effects may overlap.
- Future research: The taxonomy is a stepping stone for studying prompt engineering and motivates research on the ethical, societal, social, and co-creative dimensions of AI-generated creative work.These areas are identified as key directions for future research.