Source-linked AI summary
Prompting AI Art: An Investigation into the Creative Skill of Prompt Engineering
Jonas Oppenlaender, Rhema Linder, Johanna Silvennoinen
TL;DR
As text-to-image systems make digital image creation accessible, the paper asks whether prompt engineering is an intuitive ability or an acquired creative skill. Across three studies with crowdsourced participants, it examines prompt evaluation, writing, and refinement, finding descriptive competence but inadequate style-specific vocabulary and limited image improvement. The authors therefore characterize prompt engineering as a non-intuitive skill requiring acquisition through practice and learning, while acknowledging limitations in assessing generated-image quality.
Problem
The paper addresses limited evidence about whether prompt engineering is intuitive or an acquired skill requiring practice and learning.
Method
The paper investigates prompt engineering for AI art through three studies examining whether crowdsourced participants can evaluate, write, and improve prompts and resulting images.
Results
Participants could craft descriptive prompts but lacked the special vocabulary used in AI art communities, and only a minority improved their revised images.
Takeaways & Limitations
Prompt engineering is characterized as a new, non-intuitive skill that must be acquired before it can be applied meaningfully.
Takeaways & Limitations
The exploratory studies lacked direct measurements of the actual quality of generated images.
Abstract
from arXiv · showhide
We are witnessing a novel era of creativity where anyone can create digital content via prompt-based learning (known as prompt engineering). This paper investigates prompt engineering as a novel creative skill for creating AI art with text-to-image generation. In three consecutive studies, we explore whether crowdsourced participants can 1) discern prompt quality, 2) write prompts, and 3) refine prompts. We find that participants could evaluate prompt quality and crafted descriptive prompts, but they lacked style-specific vocabulary necessary for effective prompting. This is in line with our hypothesis that prompt engineering is a new type of skill that is non-intuitive and must first be acquired (e.g., through means of practice and learning) before it can be used. Our studies deepen our understanding of prompt engineering and chart future research directions. We conclude by envisioning four potential futures for prompt engineering.
1 Introduction
The introduction frames prompt engineering as a potentially non-intuitive creative skill whose acquisition matters for AI art, education, and human-AI interaction. Three studies examine whether participants can evaluate, write, and refine prompts, finding descriptive ability but insufficient style-specific vocabulary and limited improvement.
- Text-to-image practitioners use prompt modifiers to influence generated images’ quality and artistic style, while their effects may be unintuitive to laypeople.
- Prompt engineering is presented as a largely unexplored question of whether the skill is intuitive or must be acquired through practice and learning.
- Prompt engineering matters for education because curricula decisions depend on whether it has a learning curve and requires substantial learning.
- The paper investigates prompt engineering in three studies with 227 crowdsourced participants.
- Study 1 found that participants could grasp what makes a good prompt, but later studies found rich descriptive language without the style-specific vocabulary needed for effective prompting.
- Participants were not able to significantly improve artwork quality in the follow-up study, supporting the view that prompt engineering is non-intuitive and must first be acquired.
2 Related Work
Related work describes text-to-image generation, community-developed prompt modifiers, and prior research on applying knowledge in practice. The paper defines prompt-engineering skill as using language and prior knowledge, including modifiers, to guide generative models toward desired outputs.
- 2.1 Text-to-Image Generation with Deep Learning: Text-to-image generation uses text descriptions as input to deep-learning systems that synthesize images, including through text-conditional models.
- 2.2.1 The engineering character of prompting: Prompt engineering is commonly practiced through systematic trial and error to find terms that describe intended outputs and anticipate web-based descriptions of them.
- 2.2 Prompt Modifiers: Prompt modifiers are keywords or phrases that alter image style, quality, or both, including quality boosters such as “8k” and “trending on artstation.”
- 2.2 Prompt Modifiers: Community resources and prior studies provide guidance on subject and style keywords, while users without modifier knowledge may need brute-force trial and error.
- 2.3 Prompt Engineering as a Skill: The paper defines skill in prompt engineering as effectively using language and prior knowledge to craft prompts that guide generative models toward desired outputs.
- 2.3 Prompt Engineering as a Skill: This skill includes both relevant language syntax and strategic use of prompt modifiers to refine or redirect generative outputs.
- 2.4 Prior Work on Applying Skill in Practice: The paper relates prompt engineering to research on how people acquire skills, rewrite failed queries, and improve practice.
3 Study 1: Understanding Prompt Engineering
Study 1 tested whether participants could understand prompt engineering by rating AI-generated artworks and prompts separately, then comparing ratings across modalities. Participants distinguished high- from low-quality prompts and artworks, with prompt ratings showing a larger quality span and correlating positively with artwork ratings.
- Study design: The within-subject experiment separately measured aesthetic ratings for 20 AI-generated artworks and 20 textual prompts.Participants rated prompts by imagining the images they would generate, using the same 5-point scale as for artworks.
- Study design: The study treats prompt engineering as a practice in which users write prompts, inspect generated images, and assess resulting quality.Rich descriptive language and multiple prompt modifiers were presented as features associated with higher-quality generated artworks.
- Research materials: The study focused on subjective aesthetic appeal rather than establishing a universal image-quality standard or measuring prompt fidelity.The authors note that future studies could include prompt fidelity as an additional quality dimension.
- Visual and prompt ratings: Prompt and artwork ratings differed significantly across study conditions, with pairwise comparisons generally yielding p < 10^-4.The omnibus comparison was χ2 = 231.4, p < 10^-15, df = 3; Artwork–High versus Prompt–High was the stated exception to the pairwise threshold.
- Visual and prompt ratings: Participants could differentiate high- from low-quality artworks and prompts.Artwork–High averaged µ = 3.70 versus Prompt–Low µ = 3.39, while Prompt–High averaged µ = 3.87 versus Prompt–Low µ = 2.78.
- Connection between visual image and prompt quality: Ratings of prompts and corresponding artworks showed a significant positive correlation, indicating that higher-rated prompts were more likely to correspond to higher-rated artworks.The reported correlation was positive and significant at p < 10^-15.
4 Study 2: Writing Prompts
Study 2 examined whether crowdsourced participants could write effective prompts for generating digital artworks. Participants generally produced rich descriptive language, but rarely used style-specific prompt modifiers, leaving generated image styles largely uncontrolled.
- Study 2: Writing Prompts: The study asked crowdsourced participants whether they could create effective text-to-image prompts for digital artworks.Style information was treated as a possible indicator of understanding how effective prompts are formulated.
- Study design: Each participant wrote three prompts intended to maximize the visual attractiveness and aesthetic qualities of generated artworks.The task instructions did not mention prompt modifiers or prime participants with a specific artistic style.
- Study design: The study collected prompts without showing participants the generated images, preserving the prompts for a later follow-up study.Instructions emphasized that there was no right or wrong answer, while task validity depended on following the instructions.
- Participant recruitment: The final dataset contained 375 prompts written by 125 unique participants after rejecting invalid or gamed responses.Ten tasks were rejected and republished, followed by removal of twelve additional responses.
- Analysis: Prompt analysis combined manual coding of AI-art keywords and phrases with quantitative measures of descriptive language and lexical diversity.The quantitative measures included token and type counts and the Type-Token Ratio, while prompt modifiers were coded for presence or absence.
- On the use of descriptive language: Participants generally used rich descriptive language when describing the intended image content.The analysis quantified descriptive language using prompt length, unique-word counts, and Type-Token Ratio.
- On the use of prompt modifiers: Only a few participants included style information, and participants rarely used artistic styles, artist names, genres, media, or techniques.Because style modifiers were largely absent, generated styles were mainly determined by descriptive language and could diverge from participant intent.
5 Study 3: Improving Prompts
Study 3 tested whether participants could improve their previously generated artworks by revising prompts after viewing outputs. Participants made many textual changes, but style modifiers remained rare and image-quality improvements were inconsistent.
- Study 3: Improving Prompts: Study 3 tested whether prompt engineering was applied intuitively or required expertise, practice, and knowledge of prompt modifiers.The authors hypothesized that participants would not significantly improve artworks after only limited interaction with the system if the skill were learned.
- Study design: The same participants reviewed five images for each of their three earlier prompts and rewrote those prompts to improve the results.Each revision interface included a field pre-filled with the previous prompt and an optional field for negative terms.
- Study design: Negative terms were introduced as a possible improvement tool for controlling image subjects and quality.The study explained that adding “zebra” negatively to a pedestrian-crossing prompt could remove stripes and produce a plain road.
- Analysis: The analysis measured added and removed tokens, Levenshtein distance, and eight qualitative categories of prompt change.Categories included adjectives/adverbs, subjects, prepositions, paraphrasing or synonyms, reordering, cardinal numbers, simplification, and prompt modifiers.
- Participants’ revised prompts: 11 prompts (7.33%) were unchanged and lacked a negative term, including six cases where participants pasted random text snippets.The prompt-revision data therefore included both non-revisions and apparently irrelevant additions.
- Participants’ revised images: About half of image sets remained unchanged, 15% worsened, and one third improved after prompt revision.Revised sets were often in the same or a very similar style because participants rarely used style modifiers.
6 Discussion
The studies show that participants could judge prompt quality and write descriptive prompts, but generally lacked style-specific vocabulary and struggled to improve generated images.
- 227 participants took part in three studies assessing prompt evaluation, writing, and improvement.
- Participants could assess prompt quality, with performance increasing alongside art experience and interest.
- Participants wrote creative prompts using rich descriptive language that sometimes produced beautiful or interesting images.
- Participants rarely used specific keywords or modifiers associated with style and image-quality control.
- Only a minority improved revised images, while most images remained about the same quality.
- An overwhelming majority left image style to chance despite being instructed to create artworks.
- These findings reveal a gap between promising crowd engagement with AI art generation and effective prompt engineering.
6.1 Prompt Engineering as a Non-Intuitive Skill
The paper argues that effective prompt engineering is non-intuitive because it requires learned knowledge of keywords, modifiers, and broader creative-tool workflows.
- Prompt engineering is presented as a skill that may require practice, learning, and specialized training rather than effortless intuitive use.
- Participants produced descriptive prompts but generally lacked style modifiers and community-specific vocabulary needed to control generated images.
- The studies empirically support viewing prompt engineering as non-intuitive or potentially specialist.
- A single interaction with the model did not produce learning effects, consistent with the need for acquired skill.
- Text-to-image generation is difficult to control because initial outputs are highly random and often require several iterations.
- Prompt engineering extends beyond text to a creative-tool ecosystem including image editors, image-to-image generation, inpainting, outpainting, and upscaling.
6.2 On the Future of Creative Production with Prompt Engineering
The paper speculates that prompt engineering may become an expert, everyday, obsolete, or personal-signature skill. Its future depends on training needs, tool development, and users’ ability to control or curate outputs.
- 6.2.1 Prompt engineering as an expert skill: Prompt engineering may become an expert skill requiring subject-matter knowledge, specialized modifiers, and mastery of complex workflows.
- 6.2.1 Prompt engineering as an expert skill: If prompt engineering becomes highly skilled, extensive training could make it exclusive to a narrow privileged group.
- 6.2.1 Prompt engineering as an expert skill: Participants’ prompts notably lacked the subject-specific keywords, modifiers, and combinations needed to control outputs effectively.
- 6.2.2 Prompt engineering as an everyday skill: Prompt engineering could become an everyday skill because most participants wrote creative, detailed prompts and improving systems may require fewer modifiers.
- 6.2.3 Prompt engineering as an obsolete skill: Prompt engineering could become obsolete as models improve their understanding of user intent and automatically rewrite prompts.
- 6.2.4 Prompt engineering as personal signature or curation skill: In this possible future, prompting would remain useful for finishing, optimizing, and personalizing otherwise accessible generative results.
- 6.2.4 Prompt engineering as personal signature or curation skill: Prompt engineering could persist as a personal-signature or curation skill because users can recognize prompt quality and distinctive prompting styles.
- 6.2.5 Review and outlook: Prompt engineering is described as a perishable skill because model updates require practitioners to adopt new techniques, while curricula have not adapted.
6.3 Future Work
Future work should reassess prompt engineering as models and user familiarity change, separate interactive learning from prompting skill, and examine the broader set of creative practices involved.
- The study used naive participants recruited in mid-2022, a sample that is difficult to reproduce because image generators are now more familiar.
- Because text-to-image interaction includes repeated prompting after observing outputs, future research must disentangle interactive learning from underlying prompt-engineering skill.
- Future studies should investigate how quickly prompt engineering is learned, its components, learning curve, and time required for mastery.
- The paper studied textual prompting, while practical text-to-image generation also includes visual prompting and configuration parameters.
- Future research should consider the complexity and diversity of skills across the entire creative process rather than textual prompts alone.
- Future work could examine AI systems that adapt to humans, understand user intent, and foster more intimate relationships with users.
6.4 Limitations
The exploratory studies face limitations in aesthetic assessment, system choice, task design, and the absence of long-term learning evidence. These constraints narrow how broadly the findings about prompt engineering can be interpreted.
- Task design: Study 1 required participants to visualize and evaluate the potential outcome of written prompts, making the task particularly challenging.The authors acknowledge that this difficulty may constrain interpretation of participants’ performance.
- Aesthetic evaluation: Aesthetic quality assessment is subjective, and the studies lacked direct measurements of the generated artworks’ actual quality.The authors recommend more concrete and objective criteria for evaluating prompt–artwork pairs.
- Generation system: Latent Diffusion balanced reproducibility and performance, but it may respond differently to style keywords than CLIP-guided systems.Other tested systems were either nondeterministic or insufficiently advanced to produce recognizable results at the time.
- Study setup: The studies did not provide interactive feedback, and participants generally did not use prompt modifiers or see second-round images.This setup targeted intrinsic participant skill but limited conclusions about adaptation to generative systems.
- Instruction design: Study 3 instructions did not explicitly direct participants to modify stylistic elements or prompt modifiers, potentially favoring substantive prompt changes.More directed instructions could clarify how novices perceive prompt modifiers.
- Learning dynamics: A long-term field study is needed to investigate prompting-skill learning, whereas this research provides only a first indication from a one-time modification experiment.The authors therefore treat the learning conclusion as preliminary.
7 Conclusion
This paper examines prompt engineering as a creative skill in AI art through three studies of lay participants’ prompt evaluation, writing, and improvement. Participants wrote rich descriptive prompts but lacked AI-art-specific vocabulary, supporting the view that prompting is non-intuitive and may require learning; the authors discuss four possible futures.
- 7 Conclusion: The paper investigates whether people unfamiliar with text-to-image generation can recognize, write, and improve prompts without repeated system feedback.The three studies address prompt evaluation, prompt creation, and prompt refinement.
- 7 Conclusion: Participants were creative and wrote prompts in rich descriptive language, but lacked the specialized vocabulary used in AI-art communities.This distinction separates general descriptive ability from familiarity with prompt modifiers and style terminology.
- 7 Conclusion: The findings indicate that prompting is non-intuitive and not an innate skill users can apply without first learning about it.The conclusion is presented as an indication rather than definitive evidence from long-term training.
- 7 Conclusion: The paper discusses the importance of its findings and speculates on four potential futures for prompt engineering.It also situates generative AI as a source of future research opportunities in HCI.
A.1 Images with High Aesthetic Appeal
The appendix presents examples of images with high aesthetic appeal alongside their associated prompts. The examples include varied subjects and style-related wording.
- A.1 Images with High Aesthetic Appeal: The high-appeal examples include scenes, interiors, portraits, characters, and waves paired with descriptive prompt fragments.Examples reference subjects such as Vikings, a sanctuary, a world-war portrait, and an eclectic interior.
- A.1 Images with High Aesthetic Appeal: The examples demonstrate that high-appeal prompt fragments vary substantially in subject matter and descriptive formulation.The appendix is illustrative rather than a reported quantitative comparison.
- A.1 Images with High Aesthetic Appeal: Several high-appeal prompts combine subject descriptions with style or rendering terms such as matte painting, Ghibli, octane, and 8k.These examples also include phrases associated with artistic communities and image-generation practice.
A.2 Images with Low Aesthetic Appeal
The appendix presents examples of images with low aesthetic appeal and their associated prompt fragments. The examples span social, workplace, historical, and artistic subjects.
- A.2 Images with Low Aesthetic Appeal: The low-appeal examples include references to bias, Asterix, workplace dialogue, ceiling height, China buying Russia, and green-screen effects.These fragments cover both textual scenarios and visual-art concepts.
- A.2 Images with Low Aesthetic Appeal: The appendix provides qualitative examples of low-appeal outputs without reporting a numerical comparison among them.Its role is to show the prompt examples associated with the category.
- A.2 Images with Low Aesthetic Appeal: The low-appeal prompt fragments vary from short labels to longer quoted dialogue and descriptive phrases.The examples include “we can do it!”, office conversation, and an “amazing green screen” effect.