Source-linked AI summary
Design Guidelines for Prompt Engineering Text-to-Image Generative Models
Vivian Liu, Lydia B. Chilton
TL;DR
Text-to-image prompting is open-ended but can force users into brute-force trial and error, while rigorous guidance for visual prompt engineering remains limited. The paper conducts five experiments on subject-and-style prompts and model parameters, using 5493 generations across diverse subjects and styles. It synthesizes empirical design guidelines for prompt wording, seed selection, optimization length, and subject–style choices.
Problem
Open-ended text interaction makes image-generation iteration potentially random and unprincipled, while visual prompt engineering has received less rigorous study.
Method
The paper conducts five experiments examining prompt permutations, random seeds, optimization length, style keywords, and subject–style keywords in a text-to-image framework.
Results
The experiments identify empirical guidelines: emphasize subject and style keywords, try multiple seeds, use 100–500 optimization iterations for fast iteration, and account for subject–style abstractness.
Takeaways & Limitations
The paper provides design guidelines for producing better outcomes from text-to-image generative models.
Takeaways & Limitations
Most experiments studied one prompt, “SUBJECT in the STYLE of,” with VQGAN+CLIP; other prompts, models, and modifiers remain for further research.
Abstract
from arXiv · showhide
Text-to-image generative models are a new and powerful way to generate visual artwork. However, the open-ended nature of text as interaction is double-edged; while users can input anything and have access to an infinite range of generations, they also must engage in brute-force trial and error with the text prompt when the result quality is poor. We conduct a study exploring what prompt keywords and model hyperparameters can help produce coherent outputs. In particular, we study prompts structured to include subject and style keywords and investigate success and failure modes of these prompts. Our evaluation of 5493 generations over the course of five experiments spans 51 abstract and concrete subjects as well as 51 abstract and figurative styles. From this evaluation, we present design guidelines that can help people produce better outcomes from text-to-image generative models.
1 INTRODUCTION
Text-to-image models offer open-ended visual generation, but prompt design can become unprincipled trial and error. The paper therefore studies prompt wording and model settings for subject-and-style prompts.
- Text-to-image generation gives users broad creative possibilities through free-form prompts.
- Open-ended prompting can make iteration feel random because users must search for new prompts when generation quality is poor.
- The paper systematically studies prompts structured as “SUBJECT in the style of STYLE.”
- Five experiments examine prompt phrasing, random initialization, optimization length, styles, and subject–style combinations.
2 RELATED WORK
Related work frames generative models as tools for exploring large design spaces while highlighting the need for interpretable, semantically meaningful controls. The paper extends prompt-engineering research toward rigorous visual-generation studies.
- Generative methods provide large design spaces that researchers seek to support during ideation and iteration.
- Many AI-based approaches lack meaningful and interpretable user controls despite producing many generations.
- Creativity-support systems such as AttriBit, CoCoCo, and Geppeto connect generated outputs with semantic goals.
- Prompt-engineering research has largely focused on text generation, including prompt shapes, answer engineering, and task-specific templates.
- Visual prompt engineering has received less rigorous study than text prompt engineering.
- Prior visual prompting practice has been informal and ad hoc, including community-developed keywords and the “X in the style of” template.
- Studying prompts can help probe multimodal models and develop better mental models of their knowledge distributions.
3 EXPERIMENT 1. PROMPT PERMUTATIONS
Experiment 1 tests whether rephrasing subject-and-style prompts changes generation quality. Across nine permutations, the study finds no statistically meaningful quality difference and recommends emphasizing subject and style keywords.
- 3 EXPERIMENT 1. PROMPT PERMUTATIONS: The experiment asks whether different phrasings of the same prompt produce better or worse generations.
- 3.1 Methodology: The study used VQGAN+CLIP with 12 subjects, 12 styles, 256x256 images, and 300 optimization steps.
- 3.1 Methodology: Across 144 subject–style combinations, nine prompt permutations generated 1296 images for comparison.
- 3.1 Methodology: The tested permutations varied mediums, word order, punctuation, function words, and connective phrasing around subject and style keywords.
- 3.2 Annotation Methodology: Two media-arts annotators rated randomly arranged 3x3 grids for significantly better or worse outlier generations.
- 3.3 Results: 71.3% agreement across 1296 generations accompanied a Cohen’s kappa of 0.0013.The authors attribute the low kappa to subjectivity in identifying outliers and proceed using the higher agreement value.
- 3.3 Results: 0.354 and p-value 0.55 indicate that prompt permutations judged as outliers were insignificant relative to non-outliers.
- 3.3 Results: The authors found no significant difference among the nine permutations and recommend focusing on subject and style keywords rather than connecting words.
4 EXPERIMENT 2. RANDOM SEEDS
Experiment 2 tests whether random seeds affect generation quality when the prompt is held constant. The results show significant quality outliers, motivating multiple-seed sampling during prompt engineering.
- Because generative models depend on stochastic initialization, the experiment asks whether identical prompts produce different-quality generations across seeds.
- The study generated 1296 images from 12 subjects, 12 styles, and nine seeds using “SUBJECT in the style of STYLE” prompts.
- Annotators compared randomly arranged 3x3 grids containing nine seed-based generations for each subject–style combination.
- p-value <0.01 shows that generations judged as outliers significantly outnumbered those judged non-outliers.
- Inter-rater reliability was 0.13, indicating slight agreement that the authors attribute to the task’s subjective judgments.
- The authors conclude that seed choice can significantly vary generation quality.
- The resulting guideline recommends generating between 3 to 9 different seeds to represent what a prompt can return.
5 EXPERIMENT 3. LENGTH OF OPTIMIZATION
Experiment 3 tested whether longer optimization produces better generations and found that preferred outputs generally came from shorter runs. The authors recommend 300 iterations as a practical default, while noting that 100 iterations may not yet manifest the subject.
- Experiment 3 tested whether optimization length correlates with better-evaluated generations.The study varied iteration counts while evaluating intermediate generations.
- 72 rows of generations were evaluated at iteration steps from 100 through 1000, with annotators selecting their preferred image from each set of 10.The experiment used 6 subjects across 12 styles, with a constant seed and one prompt permutation.
- p-value=0.01 showed significant differences among preferred iteration steps, with 200, 100, and 500 iterations chosen most often.Annotator agreement was fair, with Cohen’s kappa=0.33.
- Lower iteration counts of 100–500 tended to be preferred over higher values, indicating that more iterations did not necessarily produce more desirable generations.Figure 5 reports the preference frequency and significance of this difference.
- 300 iterations is suggested as a default because 100 iterations may not yet manifest the subject, especially when the style is abstract.The authors use 300 iterations in subsequent experiments.
6 EXPERIMENT 4. TESTING A BREADTH OF STYLES
Across 51 styles, VQGAN+CLIP reproduced salient visual properties such as color, texture, technique, composition, perspective, and recognizable motifs, but performance varied substantially by style. Failures commonly arose from semantic misinterpretation, incomplete style capture, incongruent elements, default motifs, and difficulty representing symbolic or culturally nuanced meanings.
- Results: Mean ratings varied across the 51 tested styles, with the model performing better on some styles than others.The evaluation aggregated ratings across subjects for each style.
- Success modes: Successful generations often captured salient color schemes, textures, and style-consistent visual techniques.Examples include signature palettes in cyberpunk and glitch art, ink-wash textures, and characteristic lines or brush strokes.
- Success modes: The model also reproduced style-specific primitives, spatial patterns, lighting, and perspective across multiple artistic styles.Examples include dots in Pointillism, deconstructed shapes in Cubism, Op art patterns, and distinct two-dimensional or three-dimensional lighting.
- Success and failure modes: Recognizable motifs could evoke a style, but sometimes reflected related meanings from decor, architecture, or sculpture rather than the intended visual-art style.Baroque furniture and Neoclassical architecture or sculpture appeared in place of the intended painting traditions.
- Failure modes: Failures included prompt misunderstandings, incomplete style capture, style-incongruent elements, and defaulting to unconvincing motifs.Concrete subjects could retain photorealistic textures even when the requested style was abstract or sketch-like.
- Failure modes: Symbolic or culturally nuanced styles remained difficult to represent visually, illustrated by poor performance for Dadaism and Bauhaus.Dadaism received a mean subjective rating of 1.42, while the passages link these difficulties to abstraction and cultural knowledge.
6.6 Results and Discussion of Partitions
Partition analyses found that figurative styles outperformed abstract styles, digital styles underperformed modern and premodern styles, and Western and non-Western styles did not differ significantly. The authors therefore conclude that users can try styles broadly, while recognizing that misinterpretation and other style-dependent failures affect outcomes.
- Abstract versus figurative: Abstract styles averaged 2.63, while figurative styles averaged 3.16; the difference was significant at p < 0.01.The comparison used 33 specific styles after excluding general mediums and Internet aesthetics.
- Abstract versus figurative: The authors’ hypothesis that abstract styles would perform better was not generally supported, because abstract styles showed a wide range of failure modes.Some abstract styles benefited from tolerance for deconstruction, but others suffered from misinterpretation and incomplete style capture.
- Abstract versus figurative: Top-performing figurative styles included Ukiyo-e, Impressionism, documentary photography, and cyberpunk, while lower-performing examples showed incomplete capture or unconvincing motifs.Kerala mural style and art deco are given as examples of these failure modes.
- Western versus non-Western: Western styles averaged 2.92 and non-Western styles 2.95, with no significant distribution difference (p-value: 0.377).The comparison used a Mann-Whitney test.
- Time-period partitions: Digital, modern, and premodern styles received aggregate ratings of 2.41, 2.83, and 3.11, respectively, with differences significant at p-value < 0.001.Digital styles performed worst, followed by modern and then premodern styles.
- Design guideline: The design guideline is to try styles broadly, because many perform well when they are not prone to misinterpretation or other identified failure modes.The guideline follows the observed breadth of style performance rather than recommending a restricted style set.
7 EXPERIMENT 5: INTERACTION BETWEEN SUBJECT AND STYLE
Experiment 5 examined how subject and style interact in 1,581 generated images, finding that concreteness, stylistic compatibility, and relevance shaped generation quality. The results support choosing subjects that complement the selected style in abstractness and relevance.
- Experiment setup: 1,581 images crossed 51 subjects with 31 styles to evaluate subject–style interaction across abstractness and stylistic diversity.Two domain-knowledgeable annotators rated the coherency of subject and style within each image.
- Subject effects: Top-ten subjects were all concrete, with average concreteness 4.47; ocean, forest, house, eye, and bird ranked among the strongest.Subject concreteness had a Pearson correlation of 0.62 with generation quality, indicating a moderate-to-strong positive association.
- Subject–style interaction: Aggregate rankings increased from abstract-abstract (3.05) and abstract-concrete (3.17) to figurative-abstract (3.49) and figurative-concrete (3.54).Two-way ANOVA found significant effects for both factors and their interaction, with p-values well below 0.01.
- Success modes: Abstract subjects succeeded when the model represented them through recognizable symbols or when subject and style components matched and blended.Examples include hearts producing heart symbols, freedom producing American flags, and intelligence appearing as a fractal brain.
- Success modes: Concrete subjects succeeded when they emerged from a style without disrupting its characteristic visual treatment.Rain retained High Renaissance line qualities, while a car used Unreal Engine depth-of-field and scene-lighting effects.
- Failure modes: A recurring failure was that abstract subjects did not come through, while incompatible subjects such as website could be dropped from the image.Some style combinations nevertheless produced nuanced interpretations of difficult subjects such as progress.
- Design guideline: When selecting a subject, choose one that complements the chosen style in level of abstractness and relevance.The study also observed grotesque or inflammatory imagery, including repeated body-part motifs and unsettling details without trigger warnings.
8 DISCUSSION
The discussion distills empirical findings into practical prompt-engineering guidelines while highlighting limitations in style and subject interpretation, model control, and the study’s scope. It also emphasizes that generated images may misrepresent artistic traditions or contain offensive content.
- Design guidelines: The authors condense their findings into default parameters and methods for end users interacting with text-to-image models.
- Design guidelines: Focus on subject and style keywords, generate 3 to 9 seeds, and use 100 to 500 optimization iterations for fast iteration.The authors report no significant correlation between iteration length and user satisfaction.
- Design guidelines: Any style can be tried, including niche styles, but style keywords prone to misinterpretation should be avoided.The frameworks captured a broad range of style information and could perform surprisingly well on niche styles.
- Design guidelines: Subjects should complement styles in abstractness and be interpretable or relevant to the chosen style.
- Design guidelines: Users should receive trigger warnings for pareidolia and offensive content because the models do not acknowledge the possibility of offensive content.
- Implications of borrowing styles: Style keywords can produce shallow summaries, homonym-based misinterpretations, stereotypes, and unwanted noise in artistic traditions.Ukiyo-e generations, for example, often used a limited beige, black, and muted-primary palette despite the style’s broader historical range.
- Limitations and future work: The study’s scope is limited by its focus on text conditioning and, for most experiments, one prompt and the VQGAN+CLIP framework.The authors identify image conditioning, intermediate steering controls, other prompts, and other models as future research directions.
- Limitations and future work: Further work should examine whether generated styles convey conceptual values and messages rather than only surface-level techniques and color palettes.
9 CONCLUSION
The paper presents five experiments examining prompt permutations, random seeds, optimization length, style keywords, and subject–style interactions. It reports category-dependent differences in generation quality and summarizes success and failure modes as design guidelines.
- Five experiments examine prompt permutations, random seeds, optimization length, style keywords, and subject–style keywords.
- Generation quality differed significantly across categories of style and across subject–style combinations.
- The authors summarize successful and failed generations through qualitative analysis and design guidelines.
A EXPERIMENT 1. PROMPT PERMUTATIONS
Experiment 1 tests whether reordering prompt components, adding function words, or including medium terms changes text-to-image generation quality. The tested permutations vary wording while retaining subject and style content.
- The experiment tests whether different phrasings and prompt permutations affect text-to-image generation quality.
- One tested form inserts a medium between the subject and style, motivated by CLIP authors’ claim that medium words may improve generations.The example contrasts “a painting of a dog in the Cubism style” with “dog in the Cubism style.”
- Another permutation places the style and medium before the subject, reflecting a reordering and the model’s reported noun-centered supervision.
- Additional permutations use verbs or function words, including “subject made/done/verb in the style” and “subject with a style style.”
B EXPERIMENT 3. RANDOM SEEDS
This experiment broadens the subject set across abstract and concrete concepts to examine how subject keywords behave in text-to-image prompting. The listed subjects span emotions, ideas, objects, actions, natural elements, and people.
- The study expands its subjects across a broad range of concreteness values, including abstract concepts and concrete entities.
- Abstract subjects include love, progress, relaxation, loyalty, beauty, freedom, chaos, success, nostalgia, intelligence, and fear.
- Concrete subjects include cars, houses, apples, mountains, oceans, forests, flowers, fish, birds, snakes, boys, women, eyes, and computers.
C EXPERIMENT 5. SUBJECTS
A pilot for Experiment 5 found that subject-only prompts were too underconstrained to produce evaluable generations. Adding an aesthetic grounding was therefore necessary for evaluation.
- Subject-only keyword dimensions were insufficient in the Experiment 5 pilot.
- The underconstrained generations were too poor to evaluate.
- Without aesthetic grounding, the generations lacked a basis for evaluation.