Source-linked AI summary

Large-scale Text-to-Image Generation Models for Visual Artists' Creative Works

Hyung-Kwon Ko, Gwanmo Park, Hyeon Jeon, Jaemin Jo, Juho Kim, Jinwook Seo

arXiv:2210.08477v3cs.HC

TL;DR

Visual artists lacked research on how to adopt rapidly advancing LTGMs in creative work. The paper combines a systematic review of 72 papers with interviews of 28 artists across 35 visual-art domains. It identifies automation, exploration, and mediation as major roles, while reporting difficulties with current integration and proposing four design guidelines.

  • Problem

    Little research examined how visual artists could leverage LTGMs in creative work beyond technical issues such as prompt engineering.

  • Method

    The study combines a systematic literature review of 72 papers with interviews of 28 visual artists covering 35 distinct visual-art domains.

  • Results

    LTGMs support automation, exploration, and mediation, but visual artists found them hard to incorporate into creative work in their current form.

  • Takeaways & Limitations

    The paper proposes four guidelines for LTGM interfaces: variability specification, domain-specific customization, multimodal controllability, and prompt-engineering assistance.

  • Takeaways & Limitations

    Artists reported that LTGMs generate predictable images and struggle with novel, philosophical, storytelling, or sophisticated interpretations.

Abstract

from arXiv · show

Large-scale Text-to-image Generation Models (LTGMs) (e.g., DALL-E), self-supervised deep learning models trained on a huge dataset, have demonstrated the capacity for generating high-quality open-domain images from multi-modal input. Although they can even produce anthropomorphized versions of objects and animals, combine irrelevant concepts in reasonable ways, and give variation to any user-provided images, we witnessed such rapid technological advancement left many visual artists disoriented in leveraging LTGMs more actively in their creative works. Our goal in this work is to understand how visual artists would adopt LTGMs to support their creative works. To this end, we conducted an interview study as well as a systematic literature review of 72 system/application papers for a thorough examination. A total of 28 visual artists covering 35 distinct visual art domains acknowledged LTGMs' versatile roles with high usability to support creative works in automating the creation process (i.e., automation), expanding their ideas (i.e., exploration), and facilitating or arbitrating in communication (i.e., mediation). We conclude by providing four design guidelines that future researchers can refer to in making intelligent user interfaces using LTGMs.

1 INTRODUCTION

LTGMs generate high-quality images from multimodal inputs, but their rapid development has left visual artists uncertain about integrating them into creative work. This study combines a literature review and artist interviews to characterize LTGMs’ roles and propose interface guidelines.

  • LTGMs and visual artists: LTGMs use text or images as multimodal input to generate high-quality images in a zero-shot fashion.DALL-E was trained on 250 million text-image pairs and showed generalizable performance on downstream tasks.
  • Research motivation: Prior research emphasized LTGMs’ technical properties, while little research examined their applicability to visual artists’ creative work.
  • Study overview: The study reviews 72 system and application papers and interviews 28 visual artists across varying domains.The interview study uses three themes identified in the literature review: user, task, and role.
  • Findings: LTGMs support automation of creation, exploration of ideas, and mediation in communication, but artists find them difficult to incorporate in their current form.
  • Design implications: The paper proposes four guidelines: specify variability, customize models by domain, improve multimodal controllability, and assist prompt engineering.

2 BACKGROUND AND RELATED WORK

LTGMs have become increasingly capable and accessible, creating opportunities for visual-artist support while raising questions about their broader impact on creative practice. The paper therefore shifts attention from technical properties toward how artists may leverage these models.

  • Scope: The background introduces LTGMs and their democratization before examining AI models as tools for visual artists’ creative work.
  • LTGM capabilities: DALL-E contains 12 billion parameters and demonstrated image-generation performance on unseen datasets using quality metrics and human evaluators.
  • Democratization: Open datasets, source code, model weights, prompt marketplaces, image search engines, and brushing interfaces have expanded public access to LTGMs.
  • Related work: Existing AI systems support visual domains including graphic design, UI design, webtoons, digital art, and new media art.
  • Research gap: Prior work discussed applications and implications, but little research examined LTGMs’ potential to change how future visual artists work.

3 SYSTEMATIC LITERATURE REVIEW

The systematic literature review broadly collected HCI publications on generative models, narrowed them through exclusion criteria, and selected 72 papers for detailed analysis. The review focuses on how generative models support human tasks rather than on technical model development alone.

  • Corpus construction: The review examined 72 system and application papers from six prominent HCI venues.The venues were CHI, UIST, IUI, DIS, CSCW, and C&C.
  • Scope: Because few HCI papers directly focused on LTGMs, the review included broader generative-model research as a related category.
  • Search strategy: Search terms included generate, generation, generative, and generating to identify relevant papers beyond the exact phrase “generative model.”
  • Exclusion criteria: The review excluded papers misusing generative terminology, qualitative studies, non-neural generative models, technical applications unrelated to human impact, and several publication types.
  • Exclusion criteria: The final analysis excluded duplicate or substantially overlapping publications by the same authors.
  • Screening procedure: The search began with 1,494 papers, deduplication left 1,429, 64 passed title-and-abstract screening, and 8 additional papers produced the final set of 72.

3.2 Analysis Procedure

The analysis procedure transformed the reviewed literature into a coding framework through independent, iterative qualitative analysis. The resulting framework organized publications by users, tasks, and generative-model roles.

  • Coding framework: The researchers used a question rubric to examine each paper’s model role, research questions or contributions, target users, and supported tasks.
  • Analysis procedure: Two authors independently generated low-level codes, revised them collaboratively through three iterations, and four authors clustered codes into themes.

3.3 Results

The literature review organized generative-model research around users, tasks, and model roles. It identified broad user and task coverage, with automation, exploration, mediation, co-work, and representation as the main roles.

  • Themes: The review identified 9 user codes, 8 task codes, and 5 generative-model role codes.The themes were user, task, and role.
  • User: General public or unspecified users appeared in 33 publications (45.8%), followed by designers in 13 publications (18.1%).Artists were targeted in 5 publications (6.9%).
  • Task: Writing and design were each targeted in 13 publications (18.1%), while everyday-work appeared in 11 publications (15.3%).Drawing appeared in 10 publications (13.9%).
  • Role: Exploration was the most frequent role at 32 publications (44.4%), followed by automation at 29 publications (40.3%).Mediation accounted for 10 publications (13.9%), while co-work and representation were also identified.

4 VISUAL ARTIST INTERVIEWS

The authors interviewed 28 professional visual artists across varied occupations and domains to examine how they might adopt LTGMs. The semi-structured interviews combined participant experience, LTGM exposure, and discussion of creative-work applications.

  • Research questions: The interviews addressed which artist subgroups would use LTGMs, which tasks they would support, and what roles LTGMs would play.These questions were framed as RQ1, RQ2, and RQ3.
  • Participants: The study included 28 visual artists covering 35 unique visual art domains.Participants had at least one year of work experience and had not previously experienced LTGMs.
  • Procedure: Remote semi-structured interviews lasted 67 minutes on average, ranging from 52 to 88 minutes.Sessions included demographics, experiencing DALL-E, and discussing LTGM-supported creative work.
  • Procedure: Participants answered questions about jobs, recurring tasks, inspiration, references, communication, and collaboration.The interview also asked about LTGM adoption, differences from previous tools, unsupported work, desired functionality, and changes to working paradigms.
  • Analysis: Analysis combined deductive coding from the literature-review themes with inductive coding refined through discussion until consensus.Dovetail was used to organize and aggregate the transcribed data.

5 FINDINGS

The findings adapted the literature-review codebook to visual artists by subdividing user and task categories into more detailed characteristics. The authors also reported four limitations of LTGMs for artists’ workflows.

  • Findings framework: The interview analysis categorized findings using three themes: user, task, and role.The visual-art focus required modifying the codebook used in the literature review.
  • Findings framework: The literature review’s artist and designer user codes, and drawing and design task codes, were subdivided into more detailed characteristics.This adaptation supported closer analysis of visual artists’ use cases.
  • Limitations: The authors presented four limitations that LTGMs cannot properly support in visual artists’ workflows.The passage identifies these as limitations of current workflow support without enumerating them here.

5.1 Potential of LTGMs

Artists described LTGMs as useful for reference search, ideation, fast visual communication, unconventional exploration, prototyping, and fine-art justification. These uses emphasize exploration and mediation across creative and collaborative work.

  • Image reference search: 12 of 28 interviewees viewed LTGMs as new image-reference search tools for learning, inspiration, and realizing imagination.Artists praised LTGMs as fast, convenient, and able to generate many unique, high-quality images.
  • Real-time visual communication: 20 artists considered LTGMs beneficial for real-time visual communication in top-down workplaces, client services, or cross-domain collaboration.Fast prototyping was especially valued for verifying and improving work with supervisors.
  • Real-time visual communication: LTGMs could generate starting pieces for discussions when art directors’ directions were vague or when teams explored alternatives.Concept artists described using generated pieces to begin initial conversations and branch selected concepts further.
  • Real-time visual communication: LTGMs could help clients and artists communicate ambiguous or paradoxical requirements through visual materials.One product designer said they could reduce the time needed for the overall communication process.
  • Unconventional creation: 14 artists said LTGMs could support unconventional experimentation, with 8 describing them as less constrained by human creative biases.Artists wanted to use unexpected outputs to leave their preferences and comfort zones.
  • Prototyping: LTGMs could provide low-fidelity prototypes for artists with limited tool expertise or knowledge of related art fields.This was identified by 9 visual artists as useful for reducing barriers to acquiring new software skills.
  • AI-art justification: In fine art, LTGMs could create artworks and provide justification for subjective concepts through AI-generated results.One artist proposed iterative generation of many figures and pairings for a painting project.

5.2 Limitations of LTGMs

Artists identified limitations in LTGMs’ predictability, weak domain-specific understanding, lack of personalization, and dependence on text prompting. These constraints made outputs impractical, insufficiently expressive of artists’ identities, or unable to support genuinely novel ideas.

  • 5.2.1 LTGMs only generate predictable images.: 5 visual artists described LTGMs as predictable machines that could not provide the unexpected images needed for inspiration and philosophical interpretation.P8 expected philosophical reasoning from “why is an apple red?” but received an ordinary red apple.
  • 5.2.1 LTGMs only generate predictable images.: 11 visual artists doubted that LTGMs could generate images requiring domain-specific understanding, including user experience, detailed design requirements, and constructability.Artists viewed LTGMs as capable of rough artifact sketches but unable to incorporate all practical requirements.
  • 5.2.2 LTGMs Do Not Support Personalization.: Artists worried that delegating all work to LTGMs would prevent them from adding personal identity or producing artifacts that reflected their own styles.P17 feared generated images would carry recognizable LTGM characteristics and appear stale.
  • 5.2.3 Text Prompting Restrains Creativity.: 9 visual artists said text prompting prevents novel image generation when new concepts lack established words or artistic processes cannot be translated into sentences.P8’s provisional brand name “Mohei” produced an irrelevant image, while P10 found text inadequate for describing abstract painting processes.
  • 5.2 Limitations of LTGMs: 6 artists found learning to use LTGMs burdensome, while 10 considered them inefficient when they already had specific goals and requirements.Repeated prompting could require substantial effort without guaranteeing a satisfactory result.

6 DESIGN GUIDELINES

The authors propose interface guidelines that address LTGMs’ limited variability, domain understanding, controllability, and text-prompt usability. The guidelines tailor generation to artists’ objectives while supporting richer interaction and customization.

  • Variability level specification: Variability controls should offer Lookup, Inspiration, and Reinterpretation levels matched to artists’ motivations and objectives.Lookup favors logical connections useful in applied art, while Reinterpretation is suggested for fine-art contexts.
  • Model customization: Model customization should reflect domain-specific priorities, such as fine-tuning on contemporary art images to support philosophical meanings.The guideline uses LTGMs’ scalable computational capacity to adapt image generation to downstream domains.
  • Multi-modal controllability: Interfaces should give artists more control because some creative situations cannot be generated adequately with only a few sentences.The authors connect controllability to artists’ serious engagement with interaction between themselves and the artifact.
  • Multi-modal controllability: Multi-modal inputs such as hand gestures, voice, and rough drawings should complement text when imagery is vivid or abstract.A rough drawing could be converted into a more complete and sophisticated image.
  • Prompt engineering: Prompt-engineering tools should help artists express intended images and reduce exhaustion caused by structuring prompts for the machine.The guideline responds to difficulties finding proper expressions and to outputs that fail to meet expectations.

7 LIMITATIONS AND FUTURE WORK

The study identifies limitations in participant sampling and scope, including unexamined social impacts and concerns among people affected by rapid technological change.

  • Most interview participants were in their 20s and 30s, potentially limiting the practical shortcomings identified for real-world adoption.The authors suggest that recruiting more senior visual artists with over 30 years of domain experience might reveal additional shortcomings.
  • The study did not examine the social impacts of LTGMs, including possible use of artworks without permission in training datasets.
  • Some people outside the study’s focus expressed concerns about rapid technological advancement, including serious fears of job loss.
  • Future work will study polarization in adopting new technologies and possible HCI solutions for people needing support.

8 CONCLUSION

The paper investigates how visual artists would adopt LTGMs in creative work through a literature review and interviews. It reports that LTGMs can support creation, ideation, and communication in diverse ways.

  • The study examines how visual artists would adopt LTGMs to support their creative works.
  • The authors reviewed 72 generative-model papers in HCI and interviewed 28 visual artists across 35 unique visual art domains.
  • LTGMs can automate the creation process, expand artists’ ideas, and facilitate or arbitrate communication.
Loading 2210.08477v3…