Source-linked AI summary

FocusGen: Expanding Visual Design Exploration with a Simulated Focus Group of Persona Agents

Jaewon Choi, Helena Vasconcelos, Hyun Lee, Carolyn Zou, Tak Yeon Lee, Michael Bernstein

arXiv:2608.28001v1cs.HCcs.MA

TL;DR

Current visual exploration tools derive diversity from designers’ own inputs, limiting access to directions outside their existing conceptual frames. FocusGen uses parallel persona agents with elicited preferences to generate audience-conditioned concepts, and evaluations report greater diversity and useful discovery alongside stereotyping risks.

  • Problem

    Current text-to-image exploration tools derive diversity from designers’ own inputs, limiting exploration to what they can already articulate or envision.

  • Method

    FocusGen constructs persona agents from demographic data, procedural backstories, and interviews, then independently runs iterative visual-generation loops for each agent.

  • Results

    Persona conditioning increased visual diversity over a generic-assistant baseline, while open-ended interviews increased diversity for both human and synthetic cohorts.

  • Takeaways & Limitations

    A study with 16 creative professionals found that FocusGen supported discovering unanticipated directions, overcoming fixation, and probing audience contexts.

  • Takeaways & Limitations

    The system’s demographic grounding risks stereotyping, and its LLM-generated backstories received no formal bias audit.

Abstract

from arXiv · show

Creative professionals rarely design for themselves--they design for audiences whose preferences they must anticipate. Yet current text-to-image exploration tools derive diversity entirely from the designer's own input--their prompts, their chosen dimensions, their search queries--confining exploration to what the designer already knows to look for. We present FocusGen, an interactive system that introduces external perspectives into visual design exploration through a "virtual focus group" of simulated persona agents. In contrast to prior persona systems in which multiple agents converge as critics on a single evolving artifact, FocusGen uses personas as parallel generators: each agent--constructed from demographic data, a procedurally generated backstory, and aesthetic preferences elicited through interviews--independently drives an iterative generation loop that produces its own visual concept, transforming one design brief into a spectrum of audience-conditioned directions. With real human participants, we confirm that the iterative refinement loop produces outputs people prefer over zero-shot generation. With synthetic agents at scale, we show that persona conditioning yields higher visual diversity than a generic-assistant baseline--measured by CLIP distance and corroborated by human perceptual judgments--and that open-ended preference interviews yield more diverse outputs than structured ones for both human and synthetic cohorts, while also revealing that agent cohorts recover only part of the diversity of comparable human cohorts. A qualitative study with 16 creative professionals suggests FocusGen helps designers discover unanticipated directions, overcome fixation, and probe audience contexts--while surfacing stereotyping risks that we analyze. We position FocusGen as a divergence scaffold for early-stage ideation rather than a substitute for audience research.

1 Introduction

FocusGen addresses the perspective bottleneck in visual design exploration by introducing simulated audience perspectives through persona agents. Its parallel workflow aims to expand alternatives beyond designers’ own conceptual frames while preserving interpretable persona preferences.

  • Current text-to-image exploration is bounded by designers’ prompts, dimensions, and stated intent, limiting discovery to variations they can already articulate.
  • Real audience feedback can surface external perspectives, but focus groups, interviews, and critiques are expensive, slow, and impractical during early ideation.
  • FocusGen uses a virtual focus group of persona agents whose demographic profiles, backstories, and elicited aesthetic preferences condition independent generation loops.
  • The evaluations test whether heterogeneous persona conditioning increases output diversity, without establishing persona faithfulness or demographic representativeness.
  • The system parallelizes persona-driven generation into audience-specific visual alternatives rather than converging agents on one artifact.

2 Related Work

Related work identifies a persistent tension: existing creativity tools can reduce prompting effort yet remain tied to users’ conceptual frames, while persona systems offer scalable external perspectives. FocusGen distinguishes itself by using audience personas as parallel generators that produce many independently authored alternatives.

  • Design Fixation and AI-Induced Homogenization: Design fixation and AI-induced homogenization can narrow idea variety even when generative tools improve individual creative output.
  • User-Bounded Exploration: Prompt engineering and dimensional tools remain constrained by what users can articulate, envision, or encode in their own inputs.
  • User-Bounded Exploration: Automated optimization and iterative refinement reduce prompting burden but continue optimizing user-stated intent with a generic critic rather than a specific audience’s preferences.
  • FocusGen’s Distinction: FocusGen introduces simulated audience perspectives whose feedback may contradict or expand the designer’s framing, supporting divergent rather than convergent ideation.
  • Persona Agents: Existing persona systems provide scalable audience-like feedback, but demographic conditioning requires caution because it can reproduce representational harms.
  • Interaction Topology: Unlike predominantly convergent systems that refine one artifact, FocusGen is parallel-divergent: each persona authors its own visual concept end-to-end.
  • Contribution Boundary: FocusGen composes established persona construction, narrative grounding, and self-refinement techniques into a workflow that tests whether population diversity translates into visual diversity.

3 FocusGen: System Design

FocusGen lets designers select and interview synthetic personas, then generates and refines one visual concept per agent in parallel. The interface exposes both the resulting gallery and the reasoning history behind each concept for audience-oriented exploration.

  • FocusGen is an interactive system for introducing synthetic external audience perspectives into text-to-image design exploration.
  • Persona Initialization: The Agent Bank contains 1000 persona agents initialized with demographic data and procedurally generated backstories.
  • Designer Workflow: Designers select cohorts using age, gender, education, and income filters, then define a design objective and preference interview questions.
  • Preference Elicitation: Interview questions may be specific and orthogonal or broad and open-ended, and the choice affects resulting output diversity.
  • Persona Initialization: Each agent’s identity combines structured demographics with a first-person, procedurally generated backstory, enabling scalable construction without real participant data.
  • Preference Elicitation: Agents answer interview questions in character, and their responses become operative preferences without being claimed to mirror real group members’ preferences.
  • Iterative Image Generation: At each iteration, agents critique generated images against their Persona Profiles and revise prompts; the loop terminates after no deviations remain or N=5 iterations.
  • Feedback Exploration View: The Feedback Exploration View presents final images, agent demographics, refinement histories, critiques, interview responses, and complete profiles.

4 Evaluations

FocusGen’s evaluations test whether iterative refinement improves individual image preferences and whether persona conditioning increases visual diversity. They also document important scope limits: the technical validation does not establish persona-specific alignment, and the diversity comparison does not isolate which persona-profile components drive the gain.

  • Evaluation scope: FocusGen’s evaluation claims cover refinement preference, persona-conditioned diversity, and open-ended elicitation, while explicitly excluding proof of persona-specific alignment.The supplied evaluation overview states the three supported claims and identifies a fourth claim the studies do not support.
  • Evaluation limitations: The technical validation holds the real person’s preference profile fixed and lacks generic-critic and cross-persona controls, so it cannot distinguish persona-specific alignment from generic image improvement.The study used data from 27 Prolific participants to construct digital proxy agents.
  • Technical validation: Participants selected refined images significantly more often than an option-count-adjusted random baseline (χ2(1, N=106) = 5.20, p< .05).Sessions generated up to five images, and participants ranked the sequence through pairwise comparisons.
  • Technical validation: Users’ final chosen images scored more than twice as high as initial zero-feedback images on TrueSkill (M=15.57 versus M=7.30, p< .001).The paired-samples comparison treats TrueSkill as an alignment score for the chosen images.
  • Persona-conditioned diversity: Persona-enabled grids increased CLIP pairwise distance by 58% and dispersion by 33% on average across five tasks, with a 180% Interior Design gain.Participants also significantly preferred the persona-enabled grids as more visually diverse (χ2(2, N=468) = 86.81, p< .001).
  • Evaluation limitations: The diversity comparison establishes a system-level gain but does not identify whether semantics, prompt length, lexical heterogeneity, preference variation, or their combination caused it.The persona condition included substantially more heterogeneous conditioning text than the generic-critic baseline.

4.3 Evaluation 2: Open-Ended Interviews Expand Design Diversity

Open-ended preference interviews produced more diverse visual concepts than structured interviews for both human and synthetic cohorts. However, synthetic agents remained less diverse than comparable humans, limiting the virtual focus group as a substitute for audience research.

  • Open-ended interviews increased CLIP Distance by 30% for humans and 27% for synthetic agents.
  • Open-ended interviews increased Dispersion by 25% for humans and 33% for synthetic agents.
  • Open-ended interviews produced higher CLIP Distance in all five tasks for both cohorts.The sign test yielded p= .031, but its low power makes it a consistency check rather than primary evidence.
  • Agents mirrored the human direction of effect, but their absolute diversity was consistently lower and roughly half as large on Logo, Interior, and Food.For Food under open-ended elicitation, humans reached .21 versus .14 for agents.
  • FocusGen expands exploration beyond solo prompting and generic assistants but does not reproduce the variability of comparable human cohorts.The authors therefore frame it as a fast divergence scaffold for directions worth testing with real audiences, not a census of audience reactions.
  • Structured and open-ended conditions differed in both question format and degree of constraint, so the source of the diversity gains cannot be isolated.The actionable guidance is to use open-ended questions when divergent exploration is the goal.

5 User Study: FocusGen in Professional Workflows

The user study examined how creative professionals experience and use FocusGen’s agent-driven parallel ideation in professional workflows. It focused on the workflow’s perceived value, limitations, and use of diverse agents to explore design spaces and generate novel ideas.

  • An exploratory qualitative study investigated how 16 creative professionals experienced and evaluated agent-driven parallel ideation.
  • The study asked what value and limitations professionals perceived when incorporating agent-driven parallel ideation into their workflows.
  • A second research question examined how professionals used agents and their preferences to explore design spaces and generate novel ideas.

5.1 Participants

The study recruited 16 creative professionals with prior experience using generative AI tools.

  • The participants were 13 female and 3 male creative professionals with an average of 7.75 years of experience.
  • All participants had professional experience with generative AI tools, most commonly ChatGPT and Midjourney.

5.2 Study Task

Participants selected one of six creative ideation tasks and produced multiple distinct concepts with references and rationales to mimic early-stage creative workflows.

  • Participants chose one of six creative ideation tasks, such as fintech logo design or a coffee-festival poster.
  • They produced 3-5 distinct visual concepts with reference images and rationales.The task was designed to mimic early-stage creative workflows.

5.3 Procedure

The study used remote, fixed-order sessions in which professionals first ideated with a preferred tool, then repeated the task with FocusGen. Researchers coded interviews and submitted concepts using open coding, while treating cross-condition comparisons as uninterpretable.

  • Procedure: Participants completed two ideation tasks remotely: a baseline task with a preferred generative AI tool, followed by the same task using FocusGen.Sessions lasted approximately two hours and were recorded with IRB approval.
  • Study Design Rationale: The fixed order let participants anchor reflections in a fresh experience of their established workflow, but prevented interpretable between-condition comparisons.The conditions differed in generation backend and output volume as well as tool familiarity.
  • Analysis: Researchers independently open-coded interview transcripts and submitted concepts before collaboratively grouping codes into reported themes.The analysis combined qualitative coding with participants’ submitted design concepts.

5.4 Results

Participants found FocusGen useful for receiving many visual starting points, validating intuitions, and discovering novel, contextually grounded directions. They also valued demographic and professional-context filters, while recognizing stereotyping risks and positioning the system mainly for divergent ideation.

  • Workflow: FocusGen’s parallel architecture delivered 20–30 visual concepts at once, which participants likened to browsing a mood board.Participants initially needed a brief adjustment period to asking preference questions instead of issuing direct commands.
  • Confidence: Participants reported greater confidence, especially in unfamiliar domains, and used simulated audience outputs as a rapid sanity check for their own ideas.Confidence for unfamiliar domains was rated Q10: M=5.7, SD=1.1.
  • Novel Ideas: FocusGen generated more diverse and contextually aware ideas that pushed participants beyond conventional directions, including caffeine-molecule and coffee-drop concepts.Novel-direction ratings were Q7: M=6.1, SD=0.9 and Q9: M=4.6, SD=1.3.
  • Contextual Relevance: Preference interviews added explanations for generated designs, encouraging participants to engage in critical thinking about the rationale behind each concept.Participants rated the value of these responses Q1: M=4.9, SD=1.3 and Q2: M=4.9, SD=1.1.
  • Demographic Filters: Demographic filters surfaced unexpected conceptual directions, but participants noted that demographic conditioning could also reinforce stereotypes.Perceived demographic impact was Q6: M=4.8, SD=1.2; the Beauty Expo case associated a demographic filter with robot mascots.
  • Ideal Role: Participants viewed FocusGen as most useful when creatively stuck and for exploring diverse ideas rather than refining one concept.The reported ratings were Q10: M=5.7, SD=1.1 and Q9: M=4.6, SD=1.3.

6 Discussion

FocusGen expands design alternatives through simulated audience perspectives, but its outputs should be treated as exploratory provocations rather than faithful representations of real people. The discussion identifies ethical, validation, cultural, and convergence limitations alongside a curatorial role for designers.

  • Contribution: Evaluations indicate that persona conditioning expands design diversity beyond a generic assistant, while qualitative findings show professionals using that expanded spectrum in practice.The discussion explicitly separates demonstrated option-set diversity from untested persona faithfulness.
  • Ethical Considerations: Demographic grounding and LLM-generated backstories can reproduce stereotypes, and the proposed mitigation remains partial and unaudited.The paper contrasts these synthetic backstories with agents grounded in interviews with real people.
  • Ethical Considerations: The Beauty Expo example shows that surprising demographic associations can attract adoption, making designer curation an imperfect safeguard against stereotypes.The same surprisingness that reveals blind spots can make a stereotyped association appealing.
  • Ethical Considerations: The paper recommends provenance and friction, careful framing, bias auditing, and validation against real demographic samples before broader deployment.Until those steps occur, FocusGen should be claimed to diversify designers’ option sets, not represent audiences.
  • Creative Agency: FocusGen is positioned as an upstream ideation tool in which designers retain agency by curating, interpreting, and developing generated starting points.Participants’ reports of increased critical thinking support this framing, while longer-term ownership remains for future study.
  • Broader Implications: The system proposes polyphonic co-creation by presenting conflicting persona perspectives simultaneously instead of optimizing a single response.This extends crowd-powered design into an in silico setting.
  • Meaningful Diversity: Narrative grounding makes persona-driven novelty more actionable than stochastic variation by giving unexpected ideas a plausible semantic rationale.The paper illustrates this with an engineering agent suggesting a caffeine molecule.
  • Validation Gaps: Synthetic personas lack lived experience, leaving untested whether outputs are discriminably tied to specific personas or representative of sampled demographic groups.The proposed next steps are persona-faithfulness testing and validation against real demographic preferences.

7 Conclusion

FocusGen brings simulated audience perspectives into visual exploration by having independently conditioned persona agents generate alternatives. The conclusion presents it as a tool for discovering directions and building confidence, while reserving claims about representativeness.

  • Contribution: FocusGen constructs personas from demographic data, procedural backstories, and elicited preferences, then independently generates profile-conditioned visual concepts.The system converts one design brief into audience-conditioned alternatives.
  • Findings: In a study with 16 creative professionals, FocusGen helped participants discover unanticipated directions, overcome fixation, and build confidence through simulated audience feedback.Examples included caffeine molecules on coffee posters and robots in beauty mascots.
  • Scope: The simulated perspectives expanded designers’ option sets, but the evaluation did not establish faithfulness to the specific people or demographic groups modeled.The conclusion explicitly leaves representativeness as an open question.
  • Implication: The broader contribution is a shift from engineering prompts for one AI toward curating a diverse, persona-driven virtual focus group.The system is framed as a way to surface designers’ blind spots.
Loading 2608.28001v1…