Source-linked AI summary

AI-Augmented Brainwriting: Investigating the use of LLMs in group ideation

Orit Shaer, Angelora Cooper, Osnat Mokryn, Andrew L. Kun, Hagit Ben Shoshan

arXiv:2402.14978v2cs.HCcs.AIcs.CY

TL;DR

The paper asks how LLMs can support group ideation across idea generation and evaluation. It introduces and studies a collaborative group-AI Brainwriting framework plus an LLM evaluation engine, finding evidence that LLMs could enhance ideation and assist idea evaluation.

  • Problem

    The paper addresses limited knowledge about the merits and limitations of integrating LLMs into collaborative ideation.

  • Method

    The authors study a group-AI Brainwriting framework and compare GPT-4 idea ratings with ratings from expert and novice evaluators.

  • Results

    GPT-4 rankings showed a moderate positive linear relationship with human rankings, supporting its potential for preliminary idea filtering.

  • Takeaways & Limitations

    The findings provide evidence that LLMs could support both ideation and preliminary idea evaluation in collaborative Brainwriting.

  • Takeaways & Limitations

    The study examined novice designers using one Brainwriting process, one problem statement, and an HCI-education context.

Abstract

from arXiv · show

The growing availability of generative AI technologies such as large language models (LLMs) has significant implications for creative work. This paper explores twofold aspects of integrating LLMs into the creative process - the divergence stage of idea generation, and the convergence stage of evaluation and selection of ideas. We devised a collaborative group-AI Brainwriting ideation framework, which incorporated an LLM as an enhancement into the group ideation process, and evaluated the idea generation process and the resulted solution space. To assess the potential of using LLMs in the idea evaluation process, we design an evaluation engine and compared it to idea ratings assigned by three expert and six novice evaluators. Our findings suggest that integrating LLM in Brainwriting could enhance both the ideation process and its outcome. We also provide evidence that LLMs can support idea evaluation. We conclude by discussing implications for HCI education and practice.

1 INTRODUCTION

The paper investigates whether LLMs can enhance both divergence and convergence in collaborative Brainwriting. It proposes and evaluates a group-AI framework alongside an LLM-based idea evaluation engine.

  • Group Brainwriting can produce fewer alternatives because peer judgment, free riding, and production blocking constrain collaborative ideation.
  • Brainwriting uses parallel individual idea generation followed by sharing and elaboration, often producing more ideas than face-to-face brainstorming.
  • The study examines LLM integration in both idea generation and idea evaluation and selection.
  • The authors evaluate a collaborative group-AI Brainwriting framework with 16 undergraduate students using qualitative and quantitative methods.
  • The evaluation engine rates ideas for relevance, innovation, and insightfulness, then compares its ratings with three expert and six novice evaluators.
  • The paper reports evidence that LLM integration could enhance ideation and its outcome, while LLMs can assist users in evaluating ideas.

2 RELATED WORK

Prior work frames Brainwriting as a structured alternative to brainstorming and positions LLM-assisted ideation within broader human-AI co-creation and idea-evaluation research.

  • Structured Approaches to Ideation: Group ideation can expose participants to diverse perspectives, but systems must select and present creative and diverse ideas effectively.
  • Structured Approaches to Ideation: Brainwriting addresses barriers in group brainstorming through parallel idea generation, sharing, and subsequent elaboration.
  • Structured Approaches to Ideation: Brainwriting often produces more quality ideas than face-to-face brainstorming, although the process should be adapted to group characteristics.
  • Human-AI Co-Creation: Existing work reports opportunities and challenges when experienced designers use GPT-3 and DALL-E for creative design and ideation.
  • Human-AI Co-Creation: This study extends prior work by examining how novice designers interact with and perceive ideas co-created with LLMs.
  • Approaches for Evaluating Ideas: AI-based evaluation may increase evaluation speed and provide feedback within human-AI creative teams.
  • Approaches for Evaluating Ideas: The study applies an LLM evaluation engine to written ideas generated by teams comprising humans and another LLM.

3 COLLABORATIVE GROUP-AI BRAINWRITING FRAMEWORK DESIGN

The proposed framework places LLMs after initial human Brainwriting to support divergence, then uses structured criteria and GPT-4 to support convergence.

  • Framework Process: During divergence, participants first generate individual ideas, then review collective ideas while prompting an LLM for additions.
  • Framework Process: During convergence, groups discuss and narrow the idea list before further developing selected ideas with LLM assistance.
  • Phase 1: Brainwriting using Conceptboard: Teams use a shared Conceptboard, writing ideas in parallel and iterating until each member contributes at least six ideas.
  • Phase 2: Collaborative ideation with LLM: The divergence phase uses GPT-3 to generate additional ideas that are added to a dedicated collaborative-ideas area.
  • Phase 3: LLM-powered evaluation: GPT-4 produces three criterion ratings and explanatory text for each evaluated idea.
  • Phase 3: LLM-powered evaluation: The evaluation engine uses relevance, innovation, and insightfulness to assess early-stage ideas.

4 USER STUDY: COLLABORATIVE GROUP-AI BRAINWRITING PROCESS

A 16-student course session tested GPT-3-assisted Brainwriting for idea generation and GPT-4-assisted evaluation and selection. Students rated generated ideas highly across all three criteria.

  • Study Procedure: The study ran a 70-minute Brainwriting session in which students generated ideas independently, co-created with GPT-3, and selected ideas for further development.
  • Convergence: Students evaluated ideas using relevance, innovation, and insightfulness on a Likert scale before selecting a small final set.
  • Participants and Setting: Sixteen students worked in five teams on a tangible user-interface design problem.
  • Divergence: Human-generated ideas averaged 16.5 words, while GPT-3-generated ideas averaged 20.9 words.
  • Convergence: Mean self-ratings were 4.75 for relevance, 4.45 for innovation, and 4.45 for insightfulness.
  • Convergence: Sixty percent of rating questions received the maximum score of 5 out of 5.

5 FRAMEWORK EVALUATION

The framework evaluation examines whether LLMs enhance divergence-stage group Brainwriting and whether they can assist convergence-stage idea evaluation. It combines qualitative and quantitative analyses of ideation, solution spaces, and evaluator ratings.

  • Evaluation aims: The evaluation addresses two research questions: whether LLMs improve group Brainwriting ideation and whether they can assist idea evaluation during convergence.The first question covers ideation process and outcomes; the second concerns LLM assistance in evaluating ideas.
  • Evaluation aims: Idea quality was rated by students, a GPT-4 evaluation engine, three HCI expert reviewers, and six novice designers using the same dimensions.The reviewers rated ideas on the same dimensions as the GPT-4 engine, enabling comparison across evaluator groups.
  • Evaluation aims: Divergence was evaluated by comparing the semantic distributions of human-generated and GPT-3-generated ideas and identifying terms unique to their solution spaces.The analysis examined both semantic distributions and distinctive vocabulary in the two spaces.

5.1 Data and Methods

The study collected ideation artifacts, prompts, reflections, and evaluator ratings, then combined thematic, semantic, topic-modeling, LPA, and statistical analyses. Student reflections reported both benefits and shortcomings of GPT-3, including expanded viewpoints, prompt difficulties, and redundancy.

  • Data collection: The dataset included team ideas, GPT-3 prompts, student reflections, and novice, expert, and GPT-4 ratings.These materials supported analyses of ideation experiences, idea quality, and evaluator agreement.
  • Evaluator study: Ratings were collected from six novice designers and four expert reviewers, but results reported only the three experts who evaluated all student-produced ideas.Ideas were presented in random order without identifying whether they came from humans or GPT-3.
  • Qualitative analysis: Thematic analysis grouped student reflections and GPT-3 interaction prompts into keywords, tags, broad themes, and categories.Researchers first identified common keywords and tags before aggregating them into broader themes.
  • Semantic analysis: Semantic divergence analysis used spaCy to extract nouns and adjectives, Gensim for topic modeling, and Domain-based Latent Personal Analysis.LPA compared normalized term-frequency representations for human and GPT-3 idea spaces using symmetric KLD.
  • Semantic analysis: LPA identifies document-signature terms by comparing each document’s normalized term frequencies with corpus frequencies using symmetric KLD.Positive weights indicate rare corpus terms overused in a document, while negative weights indicate popular corpus terms that are underused or missing.

5.2.2 Ideation outcomes.

All five teams selected project ideas that incorporated GPT-3 contributions, most often by merging human-generated and GPT-3-generated ideas. The selected outcomes therefore reflected combined human–AI ideation rather than exclusively human-originated ideas.

  • Selected ideas: 3 out of 5 chosen ideas were developed by merging a human-generated idea with a GPT-3-generated idea.These ideas combined contributions from both sources during development.
  • Selected ideas: One chosen idea merged multiple human-generated ideas with multiple GPT-3-generated ideas, while another was based solely on GPT-3-generated material.The five outcomes therefore included varied patterns of human and GPT-3 contribution.

5.2.3 Exploring the Human and LLM solution spaces.

Human and GPT-3 idea spaces showed substantial semantic overlap but also distinct emphases. Human ideas tended toward broader or more abstract concepts, whereas GPT-3 ideas more often used concrete, detailed, and object-specific terminology.

  • Semantic clustering: Human-only clusters included vehicles, clothing, food and beverages, learning and information, and games and entertainment.These clusters contained terms such as bus, jacket, dining, study, Pokemon, and music.
  • Semantic clustering: GPT-3-only clusters emphasized screen and display elements, interaction controls, measurements, visual design, and work-related terms.Examples included background, buttons, gestures, diameter, shapes, signs, brainstorming, and distractions.
  • Conceptual differences: The authors characterize human-only concepts as more abstract or generalized and GPT-3-only concepts as more concrete or detail-specific.The difference concerns the level of detail in the concepts rather than a complete separation between the solution spaces.
  • LPA terminology: The ten most prevalent terms across the ideas included user, device, light, people, sound, surface, task, wrist, pillow, and day.LPA further distinguished term usage: GPT-3 favored users, device, surface, light, posture, and wrist, while human ideas favored people, wearable, screen, work, time, space, interface, day, and app.

5.2.4 Prompt analysis.

Students used GPT-3 through broad-area and solution-specific prompts, then expanded selected ideas with usage- and detail-focused follow-ups. The collaboration broadened and reshaped ideas, while also revealing recurring concerns about redundancy, creativity, and prompt-crafting difficulty.

  • Prompt analysis: Students typically began with either broad-area prompts requesting open-ended ideas or solution-specific prompts addressing a concrete problem.They combined these approaches during ideation.
  • Prompt analysis: After selecting an idea, teams used usage-focused prompts to explore context and detail-focused prompts to expand features and capabilities.Examples addressed how a device would be used and what functionalities it could provide.
  • Student experience: 44% of students reported that GPT-3 provided a unique or expanded perspective on the problem and possible solutions.Students also described the system as helping them reframe, refine, and elaborate on concepts.
  • Student experience: 50% of students later said GPT-3 contributed to their project, while 50% perceived it as helpful during ideation.The reported contributions included generating new ideas, adding characteristics, and tackling particular challenges.

5.3.1 Consistency of the GPT-4 evaluation engine.

The GPT-4 evaluation engine produced internally consistent ratings across relevance, innovation, and insightfulness. Its ratings were generally high, while human evaluators differed in distribution and agreement, and rankings showed moderate correspondence across groups.

  • Internal consistency: Fleiss’ Kappa exceeded 0.4 for GPT-4 across all three criteria, indicating consistent repeated evaluations.The criterion-specific values were 0.42 for Relevance, 0.40 for Innovation, and 0.49 for Insightfulness.
  • Human evaluators: Experts were more critical than novices, while expert and novice evaluations showed diverging opinions and medium to low internal consistency.The expert rating distributions were non-normal across all three criteria, with p < 0.001.
  • Ranking agreement: GPT-4 placed most expert-selected top ideas in the top half and three of the experts’ four lowest-ranked ideas among its bottom six.Two of the experts’ four top ideas were also rated at the top by novices, while novice and GPT-4 rankings showed high agreement at the extremes.
  • Ranking agreement: 0.556 was the Pearson correlation between expert and GPT-4 rankings, compared with 0.547 between novice and GPT-4 and 0.602 between expert and novice rankings.These coefficients indicate moderate positive relationships among the ranked lists.

5.3.3 Summary of findings for RQ2.

GPT-4 rankings showed moderate agreement with expert and novice rankings and consistently rated the teams’ chosen ideas above average. The authors therefore identify GPT-4 as a potential tool for preliminary idea filtering, while noting that its evaluations were conducted after ideation.

  • Evaluation consistency: GPT-4 evaluations were internally consistent, with Fleiss’ Kappa values exceeding 0.4 across Relevance, Innovation, and Insightfulness.Expert and novice evaluators instead had diverging opinions and medium to low internal consistency.
  • Ranking alignment: GPT-4’s ranking showed a moderate positive relationship with human rankings, with correlations of 0.556 versus experts and 0.547 versus novices.The expert–novice correlation was 0.602.
  • Implications: The authors identify GPT-4 as a potential tool for preliminary idea filtering because its rankings aligned with human judgment on high-quality ideas.This conclusion is based on consistency across human and AI evaluations.
  • Implications: None of the ideas ultimately chosen by teams received a low GPT-4 rating, suggesting that GPT-4 would not have filtered out those selected ideas.The authors also report that ideas rated low by GPT-4 were not ultimately chosen.
  • Scope: The GPT-4 evaluations were conducted only after ideation sessions ended, so teams did not receive this feedback during ideation.The filtering implications were therefore inferred from post hoc evaluation rather than an intervention with teams.

6 DISCUSSION

The discussion finds that integrating LLMs into group Brainwriting can support both idea generation and idea evaluation, while identifying limits involving creativity, prompting, bias, and generalizability.

  • Scope and approach: The framework examines LLM support for both divergent idea generation and convergent idea evaluation in collaborative group Brainwriting.The study evaluates the process, its outcomes, and LLM-based evaluation in a college-level interaction design course.
  • Divergence-stage findings: GPT-3 contributed ideas that differed somewhat from human ideas and included more technical and usage details.Students also reported that GPT-3 provided unique or expanded perspectives on the problem and possible solutions.
  • Divergence-stage findings: All 5 teams selected project ideas that either combined GPT-3-generated and human-generated ideas or were based on a GPT-3-generated idea.The findings suggest that LLM integration supported both divergent thinking and incremental, step-by-step development of solution details.
  • Limitations and future directions: About 30% of students reported that GPT-3 tended to be redundant and lacked creativity, while students also struggled to create effective prompts.The discussion proposes prompt engineering, conceptual blending, and multiple personas as possible directions for improving LLM contributions.
  • Convergence-stage findings: GPT-4 gave relatively high ratings to all ideas ultimately chosen by student teams, and no chosen idea received a low GPT-4 rating.This suggests GPT-4 feedback would not have filtered out ideas the teams considered good.
  • Limitations and future directions: The authors warn that co-creation processes might embody and amplify human social biases, although this study did not identify specific social biases in its ideas.They recommend probing for biased ideas and developing methods to filter them in future work.
  • Limitations and future directions: The study’s scope is limited to novice designers using Brainwriting for one problem statement in HCI education, without examining long-term effects.The findings may not generalize to expert designers, expert LLM users, trained prompt engineers, other innovation domains, or other educational disciplines.

7 CONCLUSION

The paper explores LLM-supported collaborative Brainwriting as one potential form of human–machine collaboration, focusing on educational ideation. Its findings suggest benefits for idea generation and evaluation, while emphasizing explainability and bias safeguards.

  • LLMs can broaden the topics explored by teams and enhance ideas generated through Brainwriting.
  • LLM-based evaluations show promise for identifying both good and poor ideas during Brainwriting.
  • Such evaluations could provide useful feedback to teams as they work through the Brainwriting process.
  • The evaluation system must provide explainable feedback and avoid propagating biases from human-generated data.
Loading 2402.14978v2…