Source-linked AI summary

Exploring Collaboration Patterns and Strategies in Human-AI Co-creation through the Lens of Agency: A Scoping Review of the Top-tier HCI Literature

Shuning Zhang, Hui Wang, Xin Yi

arXiv:2507.06000v2cs.HC

TL;DR

Human-AI co-creation research lacks a systematic synthesis linking agency configurations, control mechanisms, and contexts. This scoping review analyzes 134 HCI/CSCW papers, producing an integrated framework, operational catalog, cross-context map, and grounded guidance for future co-creative systems.

  • Problem

    Human-AI co-creation research lacks a systematic synthesis mapping agency configurations and operational control mechanisms across contexts.

  • Method

    The paper conducts a scoping review of 134 papers from prominent HCI venues and synthesizes agency patterns, control mechanisms, contexts, and applications.

  • Results

    The review produces an integrated analytical framework, a catalog of operational control mechanisms, and a cross-context map linking agency configurations to co-creative practices.

  • Takeaways & Limitations

    The synthesis provides grounded guidance for designing future co-creative systems and for CSCW discussions of trust, ethics, and long-term implications.

  • Takeaways & Limitations

    Existing agency theories provide limited models of dynamic role and control transitions, and general theories fit open-ended co-creative goals less clearly.

Abstract

from arXiv · show

As Artificial Intelligence (AI) increasingly becomes an active collaborator in co-creation, understanding the distribution and dynamic of agency is paramount. The Human-Computer Interaction (HCI) perspective is crucial for this analysis, as it uniquely reveals the interaction dynamics and specific control mechanisms that dictate how agency manifests in practice. Despite this importance, a systematic synthesis mapping agency configurations and control mechanisms within the HCI/CSCW literature is lacking. Addressing this gap, we reviewed 134 papers from top-tier HCI/CSCW venues (e.g., CHI, UIST, CSCW) over the past 20 years. This review yields four primary contributions: (1) an integrated theoretical framework structuring agency patterns, control mechanisms, and interaction contexts, (2) a comprehensive operational catalog of control mechanisms detailing how agency is implemented; (3) an actionable cross-context map linking agency configurations to diverse co-creative practices; and (4) grounded implications and guidance for future CSCW research and the design of co-creative systems, addressing aspects like trust and ethics.

1 Introduction

AI’s shift from tool to collaborative partner creates a need to understand how agency is distributed, operationalized, and negotiated in human-AI co-creation. This review addresses fragmented evidence by synthesizing agency patterns, control mechanisms, and contexts across 134 papers.

  • AI’s transition from static-task tool to collaborative partner raises challenges concerning initiative, control, and agency distribution in CSCW.
  • Existing agency frameworks and empirical studies provide fragmented or context-specific accounts and lack granular mappings of operational control mechanisms.
  • The review analyzes 134 papers from prominent HCI venues published over the past 20 years.
  • Its analytical framework integrates interaction context, agency patterns, and operational control mechanisms from system-level and user-centric perspectives.
  • The synthesis identifies agency forms, catalogs control mechanisms across interaction stages, and links recurring strategies with application domains and collaboration dynamics.

2 Related Work

Prior HCI/CSCW work offers important theories, models, and situated studies of agency, autonomy, and collaboration, but lacks an integrated cross-context account of agency patterns and operational control.

  • Agency research spans psychology, sociology, and AI, while HCI/CSCW focuses on agency, control, and autonomy in interactive human-AI partnerships.
  • Conceptual models provide high-level structures but often do not specify concrete control mechanisms for practical design across contexts.
  • Empirical studies frequently concentrate on individual settings, limiting synthesis and generalizability across co-creative contexts.
  • Agency and control are distinct: agency concerns intentional initiation, whereas control concerns the operational means of executing intentions.
  • Existing interaction principles, co-creation models, and user-centered frameworks illuminate collaboration but do not systematically connect agency patterns with control mechanisms.

3 Scope, Definitions and Methodology

The review defines human-AI co-creation and agency operationally, adopts an HCI/CSCW lens, and uses a PRISMA-ScR-informed scoping process to synthesize top-tier HCI literature. Its framework connects context, agency, control, and applications, producing structured design guidance.

  • Scope and Definitions: Human-AI co-creation is scoped as creative work in which humans and AI contribute distinct efforts toward shared goals, excluding routine automation without creative contribution.
  • Scope and Definitions: Its HCI/CSCW perspective examines how agency is designed, experienced, and implemented in interaction rather than treating agency as philosophical intentionality or autonomous capability alone.
  • Scope and Definitions: The review distinguishes agency as intentional initiation and directing outcomes from control as the concrete mechanisms used to execute intentions.
  • Methodology: The analytical framework integrates context, agency patterns and distribution, control mechanisms, and applications to synthesize fragmented literature.
  • Methodology: The scoping review follows PRISMA-ScR principles and searches the ACM Digital Library through August 2024 using co-creation and agency terms.
  • Methodology: The final corpus contains 134 papers, with qualitative research and user-centric design most common and design and system implementation the predominant contribution types.
  • Contributions: The paper contributes an integrated agency framework, a catalog of control mechanisms, a cross-context map, and grounded guidance concerning trust, ethics, and future co-creative systems.

4 Context

The review organizes co-creation contexts through six interaction stages and four broad modalities. These contexts shape how humans and AI exchange information, develop ideas, express outputs, collaborate, build artifacts, and test results.

  • Interaction stages: The six stages are perceive, think, express, collaborate, build, and test.They span information gathering, idea formation, externalization, iterative partnership, artifact construction, and evaluation.
  • Interaction stages: Perceive and think establish task understanding as humans provide inputs while AI processes information into ideas, strategies, or solutions.Inputs may be textual, visual, auditory, video-based, or embodied, while AI performs tasks such as recognition, generation, analysis, reasoning, and decision-making.
  • Interaction stages: Express, collaborate, build, and test turn ideas into shared outcomes through outputs, negotiation, construction, and evaluation.AI outputs may be textual, visual, auditory, multimodal, embodied, or decisional, while feedback can loop back to earlier stages.
  • Interaction modalities: The review classifies interaction modalities as textual, visual, auditory, and multimodal or hybrid.These modalities cover symbolic language, visual information, sound, and complex physical or digital products, interfaces, experiences, and systems.
  • Interaction modalities: Multimodal and hybrid systems combine interaction channels for physical or digital products, installations, interfaces, experiences, systems, and clinical decision support.Such systems may integrate multiple sensory channels, physical actions, and context-adaptive interaction methods.

5 Agency: Patterns and Distribution

The review identifies five AI agency patterns ranging from passive influence to proactive partnership, and analyzes agency distribution through locus, dynamics, and granularity. Together, these dimensions describe who holds authority, how it shifts, and at what level control is exercised.

  • Agency patterns: 106 papers with concrete system implementations or detailed interaction designs were used to classify AI agency patterns.The remaining 28 papers were excluded from this specific analysis because they lacked sufficient interaction detail.
  • Agency patterns: The five agency patterns are cooperative, proactive, semi-active, re-active, and passive agency.They range from AI collaboration and initiative to requested support, direct responses, and subtle influence on human dynamics.
  • Agency patterns: Cooperative and proactive agency involve AI collaboration or initiative, including identifying harms, summarizing reviews, offering perspectives, suggesting writing, and generating sound options.These patterns describe AI contributions that extend processes or introduce insights during co-creation.
  • Agency patterns: Semi-active and re-active agency provide requested support or direct responses to user actions, while passive agency subtly influences roles, common ground, authorship, bias, or group interaction.Examples include code translation, interaction-triggered features, reactive design tools, recommenders, and contextual influence in creative or educational settings.
  • Agency distribution: Agency distribution is characterized by locus, dynamics, and granularity.Locus specifies who holds primary authority; dynamics specifies whether authority is statically or dynamically allocated; granularity specifies the level of abstraction or detail.
  • Agency distribution: Finer-grained control points can enhance user agency and transparency by enabling steerable interactions and mitigating black-box issues.Granularity ranges from strategic goals and workflows to specific actions or parameter adjustments.

6 Control Mechanisms

The review organizes control mechanisms across input, process, output, and feedback, showing how systems operationalize agency through guidance, context, transparency, coordination, intervention, adaptation, and iterative refinement.

  • Framework: The control-mechanism framework uses four IPOF processes spanning human-initiated to AI-initiated methods.Three types of control mechanisms are summarized within each process.
  • Input mechanisms: Guided input interaction structures how users provide information through input optimization, interface guidance, and multimodal integration.Users may iteratively refine prompts or customize interface elements to align AI outputs with expectations.
  • Process mechanisms: Context awareness and memory retention use historical, environmental, task, and user information to support contextual understanding and collaboration continuity.Interaction histories and provenance tracking help preserve prior activity and idea evolution.
  • Process mechanisms: Transparency and explainability mechanisms expose interaction histories, progress, decisions, influencing factors, and model operations to support understanding and trust.Examples include contribution histories, progress visualizations, urgent symptom combinations, and links between actions and algorithmic processes.
  • Action mechanisms: Action coordination distributes responsibilities and authority through complementary human-AI roles, while modification and intervention let users edit outputs or control parameters and prompts.Humans commonly provide strategic or creative direction while AI handles data processing or routine tasks.
  • Output mechanisms: Adaptive scaffolding, chain-of-thought, confidence visualization, and explanatory feedback adjust assistance or expose reasoning and reliability information.These mechanisms help users evaluate outputs, interpret uncertainty, and understand AI decision-making.
  • Feedback mechanisms: Iterative feedback loops enable continuous refinement through user-directed, system-initiated, or bidirectional exchanges.Examples include real-time suggestions, generated sounds, immediate visual feedback, goal tracking, and response cycles for refinement.

7 Applications

Across seven application domains, the review maps agency configurations to distinct co-creative practices. The domains combine reactive, proactive, semi-proactive, and passive mechanisms according to goals such as authorial control, expert oversight, empowerment, and ethical alignment.

  • News: News systems combine semi-proactive and reactive agency, balancing automated cognitive support with writers’ active revision and authorial control.Provenance tracking and viewpoint-challenging or viewpoint-reinforcing exposure are additional strategies.
  • Healthcare: Healthcare systems emphasize expert-supervised reactive agency while integrating semi-proactive workflow support and participatory proactive design.Clinicians retain final decision-making authority, while systems address documentation, cultural sensitivity, therapy, and creative assistance.
  • Art: Art applications are dominated by tool-mediated passive and semi-proactive agency, with prompt engineering, interface control, and iterative refinement shaping ideation.These systems affect verbal articulation, user expectations, and mixed-initiative creative work.
  • Education & Research: Education and research systems combine curriculum-bound reactive and proactive mechanisms with semi-proactive, context-aware assistance.Examples include code error correction, interpretable pattern generation, exploration strategies, and meta-review support.
  • Entertainment: Entertainment applications use proactive and semi-proactive agency for context-aware co-creation, reactive parameter control, and collaborative editing.Examples include environmental data integration, slider-based music controls, and chatbot interfaces.
  • Software Development: Software development prioritizes value-aligned proactive strategies and harm-aware reactive control for ethical alignment and incident-linked prototyping.The reviewed examples connect agency mechanisms to ethical values and identifying potential harms.
  • Accessibility: Accessibility technologies emphasize proactive and semi-proactive agency to support empowerment, keyword-driven communication, and maximized user choice.These mechanisms are applied to motor-impaired communication and inclusive HCI methods.

8 Challenges & Directions

The review organizes co-creation challenges across collaborative experience, infrastructure, and social implications. It highlights unresolved tensions involving ownership, trust, privacy, interoperability, equity, ethics, and long-term changes in human-AI collaboration.

  • Challenge framework: The review groups challenges into collaborative experience, collaboration infrastructure, and social implications, spanning creativity, ownership, trust, security, interoperability, equity, ethics, and future paradigms.These directions are intended to guide responsible human-AI collaborative-system design.
  • Collaborative experience: AI-assisted co-creation complicates authorship and intellectual property because distinguishing human input from AI contributions remains difficult.The review emphasizes preserving user control, perceived ownership, and human agency while legal, ethical, and experiential questions remain unresolved.
  • Collaborative experience: Trust depends on transparent decisions and mechanisms for user influence or correction, yet ambiguous AI roles and incomprehensible reasoning can undermine confidence.The review frames the design task as balancing AI autonomy with sufficient human oversight.
  • Collaboration infrastructure: Protecting privacy and security requires encryption, access controls, and clear user control over data usage, while interoperability remains difficult across systems, workflows, interfaces, and standards.The review identifies inconsistent performance, efficiency, usability, and limited controllability as practical interoperability concerns.
  • Social implications: Socially responsible co-creation must mitigate inequity, job displacement, harmful content, and disruption to creative economies.The challenge is ensuring that AI benefits society broadly and equitably rather than reinforcing existing harms.
  • Social implications: Ethical risks include inaccuracies, academic-integrity violations, dominant views, confirmation bias, and over-reliance, while limited long-term analysis constrains preparation for increasing AI autonomy.The review also notes that AI-mediated communication may alter interaction and self-perception, and that ideation systems may encourage blind acceptance of AI.

9 Flow Analysis, Case Analysis and Discussions

The Sankey analysis identifies recurring human-AI co-creation pathways across modalities, agency patterns, control mechanisms, and application domains. It also connects these patterns to design opportunities, ethical concerns, and longer-term theoretical and societal implications.

  • Pathway Analysis: Textual interaction and co-operative agency form a prevalent pathway in writing and software development, supported by feedback loops and transparency mechanisms.The pathway commonly involves perceive, collaborate, and build stages with shared or dynamically negotiated control.
  • Pathway Analysis: Visual co-creation typically moves from reactive AI responses to co-operative refinement through guided input and iterative feedback.Control may shift between human guidance and AI generation at fine granularity.
  • Pathway Analysis: Multimodal software development combines co-operative or semi-active agency with guided input, action coordination, and iterative feedback.These mechanisms support shared or dynamically negotiated control across complex interaction stages.
  • Design Implications: Frequent Collaborate stages require robust iterative feedback loops and transparency, while high AI initiative may require finer control granularity or dynamic agency allocation.The review links control-mechanism choices to agency patterns and interaction stages.
  • Discussion: The taxonomy supports systematic comparison of co-creative systems while highlighting ethical, privacy, labor, accountability, and bias concerns associated with increasing AI autonomy.It also identifies sparse auditory interaction as an opportunity for innovation in assistive technology and music co-creation.
  • Theoretical Connections: The synthesis provides empirical support for distributed, situated, and dynamically negotiated agency theories, while showing that raw AI capabilities alone are insufficient for effective co-creation.The catalog of control mechanisms connects user intention to system execution and supports agency preservation.

10 Conclusion

The review maps human-AI co-creation in 134 papers from premier HCI venues through an agency-centered framework. It catalogs control mechanisms and derives implications for designing co-creative systems with attention to collaboration, trust, ethics, and societal effects.

  • Conclusion: The review analyzed 134 papers and integrated context, agency patterns and distribution, control mechanisms, and practical applications.It systematically mapped HCI literature on human-AI co-creation through the lens of agency.
  • Conclusion: The catalog shows that transparency and feedback loops help shape collaboration, user experience, and trust.These mechanisms operationalize how agency is implemented in co-creative systems.
  • Conclusion: The synthesis offers grounded guidance for future CSCW research and co-creative-system design concerning ethical, societal, and long-term implications.

A PRISMA-ScR Guided Literature Review Process and Justification

The review followed PRISMA-ScR and used a targeted search and staged screening process to identify relevant peer-reviewed HCI studies of human-AI co-creation and agency. Eligibility emphasized concrete HCI contributions and direct human-AI interaction.

  • Review Process: The literature review followed PRISMA-ScR guidelines, with the selection process documented in Figure 9.
  • Search Strategy: Searches targeted co-creation-related terms and agency in titles and content across CHI, UIST, CSCW, Ubicomp, IUI, and DIS.The query included co-creation, co-writing, co-drawing, and co-design, while related terms produced no additional eligible papers.
  • Screening Criteria: Title and abstract screening excluded records outside human-AI co-creation and agency or lacking peer-reviewed full-paper status.
  • Eligibility Criteria: Full-text eligibility required collaborative human-AI co-creation in HCI and excluded papers focused only on conceptual discussion without concrete HCI techniques, designs, or systems.The criteria prioritized operational agency and implemented control mechanisms.
  • Eligibility Criteria: Studies involving only human-human or AI-AI interaction were excluded to preserve focus on human-AI agency distributions and control challenges.

B Coding Criteria and Theoretical Foundations

The coding framework combined established theories and criteria to classify research methods, contribution types, and interaction stages. Interaction stages were coded using a six-stage model derived from prior work.

  • Coding Foundations: The review stated that theories and criteria guided the coding of the included papers.
  • Coding Criteria: Research methods were categorized as quantitative, qualitative, theoretical or evaluation framework, mixed methods, or user-centered research.Research contributions were classified as comparison study, concept generation, design, system implementation, or interview.
  • Coding Criteria: Interaction stage was coded across perceive, think, express, collaborate, build, and test.The six-stage scheme was derived from prior work and used to characterize co-creative interaction context.

C The Taxonomy of Papers

The taxonomy organizes the reviewed papers across interaction context, agency, control, application, and challenge dimensions. Its tables separately classify interaction stages, modalities, agency patterns and distributions, control mechanisms, applications, and challenges or directions.

  • Context: The taxonomy classifies papers by interaction stages and modalities, capturing the contexts in which human-AI co-creation occurs.
  • Agency: It distinguishes agency patterns from agency distributions, separating forms of agency from how agency is allocated.
  • Control: A dedicated taxonomy of control mechanisms captures the strategies used to manage agency in the reviewed systems.
  • Applications: The review maps the application domains represented in the surveyed papers.
  • Challenges & Directions: It also catalogs reported challenges and directions for future work.
Loading 2507.06000v2…