Source-linked AI summary

The Metacognitive Demands and Opportunities of Generative AI

Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, Sean Rintel

arXiv:2312.10893v3cs.HC

TL;DR

GenAI’s broad capabilities create usability challenges in prompting, evaluating outputs, and deciding how to incorporate AI into workflows. The paper applies metacognition to analyze these demands and draws on cognitive science, user studies, interventions, prototypes, and human-AI interaction research. It concludes that systems can address the demands by improving users’ metacognition and reducing required metacognitive work through explainability and customizability, while recognizing that added support may increase cognitive load.

  • Problem

    GenAI’s flexibility and generality create usability challenges around prompting, output evaluation, reliance, and workflow automation, without a coherent cognition-grounded understanding.

  • Method

    The paper uses metacognition to analyze GenAI interactions, drawing on psychology, cognitive science, user studies, metacognitive interventions, prototypes, and human-AI interaction research.

  • Results

    The paper identifies metacognitive demands in GenAI use and proposes improving users’ metacognition or reducing system demands through explainability and customizability.

  • Takeaways & Limitations

    Metacognition offers a framework for understanding GenAI usability and for developing human-AI interaction research and design directions.

  • Takeaways & Limitations

    Greater customizability and metacognitive support may increase cognitive load, and the appropriate balance between too much and too little customizability remains unresolved.

Abstract

from arXiv · show

Generative AI (GenAI) systems offer unprecedented opportunities for transforming professional and personal work, yet present challenges around prompting, evaluating and relying on outputs, and optimizing workflows. We argue that metacognition$\unicode{x2013}$the psychological ability to monitor and control one's thoughts and behavior$\unicode{x2013}$offers a valuable lens to understand and design for these usability challenges. Drawing on research in psychology and cognitive science, and recent GenAI user studies, we illustrate how GenAI systems impose metacognitive demands on users, requiring a high degree of metacognitive monitoring and control. We propose these demands could be addressed by integrating metacognitive support strategies into GenAI systems, and by designing GenAI systems to reduce their metacognitive demand by targeting explainability and customizability. Metacognition offers a coherent framework for understanding the usability challenges posed by GenAI, and provides novel research and design directions to advance human-AI interaction.

1 INTRODUCTION

GenAI’s flexibility, generality, and originality create usability challenges that require users to monitor and control their thinking and behavior. The paper uses metacognition to interpret these challenges and proposes support strategies, explainability, and customizability as design directions.

  • Motivation: GenAI can transform work through flexibility, generality, and originality, but these properties complicate human-centered system design.User studies identify challenges involving prompting, evaluating and relying on outputs, and choosing automation strategies.
  • Metacognitive lens: Metacognition provides a cognitive framework for understanding GenAI usability challenges and identifying new research and design opportunities.The paper defines metacognition as the ability to monitor and control one’s thought processes.
  • Metacognitive demands: Working with GenAI demands goal formulation, task decomposition, output evaluation, prompting adjustment, and decisions about whether and how to automate workflows.These demands involve metacognitive monitoring and control across local interactions and higher-level workflow decisions.
  • Design directions: The paper proposes improving users’ metacognition through support strategies for planning, self-evaluation, and self-management integrated into GenAI systems.This direction draws on evidence that metacognitive abilities can be taught and on recent HCI work.
  • Design directions: A complementary direction is reducing GenAI’s metacognitive demand through task-appropriate explainability and customizability.The paper suggests explainability can offload metacognitive processing to the system and that GenAI’s flexibility, generality, and originality can inform design solutions.
  • Paper approach: The paper organizes its analysis around metacognition, GenAI’s prompting and output challenges, automation strategy, and interventions addressing these demands.It draws on psychology, cognitive science, metacognitive interventions, GenAI prototypes, and human-AI interaction research.

2 WHAT IS METACOGNITION?

Metacognition concerns how people understand, monitor, and control their own thinking. The paper distinguishes knowledge and experiences from monitoring and control, while emphasizing their interdependence and context dependence.

  • Metacognition includes explicit knowledge about abilities, strategies, reasoning, decisions, and beliefs, alongside directly experienced or implicit cues about cognitive processing.
  • Metacognitive knowledge and experiences influence each other, as experiences can become knowledge and knowledge can be retrieved during later experiences.
  • Monitoring assesses one’s thinking, whereas control guides it; relevant GenAI abilities include self-awareness, confidence adjustment, metacognitive flexibility, and task decomposition.
  • Confidence adjustment supports decision-making and reasoning, with well-adjusted confidence matching one’s abilities and distinguishing correct from incorrect performance.
  • Monitoring and control are reciprocal: assessments can change strategies, while strategy changes provide feedback that alters performance assessment.
  • Whether metacognition is domain-general or domain-specific remains debated, and the demands of a situation are likely context-dependent.

3 THE METACOGNITIVE DEMANDS OF GENERATIVE AI

GenAI changes work across domains while requiring users to manage additional metacognitive demands. These demands involve prompting, evaluating and relying on outputs, and deciding how to automate workflow tasks.

  • GenAI introduces greater metacognitive demand when users prompt systems, evaluate and decide whether to rely on outputs, and determine how to automate workflows.
  • Metacognitive demand contributes to cognitive load but is distinct from total mental effort, which also includes non-metacognitive aspects of task processing.
  • The paper characterizes GenAI work as imposing high monitoring and control demands, while noting that interface design and interaction modes can change their type and extent.

3.1 Prompting generative AI systems

Prompting GenAI requires users to make goals explicit, decompose tasks, evaluate outputs, adjust confidence, and flexibly revise strategies. These demands vary with interface design, user expertise, task context, and model behavior.

  • 3.1.1 Prompt formulation: self-awareness and task decomposition.: Prompt formulation requires self-awareness of task goals and decomposition of tasks into smaller sub-tasks that can be verbalized as prompts.
  • 3.1.1 Prompt formulation: self-awareness and task decomposition.: Many GenAI systems require users to specify implicit intentions, such as an email recipient and tone, and to express multi-step tasks as discrete instructions.
  • 3.1.1 Prompt formulation: self-awareness and task decomposition.: Non-experts may struggle to begin prompting, choose effective instructions, or distinguish socially appropriate prompting from prompting that works effectively.
  • 3.1.1 Prompt formulation: self-awareness and task decomposition.: Diegetic prompting can be easier but offers less control, whereas non-diegetic prompting may be harder and slower while potentially encouraging self-awareness and task decomposition.
  • 3.1.1 Prompt formulation: self-awareness and task decomposition.: These prompting demands vary with interaction mode, task context, domain, usage intention, and users’ self-awareness and task-decomposition abilities.
  • 3.1.2 Prompt iteration: confidence adjustment and metacognitive: Prompt iteration requires evaluating outputs, adjusting confidence in one’s prompting ability, and flexibly changing prompts, task decomposition, or retry strategies.
  • 3.1.2 Prompt iteration: confidence adjustment and metacognitive: GenAI non-determinism makes iteration difficult because changing one prompt aspect can unintentionally alter another output aspect, risking derailment from task goals.
  • 3.1.2 Prompt iteration: confidence adjustment and metacognitive: Model flexibility creates a fuzzy abstraction matching problem, making it difficult to discern system capabilities and align intent and prompting accordingly.

3.2 Evaluating and relying on generative AI outputs

Evaluating and relying on GenAI outputs creates substantial metacognitive demands because users must judge extensive, novel content despite uncertain confidence, complex failure modes, and limited objective quality measures.

  • Confidence adjustment: Users need well-calibrated confidence that matches evaluation performance and well-resolved confidence that discriminates correct from incorrect outputs.Calibration concerns overall accuracy of confidence, whereas resolution concerns discrimination between correct and incorrect outputs.
  • Confidence adjustment: Evidence from GenAI user studies suggests confidence affects reliance, but direct studies measuring and manipulating self-confidence in GenAI interactions remain missing.Confident programmers actively questioned confusing AI code, while others avoided deep review or were uncertain how to address errors.
  • Challenges of output evaluation: Automated suggestions can make evaluation harder because users must infer the system’s intent, even without explicitly prompting it.Novices may be especially vulnerable when they lack sufficient domain or GenAI knowledge.
  • Challenges of output evaluation: GenAI’s extensive novel outputs make quality evaluation more cognitively effortful than evaluating autocomplete or intelligent code-completion suggestions.Outputs may include entire emails, presentations, or software, increasing the length and complexity of evaluation.
  • Confidence adjustment: GenAI’s speed, fluency, and ease of generation can misleadingly increase confidence in outputs and evaluation ability, reducing deliberate processing effort.These cues may influence metacognitive control by shaping users’ confidence and subsequent evaluation behavior.
  • Challenges of output evaluation: GenAI’s multiple, non-intuitive failure modes can introduce subtle errors and require expertise distinct from existing domain expertise.Developers may need new craft practices for debugging AI-generated code.
  • Confidence adjustment: Objective confidence adjustment is difficult when generated content has subjective, diffuse, or indirect benefits, such as an LLM-generated email or ideation output.Even subjective workflows require appropriate reliance strategies grounded in well-adjusted confidence.

3.3 Automation strategy and generative AI workflows

GenAI’s broad applicability creates a higher-level metacognitive demand: users must decide whether, how, and how much to integrate it into workflows. Effective automation requires self-awareness, confidence, and flexibility because reliance can restructure tasks and produce unproductive outcomes.

  • Automation strategy: GenAI’s generality requires users to assess whether, how, and how much generated content should be integrated compared with conventional approaches.This is a workflow-level automation strategy distinct from local prompting and output evaluation.
  • Automation strategy: Users need self-awareness of GenAI’s applicability and impact, confidence in manual versus AI-assisted performance, and flexibility to adapt workflows.These abilities support decisions about when and how GenAI should be incorporated.
  • Workflow evidence: GenAI changes workflows across domains: coding tools alter work practices, data scientists identify workflow integration as a control lever, and writing shifts time from drafting to editing.The reported evidence spans programming, data science, and writing, though usability implications remain underexplored in writing.
  • Workflow evidence: Users face switching costs and task restructuring associated with automation, while GenAI’s broad applicability creates additional decisions about where automation belongs.These concerns connect GenAI workflows with longstanding human-automation challenges.
  • Automation strategy: Inappropriate reliance on GenAI may reduce productivity, increase error risk, or contribute to de-skilling, making task-level cost-benefit judgments necessary.The attention investment problem frames automation as weighing saved attention against the attention required to implement it.
  • Workflow evidence: Early programming studies report unproductive interaction patterns, including repeatedly editing abandoned Copilot suggestions or trying to coerce a correct suggestion.These patterns suggest potential over-reliance among some novices.
  • Limitations and future research: Evidence remains limited because existing workflow reports use short study contexts and focus mainly on programming rather than realistic, cross-domain workflows.Further research should examine self-awareness and confidence outside programming.
  • Cognitive offloading: Cognitive offloading shifts ideation, memory retrieval, and reasoning partly from users to GenAI, and lower self-confidence is associated with greater use of external aids.Psychological research on reminders and information search provides relevant evidence for studying GenAI reliance.

4 ADDRESSING THE METACOGNITIVE DEMANDS OF GENERATIVE AI

The paper proposes two complementary ways to address GenAI’s metacognitive demands: strengthen users’ metacognition and reduce systems’ demands through explainability and customizability. It frames planning, self-evaluation, and self-management as support strategies while identifying design choices that can offload metacognitive processing.

  • Design space: GenAI’s metacognitive demands can be addressed by improving users’ metacognition or reducing systems’ metacognitive demand.The paper treats these approaches as complementary, while acknowledging that the distinction is not clean-cut.
  • Supporting users: Evidence across ages, tasks, time scales, and learning settings suggests that metacognition can be improved and that targeted support can improve performance on demanding tasks.The cited evidence includes children and adults, lecture comprehension, mathematical reasoning, and immediate or delayed effects.
  • Research directions: Table 2 organizes open research questions for understanding GenAI’s metacognitive demands.The paper also summarizes these demands and research directions through figures and tables.
  • Reducing system demands: Task-appropriate explainability may offload metacognitive processing from users to GenAI systems.The paper proposes augmenting existing explainability approaches by explicitly considering metacognition.
  • Reducing system demands: Customizability can reduce metacognitive demand by surfacing the explicit and implicit parameters users can control through settings and prompting strategies.Current GenAI systems expose many parameters, but appropriate ways to surface them remain a design challenge.
  • Supporting users: Metacognitive support strategies integrated into GenAI systems can target users’ planning, self-evaluation, and self-management.The paper draws on intervention research and prototype studies to identify possible strategies and research directions.

4.1 Improving user metacognition

Metacognitive support can improve how users plan, evaluate, and manage work with GenAI. Proposed interventions include task decomposition, reflective prompts, interactive guidance, and context-sensitive control over when support appears.

  • Planning: Planning helps users define goals, decompose tasks, and translate intentions into explicit prompts for GenAI.Prompt chaining operationalizes this by mapping targeted subtasks to sequential LLM steps.
  • Planning: Feedforward can warn users when vague prompts are unlikely to achieve their goals and guide them toward more effective prompting.Because feedforward can also increase cognitive load, its complexity must be balanced against users’ needs.
  • Self-evaluation: Self-evaluation prompts encourage reflection on goals, strategies, confidence, and outputs, supporting more effective prompting and detection of hallucinations.Prior outputs, think-aloud prompts, and critical questions are proposed ways to promote real-time self-awareness.
  • Self-evaluation: Interactive GenAI systems can guide users through problem-solving steps rather than simply provide solutions, but intuitive interfaces may inflate confidence without improving accuracy.Periodic checks that challenge users’ assumptions can help counter this risk.
  • Self-management: Self-management support can adapt prompts and content to users’ workflow, while allowing users to schedule interventions or personalize their frequency and timing.Hypothetical interfaces use prior interactions, user preferences, and confidence settings to tailor support or let users dismiss it.

4.2 Reducing metacognitive demands

GenAI systems can reduce users’ metacognitive demands through explainability and customizability. These design choices can offload monitoring and control, but the appropriate amount of user control depends on task, expertise, and context.

  • Explainability: Explainability can partly offload metacognitive processing by providing contextual and performance information alongside GenAI inputs and outputs.This information can help users adjust confidence in prompting, output evaluation, and automation strategy.
  • Explainability: Explanations about capabilities, output quality, and effective prompting can help users choose between manual work and GenAI assistance and revise prompts.These explanations address local prompting and evaluation as well as broader workflow decisions.
  • Explainability: Interactive explanations can be augmented with self-evaluation interventions that encourage users to reflect on mental models and confidence.Interactivity is presented as especially relevant because GenAI systems are flexible, general, and original.
  • System customizability: Increasing customizability can raise demands for self-awareness, confidence calibration, task decomposition, and metacognitive flexibility, especially for novice users.In one cited study, half of users did not adjust model parameters despite many producing insecure code with an AI assistant.
  • System customizability: Customizability can also support experienced users by allowing task-appropriate control over diversity, factuality, and output presentation.The paper identifies the balance between too much and too little customizability as an open research question.

4.3 Managing cognitive load while addressing metacognitive demands

Metacognitive support may improve interaction with GenAI while also adding processing demands. The paper therefore frames cognitive load as a design trade-off requiring adaptive interventions and further study.

  • Cognitive-load trade-offs: Metacognitive interventions can increase cognitive load by adding self-reflective prompts, task sub-goals, or model explanations.The authors note that some studies find no increase in overall cognitive load, but the relationship remains under active investigation.
  • Cognitive-load trade-offs: Explainability may reduce the load of monitoring and control while increasing the load of processing explanations.Some interactive explanations have been found to increase cognitive load, so a net reduction remains a hypothesis.
  • Adaptation over time: Users may gradually internalize metacognitive strategies, explanations, and settings, potentially reducing reliance on external prompts over time.Future research should examine how interventions and their intensity should adapt or fade as users learn.
  • Seamful design: The paper argues that some added effort may be justified when metacognitive support and explanations are well designed and promote reflection.This motivates seamful design rather than assuming interfaces should always minimize visible complexity.

5 CONCLUSION

Metacognition provides a framework for understanding how GenAI changes users’ cognitive responsibilities and for designing more human-centered interaction. GenAI interaction may also advance foundational research on metacognition.

  • Conclusion: As cognition is increasingly offloaded to GenAI, users face greater demands to understand, guide, and evaluate their interaction with these systems.The paper connects this challenge to the idea of meta-literacy for using computational tools effectively.
  • Conclusion: Combining metacognition with GenAI’s flexibility, generality, and originality offers a basis for research toward a collaborative relationship between people and computational agents.The paper also identifies GenAI interaction as a paradigm for interdisciplinary study of metacognition.
Loading 2312.10893v3…