Source-linked AI summary

When AI Becomes Hard to Understand: Cognitive Demands in Real-World Human-AI Conversations

Yingcan Carol Wang, Iman Munire Bilal, Qamar Zaman

arXiv:2609.17301v1cs.HC

TL;DR

Generative AI is increasingly used for complex financial and health decisions, but evidence is limited on when responses become difficult to process in real-world dialogue. The paper analyses more than 84,000 conversations using repair behaviours as indicators and finds that difficulty depends on combinations of response characteristics rather than any single feature. It proposes a conversational complexity budget to guide response configuration for particular users, tasks, and interactions.

  • Problem

    Evidence is limited on how multiple dimensions of generated responses combine to affect conversational difficulty during natural human–AI interaction.

  • Method

    The paper analyses more than 84,000 naturalistic Finance and Health conversations, using repeated prompting and clarification following misunderstanding as behavioural indicators.

  • Results

    Conversational difficulty cannot be explained by individual response characteristics; greater lexical diversity was associated with less repeated prompting in shorter responses, with this association weakening as responses lengthened.

  • Takeaways & Limitations

    The conversational complexity budget frames comprehensibility as dependent on how response demands combine with the user and task.

  • Takeaways & Limitations

    Repeated prompting and clarification are behavioural indicators rather than direct measures of cognitive load, and the observational design requires future validation against established metrics.

Abstract

from arXiv · show

Generative AI increasingly supports complex financial and health decisions, yet we know little about when its responses become difficult to process in real-world dialogue. We analyse more than 84,000 ChatGPT and Gemini conversations, using repeated prompting and clarification following misunderstanding as behavioural indicators of cognitive difficulty. We find that response characteristics such as length, readability and lexical diversity do not have fixed relationships with conversational difficulty; instead, their relationships depend on how they combine. Most notably, greater lexical diversity was associated with less repeated prompting in shorter responses, but this association weakened as response length increased, a pattern that replicated across financial and health conversations. We propose a conversational complexity budget to conceptualise these interdependencies: the demands associated with one response characteristic may depend on those accompanying it. The resulting design challenge is how to configure response complexity for the particular user, task and interaction.

1 Introduction

The paper studies when naturally occurring AI conversations become difficult to process, focusing on how user goals and response characteristics jointly relate to conversational breakdown. Across Finance and Health, it uses repeated prompting and clarification as behavioural indicators and proposes a conversational complexity budget.

  • The study analyses more than 84,000 naturally occurring Finance and Health conversations using repeated prompting and clarification following misunderstanding as behavioural indicators.
  • Response length, readability, lexical diversity, information density, and formatting did not have fixed relationships with conversational difficulty when considered independently.
  • Across Finance and Health, research, comparison, personalised analysis, and learning intents showed elevated breakdown, whereas simple retrieval showed lower breakdown.
  • Greater lexical diversity was associated with less repeated prompting in relatively short responses, but this association weakened as responses became longer in both domains.
  • The paper proposes a conversational complexity budget in which response demands combine, making comprehensibility dependent on the user and task rather than the response alone.

2 Related Work

Prior work frames cognitive demand as dependent on limited processing capacity, task complexity, presentation, and user knowledge. This paper addresses limited evidence on how these factors combine during natural human–AI conversation by studying observable repair behaviours and response characteristics.

  • 2.1 Cognitive load in human–AI interaction: Cognitive Load Theory treats processing capacity as limited and distinguishes inherent material complexity from avoidable demands introduced by presentation.
  • 2.1 Cognitive load in human–AI interaction: In conversational interfaces, users process dynamically generated information while maintaining context and deciding what to ask or do next.
  • 2.3 Conversational repair: Repeated prompting and clarification provide observable, scalable indicators of conversational difficulty, but they are not direct measures of cognitive load.
  • 2.5 Research gap: Existing research leaves comparatively little evidence about how multiple dimensions of generated responses operate in combination during natural conversation.
  • 2.5 Research gap: The paper examines how observable breakdown relates to user intent, substantive domain, and linguistic and structural response characteristics.

3 Data overview and processing

The study extends a naturalistic ChatGPT and Gemini corpus and focuses on Finance and Health conversations. It segments long histories into topic-coherent units and applies privacy-preserving, aggregate analysis to the resulting datasets.

  • The source corpus records user messages, AI responses, and timestamps and was extended to July–October 2025, covering 2,626 users and more than 1.1 million prompts.
  • 3.2 Conversation segmentation and processing: Longer histories were segmented into topically coherent subchats using lexical overlap between consecutive prompts, yielding approximately 395,000 conversation-level units.
  • 3.3 Domain selection and identification: Finance and Health were selected because they are high-stakes domains where AI increasingly supports information seeking and decision-making.
  • The analytical datasets contain approximately 43,100 Finance conversations from 1,745 users and 41,500 Health conversations from 1,558 users.
  • The data were voluntarily contributed, de-identified, access-controlled, and reported only in aggregate.

4 Methodology

The methodology classifies domain content and user intent, constructs two conversational repair measures, characterises AI-response features, and models their associations with difficulty. It combines multi-label taxonomies with rule-based and learned classification procedures.

  • The methodology models how response characteristics, user intent, and substantive context are associated with two repair behaviours indicating conversational difficulty.
  • 4.1 Conversation classification: Conversations receive non-mutually exclusive content and intent labels because one conversation may involve multiple topics and user goals.
  • 4.2.2 Clarification Following AI-Induced Misunderstanding: The second repair measure captures explicit clarification requests following apparent AI-induced misunderstanding.
  • 4.2.1 Fragmented Repeated Prompts: Fragmented repeated prompting is detected by comparing each eligible prompt with preceding prompts using character-sequence, Jaccard, and containment similarities.

4.3 AI Linguistic Measures

The study characterizes AI responses using linguistic and structural measures, then models how these features relate to conversational repair indicators of cognitive difficulty.

  • Measures: Five response measures captured readability, length, formatting density, lexical diversity, and information density.Readability was measured with Flesch Reading Ease, lexical diversity with MTLD, and information density as the ratio of content to function words.
  • Analysis windows: Temporal windows included all eligible turns before the first repeated prompt or clarification request, excluding the triggering event and subsequent turns.This design ensured that predictors preceded the corresponding repair outcome.
  • Aggregation: Turn-level measures were averaged without weighting so that each conversational turn contributed equally rather than longer responses dominating the estimates.Response length therefore represented average words per turn, not total text volume in a conversation window.
  • Measurement scope: Several measures have constrained interpretations: readability captures surface textual difficulty, formatting detects explicit Markdown-like structure, and information density is only a lexical proxy.These measures do not directly capture conceptual or clinical complexity, all visual organization, factual accuracy, or clinical informativeness.
  • Statistical modeling: Population-averaged logistic regression with generalized estimating equations modeled repair outcomes while controlling for user, task, content, and model-provider variation.Models included interactions between response length and other AI-response characteristics and clustered observations by user.
  • Outcomes: The analysis used fragmented repeated prompting and clarification after AI-induced misunderstanding as binary indicators of conversational difficulty.Predictors were restricted to interaction turns preceding the relevant repair event.

5 Results

Conversational difficulty was not governed by a simple decision-authority gradient or any single response feature. Across Finance and Health, the strongest evidence showed that response characteristics interacted, especially lexical diversity with response length.

  • Conversational breakdown and task context: Substantive content was more consistently associated with breakdown in Health than Finance, including disease-focused, psychological, social, and functional concerns.Business Finance was the clearest Finance exception, with higher odds of both repeated prompting and clarification.
  • Single-feature associations: Main-effects models provided little evidence that individual response features had fixed relationships with difficulty; lexical diversity and response length sometimes showed reduced odds, while readability and information density were often nonsignificant.In Finance, lexical diversity and response length were associated with lower odds of repeated prompting, whereas Markdown density was associated with higher odds.
  • Response-characteristic interactions: Lexical diversity interacted significantly with response length in both Finance and Health, with its association with repeated prompting weakening as responses grew longer.In Finance, trajectories converged and crossed at approximately 600 words; the interaction remained significant across alternative Health specifications.
  • Response-characteristic interactions: In Finance, lower information density predicted less clarification for longer responses, whereas higher information density predicted more clarification; trajectories crossed at approximately 150–200 words.The crossover is specific to the fitted model and observed data, not a general threshold for conversational difficulty.
  • Response-characteristic interactions: Response readability also interacted with length: misunderstanding decreased with length for less readable responses but increased with length for highly readable responses.This pattern distinguishes surface reading ease from overall comprehensibility in longer responses.

6 Discussion

Across Finance and Health, conversational difficulty depended on user intent and combinations of response characteristics rather than any single task or response feature. These patterns motivate a conversational complexity budget and adaptive response design, while remaining observational and limited in scope.

  • 6 Discussion: Across more than 84,000 Finance and Health conversations, breakdown differed by user intent and response configuration rather than following a simple decision-authority or single-feature gradient.Higher-order research, comparison, and personalised-analysis intents generally showed more breakdown than simple retrieval, but decision authority and cognitive demand remained distinct dimensions.
  • 6.1 Response complexity and conversational difficulty: Greater lexical diversity was associated with less repeated prompting in shorter responses, but this relationship weakened as response length increased in both Finance and Health.This replicated interaction shows why lexical diversity does not have a fixed relationship with conversational difficulty.
  • 6.1 Response complexity and conversational difficulty: Response length, readability, information density, and formatting had conditional relationships with difficulty, so easier-to-read or longer responses were not uniformly easier to understand.In Finance, information density and response length interacted, while surface readability also changed its relationship with misunderstanding as responses became longer.
  • 6.2 Defining a conversational complexity budget: A conversational complexity budget conceptualises how demands accumulate across response characteristics, while remaining a theoretical heuristic because underlying cognitive capacity was not directly measured.The same characteristic may help when other demands are modest but become burdensome as demands accumulate.
  • 6.3 Evaluating the user–task–response configuration: Preliminary evidence suggests response effects depend on the user’s input: easier-to-read AI responses were associated with more misunderstanding when prompts were harder to read.This relationship was much weaker for easier-to-read prompts, but prompt readability should not be treated as a measure of user literacy.
  • 6.4 Beyond Universal Simplicity: Designing Adaptive Responses: The design implication is to configure response demands for particular users, tasks, and interactions rather than optimise simplicity along one dimension.Repeated prompting and clarification may provide observable signals for adaptation, but the observational study does not establish which adaptive strategies reduce difficulty.
  • 6.6 Limitations and future work: The study cannot establish causal effects because its repair indicators are behavioural proxies for cognitive difficulty, and its linguistic measures cover selected complexity dimensions in Finance and Health only.Future experiments should manipulate response characteristics and validate repair behaviours against established cognitive-load metrics.

7 Conclusion

Across more than 84,000 real-world conversations, conversational difficulty could not be explained by individual response characteristics alone; it depended on how characteristics combined across user intents. The authors propose a conversational complexity budget that shifts attention toward the user–task–response configuration and motivates adaptive response design.

  • Across more than 84,000 real-world conversations, conversational difficulty was associated with combinations of response characteristics rather than length, readability, or lexical diversity alone.
  • The conversational complexity budget frames complexity as manageable under some conditions but more difficult as other demands accumulate.
  • The findings shift attention from responses alone toward the broader user–task–response configuration in which complexity is formed.
  • Adaptive systems could configure the form and amount of complexity to users’ goals and adjust when conversational breakdown emerges.

A Context taxonomy for Finance

The Finance context taxonomy is presented as a structured category system for Human–AI conversations.

  • Table 5 structures the Finance taxonomy into categories and subcategories.

B Context taxonomies for Health

The Health context taxonomies organize Human–AI conversations into general-health and disease-specific categories.

  • Tables 6 and 7 present separate taxonomies for General Health and Disease-specific Health conversations.

C Intent taxonomy for Finance and Health

The intent taxonomy organizes Finance and Health Human–AI conversations by user intent and decision authority.

  • Table 8 defines user-intent categories with examples and associated decision-authority levels.

D Training corpus for Clarification classification

Table 9 presents the proportions of organic and synthetic data for each clarification-classification label.

  • Table 9 reports the proportion of organic and synthetic data for each clarification-classification label.
  • The table concerns the training corpus used for clarification classification.
  • The manuscript was submitted to ACM.
Loading 2609.17301v1…