Source-linked AI summary
Probing How Users Interact with Turn-Level Design Frictions for AI Chatbots
Helen Weixu Chen, Katy Ilonka Gero
TL;DR
AI chatbot interaction can make users passive consumers by enabling minimal prompts to produce ready-to-adopt responses, raising questions about preserving meaningful involvement. This paper examines design friction as interaction constraints and finds increased ownership alongside consistent interaction costs, with selective memory benefits.
Problem
AI chatbot interaction can position users as passive consumers who move from minimal prompts directly to ready-to-adopt responses, motivating ways to preserve meaningful involvement.
Method
The paper studies intentional design friction in chatbot exchanges and evaluates its effects on ownership, recall, and recognition using interaction logs and semi-structured interviews.
Results
All six friction probes increased perceived ownership relative to ChatGPT, while memory effects were selective and every friction increased interaction costs.
Takeaways & Limitations
Design friction may need to be tailored to when and how it can increase involvement productively.
Takeaways & Limitations
The recall difference cannot be attributed to requiring an initial contribution alone because the two probes differed in several ways.
Abstract
from arXiv · showhide
AI chatbots can help people write faster, but they can also encourage overreliance by making it easy to turn minimal input into usable text. We study turn-level design friction: intentional constraints added to each chatbot exchange that slow, limit, or redirect how users request, access, or use model responses. We designed six friction probes, organized around three mechanisms: eliciting user contribution, restricting access to generated content, and reshaping system output. In a within-subject study with 24 participants, all six probes increased workload, task duration, and perceived ownership relative to a conventional AI chatbot, while their effects on recall and recognition were more selective. We further found that participants adapted to friction in different ways, and that the same constraint could support or obstruct involvement depending on users' goals and workflows.
1 Introduction
The paper examines how turn-level design friction can preserve meaningful user involvement in AI-assisted writing without discarding practical assistance. It introduces six probes across three mechanisms and evaluates their effects on involvement, costs, memory, and workflow adaptation.
- Turn-level design friction embeds constraints within individual chatbot exchanges, changing how users request, access, interpret, or use generated responses.
- The six probes represent eliciting user contribution, restricting access, and reshaping system output, with two probes for each mechanism.
- A within-subject study with 24 participants used short writing tasks requiring participants to request, interpret, and incorporate AI-generated material.
- The mixed-methods evaluation measured workload, task duration, psychological ownership, recall, recognition, interaction logs, and interview responses.
- All six probes increased workload, task duration, and psychological ownership relative to a conventional AI chatbot, while memory and workflow effects varied.
- Friction supported involvement when effort remained directed toward formulating, interpreting, or assembling responses, but became obstructive when users mainly managed the constraint or interface.
- Participants adapted by changing how they distributed work between themselves and the system, motivating further study in longer, more naturalistic settings.
2 Background and Related Work
Prior work shows that effortless AI assistance can reduce users’ cognitive contribution and shape their judgment, but structured interaction can preserve reflection and involvement. This paper positions friction as a reusable design resource for conversational turns rather than merely a usability defect.
- LLM assistance can reduce users’ mental work and encourage reliance on model recommendations, including incorrect ones, rather than independent evaluation.
- In AI-assisted writing, limited creative decisions and dominant chatbot involvement have been associated with weaker expressive ownership and poorer immediate recall.
- The effects of AI assistance depend on how interaction is structured: independent work before assistance and prompts for analysis, explanation, and decision-making can support reasoning.
- Digital wellbeing and social-media interventions use delays, lockouts, feed redesign, and additional actions to reduce distraction and support intentional use, while introducing annoyance and frustration.
- In AI-assisted work, friction has been used to preserve judgment, reflection, or contribution through cognitive forcing, output cues, explanation, prediction, tracing, and structured revision.
- Existing interventions are often tied to particular tasks or behaviors, whereas this paper explores friction as a reusable layer applied at individual conversational turns.
3 Friction Design
The authors derive three recurring friction mechanisms from prior HCI designs and use them to develop six probes for chatbot interactions. They retain interventions that redirect effort toward the task, reasoning, or model content rather than creating obstruction.
- Mechanism synthesis: A synthesis of 24 recent HCI papers identified 73 unique designs organized mainly around eliciting user contribution, restricting access, and reshaping system output.
- Mechanism synthesis: The three mechanisms describe added user information or reasoning, delayed or conditional access, and altered system output that changes interpretation or use.
- Probe selection: The researchers selected six probes, assigning two distinct probes to each mechanism while giving every probe one primary mechanism for clearer comparison.
- Probe selection: They treated friction as productive when it redirected effort toward the task, users’ reasoning, or model-generated content instead of merely increasing difficulty or unpleasantness.
- Eliciting user contribution: Think Multiplier requires a 20-word prompt, rejects repeated padding, and caps the reply at around half the input length so users remain primary authors.
- Eliciting user contribution: User Echo requires users to draft a tentative response before receiving model feedback and additional ideas.
- Restricting access: Hold Reveal uses a two-handed key action to reveal blurred output while making selection or copying impractical, linking physical effort to access.
- Restricting access: Timer Lockout hides output until users slide a timer, then reveals it for 15 seconds before re-locking; extended delays were discarded as potentially punitive.
4 Experimental Method
The study used a within-subject design in which participants completed writing, memory, recall, and recognition tasks across frictional and baseline interfaces. Measures combined self-report, behavioral memory tests, interviews, and interaction logs.
- Study procedure: Participants completed short writing tasks with LLM-supported interfaces, followed by surveys, interviews, memory refresh, recall, and recognition procedures.
- Research questions: The study addressed users’ involvement, interaction costs, and workflow adaptation through three research questions.
- Participants: The sample included 24 participants aged 18–44, with balanced self-identified gender counts and extensive prior experience using generative AI for writing.
- Experimental design: Each participant experienced all eight interface conditions, including six frictional interfaces, ChatGPT, and Write by Yourself, with order counterbalanced using a balanced Latin square.
- Writing task: Writing responses had to contain 50–80 words and meet coherence and grammatical-completeness requirements before participants could proceed.
- Writing task: The writing interface combined an LLM interaction area, writing board, topic display, real-time word-count check, and evaluation-page control.
- Measures: Workload was measured with NASA-TLX subscales, while ownership was assessed using a 5-point personal-ownership item.
- Memory measures: A 50-problem arithmetic task preceded recall to reduce recency effects, and recognition tested identification of the exact written sentence among generated foils.
5 Results
Across the study, turn-level frictions increased workload, ownership, and task duration relative to ChatGPT, while recall and recognition effects varied by interface. Participants adapted their interaction patterns, and their responses to friction differed according to their goals and workflows.
- Workload: All friction conditions significantly increased workload relative to ChatGPT, with higher Mental Demand and Effort across conditions.Frictions also increased frustration, while Performance showed no significant differences; Physical Demand was higher for most conditions.
- Ownership: All friction interfaces significantly increased ownership relative to ChatGPT’s mean of 2.00, with Write by Yourself producing the highest ownership ratings.Ownership increases ranged from 1.00 for Hold Reveal to 2.71 for Write by Yourself.
- Task Duration: All friction conditions took significantly longer than ChatGPT’s 3.15-minute average, with Think Multiplier showing the largest increase of 3.41 minutes.All frictions also had higher average task durations than Write by Yourself.
- Recall and Recognition: Most interfaces improved free-recall performance over ChatGPT’s mean of 0.38, with Elenchus showing the largest gain of ΔM= 0.40 (p= .002).User Echo was the exception.
- Recall and Recognition: Recognition improved significantly only for Less is More (ΔM= 0.38) and Elenchus (ΔM= 0.33), with no significant differences for the remaining conditions.ChatGPT’s average recognition accuracy was 0.63.
- Adaptation: Participants adapted by distributing writing across more turns, decomposing requests, developing intermediate positions, and narrowing broad prompts.Less is More averaged 6.04 turns, while User Echo, Elenchus, and Think Multiplier also involved more turns than ChatGPT.
6 Discussion
The discussion presents design friction as a mixed, goal-dependent intervention: all six probes increased workload, task duration, and ownership, but memory effects were selective and user responses varied. Productive friction therefore depends on where effort is placed, how it fits users’ workflows, and whether it preserves meaningful involvement beyond initial contribution.
- User involvement and outcomes: All six friction probes increased perceived ownership relative to ChatGPT, while effects on recall and recognition were selective.Most probes improved recall, but only Elenchus and Less is More significantly improved recognition over ChatGPT.
- Interaction costs and adaptation: Every friction increased workload and task duration, while participants adapted rather than simply complying with the constraints.Adaptations included decomposing tasks, bringing ideas earlier, compressing responses, or working around access restrictions.
- Fit and user goals: The same constraint could support or obstruct involvement depending on users’ goals, workflows, and timing.Socratic questioning could prevent premature outsourcing when users lacked a position, but interfere when users already had an argument and wanted supporting information.
- Scope of turn-level design: Turn-level interventions are lightweight and modular, but applying the same friction across turns can overlook changing user goals.This scope boundary supports designs that adapt friction as the task evolves.
- Adaptive interaction design: Turn-level friction is better treated as a selectable repertoire than as a uniform rule applied throughout an interaction.Participants described systems that switch between reflective questions and direct assistance or let users turn friction on and off by task.
- Preserving meaningful work: Requiring an initial contribution may not preserve involvement when the model subsequently produces complete prose.User Echo and Think Multiplier both elicited contribution, but only Think Multiplier significantly improved recall, and the probes differed in several ways.
- Where to place friction: Friction may be more useful when placed in interpreting, formulating, or assembling responses than around access alone.Access restrictions sometimes reduced help-seeking or distracted participants with maintaining access rather than developing ideas.
7 Limitations and Future Work
The study’s goal was to examine turn-level friction in AI chatbot interactions, using short tasks as its research setting.
- The study focused on turn-level friction in AI chatbot interactions.
- Short writing tasks were used to investigate chatbot interactions.
- The research examined friction at the level of individual chatbot turns.
Writing tasks were short and bounded.
The study used short, bounded writing tasks, limiting how broadly its findings should be interpreted. Longer and more varied task contexts may reveal different patterns of chatbot use and friction effectiveness.
- Short writing tasks provided a bounded setting for evaluating friction across chatbot interactions.Participants could request, interpret, and incorporate model output under each condition.
- Friction effectiveness was not universal; its value depended on users’ goals and workflows.
- Participants used the chatbot differently across task stages and had different expectations for support.
- A friction productive during ideation could obstruct information seeking.
- Future work should test turn-level friction in longer tasks with more diverse chatbot use.
Static rather than adaptive friction.
The prototypes applied friction in fixed ways, even as users’ intentions and needs changed. This rigidity could turn constructive friction into frustration, motivating more adjustable designs.
- The prototypes applied friction rigidly, regardless of users’ intentions or changing needs.
- Fixed constraints sometimes caused frustration when participants shifted between deeper engagement and efficient information access.
- Future systems could use more legible, adjustable friction, including lighter-weight nudges.
- Usage Mirror suggests self-monitoring as one possible direction for future friction design.
- Most frictions could be avoided when users were motivated enough to turn them off.
Optional frictions and user workarounds.
Users could often avoid friction by disabling it, so future work should examine constraints that users voluntarily keep enabled. User-controlled activation over time is one proposed direction.
- Users could avoid most frictions when sufficiently motivated to turn them off.
- One alternative is designing frictions that users are more likely to leave enabled voluntarily.
- Study Mode is presented as an example of a user-controlled, step-by-step learning experience.
- Future work should investigate when and why users activate optional frictions over days or weeks.
8 Conclusion
The study evaluated six turn-level friction probes in writing with 24 participants, finding increased ownership, workload, and task duration, alongside design-dependent memory effects and workflow adaptation.
- Six turn-level friction probes were evaluated in a within-subject writing study with 24 participants.The probes examined how friction affects user involvement, interaction costs, and workflow adaptation.
- Friction consistently increased perceived ownership, workload, and task duration compared with conventional chatbot interaction.
- Memory outcomes varied across friction designs rather than following a uniform pattern.
- Participants adapted their workflows in response to different forms of friction.
- Productive friction depends on preserving meaningful human work when it matters, not simply making AI interaction harder.
A Exploratory Corpus of Prior Friction Designs
The exploratory corpus organizes prior friction designs into recurring interaction mechanisms and uses that organization to derive a structured design-mechanism framework.
- The corpus contains 73 unique designs from 24 papers published within the past five years.
- Four designs instantiated two mechanisms, producing 77 design-mechanism entries from the 73-design corpus.
- The resulting corpus is presented in Table A.1 as the basis for deriving the three friction mechanisms.
- Examples include interfaces that require users to write more, explain or repair generated code, delay or selectively reveal outputs, and interrupt continued use.
D Recall Scoring Pipeline
The appendix presents materials for recall scoring and per-intervention comparisons across workload, NASA-TLX dimensions, ownership, task duration, recall, and recognition.
- Table D.2 presents recall scoring examples.
- Table E.3 compares overall workload scores for each intervention with the ChatGPT control.
- Table E.4 reports NASA-TLX dimension-level effects across interventions versus ChatGPT.
- Table E.5 compares ownership scores and Table E.6 compares task duration for each intervention against ChatGPT.
- Recognition accuracy is shown by interface as mean binary scores with standard-error bars.