Source-linked AI summary
RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces
Nikhil Wani
TL;DR
Conversational AI interfaces can burden users with repeated re-prompting and context management, producing fatigue and cognitive load. RecalibrateGPT addresses this with five cross-turn operators and single-click controls over the full conversation history, and its pilot evaluation reports lower workload and high usability.
Problem
Conversational AI interfaces often require repeated re-prompting and context management, creating fatigue, cognitive load, and possible task abandonment.
Method
RecalibrateGPT uses five cross-turn operators to recalibrate responses from the full conversation history through a structured panel and single-click AssistiveButton interaction.
Results
NASA-TLX was M=2.7 versus 5.4, while SUS was M=86.5; Anchor, Replay, and Delta were rated most useful.
Takeaways & Limitations
The findings suggest conversational AI fatigue includes an interaction-flow cost that interfaces can remove, beyond model-quality issues.
Abstract
from arXiv · showhide
Large language models are powerful, but their interfaces often devolve into a type $\rightarrow$ read $\rightarrow$ retype loop, creating conversational AI fatigue, cognitive load, and eventual task abandonment. To mitigate this, we present RecalibrateGPT, a system introducing five cross-turn operators (Anchor, Replay, Delta, Scope, and Steer) that each target a distinct fatigue type, recalibrating LLM responses through a structured panel by acting on the full conversation history with a single click. Users invoke these operators through the AssistiveButton in one of three operator palette layouts: Vertical, Arc, or Tablet. We conducted two pilot studies with the same 12 advanced LLM users. An initial formative qualitative study identifies a taxonomy of four fatigue types (retyping, scanning, decision paralysis, and context drift) and derives two design objectives for RecalibrateGPT. A follow-up quantitative evaluation finds it reduces perceived cognitive workload by half (NASA-TLX = 2.7) at high perceived usability (SUS = 86.5), suggesting AI fatigue is not just a model-quality issue but an interaction-flow cost that interfaces can remove.
1 Introduction and Related Work
RecalibrateGPT frames conversational AI fatigue as an interaction-flow problem in addition to a model-quality issue. It replaces repeated full re-prompting with one-click, cross-turn response calibration.
- Type → read → retype loops make users restate goals, correct drift, track changing context, and manage growing conversation history.The resulting fatigue is associated with cognitive load, frustration, decision paralysis, and possible abandonment.
- Full re-prompting increases per-token inference cost and compounds the burden of every additional turn.
- RecalibrateGPT introduces five cross-turn operators that act on full conversation history through a structured panel with one click.The AssistiveButton offers Vertical, Arc, and Tablet operator palettes.
- Prior systems largely manipulate a single-turn response, leaving session-level conversational fatigue insufficiently addressed.The related-work gap is described across reusable prompt controls, response navigation, editing surfaces, and prompt pipelines.
2 Study 1: Formative Study
A formative study with 12 advanced LLM users identified four recurring conversational fatigue types and translated them into two design objectives for RecalibrateGPT.
- The study recruited 12 advanced LLM users with sustained, multi-platform, and high-stakes information-seeking experience.Participants completed an online survey about interaction frustrations and retyped corrections.
- Participants produced four fatigue themes: retyping, scanning, decision paralysis, and context drift.These themes describe repeated constraints, rereading, uncertainty about choices, and fighting the interface instead of solving the problem.
- The findings motivated DO1, multi-turn response calibration, by grounding cross-turn operators in recurring interaction breakdowns.
- DO2, single-click corrective control, motivated the AssistiveButton and three-way toggle to reduce retyping burden.
3 RecalibrateGPT System Design
RecalibrateGPT implements five operators that target distinct fatigue types while recalibrating responses from full conversation history. Users access them through a shared AssistiveButton and selectable palette layouts.
- Anchor, Replay, Delta, Scope, and Steer target context drift, scanning, retyping, retyping, and decision paralysis, respectively.Their outputs include goal reorientation, session digests, semantic differences, subtopic expansion, and follow-up questions.
- The three-way toggle selects Vertical, Arc, or Tablet palette layouts for repeated use, rapid micro-corrections, or persistent single-click access.
- The AssistiveButton is the shared entry point that reads the selected mode and renders its corresponding operator palette.Its placement adapts across Vertical, Arc, and Tablet layouts.
- The backend uses conversation history, the original goal, the latest response, embeddings, cosine similarity, KL divergence, and structured summaries to implement the operators.Replay maps history to established facts, open questions, and a next step.
4 Study 2: User Feedback on RecalibrateGPT
A within-subjects pilot compared standard chat with RecalibrateGPT for the same 12 participants across five operators. Participants preferred RecalibrateGPT as less fatiguing and rated it highly usable, though the authors characterize the findings as directional.
- The follow-up pilot used the same 12 participants and compared standard chat with RecalibrateGPT side-by-side for each operator.Participants viewed RecalibrateGPT in their preferred palette layout and rated both interfaces on NASA-TLX.
- NASA-TLX: M=2.7 vs. 5.4, with participants consistently selecting RecalibrateGPT as less fatiguing.
- SUS: M=86.5, indicating high perceived usability in the pilot.
- Anchor (n=4), Replay (n=3), and Delta (n=3) were rated most useful.
- Given the pilot scale, the authors report the findings as directional feasibility evidence rather than generalizable effects.
5 Conclusion and Future Work
RecalibrateGPT uses five cross-turn operators and single-click interactions over full conversation history to target conversational AI fatigue. In a pilot, it halved perceived cognitive workload while maintaining high perceived usability.
- Five cross-turn operators act on the full conversation history to target conversational AI fatigue through calibrated responses and single-click interactions.
- Future work will investigate proactive interfaces that surface operators as fatigue emerges and validate them in larger-scale studies.
A. Appendix
The appendix includes a figure showing two healthcare conversations calibrated with Scope and Steer in the Vertical palette. It also provides supplementary material for completeness while directing results to the main paper.
- Figure A depicts two drifting healthcare conversations calibrated by Scope and Steer in the Vertical palette layout.
- The appendix states that its summary and figures are provided for completeness, while results are reported in the main paper.
A.1 Study 2 Procedure
Study 2 used an online, three-phase within-subjects pilot with 12 participants evaluating five operator-specific healthcare-interface comparisons. Participants selected the less fatigue-inducing interface, rated workload, and then assessed usability and provided feedback.
- Study setup: Sessions lasted approximately 35 minutes online, used three phases, received ethics review, and obtained informed consent.
- Phase 1: Onboarding: Phase 1 briefed participants, re-verified inclusion criteria, and introduced the scenario during a five-minute onboarding.
- Phase 2: Paired Evaluation: Phase 2 used a within-subjects design with 12 participants comparing standard-chat and RecalibrateGPT interfaces in five operator-specific Type 2 diabetes evaluations.
- Phase 2: Paired Evaluation: Participants selected the less fatigue-inducing interface and rated both conditions on six averaged NASA-TLX dimensions using parallel 7-point scales.
- Phase 3: Usability and Debrief: After the comparisons, participants completed the 10-item SUS and a semi-structured interview about retyping, scanning, decision-making, and context maintenance.
A.2 Participant Feedback
All participants preferred RecalibrateGPT as less fatigue-inducing, reported lower workload, and rated its usability highly. Anchor, Replay, and Delta were most frequently identified as useful.
- All 12 participants selected RecalibrateGPT as the less fatigue-inducing interface.
- NASA-TLX averaged 2.7 with RecalibrateGPT versus 5.4 for standard chat.
- SUS averaged 86.5, indicating high perceived usability for the RecalibrateGPT panel.
- Anchor, Replay, and Delta were most frequently identified as useful during debrief, with counts of 4, 3, and 3 respectively.
A.3 Measurement Instruments
The study used unweighted NASA-TLX and SUS to measure perceived workload and system usability. NASA-TLX covered six workload dimensions on a 7-point scale, while SUS used 10 Likert items scored 0–100.
- NASA-TLX: NASA-TLX measured perceived workload across six dimensions using an unweighted 7-point scale.Ratings ran from 1 (Very Low) to 7 (Very High), with reversed Performance anchors so higher values consistently indicated greater workload.
- NASA-TLX: The NASA-TLX dimensions covered mental demand, temporal demand, performance, effort, and frustration, with the instrument describing task pace, success, work, and negative affect.The supplied instrument passages explicitly define temporal demand, performance, effort, and frustration; mental demand is named in the six-dimension description.
- SUS: System usability was assessed with the 10-item System Usability Scale using a 5-point Likert response format and a 0–100 score.The items addressed complexity, ease of use, technical support, integration, inconsistency, learnability, cumbersomeness, confidence, and preparation required before use.