Source-linked AI summary

What Are You Doing? Effects of Intermediate Feedback from Agentic LLM In-Car Assistants During Multi-Step Processing

Johannes Kirmayr, Raphael Wennmacher, Khanh Huynh, Lukas Stappen, Elisabeth André, Florian Alt

arXiv:2602.15569v2cs.HC

TL;DR

Agentic assistants must communicate progress during extended, multi-step tasks without creating excessive distraction or cognitive load, particularly in driving contexts. This paper studies feedback timing and verbosity through a controlled mixed-methods in-car study with N=45, finding that intermediate feedback improved several user outcomes and that users preferred adaptive verbosity.

  • Problem

    Agentic systems’ extended processing and expanded information volume leave open how feedback should balance responsiveness, trust, grounding, and task load, especially during driving.

  • Method

    A controlled mixed-methods study with N=45 compared intermediate planned-step and result feedback with silent operation across task durations and single- versus dual-task contexts.

  • Results

    Intermediate feedback improved perceived speed, user experience, and trust while reducing task load compared with end-only delivery across task complexities and interaction contexts.

  • Takeaways & Limitations

    Feedback should begin transparently, become more concise as reliability is demonstrated, and expand again for novel, ambiguous, or high-stakes requests.

  • Takeaways & Limitations

    The study used a standardized lane-keeping simulation and recruited participants from a single automotive company, limiting generalization to real-world driving and broader populations.

Abstract

from arXiv · show

Agentic AI assistants that autonomously perform multi-step tasks raise open questions for user experience: how should such systems communicate progress and reasoning during extended operations, especially in attention-critical contexts such as driving? We investigate feedback timing and verbosity from agentic LLM-based in-car assistants through a controlled, mixed-methods study (N=45) comparing planned steps and intermediate results feedback against silent operation with final-only response. Using a dual-task paradigm with an in-car voice assistant, we found that intermediate feedback significantly improved perceived speed, trust, and user experience while reducing task load - effects that held across varying task complexities and interaction contexts. Interviews further revealed user preferences for an adaptive approach: high initial transparency to establish trust, followed by progressively reducing verbosity as systems prove reliable, with adjustments based on task stakes and situational context. We translate our empirical findings into design implications for feedback timing and verbosity in agentic assistants, balancing transparency and efficiency.

1 Introduction

Agentic assistants create open questions about how to communicate progress, timing, and detail during extended tasks, especially when users are also driving. This study examines these issues and finds benefits from intermediate feedback alongside preferences for adaptive verbosity.

  • Agentic systems autonomously decompose requests, invoke tools, and synthesize results across extended processing periods.
  • Driving contexts make feedback design safety-critical because excessive information can distract or overload users, while silence can undermine trust and perceived responsiveness.
  • Existing systems range from nearly silent operation to verbose step-by-step narration, demonstrating a lack of shared principles for agentic feedback.
  • The mixed-methods study with N=45 participants investigates feedback timing, interaction context, and adaptive verbosity for an LLM-based in-car assistant.
  • Intermediate planned-step and result feedback improved perceived speed, user experience, and trust while reducing task load versus silent processing with final-only responses.
  • Interviews favored detailed initial transparency that becomes more concise with demonstrated reliability, then expands for novel, ambiguous, or high-stakes requests.

2 Related Work

Related work frames agentic feedback as a balance among transparency, responsiveness, grounding, trust, cognitive resources, and distraction. In-car interaction intensifies these trade-offs because voice interaction can reduce—but not eliminate—driving-related interference.

  • Minimal and verbose agentic feedback embody competing priorities between uninterrupted workflow and transparency, with intermediate detail potentially distracting trusted expert users.
  • Grounding communication requires continuously updating shared knowledge about both interaction content and process.
  • Response delays can reduce satisfaction, increase frustration, and lead voice-interface users to infer failure or error during waiting.
  • Trust in human–AI interaction depends on users understanding what they can expect when delegating complex requests under uncertainty and vulnerability.
  • Human oversight includes passive monitoring and active intervention, enabling error detection and correction while user control can increase trust.
  • Auditory-vocal interaction generally interferes less with vehicle control than visual-manual interaction, but even basic voice tasks increase cognitive workload versus undistracted driving.

3 Research Questions

The paper asks how agentic assistants should time, contextualize, and vary the detail of feedback during long-running tasks. These questions target responsiveness, trust, workload, distraction, and adaptation over time.

  • RQ1 examines whether feedback during task execution versus only at completion changes perceived waiting time, experience, trust, and cognitive workload.
  • RQ2 examines how longer processing times and active driving shape preferences for when and how much feedback users receive.
  • RQ3 examines how feedback verbosity should adapt with familiarity and situational context while balancing information, distraction, and trust.

4 User Study

The user study used a controlled, mixed-methods in-car setup to compare feedback timing across task duration and interaction contexts. Participants completed a within-subject factorial experiment with questionnaires and follow-up interviews.

  • Participants interacted through a voice interface and center-console graphical display while completing stationary or lane-keeping driving conditions.
  • The apparatus combined an external speaker, tablet display, and mouse-controlled lane-keeping simulation to manipulate cognitive load consistently.
  • Forty-five participants completed preparation, eight experimental tasks with interleaved questionnaires, and post-task interviews in approximately 60-minute sessions.
  • The quantitative study used a within-subject 2×2×2 design varying feedback timing, task duration, and interaction context.
  • No Intermediate feedback acknowledged the request and then remained silent, whereas Planning & Results provided intermediate updates before the final response.
  • Task duration varied between medium tasks with 3 intermediate steps and long tasks with 6 intermediate steps, with PR updates presented at fixed 5-second intervals.
  • The study measured perceived speed, task load, user experience, and trust, with detailed duration effects analyzed only for perceived speed.

4.4 Qualitative Study

The qualitative study used semi-structured interviews and thematic analysis to examine adaptive feedback preferences, alongside a controlled participant sample and multimodal study setup. Limitations constrain generalization to real-world driving, longitudinal adaptation, and modality combinations.

  • Interview procedure: Interviews addressed desired feedback amount, uncertainty communication, and behaviors that foster long-term use.Follow-up prompts were used to clarify participants’ responses.
  • Study and analysis: 45 participants from a single automotive company completed semi-structured interviews analyzed through thematic analysis.The sample included adults aged 18–64 with varied familiarity with LLMs, voice assistants, and the company’s in-car assistant.
  • Qualitative analysis: Two researchers independently open-coded 20% of transcripts before refining conceptual boundaries and identifying five themes.The analysis distinguished external real-time adaptation from internal adaptations involving task ambiguity and novelty.
  • Limitations: The qualitative analysis and participant sample were limited by single-company recruitment, simulated driving, self-reported adaptation, and fixed feedback intervals.The final codebook and theme structure were made available online.
  • Experimental scope: The driving context used a standardized lane-keeping task, while feedback was delivered simultaneously through voice and visual channels.These choices established a controlled baseline but limited exploration of real-world variability and modality combinations.

5 Results

Intermediate planned-step and result feedback outperformed silent processing with final-only delivery across perceived speed, task load, user experience, and trust. Benefits generally held across task durations and interaction contexts, while interviews supported context- and reliability-sensitive adaptation of verbosity.

  • Perceived speed: Intermediate feedback buffered the negative effect of longer tasks on perceived speed, especially in the stationary single-task condition.The Timing × Duration interaction was significant (p=.049).
  • Task load: Task load was lower with intermediate feedback than final-only feedback, with a small effect (dz=−0.26).The reduction was primarily driven by Frustration; Mental and Temporal Demand did not differ significantly.
  • User experience: Intermediate feedback improved the overall user experience composite with a medium effect (dz=0.54).Attractiveness, Dependability, and Risk Handling all improved, with the strongest effect for Risk Handling (dz=0.60).
  • User trust: Intermediate feedback increased trust with a small effect (dz=0.38), improving Reliability and Trustworthiness but not Confidence.No significant order effect or Timing × Order interaction was observed.
  • Adaptive feedback: Interviewees preferred verbosity to decrease as systems proved reliable, but increase for ambiguity, novelty, and high-stakes actions.Preferences diverged for media and passengers, motivating user controls such as mute or dampening commands.

6 Discussion & Implications

Intermediate, content-rich feedback improves responsiveness, trust, user experience, and task load, especially for longer tasks. Users favor verbosity that adapts to demonstrated reliability, task characteristics, and situational context, although transfer beyond compatible dual-task settings and short execution times remains limited.

  • Feedback Timing: Intermediate updates improved perceived speed, user experience, and trust while reducing frustration and task load compared with final-only feedback.Reported effects were large for perceived speed, medium for user experience, and small for trust, frustration, and task load.
  • Feedback Timing: Intermediate updates buffered perceived-speed declines on longer, more complex tasks, with stronger effects when the assistant was the primary task.Benefits also appeared during dual-task interaction, but moderation effects were weaker.
  • Feedback Content: Content-bearing updates preserve grounding, maintain trust, and distribute cognitive effort more evenly than progress-only cues.The benefit depends on the feedback channel being available without becoming interruptive.
  • Long-term Adaptation: Users preferred verbosity to begin high and decrease as repeated successful interactions demonstrated system reliability.The findings connect reduced detail with learned trust rather than with a fixed preference for silence or verbosity.
  • Situational Adaptation: Users sought more transparency for novel, ambiguous, or high-stakes tasks, while external distractions produced divergent preferences best addressed through user controls.The authors recommend increasing verbosity for demanding task factors, reducing it for routine or low-stakes requests, and offering overrides when context preferences differ.
  • Applicability Across Domains: The implications may apply beyond in-car assistants when primary and secondary activities use different cognitive channels, but transferability requires further investigation.Intermediate updates may interfere when both tasks share the same channel, and the study does not establish a uniform policy for those settings.
  • Temporal Scope: The design implications target execution times from multiple seconds to one minute rather than multi-minute deep-agent operations.Sustained feedback over much longer periods may overwhelm users and require a transition to background processing.
  • Future Design and Technical Challenges: Situational verbosity adaptation remains technically constrained because ambiguity detection is difficult and current LLM confidence assessments are poorly calibrated.Novelty and stakes may be estimable in some settings, whereas reliable ambiguity detection remains an open challenge.

7 Conclusion

The paper addresses how agentic assistants should communicate during long-running tasks, especially when users’ cognitive load is constrained. A controlled mixed-methods in-car study finds that intermediate updates improve key user outcomes, while adaptive verbosity should vary with reliability and situational demands.

  • The study examines how agentic assistants should communicate progress and manage information load during long-running tasks.
  • Intermediate informative updates improved trust, perceived speed, and user experience while reducing task load.
  • Feedback should begin transparently, become more concise as reliability is demonstrated, and expand again for ambiguous, novel, or high-stakes tasks.
  • The findings inform adaptive-feedback design beyond driving where intermediate communication can use a compatible cognitive channel.

A Survey

The survey section introduces questions translated from German into English.

  • The following survey questions were translated from German into English.

A.1 Demographics & Technical Familiarity

The survey measures demographic characteristics and participants’ familiarity with LLMs, voice assistants, and the company’s in-car voice assistant.

  • Participants report age, gender identity, familiarity with LLMs, familiarity with voice assistants, and familiarity with the company’s in-car voice assistant.

A.2 Questionnaires

The questionnaires measured perceived speed, user experience, task load, trust, and perceived product risks using standardized rating scales. They also assessed participants’ confidence, reliability judgments, and the demands they experienced during the task.

  • A.2.1 Perceived Speed.: Perceived speed was rated from very slow (1) to very fast (7).
  • A.2.2 User Experience - UEQ+ subset.: User experience included attractiveness ratings from damaging (-3) to not damaging (3), with importance judgments on a 1–7 scale.
  • A.2.2 User Experience - UEQ+ subset.: User experience also assessed annoyance, pleasantness, friendliness, and overall goodness using bipolar -3 to 3 scales.
  • A.2.2 User Experience - UEQ+ subset.: Dependability covered predictability, supportiveness, security, and whether the product met expectations, each paired with importance ratings.
  • A.2.2 User Experience - UEQ+ subset.: Risk handling measured whether application errors and risks felt threatening, hazardous to health, or harmless on bipolar scales.
  • A.2.3 Task Load - NASA-RTLX subset.: Task load was assessed across six dimensions, including mental demand, temporal demand, and frustration, on low (0) to high (100) scales.
  • A.2.4 User Trust - S-TIAS.: Trust measures asked participants to rate confidence in the assistant, system reliability, and trust from not at all (1) to extremely (7).

A.3 Semi-structured Interview Questions

The semi-structured interviews examined preferred verbal feedback, uncertainty handling, and behaviors that could foster long-term trust in the assistant.

  • Participants were asked how much verbal feedback they would like from the system.
  • Participants discussed whether the system should notify them about uncertainty or decide autonomously, including preferred communication methods.
  • Interview questions explicitly situated feedback preferences within driving conditions, passengers, music, and other distractions.
Loading 2602.15569v2…