Source-linked AI summary

Understanding Large-Language Model (LLM)-powered Human-Robot Interaction

Callie Y. Kim, Christine P. Lee, Bilge Mutlu

arXiv:2401.03217v1cs.ROcs.HC

TL;DR

The paper addresses limited evidence about the design requirements for integrating LLMs with robots, particularly how those requirements vary from text and voice interaction and across tasks. It compares text, voice, and robot agents in a 32-participant study spanning four conversational tasks. LLM-powered robots elicited expectations for sophisticated non-verbal cues and were favored for connection-building and deliberation, but faced communication difficulties and could induce anxiety.

  • Problem

    Research has limited evidence about the distinctive design requirements for using LLMs with robots and which tasks most benefit from their capabilities.

  • Method

    A user study with 32 participants compared text, voice, and robot agents across choose, generate, execute, and negotiate tasks.

  • Results

    LLM-powered robots elicited expectations for sophisticated non-verbal cues and were more favored for connection-building and social deliberation, but less preferred amid communication difficulties and anxiety.

  • Takeaways & Limitations

    Robot embodiment and task context should inform design recommendations for LLM-powered robots and the LLMs used with them.

  • Takeaways & Limitations

    The study did not compare the LLM-powered robot with a non-LLM-powered robot, limiting conclusions about the unique effects of integrating LLMs into robots.

Abstract

from arXiv · show

Large-language models (LLMs) hold significant promise in improving human-robot interaction, offering advanced conversational skills and versatility in managing diverse, open-ended user requests in various tasks and domains. Despite the potential to transform human-robot interaction, very little is known about the distinctive design requirements for utilizing LLMs in robots, which may differ from text and voice interaction and vary by task and context. To better understand these requirements, we conducted a user study (n = 32) comparing an LLM-powered social robot against text- and voice-based agents, analyzing task-based requirements in conversational tasks, including choose, generate, execute, and negotiate. Our findings show that LLM-powered robots elevate expectations for sophisticated non-verbal cues and excel in connection-building and deliberation, but fall short in logical communication and may induce anxiety. We provide design implications both for robots integrating LLMs and for fine-tuning LLMs for use with robots.

1 INTRODUCTION

This study examines how LLM-powered robots should be designed differently from text- and voice-based agents across conversational tasks and contexts. It compares four task types and finds that robot embodiment raises expectations for non-verbal communication while producing both social benefits and communication challenges.

  • The study investigates perceptions and preferences for LLM-powered robots compared with text- and voice-based agents.It focuses on design requirements tailored to robot embodiment, tasks, and contexts.
  • Users engaged with four conversational tasks: choose, generate, execute, and negotiate.The study uses these task categories to assess where LLM-powered robots may be most suitable.
  • LLM-powered robots elicit expectations for sophisticated non-verbal cues and are preferred for connection-building and social deliberation.These advantages are associated with the robot’s embodied interaction capabilities.
  • Robots are less preferred when rich social capabilities produce verbose responses, logical or communication errors, or anxiety.The findings identify both interaction benefits and risks for LLM-powered robot design.
  • The paper provides design implications for developing LLM-powered robots and adapting LLMs for future human-robot interaction.

2 RELATED WORK

Prior work frames robots as embodied social agents that connect physical environments with LLM capabilities. Research has applied LLMs to conversational robots, companionship, and task-oriented interaction, but these applications also expose practical challenges.

  • Physical embodiment can provide communication channels including gesture, posture, gaze, facial expression, proxemics, and social touch.Prior studies associate physically embodied robots with higher engagement, enjoyment, trust, and empathy than several other agent forms.
  • Robots connect tangible environments and LLMs, allowing sensor-based environmental inference alongside semantic comprehension and flexible dialogue.These capabilities support applications such as task planning and human-robot collaboration.
  • Existing conversational-robot studies use LLMs for facial expressions, personalized companionship, and improved well-being-related interaction.The cited work also examines open-domain dialogue challenges with older adults.

3.1 Embodiment Design

The study compares three LLM-equipped agent embodiments: text, voice, and a physically embodied social robot. The conditions use consistent underlying models while differing in interaction modality and embodiment.

  • All three agents used GPT-3.5 or text-davinci-003 without fine-tuning, with identical temperature and token settings.Pre-prompts specified the four study tasks.
  • The text agent exchanged prompts through typed text using the GPT model’s API.
  • The voice agent used spoken interaction while the participant and robot were separated by a screen.Speech was converted to text before being forwarded to the GPT model.
  • The same robot was used in the voice and robot conditions to keep voice interactions consistent across embodiments.
  • Pepper provided animated gestures, text-to-speech, and face recognition through the Pepperchat system.The robot used a minimalist design emphasizing basic embodiment.

3.2 Task Design

The task design is based on the Group Task Circumplex Model, which organizes tasks along conflict-to-cooperation and conceptual-to-behavioral dimensions. The study uses four task categories to structure conversational interactions.

  • The task circumplex spans conflict-based to cooperative and conceptual to behavioral dimensions.
  • Negotiate tasks resolve conflicts among viewpoints, interests, and motives.
  • Execute tasks involve carrying out a plan or performance.

generate:

The study organizes conversational interaction into four task types—generate, choose, negotiate, and execute—each requiring a distinct participant-agent collaboration.

  • generate: Generate tasks involve collaboratively creating ideas or plans, exemplified by taking turns adding sentences to construct an imaginary story.Participants developed characters, settings, obstacles, solutions, and climaxes through shared story construction.
  • choose: Choose tasks require selecting a solution or plan from alternatives, exemplified by discussing practical items for a ski, beach, or camping trip.Participants discussed item criteria with the agent before finalizing their selections.
  • negotiate: Negotiate tasks involve resolving conflicting viewpoints, interests, or motives, exemplified by bargaining over the price of a second-hand item.The seller sought the highest price while the buyer sought the lowest, within a fixed minimum-price constraint.
  • execute: Execute tasks involve carrying out a plan or performance, exemplified by an agent instructing participants to prepare a beverage.Participants followed instructions and could ask questions when confused.

4 USER STUDY

The user study compared text, voice, and robot agents across four tasks using a mixed-factorial design, questionnaires, interviews, behavioral measures, and thematic analysis.

  • Study design: 32 participants were randomly assigned to generate, choose, execute, or negotiate and interacted with text, voice, and robot agents in counterbalanced order.Task was a between-subjects factor, while agent embodiment was a within-subjects factor.
  • Study design: After each agent interaction, participants completed questionnaires and a semistructured interview about their experience.Sessions were conducted in person and audio- and video-recorded.
  • Measures: Perceptions were measured with modified Godspeed scales covering animacy, anthropomorphism, likeability, intelligence, and safety.The scales used five-point ratings; interaction satisfaction was measured separately on a seven-point scale.
  • Measures: The safety subscale initially showed Cronbach’s α = −0.27 because of a miscoded item, then reached Cronbach’s α = 0.72 after recoding.The still-surprised item was flipped to align its semantic direction with the other safety items.
  • Measures: Behavioral measures included participants’ total input tokens and two interaction-failure categories: technical errors and LLM hallucinations.Token counts indexed dialogue-input length, while failures included interruptions, inaccurate ASR transcription, and nonsensical or unfaithful responses.
  • Analysis: Factorial repeated-measures ANOVA tested task and embodiment effects, followed by Tukey HSD for significant pairwise differences and thematic analysis for qualitative data.Tukey HSD was used to control Type I error across comparisons.

5 RESULTS

Results varied by task and embodiment: robots supported engagement, learning, and rapport in execution and negotiation, but communication and logic problems reduced their suitability for choice and generation.

  • Quantitative measures: F(2, 56) = 14.30, p < .001: embodiment significantly affected input-prompt length, with longer prompts for text than voice or robot agents.Input length also differed by task and embodiment, F(6, 56) = 4.25, p = .001.
  • Execute: In execution, contextual conversation supported concise instructions, follow-up answers, engagement, and learning despite minimal failures to respond logically.Participants frequently asked for guidance while preparing a beverage.
  • Execute: Robot gaze, head tilts, eye contact, waiting, encouragement, and physical presence increased some participants’ receptiveness, focus, immersion, and companionship during execution.Six participants preferred the robot for efficiency and engagement, while four specifically valued its social cues.
  • Negotiate: In negotiation, contextual understanding supported coherent discussion, while the robot’s gaze, facial expressions, body movements, and physical presence helped establish rapport and trust.Four participants found the robot most effective for connection and rapport, although text was more convenient for some information exchange.
  • Choose: In choice, text enabled more accurate and expeditious information exchange, whereas robot interactions were inefficient and sometimes produced logical inconsistencies.Participants reported unusual recommendations and time-consuming discussions while iteratively validating item selections.
  • Generate: In generation, extensive prompts and real-time communication demands created barriers to creative collaboration, including difficulty expressing ideas spontaneously and processing the agent’s contributions.The creative, turn-taking process required participants to formulate personal, situation-specific prompts and consider how to incorporate suggestions.

6 DISCUSSION

The discussion argues that LLM-powered robots create heightened expectations for non-verbal behavior and fit tasks differently, while also introducing hallucination risks and study-design limitations.

  • Discussion: LLM-powered robots elicited expectations for sophisticated non-verbal cues and were favored in execution and negotiation but less favored in choice and generation.The latter tasks involved communication difficulties and potential anxiety during collaboration.
  • Combining LLMs with non-verbal cues: Users’ expectations for rich non-verbal cues arose from the robot’s advanced language capabilities, not solely from its physical form.The authors recommend aligning gaze, gestures, behaviors, and facial expressions with spoken interaction.
  • Task characteristics: Task characteristics motivate customization and fine-tuning: existing LLMs may suit execute and negotiate, whereas choose and generate require adaptation.Suggested adaptation includes simplifying rich social descriptions to improve efficiency, intuitiveness, and directness.
  • Robot design: LLMs can help robots adapt to broader user requests and preferences, supporting personalized experiences through iterative and engaging interactions.These capabilities may substitute for or complement traditionally challenging robot-design tasks.
  • Robot design: LLM integration can produce context deviation, hallucinations, unexpected behavior, and mismatches between situational context and intended robot personality.The authors recommend explicit action boundaries supported by curated data, verification, human review, and fine-tuning.
  • Limitations: The study’s interpretation is limited by the absence of a non-LLM robot comparison and by high variance in subjective quantitative measures associated with the sample size.These constraints limit assessment of LLM-specific effects and generalizability.

7 CONCLUSION

The study identifies design requirements for LLM-connected robots and the conversational tasks where they excel. LLM-equipped robots raise expectations for non-verbal cues and support connection building and deliberation, while presenting communication and social-pressure challenges.

  • LLM-equipped robots enhance user expectations for sophisticated non-verbal cues.
  • They excel in connection building and deliberation across conversational tasks.
  • They face challenges involving communication difficulties and social pressure.
Loading 2401.03217v1…