Source-linked AI summary

A Review of Verbal and Non-Verbal Human-Robot Interactive Communication

Nikolaos Mavridis

arXiv:1401.4994v1cs.ROcs.CL

TL;DR

Human-robot communication research must move beyond limited command-response interaction toward fluid verbal and non-verbal communication grounded in physical contexts. The paper synthesizes this area through ten desiderata and concludes that, despite substantial progress, many sub-problems remain unsolved and may benefit from massive real-world data.

  • Problem

    Many systems still lack integrated purposeful speech and motor planning, while connecting linguistic symbols to the physical world remains a central challenge.

  • Method

    The paper reviews verbal and non-verbal human-robot communication, organizes research around ten desiderata, and examines relevant work in detail.

  • Results

    The review finds that research has moved beyond canned commands and responses but remains far from fluid, natural verbal and non-verbal human-robot communication.

  • Takeaways & Limitations

    Future progress is framed around unresolved sub-problems including symbol grounding, integrated speech-motor planning, conversational turn-taking, and learning from massive real-world data.

Abstract

from arXiv · show

In this paper, an overview of human-robot interactive communication is presented, covering verbal as well as non-verbal aspects of human-robot interaction. Following a historical introduction, and motivation towards fluid human-robot communication, ten desiderata are proposed, which provide an organizational axis both of recent as well as of future research on human-robot communication. Then, the ten desiderata are examined in detail, culminating to a unifying discussion, and a forward-looking conclusion.

I. INTRODUCTION: HISTORICAL OVERVIEW

Conversational robots with natural-language abilities emerged in the 1990s, after a much longer history of robots and automata. Early systems motivated a research agenda centered on more natural, flexible, and broadly applicable human-robot communication.

  • Historical emergence: Conversational robots such as MAIA, RHINO, and AESOP began appearing in the early 1990s, despite earlier traditions of automata and speech synthesis.The systems served different domains, including object delivery, museum guidance, and surgery.
  • Early limitations: Early conversational robots accepted only small sets of simple canned commands and produced canned responses.Their interaction capabilities ranged from Polly’s predetermined tours to TJ’s keyboard-mediated responses and RHINO’s fixed tours.
  • Early limitations: These systems generally handled requests rather than diverse speech acts, supported mainly human-initiative dialogue, and lacked situated language and affective speech.Their verbal interaction was not flexibly mixed initiative and rarely addressed physical situations or emotion-carrying prosody.
  • Early limitations: Early robots had almost no non-verbal communication, lacking recognition or production of gestures, gait, facial expressions, and head nods.Their dialogue systems also operated largely as stimulus-response or stimulus-state-response systems without purposeful speech generation integrated with motor planning.
  • Early limitations: No real offline or on-the-fly learning occurred in these systems, so their verbal behaviors had to be prescribed.This limitation extended the dependence on designer-specified behavior.
  • Research agenda: The paper turns these shortcomings into ten desiderata, reviews research addressing them, gives special attention to symbol grounding, and discusses open problems.It motivates natural communication through first-principles reasoning and application domains.

III. DESIDERATA - WHAT MIGHT ONE NEED FROM A CONVERSATIONAL ROBOT?

The paper frames conversational-robot design around desiderata that extend beyond simple command execution. These requirements target richer speech, initiative, situatedness, embodiment, planning, learning, and broader interaction capabilities.

  • Framework: The proposed desiderata are explicitly provisional rather than exhaustive or fully orthogonal, serving as a starting point for organizing the field.Their ordering is illustrative and partially builds key points across the discussion.
  • Framework: The ten desiderata cover simple-command flexibility, multiple speech acts, mixed initiative, situated language and symbol grounding, affect, non-verbal communication, planning, learning, online services, and miscellaneous abilities.Together they define a broad agenda for conversational robots.
  • Starting point: Traditional conversational robots typically assigned humans the master role and robots the servant role, restricting interaction largely to simple motor-command requests.Their dialogue commonly mapped human requests directly to robot motor or verbal responses.
  • Starting point: Classic command systems were primarily single-initiative, offered limited robot speech, and handled only RequestForMotorAction speech acts.They were also inflexible to alternative command formulations and often used designer-chosen word-to-response mappings.
  • Open scope: The paper notes that systematic command-oriented research remains important while leaving many other aspects of natural language and robotics unaccounted for.The extent to which those aspects can be integrated remains open.

B. Multiple speech acts

Multiple-speech-act interaction extends robots beyond motor-action requests toward information exchange, assertions, declarations, and indirect meanings. This requires interpreting utterance function and using linguistic and paralinguistic evidence for classification.

  • Speech-act scope: Speech-act analysis focuses on an utterance’s function and purpose rather than only its content and form.The paper introduces speech acts as utterances with performative functions in communication.
  • Speech-act scope: Beyond motor-action requests, conversational robots may need to handle requests for information, informs, and declarations.Examples include asking an object’s size, reporting a red object’s location, and assigning a name to a doll.
  • Internal responses: Assertive and declarative utterances primarily require changes to the robot’s mental or situation model rather than external motor or verbal actions.An inform can create a mental token for an object, while a declaration can change its name label.
  • Indirect speech acts: Indirect speech acts require the robot to infer an intended directive from an utterance that is formally assertive.“It is quite hot in this room” can function as a polite way of requesting that someone open the window.
  • Classification: Speech-act classification can use paralinguistic information such as prosodic features in addition to linguistic information.The paper identifies classification after hearing an utterance as a distinct problem.

C. Mixed Initiative Dialogue

Mixed-initiative dialogue allows either participant to drive different parts of an interaction rather than assigning initiative permanently to the human or robot. The paper contrasts robot-initiative examples with a dialogue in which initiative changes hands.

  • Robot initiative: Natural interaction can support robot-initiative dialogue, in which the robot starts exchanges and expects simple human responses.FaceBots exemplifies this pattern with robot questions and updates requiring mainly yes-or-no replies.
  • Mixed initiative: BIRON illustrates limited mixed initiative by combining robot openings, human questions, and robot responses within the same conversation.The interaction includes repeated repair attempts before the robot identifies the object being shown.
  • Mixed initiative: In the BIRON dialogue, initiative reverses at multiple points as robot- and human-led exchange pairs follow one another.The opening is robot-initiative, while the subsequent identity question is human-initiative.
  • Related systems: The state of the art includes additional humanoid and conversational systems, as well as workshops focused on mixed initiative.The paper points readers to the Karlsruhe Humanoid and Bielefeld’s Biron and Barthoc systems.

D. Situated Language and Symbol Grounding

Situated language requires connecting linguistic processing to perception, action, and procedural semantics rather than treating language as an isolated speech-in/speech-out module. The paper presents empirical meaning models as an alternative to designer-defined semantics.

  • D. Situated Language and Symbol Grounding: Traditional command-only systems normatively define utterance meanings, whereas empirical approaches model vocabulary and action semantics from user observations.Examples include wizard-of-Oz collection of users’ vocabulary and context-dependent interpretations of commands such as “go left” and “go right”.
  • D. Situated Language and Symbol Grounding: A robot’s local environment and situational context are crucial to interpreting an utterance’s procedural semantics.The paper links language meaning to the concrete context in which the robot perceives and acts.
  • D. Situated Language and Symbol Grounding: Robotic language systems cannot be isolated from perception and action because utterances must connect to what the robot senses and does.The paper contrasts this integrated requirement with black-box conversational systems lacking links to physical sensing and behavior.
  • D. Situated Language and Symbol Grounding: Situated-language research starts with concrete here-and-now language, following a proposed progression from simpler to wider linguistic abilities.This approach is presented as partially mimicking human linguistic development and is exemplified by work in.

1) Situated Language:

Situated language concerns concrete physical events and objects in the here-and-now, with proposed levels extending toward past, imagined, and abstract entities. Current robots can handle the first four levels, but abstract concepts remain beyond the state of the art.

  • 1) Situated Language:: Situated language is concrete language about the physical here-and-now rather than abstract, past, or imagined events.The paper proposes it as an initial target for robots before broader linguistic abilities.
  • 1) Situated Language:: The proposed progression moves from immediately sensed objects to remembered, predicted, and finally abstract entities.The first level restricts reference to concrete things currently accessible to the senses; later levels relax these restrictions.
  • 1) Situated Language:: Abstract things are entities not directly connected to the senses, such as “threeness” instantiated across auditory and visual examples.The paper characterizes abstract entities as built upon first-order sensed entities and their relationships.
  • 1) Situated Language:: Robots and methodologies currently handle basic language across the first four detachment stages, while the fifth stage remains out of reach.Computational analogy-making is identified as a possible starting point for extending systems toward abstract domains.
  • 1) Situated Language:: Natural-language robotics must connect language to sensory data and internalized mental models, while addressing semantics and meaning beyond phonology and syntax.Mental models support language detached from the immediate here-and-now.

2) Symbol Grounding:

Symbol grounding connects linguistic symbols to external referents through sensory and cognitive processes, including bidirectional links for understanding and production. The paper also highlights unresolved challenges in meaning spaces, composition, and large-scale data acquisition.

  • 2) Symbol Grounding:: Full grounding links external referents, sensory streams, internal symbols, and produced or heard utterances in both directions.The first direction is called half grounding; full grounding includes the reverse path from heard utterance through symbol and sensory expectation to the referent.
  • 2) Symbol Grounding:: Grounding research addresses spatial relations and personal pronouns through empirical models, referring expressions, real-robot implementations, and learned semantics.Examples include “to the left of”, “inside”, “my left”, and pronoun meanings learned through examples.
  • 2) Symbol Grounding:: The literature includes proposed solutions, language-evolution accounts, grounded ontologies, and robot-usable lexica with computational grounding models.POETICON and POETICON II are identified as projects pursuing grounded ontologies and robot-usable lexica.
  • 2) Symbol Grounding:: Grounding must accommodate different target meaning spaces, including sensory, sensorimotor, and teleological spaces.The paper also identifies semantic composition, such as deriving “dark red” from “dark” and “red”, as insufficiently addressed.
  • 2) Symbol Grounding:: Large-scale grounding models require substantial empirical data, making online collection from non-experts desirable.The paper cites a modified computer game as an example of collecting referential and functional meaning models.
  • 2) Symbol Grounding:: Communication between partners with different meaning models is possible when their models overlap sufficiently.The paper frames shared meaning as the minimally informative basis for communication across differing lexical semantics.

3) Meaning Negotiation:

The paper situates affective and non-verbal communication within the broader challenge of fluid human-robot interaction. It surveys emotional speech, facial-expression systems, and related affect-enabled technologies.

  • 3) Meaning Negotiation:: Affective interaction matters because affect is intertwined with learning, persuasion, and empathy in human interaction.For speech, affect appears in both semantic or pragmatic content and prosody.
  • 3) Meaning Negotiation:: Early affective human-robot work includes Kismet’s interactive emotion and drive system with multiple facial expressions.The paper places this work alongside affective virtual avatars and emotional speech research.
  • 3) Meaning Negotiation:: Robotic affective systems have explored emotion recognition and generation through children’s speech interactions, storytelling, and facial-expression recognition.Reported tools include SHORE, FaceAPI, and systems used as aids for autistic children.
  • 3) Meaning Negotiation:: Affect-enabled text-to-speech and speech recognition support broader efforts to build real-world robotic systems with affective communication.The paper points readers to general reviews and notable developments in affective speech-enabled robots.

F. Motor corellates of speech and non-verbal communication

Natural human-robot communication requires speech to be coordinated with non-verbal behaviors such as lip movements, gestures, and gaze. These cues support speech naturalness, recognition, reference resolution, collaboration, and potentially communication training.

  • Lip synchronization: Lip movements should accompany speech production because verbal communication does not occur independently of non-verbal signs.Soundtrack-only lip synchronization can produce unnatural movements, while lip information can also improve speech recognition under low signal-to-noise conditions.
  • Gestures: Deictic gestures connect spoken indexicals such as “this one!” to objects in the environment.Such gestures have been used in human-robot interaction from virtual avatars to navigating robots.
  • Gaze and shared attention: Eye gaze coordinates collaborative tasks, improves efficiency and robustness in teamwork, and helps disambiguate referring expressions.Gaze imitation and shared-attention models have been studied for these functions.
  • Theory of mind: Eye-gaze observation can help robots model other agents’ mental content and functions, supporting purposeful speech generation and referring expressions.The paper links these abilities to progressively developing forms of theory of mind in children.
  • Applications: Specially designed robots could potentially help children with Autism Spectrum Disorders improve communication skills.The paper presents this as a possibility associated with research on theory-of-mind deficiencies and robot interaction.

G. Purposeful speech and planning

Fluid communication requires robots to plan purposeful speech and motor behavior together rather than map stimuli directly to responses. The paper also identifies learning, internet-based resources, and shared robotic experience as routes toward more capable systems.

  • Purposeful behavior: Existing dialogue systems lack explicit modeling of purposeful behavior toward goals, even when they support situated language and multiple speech acts.They can be viewed as stimulus-response or stimulus-state-to-response maps.
  • Integrated planning: Motor planning and speech planning are rarely integrated, despite real-world actions often being interchangeable and unable to remain isolated.The paper calls for mixed speech- and motor-act planning and planners directed toward agent and object interactions.
  • Learning: Learning for human-robot communication must address both when learning occurs and what linguistic and non-verbal knowledge should be learned.Relevant layers include phonological, morphological, syntactic, semantic, pragmatic, and dialogic knowledge, alongside grounded meaning and adjustable non-verbal models.
  • Learning: Real-world communicative learning may require many robots operating for many hours daily, making direct data collection unrealistic for current systems.The paper therefore points to specially designed crowdsourced online games as an alternative avenue for acquiring data.
  • Online resources: Internet connectivity can offload part of a robot’s intelligence to online information, services, and distributed human or machine sensing, actuation, and processing.Examples include Facebots’ shared online memories and RoboEarth’s shared repository for robot experience.

J. Miscellaneous abilities

Fluid human-robot communication also requires handling multiple partners, languages, and modalities. These abilities support group interaction, broader participation, translation, mediation, and integration into human social environments.

  • Multiple conversational partners: Multiple-partner conversation requires turn-taking, overlapping-speech recognition, speaker identification, and sound-source localization.Gaze and head movements also contribute to turn-taking and participant-role assignment.
  • Multilingual capabilities: Multilingual robots could communicate with more people and act as translators or mediators in multicultural settings.The paper notes progress in non-embodied multilingual dialogue and virtual avatars, but limited implemented real-world multilingual physical robots.
  • Multimodal natural language: Robots can generate and recognize natural language across acoustic, written, and sign-language modalities.Relevant technologies include sign-language systems, optical character recognition, and handwriting recognition.
  • Multimodal natural language: Written and electronic communication abilities could help robots integrate more fluidly into human societies and use services offered through human networks.The paper also envisions robots helping maintain and improve the social networks in which they reside.

IV. DISCUSSION

The review finds substantial progress beyond early canned-command systems, but fluid verbal and non-verbal human-robot communication remains far from achieved. It highlights open research directions and argues that massive real-world data may drive further progress.

  • Discussion: Although research has moved beyond canned commands and responses, fluid and natural verbal and non-verbal human-robot communication remains a distant goal.The discussion frames this gap as the central outcome of examining the ten desiderata.
  • Discussion: Open directions include grounded semantics, online meaning negotiation, affective interaction, mixed speech-motor planning, data acquisition, and online resource use.These areas are presented as promising avenues for future projects.
  • Discussion: Massive real-world data could drive further data-driven models, but obtaining it requires robots deployed in interactive settings with extended operation.The paper mentions shopping malls, reception areas, museums, and companion applications as possible settings.
  • Conclusions: The review organizes verbal and non-verbal research around ten desiderata after examining historical development and motivations for fluid communication.It concludes that significant progress has occurred across many fronts while many sub-problems remain unsolved.
Loading 1401.4994v1…