Source-linked AI summary

Symbol Emergence in Robotics: A Survey

Tadahiro Taniguchi, Takayuki Nagai, Tomoaki Nakamura, Naoto Iwahashi, Tetsuya Ogata, Hideki Asoh

arXiv:1509.08973v1cs.AIcs.CLcs.CVcs.RO

TL;DR

The paper addresses how humans and robots can form adaptive symbol systems for embodied, socially situated communication. It introduces SER as a constructive framework in which symbol systems self-organize through physical and semiotic interaction, and surveys unsupervised methods for multimodal categorization, word discovery, and double articulation analysis. SER’s supported direction emphasizes embodied evaluation and identifies substantial time and cost for learning communication through situated interaction.

  • Problem

    Existing symbol-system approaches do not adequately explain the dynamic and emergent properties of human symbol systems, although understanding them matters for long-term human–robot communication.

  • Method

    The paper introduces SER, a constructive framework using embodied computational models and unsupervised learning from multimodal sensory–motor information.

  • Results

    SER surveys research showing how robots can acquire embodied meanings and words through multimodal categorization, word discovery, and double articulation analysis without supervision.

  • Takeaways & Limitations

    SER frames robotic symbol systems as socially self-organized through physical and semiotic interaction and calls for evaluation in embodied, contextual, collaborative tasks.

  • Takeaways & Limitations

    Learning the knowledge required for communication and collaboration through situated embodied interaction requires very large costs and time.

Abstract

from arXiv · show

Humans can learn the use of language through physical interaction with their environment and semiotic communication with other people. It is very important to obtain a computational understanding of how humans can form a symbol system and obtain semiotic skills through their autonomous mental development. Recently, many studies have been conducted on the construction of robotic systems and machine-learning methods that can learn the use of language through embodied multimodal interaction with their environment and other systems. Understanding human social interactions and developing a robot that can smoothly communicate with human users in the long term, requires an understanding of the dynamics of symbol systems and is crucially important. The embodied cognition and social interaction of participants gradually change a symbol system in a constructive manner. In this paper, we introduce a field of research called symbol emergence in robotics (SER). SER is a constructive approach towards an emergent symbol system. The emergent symbol system is socially self-organized through both semiotic communications and physical interactions with autonomous cognitive developmental agents, i.e., humans and developmental robots. Specifically, we describe some state-of-art research topics concerning SER, e.g., multimodal categorization, word discovery, and a double articulation analysis, that enable a robot to obtain words and their embodied meanings from raw sensory--motor information, including visual information, haptic information, auditory information, and acoustic speech signals, in a totally unsupervised manner. Finally, we suggest future directions of research in SER.

1. Introduction

The paper introduces symbol emergence in robotics (SER) as a constructive approach to adaptive, socially self-organized symbol systems for long-term human–robot interaction. It surveys bottom-up research on multimodal categorization, language acquisition, and related processes grounded in sensory–motor information.

  • SER addresses the challenge of developing robots that communicate and collaborate naturally with people over long-term interaction.
  • SER models symbol systems as socially self-organized through physical and semiotic interactions between people and developmental robots.
  • Figure 1 organizes SER around robots communicating semantically and interacting collaboratively with people while becoming part of an emergent symbol system.
  • SER requires bottom-up learning from sensory–motor information rather than top-down intelligence design, with multimodal concept formation and autonomous speech-based language acquisition as fundamental topics.
  • The paper surveys multimodal categorization, word discovery, and double articulation analysis as core components of SER research.

2. Background

The background contrasts traditional discrete, static, and top-down symbol systems with embodied approaches that ground meaning in sensory–motor interaction. SER extends these approaches by computationally studying emergent symbol systems as dynamic, social, and interdisciplinary phenomena.

  • Traditional symbolic approaches assume discrete, deterministic, and static representations, whereas human symbol systems are treated as grounded, dynamic, and socially shaped.
  • Physical symbol system approaches and subsumption-based robotics have been criticized for struggling to explain how autonomous systems acquire human-like language through interaction.
  • The symbol grounding problem remains unresolved because its physical-symbol-system framing emphasizes grounding while largely ignoring the dynamic and social characteristics of symbols.
  • SER distinguishes its emergent-symbol-system approach by considering multiple symbol types together rather than grounding only predesigned symbols.
  • SER integrates computational models, embodied cognition, semiotics, linguistics, robotics, artificial intelligence, developmental psychology, and cognitive science.

3. Emergent Symbol Systems

Emergent symbol systems are dynamic, self-organized through physical and semiotic interactions among humans and robots. SER therefore treats robots as adaptive participants that acquire language and collaboration abilities bottom-up.

  • 3.1 Semiosis and umwelt: In Peircean semiotics, a symbol is a dynamic interpretive process linking sign, object, and interpretant rather than a static sign.
  • 3.3 Emergent symbol systems: Internal representations self-organize through sensory–motor interaction, while semiotic communication organizes human symbol systems from the bottom up under higher-level constraints.
  • 3.3 Emergent symbol systems: A symbol system forms a micro-macro loop in which societal rules constrain communicating agents and agent interactions sustain the emergent system.
  • 3.3 Emergent symbol systems: SER defines an emergent symbol system as socially self-organized through physical and semiotic interactions between people and robots.
  • 3.3 Emergent symbol systems: SER aims to develop autonomous robots that acquire language and learn to communicate and collaborate with humans through unsupervised, bottom-up development.

4. Multimodal Categorization

Multimodal categorization grounds robot concepts in unsupervised sensory–motor interaction rather than visual recognition alone. Bayesian models integrate modalities, estimate latent category structure, and produce object, hierarchical, attribute, and word-linked concepts.

  • 4.1 Multimodal categorization: Multimodal categorization integrates sensory modalities because human symbol systems arise from multimodal sensory–motor experience and cross-modal prediction.
  • 4.2 Computational models for multimodal object categorization: Robots formed human-like object categories from visual, auditory, and haptic information in a fully autonomous home-environment experiment.
  • 4.3 Estimating latent structure in multimodal categories: MHDP extends multimodal Bayesian clustering to estimate the number of categories adaptively rather than fixing it in advance.
  • 4.2 Computational models for multimodal object categorization: MLDA can jointly cluster multimodal information and words, allowing a robot to estimate category labels without supervision.
  • 4.3 Estimating latent structure in multimodal categories: Hierarchical MLDA enables nested concepts such as “plastic bottle” as a subcategory of “water container.”
  • 4.3 Estimating latent structure in multimodal categories: BoMLDA, BoMHDP, and IMoMHDPs organize concepts around selected modalities and support categories for attributes such as softness, hardness, and color.

5. Word Discovery

Word discovery is a fundamental challenge in language acquisition because robots must segment continuous speech and learn word inventories without supervision. SER studies extend this problem to raw multimodal interaction, but robust real-world discovery remains limited.

  • 5. Word Discovery: Unsupervised word discovery requires children and robots to segment continuous speech while learning word inventories and phoneme representations.Unlike automatic speech recognition, children must acquire both linguistic and acoustic models from speech signals without supervision.
  • 5. Word Discovery: Infants can segment fluent speech using distributional cues alone, although distributional, co-occurrence, and prosodic cues jointly contribute to word discovery.Saffran’s results reported this ability at eight months, with distributional cues also reported by seven months.
  • 5. Word Discovery: Robotic systems have acquired lexicons from raw multimodal sensor data and learned speech units, lexicons, grammar, and interpretation through embodied communication.These approaches avoid human transcription or labeling and can integrate speech, visual, and behavioral information online and incrementally.
  • 5. Word Discovery: Word-inventory management is a model-selection problem because robots face finite memory despite potentially infinite words, motivating Bayesian nonparametric language models.Hierarchical Pitman–Yor process models assign probabilities to infinitely many possible words within a Bayesian framework.
  • 5. Word Discovery: Robust word discovery remains challenging because current methods work in limited situations and must improve to match children’s real-world learning.Methods increasingly learn from acoustic data without preexisting models, but phoneme errors and practical constraints remain obstacles.

6. Double Articulation Analysis

Double articulation analysis models time series as two hierarchical layers: meaningless low-level segments combine into meaningful words or chunks. SER applies this structure to speech, motion, driving behavior, and sensory–motor data, while integrated models address segmentation errors.

  • 6. Double Articulation Analysis: Double articulation enables infinite meaningful combinations from a small set of meaningless low-level units, such as phonemes forming words.Speech can be decomposed into words and then into phonemes or letters, separating sign-bearing units from lower-level elements.
  • 6.2 Segmentation of human bodily motion: Motion segmentation distinguishes short-term physically elemental segments from longer semantically meaningful chunks, paralleling phonemes and words in speech.A semantic action such as pitching can comprise several dynamically elemental motions that lack individual names.
  • 6.2 Segmentation of human bodily motion: The double articulation analyzer combines sticky HDP-HMM and NPYLM sequentially to infer latent segments and words without supervision.The method explicitly represents two hierarchical levels corresponding to letters or phonemes and words or segments.
  • 6.3 Modeling driving behavior data: Driving behavior exhibits double articulation: meaningful activities consist of sequences of physically elemental actions, and DAA-based prediction has outperformed conventional methods.The analysis labels semantically elemental states driving words and physically elemental states driving letters, with effectiveness verified on large-scale data.
  • 6.5 Direct word discovery from speech signals: A nonparametric Bayesian double articulation analyzer integrates the two layers generatively and can discover complete word lists directly from vowel speech without supervision.It was introduced to address degradation caused when sequential segmentation errors propagate into subsequent unsupervised chunking.
  • 6. Double Articulation Analysis: The same computational structure can analyze speech, human behavior, and sensory–motor time series, suggesting shared processes relevant to emergent symbol systems.In these cases, first-layer elements are meaningless while second-layer words or chunks are meaningful; hierarchical recurrent models also capture contextual articulation.

7. Further Topics

Further SER topics address context-sensitive communication, autonomous exploration, active perception, and semantic composition. These topics extend symbol emergence from acquiring units to interpreting, generating, and combining meanings in embodied interaction.

  • 7. Further Topics: Mutual beliefs constrain both utterance interpretation and speech generation by linking language to shared physical and social context.A statement such as “This coffee is cold” can function as a request for hotter coffee depending on the shared situation.
  • 7. Further Topics: Embodied robots need autonomous exploration, knowledge acquisition, and communication, making active perception and active learning central to lifelong development.SER connects these capabilities with information-seeking behavior and decision-making during interaction.
  • 7. Further Topics: Information-gain criteria provide a computational treatment of curiosity for designing intrinsic motivation and explanatory behavior.The cited work identifies the effectiveness of the information-gain criterion in autonomous robot development.
  • 7.3 Compositionality and semantics: Double-articulation hierarchies can generate meaningful sentences, but SER must explain how compositions of words acquire adequate meanings for embodied situated agents.This motivates bottom-up computational models of semantic composition rather than only word segmentation.
  • 7.3 Compositionality and semantics: Robotic systems and neural architectures have demonstrated distributed emergence of compositionality while jointly learning sentences and behaviors.Related work studies compositionality in language–behavior association learning and distributional semantic representations.

8. Conclusion

The conclusion frames SER as a constructive approach for robots to adapt to human symbol systems through embodied and semiotic interaction. It identifies real-world evaluation and the cost of situated learning as central challenges while outlining SER’s promise for long-term human–robot interaction.

  • Understanding dynamically changing symbol systems is important for robots intended to communicate and collaborate smoothly with human users.
  • SER is a constructive approach in which emergent symbol systems are socially self-organized through semiotic and physical interactions between humans and developmental robots.
  • Real-world collaborative tasks and embodied cognition should guide evaluation because document-only studies do not capture situated human–robot interaction.
  • Situated and embodied learning of the knowledge needed for communication and collaboration requires very large costs and time, motivating simulators and cloud-based interactions.
  • SER remains an emerging but promising field expected to advance long-term human–robot interaction and understanding of human intelligence.
Loading 1509.08973v1…