Source-linked AI summary
Human-robot conversation with multiple participants in noisy public spaces
Divesh Lala, Yogeeswaran Muthukumaran, Vincent Fernandes, Kazushi Kato, Shota Fujiki, Zihao Chi, Masaya Iwasaki, Taiken Shintani, Megumi Kawata, Kazuki Sakai, Koji Inoue, Yuicihiro Yoshikawa, Tatsuya Kawahara
TL;DR
Noisy public spaces make speech recognition and remote operation difficult for multi-party robot and avatar conversations. The paper presents a shared microphone-array audio system, demonstrates it with ERICA and mobile Teleco robots at the 2025 World Expo, and reports robust proof-of-concept operation alongside clear limitations in true autonomous multi-party turn-taking.
Problem
Multi-party robot and avatar conversations in noisy public spaces require speech signals that support recognition and remote operators, but microphone arrays alone struggle with noise and source separation.
Method
The paper uses one microphone-array audio architecture with speech enhancement, channel selection, and local incremental recognition across attentive-listening and mobile Teleco scenarios.
Results
The systems were demonstrated as proof-of-concept at the 2025 World Expo and were reported to recognize conversational speech, maintain conversation flow, select appropriate channels, and provide clear audio to operators.
Takeaways & Limitations
Enhanced audio channels can support noisy multi-party robot dialogue and provide remote operators with clearer conversation audio than raw audio.
Takeaways & Limitations
The dialogue system is not truly multi-party because autonomous turn-taking and addressee detection remain unresolved, and some users’ speech was missed when they spoke quietly or turned away.
Abstract
from arXiv · showhide
For noisy real-world environments such as those in open public spaces, spoken dialogue systems for both autonomous robots and avatars should be carefully designed to provide enhanced speech signals. These signals can be used either for speech recognition or, in the case of an avatar system, transmitted as clean speech to a remote operator. This work proposes an audio system that can be used for both these scenarios and was demonstrated as a proof-of-concept at the 2025 World Expo in Osaka. The first scenario is an attentive listening system with the android ERICA, and the second is a conversation support system with mobile Teleco robots, with one of them acting as an avatar for a remote operator. Both systems feature multi-party conversation and use a single multi-channel microphone array. We describe how our audio system not only enhances the speech of multiple speakers in a noisy environment, but provides a form of spatial audio which allows for more immersiveness in avatar-based conversational interactions.
I. INTRODUCTION
Spoken dialogue systems remain challenged by noisy public environments and multi-party interaction, where diverse users, overlapping speech, and complex turn-taking complicate recognition. The paper presents an audio system demonstrated in two such human-robot dialogue scenarios at the 2025 World Expo in Osaka.
- Motivation: Public environments expose conversational robots to background noise and diverse user behaviors beyond controlled laboratory conditions.Children and elderly users may interact differently, including speaking more simply, speaking slowly, or repeating themselves.
- Motivation: Multi-party conversation involving three or more participants introduces overlapping utterances, interruptions, and more complex turn-taking than dyadic dialogue.Situated multi-party conversation is common in everyday life and should be accommodated by social robots.
- Motivation: Cybernetic avatars must support immersive multi-party interaction in noisy environments by helping remote operators focus on conversational speech.This requires replicating the human ability to filter background noise and attend to particular speech stimuli.
- Audio challenge: A single multi-channel microphone array can extract multiple speakers, but noise from all directions and source separation remain difficult for recognition and remote operation.The paper identifies this as an incompletely addressed technical challenge in spoken dialogue research.
- Contribution: The paper presents an audio system for human-robot multi-party dialogue in high-noise environments.It addresses the stated need for enhanced speech signals in public conversational settings.
- Demonstrations: The system was demonstrated in Osaka at the 2025 World Expo with three participants in each scenario: ERICA with two humans, and two mobile Teleco robots with one remote operator.The Teleco setup included one tele-operated robot acting as a cybernetic avatar and one fully autonomous robot; both systems were implemented in Japanese.
II. RELATED WORK
Prior work has addressed multi-party interaction, microphone-array speech extraction, and public robotics separately, but mobile multi-party human-robot conversation in public remained insufficiently addressed. This work combines these areas with a shared low-cost microphone-array audio architecture for speech enhancement and recognition.
- Existing research: Multi-party robot conversation is an increasing research focus because dyadic conversation is comparatively well handled by LLMs.Hands-free systems commonly use microphone arrays to extract the current speaker’s speech.
- Existing research: Public-interaction research has largely involved stationary robots or dyadic conversations, while mobile robots handling multi-party interaction remain relatively rare.The paper aims to perform audio processing while a robot moves in a public space.
- Research gap: Prior systems used visual speaker identification without simultaneous speech, or microphone-array separation in controlled environments; this work combines public, mobile, and multi-party settings.The authors state that such a system had not previously been addressed to their knowledge.
- Audio architecture: The audio architecture uses one low-cost four-channel ReSpeaker circular microphone array whose channels exhibit substantial audio bleeding.The array can operate as a standalone device or be integrated into a robot, but simultaneous speech recognition is difficult without further processing.
- System goal: The system seeks to extract clean speech from multiple people in noise for speech recognition and remote-operator audio.These two uses support autonomous dialogue and cybernetic-avatar interaction, respectively.
- Speech enhancement: A learned system separates the array input into multiple enhanced channels with configurable listening directions and voice-activity enhancement under substantial background noise.Enhanced audio is produced from multiple channels simultaneously.
- Speech recognition: A local low-latency Japanese incremental recognizer, trained on noisy speech and lightweight enough for concurrent execution, provides noise-robust recognition from multiple people.The recognition models run locally on one computer.
IV. DIALOGUE SCENARIOS
The paper evaluates the audio system in two multi-party dialogue scenarios with different physical arrangements and interaction requirements. ERICA performs attentive listening with two humans, while the second scenario uses mobile Teleco robots and requires continuous speech processing for interaction.
- Scenario overview: The two dialogue scenarios were designed with different requirements and are compared as distinct system configurations.The paper uses both scenarios to test the audio system for multi-party conversation.
- Attentive Listening with ERICA: ERICA conducts multi-party attentive listening with two human participants, using a microphone array positioned on a tripod between them.The android produces realistic facial expressions, speech, and upper-body gestures.
- Attentive Listening with ERICA: During attentive listening, ERICA asks participants to speak individually, provides verbal backchannels, nods, and short responses, then supports discussion and summarizes it.Short responses include elaborating or empathetic questions generated by GPT 4.1.
- Attentive Listening with ERICA: The scenario supports exchanges in which participants speak to one another or ERICA, requiring continuous simultaneous recognition and separated speaker histories for LLM input.A voice activity projection model predicts when ERICA should take the turn from the current speaker.
- Attentive Listening with ERICA: The seated triangular setup places participants approximately 30° to either side of ERICA, with enhanced channels directed toward their positions.The fixed arrangement also supports gaze behavior without visual information.
B. Conversation Support with Teleco
The Teleco demonstration used one robot as a remotely operated avatar and another as an autonomous conversational support robot in a three-party interaction. An operator GUI supported avatar control and commands that prompted T-AUTO to contribute to the conversation.
- System roles: T-OP is a mobile robot avatar controlled remotely through a computer, while T-AUTO operates autonomously as conversational support.T-OP streams camera video and audio to the operator, who controls its viewpoint and speech; T-AUTO receives support commands from the operator.
- Operator interface: The operator GUI displayed T-OP’s camera feed and provided controls for viewpoint adjustment, speech, and commands to T-AUTO.The operator could raise or lower T-OP’s pole, rotate it, and use buttons to request autonomous conversational support.
- Demonstration setup: The Teleco robots were visually distinguished by their roles: the autonomous robot wore a blue vest and the tele-operated robot wore an orange vest.This role distinction was shown in the public demonstration figure.
- Interaction scenario: The demonstration placed one human participant, T-OP, and T-AUTO in a triangular F-formation for a casual group conversation.Both robots initially approached the participant from approximately four meters away.
- Conversational support: T-AUTO could be asked to request more details, suggest something different, or support the operator during the conversation.After a command, GPT 4.1 generated an utterance from the dialogue history and request type, which T-AUTO delivered at an appropriate time.
V. REQUIREMENTS FOR MOBILE TELECO AVATAR SYSTEM
The mobile Teleco avatar requires audio processing to remain effective while its microphone array moves with the robot. The system therefore handles operator-voice echo, tracks participants, and updates the listening direction dynamically.
- Audio requirements: Because the microphone array sits on T-OP’s head, the operator’s nearby loudspeaker voice can enter all channels and disrupt speech recognition.Specialized software was used for microphone-array echo cancellation, which removed the operator’s voice from the recognition input.
- Participant tracking: The system tracks Teleco positions with SLAM and human position with a depth camera and YOLO vision model.Teleco orientation combines rover, waist-joint, and neck-joint yaw values, while the camera observes the environment from approximately two meters high.
A. Audio channel selection
The audio-channel selector chooses microphone sectors containing nearby humans and updates the listening direction as T-OP moves. Enhanced signals support recognition and remote listening, while spatial panning preserves conversational directionality.
- Mobile extension: The mobile extension is the ability to select audio channels for a moving robot in a dynamic, unpredictable environment.The implementation depends on tracking the positions of all participants.
- Channel representation: The selector represents four channels with principal listening directions that rotate when T-OP changes orientation.With an initial facing direction of 0°, the channel directions are 0°, 45°, 180°, and 315°.
- Channel selection: When a human enters a channel sector, the corresponding channel is used for speech recognition; after T-OP turns 50° right, the selected channel changes from A1 to A4.The figure illustrates how channel selection follows the robot’s facing direction and participant location.
- Sector selection: For each channel, the system creates a real-time adjustable sector with a 60° angle and 2-meter radius, then selects sectors containing the closest human.If no nearby user is present, the system performs no listening; overlapping sectors retain the current channel or choose the first candidate.
- Speech enhancement: The system extracts user speech for recognition and transmits enhanced signals to the remote operator, including when the user moves while speaking.Multiple speakers can be handled when they occupy different listening channels.
- Spatial audio: Relative user position is transmitted with the audio so the operator hears each speaker panned from the appropriate direction.When T-AUTO speaks, its voice is presented from its relative direction to T-OP, preserving the spatial relationship between participants.
VI. SUPPORTING THE SENSE OF AGENCY
The Teleco system supports operator agency through direct control of the avatar and indirect influence over an autonomous support robot. Spatially enhanced audio is intended to make noisy multi-party interaction more immersive and less cognitively demanding.
- System objective: The system is intended to enable multi-user speech recognition in highly noisy environments while supporting the operator’s sense of agency.These aims combine robust audio processing with control of a robotic avatar and interaction with an autonomous support robot.
- Avatar agency: The operator can control T-OP’s speech and viewpoint while spatial audio supports immersion in the avatar interaction.The stated goal is to reduce the cognitive load of listening to conversation in a noisy environment.
- Shared agency: T-AUTO gives the operator influence over another participant without requiring full control of its behavior.The paper frames this arrangement as shared agency because fully controlling multiple conversational participants would be difficult.
VII. RESULTS
The Expo demonstration could not support formal subjective evaluation because of privacy constraints. Instead, speech-recognition functionality was pilot-tested in a controlled laboratory environment.
- Privacy issues prevented formal subjective evaluation at the Expo itself.
- Speech-recognition functionality was evaluated through pilot testing in a controlled laboratory environment.
- The testing approach separated the live Expo demonstration from the laboratory speech-recognition pilot.
A. Pilot testing
Pilot testing measured character error rate under microphone, robot-motion, and simulated-noise conditions. The speech-recognition system performed reasonably well and was comparable to a hand-held microphone, with similar performance observed for other users.
- Pilot testing: Character error rate was measured across hand-held microphone, stationary Teleco, rotating Teleco, and simulated-noise conditions.Teleco users spoke from approximately one meter, corresponding to casual conversation distance.
- Pilot testing: The system’s character error rate was reasonably well-performing and comparable to a hand-held microphone.
- Pilot testing: Similar speech-recognition performance was observed for users beyond the single user included in the pilot tests.The pilot tests themselves used one user.
B. Live expo demonstration
The systems operated simultaneously for six days in a busy, reverberant Expo pavilion serving thousands of visitors. Enhanced audio supported speech recognition, channel selection, conversation flow, and remote operators’ hearing.
- Live Expo demonstration: Across six days, approximately 7300 people entered the pavilion, and ERICA recorded 513 interactions.Interactions with the Teleco robots were not logged.
- Live Expo demonstration: The pavilion presented high reverberation and noise from visitors, other robot demonstrations, background speech, and occasional heavy rainfall.
- Live Expo demonstration: Conversational speech was successfully recognized in the majority of cases, with coherent responses maintaining conversation flow in both scenarios.
- Live Expo demonstration: Teleco selected audio channels according to participant position and orientation, while operators reported clearly hearing conversations despite high noise.
C. Limitations
The implementation did not yet provide fully autonomous multi-party dialogue, and microphone pickup failures occurred under specific speaking behaviors and voice levels. The conclusion reports that one shared array supported both demonstrated scenarios despite their differing configurations.
- Limitations: The dialogue system was not truly multi-party because interaction remained sequential or depended on operator-commanded interjections.Smooth autonomous turn-taking and addressee detection remain necessary for multi-party conversation.
- Limitations: Speech was sometimes missed when users spoke too quietly or turned toward T-AUTO.These behaviors prevented the microphone array from processing the speech correctly and could cause conversation stagnation.
- Scope of implementation: The same microphone-array audio processing accommodated differing human-robot distributions, mobility, and microphone locations across both scenarios.The demonstrations were proof-of-concept systems at the 2025 World Expo in Osaka.