Source-linked AI summary

TEIDAN: A Multilingual Multiparty Dialogue Corpus

Taiga Mori, Koji Inoue, Mikey Elmers, Divesh Lala, Tatsuya Kawahara

arXiv:2609.00802v1cs.CLcs.HC

TL;DR

Spoken multiparty systems need resources that capture how people coordinate turns and recipients in groups, but comparable spontaneous face-to-face triadic corpora remain limited. TEIDAN addresses this gap with a Japanese-English multimodal corpus using shared recording and transcription conventions, and preliminary analyses show cross-language differences in turn organization while highlighting important scope limits.

  • Problem

    Existing corpora provide limited evidence for comparing spontaneous face-to-face triadic interaction across languages, despite the importance of group turn-taking for dialogue systems and embodied agents.

  • Method

    TEIDAN constructs a multilingual multimodal corpus of Japanese and English three-party discussions with shared IPU-based transcripts and multimodal recordings.

  • Results

    Preliminary analyses find substantial pooled-corpus differences between English and Japanese in IPUs per pseudo-turn and mean pseudo-turn duration.

  • Takeaways & Limitations

    TEIDAN supports cross-linguistic and multimodal study by linking corpus-level statistics with close sequential analysis of interaction.

  • Takeaways & Limitations

    Language and culture cannot be isolated as independent causal factors because the portions also differ in participant relationships, including familiarity.

Abstract

from arXiv · show

Multi-party interaction is a central setting for human communication and a necessary target for human-agent interaction systems that must participate in group conversation. Yet available corpora often focus on meetings, task-oriented interaction, text-based interaction, or acted scenarios, and fewer resources support cross-linguistic comparison of spontaneous face-to-face triadic discussion. This paper presents TEIDAN, a multilingual multimodal corpus that currently consists of Japanese and English three-party conversations. TEIDAN records groups of three participants discussing open-ended topics with individual pin microphones, a microphone array, and participant-facing cameras, and provides IPU-based transcripts for both language portions. Earlier studies used subsets of the Japanese portion for task-specific benchmarks in multi-party dialogue modeling; in contrast, this paper presents TEIDAN as a corpus resource spanning both Japanese and English, with planned expansion to additional languages. We describe the collection design, participants, recording setup, transcription format, and corpus statistics, and provide preliminary analyses to illustrate how TEIDAN can support research on turn-taking, addressee recognition, and multimodal grounding in human-human and human-agent interaction.

1 Introduction

TEIDAN addresses the limited availability of natural, multimodal, multilingual corpora for spontaneous face-to-face three-party interaction. It presents Japanese and English data as a comparable corpus resource for studying turn-taking and related phenomena.

  • Existing multiparty resources often target text, meetings, situated tasks, teams, games, or human-robot interaction rather than natural multimodal face-to-face conversation.
  • Prior Japanese TEIDAN subsets supported addressee recognition, next-speaker prediction, and voice activity projection benchmarks but lacked a standalone multilingual corpus account.
  • TEIDAN currently contains 69 sessions, 57 unique participants, and 9 hours 46 minutes 57 seconds of multimodal interaction across Japanese and English.
  • The corpus uses a shared design to support comparison of interactional properties across languages, cultures, and participant relationships, with planned expansion to additional languages.
  • Its contributions include corpus documentation, an English extension, recording and transcription descriptions, statistics, and preliminary analyses of turn-taking, addressee recognition, next-speaker prediction, and multimodal grounding.

2 Method

TEIDAN records spontaneous triadic opinion discussions using comparable multilingual protocols, multimodal capture, and shared IPU-based transcription conventions. Japanese and English portions share core design features while differing in topics, participants, and language-specific transcript details.

  • 2.1 Corpus Design: Each session records three participants discussing broad open-ended prompts without a strict task objective or single correct answer.
  • 2.1 Corpus Design: The English corpus shares three discussion themes with the Japanese corpus: island, travel, and life.
  • 2.2 Participants: Japanese participants were acquainted students and faculty from the same laboratory, with some individuals appearing in multiple groups.
  • 2.2 Participants: English participants were native or near-native speakers recruited from 13 nationalities and met one another for the first time at recording.
  • 2.3 Recording Setup: Both portions use individual pin microphones, a microphone array, and participant-facing cameras to capture separate speech, shared acoustics, and faces.
  • 2.4 Transcription: An IPU is a speech segment bounded by pauses of at least 200 ms, with timing and speaker labels recorded for each unit.

3 Statistics

The statistics section reports the size and duration of the Japanese and English portions, which together form the current TEIDAN corpus.

  • 36 Japanese sessions comprise 3 hours 39 minutes 10 seconds of data from 12 groups and 24 unique participants.
  • The English portion contains 11 groups, 33 unique participants, and 33 sessions.

4 Preliminary Analysis

Preliminary analyses reveal distinct turn-organization patterns across TEIDAN’s Japanese and English portions, alongside cooperative multimodal practices in Japanese conversation. The comparisons are descriptive and illustrate the corpus’s value for studying multilingual multiparty interaction.

  • Quantitative analysis: 10,097 vs. 7,348 substantive IPUs, while pseudo-turn counts remain similar at 4,851 vs. 4,748 for English and Japanese.English pseudo-turns average 2.08 IPUs versus 1.55 in Japanese, and average 4.65 versus 2.95 seconds.
  • Quantitative analysis: The pooled language comparisons are corpus-level descriptive differences rather than inferential tests of population-level differences.IPUs and pseudo-turns within the same group are interactionally related, so the statistics should not be treated as independent-sample evidence.
  • Qualitative analysis: English conversations often feature long opinion turns, with speakers holding the floor for 187.956 and 172.594 seconds in the reported examples.Other participants rarely take the floor during these opinion statements, although brief backchannels may occur.
  • Qualitative analysis: Across the Japanese examples, gaze shifts, beat gestures, agreement displays, and syntactic continuation jointly shape addressee-oriented participation.A shifts gaze toward B and produces beat gestures when soliciting B’s agreement, after earlier interaction with C.
  • Qualitative analysis: Japanese speakers often develop opinions incrementally while seeking agreement or empathy, using polite endings and syntax to project continuation.These practices create space for others to respond while allowing the original speaker to resume the turn.
  • Qualitative analysis: Japanese listeners may take brief turns at non-TRP positions as cooperative actions that support the current speaker’s ongoing activity.Compact examples or confirmation-like utterances can support the speaker’s claim and help a third participant understand it; B’s nodding provides multimodal evidence of understanding.

5 Discussion

TEIDAN’s preliminary analyses reveal substantial Japanese–English differences in turn organization while showing how multimodal, sequential evidence can support cross-linguistic investigation. These findings are descriptive rather than causal because the language portions also differ in participant relationships.

  • The English and Japanese portions differ clearly in IPUs per pseudo-turn and mean pseudo-turn duration.The pooled corpus comparison shows a substantial difference, but does not quantify between-group variation or support population-level inference.
  • English examples show extended floor-holding during opinion statements, whereas Japanese examples show compact listener contributions embedded in ongoing talk.
  • Japanese excerpts depict speakers building opinions incrementally while listeners offer compact responses, candidate understandings, or supportive examples before turn completion.The shorter pseudo-turns do not simply indicate that Japanese participants speak less.
  • The observations are compatible with accounts emphasizing high-context communication and interdependent attention to social relations in Japanese interaction.
  • Language and culture cannot be isolated as independent causal factors because Japanese participants were acquainted laboratory members whereas English participants were strangers.Participant relationships may also affect listener entry into ongoing turns and speaker floor-holding duration.

6 Conclusion

The paper presents TEIDAN as a corpus-level multilingual resource for spontaneous Japanese and English three-party dialogue. Its preliminary analyses demonstrate the corpus’s potential for studying turn-taking and multimodal listener behavior, while future work will broaden annotations and language coverage.

  • TEIDAN is a multilingual multimodal corpus of spontaneous three-party dialogue in Japanese and English.The paper documents collection protocols, participant characteristics, recording setup, transcription format, and statistics across both portions.
  • The corpus currently contains 69 sessions of open-ended face-to-face discussion with individual audio, array audio, video, and IPU-based transcripts.
  • Preliminary analyses link corpus-level statistics with close sequential analysis of turn-taking and multimodal listener behavior.
  • The analyses demonstrate corpus potential rather than definitive claims about language or culture, partly because participant relationships differ across portions.Future work will add annotations for addressees and transition-relevance places, expand to additional languages, and consider privacy-respecting release conditions.

Safe and Responsible Innovation Statement

TEIDAN’s human-participant recordings require privacy- and consent-aware handling, including compliance with individual redistribution restrictions.

  • Corpus release must respect participant-specific opt-outs restricting redistribution of video or both audio and video.Public examples should be anonymized where possible, and uses beyond the stated research purpose should be separately disclosed.
  • Models using TEIDAN should be examined for differences across language groups to avoid reinforcing overgeneralized assumptions about conversational behavior.
Loading 2609.00802v1…