Source-linked AI summary

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

Freeman Jiang, Ramon Sanabria, Soham Deshmukh, Bandhav Veluri, Simon Michael Vuch Williams, Elliott K. Suen, Garreth Lee, Kevin Yoonho Choi, Takuya Umeki, Riku Kubo, Sathvik Udupa, Chien-yu Huang, Shih-Yun Shan Kuan, Zhuoyan Tao, Satyapriya Krishna, Sefik Emre Eskimez, Yu Tsao, Hung-yi Lee, Shinji Watanabe

arXiv:2608.25218v1eess.AScs.CL

TL;DR

Turn-taking evaluation lacks consistent, linguistically grounded protocols and hand-annotated data spanning diverse conversation types. TURNBENCH addresses this with a triple-annotated, multi-domain corpus and unified evaluation of end-of-turn and interruption detection; across 14 systems, end-of-turn recall is stable by type, but interruption false positives vary strongly with conversation style, and no system is simultaneously fast, selective, and high-recall.

  • Problem

    Turn-taking evaluation lacks a consistent, linguistically grounded protocol and hand-annotated data covering diverse conversation types.

  • Method

    TURNBENCH combines a 30-hour triple-annotated dyadic corpus with a standardized protocol for evaluating end-of-turn and interruption detection across conversation types and heterogeneous systems.

  • Results

    Across 14 systems, end-of-turn recall is largely stable across conversation types, whereas interruption false positives are type-dependent; no system is simultaneously fast, selective, and high-recall.

  • Takeaways & Limitations

    TURNBENCH provides a reproducible public benchmark for comparing heterogeneous turn-taking systems under shared annotations and conversation-type conditions.

  • Takeaways & Limitations

    The corpus is limited to English, studio-recorded, dyadic dialogues, and the authors lack a methodology for evaluating interruption in full-duplex models.

Abstract

from arXiv · show

Speakers in natural conversation take turns speaking and listening, deciding in real time when to take, hold, or yield the floor. However, turn-taking evaluation remains limited due to the lack of a consistent, linguistically grounded evaluation protocol and hand-annotated data covering diverse conversation types. To address this, we present TurnBench, a multi-domain benchmark that pairs a 30-hour, hand-labeled corpus of dyadic human conversation with a standardized evaluation protocol for end-of-turn and interruption detection. We set conversation type as a controllable experimental variable, covering six distinct interaction styles, and triple-annotate each conversation. Benchmarking 14 heterogeneous turn-taking systems, we find end-of-turn recall stable across types, while interruption false positives are strongly type-dependent and concentrated in backchannel-dense interaction styles. Although in smooth floor transfers human listeners begin speaking a median 151 ms before the current turn ends, no current system performs equivalently without incurring excessive false positives. We release our corpus, a 104-hour training set, and a public leaderboard with an interactive dataset viewer at https://turnbench.sesame.com

I. INTRODUCTION

TURNBENCH addresses missing diverse, hand-annotated resources and inconsistent evaluation conventions for spoken turn-taking. It introduces a unified benchmark spanning conversation types and evaluates heterogeneous systems on end-of-turn and interruption detection.

  • Existing full-duplex models expose no direct turn-taking decisions, while available corpora lack annotations or cover only one conversational register.
  • Inconsistent definitions of backchannels, fillers, and interruptions make results across existing benchmarks difficult to compare.
  • TURNBENCH provides dyadic conversations balanced across six types, conversation-analysis-grounded event labels, and one scoring protocol for disparate implementations.
  • The benchmark includes a 30-hour, triple-annotated corpus, a 104-hour hand-labeled training set, and a public platform with leaderboard and dataset viewer.
  • 14 systems are benchmarked, with end-of-turn recall roughly stable across types but interruption false positives most frequent in casual, backchannel-rich conversation.

II. RELATED WORK

Prior resources differ in conversational coverage, linguistic grounding, and evaluation scope. TURNBENCH combines conversation-analysis-based taxonomy, six conversation types, and direct comparison of heterogeneous systems.

  • Conversation-analysis concepts distinguish transition-relevance places, where speaker change is legitimate, from backchannels, which signal attention without claiming the floor.
  • TURNBENCH compares corpora and benchmarks across dialogue phenomena, register count, linguistic grounding, and whether model evaluation is included.
  • Existing benchmarks use broader or idiosyncratic categories, while some resources cover only a single register or lack reliable end-of-turn and backchannel performance.
  • TURNBENCH covers six conversation types and enables direct comparison of open- and closed-source systems through a unified protocol.

III. CORPUS CONSTRUCTION

TURNBENCH treats conversation type as a controllable variable and constructs a roughly balanced corpus across six interaction styles. These styles are intended to elicit different human and system turn-taking behavior.

  • Conversation type is set at recording time as a manipulable variable hypothesized to produce measurable behavioral differences in humans and systems.
  • The corpus assigns each session one of six interaction styles: Casual, Task-Oriented, Instructional, Collaborative, Argumentative, or Narrative.
  • Casual dialogue features topic drift, humor, and many backchannels, whereas Task-Oriented dialogue centers on goal-directed exchange and clarifications.
  • Instructional dialogue uses asymmetric expert–learner turns, while Collaborative dialogue involves shared reasoning with frequent overlap and cooperative interruptions.
  • Argumentative dialogue contains structured disagreement with longer turns and fewer backchannels, whereas Narrative dialogue pairs storytelling with an active listener and little floor competition.
  • 13–21% of the corpus belongs to each conversation type, yielding roughly balanced coverage.

B. Event Taxonomy

The corpus preserves fine-grained, conversation-analysis-grounded annotations while mapping them to a smaller canonical taxonomy for evaluation. Recordings use controlled studio conditions and quality filtering.

  • Annotators label time-localized events using 17 fine-grained categories grounded in conversation-analysis literature.
  • For evaluation, the 17 labels are collapsed into seven canonical categories, while the fine annotations remain available in the released data.
  • All sessions were recorded in a professional studio using separate sound-isolated booths and dual-channel, high-resolution audio.
  • Recordings below a minimum quality threshold were discarded before annotation, and released audio is provided per speaker in FLAC format.

D. Sessions and Participants

TURNBENCH records approximately 30 hours of dual-channel dyadic speech and labels VAD-derived segments through independent annotation and deterministic consensus filtering.

  • Sessions and Participants: The corpus contains approximately 30 hours across 154 dialogues, recorded by 106 voice actors in 53 pairs.Sessions average approximately 11.7 minutes and begin with an unrecorded warm-up establishing the conversation type and topic.
  • Sessions and Participants: Three independent annotators assign each segment one fine-grained label after reviewing it in conversational context.Annotators prioritize audio over automatic transcripts when the two conflict and complete training with examples and a qualification test.
  • Sessions and Participants: Gold segments require a two-of-three canonical-label majority with endpoints agreeing within ±200 ms.Gold endpoints are medians across agreeing annotators; segments without a majority become excluded intervals.
  • Sessions and Participants: 0.77–0.80 pairwise Cohen’s κ and 0.78 Fleiss’ κ indicate high frame-level agreement after canonical mapping.Event-onset boundary F1 is 0.94–0.96 within ±200 ms, and 85.8% of annotator events survive gold filtering.
  • Sessions and Participants: The turn view yields 8,197 EOT anchors and 4,254 mid-turn pause negatives, while the label view yields 1,151 INT anchors.The label view is also scored against 17,511 BACKCHANNEL and NONCONTENT negatives.

B. Dataset Statistics

Conversation types show distinct turn-taking dynamics, and TURNBENCH supplements the benchmark with speaker-disjoint training and evaluation resources.

  • Dataset Statistics: Retention is about 88–89% for Instructional, Narrative, and Task-Oriented types versus about 84–85% for Casual, Collaborative, and Argumentative types.Retention is lower in the higher-overlap types.
  • Dataset Statistics: Argumentative dialogues contribute 377 consensus interruption events, while Casual and Collaborative dialogues show the fastest exchange and highest overlap.Instructional and Narrative dialogues have the longest turns and fewest interruptions.
  • Dataset Statistics: Human turn transfers have a median offset of −281 ms, or −151 ms when floor-taking interruptions are excluded.After a floor-taking interruption begins, the interrupted speaker continues for a median 1.48 s before ceding the floor.
  • Dataset Statistics: The training set provides approximately 104 hours of hand-labeled full-duplex dialogue and is speaker-disjoint from the benchmark.It uses the same protocol for baselines reported under the TURNBENCH-trained condition.
  • Dataset Statistics: The benchmark corpus uses a 25/75 speaker-disjoint, type-balanced dev/test split, with 38 public dev dialogues and 116 unlabeled test dialogues.Per-annotator tracks and metadata are also released.

V. THE TURNBENCH BENCHMARK

TURNBENCH evaluates end-of-turn and interruption detection as separate boundary-detection tracks using distinct gold views and event definitions.

  • End-of-Turn (EOT): The EOT track matches submitted timestamps against turn segment-ends where the conversational floor leaves the speaker.Because VAD splits turns into segments, many segment-ends are mid-turn pauses rather than true EOTs.
  • End-of-Turn (EOT): A positive EOT occurs when the floor passes to the other speaker or the speaker’s turn is final; same-speaker resumption defines a negative mid-turn pause.Negative spans are bounded so a real EOT cannot fall inside them.
  • Interruption (INT): The INT track defines an interruption as a listener’s mid-turn floor entry that takes the floor from the current speaker.The gold anchor is the interrupter’s onset on the interrupter’s channel.
  • Interruption (INT): Positive INT events are consensus floor-taking INTERRUPTION onsets, while BACKCHANNEL and NONCONTENT spans are negatives.Other listener events such as LAUGHTER are neither positives nor negatives.
  • Interruption (INT): Non-floor-taking interruptions and non-consensus floor-taking interruptions are excluded intervals rather than negatives.At onset, a non-floor-taking attempt is indistinguishable from a real interruption, so firing is neither rewarded nor penalized.

C. Evaluation Protocol

The protocol scores causal event timestamps against fixed positive windows and scored negative spans, reporting recall, false-positive rate, and latency at defined operating points.

  • Evaluation Protocol: For each gold event at time t, predictions within [t − τpre, t + τmax] are matched using τpre = 0.25 s and τmax = 3.0 s.The earliest unclaimed prediction in the window counts as a true positive.
  • Evaluation Protocol: Predictions in excluded intervals or outside scored positive and negative spans are ignored, while firing inside a negative span counts as at most one false positive.FPR therefore measures firing on scored negative spans.
  • Evaluation Protocol: The benchmark reports recall, false positive rate, and signed latency at the 10th, 50th, and 90th percentiles.A negative latency means the model committed before the gold boundary.
  • Evaluation Protocol: The leaderboard ranks test recall subject to a 0.15 false positive rate ceiling, while the dev budget is 0.1.Submissions exceeding the test ceiling rank below all qualifiers.
  • Evaluation Protocol: The study benchmarks 14 rule-based, academic, and deployed commercial systems using single- and dual-channel inputs.Accessible training pipelines are fine-tuned on TURNBENCH training data, whereas published models and commercial tools are scored as-is.

A. Rule-Based Heuristics

Rule-based and endpointing baselines detect turn events from acoustic activity or model outputs, with some systems adding semantic or causal context.

  • Rule-Based Heuristics: RMS VAD commits end-of-turn on speaker silence and interruption when the listener becomes active during the speaker’s turn.It uses fixed channel-energy thresholds and no linguistic information.
  • Rule-Based Heuristics: OpenAI Server VAD uses silence duration alone, whereas Semantic VAD waits longer when linguistic content suggests the turn is unfinished.
  • Rule-Based Heuristics: Kyutai SVAD combines streaming ASR with a semantic end-of-turn head, while SmartTurn v3 emits per-chunk turn-completion probabilities.
  • Rule-Based Heuristics: VAP predicts future voice activity per speaker and derives end-of-turn and interruption decisions from floor-hold probabilities.
  • Rule-Based Heuristics: WavLM variants operate causally with left context, while a per-channel ESPnet variant processes each speaker channel separately instead of mixed mono input.The WavLM-Large predictor is also evaluated at 25 Hz over 4-second current-anchored windows with 0 ms effective lookahead.

D. Full-Duplex Models

Full-duplex systems are evaluated through their produced audio, while continuous-output models use dev-set thresholding to produce causal event timestamps. Results expose trade-offs among latency, false positives, and recall.

  • D. Full-Duplex Models: Full-duplex evaluation streams one speaker’s channel into a live model session and detects end-of-turn decisions from the model’s recorded speech onsets.Gemini 3.1 Live and Moshi do not expose turn-taking labels.
  • Evaluation: The benchmark freezes the highest-recall threshold whose dev false-positive rate stays within 0.1, then evaluates that operating point on the test set.A common rising-edge commit rule with a 2-second refractory period converts continuous outputs into event times.
  • Results: VAP achieves 0.845 EOT recall at 0.055 FPR with 368 ms median latency and 0.945 INT recall at 0.107 FPR with 994 ms median latency.These are the strongest reported in-budget operating points on both tracks.
  • Results: Casual INT FPR exceeds Argumentative INT FPR for every model, showing that interruption false positives vary by conversation type.
  • Discussion: Acoustic-only RMS VAD and OpenAI Server VAD reach high recall at false-positive rates far above budget, unlike linguistically informed systems that remain in budget.The comparison isolates endpointing logic as the difference between the two OpenAI VAD modes.
  • Discussion: Interruption detection trades speed for selectivity: SmartTurn v3 commits at 159 ms but recovers 0.107 of interruptions, whereas delayed systems reach 0.87–0.95 recall at 0.05–0.11 FPR.VAP, Mimi-EP, and WavLM-Large delay beyond speech onset to discriminate interruptions from backchannels.

VIII. CONCLUSION

TURNBENCH provides a conversation-analysis-grounded benchmark and public resources for evaluating end-of-turn and interruption detection. Its conclusions are bounded by the English, studio-recorded, dyadic corpus and majority-consensus filtering.

  • VIII. CONCLUSION: TURNBENCH combines 30 hours of triple-annotated dyadic speech, a 104-hour training set, and a reproducible protocol for two turn-taking detection tasks.
  • VIII. CONCLUSION: Across 14 heterogeneous systems, end-of-turn recall is largely stable across conversation types, while interruption false positives are not.
  • VIII. CONCLUSION: No system is simultaneously fast, selective, and high-recall under the benchmark’s evaluation conditions.
  • VIII. CONCLUSION: The released corpus, scorer, leaderboard, and held-out test labels support reproducible evaluation while reducing contamination risk.
  • IX. LIMITATIONS AND FUTURE WORK: The corpus is limited to English, studio-recorded, dyadic dialogues, and majority-consensus filtering drops events without sufficient agreement.The paper identifies multilingual, non-studio, full-duplex interruption evaluation as future work.
Loading 2608.25218v1…