Source-linked AI summary
Do What I Say: A Spoken Prompt Dataset for Instruction-Following
Maike Züfle, Sara Papi, Fabian Retkowski, Szymon Mazurek, Marek Kasztelnik, Alexander Waibel, Luisa Bentivogli, Jan Niehues
TL;DR
Existing SLLM instruction-following benchmarks mostly use text prompts, limiting realistic evaluation of spoken interactions across styles, tasks, and languages. DOWIS provides multilingual parallel spoken and written prompts for nine tasks and benchmarks SLLMs, finding that text prompts outperform spoken prompts for text-output tasks while spoken prompts perform comparably or better for speech-output tasks.
Problem
Existing speech instruction-following benchmarks mainly use text prompts and have limited coverage of prompt styles, tasks, languages, and cross-lingual settings.
Method
DOWIS provides human-recorded parallel spoken and written prompts across nine tasks, 11 languages, and five styles for combination with existing benchmarks.
Results
Text prompts outperform spoken prompts for text-output tasks, whereas spoken prompts perform comparably or better for speech-output tasks; informal prompts consistently underperform.
Takeaways & Limitations
Text-only evaluation can overestimate SLLM capabilities, so realistic evaluation should account for prompt modality, style, language, and task type.
Abstract
from arXiv · showhide
Speech Large Language Models (SLLMs) have rapidly expanded, supporting a wide range of tasks. These models are typically evaluated using text prompts, which may not reflect real-world scenarios where users interact with speech. To address this gap, we introduce DoWhatISay (DOWIS), a multilingual dataset of human-recorded spoken and written prompts designed to pair with any existing benchmark for realistic evaluation of SLLMs under spoken instruction conditions. Spanning 9 tasks and 11 languages, it provides 10 prompt variants per task-language pair, across five styles. Using DOWIS, we benchmark state-of-the-art SLLMs, analyzing the interplay between prompt modality, style, language, and task type. Results show that text prompts consistently outperform spoken prompts, particularly for low-resource and cross-lingual settings. Only for tasks with speech output, spoken prompts do close the gap, highlighting the need for speech-based prompting in SLLM evaluation.
1. Introduction
Existing spoken instruction-following benchmarks are limited in naturalness, reusability, task coverage, and language scope. DOWIS addresses these gaps with native-speaker parallel spoken and textual prompts decoupled from task inputs, revealing modality-dependent SLLM performance.
- Limitations of Existing Benchmarks: Current spoken instruction-following benchmarks often use costly human recordings, while text prompts are more straightforward to generate.Spoken prompt collection requires human recording, whereas text prompts are commonly generated by LLMs or written by hand.
- Limitations of Existing Benchmarks: Existing benchmarks use text-to-speech instructions, cover only English and Chinese, concatenate task inputs, and cannot be reused with other datasets.They also emphasize general instruction-following and reasoning rather than task-specific, cross-lingual evaluation.
- DOWIS: DOWIS is a multilingual dataset of native-speaker parallel spoken and textual prompts that keeps instructions decoupled from task inputs.This design enables pairing with existing benchmarks while preserving naturalness and linguistic diversity.
- DOWIS: DOWIS spans 9 tasks and 11 languages, providing 10 prompt variants per task-language pair across five styles.It is designed to support realistic evaluation of SLLMs under spoken instruction conditions and to pair with existing benchmarks.
- Findings: For text-output tasks, text prompts significantly overestimate SLLM performance relative to spoken prompts, whereas spoken prompts perform on par or better for speech-output tasks.The evaluation covers Phi-4 Multimodal and Qwen2.5-Omni; speech-output examples include text-to-speech synthesis and speech-to-speech translation.
2. The DOWIS Prompt Dataset
DOWIS is a multilingual dataset of parallel spoken and textual instructions for nine speech and language tasks across 11 languages. It supports realistic, multifaceted evaluation by pairing its prompts with existing downstream benchmarks and recording prompts in simulated meeting scenarios.
- Dataset Overview: DOWIS provides parallel spoken and textual instructions for nine speech and language processing tasks across 11 languages.The dataset is designed to combine with existing downstream task benchmarks for instruction-following speech-model evaluation.
- Prompt Collection and Translation: Researchers created 10 English prompts per task by collecting two basic prompts and two rephrasings for each of four styles.The styles include formal, informal, detailed, and short, alongside the basic style.
- Prompt Collection and Translation: Native speakers translated the prompts into 10 additional languages, producing 990 unique text prompts.The total is calculated as 10 prompts × 9 tasks × 11 languages.
- Recording: Nineteen native or highly proficient speakers per language recorded the 90 prompts using phones or laptops to simulate realistic meetings.Audio was converted to .wav and trimmed using loudness-based voice activity detection with 10 ms windows and a −40 dBFS threshold.
- Statistics: The completed dataset contains 3h and 17m of audio across 11 languages, with most prompts averaging 4–5 seconds and ACHAP prompts averaging 16 seconds.ACHAP prompts are longer because they contain detailed formatting instructions.
3. Experiments
Experiments evaluate Qwen and Phi across five prompt styles and both text and speech modalities, using task-specific datasets and standard metrics. Speech-output tasks are assessed for both content accuracy and speech quality.
- Models: The study evaluates Qwen2.5-Omni-7B and Phi-4-multimodal-instruct with default inference settings on a single NVIDIA A100-SXM4-40GB GPU.Both models use batch size 1.
- Models: Each task uses five prompt styles—basic, formal, informal, short, and detailed—in both text and speech modalities.The styles are evaluated across the task suite, including cross-lingual tasks.
- Data: FLEURS supplies data for ASR, MT, ST, S2ST, and TTS across the listed language directions, with speech outputs restricted to English because Qwen supports only English speech generation.TSUM, SSUM, SQA, and ACHAP use separately described evaluation data, but the supplied passage is truncated before those details.
- Evaluation: ASR uses Word Error Rate (WER), while MT and ST use CometKiwi quality estimation without reference translations.CometKiwi is described as correlating well with human judgments for speech and text translation.
- Evaluation: TSUM, SSUM, and SQA use normalized BERTScore, while TTS and S2ST outputs are transcribed with whisper-large-v3 for content evaluation.TTS reports WER and S2ST reports CometKiwi after transcription.
- Evaluation: UTMOS is additionally reported for TTS and S2ST to assess speech quality.Content accuracy and speech quality are therefore evaluated separately for speech-output tasks.
4. Analysis
The analysis compares SLLM performance across text and speech prompts, then examines how prompt types affect results. Text prompts generally outperform speech, while informal and short prompts are consistently most challenging.
- General Trends: For tasks with text output, text instructions consistently outperform speech prompts, with only small differences for TSUM and SQA using Qwen.Phi sometimes exhibits severe failures with speech prompts.
- Female vs. Male Audio Prompts: Gender effects are inconsistent: Qwen prefers male prompts for TSUM and SSUM but female prompts for TTS, MT, ST, and S2ST.These comparisons use languages where both genders recorded all prompts.
- Female vs. Male Audio Prompts: 12% WER for both genders in TSUM coexists with BERTScore 43.88 vs. 42.93, suggesting intelligibility is not the primary driver of gender differences.The authors suggest speaker-related biases may instead contribute.
- Speech Prompts’ Interplay with Languages: Qwen shows a strong preference for text in ASR for Czech, Dutch, Portuguese, and Swedish, and in MT and ST for Czech, Dutch, and Swedish.For Czech and Swedish ASR, text-prompt WER already exceeds 100; Dutch and Portuguese ASR text WER is 31–37, while MT and ST text COMET reaches 76–82.
- General Trends: Informal and short prompts are consistently the most challenging across tasks, whereas formal and detailed prompts generally perform well.The results suggest models respond better to structured and explicit instructions.
- Speech Prompts’ Interplay with Prompt Types: For ASR, MT, and ST, text prompts outperform speech across all prompt types; for TTS, formal and detailed prompts perform better with speech.The TTS pattern is more nuanced because basic, informal, and short instructions perform differently.
5. Conclusion
DOWIS is a human-recorded parallel spoken-textual prompt dataset for realistic, comprehensive evaluation of spoken instruction-following in SLLMs. The conclusion emphasizes that text-only evaluation is overly optimistic, while prompt style, modality, language, type, and speaker gender affect performance.
- Dataset contribution: DOWIS introduces the first human-recorded parallel spoken-textual prompt dataset for evaluating spoken instruction-following in SLLMs.It covers nine tasks, 11 languages, and five prompt styles, and can be combined with task-specific benchmarks.
- Dataset contribution: DOWIS covers nine tasks, 11 languages, and five prompt styles for more realistic and comprehensive SLLM evaluation.The dataset is designed to pair easily with any task-specific benchmark.
- Findings: Informal prompts consistently underperform across tasks, while model performance varies across prompt types and speaker gender.The paper attributes informal prompts’ lower performance likely to their more colloquial nature.
- Implications: Text-based evaluation alone paints an overly optimistic picture of model capabilities, with prompt style, modality, and language affecting instruction-following performance.The conclusion presents DOWIS as a resource for multifaceted SLLM evaluation.