Source-linked AI summary
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
Dingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang, Xueyuan Chen, Tianhua Zhang, Helen Meng
TL;DR
Fine-grained perception and complex reasoning in natural speech remain insufficiently evaluated because spoken language contains semantic, paralinguistic, and phonological information beyond text. MMSU addresses this gap with a linguistically grounded benchmark of authentic speech across diverse perception and reasoning tasks, and its evaluation exposes substantial weaknesses in current SpeechLLMs.
Problem
Fine-grained perception and complex reasoning in natural speech remain largely unexplored, while existing benchmarks omit important spoken phenomena and rely heavily on synthetic audio.
Method
MMSU is a linguistically grounded benchmark using primarily real-world audio, expert review, and 47 tasks spanning spoken-language perception and reasoning.
Results
Evaluation reveals widespread phonological-perception challenges, difficulty handling complex reasoning, and specific subtask deficiencies across SpeechLLMs.
Takeaways & Limitations
MMSU provides a systematic framework for assessing multiple facets of spoken-language understanding and identifying areas for targeted SpeechLLM improvement.
Abstract
from arXiv · showhide
Speech inherently contains rich acoustic information that extends far beyond the textual language. In real-world spoken language understanding, effective interpretation often requires integrating semantic meaning (e.g., content), paralinguistic features (e.g., emotions, speed, pitch) and phonological characteristics (e.g., prosody, intonation, rhythm), which are embedded in speech. While recent multimodal Speech Large Language Models (SpeechLLMs) have demonstrated remarkable capabilities in processing audio information, their ability to perform fine-grained perception and complex reasoning in natural speech remains largely unexplored. To address this gap, we introduce MMSU, a comprehensive benchmark designed specifically for understanding and reasoning in spoken language. MMSU comprises 5,000 meticulously curated audio-question-answer triplets across 47 distinct tasks. To ground our benchmark in linguistic theory, we systematically incorporate a wide range of linguistic phenomena, including phonetics, prosody, rhetoric, syntactics, semantics, and paralinguistics. Through a rigorous evaluation of 14 advanced SpeechLLMs, we identify substantial room for improvement in existing models, highlighting meaningful directions for future optimization. MMSU establishes a new standard for comprehensive assessment of spoken language understanding, providing valuable insights for developing more sophisticated human-AI speech interaction systems. MMSU benchmark is available at https://huggingface.co/datasets/ddwang2000/MMSU. Evaluation Code is available at https://github.com/dingdongwang/MMSU.
1 INTRODUCTION
MMSU addresses gaps in spoken-language evaluation by combining fine-grained acoustic coverage, expert-reviewed authentic data, and broad assessment of perception and reasoning. Evaluation reveals persistent weaknesses in current SpeechLLMs, especially for phonological perception and complex reasoning.
- Motivation: SpeechLLM evaluation remains limited because existing benchmarks overlook many authentic spoken phenomena, rely heavily on synthetic audio, and lack systematic linguistic grounding.Overlooked phenomena include disfluencies, sarcasm, prosody variations, non-verbal sounds, mispronunciations, puns, and code-switching.
- MMSU: MMSU spans 47 perception and reasoning skills, providing comprehensive coverage of spoken-language understanding.The overview presents the benchmark as integrating acoustic features, quality assurance, and broad task coverage.
- MMSU: MMSU captures fine-grained acoustic information, including non-verbal sounds, accents, emotions, prosody, and intonation variations.The benchmark is designed to assess speech signals that convey meaning beyond literal textual content.
- MMSU: MMSU uses primarily real-world audio and expert review of tasks and questions to support acoustic authenticity, accuracy, and representativeness.Its data comes from open-source datasets and professional studio recordings rather than relying primarily on synthetic speech.
- Findings: Evaluation across 22 SpeechLLMs reveals widespread phonological-perception challenges, difficulty with complex reasoning, and deficiencies in specific subtasks.These findings identify targeted areas for future SpeechLLM improvement.
2 RELATED WORK
Prior SpeechLLM benchmarks cover general audio, dialogue, reasoning, or selected paralinguistic abilities, but provide limited depth on comprehensive spoken-language understanding. MMSU extends this landscape with a taxonomy spanning perception and reasoning across linguistic and paralinguistic domains.
- Existing benchmarks: Existing benchmarks evaluate instruction-tuned speech models, open-ended audio interaction, dialogue skills, general audio reasoning, or selected paralinguistic information.Examples include Dynamic-SUPERB, AIR-Bench, VoiceBench, ADU-Bench, MMAU, and SD-Eval.
- Existing benchmarks: Prior benchmarks often provide limited depth in spoken-language understanding or emphasize semantic speech content while underrepresenting diverse acoustic phenomena.The related-work discussion identifies insufficient attention to the acoustic features that characterize speech phenomena.
3 MMSU BENCHMARK
MMSU is a linguistically grounded benchmark for spoken-language perception and reasoning, designed to cover authentic acoustic and linguistic phenomena across a broad task hierarchy. It contains 5,000 expert-annotated MCQs spanning 47 tasks, with balanced coverage of perception and reasoning capabilities.
- Overview: MMSU assesses the full spectrum of spoken-language understanding and complex reasoning through 5,000 expert-annotated MCQs across 47 tasks.The benchmark includes 24 perception tasks and 23 reasoning tasks.
- Benchmark structure: Its three-level hierarchy separates perception from reasoning, then organizes both dimensions into linguistic and paralinguistic categories and finer linguistic subfields.Perception extracts basic audio information, whereas reasoning integrates contextual semantics with acoustic information for deeper interpretation.
- Benchmark structure: MMSU task design draws on phonetics, prosody, rhetoric, syntactics, semantics, and paralinguistics to represent authentic spoken-language phenomena.The design is guided by linguistic theory to support theoretical soundness and practical relevance.
- Data construction: The four-stage construction process covers linguistic task design, question and option collection, authentic audio recording, and manual review with quality control.Real-world recordings are prioritized, while professional voice actors provide targeted phonology-related recordings where open-source coverage is lacking.
- Statistics: 2,580 perception questions and 2,420 reasoning questions comprise the 5,000-question benchmark, whose distribution is balanced across tasks.Within reasoning, semantics and phonology account for 22.16% and 19.54%, respectively.
- Comparison with previous benchmarks: Compared with prior benchmarks, MMSU offers broader acoustic and linguistic coverage and deeper reasoning that integrates paralinguistic, phonetic, and semantic information.The comparison identifies MMSU as systematically incorporating linguistically grounded phenomena into spoken-language understanding evaluation.
4 EXPERIMENTS
The experiments systematically evaluate 22 models on multiple-choice audio questions using standardized prompts and balanced answer positions. Human performance is also measured on a 1,000-instance sample as a reference baseline.
- Evaluation strategy: Each model answers four-option audio-question items with randomly ordered and balanced choices under the same optimized instruction-following prompts.This strategy is intended to ensure fairness and minimize prompt-induced variance.
- Human evaluation: 15 students evaluate a random sample of 1,000 instances, and their average score serves as the human reference baseline.Evaluators receive the same instructions used in model evaluation.
5 RESULTS AND DISCUSSION
MMSU exposes substantial weaknesses in current SpeechLLMs, especially fine-grained acoustic and phonological perception, despite some competitive models and robustness under noise. Error and task analyses show that performance varies widely across capabilities, with persistent failures in complex reasoning and specialized speech understanding.
- Main results: 89.72% human accuracy exceeds Gemini-1.5-Pro’s 60.68%, revealing a substantial performance gap on MMSU.The benchmark therefore remains challenging even for the best-performing evaluated model.
- Main results: 60.57% and 59.28% accuracies make Qwen2.5-Omni-7B and Kimi-Audio competitive with proprietary models.Qwen2.5-Omni-7B trails Gemini-1.5-Pro by only 0.11%, while GPT-4o-Audio reaches 56.38%.
- Main results: Current models struggle more with fine-grained acoustic perception than humans, whose perception accuracy exceeds reasoning accuracy, 91.24% versus 86.77%.Models perform relatively better on semantic reasoning but remain weak on subtle acoustic and non-verbal signals.
- Main results: 53.60% is the best model accuracy on phonology-related perception tasks, despite substantially higher semantic-task scores.The weakness extends to reasoning tasks involving phonological cues, including rhythm, prosody, and pronunciation.
- Task-specific analysis: Near-homophone, consonant, vowel, and syllable perception are generally weak, whereas speech grounding and gender prediction are stronger.Simpler, clearer reasoning tasks outperform sarcasm detection, couplet matching, and background scene recognition.
- Performance under noisy conditions: Noise causes only minor performance drops, with Gemini-1.5-Pro and Qwen2.5-Omni remaining relatively stable under stronger corruption.Noise-Level 1 uses half the original waveform amplitude; Noise-Level 2 uses equal amplitude.
- Error analysis: Perceptual errors dominate sampled failures, while models also show persistent complex-reasoning and domain-knowledge limitations.The analysis sampled 300 mispredictions per model and reports GPT-4o-Audio’s answer-extraction errors at 14.7%.
6 CONCLUSION
MMSU is a comprehensive benchmark for spoken-language understanding and reasoning, integrating linguistic theories across 47 tasks and 5,000 audio samples. Evaluation across 22 models found that even the best model reached only 60.68% accuracy, while a leaderboard was planned for ongoing comparison.
- MMSU covers 47 tasks and 5,000 curated audio samples across phonetics, prosody, rhetoric, syntax, semantics, and paralinguistics.
- 60.68% accuracy was achieved by the best-performing model among 22 evaluated open-source and proprietary models.
- The planned leaderboard is intended to provide a consistent platform for accessing and comparing model performance.
Supplementary Material
LLMs supported MMSU data construction and manuscript polishing, while human review was used to check generated data.
- LLMs generated task questions and multiple-choice options during data construction.
- All generated data underwent human review for reliability and correctness.
- LLMs were used for copyediting and phrasing to improve clarity and fluency.
B DATA SOURCES
MMSU draws on open-source speech, emotion, reasoning, multilingual, accent, child-language, and synthetic-speech resources. These datasets support tasks spanning linguistic and audio phenomena.
- SLURP provides approximately 72,000 English home-assistant recordings across 18 domains for semantic understanding tasks.
- The source set includes corpora for code-switching, spontaneous speech, logical reasoning, multilingual speech, accents, speech generation, and child language.
- FoR contains more than 195,000 human and computer-generated utterances for fake-audio detection.
- RAVDESS supplies 7,356 files performed by 24 actors and covering seven emotions in speech.
- Figure 6 presents the data-volume distribution across MMSU tasks.
C MMSU DATA DISTRIBUTION
MMSU distributes data relatively evenly across its 47 tasks and combines open-source, custom-recorded, and synthetic audio sources.
- Approximately 90 to 120 samples occur per task across the 47-task benchmark.This distribution supports representation of semantics, syntax, phonetics, sociolinguistics, and paralinguistics.
- 76.74% of MMSU data comes from open-source audio, 13.44% from custom recordings, and 9.82% from Azure TTS.
- Table 5 summarizes the audio-source distribution of MMSU.
D TASKS DETAILS
MMSU defines 47 tasks spanning perception and reasoning across acoustic, linguistic, semantic, prosodic, and paralinguistic dimensions. The tasks cover both basic speech-feature extraction and context-dependent interpretation.
- Task organization: MMSU organizes its 47 tasks around perception and reasoning abilities within a three-level linguistic hierarchy.Perception extracts basic audio information, while reasoning involves more complex interpretation.
- Linguistic grounding: Tasks are linked to linguistic subfields including phonetics, prosody, semantics, and paralinguistics.This classification connects task design to linguistic descriptions of speech production, perception, and interpretation.
- Perception: Perception tasks include volume, stress, emotion, speaker identity, intonation, content, vocal range, gender, duration, sounds, disfluencies, speakers, turns, and prolonged sounds.These tasks target acoustic, phonological, semantic, and paralinguistic properties of speech.
- Spoken phenomena: The benchmark explicitly includes non-verbal sound detection, disfluency detection, speaker counting, dialogue-turn counting, and prolonged-sound perception.These tasks address phenomena often present in spontaneous spoken language.
- Reasoning: Reasoning tasks cover stress, continuation, emotional context, dialogue, background scenes, intent, synthetic speech, and causal interpretation.They require interpreting spoken content, context, or speaker characteristics beyond isolated signal recognition.
D.2 TASK EXAMPLES
The task examples instantiate MMSU’s coverage through multiple-choice questions about speech content, acoustic variation, linguistic structure, and contextual reasoning. Examples pair transcribed or described audio with targeted answer choices.
- Acoustic perception: Examples test acoustic patterns such as volume, vocal range, pitch, speed, stress, pauses, and intonation through ordered or categorical choices.Several examples present the same speaker or segment with three varying acoustic levels or patterns.
- Linguistic perception: Other examples assess content and linguistic perception through transcription, syllable counting, consonant-vowel identity, near-homophones, and prolonged sounds.Questions ask models to select heard words, sound properties, syllable counts, or elongated words.
- Paralinguistic tasks: Paralinguistic examples ask about gender, accents, non-verbal sounds, speaker identity, and the number of speakers or dialogue turns.The examples use descriptions such as a child’s voice, an Indian accent, a cry, or multiple speakers.
- Conversational phenomena: Disfluency and dialogue examples require identifying filled pauses, discourse markers, restarts, speakers, and turn transitions.The supplied dialogue and spontaneous transcript illustrate how these phenomena are embedded in natural speech.
- Reasoning tasks: Reasoning examples require interpreting speech acts, emotional context, references, stress implications, causal relations, and coherent continuations.Questions connect utterance meaning or audio context to scenarios, referents, consequences, or likely continuations.
E ERROR CASES ANALYSIS
The error analysis separates model and human error sources and shows that perceptual errors dominate model failures. It also identifies instruction-following, knowledge, refusal, and evaluator-attention issues.
- Error taxonomy: Model error categories include perceptual errors, reasoning errors, lack of knowledge, answer rejection, and answer extraction errors.Distraction and difficulty in answering are categorized as human rather than model errors.
- Knowledge limitations: Some errors reflect insufficient knowledge or context, including difficulty with intonation across English accents.The accent example attributes an incorrect answer to lacking intonation knowledge of different English accents.
- Response and evaluation issues: Models sometimes refuse to answer or fail to follow the required answer format, while evaluators can lose concentration.These issues affect answer extraction and human evaluation separately.
F.3 HUMAN EVALUATION
Human evaluation uses sampled MMSU entries, while GPT-based prompts generate questions, options, summaries, translations, and continuations. The prompt designs emphasize plausible distractors and format-controlled outputs.
- Human evaluation: Human evaluation recruits 15 students to answer a randomly sampled set of 1,000 MMSU entries through a dedicated review interface.Each evaluator listens to an audio clip and selects the answer corresponding to its question.
- Code-switching prompts: The code-switching prompt requests a challenging multiple-choice question requiring careful understanding of both languages and contextual clues.It specifies one correct answer and three plausible incorrect options containing explicit errors.
- Other prompt designs: Additional prompts generate emotional scenarios, translations, and coherent continuations under explicit length, relevance, and output-format constraints.The continuation prompt limits responses to 50 words and requires matching the input’s style and context.
- Idiom prompts: The idiom prompt generates a multiple-choice question whose first option is the figurative interpretation and whose remaining options are plausible literal or otherwise incorrect readings.The design tests distinction between idiomatic and superficial interpretations.
- Summarization prompts: The summarization prompt requires four concise options, with the first as the accurate summary and the others containing specified error types.Errors may concern the main idea, details, causality, or sentiment.