Source-linked AI summary
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
Sonal Kumar, Šimon Sedláček, Vaibhavi Lokegaonkar, Fernando López, Wenyi Yu, Nishit Anand, Hyeonggon Ryu, Lichang Chen, Maxim Plička, Miroslav Hlaváček, William Fineas Ellingwood, Sathvik Udupa, Siyuan Hou, Allison Ferner, Sara Barahona, Cecilia Bolaños, Satish Rahi, Laura Herrera-Alarcón, Satvik Dixit, Siddhi Patil, Soham Deshmukh, Lasha Koroshinadze, Yao Liu, Leibny Paola Garcia Perera, Eleni Zanou, Themos Stafylakis, Joon Son Chung, David Harwath, Chao Zhang, Dinesh Manocha, Alicia Lozano-Diez, Santosh Kesiraju, Sreyan Ghosh, Ramani Duraiswami
TL;DR
Comprehensive evaluation of audio intelligence remains underserved despite its importance for generally intelligent AI systems. MMAU-Pro addresses this gap with a broad, expert-annotated benchmark spanning diverse audio skills and challenging reasoning settings, and evaluations show that leading models still struggle, especially on complex multi-audio, spatial, and open-ended tasks.
Problem
Audio intelligence evaluation remains underserved, while progress toward artificial general intelligence is incomplete without strong capabilities across diverse and complex audio.
Method
MMAU-Pro is a human-involved benchmark-construction pipeline with 5,305 expert-annotated question–answer pairs covering 49 skills across speech, sounds, music, mixtures, long audio, spatial reasoning, and multi-audio analysis.
Results
Leading models achieve only moderate accuracy on single-modality tasks, typically 20–50% on mixed modalities, rarely surpass 40% on spatial audio, and remain below 30% on multi-audio reasoning.
Takeaways & Limitations
Multi-audio, spatial reasoning, free-form answering, instruction following, and nuanced temporal, prosodic, and emotional understanding remain clear directions for model and benchmark development.
Takeaways & Limitations
The benchmark does not explore all dimensions of audio processing and reasoning ability, and fully benchmarking human auditory capabilities remains a work in progress.
Abstract
from arXiv · showhide
Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly ``from the wild" rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems' progression toward audio general intelligence. The benchmark and code is available at https://sonalkum.github.io/mmau-pro.
Introduction
MMAU-Pro addresses the limited coverage of auditory-intelligence evaluation with a broad benchmark spanning core audio domains, mixtures, and challenging reasoning settings. It combines expert annotation, real-world audio, diverse skills, and model evaluation to expose persistent weaknesses in current systems.
- MMAU-Pro targets an underserved gap in evaluating audio intelligence as a component of general intelligence.
- Questions require deliberate multi-hop reasoning and include both multiple-choice and open-ended formats, while audio is drawn from the wild.
- The benchmark contains 5,305 expert-annotated question–answer pairs spanning 49 auditory-intelligence skills.
- It covers speech, environmental sounds, music, mixtures, long-form audio up to ten minutes, multi-audio and spatial reasoning, voice-chat comprehension, and instruction following.
- 59.2% accuracy for Gemini 2.5 Flash and 51.7% for Audio Flamingo 3 show that leading models still face substantial challenges.
Related Work
Prior audio benchmarks established broad and specialized evaluation settings but remain shallow, fragmented, or limited in scale, data sources, and task coverage. MMAU-Pro addresses these gaps by jointly testing underexplored dimensions relevant to real-world audio intelligence.
- Earlier benchmarks evaluated speech, sounds, music, reasoning, spoken language, instruction following, or hallucinations, but generally focused on narrower settings.
- Existing evaluations remained shallow and fragmented despite some progress on multi-audio and spatial-audio tasks.
- No existing benchmark systematically combined long-form audio up to ten minutes, multi-audio, spatial, open-ended, instruction-following, and multicultural scenarios.
- MMAU-Pro was designed to provide comprehensive coverage of these underexplored dimensions for real-world audio general intelligence.
The MMAU-Pro Benchmark
MMAU-Pro is a human-curated benchmark built to test diverse audio skills and realistic reasoning conditions beyond conventional short, single-audio multiple-choice evaluation. Its design combines expert validation, in-the-wild recordings, long-duration categories, open-ended responses, and specialized audio scenarios.
- Benchmark Design: MMAU-Pro contains 5,305 expert-annotated question–answer pairs covering 49 skills across speech, sounds, music, and combinations.
- Benchmark Design: Questions require multi-step reasoning, are authored and validated by domain experts, and use in-the-wild audio except for the spatial subset.
- Evaluation Formats: The benchmark includes open-ended responses and multiple-choice questions with up to ten options, reducing reliance on random guessing.
- Data Curation, Annotation and Validation: Its curation pipeline covers task design, expert allocation, manual audio and QA creation, expert distractors, independent verification, and final review.
- Task Coverage: MMAU-Pro explicitly evaluates long-form comprehension, multi-audio reasoning, multicultural music, spatial audio, voice QA, and instruction following.
Experimental Setup
The evaluation covers large audio-language models, cascaded captioning or transcription systems, and both multiple-choice and open-ended outputs. It uses embedding similarity for answer-choice selection and an LLM judge for open-ended response quality.
- Large Audio-Language Models are evaluated on long- and short-form reasoning, spatial understanding, multicultural music, and multi-audio comparisons.
- Cascaded systems first generate captions for sound and music or transcripts for speech, then pass these with questions to text-only LLMs.
- For multiple-choice questions, the system selects the answer whose pretrained-transformer embedding is most similar to the model’s output embedding.
- Open-ended responses are judged on correctness, relevance, completeness, clarity, and an overall score, with 1–5 scores converted to percentages.
Results and Discussion
MMAU-Pro exposes persistent weaknesses in audio-language models as tasks move beyond single-domain understanding toward mixed, spatial, multi-audio, open-ended, and instruction-following settings. Performance also varies substantially by skill, question phrasing, musical culture, and whether models can rely on acoustic evidence.
- Cross-modal performance: 30–60% accuracy is typical on core Sound, Music, and Speech tasks, while Qwen2.5-Omni-7B rarely exceeds 65% on Music or Speech.These results indicate that foundational audio understanding remains challenging even for strong open-source models.
- Cross-modal performance: 20–50% accuracy is typical on mixed-modality tasks, and top models rarely surpass 40% on spatial audio understanding.The reported degradation suggests difficulty fusing information across multiple audio streams.
- Reasoning and instruction following: 71.7% is Gemini 2.5 Flash’s score on voice-chat reasoning and STEM knowledge inference, compared with 60% for Qwen2.5-Omni-7B; smaller or less instruction-tuned models often fall below 50%.Most models nevertheless score between 25% and 60% on these tasks.
- Benchmark sensitivity: Up to 5% gains from LLM-generated question rephrasing show that Phi4-MM-Instruct and Qwen2.5-Omni-3B are sensitive to surface-level language cues.Phi4-MM-Instruct rises from 37.5% on originals to 41.3% on GPT4o rewrites, while Qwen2.5-Omni-3B rises from 43.4% to 48.6%.
- Benchmark sensitivity: 52.2% to 30.6% is Qwen2.5-Omni-7B’s drop when clean audio is replaced with noise, while AF3 improves from 47.2% noise-only to 51.7% with actual audio.The clean-versus-noise gap indicates that benchmark performance depends on acoustic input rather than language priors alone.
Conclusion, Limitations and Future Work
MMAU-Pro broadens holistic audio-intelligence evaluation with diverse, demanding tasks, while showing that leading models still struggle across categories. The authors acknowledge incomplete coverage and outline extensions for languages, environments, interaction, instruction following, and culturally grounded reasoning.
- Conclusion: MMAU-Pro evaluates 5,305 expert-annotated QA pairs across 49 skills spanning speech, sounds, music, and combinations.
- Conclusion: Its innovations include long-audio understanding, multi-audio reasoning, spatial audio comprehension, multicultural music understanding, and instruction following.
- Limitations: Even the strongest evaluated open and proprietary LALMs struggle across several categories.These tasks require advanced perception, contextual understanding, and complex reasoning.
- Limitations: The authors do not claim to explore all dimensions of audio processing and reasoning.
- Future Work: Future work will expand languages and low-resource acoustic environments, add interactive streaming tasks, refine instruction-following evaluation, and improve paralinguistic and culturally grounded reasoning metrics.
A Curating the Audio Set
The audio set combines in-the-wild recordings from everyday videos, spoken-content sources, sound effects, and full musical tracks across cultures and genres.
- A Curating the Audio Set: In-the-wild audios are sourced from varied materials to reduce contamination with large audio-language models.
- A Curating the Audio Set: Sound data includes YouTube videos of everyday real-world tasks and Adobe Sound Effects recordings.
- A Curating the Audio Set: Speech and mixed-speech clips come from TV, reality shows, podcasts, movies, YouTube Shorts, and the CASPER dataset.
- A Curating the Audio Set: Music data consists of full tracks spanning multiple cultures and genres.
B Annotation Guidelines
Annotation procedures require audio-only solvability, complete listening, valid audio inputs, English questions, and task labeling, while rephrasing prompts preserve meaning without adding context.
- B Annotation Guidelines: Annotators filter videos that can be used without visual cues and listen to the complete audio before writing QA pairs.
- B Annotation Guidelines: Questions must contain one or more non-corrupt audios, be written in English, and carry a task-type tag.
- B Annotation Guidelines: The rephrasing prompt requires clear, concise wording that preserves the original question’s meaning and context.
- B Annotation Guidelines: The rephrased question must not add information or context absent from the original.
- B Annotation Guidelines: When a prior rephrasing is supplied, the new version must avoid matching it while retaining the original meaning.
Prompt used for open-ended Evaluation
Open-ended responses are evaluated by an expert-evaluator prompt that scores factual accuracy, question alignment, coverage, clarity, and an overall average in a fixed format.
- Prompt used for open-ended Evaluation: The evaluation prompt compares a model response with the question and reference answer.
- Prompt used for open-ended Evaluation: Responses receive 1–5 scores for correctness, relevance, completeness, and clarity.
- Prompt used for open-ended Evaluation: Each criterion requires a score and a brief one- to two-sentence justification.
- Prompt used for open-ended Evaluation: The output format includes four criterion scores, justifications, an overall average score, and an overall assessment.
D Open-ended QA Evaluation Ablations
The evaluation compares LLM-as-a-judge scores with human correctness ratings for open-ended audio question-answer responses on MMAU and MMAR.
- Human annotators rated open-ended answers for correctness on a 1 to 5 scale.Each answer was rated by at most five humans, whose scores served as the gold standard.
- The study used several large audio language models to generate open-ended responses.Responses were collected on the MMAU test mini and MMAR benchmarks.
- Table 6 reports Spearman’s ρ and Kendall’s τ correlations between human scores and LLM-as-a-judge evaluations.The correlations assess agreement for answer correctness.
E Skill wise examples and their definitions
The appendix presents examples and brief definitions for MMAU-Pro skills across music, sound, and speech, covering both perceptual and reasoning tasks.
- Skill-wise examples and definitions: Tables 7–12 provide examples of skills under the reasoning and perception categories.The tables also include brief descriptions of these skills.
- Music skills: Music skills are organized into perceptual and reasoning categories with domains, definitions, and example questions.Table 7 focuses on multiple-choice perceptual music questions, while Table 8 covers multi-step reasoning in music.
- Sound skills: Sound skills are divided into perceptual and reasoning categories with domains, definitions, and example questions.Tables 9 and 10 describe sound perception tasks and deep reasoning over sound, respectively.
- Speech skills: Speech skills include separate perceptual and reasoning tables containing domains, definitions, and example questions.Table 11 addresses speech perception, whereas Table 12 addresses speech reasoning tasks.