Source-linked AI summary
Massive Sound Embedding Benchmark (MSEB)
Georg Heigold, Ehsan Variani, Tom Bagby, Cyril Allauzen, Ji Ma, Shankar Kumar, Michael Riley
TL;DR
MSEB addresses the lack of holistic, standardized evaluation for audio embeddings across the capabilities required by intelligent multimodal systems. It provides an extensible benchmark with eight tasks, diverse datasets including SVQ, and model-agnostic evaluation, finding substantial performance headroom and gaps between sound-based methods and text-based references.
Problem
Audio intelligence spans many capabilities, but existing evaluation is fragmented and lacks a standardized way to measure broad task performance, efficiency, and general-purpose representations.
Method
MSEB combines eight auditory super tasks, diverse datasets including the multilingual SVQ resource, and an extensible model-agnostic library for standardized evaluation.
Results
MSEB experiments reveal substantial performance headroom, with sound-based approaches lagging text-based counterparts across tasks and languages and varying strongly across sound types, languages, and noise conditions.
Takeaways & Limitations
MSEB provides a common platform for assessing and advancing robust sound representations in real-world multimodal applications.
Abstract
from arXiv · showhide
Audio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, each task involves transforming a raw audio signal into a meaningful 'embedding' - be it a single vector, a sequence of continuous or discrete representations, or another structured form - which then serves as the basis for generating the task's final response. To accelerate progress towards robust machine auditory intelligence, we present the Massive Sound Embedding Benchmark (MSEB): an extensible framework designed to evaluate the auditory components of any multimodal system. In its first release, MSEB offers a comprehensive suite of eight core tasks, with more planned for the future, supported by diverse datasets, including the new, large-scale Simple Voice Questions (SVQ) dataset. Our initial experiments establish clear performance headrooms, highlighting the significant opportunity to improve real-world multimodal experiences where audio is a core signal. We encourage the research community to use MSEB to assess their algorithms and contribute to its growth. The library is publicly hosted at github.
1 Introduction
MSEB addresses the fragmented evaluation of auditory intelligence by benchmarking how audio embeddings support capabilities from perception to reasoning. It contributes a broad, extensible benchmark suite, a multilingual dataset, and evidence of substantial remaining performance headroom.
- Motivation: Auditory intelligence requires transforming raw audio into meaningful embeddings that support transcription, segmentation, retrieval, reasoning, and clustering.These capabilities span both low-level perception and higher-level cognitive functions.
- Motivation: Existing audio embedding technologies are fragmented across speech, speaker verification, and general sound classification, limiting assessment of generalized machine audition.Speech systems commonly use discrete textual embeddings, while monitoring systems often use dense continuous vectors.
- Benchmark gap: MSEB fills the need for a standardized benchmark measuring performance across tasks, compression rates, computational complexities, and general-purpose audio representations.The benchmark is positioned as an audio counterpart to broad embedding benchmarks for text and images.
- Contributions: MSEB introduces eight tasks, an extensible open-source library, and the large-scale SVQ dataset spanning 26 locales and 17 languages.SVQ is designed to support evaluation across multiple tasks from one source.
- Contributions: MSEB emphasizes real-world applicability, embedding efficiency, user-centric retrieval and reasoning, and evaluation without task-specific fine-tuning.The benchmark is intended as a dynamic platform for community contributions.
2 Massive Sound Embedding Benchmark (MSEB)
MSEB operationalizes auditory evaluation through diverse datasets, eight application-oriented super tasks, standardized metrics, and a model-agnostic execution library. Its tasks cover information access, core perception, organization, and generation while supporting efficient bulk inference and leaderboard evaluation.
- Framework: MSEB evaluates multimodal systems through tasks that map curated dataset inputs to required prediction formats, using audio with optional information from other modalities.The benchmark is designed to evaluate auditory capabilities regardless of whether sound is processed alone or within a multimodal system.
- Datasets: Its datasets prioritize public access, nontrivial complexity, large scale, verifiable labels, and diverse sound coverage.These criteria are intended to support statistically reliable and broadly relevant evaluation.
- Datasets: The initial release includes four datasets covering spoken queries, multilingual spoken language understanding, general sound events, and bioacoustics.The release currently uses sound and text modalities, with expansion to other modalities planned.
- Datasets: SVQ contains over 177k spoken queries across 26 locales and 17 languages, four acoustic environments, rich textual metadata, speaker information, and temporal annotations.The collection is released as a single undivided dataset.
- Tasks: The eight super tasks progress from Information Access through Core Perception to Organization & Generation, covering retrieval, reranking, reasoning, classification, transcription, segmentation, clustering, and reconstruction.Each super task contains task variations with strict input and output specifications aligned to realistic applications.
- Information Access: Retrieval searches a knowledge index for relevant documents, passages, or spans, including same-language and cross-language variants across six subtasks.SVQ supports retrieval evaluation across 26 locales and four recording environments.
- Information Access: Reranking orders acoustically plausible text hypotheses by relevance to the spoken signal, while reasoning extracts an answer span or identifies that no answer exists.These tasks target ambiguity resolution and precise question answering from provided multimodal context.
3 Experiments
Across eight super tasks, sound-based systems consistently trail text-input oracles, revealing substantial headroom and strong dependence on language, noise, and task conditions. Results also show that specialized encoders can work well for particular domains, while transcription-dependent pipelines remain brittle.
- Overview: Sound-based approaches significantly lag behind text-input oracles across all tasks, leaving substantial opportunity to improve auditory representations.The comparison uses sound as the typical input and ground-truth transcripts as an oracle scenario.
- Retrieval and multilinguality: ASR WER does not perfectly predict retrieval MRR, while cross-language performance gaps persist for both sound and text modalities.Passage retrieval shows a relatively smaller cross-language gap than full-page retrieval.
- Transcription: ASR quality varies widely, with WER ranging from 7.7% for Arabic in clean conditions to over 100% for Malayalam.Noise also affects languages differently; English WER rises from 1.6% in clean conditions to 3.8% with background speech.
- Clustering: Domain-appropriate encoders achieve strong clustering results, including V-measures up to 0.53 for BirdSet and 0.68 for FSD50K.Raw spectrograms also reach 0.97 V-measure for speaker identification in te-IN, matching or exceeding pretrained models in many locales.
- Reconstruction: Audio reconstruction remains highly non-universal, with FAD ranging from approximately 15k for English speech to approximately 778k for environmental sounds.English traffic-noise reconstruction reaches approximately 490k FAD, showing further degradation in noisy conditions.
4 Conclusion
MSEB provides a broad, flexible platform for evaluating sound capabilities across diverse tasks, datasets, and model architectures. Initial experiments reveal substantial performance headroom and motivate continued community contributions.
- Conclusion: MSEB combines eight diverse super tasks, four large-scale datasets, and a model-agnostic library for evaluating intelligent systems across the sound spectrum.The library supports architectures ranging from conventional models to LLMs.
- Conclusion: Initial experiments reveal substantial performance headroom, motivating further research and community contributions to the benchmark.The benchmark is intended as a flexible platform for continued evaluation and growth.
A SVQ Dataset
SVQ is an open-source dataset supporting MSEB tasks, with multilingual recordings, speaker metadata, and a single undivided release that preserves data volume but complicates independent splitting.
- SVQ is an open-source collection curated to support a variety of MSEB tasks.
- Splits: The data is released as one undivided evaluation set because speaker- and text-disjoint splits would reduce recordings by around 40%.Users who train on SVQ must design their own splits and balance data volume against strict speaker and text disjointness.
- Linguistic Composition: The dataset contains 26 language locales representing 17 languages, with multiple locales for languages including English, Arabic, Bengali, and Urdu.
- Speaker Information: SVQ provides anonymous speaker identifiers, self-reported gender, and age for 700 participants, with a generally balanced female and male composition.
- Speaker Information: Participants are primarily adults: the overall median age is 29 years, with ages ranging from 18 to 71.
B Retrieval
MSEB retrieval spans page, passage, and span granularities in both in-language and cross-language settings, revealing substantial and language-dependent headroom between ASR-based systems and stronger references.
- Task Definition: MSEB defines page, passage, and span retrieval tasks, each evaluated in in-language and cross-language variants.
- Indexes: Retrieval indexes vary substantially: passage indexes contain 271,711 or 112,426 passages, while full page indexes range from thousands to millions of pages.
- Page Retrieval: Pages exceeding the 8,000-token context limit are chunked, and their page representation is the mean of the chunk embeddings.
- Results: Headroom varies with language, language-match condition, and index size; it can persist even with nearly perfect text embedders.
- Results: Whisper word error rates range from 10% in English to 100% in Malayalam, demonstrating strong language variation.
- Results: Gemini Embedding consistently outperforms Gecko, while substantial headroom remains between ASR-based models and ground-truth transcripts.
C Reranking
MSEB query reranking constructs acoustically and semantically challenging candidates, and cascade results show language-dependent headroom relative to text-based references.
- Task Design: The reranking task tests acoustic discrimination using phonetically similar queries and semantic discrimination using plausible distractors.
- Task Design: Candidates are augmented with the ground-truth query, deduplicated, sorted by increasing WER, and truncated to at most 250 items.
- Results: The text-based baseline achieves optimal performance, such as MRR=1, because the ground-truth query is compared against itself.
- Results: Reranking headroom varies across languages, with Gemini Embedding consistently outperforming Gecko and ASR quality limiting downstream performance.
D Reasoning
MSEB reasoning asks models to extract answer spans from provided contexts or identify when no answer exists, with passage- and page-level variants. Initial baselines include generative and retrieval approaches, but retrieval performs poorly without task-specific fine-tuning.
- Task Definition: Reasoning requires predicting an answer span from a context or indicating that no answer is present.
- Task Definition: MSEB distinguishes passage-level reasoning from page-level reasoning according to the supplied context.
- Baselines: Baseline experiments compare a generative Gemma 3 approach with a retrieval approach using Gemini embedding.
- Baselines: The retrieval approach uses a 0.8 threshold to distinguish real answers from “No Answer” predictions.
- Results: Retrieval performs poorly, reportedly because Gemini embedding and Gecko were not fine-tuned for this task.
E Classification
MSEB evaluates classification across speech, bioacoustic, and environmental domains using zero-shot and specialized encoders. Results show strong domain-specific performance alongside promising generalist capabilities.
- Classification experiments span human speech intent, environmental sounds, and bioacoustic signals.
- Gemini embeddings achieve strong zero-shot intent-classification performance across many Speech-MASSIVE languages without task-specific training.
- Perch achieves 0.66 mAP on NBP, 0.54 mAP on HSN, and 0.53 mAP on POW in BirdSet bioacoustics.
- CLAP achieves 0.43 mAP on FSD50K environmental sounds in a zero-shot setting.
F Transcription
The transcription evaluation measures Whisper Large v3 across 26 SVQ locales and four acoustic environments. Results reveal large language disparities and environment-specific degradation, leaving substantial headroom for universal ASR.
- Whisper Large v3 is evaluated across 26 SVQ locales and clean, background-speech, traffic-noise, and media-noise conditions using WER and SER.
- High-resource English and Russian achieve low WERs, while Malayalam and Bengali exhibit extremely high error rates.
- The language disparities indicate substantial remaining headroom for truly universal ASR capabilities.
- Clean audio consistently yields the lowest errors, but traffic noise or background speech causes greater degradation depending on the locale.
G Segmentation
The segmentation cascade transcribes audio, selects salient transcript terms, and localizes them temporally. Its performance varies sharply by locale, with ASR quality forming a central bottleneck alongside temporal alignment challenges.
- The cascade uses Whisper Large v3 transcription, Wikipedia-derived IDF scores to select three salient terms, and word-level timestamps for localization.
- NDCG reaches approximately 0.80 for English locales but is near zero for Malayalam and Bengali.
- Content Accuracy is often the primary bottleneck because ASR errors prevent correct salient-term identification and drive strict accuracy to zero.
- Korean illustrates a separate timing problem: content can be identified while temporal alignment remains difficult, reducing strict overall accuracy.
- WER and NDCG show a clear negative correlation, confirming the cascade’s dependence on ASR quality.
H Clustering
The clustering evaluation tests unsupervised discovery of acoustic structure across bioacoustic, environmental, and speech domains. Specialized encoders excel in matching domains, while generalist and simple baselines can remain highly competitive.
- Clustering uses V-Measure to assess homogeneity and completeness across bioacoustic, environmental, and human-speech domains.
- Speech clustering evaluates Speaker ID, Gender, and Age across 26 SVQ locales using self-supervised models and the Spectrogram baseline.
- Perch dominates bioacoustic clustering, exceeding 0.5 V-Measure on complex BirdSet soundscapes such as NBP.
- CLAP acts as a strong generalist, reaching 0.68 V-Measure on FSD50K while remaining respectable on BirdSet.
- A raw Spectrogram baseline is highly competitive with HuBERT and Wav2Vec2 for speaker clustering, especially Speaker ID.