Source-linked AI summary
How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
Ke-Han Lu, Szu-Wei Fu, Chao-Han Huck Yang, Zhehuai Chen, Sung-Feng Huang, Chih-Kai Yang, Yi-Cheng Lin, Chi-Yuan Hsiao, Wenze Ren, En-Pei Hu, Yu-Han Huang, An-Yu Cheng, Cheng-Han Chiang, Yu Tsao, Yu-Chiang Frank Wang, Hung-yi Lee
TL;DR
The paper asks how much auditory knowledge LLMs acquire from text-only pre-training and how that knowledge relates to LALM performance. It compares direct AKB-2000 probing, cascade reasoning over audio captions, and audio-grounded fine-tuning, finding substantial family variation and strong text-only/audio-grounded correlation. It further identifies phonological reasoning as a limitation of text-only pre-training and treats backbone choice as a first-order LALM design decision.
Problem
How much auditory knowledge LLMs encode through text-only pre-training, and how that knowledge translates to multimodal performance, remains unclear.
Method
The study evaluates LLMs with AKB-2000, cascade reasoning over captioned audio, and DeSTA-based audio-grounded LALM fine-tuning.
Results
Auditory knowledge varies substantially across model families, and text-only performance is strongly correlated with audio-grounded performance.
Takeaways & Limitations
AKB-2000 and cascade evaluation provide lightweight proxies for selecting LLM backbones, making backbone choice a first-order LALM design decision.
Takeaways & Limitations
LLMs consistently struggle with phonological tasks, indicating that addressing this limitation may require phonology-aware training beyond standard text corpora.
Abstract
from arXiv · showhide
Large language models (LLMs) have been widely used as knowledge backbones of Large Audio Language Models (LALMs), yet how much auditory knowledge they encode through text-only pre-training and how this affects downstream performance remains unclear. We study this gap by comparing different LLMs under two text-only and one audio-grounded setting: (1) direct probing on AKB-2000, a curated benchmark testing the breadth and depth of auditory knowledge; (2) cascade evaluation, where LLMs reason over text descriptions from an audio captioner; and (3) audio-grounded evaluation, where each LLM is fine-tuned into a Large Audio Language Model (LALM) with an audio encoder. Our findings reveal that auditory knowledge varies substantially across families, and text-only results are strongly correlated with audio performance. Our work provides empirical grounding for a comprehensive understanding of LLMs in audio research.
1. Introduction
This work examines how much auditory knowledge text-only LLMs encode and how that knowledge relates to their use as LALM backbones. It evaluates direct auditory knowledge, cascade reasoning, and audio-grounded LALM performance across diverse models.
- Existing LALMs rarely justify their LLM backbone choice based on the backbone’s auditory knowledge.Different training corpora and recipes may produce substantially different auditory understanding.
- The study evaluates direct AKB-2000 probing, cascade reasoning over audio captions, and audio-grounded LALMs.AKB-2000 spans six audio-knowledge categories, while cascade evaluation feeds captioned audio descriptions to text-only LLMs.
- Auditory knowledge varies substantially across model families, with Qwen consistently outperforming Llama in most evaluated settings.The evaluation covers 12 open-weight LLMs from four families and five proprietary baselines.
- Over 10% absolute performance difference can result from changing only the base LLM under an identical LALM training recipe.This isolates the backbone choice as a substantial contributor to downstream performance.
- Text-only evaluation correlates strongly with audio-grounded evaluation, while phonological tasks remain a consistent weakness.A simple caption-based cascade can match or surpass several state-of-the-art end-to-end LALMs.
2. Related Work
Audio understanding systems use either end-to-end LALMs or modular cascades, but both rely on an LLM’s auditory knowledge. Existing system-level benchmarks confound encoder, data, and backbone effects, motivating broader direct and cross-setting evaluation.
- End-to-end LALMs connect audio encoders to LLM backbones, whereas modular systems pass audio-to-text representations to text-only LLMs.Cascades improve interpretability and avoid multimodal training costs but depend on caption granularity.
- Both paradigms assume that the underlying LLM has sufficient auditory knowledge for downstream reasoning.How much text-only pre-training encodes this knowledge and how it transfers to multimodal performance remains open.
- Holistic audio benchmarks conflate audio encoding quality, training-data coverage, and the LLM’s internal knowledge.Observed performance gaps therefore do not directly identify which component is responsible.
- Prior auditory-knowledge studies mainly probe basic sound events and coarse acoustic properties.Representation probing, retrieval or generation augmentation, and direct question answering leave broader domains underexamined.
- This work broadens probing, tests reasoning over captioned real-audio questions, and links text-based auditory knowledge to audio-grounded capability.The three dimensions jointly address breadth, application, and transfer across model families.
3. Method
The method combines a broad direct knowledge benchmark, cascade reasoning over generated audio descriptions, and DeSTA-based audio-grounded fine-tuning. These evaluations separately probe auditory knowledge, its application, and its transfer to audio input.
- AKB-2000 is a 2,000-question multiple-choice benchmark testing factual and common-sense auditory knowledge.Its taxonomy covers six categories and 48 subcategories spanning major audio-research domains.
- AKB-2000 questions are generated with proprietary LLM assistance and retained only after two audio-background annotators agree.Annotators assess correctness, clarity, and distractor plausibility.
- Cascade evaluation asks LLMs to answer MMAU and MMAR questions using detailed textual descriptions generated from audio.Descriptions include acoustic properties, sound sources, temporal structure, spoken content, and speaking style.
- AKB-2000 measures encoded auditory knowledge, whereas cascade evaluation measures applying that knowledge to real-audio questions represented as text.The two text-only evaluations therefore serve complementary roles.
- Audio-grounded evaluation fine-tunes each LLM into an LALM with an audio encoder using DeSTA self-distillation.Textual metadata first elicits targets from the LLM, then raw audio replaces that metadata as input.
- The DeSTA setup lets the backbone influence both generated supervision and the resulting model’s retained knowledge and generation style.This creates data-side and model-side pathways through which the backbone shapes the LALM.
4. Experimental Setup
The experiments compare diverse LLM families and scales, then fine-tune selected backbones into LALMs under matched conditions. Correlation analysis and standardized benchmarks assess relationships between text-only and audio-grounded performance.
- The study evaluates 12 open-weight instruction-tuned LLMs from Qwen, Llama, Phi, and OLMo across 4B–14B parameter scales.Multiple generations are included for Qwen and Llama to examine variation across model families and generations.
- Eight open-weight LLMs are fine-tuned for audio-grounded evaluation to enable cross-family comparisons at matched parameter scales.The selected models include Qwen, Llama, Phi, and OLMo variants.
- All LALMs use the DeSTA framework and identical training data and training recipe, with only generated responses differing across backbones.The source data includes 404 hours of speech, 329 hours of sound events, and 144 hours of music.
- The audio-grounded architecture uses Whisper-large-v3, a six-layer Q-Former, and frozen audio-encoder and LLM parameters.Only the modality connector is trainable, making the setup a stricter test of pre-existing auditory knowledge.
- Figure 2 summarizes Pearson correlations across five text-only and audio-grounded metrics in a heatmap.A white line separates the text-only block from the audio-grounded block.
- AKB-2000 contains 2,000 four-option questions with uniform answer distribution, giving random chance a 25% baseline.MMAU test-mini and MMAR each contribute 1,000 questions, with speech, sound, and music categories reported.
5. Results
Results show substantial variation in auditory knowledge across LLM families, with text-only evaluations closely tracking audio-grounded performance. Phonological reasoning remains a systematic weakness, while backbone choice and audio-text processing bottlenecks strongly affect LALM outcomes.
- Overall trends: Text-only model rankings are highly consistent across AKB-2000 and cascade evaluation, with within-text correlations reaching 0.94.Correlations between text-only and audio-grounded metrics are also strong, ranging from r = 0.71 to r = 0.82.
- Auditory knowledge benchmark: Qwen2.5-7B scores 80.70% on AKB-2000 versus 68.10% for Llama-3.1-8B at comparable scale.Proprietary models exceed 94% on AKB-2000, while open-weight scores range from 45.90% to 86.35%.
- Auditory knowledge benchmark: Phonetic accuracy trails other AKB-2000 categories by 10–15 percentage points across every model family.The hardest subcategories largely require pronunciation, prosody, or phonological reasoning that is not directly observable in written text.
- Cascade evaluation: Captioner quality remains a critical cascade bottleneck: Gemini-2.5-Pro captions achieve 70.90% on MMAU and 71.80% on MMAR, but top-tier systems plateau near 70%.Caption recognition errors propagate directly to the downstream LLM.
- Audio-grounded evaluation: With identical components except for the backbone, fine-tuned LALMs differ by over 10 points on MMAU and 8 points on MMAR.Qwen2.5-7B and Qwen3-14B reach 66.6% and 66.2% on MMAU, respectively, matching or surpassing DeSTA2.5-Audio at 66.0%.
- Audio-grounded evaluation: Cascade systems can match or surpass end-to-end LALMs on MMAR, including 62.0% for Qwen3-8B versus 58.6% for Audio Flamingo 3 and 56.7% for Qwen2.5-Omni.The comparison points to an audio-text alignment bottleneck in end-to-end systems, where audio encoders may lose fine-grained details preserved by captioners.
6. Conclusion
The paper’s holistic evaluation shows that auditory knowledge in LLM backbones shapes LALM performance, making backbone choice a first-order design decision. It also compares cascade and audio-grounded performance across Sound, Music, and Speech domains.
- Auditory knowledge in LLM backbones shapes every stage of audio understanding and makes backbone selection a first-order LALM design decision.
- Cascade and audio-grounded accuracy are compared for 8 fine-tuned LALMs across Sound, Music, and Speech domains.
- The cascade pipeline performs comparably to, or surpasses, recent state-of-the-art LALMs, suggesting that many end-to-end systems do not fully use their LLM backbones’ capabilities.
7. Generative AI Use Disclosure
Generative AI tools supported manuscript proofreading and candidate-question curation for AKB-2000, while humans verified the included questions and authored the scientific work.
- Generative AI tools assisted with manuscript proofreading and fluency improvement, and helped curate candidate AKB-2000 questions under human-authored guidelines.
- Human annotators verified all AKB-2000 questions before inclusion, while the authors retained responsibility for scientific content, design, analysis, and conclusions.