Source-linked AI summary

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey

Chih-Kai Yang, Neo S. Ho, Hung-yi Lee

arXiv:2505.15957v4eess.AScs.AIcs.CLcs.SD

TL;DR

LALM benchmarks have proliferated without a structured organization, limiting systematic evaluation of their broad auditory capabilities. This survey synthesizes the evaluation literature into a four-part taxonomy, identifies challenges and future directions, and concludes that substantial gaps remain across auditory processing and knowledge domains.

  • Problem

    Existing LALM evaluations are fragmented, while current benchmarks reveal gaps across auditory processing, knowledge, dialogue, and responsible deployment.

  • Method

    The paper conducts a comprehensive survey and organizes LALM evaluation frameworks into four objective-based categories.

  • Results

    The survey provides a structured taxonomy and identifies persistent challenges across LALM evaluation, including limited robustness and inconsistent knowledge performance.

  • Takeaways & Limitations

    The taxonomy offers guidelines for selecting benchmarks while highlighting the need for broader and more robust LALM evaluation.

  • Takeaways & Limitations

    The taxonomy is based on existing frameworks and benchmarks, does not cover all real-world auditory tasks, and emphasizes advanced benchmarks over traditional subjective assessments.

Abstract

from arXiv · show

With advancements in large audio-language models (LALMs), which enhance large language models (LLMs) with auditory capabilities, these models are expected to demonstrate universal proficiency across various auditory tasks. While numerous benchmarks have emerged to assess LALMs' performance, they remain fragmented and lack a structured taxonomy. To bridge this gap, we conduct a comprehensive survey and propose a systematic taxonomy for LALM evaluations, categorizing them into four dimensions based on their objectives: (1) General Auditory Awareness and Processing, (2) Knowledge and Reasoning, (3) Dialogue-oriented Ability, and (4) Fairness, Safety, and Trustworthiness. We provide detailed overviews within each category and highlight challenges in this field, offering insights into promising future directions. To the best of our knowledge, this is the first survey specifically focused on the evaluations of LALMs, providing clear guidelines for the community. We will release the collection of the surveyed papers and actively maintain it to support ongoing advancements in the field.

1 Introduction

LALMs extend language models with auditory inputs and are expected to support increasingly complex auditory, reasoning, and interactive capabilities. Because existing benchmarks are fragmented, this survey organizes LALM evaluation and identifies challenges and future directions.

  • LALMs process auditory and/or textual inputs, including speech, audio, and music, while supporting auditory processing, multimodal reasoning, and interaction.
  • Evaluation expectations have expanded from speech recognition to audiogrounded reasoning and interactive dialogue across diverse input and output modalities.
  • Existing LALM benchmarks remain fragmented, making it difficult to select suitable evaluations and track field-wide progress.
  • The survey introduces a taxonomy covering four evaluation categories: auditory processing, knowledge and reasoning, dialogue-oriented ability, and fairness, safety, and trustworthiness.
  • The authors present the first comprehensive survey of LALM evaluations, propose structured guidelines, and identify challenges and future directions.

2 Taxonomy of Evaluation Frameworks for Large Audio-Language Models

The survey organizes LALM evaluation frameworks by objective into four categories spanning auditory processing, knowledge and reasoning, dialogue, and responsible deployment. Because benchmarks can be multidimensional, some appear in multiple categories.

  • The taxonomy groups evaluations into General Auditory Awareness and Processing, Knowledge and Reasoning, Dialogue-oriented Ability, and Fairness, Safety and Trustworthiness.
  • General Auditory Awareness and Processing: General Auditory Awareness and Processing covers foundational capabilities such as speech recognition and audio captioning.
  • Knowledge and Reasoning: Knowledge and Reasoning evaluates knowledge acquisition and advanced reasoning skills.
  • Dialogue-oriented Ability: Dialogue-oriented Ability addresses natural conversation through affective and contextual interaction, dialogue management, and instruction following.
  • Fairness, Safety and Trustworthiness: Fairness, Safety and Trustworthiness examines bias, toxicity, and reliability for ethical and safe deployment.
  • Benchmarks may span multiple categories when they contain separate tasks or require several integrated capabilities.

3 General Auditory Awareness and Processing

This section evaluates LALMs’ awareness of acoustic and paralinguistic cues alongside foundational speech, audio, and music processing. The surveyed benchmarks expose persistent gaps in fine-grained perception and universally robust performance.

  • 3.1 Auditory Awareness: LALMs are evaluated on acoustic cues such as emotion, prosody, environmental sounds, speaker changes, and mismatches between tone and semantic content.
  • 3.1 Auditory Awareness: Speech-to-speech paraphrasing and translation benchmarks test whether models preserve fine-grained prosodic emphasis.
  • 3.1 Auditory Awareness: Current benchmarks reveal substantial gaps in fine-grained auditory awareness and modeling of subtle acoustic and paralinguistic information.
  • 3.2 Auditory Processing: Foundational auditory processing benchmarks cover speech recognition, audio classification, and music analysis across speech, audio, and music modalities.
  • 3.2 Auditory Processing: 180 tasks make Dynamic-SUPERB Phase-2 the largest evaluation suite described for LALMs’ general processing abilities.
  • 3.2 Auditory Processing: Task-specific metrics include word error rate and BLEU, while LLM-as-a-judge supports scalable evaluation of open-ended responses.
  • 3.2 Auditory Processing: Despite progress in some areas, current LALMs lack universally robust performance across auditory-processing tasks.

4 Knowledge and Reasoning

Knowledge and reasoning evaluations span linguistic knowledge, world knowledge, content-based reasoning, and acoustic-based reasoning. Across these areas, benchmarks reveal limited expertise, domain inconsistency, and difficulty integrating auditory cues with broader knowledge.

  • Knowledge and reasoning evaluations cover linguistic knowledge, world knowledge assessment, and reasoning as complementary capabilities.
  • 4.1 Linguistic Knowledge: Linguistic knowledge benchmarks test lexical knowledge, syntax, and semantic coherence through paired spoken samples.
  • 4.2 World Knowledge Assessment: World knowledge benchmarks assess auditory expertise, including music and medical sound diagnosis, alongside commonsense and factual knowledge.
  • 4.2 World Knowledge Assessment: LALMs show limited auditory expertise and inconsistent performance across domains, often declining outside specialized areas.
  • 4.3 Reasoning: Content-based reasoning evaluates semantic understanding, while acoustic-based reasoning uses speaker traits, environmental sounds, and other acoustic features.
  • 4.3.1 Content-based Reasoning: Current models struggle with content-based reasoning and show instability across speaking styles, even with chain-of-thought prompting.
  • 4.3.2 Acoustic-based Reasoning: Cross-auditory reasoning benchmarks find that LALMs often neglect non-speech cues when combining speech and environmental sounds.
  • 4.3.2 Acoustic-based Reasoning: Multi-hop evaluations show difficulty combining auditory information with stored knowledge even when both are separately extracted and known.

5 Dialogue-oriented Ability

Dialogue-oriented evaluations assess LALMs’ conversational naturalness and controllability through affective and contextual interaction, full-duplex dialogue management, and instruction following. Existing results identify weaknesses in dynamic spoken interaction and instruction adherence.

  • Scope: Dialogue-oriented ability targets affective and contextual interaction, fluent dialogue management, and precise instruction following as integrative skills for natural, controllable human-AI interaction.
  • Conversational Ability: Affective and contextual evaluations commonly use half-duplex, turn-by-turn conversations that test responses to emotional tone, speaking style, speaker traits, and underspecified intent.
  • Conversational Ability: Full-duplex evaluations test real-time turn-taking, backchanneling, interruptions, overlaps, and related timing behaviors in dynamic dialogue.
  • Conversational Ability: Automatic metrics in Talking Turns and Full-Duplex-Bench include human-dialogue reference models and response latency, but heuristic-based evaluation may be inaccurate.
  • Conversational Ability: LALMs struggle with full-duplex management, particularly interruptions and seamless turn transitions, revealing limitations in dynamic spoken interaction.
  • Instruction Following: Instruction-following evaluations modify existing benchmarks, convert text benchmarks into speech, or create dedicated datasets, and measure constraints such as length, format, action, style, and content.
  • Instruction Following: LALMs show significant instruction-following gaps relative to their LLM backbones, indicating catastrophic forgetting during adaptation to auditory modalities.

6 Fairness, Safety, and Trustworthiness

Fairness, safety, and trustworthiness evaluations examine social bias, harmful outputs, and auditory hallucinations in LALMs. Existing findings show biases, spoken-input safety weaknesses, and object-identification hallucinations, while coverage remains incomplete.

  • Fairness and safety evaluation addresses societal bias, harmful content, misinformation, and reliable deployment of LALMs.
  • Fairness and Bias: Content-triggered gender bias is evaluated across translation, coreference, sentence continuation, and question answering tasks.
  • Fairness and Bias: Acoustic-triggered bias benchmarks remove explicit demographic indicators from content and vary synthesized speakers’ gender and age.
  • Fairness and Bias: Social biases may be inherited from training data or LLM backbones, while current benchmarks cannot cover all societal factors.
  • Safety: Safety evaluations use malicious spoken queries and adversarial manipulations including fictional scenarios, silence, noise, accents, and audio edits.
  • Safety: LALMs often accept malicious spoken inputs that they refuse textually, show safety degradation relative to LLM backbones, and can be bypassed by jailbreaking methods.
  • Hallucination: Auditory-induced hallucination evaluations test whether models falsely identify absent objects or events through discriminative and generative tasks.
  • Hallucination: Overrepresented objects and frequent object-event co-occurrences in training data can increase false predictions of absent auditory content.

7 Challenges and Future Directions

The survey identifies evaluation gaps involving contamination, linguistic and communicative diversity, auditory-specific safety, and personalization. It calls for broader data coverage, contamination mitigation, and joint assessment of safety and helpfulness.

  • Data Leakage and Contamination: Reliance on existing auditory corpora raises data-leakage concerns because models may have encountered benchmark data during training.
  • Data Leakage and Contamination: Custom data collection and methods to detect and mitigate contamination are needed for more reliable LALM evaluation.
  • Inclusive Evaluation Across Linguistic, Cultural, and Communication Diversity: Many benchmarks underrepresent low-resource languages, code-switching, cultural variation, and speech disorders, limiting coverage of human communication diversity.
  • Inclusive Evaluation Across Linguistic, Cultural, and Communication Diversity: Future evaluations should incorporate linguistic, cultural, and communicative diversity to support fair and broadly applicable LALMs.
  • Auditory-Specific Safety: Auditory-specific safety should assess tone, emotion, voice quality, noise, and user comfort alongside content harmlessness.
  • Unified Evaluation of Harmlessness and Helpfulness: Post-training that improves harmlessness can reduce helpfulness, causing refusals even when queries present no safety or privacy issue.
  • Unified Evaluation of Harmlessness and Helpfulness: Because harmlessness benchmarks rarely include helpfulness, joint evaluation is needed to understand and balance their trade-offs.
  • Personalization: LALM personalization remains underdeveloped because models must adapt to user knowledge, voices, speaking habits, and preferred response styles.

8 Conclusion

The survey reviews LALM evaluation frameworks and proposes a taxonomy organized around four research areas. It synthesizes challenges and future directions to provide guidance for advancing evaluation.

  • The survey proposes a taxonomy organizing LALM evaluation progress into four research areas reflecting diverse model capabilities.
  • It highlights data contamination, inclusivity, auditory-specific safety, and personalization as challenges and future directions.

Limitations

The survey’s taxonomy is bounded by existing evaluation frameworks and benchmarks, and it emphasizes current specialized benchmarks over traditional subjective assessments of generated-audio quality.

  • Scope: The taxonomy does not cover all possible real-world auditory tasks because it is based on existing evaluation frameworks and benchmarks.Its scope may need updating as new LALM capabilities and applications emerge.
  • Evaluation coverage: The survey primarily focuses on current benchmarks for evaluating LALM performance across various aspects.This focus defines the paper’s coverage of evaluation evidence.
  • Evaluation coverage: Traditional methods such as Mean Opinion Score for speech-generation quality receive less emphasis because the survey targets advanced and specialized benchmarks.The authors note that subjective assessments remain valuable in certain applications but fall outside the paper’s scope.

A Detailed Categorization of the Surveyed Papers

The surveyed resources cover auditory tasks and their datasets, while the current LALM evaluation literature predominantly emphasizes auditory processing. The survey organizes benchmark features and metrics to guide benchmark selection for different use cases.

  • Evaluation focus: Current LALM evaluations predominantly center on auditory processing tasks.The survey cautions that auditory processing should not be the sole consideration for real-world model evaluation.
  • Dataset coverage: Table 1 covers commonly used datasets across audio, speech, and music processing tasks.These resources are widely adopted in academic and industrial research.
  • Evaluation focus: A broader evaluation scope is needed to understand LALMs’ potential and shortcomings more fully.The survey presents diversity beyond auditory processing as important for real-world applications.
  • Benchmark selection: The survey summarizes benchmark features and evaluation metrics to help researchers select resources suited to their use cases.These details are provided in Tables 5 and 6.

C Examples of General Auditory Processing Tasks and Resources

The survey catalogs auditory-processing resources and evaluation practices while highlighting dataset reuse, synthetic-data roles, dialogue dynamics, and end-to-end versus cascaded comparisons.

  • Auditory processing resources: Table 1 lists representative auditory processing tasks and associated resources, while these foundational tasks can be adapted for LALM evaluation.
  • Dataset reuse: Frequent reuse of certain corpora reinforces data-contamination concerns, particularly when web-scraped data enter training without rigorous filtering.
  • Real and synthetic data: Synthetic audio supports controlled conditions that are difficult to obtain in real settings and consistent verbalization of task instructions or dialogues.
  • Dialogue dynamics: Full-duplex evaluation covers turn-taking, backchanneling, speaker interruptions, and overlaps in real-time dynamic dialogue.
  • E2E versus cascaded systems: Cascaded systems generally perform better on content- and semantics-based tasks, whereas E2E LALMs show advantages on benchmarks requiring fine-grained auditory perception.
  • Evaluation metrics: Common evaluation metrics include MOS for subjective auditory quality, CER and WER for speech-text consistency, ROUGE and BERTScore for semantic overlap, and LLM-as-a-judge for flexible criteria.
Loading 2505.15957v4…