Source-linked AI summary
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook
Kaiwen Luo, Zhenhong Zhou, Leyan Wang, Liang Lin, Tianyu Shao, Yuanhe Zhang, Yang Xiao, Yuxuan Li, Miao Yu, Kailin Lyu, Jiaming Zhang, Li Sun, Songze Li, Yueming Wu, Ting Dang, Xiaojun Jia, Dongrui Liu, Kai Li, Rohan Kumar Das, Siyuan Liang, Xinfeng Li, Qiankun Li, Jing Chen, Xingjun Ma, Kun Wang, Junhao Dong, Deqing Zou, Yu Cheng, Xia Hu, Zhigang Zeng, Sen Su, Yang Liu, Yu-Gang Jiang, Philip S. Yu, Yew-Soon Ong
TL;DR
LALMs’ integration of continuous audio with language creates complex safety and alignment challenges, while systematic trustworthiness frameworks remain underdeveloped. This survey examines their mechanisms and risks across six pillars, finding that offensive research is significantly more advanced than defensive mechanisms.
Problem
LALMs’ continuous acoustic signals introduce complex safety and alignment challenges, while robust defenses lack clear safe boundaries and standardized evaluation.
Method
The survey examines LALM structures and alignment techniques, then analyzes trustworthiness across hallucination, robustness, safety, privacy, fairness, and authentication.
Results
The analysis reveals a significant developmental asymmetry: offensive research has advanced substantially, while defensive mechanisms remain limited and reactive.
Takeaways & Limitations
The survey advocates layered defenses, causal auditory world modeling, and intrinsic representation engineering for more trustworthy audio intelligence.
Takeaways & Limitations
Existing defenses focus primarily on jailbreak prevention, with little coverage of backdoors, bias, or multimodal privacy risks.
Abstract
from arXiv · showhide
Advances in Large Language Models (LLMs) have paved the way for Multimodal Large Language Models (MLLMs). Among these, Large Audio Language Models (LALMs) are essential for realizing universal auditory intelligence. Despite their remarkable performance, the escalation of LALMs' capabilities has significantly outpaced the development of systemic frameworks to ensure their trustworthiness. This survey provides a comprehensive investigation into the endogenous mechanisms of LALMs, detailing the architectural innovations and alignment algorithms that facilitate emergent reasoning. Specifically, we analyze how the transition to unified end-to-end frameworks and the integration of continuous acoustic signals expand the attack surface. To rigorously evaluate the risks within these paradigms, we establish a comprehensive taxonomy of trustworthiness, categorizing critical vulnerabilities such as cross-modal jailbreaking, latent acoustic backdoors, and biometric privacy leakage. We review the state-of-the-art LALMs through six analytical pillars: hallucination, robustness, safety, privacy, fairness, and authentication. The pronounced imbalance between a mature offensive landscape and underdeveloped defenses highlights persistent trustworthiness gaps and multidimensional risks in audio-centric intelligence. Finally, we propose a roadmap advocating for ``Defense-in-Depth'' architectures, causal auditory world modeling, and intrinsic representation engineering to support the development of more reliable and trustworthy audio intelligence. Our project has been uploaded to GitHub https://github.com/Kwwwww74/Awesome-Trustworthy-AudioLLMs.
Abstract
The survey positions Large Audio Language Models as essential to universal auditory intelligence while highlighting that their rapid capability growth has outpaced systemic trustworthiness frameworks. It investigates LALMs’ internal mechanisms and architectural developments.
- LALMs are presented as essential for realizing universal auditory intelligence.
- Their capability growth has significantly outpaced the development of systemic frameworks for ensuring trustworthiness.
- The survey investigates LALMs’ endogenous mechanisms and details their architectural innovations.
1 Introduction
This survey examines LALMs’ internal mechanisms, alignment techniques, and expanding trustworthiness risks as audio becomes integrated into unified multimodal systems. It organizes vulnerabilities and evaluations systematically, identifies defensive shortcomings, and proposes a future framework for intrinsically trustworthy audio intelligence.
- Motivation: The integration of language and audio creates complex safety and alignment challenges beyond those encountered by text-focused language models.LALMs introduce an intricate risk landscape through the added audio modality.
- Research gap: Existing surveys often treat safety and ethics as peripheral and lack a systematic taxonomy connecting security threats with safety mechanisms.Evaluation-focused literature provides behavioral assessment frameworks but does not systematically organize underlying threats and defenses.
- Survey contributions: The survey analyzes LALM internal structures and alignment techniques that support emergent logical reasoning in unified auditory-intelligence models.This provides a technical foundation for understanding LALM evolution.
- Survey contributions: It classifies critical trustworthiness vulnerabilities, including cross-modal acoustic jailbreaks, latent acoustic backdoors, and biometric privacy leakage.The review evaluates leading models across hallucination, robustness, safety, privacy, fairness, and authentication.
- Future framework: The survey finds offensive research substantially ahead of limited, reactive defenses and advocates layered defense, causal auditory world modeling, and intrinsic representation engineering.These directions aim to achieve intrinsically trustworthy audio intelligence.
2 Endogenous Mechanisms of LALMs
LALMs process audio through an acoustic encoder, alignment projector, and LLM backbone, while evolving representations, optimization, and reasoning mechanisms support multimodal understanding and emergent auditory cognition. The section also identifies causal auditory world modeling and intrinsic representation engineering as directions for extending these capabilities toward trustworthy systems.
- Architectural Design: LALMs use a three-component pipeline comprising an acoustic encoder, alignment projector, and LLM backbone to translate raw acoustic signals into semantic representations.The encoder supports sensory perception, the projector bridges modalities, and the backbone supplies reasoning capacity shaped partly by auditory knowledge from text pre-training.
- Optimization and Alignment: Alignment and optimization methods improve cross-modal learning by addressing modality bias, gradient conflicts, continuous-stream overhead, and long-context dependency capture.These methods include Mixture of Experts adapters, segmentwise pruning, attention rebalancing, audio contribution-aware post-training, extended-context mechanisms, and temporal-gap bridging.
- Emergent Reasoning: Audio Chain-of-Thought and audio-interleaved frameworks generate intermediate reasoning trajectories and embed reasoning steps within multimodal processing to support complex auditory deduction.The section also considers hidden-state nudging and maintaining reasoning efficiency while listening to continuous audio.
- Emergent Reasoning: Reinforcement learning and process-oriented rewards incentivize logically valid multi-step deductions, including consistent reasoning in affect-rich tasks through emotion-rule-based reinforcement learning.These approaches are presented as primary drivers for scalable and consistent reasoning behavior.
- Future Directions: Because multimodal advances expand the attack surface, intrinsic representation engineering is proposed to ground emergent capabilities in trustworthiness.This direction connects endogenous capability development with the need for trustworthy internal representations.
3 Taxonomy of Trustworthiness
The taxonomy organizes LALM trustworthiness into six analytical pillars: hallucination, robustness, safety, privacy, fairness, and authentication. It distinguishes failures of auditory evidence use, behavioral stability, policy compliance, information exposure, equitable performance, and audio identity or integrity verification.
- Hallucination: Hallucination concerns unsupported outputs and includes perceptual fabrication, grounding or attribution failure, and modality neglect.These categories distinguish failures of auditory perception from failures of cross-modal grounding and evidence use, requiring different diagnostic protocols.
- Robustness: Robustness measures behavioral stability across environmental, interactional, temporal, and adversarial variations that preserve task-relevant semantics.The taxonomy separates benign distribution shifts from intentional attacks and distinguishes input-level, long-context, and interaction-level reliability.
- Authentication: Authentication evaluates speaker identity, audio genuineness, and manipulation localization or attribution, which require separate assessment.Successful speaker recognition does not necessarily imply reliable spoof detection or manipulation localization.
- Privacy: Privacy addresses inference, retention, or disclosure of linguistic, speaker, paralinguistic, and environmental information through identity, attribute, contextual, and interactional leakage.Audio creates a broad privacy surface because it jointly encodes these information types.
- Fairness: Fairness concerns equitable behavior across population groups, languages, acoustic conditions, and interaction formats, including demographic, linguistic, paralinguistic, and structural biases.Intersecting biases can make aggregate performance conceal substantial disparities among demographic, linguistic, and acoustic subgroups.
- Safety: Safety evaluates whether behavior remains within policy boundaries under benign or malicious interaction across semantic, paralinguistic, linguistic, signal-level, and training-time attacks.Evaluation should distinguish under-refusal from over-refusal rather than treating stronger refusal behavior as invariably safer.
4 Safety Challenges in LALMs
LALMs expand the attack surface beyond text through continuous audio, paralinguistic cues, and multimodal interactions, creating risks spanning manipulation, jailbreaking, backdoors, privacy, and bias. Existing defenses remain immature and largely jailbreak-focused, motivating audio-aware alignment and broader evaluation frameworks.
- Adversarial manipulation: Continuous audio enables imperceptible perturbations and naturally occurring noise to hijack latent representations without changing human-perceived semantics.Text-only safety paradigms are inadequate because acoustic realizations can obscure malicious intent through benign paralinguistic patterns.
- Jailbreaking: Cross-modal jailbreaks exploit non-semantic speech attributes and acoustic perturbations to bypass text-centric safety filters and degrade safety alignment.Jailbreak-AudioBench, JALMBench, HIN, and WhisperInject expose attack surfaces and adversarial audio frameworks not covered by textual alignment.
- Backdoors: Training-time data poisoning can associate imperceptible frequency patterns, background noises, or subtle acoustic features with attacker-chosen behaviors.These backdoor triggers compromise LALM integrity during training rather than exploiting only inference-time behavior.
- Privacy: LALMs can leak speaker gender, age, health status, identity, location, and private background conversations from acoustic mixtures.HearSay, audio geo-localization studies, and SH-Bench demonstrate risks to users and non-consenting third parties.
- Bias: Voice-derived accent, dialect, age, and other demographic cues can drive discriminatory behavior, including biased clinical decisions in healthcare.MedVoiceBias shows that demographic information inferred from voice may influence decisions instead of task-relevant medical evidence.
- Defensive gaps and directions: Defenses remain rudimentary and primarily jailbreak-focused, while continuous audio complicates safe-boundary definition and standardized evaluation across threats.Audio-aware multimodal preference signals, endogenous representation optimization, and complementary LALM-based detectors are proposed, but detectors remain costly and vulnerable to subtle artifacts.
5 Evaluation · 5.1 Fidelity and Grounding · 5.2 Stability and Robustness
The evaluation framework organizes trustworthy LALM assessment around fidelity, stability, and alignment, emphasizing whether outputs remain grounded in acoustic evidence and consistent under perturbations. Across these dimensions, benchmarks expose cross-modal hallucination, text-driven reasoning, long-context degradation, and fragile interaction behavior.
- 5 Evaluation: Trustworthy LALM evaluation is organized into fidelity, stability, and alignment, measuring acoustic grounding and behavioral consistency under varied conditions.The framework treats fidelity and grounding as cognitive trust, while stability and robustness assess consistency across temporal, instructional, acoustic, and conversational changes.
- 5.1 Fidelity and Grounding: Hallucination benchmarks show that LALMs frequently fail to ground language outputs in acoustic reality, including fabricated events, misinterpreted properties, and linguistic-prior dependence.HalluAudio evaluates speech, environmental sound, and music with over 5K human-verified QA pairs and adversarial or mixed-audio conditions.
- 5.1.1 From Classification to Disentanglement: Fine-grained benchmarks reveal that models can recognize acoustic or musical content while failing to localize the evidence supporting their answers, especially in overlapping or temporally complex scenes.WESR separates ASR errors from event-localization failures, while MMAU, MMAU-Pro, and MusTBENCH examine disentanglement and temporal grounding.
- 5.1.2 From Textual Bias to Genuine Listening: Under conflicting text and audio, LALMs exhibit textual dominance: accuracy drops while confidence remains high, and open-ended audio captioning reaches only 63.19 F1 on BRACE-Hallucination.MCR-BENCH measures textual bias with Text Influence Rate, while S2S-Arena finds paralinguistic generation harder than understanding.
- 5.1.3 From Fact Retrieval to Personal Alignment: Higher-level evaluations show uneven cognitive profiles, weak phonological and paralinguistic understanding, and reasoning that remains largely text-driven rather than genuinely multimodal.AudioProcessBench shows that correct final answers can conceal hallucinated events, faulty grounding, or invalid inference, while speech dialogue models struggle when acoustic-prosodic cues are needed.
- 5.2.2 Interaction and Instruction Robustness: Interactional robustness is fragile: structured-output compliance can fall below 50%, option permutations can change accuracy by up to 24%, and acoustic or conversational variation disrupts instruction following.Evaluations also identify fluent speaking alongside weaker audio understanding, missed turn yields, excessive interruptions, scarce backchannels, and near-random backchannel comprehension.
- 5.2.1 Temporal Robustness in Long-Form Context: Temporal robustness is limited by long-context collapse: some ChronosAudio tasks drop by over 90% as duration increases, while attention-based processing incurs quadratic cost.AudioMarathon finds that token pruning and sparse attention can recover retrieval performance but often fail to restore high-fidelity reasoning, producing a restorative ceiling.
5.3 Safety and Alignment · 5.4 Future Horizons of LALMs’ Evaluation
LALM safety evaluation must address audio-specific jailbreaks, backdoors, privacy leakage, fairness risks, and speaker impersonation, while future evaluation should move beyond static behavioral snapshots toward causal, dynamic, intrinsic, and mechanistic verification.
- 5.3 Safety and Alignment: Audio introduces adversarial carriers, backdoor triggers, biometric identifiers, and demographic bias, expanding LALMs’ safety and alignment risks beyond text-only systems.Existing work examines defensive security against jailbreaks and acoustic backdoors alongside privacy, fairness, and authentication risks.
- 5.3.1 Jailbreaks and Acoustic Backdoors: Audio jailbreaks achieve higher success rates than text attacks, while prompt-level defenses reduce but do not eliminate vulnerabilities and can reduce utility.JALMBench compares textual and auditory jailbreaks at large scale, and Jailbreak-AudioBench further studies simple audio edits.
- 5.3.1 Jailbreaks and Acoustic Backdoors: Optimized time-stretching and fading perturbations lower refusal rates while preserving semantic transcription similarity, exposing weaknesses in clean ASR-like safety representations.AJailBench indicates that adversarial information embedded in non-standard acoustic patterns can evade current safety mechanisms.
- 5.3.1 Jailbreaks and Acoustic Backdoors: Background noise, prosody, emotion, and speaking rate can act as latent acoustic backdoors, with highly effective triggers implantable using small amounts of poisoned data.These socially natural conditions can become hidden switches for malicious outputs.
- 5.3.2 Privacy, Fairness, and Authentication: Speech models can infer sensitive attributes such as gender, socioeconomic status, and health conditions from short speech clips without consent, creating a personalization–privacy tension.HearSay reports high-accuracy recovery of sensitive information from voice.
- 5.3.2 Privacy, Fairness, and Authentication: Acoustic identity cues make bias more implicit than in text-only settings, while multilingual evaluations show sensitivity to language and option ordering.Privacy leakage can amplify fairness risks when responses condition on inferred demographic attributes.
- 5.3.2 Privacy, Fairness, and Authentication: Open-source LALMs remain vulnerable to identity-verification bypass and voice-cloning spoofing, requiring defenses against synthetic-voice impersonation alongside refusal policies.AudioTrust evaluates verification bypass and voice-cloning spoofing, while DailyTalkEdit treats anti-spoofing as holistic generation involving spoofing methods, manipulated attributes, and semantic impacts.
- 5.4 Future Horizons of LALMs’ Evaluation: Future trustworthiness evaluation should replace static behavioral snapshots with structural verification through causal auditory world modeling, agent-based dynamic red-teaming, intrinsic representation engineering, and mechanistic interpretability.The proposed shifts target counterfactual physical reasoning, adaptive attack-defense curves, residual biometric leakage, and neural-circuit-based failure detection.
6 Outlook and Conclusion
The survey frames LALMs as undergoing a structural and cognitive transformation beyond empirical performance scaling. It organizes the outlook around intrinsic mechanisms, multimodal safety, and rigorous evaluation while highlighting the expanded attack surface of unified multimodal systems.
- LALMs are transitioning from empirical performance scaling toward a structural and cognitive transformation.
- The outlook organizes future research around intrinsic mechanisms, multimodal safety, and rigorous evaluation.Figure 7 synthesizes these three dimensions into a roadmap of emerging directions.
- LALMs have evolved from task-specific cascaded systems toward unified multimodal generative frameworks with emergent reasoning abilities.Cross-modal alignment and reinforcement learning strategies contributed to these capabilities.
- These architectural and alignment advances have simultaneously introduced a complex, high-dimensional attack surface.
7 Broader Impact Statement
LALMs could improve accessibility and several human-centered applications, but their audio inputs introduce risks beyond text-only systems because they may encode sensitive personal and environmental information. The survey organizes evidence on these risks and defenses while limiting attack descriptions to support scientific comparison and risk assessment without implementation-level instructions.
- LALMs may benefit accessibility, human–computer interaction, education, healthcare, and multilingual communication.
- Audio signals can encode speaker identity, demographic attributes, emotional state, health conditions, and environmental context, expanding risks beyond text-only systems.
- The survey organizes evidence on LALM risks and defensive and evaluation strategies while presenting attacks at the level needed for scientific comparison and risk assessment.Systematizing attack methods may also make the threat landscape more accessible to malicious actors.
A Appendix · A.1 Release-Date Convention
The survey orders works by earliest publicly verifiable release date rather than cited publication year, applying this convention consistently to its tables and timeline analyses. It defines source-specific release-date rules, preserves original chronological placement despite later versions, and records verification through June 30, 2026.
- A.1 Release-Date Convention: Works are organized by their earliest publicly verifiable release date, not the bibliography’s cited publication year.This convention supports consistent chronological comparisons across the survey.
- A.1 Release-Date Convention: The release-date convention is applied consistently to Tables 1–5 and timeline analyses in Figures 1 and 5.
- A.1 Release-Date Convention: For papers and surveys, the date is the initial arXiv submission when available.
- A.1 Release-Date Convention: Without an arXiv version, the date is the earliest verifiable date from an official technical report, project page, or public repository.
- A.1 Release-Date Convention: For models, systems, and benchmarks, release dates correspond to their first public availability or earliest public release through the specified official sources.
- A.1 Release-Date Convention: Later revisions, repository updates, and peer-reviewed publications do not change chronological placement, so release and bibliographic years may differ.
- A.1 Release-Date Convention: The literature search and release records were last updated and verified on June 30, 2026.
A.2 Detailed Comparison of Large Audio Language Models
Table 4 provides a detailed comparison of audio-capable large language models from 2022 to 2026. It covers model specifications, modalities, languages, audio representations, and publicly disclosed pre-training data, while excluding image- and video-capable models.
- Comparison criteria: Table 4 compares release date, parameter scale, full-duplex capability, supported modalities, base LLM, supported languages, audio input representation, and disclosed pre-training data scale.A dash denotes information that was not publicly reported or could not be reliably verified from technical documentation.
- Scope: The comparison covers large language models with audio modality and excludes models supporting image or video modalities.The table summarizes models from 2022 to 2026.
- Notation: “Lang.” denotes language, “Input Repr.” denotes input representation, and “Contin.” denotes continuous representation.“EN” means English, “CN” means Chinese, and “Multi.” indicates support for more than two languages.
A.3 Detailed Comparison of LALM Evaluation Benchmarks
Table 5 details LALM evaluation benchmarks by their original metrics across general capabilities and trustworthy dimensions. Coverage is marked only when benchmarks provide explicit evaluation examples, annotations, or metrics for a dimension.
- Benchmark dimensions: Table 5 covers perception, reasoning, and interaction for general capabilities, alongside hallucination, privacy, authentication, safety, robustness, and fairness for trustworthiness.General dimensions are abbreviated PE, RE, and IN; trustworthy dimensions are H, P, A, S, R, and F.
- Coverage criterion: Benchmark coverage requires explicit evaluation examples, annotations, or metrics for the corresponding dimension.Mentioning a dimension only as motivation or future work does not qualify as coverage.
- Coverage criterion: The same coverage criterion is applied consistently in both the compact and detailed benchmark tables.This ensures that reported check marks reflect actual evaluation support rather than stated research motivation.