Source-linked AI summary
Auditory Illusion Benchmark for Large Audio Language Models
Hayoon Kim, Eunice Hong, Kyogu Lee
TL;DR
LALM research has largely evaluated objective audio tasks, leaving open whether these models reproduce human-like auditory misperceptions. AIB addresses this gap with a ten-illusion benchmark paired with controlled human studies. Results show that models often follow the physical signal on low-level illusions, become more human-like when linguistic or musical priors matter, and still fail to match the human perceptual profile.
Problem
Existing audio benchmarks emphasize objective signal-level tasks, leaving whether LALMs internalize human-like auditory perceptual biases insufficiently studied.
Method
AIB evaluates LALMs on ten auditory illusions across music, sound, and speech, using matched controls and controlled human listening studies for direct comparison.
Results
Illusions involving linguistic or musical priors elicited the most human-like responses, but no model matched humans; alignment was at best an ISI of 0.505.
Takeaways & Limitations
Auditory illusions provide a complementary axis for evaluating perceptual alignment and probing which human-like auditory mechanisms LALMs capture.
Takeaways & Limitations
The benchmark classifies illusions by their dominant explanatory mechanism rather than imposing a strict dichotomy.
Abstract
from arXiv · showhide
Perceptual illusions have long served as crucial probes into human cognition, revealing biases and limitations of perception. In the auditory domain, such illusions provide a unique lens for testing whether Large Audio Language Models (LALMs) replicate human perceptual tendencies. Despite their importance, most benchmarks focus on visual illusions or general audio tasks, leaving auditory illusions underexplored. To this end, we present AIB, the first auditory illusion benchmark for LALMs, covering ten representative illusions across music, sound, and speech, each annotated for the presence of knowledge-based priors. Our methodology pairs model evaluation with controlled human listening studies, enabling direct comparison of responses. Results show systematic differences: while most LALMs remain signal-faithful on low-level acoustic illusions, several exhibit more human-like responses when linguistic or musical priors are involved, although no model matches the human perceptual profile. These findings highlight the current limitations of LALMs as cognitive models. By establishing auditory illusions as a rigorous testbed, our work offers a new perspective for probing neural black-box models and advancing understanding of auditory cognition. AIB is publicly available at https://github.com/gillosae/aib.
1. INTRODUCTION
Auditory illusions expose how perception can diverge from acoustic signals, but existing LALM benchmarks rarely test whether models share these human-like biases. AIB addresses this gap by combining ten illusion types with controlled human comparisons, revealing systematic differences between models and people.
- Auditory illusions reveal how the brain combines acoustic signals and prior knowledge to form percepts that diverge from physical stimuli.
- Existing LALM evaluations emphasize tasks with clear ground truth, while their susceptibility to human-like auditory illusions remains poorly understood.
- AIB is the first auditory illusion benchmark for LALMs, spanning ten representative illusions across music, sound, and speech.
- The benchmark pairs model evaluations with controlled human listening studies to enable direct comparison of perceptual responses.
- LALMs differ systematically from humans: several respond more human-like when linguistic or musical priors are involved, but no model matches the human perceptual profile.
- AIB extends LALM evaluation beyond recognition accuracy toward perceptual alignment and provides human baselines alongside empirical analyses.
2. RELATED WORK
Prior work established auditory illusions as probes of perception and developed increasingly capable audio-language benchmarks. However, auditory illusion benchmarks for testing human-like LALM perception remained absent.
- Auditory Illusions: Classic auditory illusions have been used to study auditory mechanisms, perceptual variation, and the roles of physiology and linguistic background.
- Large Audio Language Models: LALMs combine multimodal encoders with large-scale training to support recognition, translation, multitask understanding, and in-context dialogue about audio.
- Auditory or Illusion Benchmarks: Audio benchmarks have expanded from recognition to reasoning and instruction following across speech, music, and environmental audio.
- Auditory or Illusion Benchmarks: Vision-language benchmarks have tested illusion susceptibility, but no systematic auditory benchmark had examined whether LALMs perceive illusions like humans.
3. BENCHMARK DATASET
AIB constructs a controlled auditory-illusion dataset spanning domains and causal mechanisms, with matched illusion and control stimuli for comparing human judgments and model predictions.
- Overview: The AIB dataset contains 10,385 illusion stimuli and 4,444 control stimuli across music, speech, and general sound.
- Overview: Each illusion is categorized as physics-based or physics+knowledge-based according to its dominant mechanism.
- Overview: Matched stimulus and control sets enable controlled comparisons between human judgments and model predictions.
- Data Curation and Illusion Classification: Physics-based illusions arise from acoustic properties or basic auditory physiology, including frequency interactions, masking, and temporal adaptation.
- Data Curation and Illusion Classification: Physics+Knowledge-based illusions additionally depend on prior knowledge or expectations that impose linguistic, musical, or contextual structure on incomplete input.
- Data Curation and Illusion Classification: The classification reflects dominant explanatory mechanisms rather than a strict dichotomy, especially in ambiguous cases.
4. EXPERIMENTS
The experiments convert perceptual judgments into binary or ternary questions and compare LALM responses with human reference distributions using complementary alignment metrics. Human baselines come from controlled listening studies with absolute-pitch participants.
- Evaluation Setup: Models received the same stimuli as human listeners, with perceptual judgments reformulated as binary or ternary multiple-choice questions.
- Metrics: Human-Likeness Accuracy measures agreement with the dominant human response, while Reality Alignment measures agreement with the physical ground truth.
- Metrics: Illusion Susceptibility Index is defined as the difference between Human-Likeness Accuracy and Reality Alignment.
- Human Listening Study: Twenty participants with absolute pitch completed randomized, feedback-free illusion and control trials to establish perceptual baselines.
- Human Listening Study: Majority voting aggregated participant responses into robust ground-truth distributions for each illusion.
5. RESULTS
Across auditory-illusion domains, LALMs separate into susceptible, literal, and domain-dependent regimes, but none combines human-like susceptibility with reliable control performance. Prompt wording can also substantially change measured susceptibility, making model–prompt pairing part of the evaluation.
- Overall evaluation: 0.455: Audio Flamingo 3 achieves the best average ISI, still less than half the 1.0 human reference.Control accuracy remains essential because a degenerate always-illusion strategy reaches average ISI 0.799 while failing Physics controls almost entirely.
- Physics-based illusions: On physics-based illusions, MuLLaMa and Audio Flamingo 3 are the only models sharing the human-like percept, but both have below-chance control accuracy.The remaining models follow the physical signal; the most signal-faithful models also have the highest control accuracy.
- Physics+Knowledge illusions: On knowledge-driven illusions, Audio Flamingo 3 reaches the highest ISI at 0.505 and Gemini 3.1 Pro the highest HLA at 0.707.Gemini 3.1 Pro, Kimi-Audio-Instruct, and Fun-Audio-Chat reverse from signal-faithful acoustic responses to more human-like responses when linguistic or musical expectations apply.
- Prompt sensitivity: Wording alone shifts ISI by as much as 0.6, showing that instruction-tuned susceptibility depends on the model–prompt pair rather than the checkpoint alone.Under refined prompts, Qwen2-Audio-Instruct becomes purely physical with perfect Physics control accuracy.
- Model regimes: Across domains, models form susceptible, literal, and domain-dependent regimes, yet no model combines high susceptibility with reliable controls.Same-scale models span the full ISI range, indicating that scale alone is insufficient for matching the human profile.
6. DISCUSSION
The benchmark shows that current LALMs capture some low-level auditory regularities but diverge from humans on knowledge-driven illusions and often behave unstably. AIB therefore provides a complementary way to assess cognitive alignment, with the preferred degree of illusion susceptibility depending on the application.
- Findings: Current LALMs reproduce some low-level regularities but diverge from human perception on knowledge-driven illusions, with alignment reaching at most ISI 0.505.Several models reverse their physics-based behavior under linguistic or musical priors, while many remain unstable even on controls.
- Application scope: Higher illusion alignment may support natural and empathetic interaction, whereas lower susceptibility may be preferable for safety-critical or precision-measurement applications.The benchmark is intended to position models along this application-dependent spectrum rather than define universally desirable susceptibility.
- Implications: AIB complements recognition benchmarks by probing whether audio models reflect human perceptual biases rather than merely adhere to the physical signal.The conclusion frames auditory illusions as diagnostic tools for evaluating perceptual alignment and human-likeness.