Source-linked AI summary

Effects of HRTF Augmentation on Predicted Spatial Release from Masking in Music

Jack Webb, Christophe Lesimple, Volker Kuehnel, Lorenzo Picinali

arXiv:2608.28422v1eess.AS

TL;DR

Listeners with hearing loss may struggle to separate instruments in complex mixtures, and spatial-release benefits for music are poorly understood. The paper augments individual HRTFs to increase spatial-cue salience and evaluates the approach with auditory models. Augmented HRTFs increase predicted musical SRM, but moderate hearing loss substantially reduces the benefit and simulated hearing-aid processing does not restore it to normal-hearing levels.

  • Problem

    The role of spatial release from masking in music scene analysis, especially with hearing loss, remains poorly understood.

  • Method

    The study augments direction-dependent spectral components of individual HRTFs and evaluates musical SRM with computational auditory models.

  • Results

    Augmented HRTFs produced elevated predicted SRM relative to individual HRTFs; the advantage persisted under moderate hearing loss but was substantially reduced.

  • Takeaways & Limitations

    HRTF augmentation may improve predicted musical-instrument separability, but hearing loss limits the size of the predicted benefit.

  • Takeaways & Limitations

    The models omit auditory scene-analysis cues relevant to music and approximate hearing-loss effects, limiting direct behavioural interpretation.

Abstract

from arXiv · show

Separating individual musical instruments within a complex mixture of sounds poses a persistent challenge for listeners with hearing loss. Although spatial separation of sources improves speech recognition in this population, the potential benefits of spatial cue enhancement for music perception remain largely unexplored. This paper introduces a method to increase spatial cue salience through the augmentation of individual head-related transfer functions (HRTFs). Auditory model analyses indicate that augmented HRTFs may enhance the separability of musical instruments relative to individual HRTFs. Predicted benefits persist when moderate sensorineural hearing loss is modelled, though they are substantially reduced. Simulated hearing aid processing does not restore these benefits to normal-hearing levels.

1. Introduction

Music scene analysis is difficult for listeners with hearing loss, while the role of spatial release from masking in separating musical instruments remains poorly understood. This study proposes augmenting HRTF-based spatial cues and evaluates predicted effects under normal hearing and hearing loss.

  • Music listening with hearing loss is complicated by hearing-aid signal distortions and reduced acuity in music perception tasks.
  • Music scene analysis requires discerning a single instrument from competing musical sounds, but the role of spatial release from masking remains elusive.
  • Prior spatial manipulation evidence in music is limited to a small but inconsistent stereo-width benefit induced by frequency-independent interchannel level differences.
  • The study aims to augment interaural cues in individual HRTFs, estimate their contribution to musical SRM, and assess hearing-loss effects.
  • Because the study uses computational models of SRM and hearing loss, its results are predictions intended to support future behavioural validation.

2. Methods

The method augments direction-dependent spectral components of individual HRTFs using PCA-based residual contrast, then predicts musical SRM across rendering, geometry, instrument, and hearing-profile conditions. The analysis combines auditory models with systematically rendered musical mixtures.

  • 2.1. HRTF Augmentation: The method scales direction-dependent spectral components of individual HRTFs to enhance spatial cues while retaining the listener-specific broader transfer structure.
  • 2.1. HRTF Augmentation: PCA decomposes the HRTF side channel across listeners and horizontal directions, after which residual deviations from each listener’s mean score vector are magnified.
  • 2.1. HRTF Augmentation: The augmented representation uses the first six principal components, reconstructs the side spectrum, and rebuilds augmented DTFs from the unchanged mid component.
  • 2.1. HRTF Augmentation: Only the contralateral ear is augmented and midline positions remain unchanged, magnifying frequency-dependent ILDs while preserving broader monaural structure.
  • 2.2. Stimulus and Rendering Conditions: The musical dataset contained 2-second mixtures with one target and three maskers, rendered under diotic, cue-based, individual-HRTF, and three augmented-HRTF conditions.
  • 2.2. Stimulus and Rendering Conditions: Targets covered guitar, piano, synthesiser, and drums across two horizontal-plane geometries, with eight samples per category repeated 25 times using unique listener HRTFs.
  • 2.2.2. Auditory Model Analysis: SRM was estimated with vicente2020 and bischof2023 using better-ear SNR and binaural unmasking, with bischof2023 better suited to non-stationary music signals.
  • 2.2.2. Auditory Model Analysis: vicente2020 compared normal hearing, moderate sensorineural hearing loss, and hearing loss with simulated WDRC processing.

3. Results

Augmented HRTFs increased predicted spatial release from masking under normal hearing, while moderate hearing loss reduced both overall SRM and augmentation gains. Simulated WDRC restored effective SNR but did not systematically restore SRM.

  • Normal hearing: 10.31 dB SRM was predicted for a frontal target with HRTF4.0, compared with 6.97 dB for individual HRTFs.All conditions produced an SRM benefit relative to the diotic baseline.
  • Normal hearing: 3.08 dB was the mean N0 SRM increase from individual HRTFs to HRTF4.0 across target locations.The increase comprised a 3.35 dB BE SNR increase and a 0.27 dB BU reduction.
  • Model comparison: 59% versus 29%: adding ITD to ILD increased SRM more under bischof2023 than under vicente2020.Under bischof2023, ILD+ITD also exceeded HRTF3.0 in SRM.
  • Hearing loss: 7.60 dB versus 9.37 dB: unaided N3 hearing loss reduced mean SRM across conditions relative to N0.Unaided N3 also reduced access to BU by an average of 0.65 dB across conditions.
  • Hearing loss: 1.05 dB versus 3.08 dB: the HRTF4.0 gain over individual HRTFs was smaller with unaided N3 than with N0, despite remaining significant.The N3 comparison had p < 0.001.
  • Hearing-aid simulation: WDRC raised mean effective SNR from −6.84 to −2.76 dB but did not systematically change SRM across conditions.The HRTF4.0-versus-individual-HRTF SRM gain remained unchanged from unaided N3 (Δ = +0.01 dB, p = 0.81).
  • Target instrument: Under N0, HRTF4.0 SRM gains differed by target instrument: synthesiser targets gained 3.39 dB versus 2.80 dB for drums.The difference was significant after Holm adjustment (p = 0.007); no significant instrument differences emerged under N3 or aided N3.

4. Discussion

Augmented HRTFs produced greater predicted spatial release from masking than individual HRTFs, but hearing loss substantially reduced this advantage and hearing-aid processing did not restore normal-hearing performance.

  • Spatial cue contributions: Augmented HRTFs yielded greater predicted SRM than individual HRTFs for normal-hearing listeners.The predicted advantage was associated largely with elevated binaural-ear SNR.
  • Spatial cue contributions: Predicted SRM increased when ITDs were added to ILDs, although the relevance of ITD-based equalisation-cancellation mechanisms for music remains unresolved.The ILD+ITD condition produced greater predicted SRM than the ILD condition, differing from a prior speech-onspeech result.
  • Spatial cue contributions: Predicted SRM depended on the frequency distribution of spatial cues, with low frequencies potentially contributing more in music than speech-weighting functions assume.Music’s more variable long-term spectra may make low-frequency SRM behaviorally relevant.
  • Hearing-loss effects: Under an unaided N3 loss, the augmentation benefit persisted but was largely diminished because elevated internal noise rendered higher-frequency ILDs inaudible.The primary SRM determinant shifted from head-shadowing advantages toward target audibility, while HRTF spectral colouration provided brief target-energy glimpses above the noise floor.
  • Hearing-loss effects: Hearing-aid processing partially restored audibility and the normal-hearing condition ordering, but average SRM remained below normal-hearing levels.Residual hearing-loss effects and interaural distortions from unlinked bilateral WDRC processing likely contributed to the remaining deficit.
  • Limitations: The modelling approach omits several music-specific scene-analysis cues and higher-level processes, so behavioural SRM may diverge from these predictions.The hearing-loss model also approximates hearing impairment without explicitly modelling auditory-filter broadening or reduced temporal fine-structure processing.

5. Conclusion

The study introduced HRTF spectral-cue augmentation to increase spatial release from masking for target instruments in music. Model predictions showed a benefit over individual HRTFs, but hearing loss substantially reduced it and hearing-aid processing did not meaningfully recover the gain.

  • Augmented HRTFs increased predicted SRM for target instruments compared with individual HRTFs.
  • The predicted augmentation benefit persisted under moderate sensorineural hearing loss but was substantially reduced.
  • Hearing-aid processing partially restored overall audibility but did not meaningfully recover augmentation-related SRM gains.
Loading 2608.28422v1…