Source-linked AI summary

Not all generalisation failures can be bought back: four boundaries in affective audio modelling

Jingyi Zhang, Xiaotong Yao

arXiv:2608.27674v1eess.AS

TL;DR

Affective audio mappings are usually evaluated within the corpus used for fitting, leaving unclear whether failures outside that setting reflect insufficient resources or missing information. The paper pushes one mapping across four application boundaries and measures target ceilings, surviving performance, and recovery costs. It finds that generalisation failures divide into budget-limited and information-limited cases with different remedies.

  • Problem

    Affective audio mappings are evaluated mainly within their fitting corpus, leaving their performance across changed material, measurements, and listeners insufficiently characterized.

  • Method

    The study pushes one mapping across four boundaries and reports each target’s ceiling, surviving fraction, and target-side observation cost.

  • Results

    The paper finds two kinds of failure: some gaps close with target observations, whereas others persist because the input does not appear to contain the required information.

  • Takeaways & Limitations

    “The model does not generalise” requires distinguishing budget-limited failures from information-limited failures because the remedies are mutually exclusive.

  • Takeaways & Limitations

    The physiological analysis does not relate naturalistic soundscapes to peripheral physiological measures and does not cover respiratory or actigraphic signals.

Abstract

from arXiv · show

Models mapping acoustic properties onto affective response underpin applications from music recommendation to sound design, yet are evaluated almost entirely within the corpus they were fitted on. When one fails outside it, the standard response -- more data, or a larger model -- assumes every failure is a shortage of resources. We show it is not, and that the alternative calls for the opposite remedy. Using four corpora of rated sound, four pretrained representations and three corpora of physiological recording, we pushed one mapping across four boundaries an application must cross: to new material, to edited audio, to a sensor in place of a self-report, and to an individual listener. At each we report the ceiling the target permits, the fraction surviving the crossing, and the price in target-side observations of closing the gap. Within a corpus, prediction reaches 84% of the ceiling set by inter-listener agreement. A same-domain corpus swap costs a fifth of that, and a hundred target labels return two-thirds of the loss. Crossing between music and environmental sound costs four-fifths to all of it, and four pretrained representations recover none of it. Against physiological response no information source we constructed exceeds a third of the attainable ceiling. "The model does not generalise" is therefore two diagnoses, not one, with mutually exclusive remedies; treating the second as the first is the more expensive mistake.

Introduction

The paper asks how much affective audio mappings survive four application-relevant changes and what it costs to recover what is lost. It distinguishes budget-limited failures, which target observations can reduce, from information-limited failures, which require information absent from the input.

  • Introduction: Within-corpus accuracy does not answer how much performance survives changes in material, response measurement, or listener.The relevant quantities are the target ceiling, surviving fraction, and target-side observations needed to recover losses.
  • Introduction: The standard remedies—more data or a larger representation—assume failures reflect resource shortages, but some targets require information absent from the input.These two diagnoses imply mutually exclusive remedies.
  • Introduction: Applications must evaluate mappings beyond their fitting corpus: on new material, edited audio, sensor responses, and individual listeners.The paper measures performance survival and recovery cost across four boundaries rather than relying on within-corpus accuracy.
  • Introduction: The study tests one mapping across four boundaries while fixing criteria in advance and reporting minimum detectable effects for null results.Null distributions are also reported with their behavior when the true effect is zero.
  • Introduction: The main result is that generalisation failure separates into budget-limited gaps that target observations can close and information-limited gaps that representation changes do not move.Transfer across corpora contains both kinds within a single experiment.
  • Introduction: The proposed distinction is falsifiable because a boundary closed by both additional observations and a representation change would disprove the dichotomy.The study includes a case where an information-limited gap is successfully filled, preserving the framework’s testability.

Results

Within-corpus prediction approaches the annotation ceiling, but generalisation depends sharply on the boundary crossed. Same-domain transfer is partly recoverable with target labels, whereas cross-domain transfer remains largely inaccessible to the tested representations.

  • Within-corpus ceiling: 84.4% of the attainable correlation (95% CI 81.5–87.0%) is reached within corpus after correcting ρ = 0.817 by annotation reliability.The corresponding models reach about 71% of reproducible variance, leaving roughly 29% unexplained by better annotation.
  • Corpus transfer: Same-domain corpus transfer preserves mean ρ = 0.592, whereas cross-domain transfer falls to mean ρ = 0.054.Relative to within-corpus performance, the swap costs about a fifth within domain and four-fifths to all performance across domains.
  • Domain boundary: ρ = 0.52 remains within environmental-sound categories, while cross-domain transfer reaches ρ = 0.05, isolating the collapse to the domain boundary.Transfer into music is uniformly low at ρ ≤0.09, whereas music-to-environment transfer varies with the source corpus’s atmospheric content.
  • Corpus transfer: 20–31% of within-corpus performance is lost for same-domain transfer, compared with 84–95% for cross-domain transfer.The 59–74 percentage-point gap shows that representations transfer between datasets without reliably crossing the sound-domain boundary.
  • Representations: AST raises cross-domain transfer from ρ = 0.047 to 0.125, but this is still only a sixth of its within-corpus performance.CLAP’s lower 57.3% barrier is associated with pretraining overlap, while its music-to-music transfer is not generally stronger than AST’s.
  • Recoverability: 100 target-side labels recover 64% of the same-domain transfer gap, while subspace alignment recovers −5%.The label route returns 20%, 34% and 50% of the gap at 10, 25 and 50 labels, respectively.
  • Descriptor stability: Within-corpus significance identifies temporal-variability descriptors in 17% of cases, versus 60–67% for routes probing generalisation or model multiplicity.The descriptors most associated with within-corpus prediction are precisely those that fail to transfer across the boundary.

The falsification test: an information-limited gap that a representation does close

The falsification test asks whether an information-limited gap can be closed by changing the representation rather than buying target-side observations. Representation changes helped some affective dimensions, but four pretrained representations did not close the cross-domain barrier, and the evidence supports only descriptor-level explanations.

  • Selective representational gains: A representation change improved prediction selectively: happiness gained +0.132, versus +0.045 averaged across the other seven axes.Happiness, valence, and sadness gained most, while energy and anger gained least.
  • Selective representational gains: 39 tonal descriptors reached ρ = 0.580 for happiness, exceeding the 122 spectro-temporal descriptors at 0.474.The deficit was representational rather than dimensional for happiness.
  • Target-dependent structure: Descriptor rankings depended on the affective target: mode ranked highest for happiness, whereas consonance and roughness ranked highest for tension.The ordering was selected by the model rather than imposed.
  • The barrier remains: Adding tonal descriptors lowered music-to-environmental transfer by 0.021, accounting for roughly 4% of the barrier.The paired comparison fell from 0.592 to 0.054, with p = 0.038 uncorrected.
  • The barrier remains: Tonal descriptors contributed little to arousal, the target on which the barrier was measured, so tonality was not its explanation.Within music, the gain was +0.006 for arousal and +0.012 for the energy axis.
  • What survives editing: Smoothing spectral flux transferred direction across domains: it lowered predicted arousal in both, despite predictions themselves failing to cross.The result indicates calibration, not sign, as the transferable component; smoother audio was still treated as calmer.
  • Physiological boundary: Physiological prediction remained below one-third of the attainable ceiling: the best source reached 30.6%, while acoustic routes reached 24.5% and 21.3%.Their confidence intervals overlapped, and a random projection reached 20.6% at its 95th percentile.
  • Physiological boundary: The physiological result is evidence from three routes failing alike, not from a saturation curve, so the study does not claim unlimited target-side data could never recover the gap.A target-side learning curve on a physiological outcome is identified as the most informative missing experiment.

Discussion

The discussion separates generalisation failures into budget-limited gaps that target observations can reduce and information-limited gaps that require a different information source. Across four boundaries, within-domain corpus and individual transfer are purchasable, whereas cross-domain and physiological transfer remain sharply constrained.

  • Two diagnoses: “The model does not generalise” comprises two diagnoses with opposite remedies: some gaps close with target observations, while others reflect missing information.The distinction was specified as falsifiable in advance rather than imposed after the results.
  • Budget-limited boundaries: A same-domain corpus swap costs a fifth of attainable performance, while 100 target-side labels recover two-thirds of that loss.The first hundred labels are substantially cheaper per label than later observations.
  • Information-limited boundaries: Crossing between environmental sound and music costs four-fifths to all performance, and four pretrained representations recover none of that loss.The transfer is asymmetric: environmental sound does not transfer into music, whereas music transfers outward with atmospheric content.
  • Information-limited boundaries: No constructed predictor exceeds 31% of the attainable physiological ceiling, with acoustic components, predicted ratings, and ratings themselves stopping within a third.The authors interpret this pattern as information-limited, while noting that it is based on tested routes rather than a saturating budget curve.
  • Budget-limited boundaries: Individual fitting costs 67 observations near the application setting, 166 in music ratings, and 318 in electrodermal response.The fivefold range makes corpus and setting choice a budgeting decision rather than a convenience.
  • Physiological boundary: The physiological conclusion is a bound: the strongest EEG measure reaches ICC 0.09 versus 0.22 for the best-agreeing rating scale, with non-overlapping intervals.The common response to sound is strong; the differentiating stimulus-specific response is weak, and the estimate is the maximum selected from 156 measures.

Methods

The study compares affective-audio mappings across music and environmental-sound corpora using ratings of perceived arousal, controlled audio preprocessing, engineered descriptors, and pretrained embeddings. It also tests whether within-music transfer reflects shared acoustic structure or annotation practice.

  • Corpora: Four rating corpora comprised DEAM, PMEmo, Soundtracks Set 1, and Emo-Soundscapes, with ESC-50 added as a category control.The corpora include three music datasets and two environmental-sound datasets.
  • Targets: The prediction target was group-averaged perceived arousal, rescaled to −1–1.Soundtracks also supplied seven additional affective scales rescaled to the same range.
  • Corpus design: Soundtracks was chosen to test whether within-music transfer reflects shared acoustic structure or shared annotation tradition.It came from a different laboratory and decade, used categorical rather than continuous ratings, and balanced twelve emotions across 30 excerpts each.
  • Preprocessing: All audio was resampled to 22,050 Hz, centre-cropped to at most 30 s, and loudness-normalised to −23 LUFS before descriptor computation.Without loudness normalisation, corpus-specific mastering levels could act as identifiers and inflate apparent transfer.
  • Representations: The engineered representation combined spectro-temporal, tonal, and pitch-related descriptors, including a key-invariant chroma profile and fixed interval-consonance weights.Chroma rotation prevents models from learning arbitrary corpus-specific key distributions.
  • Representations: A contrastive language–audio model supplied 512-dimensional pretrained embeddings for a separate representation comparison.One corpus overlapped with the model’s public pretraining data, so it was reported separately from two non-overlapping corpora.

Models and evaluation

The evaluation compares six regressors under grouped within-corpus cross-validation and whole-corpus transfer, while expressing performance relative to an annotation-reliability ceiling. Physiological analyses use carefully controlled EEG-derived measures and matched artefact treatment.

  • Models: Six regressors were evaluated: ElasticNet, support vector regression, k-nearest neighbours, random forest, gradient boosting, and XGBoost.Five-fold GroupKFold kept excerpts from the same source recording out of both training and test folds.
  • Evaluation: Performance was measured with Spearman rank correlation between predicted and observed values.Cross-corpus transfer trained on one complete corpus and evaluated on another.
  • Evaluation: Single-cell results report the best of six algorithms, whereas condition comparisons retain all six as paired observations and use Wilcoxon signed-rank tests.This distinguishes best-model reporting from paired model comparisons.
  • Ceiling: The annotation-reliability ceiling was estimated by correlating random half-splits of per-rater judgements and applying the Spearman–Brown correction.Model performance was expressed as a fraction of this ceiling.
  • Physiological measures: EEG analyses used 156 trial-wise measures, including baseline-relative band power, coherence, imaginary coherency, and log-ratio asymmetry.Coherence and asymmetry were retained because the original recordings showed effects in those families.
  • Artefact handling: Main analyses applied no epoch rejection, while sensitivity analyses removed independently identified ocular components with all other parameters fixed.The matched treatment isolates changes attributable to artefact removal.

Controls preceding physiological analysis

Before interpreting physiological analyses, the study required data-bookkeeping, channel, alignment, and stimulus-response checks, then evaluated reliability and per-listener decoding with null-based procedures. One initially failed control was revised after identifying a timing mismatch and an unsuitable shifted null.

  • Validation checks: Four prespecified checks covered event bookkeeping, channel identity and spectral estimation, event-to-sample alignment, and stimulus-locked response.Interpretation was halted if any check failed.
  • Validation checks: The alignment check relied on blink-locked frontal deflections rather than a physiological hypothesis.It tested whether recorded blink times and amplitudes were correctly represented.
  • Stimulus-locked response: Control 4 initially failed because its 0–300 ms window missed a response peaking at 452 ms and its temporally shifted null altered the noise environment.A subject-wise sign-flip null avoided moving the analysis window in time.
  • Stimulus-locked response: The mis-specified criterion would have discarded a valid dataset if followed.The revised control preserved the dataset for substantive analysis.
  • Reliability: Stimulus-level reliability used both intraclass correlation across 1,240 trials and split-half reliability on stimuli heard by at least eight listeners.The two reliability approaches were reported as agreeing.
  • Individual-level decoding: Per-listener decoding used fold-fitted standardisation, eight-component principal-component reduction, and ridge regression, with significance assessed against each listener’s permuted-label null.Observed decoding was not compared directly with zero because cross-validation biases the metric downward at small sample sizes.
  • Statistical control: Multiplicity across measures was controlled with the Benjamini–Hochberg procedure.

Design simulation

The design simulation estimates how many observations are needed to detect per-person physiological associations using the same analysis pipeline as the real data, then checks the extrapolation in an independent cue-response corpus. Event-locked effects are isolated with paired random anchors and exact sign-flip tests.

  • Design simulation: Simulation cells crossed four association strengths, r ∈ {0.10, 0.20, 0.30, 0.40}, with eight observations-per-person values from 30 to 5,000.Each cell contained 500 multivariate-normal datasets with a sparse linear association injected.
  • Design simulation: Each simulated dataset was analysed with the real-data pipeline, criterion, and permutation test, with power defined as detection at p < 0.05.The simulation therefore estimates requirements under the study’s intended analysis procedure.
  • Simulation results: The extended grid resolved weak-effect requirements at 1,364 and 4,888 observations, while stronger-effect requirements changed from 271 to 314 and from 650 to 564.The two stronger requirements shifted by less than the Monte Carlo error of the coarser run.
  • Simulation limits: The simulated requirements are lower bounds because multivariate-normal data understate heavy tails, associations are linear, and within-person stationarity is assumed.The stationarity assumption is least tenable over the long recordings used in discussion.
  • Independent check: An independent corpus of 75 adults provided EEG, electrocardiogram, and pupillometry recordings during cued reaction-time tasks for an out-of-sample check.Participants contributed 240 active-task cues and 30 passive-block cues.
  • Independent check: Primary response measures were a 80–130 ms fronto-central EEG amplitude, a 1–5 s heart-rate change, and a 0.5–2.5 s pupil-diameter change.A later EEG contrast was also prespecified as a secondary measure.
  • Event-locked analysis: Real cues were paired with randomly placed anchors, analyses used within-recording differences, and sign-flipping was exact under exchangeability.Random anchors also removed a monotone within-run heart-rate decline revealed by the anchor arm.
  • Event-locked analysis: Detection used 5,000 sign flips per participant with α = 0.05 and Benjamini–Hochberg correction, while trial-count curves repeated subsampled tests 40 times per value of n.False-positive rates used sign-randomised differences, and half-sample estimates were validated out of sample.

Supplementary electrodermal corpus

The electrodermal corpus analysis was withdrawn after group-level positive controls failed, and participant-level testing found no usable subset.

  • The electrodermal corpus results were withdrawn after analysis using NeuroKit2 and standard recommendations.
  • Group-level controls failed because skin-conductance habituation ran opposite to the established effect and post-stimulus windows did not exceed random windows.
  • None of 94 analysable participants passed both participant-level controls.Habituation was in the expected direction for only 35%, so the proposed usable subset was not supported.

Reproducibility

The reproducibility record fixes scripts, seeds, parameters, and reported metrics, but does not yet pin the code or fully hash inputs and outputs. Thus, the analyses are reproducible from the scripts without supporting a stronger bit-for-bit guarantee.

  • 172 run records identify each script, random seed, parameter set, and resulting metric, while source data for every figure panel are provided.
  • The authors explicitly distinguish what the run record establishes from the stronger reproducibility claim it does not establish.
  • All 172 records came from uncommitted working trees, so recorded commits identify starting points rather than the code state that produced each number.
  • Input hashes exist for 38 of 172 analyses, output hashes are absent, and no automated test suite exists.Figure scripts assert quoted values, but that discipline covers figures only.
  • The analyses are reproducible from the scripts and public corpora at recorded parameters, but not yet bit-for-bit from a pinned commit.

Data availability

The study uses publicly available data and provides accession, licensing, and derived-table information, while noting that research usability does not convey rights to underlying recordings. Several sources retain dataset-specific licensing constraints or ambiguities.

  • All analyses use publicly available data, with accession identifiers, access dates, and licence terms documented.
  • For ds002721, conflicting CC0 and CC BY 4.0 records led the authors to follow the more restrictive terms and attribute the source.
  • The remaining listed corpora carry dataset-specific terms, including varying DEAM licences, Creative Commons material in Emo-Soundscapes, and CC BY-NC 3.0 for ESC-50.
  • Commercial excerpts in PMEmo and Soundtracks Set 1 are not made reusable by the deposits' stated licences.
  • ARAUS is used only for cross-individual analysis, and its non-commercial term restricts reuse of the material while permitting the calibration requirement estimated from it.
  • Derived tables underlying every figure panel are provided as Source Data rather than deposited as a primary dataset.The files accompany the manuscript because they are small tabular files.

Code availability

Analysis code, run records, and figure-generation scripts are publicly available, with each reported value linked to execution details. The accompanying reproducibility statement limits that guarantee because code trees are uncommitted and hashing is incomplete.

  • Analysis code, run records, and figure-generation scripts are available in the study repository.
  • Each reported value has a run record identifying its script, random seed, and parameters.
  • The reproducibility record is limited by uncommitted working trees, incomplete input hashes, and absent output hashes.
  • The availability statement should not be read as a stronger guarantee than the documented reproducibility scope supports.

Ethics

The work is a secondary analysis of de-identified public data released by the original investigators. No new human-participant data were collected or re-identification attempted.

  • The analyses used de-identified data released for public reuse by the original investigators.
  • The original dataset studies obtained informed consent and ethical approval as described in their publications.
  • No new data were collected from human participants, and no attempt was made to re-identify anyone.

Competing interests

The authors disclose involvement in developing a commercial sleep-audio product.

  • The authors are engaged in developing a commercial sleep-audio product.

Generalisation stages a

Acoustic arousal models transfer between corpora within the same sound domain but collapse across music and environmental sound. This boundary remains evident across representations and is not explained by unseen content or clip duration.

  • Every tested representation collapses at the domain boundary, with cross-domain losses of 84–95% versus 20–31% within domain.
  • Holding out environmental soundscape categories preserves within-domain generalisation at ρ = 0.52, indicating the collapse is specific to the domain boundary.
  • Truncating music clips from 30 s to 6 s leaves same-domain transfer at 0.63 and cross-domain transfer near zero.
Loading 2608.27674v1…