Source-linked AI summary
Recovering Expert Critic-Sourced Network Adjacency between Musical Artists from Acoustic Distributions: A Construct-Validity Approach
Elena Badillo-Goicoechea, Fengfeng He
TL;DR
Recommendation systems face cold-start and grounding gaps, while critic co-mentions offer an expert-sourced but externally unvalidated similarity signal. This paper tests that signal against distributional acoustic representations and finds strong, consensus-dependent recoverability, separating a sonic core from a sociological remainder.
Problem
The paper addresses the missing external validation of whether critic-sourced adjacency is grounded in musical audio rather than sociological context.
Method
It represents artists as distributions over 80 acoustic descriptors, compares them with marginal Wasserstein distances, and predicts critic adjacency in an artist-disjoint cold-start evaluation.
Results
AUC 0.767 recovers critical adjacency overall, rising monotonically to AUC 0.865 for multi-source attested edges.
Takeaways & Limitations
Critical discourse contains a reproducible sonic core alongside a non-acoustic sociological remainder, making it informative for cold-start recommendation and construct-validity research.
Takeaways & Limitations
The acoustic representation is adopted as a faithful baseline rather than empirically adjudicated against simpler alternatives, so its optimality is outside the study’s scope.
Abstract
from arXiv · showhide
Music recommendation relies primarily on two signals: user-item interactions, which fail in the cold-start regime, and intrinsic musical content, available for any recording. We argue that a third, largely untapped signal is both richer and more principled: critical adjacency, the pairwise relation established when an expert critic explicitly links two artists in long-form prose. It encodes deliberate judgments about which artists belong together. Prior work established its internal validity, showing it recovers coherent, interpretable communities and can match collaborative filtering in user-satisfaction simulations, with no user data. What has been missing is external validation: whether this critic-sourced relation is grounded in the music itself versus sociological context. We test it against acoustic content, reframing the question as one of construct validity. Representing artists as empirical distributions over 80 low-level Essentia acoustic descriptors and modeling pairwise proximity via marginal optimal-transport (Wasserstein) distances, we evaluate how far critical adjacency is sonically recoverable under a cold-start, artist-disjoint split. Our ensemble recovers these edges at out-of-sample AUC of 0.767 (95% CI 0.761-0.775). Recoverability rises monotonically with critical consensus, reaching 0.865 on multi-source attested edges. Stratified evaluations align with sociological models of genre: tightly bounded, scene-based genres show higher recoverability than broad industry umbrella terms. Critical discourse is thus a rich source of information for recommendation, decomposing into a reproducible "sonic core" and a "sociological remainder" driven by narrative positioning, subcultural context, and canonical placement. The work offers both a scalable cold-start discovery mechanism and a sociologically grounded approach to MIR and MRS research.
1 Introduction
The paper frames critic-sourced artist connections as a construct-validity question: how much critical adjacency reflects acoustic similarity versus social and narrative context. It evaluates this question with distributional audio representations and finds that recoverability increases with critical consensus.
- Motivation: Critical adjacency is an expert-articulated relation formed when professional critics co-mention artists in long-form journalism.The paper distinguishes this deliberate signal from interaction-based, raw-audio, and metadata-based recommendation signals.
- Research question: The study asks whether critic co-mentions validly measure musical similarity or instead reflect non-acoustic social context.The construct-validity framing separates sonic grounding from narrative, subcultural, political, and canonical influences.
- Implications: Testing critic graphs against audio both validates the quality of expert-articulated relations and measures how much of their signal is sonically grounded.The comparison uses audio as an objective baseline rather than treating critical adjacency as a direct proxy for acoustic similarity.
- Method: Artists are represented as distributions over 80 acoustic descriptors, with pairwise proximity modeled through marginal Wasserstein distances and a stacked ensemble.The evaluation uses a review-derived graph of 19,059 artists under an artist-disjoint cold-start split.
- Results: AUC 0.767 at weight ≥1 establishes out-of-sample acoustic recoverability of critical adjacency.Recoverability is evaluated under an artist-disjoint, cold-start setting.
- Results: AUC rises monotonically from 0.767 for single mentions to 0.865 for multi-source agreement, separating a reproducible sonic core from a sociological remainder.Higher consensus aligns with shared acoustic traits, whereas single-source edges retain more narrative and social content.
2 Related Work
Related work positions the study at the intersection of MIR construct validity, genre sociology, network-based recommendation, and criticism studies. Together, these perspectives treat similarity labels and critic relations as socially situated and potentially contestable rather than fixed ground truth.
- MIR construct validity: MIR validity frameworks distinguish construct validity from statistical, internal, and external validity, motivating an explicit test of critic co-mentions as musical-similarity measures.The framework also questions whether genre-classification accuracy establishes music similarity.
- MIR construct validity: Human similarity judgments are only moderately consistent, so imperfect AUC may partly reflect disagreement within the critic-based similarity construct.This interpretation is presented as context for the observed AUC range, not as a demonstrated cause of model error.
- MIR construct validity: Multiple-source critic relationships are more acoustically recoverable than single-source relationships, supporting consensus as evidence against treating the network as unquestioned ground truth.The related-work discussion presents this as evidence against the strongest view that musical meaning is entirely subjective and context dependent.
- Genre as a limited proxy: Genre labels generalize poorly across datasets, and human agreement across a 19-genre scheme ranges from 26% to 71%.These findings contextualize variation in within-genre AUC and limit genre as a standalone proxy for musical similarity.
- Genre as a limited proxy: Scene-based genres such as metal, hip-hop/rap, and electronic music may be more sonically coherent than broad industry-based or traditionalist categories.The explanation follows sociological genre typologies that distinguish tightly bounded scenes from broader social formations.
- Network-based recommendation: Behavioral networks face cold-start and popularity-reinforcement problems, while contextual networks can contain bias and spurious web-derived signals.Critic adjacency instead uses professional criticism and language processing to construct relationships less dependent on interactions or unmoderated online content.
- Sociology of criticism: Criticism studies explain why critic relations can encode more than acoustic resemblance: critics act as gatekeepers and cultural intermediaries involved in consecration and symbolic positioning.This makes critic-based similarity distinct from musician, label, and general web co-occurrence networks.
3 Data
The study builds a sparse critic-sourced artist graph from long-form album reviews and pairs it with low-level acoustic descriptors from AcousticBrainz. The resulting data support artist-level distributional comparisons under weakly attested network links.
- Critical adjacency graph: The critic graph contains roughly 25,000 artists and is derived from approximately 65,000 reviews across ten professional music-criticism outlets.Examples include Pitchfork, NPR, The Guardian, The Quietus, and Bandcamp Daily.
- Critical adjacency graph: Directed edges record critic co-mentions, while edge weight a_ij counts how often artist i is mentioned across reviews of artist j.The graph therefore represents the intensity and direction of review-based artist relations.
- Critical adjacency graph: The graph is sparse, with density approximately 8.6×10^-4, a median edge weight of one, and roughly four in five linked artist pairs supported by a single source.Most observed relationships are therefore weakly attested.
- Audio features: Each recording contributes 80 usable scalar Essentia descriptors spanning timbre, harmony, spectral shape, rhythm, tonal features, and loudness.The descriptors are distributed through the AcousticBrainz corpus using a single pinned extractor for cross-contributor comparability.
4 Methodology & Construct-Validity Framework
The framework validates critic-sourced adjacency by using distributional acoustic distances as predictors and a review-derived graph as supervision under artist-disjoint cold-start evaluation.
- Audio content layer: Each artist is represented as an empirical distribution over 80 Essentia descriptors rather than a mean vector.This preserves catalogue spread, skew, and multimodality in the acoustic representation.
- Audio content layer: Marginal Wasserstein distances produce an 80-dimensional vector describing feature-wise distributional differences between artist pairs.Each distance is computed from one-dimensional empirical distributions using quantile-function differences.
- Contextual target layer: Critic adjacency is constructed from artist co-mentions in review corpora and supplies the prediction target.Named-entity recognition identifies mentions and resolves them against the reviewed-artist roster.
- Prediction model: A Super Learner maps the acoustic distance vector to the probability that a critic-sourced edge exists.The ensemble stacks regularized logistic regression, random forests, gradient-boosted trees, and nearest-neighbor learning.
- Evaluation design: The evaluation partitions artists, rather than pairs, so neither artist in a test pair appears during training.This artist-disjoint split operationalizes the cold-start setting.
- Construct-validity rationale: The representation is adopted as a faithful acoustic baseline, not as an empirically established optimal audio representation.The authors argue that full marginal distributions retain information discarded by feature means.
5 Results
The ensemble recovers critic-sourced adjacency above chance in strict cold-start evaluation, with recoverability increasing as critical consensus strengthens and varying across genre organization and exposure.
- Overall recovery: 0.767 overall AUC (95% CI 0.761–0.775) was achieved on the artist-disjoint cold-start test split.Pair-random cross-validation reached AUC 0.803, but permits artists to appear in both training and test folds.
- Critical consensus: AUC rises monotonically from 0.767 at attestation weight ≥1 to 0.865 at weight ≥5.Intermediate values are 0.788 at ≥2 and 0.815 at ≥3.
- Critical consensus: Higher-consensus edges form a reproducible sonic core, whereas single-source links more often reflect sociological remainder.The remainder includes narrative framing, personas, subcultural narratives, political positioning, and canonical placement.
- Genre controls: Within every adequately powered genre, AUC remains high, indicating discrimination does not collapse when genre is held constant.The evaluation uses exogenously assigned genres independent of audio and review sources.
- Genre typology: 0.80 AUC was observed for Electronic and Hip-Hop/Rap, compared with 0.68 for Pop within genre.Several genres below the power threshold, including Jazz, Classical, and R&B, are not interpreted.
- Genre typology: 0.796 same-genre AUC exceeded 0.765 cross-genre AUC, contrary to a simple genre-detector explanation.Cross-genre recoverability remains nearly as high, consistent with acoustic affinities spanning traditional boundaries.
- Exposure robustness: Recoverability is nearly invariant across degree and mention-count exposure, while track count increases AUC from 0.679 to 0.748.The track-count gradient is interpreted as improved measurement precision from better-estimated acoustic distributions.
6 Discussion & Future Directions
The study separates critical adjacency into a sonically recoverable core and a non-acoustic sociological remainder, while proposing distributional acoustic similarity for content-based applications. Future work extends this framework toward modeling the remainder and evaluating predicted adjacency for cold-start retrieval.
- Critical adjacency integrates a sonically recoverable core with a non-acoustic sociological remainder.
- The proposed similarity measure represents artists as acoustic-feature distributions and compares them using per-feature Wasserstein distances.
- The distributional measure requires no interaction data and may generalize to items represented as feature distributions.
- Future research could explain the sociological remainder using reviewer metadata and contextual information such as geography, labels, collaborations, and interviews.
- Predicted adjacency scores could support cold-start neighbor ranking and retrieval evaluation against audio-only and behavioral baselines.
Ethics and Privacy Statement
The study uses public audio descriptors and music-publication data without personal data or human subjects, but deployment could reinforce existing coverage biases. Both the audio corpus and critical target overrepresent Anglophone and canonical artists.
- The study uses publicly available audio descriptors and a graph derived from publicly available music publications without personal data or human subjects.
- Deploying this discovery signal uncritically could reinforce coverage biases because its sources overrepresent Anglophone and canonical artists.