Source-linked AI summary
Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data
Nizam Mohammed, Abu B. S. Rahman, Dimuthu D. K. Arachchige
TL;DR
Project Qualia asks whether listening sessions contain experiential structure beyond existing metadata and artist identity. It tests this by analyzing session-derived song embeddings and artist-residual similarities, finding coherent cross-artist clusters while identifying limits in the validation model and dataset provenance.
Problem
Existing recommendation systems use proxies such as genre, tempo, artist affinity, and co-listening statistics rather than a learned representation of listening experience.
Method
Project Qualia treats consecutive songs in listening sessions as behavioral traces, collecting dense Last.fm histories and testing Song2Vec embeddings with artist-residual analysis.
Results
Residual neighborhoods formed cross-artist experiential clusters, including 2020 mainstream pop, while Frank Ocean’s tracks remained same-artist neighbors at cosine similarities of 0.775–0.920.
Takeaways & Limitations
The findings provide empirical support for testing an architecture that represents experiential structure directly rather than correcting artist-identity effects after training.
Takeaways & Limitations
The validation model ignores sequence order and session-level structure, while the residual procedure only works around artist identity after training.
Abstract
from arXiv · showhide
This report presents results from Project Qualia, an ongoing effort to determine whether experiential similarity between songs, a structure not captured by genre or metadata taxonomies, can be recovered from real listening behavior. We constructed a large-scale dataset of listening sessions, comprising 1.29 billion scrobbles collected from 9,396 users via the Last.fm API and reduced through a preprocessing pipeline to 531.6 million training scrobbles across 28.6 million sessions. On this corpus, we trained a skip-gram Word2Vec model (Song2Vec), treating each session as a sentence and each track as a token. As anticipated, the resulting embedding space was dominated by artist identity, a consequence of single-artist runs within sessions. To test for a subtler, artist-independent signal, we developed an artist-residual procedure: subtracting each artist's centroid from its tracks' embeddings and evaluating whether the remainder retained structure. Mean cross-artist cosine similarity fell from 0.2487 in raw embedding space to 0.0005 in residual space, yet 4,577 cross-artist track pairs retained cosine similarity $\ge 0.70$ in residual space, forming coherent genre- and era-based clusters, including trip-hop, 1990s grunge, 2020 mainstream pop, and cross-composer classical piano pairs at cosine similarity up to 0.95. These results confirm that the training data contains experiential structure independent of artist identity, establishing an empirical basis for an architecture designed to learn this experiential layer directly.
1. INTRODUCTION
Project Qualia asks whether listening sessions contain an experiential structure beyond existing recommendation proxies such as genre, artist metadata, and co-listening statistics. It tests whether this signal can be recovered from data before designing an architecture to represent it directly.
- 1. INTRODUCTION: Existing recommendation systems predict what listeners are likely to play next but do not capture what is consistent with their current experience.The stated proxies include genre, tempo, artist affinity, and co-listening statistics.
- 1. INTRODUCTION: The motivating interpretation is that music may have an experiential structure distinct from and irreducible to artist identity or genre classification.This interpretation motivates learning predictive representations over that structure rather than relying only on shallow input features.
- 1. INTRODUCTION: Project Qualia tests whether consecutive-song listening sessions encode an experiential structure not captured by existing metadata schemas.The report treats this as a falsifiable hypothesis about the existence of a recoverable signal, not yet about system performance.
- 1. INTRODUCTION: The project separates empirical signal validation from later architecture development, first testing whether the hypothesized structure exists in the data.The eventual architecture is framed as a second stage adapted from the JEPA framework for sequential listening data.
2. DATA
The study collected a new Last.fm corpus and characterized its identity, listening, session, and contamination patterns before full-scale preprocessing. These analyses motivated thresholds, session construction, storage choices, and filtering decisions for co-occurrence modeling.
- 2. DATA: 1,292,098,037 scrobbles from 9,396 users were collected through the Last.fm public API for the new corpus.The records were stored in a local SQLite database of approximately 222.7 GB before compression and archival.
- 2. DATA: String-based identity inflated the apparent vocabulary by approximately 4.6 times relative to MBID-based identity, while MBIDs were missing from some scrobbles.Neither identity scheme alone was sufficient for track identification.
- 2. DATA: 77.9% of unique tracks in the preliminary sample were played by exactly one user, limiting usable co-occurrence signal for those tracks.The project therefore examined vocabulary reduction and listener-count thresholds before training.
- 2. DATA: A 30-minute inter-track gap rule yielded 23.9-track mean sessions, with 53.2% containing at least eight tracks.The sample also exposed mechanically uniform timing and a BTS concentration as contamination sources requiring later handling.
- 2. DATA: Requiring a track to be heard by at least ten distinct users reduced vocabulary by 93%, from 30.4 million to 2.2 million tracks, while discarding 16.4% of scrobbles.MBID deduplication alone reduced vocabulary by only 1.7%.
- 2. DATA: 44.5% of full-corpus scrobbles were a user’s fiftieth or later play of a track, while sessions of at least 100 tracks represented 38.6% of scrobbles despite being 3.1% of sessions.These patterns informed later handling of repetition, long sessions, and automated playback.
3. METHOD
The method cleans listening sessions, trains skip-gram Song2Vec embeddings, and tests whether artist-independent experiential structure remains after subtracting artist centroids.
- Data preprocessing: The pipeline reconstructed sessions, removed automated playback, thresholded the vocabulary, collapsed repeats, capped per-user track plays, and filtered short sessions.Sessions used a 30-minute inter-track gap rule; automated playback was detected with gap-uniformity criteria, while later stages reduced repeated and user-dominated signals.
- Data preprocessing: The moderate vocabulary threshold retained 1,960,021 tracks and 84.1% of scrobbles, while mapping 15.9% of scrobbles to an out-of-vocabulary token.The threshold prioritized sufficient per-track co-occurrence signal over marginal coverage gains from looser thresholds.
- Embedding model: Song2Vec applied skip-gram Word2Vec with negative sampling to 189 session shards, treating each session as a sentence and each track as a token.The model used the standard text-embedding formulation for listening-behavior sequences, with center tracks, context tracks, and negative samples.
- Artist-residual analysis: The analysis asked whether embedding proximity reflected experiential similarity or artist identity, motivated by long single-artist runs in many sessions.The study compared raw and residual cosine similarities across random cross-artist pairs and inspected nearest-neighbor cluster composition.
- Artist-residual analysis: Artist residuals were formed by subtracting each artist’s centroid from each track embedding and L2-normalizing the remainder.Centroids were computed for artists with at least five vocabulary tracks, covering 52,445 artists and 90.7% of the vocabulary.
- Artist-residual analysis: Residual vectors from different artists were expected to align non-randomly when the embeddings captured experiential structure independent of artist identity.The procedure therefore distinguishes an identity-only space, where cross-artist residuals should be effectively random, from one retaining shared experiential structure.
4. RESULTS
Artist-residual analysis removed the dominant artist-identity component while retaining a smaller, interpretable cross-artist structure. The residual neighbors formed genre-, era-, and stylistically coherent clusters, supporting the presence of experiential signal in session co-occurrence data.
- 4.2. Residual Space Isolates a Second, Independent Signal: 0.0005 mean cross-artist cosine similarity after residualization, down from 0.2487 in raw space, indicates near-complete removal of artist-identity bias.The residual comparison used 100,000 randomly sampled cross-artist pairs with above-median residual norm.
- 4.3. Cluster Composition in Residual Space: Cross-composer classical piano pairs reached cosine similarity up to 0.946, suggesting residual alignment with tonal character and pianistic texture rather than composer identity.A Chopin–Bach pair reached 0.941, while Saint-Saëns–Grieg pairs reached 0.946 and 0.935.
- 4.3. Cluster Composition in Residual Space: Artist-internal cohesion persisted for some catalogs, while Billie Eilish instead aligned mainly with a broader late-2010s alternative-pop cluster.Frank Ocean retained all 15 neighbors within his catalog; Taylor Swift and Kendrick Lamar showed similar, less extreme patterns.
- 4.2. Residual Space Isolates a Second, Independent Signal: 4,577 cross-artist pairs reached cosine similarity ≥0.70 in residual space against a 0.0005 random baseline, forming interpretable non-artist structure.The pairs came from the ten nearest residual neighbors of 50,000 sampled tracks.
- 4.3. Cluster Composition in Residual Space: Residual neighbors recovered coherent trip-hop, 1990s alternative-rock, grunge, and 2020 mainstream-pop clusters without genre labels during training.Examples included Massive Attack–Portishead, Radiohead–R.E.M., Nirvana–Soundgarden, and The Weeknd–Doja Cat or Dua Lipa neighbors.
- 4.4. Interpretation: The validation establishes that the dataset contains cross-artist experiential signal, but Word2Vec only separates it from artist identity post hoc and cannot model whole-session structure.The result motivates a predictive architecture that represents artist identity and experiential similarity as separable components during learning.
5. DISCUSSION & LIMITATIONS
The residual analysis finds cross-artist structure in session co-occurrence data, but its interpretation and downstream usefulness remain constrained by validation, modeling, provenance, and sampling limitations.
- 5.1. What This Result Does and Doesn’t Establish: Residual clusters are genre- and era-coherent, but the analysis does not distinguish shared musical feel from acoustic or contextual similarity.Genre coherence may reflect production era, instrumentation, or tempo rather than listeners’ experiential similarity specifically.
- 5.1. What This Result Does and Doesn’t Establish: 4,577 cross-artist pairs exceeded cosine 0.70, but the evidence establishes signal existence rather than its strength or reliability for recommendation.Validation used an aggregate threshold statistic and illustrative inspection of 20 seed tracks, without a held-out similarity benchmark or downstream task.
- 5.1. What This Result Does and Doesn’t Establish: Centroid subtraction assumes an artist’s mean embedding separates identity from feel, yet cohesive catalogs did not always separate cleanly.The analysis cannot determine whether those cases reflect genuine catalog cohesion or a limitation of the linear residual method.
- 5.2. Autoplay and the Provenance of Session Structure: Autoplay-originated scrobbles cannot be identified or quantified from the collected Last.fm API data, leaving their effect on session structure unresolved.The proposed direct tests using client instrumentation or documented platform sequences were not included in the collection design.
- 5.2. Autoplay and the Provenance of Session Structure: Chopin–Bach correspondence provides limited evidence against an autoplay-only explanation, whereas the 2020 pop cluster remains especially confounded by genre radio.The analysis recommends restricting future residual analysis to clients without autoplay functionality before large-scale architecture training.
- 5.3. Population and Concentration Bias: The findings may not generalize beyond heavy Last.fm users because the sampled population is deliberately dense, curated, and unrepresentative of listeners generally.The corpus also contains concentration bias, including BTS-related listening accounting for 5.7% of the full corpus and broader K-Pop for 6.5%.
- 5.4. Model Scope: Skip-gram Word2Vec was chosen for validation convenience, but its unordered fixed-window representation omits session order and explicit identity–feel separation.These structural limitations make it unsuitable as a justified terminal model for the stated goal.
6. FUTURE WORK
Future work proposes a gated five-layer architecture for learning experiential structure directly, beginning with a JEPA-based Layer 1 and predefined evaluation criteria. The design remains planned, with unresolved modeling choices and open questions about whether the learned structure reflects experiential interchangeability rather than correlated confounds.
- 6.2. Proposed System Architecture: The proposed system has five layers, but only Layer 1 is currently specified; later layers are designed only after preceding evaluation criteria are met.No component in this section has been implemented, and the five-layer architecture is presented as planned work.
- 6.3. Proposed Layer 1 Objective and Theoretical Basis: Layer 1 adapts JEPA to sequential listening by predicting target representations at masked session positions from unmasked context.The design uses context and target encoders plus a predictor, with the target encoder updated by exponential moving average.
- 6.3. Proposed Layer 1 Objective and Theoretical Basis: The representation-space objective is intended to avoid the artist-affinity confound that Song2Vec learned from song-identity prediction and later required residual correction.Whether this objective actually prevents the confound remains an evaluation question, not an established result.
- 6.4. Open Design Questions: Layer 1 remains unresolved on representation collapse, listener-specific variation, acoustic features, encoder configuration, repeated listens, and album-block compression.The report lists candidate mitigations or design ranges for some issues, but several alternatives remain unevaluated or undecided.
- 6.5. Proposed Evaluation Protocol: Progress to Layer 2 requires passing nearest-neighbor inspection, genre-entropy analysis, and session-completion accuracy against random, popularity-weighted, and genre-matched baselines.Exceeding the genre-matched baseline is the deciding criterion, and failure triggers revision before work on higher layers.
- 6.5. Proposed Evaluation Protocol: The evaluation must also determine whether residual cross-genre clusters reflect experiential interchangeability rather than confounds such as production era or energy level.The report explicitly leaves this question open for the planned evaluation.