Source-linked AI summary
MERIT: Learning Disentangled Music Representations for Audio Similarity
Abhinaba Roy, Junyi Liang, Dorien Herremans
TL;DR
Music similarity models often entangle melody, rhythm, and timbre in one score, limiting nuanced similarity judgments. MERIT learns separate factor-specific representations using factor-controlled triplets, and its heads selectively respond to their intended dimensions across synthetic and independent real-world audio.
Problem
Existing music similarity models entangle melody, rhythm, and timbre in monolithic scores despite listeners judging similarity along individual dimensions.
Method
MERIT learns independent melody, rhythm, and timbre projection heads using factor-controlled triplets constructed with conditional generation and source-separated stems.
Results
MERIT’s intended heads respond selectively to their target factors while remaining near chance on others across synthetic tests and independent real-world audio.
Takeaways & Limitations
MERIT exposes melodic, rhythmic, and timbral similarity as three separable, addressable scores.
Takeaways & Limitations
The decomposition covers only melody, rhythm, and timbre, while timbre is represented at instrument-class granularity and under-resolves within-class variation.
Abstract
from arXiv · showhide
Current music similarity models typically compute a single, monolithic score, entangling distinct musical dimensions like melody, rhythm, and timbre. This limits user control and interpretability, making it impossible to execute nuanced queries. We introduce MERIT, a framework for learning disentangled, factor-specific music representations tailored to these three core dimensions. To overcome the lack of isolated musical variations in real-world audio, we use a novel training strategy that uses conditional audio generation and source-separated stems to strongly encourage single-factor variation in training data. Our evaluations demonstrate strong factor-wise disentanglement. Each head responds strongly to its intended perceptual dimension while remaining near chance on the others, a representational property that holds across both the synthetic training domain and independent real-world audio.
1. Introduction
MERIT addresses the multi-dimensional nature of music similarity by learning independent, interpretable scoring channels for melody, rhythm, and timbre. Its factor-controlled training pipeline and evaluation protocol target functional selectivity and generalization to real-world audio.
- Introduction: Music similarity spans melody, rhythm, and timbre, but conventional systems collapse these dimensions into a single aggregate score.Existing embeddings expose one vector per clip, leaving factor dependence fixed by pre-training and unavailable to users.
- Introduction: MERIT learns three independent projection heads for melody, rhythm, and timbre from a frozen encoder backbone.The framework seeks functional selectivity: each head should respond to its target factor while remaining invariant to the others.
- Introduction: MERIT addresses entangled real-world recordings with triplets constructed using conditional audio generation and source separation, varying only a single musical factor.This strategy enables a large-scale factor-controlled training set for learning independent boundaries.
- Introduction: The model uses multi-scale MERT features, shallow MLP projections, and Circle Loss while keeping the backbone frozen.The frozen backbone is intended to preserve foundation-model knowledge while making training fast.
- Introduction: Evaluation measures each head across all three factor test sets and uses zero-shot probes to test selectivity on independent real-world audio collections.The protocol evaluates both trained heads and monolithic representations such as CLAP and raw MERT.
2. Related Work
Prior audio and music similarity systems generally learn a single embedding or factor-specific similarity, whereas MERIT learns simultaneous melody, rhythm, and timbre similarities for full musical recordings. It uses a foundational embedding and returns the three dimensions as interpretable scores.
- General audio and music embeddings: CLAP, MuLan, and MERT produce rich music representations, but typically yield a single embedding per audio clip.CLAP and MuLan align audio with free-form text, while MERT uses masked language modelling with pitch, chroma, and beat objectives.
- Contrastive and metric learning for audio: Contrastive and metric learning has supported audio fingerprinting, speaker verification, and music tagging through objectives including triplet and Circle Loss.Most prior applications define a single similarity space.
- MERIT: MERIT extends metric learning to full musical recordings with a discriminative, shared-backbone approach and three simultaneous factor-specific spaces.The spaces are trained on independent, factor-controlled triplet datasets.
- Factor-specific retrieval: Prior factor-specific retrieval methods isolate dimensions such as melody, harmony, rhythm, or instrument characteristics rather than jointly modeling all three core musical factors.Cover song detection and MelodySim target melodic similarity, while beat and rhythm characterisation and instrument classification address complementary factors.
- MERIT: MERIT is presented as the first framework to learn melody, rhythm, and timbre similarities simultaneously while leveraging a foundational embedding and returning interpretable scores.This combination distinguishes it from prior single-space and narrowly factor-specific retrieval systems.
3. Method
MERIT learns separate melody, rhythm, and timbre similarity spaces from factor-specific triplets, using a shared frozen MERT representation and independently trained projection heads. At inference, separate factor indexes return candidates with explicit per-factor similarity scores.
- Factor-specific training data: MERIT defines separate triplet datasets for melody, rhythm, and timbre, where positives match the target factor and negatives differ on it.Each triplet is (A, Pf, N), with the anchor and positive similar on factor f but differing in other respects.
- Factor-specific training data: Melody positives preserve pitch contours through pYIN-conditioned JASCO generation, rhythm positives preserve grooves through drum-stem conditioning, and timbre positives share instrument labels without generation.The timbre dataset pairs stems from different songs with the same instrument label and uses a different-label stem as the negative.
- Shared representation: All heads use a shared frozen MERT-v1-330M encoder, extracting and pooling five transformer layers into a 5120-dimensional clip representation.The shared backbone removes encoder architecture as a confounding variable across factors.
- Factor heads: Each factor head is a shallow two-layer MLP with ℓ2-normalization, producing a 128-dimensional unit vector whose cosine similarity measures factor similarity.The three heads are Hmel, Hrhy, and Htim.
- Optimization: The three heads train independently with Circle Loss, while frozen-encoder embeddings are pre-computed and cached for head-only optimization.Circle Loss re-weights pairs by current similarity to maintain gradients on hard pairs.
- Inference and retrieval: At inference, each query is projected through all heads, separate FAISS indexes retrieve top-10 candidates per factor, and results expose Smel, Srhy, and Stim scores.The shared backbone encodes the query once before factor-specific projection and retrieval.
4. Experiments and Results … 4.3 Human Evaluation of Triplet Quality
MERIT’s experiments use factor-controlled triplets built from MoisesDB stems and JASCO-generated positives, train independent projection heads on frozen MERT embeddings, and validate triplet quality through expert ratings. Human listeners consistently rated each positive-pair type highest on its intended dimension, with varying inter-rater agreement and some perceptual coupling.
- 4.1 Datasets: Training triplets are derived from MoisesDB, a multitrack source-separation corpus providing per-song stems with instrument labels.
- 4.1 Datasets: Melody and rhythm positives use MoisesDB stems to seed JASCO synthesis, whereas timbre triplets are sampled directly from labelled MoisesDB stems.The timbre triplets use no generative step.
- 4.1 Datasets: Zero-shot evaluation uses three external corpora without JASCO outputs or MoisesDB material, including MUSDB18-HQ for instrument-class selectivity.MUSDB18-HQ provides four pure stem classes: vocals, drums, bass, and other; the mixture track is excluded.
- 4.2 Training Setup: Each projection head is trained independently with AdamW for 200 epochs while the MERT backbone remains frozen.Training uses lr = 10−3, weight decay 10−4, batch size 1024, cosine annealing to minimum lr = 10−5, and cached 5120-dim embeddings.
- 4.3 Human Evaluation of Triplet Quality: Human experts rated 10 positive pairs from each melody, rhythm, and timbre training dataset on three 0–100 similarity sliders.Each pair was rated by 15 music experts, with sliders labelled melody, rhythm, and timbre.
- 4.3 Human Evaluation of Triplet Quality: 60.0, 65.8, and 57.3 were the highest ratings for melody positives in melody, rhythm positives in rhythm, and timbre positives in timbre, respectively.The results confirm that raters assigned the highest similarity score to the dimension defining each pair.
- 4.3 Human Evaluation of Triplet Quality: Melody-positive pairs showed perceptual coupling, with rhythmic similarity rated 53.4 compared with melodic similarity rated 60.0.
- 4.3 Human Evaluation of Triplet Quality: Cronbach’s α was 0.615 for melody, 0.493 for rhythm, and 0.757 for timbre, indicating the strongest inter-rater agreement for timbre.The ratings suggest that the targeted perceptual distinctions are consistently recognizable to human listeners.
4.4 Internal Disentanglement
MERIT’s supervised heads achieve near-perfect accuracy on their intended musical factors while remaining much weaker on untargeted factors. Large cosine-distance margins and below-chance cross-factor behavior further indicate active geometric separation rather than weak or accidental performance.
- Evaluation: Triplet accuracy uses held-out melody, rhythm, and timbre triplets to test whether each head ranks the factor-matched positive closer than the negative.The test set contains 12,500 melody triplets, 12,500 rhythm triplets, and approximately 4,600 timbre triplets; chance is 50%.
- Factor-specific accuracy: The supervised heads reach 99.9% on melody, 100.0% on rhythm, and 99.6% on timbre, exceeding CLAP cosine by 21.4, 12.5, and 5.1 percentage points respectively.These are the diagonal, factor-matched test-set accuracies.
- Factor-specific accuracy: Off-diagonal performance is substantially lower, with Hmel dropping from 99.9% on melody to 58.4% on rhythm and 60.4% on timbre.Each off-diagonal evaluation tests a factor on which the head was not trained.
- Geometric evidence: The diagonal cosine-distance margins are +0.788 for Hmel and +0.862 for Hrhy, versus +0.034 and +0.048 for raw MERT on the same test sets.The much larger margins show that near-saturation accuracy reflects a strong separation effect rather than near-misses counted as wins.
- Geometric evidence: Hrhy reaches 47.7% on the melody test set, below chance, with dAP = 0.667 > dAN = 0.628 and a negative margin of −0.038.This indicates that the rhythm head places clips sharing melodic contour slightly farther apart, consistent with active suppression of the competing factor.
4.5 Zero-Shot External Probes
MERIT’s factor-specific heads generalize without fine-tuning to independently annotated, real-world audio collections unseen during training. Across probes, head selectivity tracks instrument class, rhythmic groove, and multi-factor cover-song similarity rather than absolute benchmark accuracy.
- External-probe protocol: Each head is evaluated zero-shot on three real-world audio collections unseen during training, with cross-factor selectivity assessed by which head ranks highest under each label structure.The diagnostic target is head ranking under each label structure, not absolute accuracy on any single benchmark.
- Instrument-class selectivity: Probe A tests timbre-related instrument-class selectivity using professionally separated MUSDB18-HQ stems from four pure classes: vocals, drums, bass, and other.Mixtures are excluded because they aggregate instruments and lack a single timbral identity.
- Rhythmic groove selectivity: 78.0% is the raw MERT accuracy that Hrhy exceeds by 10 percentage points on the Ballroom rhythmic-groove probe.The encoder is identical; training bends the projection toward rhythmic signature and away from competing acoustic cues.
4.6 Per-Class Selectivity Patterns
Per-class probe results reveal selective strengths and cross-factor artifacts across datasets. Timbre accuracy is highest for drums and lowest for vocals on MUSDB18-HQ, while rhythm accuracy is especially strong for vocals and meter-distinctive Ballroom classes.
- 4.6 Per-Class Selectivity Patterns: 90.7%: Htim accuracy peaks on drums in MUSDB18-HQ.Its accuracy is weakest on vocals at 70.1%, where pitched articulation introduces timbral variability across singers.
- 4.6 Per-Class Selectivity Patterns: 83.2%: Hrhy accuracy rises on vocals in MUSDB18-HQ.The passage identifies this as a known cross-factor artifact caused by the consistent rhythmic phrasing of pop vocal lines.
- 4.6 Per-Class Selectivity Patterns: 97.0%: Hrhy accuracy reaches its highest reported Ballroom Dataset value on Tango, followed by 94.4% on Viennese Waltz and 94.0% on standard Waltz.These three classes have maximally distinctive meters.
4.7 Score Fusion Strategies · 4.8 Qualitative Factor Profiles
MERIT’s factor heads can be fused to match or surpass the strongest isolated head on multiple probes, while per-factor profiles expose interpretable similarities and differences between cover pairs.
- 4.7 Score Fusion Strategies: The best single-head baseline is Htim (78.9%) for MUSDB18-HQ and Hrhy (88.0% and 69.9%) for Ballroom and Covers80, respectively.
- 4.7 Score Fusion Strategies: For Probe A and Probe C, ℓ2-normalised concatenation of [Hmel; Hrhy; Htim] matches or exceeds every isolated factor head.The result supports complementarity among the three factor projections.
- 4.8 Qualitative Factor Profiles: The qualitative analysis extracts per-factor similarities from trained heads for representative Covers80 cover pairs.Figure 4 presents MERIT scores for each cover pair as grouped per-factor profiles.
- 4.7 Score Fusion Strategies: Figure 3 compares Mean, W-Mean, Concat, and Product fusion strategies against the best single head across all three probes.W-Mean uses weights tuned on Covers80.
- 4.8 Qualitative Factor Profiles: “Let It Be” scores 0.97 each on melody, rhythm, and timbre for Beatles studio and live versions.The near-identical profile reflects a same-artist cover preserving the arrangement.
- 4.8 Qualitative Factor Profiles: “More Than Words” preserves melody (Smel = 0.93) and rhythm (Srhy = 0.95) but diverges in timbre (Stim = 0.79).The profile reflects Westlife’s pop-vocal arrangement versus Extreme’s acoustic-rock original.
4.9 Layer-Attribution Analysis
MERIT’s backbone permits direct attribution of each factor head’s reliance across MERT depths using row-normalised Frobenius norms of partitioned first-layer weights. The heads select depth patterns aligned with their intended factors: melody emphasizes deeper layers, rhythm shallower layers, and timbre spans multiple levels.
- Attribution method: Each head’s first-layer weight matrix is partitioned into five 512 × 1024 submatrices, whose Frobenius norms quantify attention to individual MERT layers.Row-normalisation converts these norms into the fraction of each head’s first-layer weight mass attending to each depth.
- Factor-specific depth patterns: Hmel concentrates on deeper MERT layer 23, consistent with pitch and melodic contour being encoded at higher abstraction levels.The layer-attribution heatmap shows the melody head’s strongest concentration at layer 23.
- Factor-specific depth patterns: Hrhy relies more on shallower MERT layers 3–6, which capture low-level temporal periodicity and rhythmic texture.This depth preference matches the rhythm head’s intended perceptual dimension.
- Factor-specific depth patterns: Htim shows a broader layer distribution, reflecting timbral identity expressed across multiple levels of spectral and acoustic representation.The timbre head therefore draws more diffusely across MERT depths than the melody or rhythm heads.
- Interpretation: The supervision signal alone selects MERT depths diagnostic of each factor, consistent with prior probing studies locating pitch and rhythm information at different encoder depths.This result indicates that factor-specific depth selection emerges from the training supervision.
5. Discussion
MERIT’s near-perfect diagonal scores reflect aligned supervision rather than conventional overfitting, but evaluation and factor definitions retain important limitations. Future work must address possible dataset correlations, expand beyond three factors, refine timbre labels, and improve JASCO conditioning fidelity.
- Evaluation: Near-100% diagonal scores reflect supervision aligned with information extractable by a shallow MLP on multi-layer MERT, not conventional overfitting.The held-out test split is folder-disjoint from training.
- Evaluation: Within-pipeline test data may inherit JASCO-borne correlations that let heads exploit shortcuts without learning the intended factor.Zero-shot probes address this concern on the evaluation side.
- Limitations: The decomposition covers only melody, rhythm, and timbre, while harmony and dynamics require additional conditioning channels.This restricts the current factorization scope.
- Limitations: MoisesDB instrument-class labels under-resolve within-class timbre variation, treating differently recorded acoustic guitars as identical positives.The limitation concerns timbre’s operationalization at instrument-class granularity.
- Limitations: JASCO’s conditioning fidelity sets a ceiling on melody and rhythm supervision quality.Improving conditioning fidelity is therefore a direct future-work target.
6. Conclusion
MERIT exposes melodic, rhythmic, and timbral similarity as three separable scores. Zero-shot probes show factor-specific behavior across instrument identity, rhythmic signatures, and cover pairs.
- MERIT represents melodic, rhythmic, and timbral similarity as three separable scores.
- On three zero-shot probes, the intended head is strongest for instrument-class identity, dance-style rhythmic signatures, and cover-pair similarity profiles.The probes use MUSDB18-HQ, Ballroom, and Covers80, respectively.