Source-linked AI summary

Vocal Music under Phoneme-Conditional Analysis

Hayoon Kim, Kyogu Lee

arXiv:2608.30823v1cs.SDcs.CL

TL;DR

The paper asks whether language-specific vocal differences are measurable and traceable to phonemes. It introduces within-song phoneme-conditional comparisons across nine languages and finds that song-level contrast profiles identify language at 85.5% balanced accuracy, while phoneme-level accumulation remains causally unresolved.

  • Problem

    Existing work leaves unresolved whether vocal differences across languages arise from specific phonemes, partly because rhythm, tone, and timbre are often studied separately or at coarse resolution.

  • Method

    Phoneme-conditional analysis compares typologically distinctive marker syllables with matched non-marker controls within the same song, holding singer, melody, and genre constant.

  • Results

    85.5% balanced accuracy was achieved in nine-way language classification with artist-grouped folds, using song-level profiles built from five acoustic dimensions.

  • Takeaways & Limitations

    Phonological structure leaves systematic measurable traces in singing, especially through voice quality and spectral tilt rather than formant transitions.

  • Takeaways & Limitations

    The nine-language chart corpus, alignment errors, lyric-quality filtering, residual accompaniment, and confounding between phonology and tradition constrain interpretation of the results.

Abstract

from arXiv · show

The vocal music of each language carries a distinctive sonic identity, even without instrumental accompaniment. We ask whether these differences are measurable and traceable to specific phonemes. To tackle this question, we introduce phoneme-conditional analysis, which isolates the acoustic effect of typologically distinctive phonemes by comparing marker syllables against matched non-marker controls within the same song, holding singer, melody, and genre constant. Across nine typologically diverse languages and thousands of songs, we measure effects along five acoustic dimensions. Song-level profiles built from these effects identify the language of an unaccompanied vocal at 85.5% balanced accuracy in a nine-way classification with folds grouped by artist; whether the separability arises by accumulation of the phoneme-local effects themselves is left open. Our findings suggest that phonological structure leaves systematic and measurable traces in how each language is sung.

1. INTRODUCTION

The paper argues that phonological structure contributes to language-specific vocal sound and proposes testing phoneme-local effects within songs. It frames this approach as a response to gaps in resolution, phonological coverage, and cross-dimensional integration.

  • Arabic and Japanese singing differ in pitch, timing, and vocal quality across genres, suggesting language-specific rather than purely conventional origins.
  • Existing research links language to rhythm, melodic contour, and text-setting, but often studies these relationships at coarse or isolated levels.
  • Three gaps motivate the study: symbolic rhythm measures miss sub-segmental timing, consonant inventories remain understudied, and acoustic dimensions are rarely integrated.
  • H1 proposes that individual phonemes physically affect the musical realization of their host syllables and can be isolated with within-song matched controls.
  • 85.5% nine-way classification accuracy was achieved across 9 languages using artist-grouped folds and song-level acoustic signatures.
  • The framework separates phoneme-local mechanics, language-wide convention, and tradition-level persistence, while testing H1 rigorously and treating H2 preliminarily.

2. RELATED WORKS

Related work shows that language–music links appear in rhythm, melody, and text setting, but vocal studies often lose these distinctions through notation or pitch-centered analysis. The paper therefore targets phoneme-local acoustic mechanisms and their aggregation into language-specific vocal identity.

  • nPVI is the standard metric for comparing linguistic and musical rhythm, but notation-based vocal analyses can weaken or reverse speech–music differences.
  • Forced-aligned phoneme durations are used to measure rhythmic variability below the coarse resolution of symbolic notation.
  • Lexical tone can constrain melodic contour, while onset consonants can displace the perceptual center of a syllable.
  • Prior corpora establish whole-utterance and whole-piece regularities, but leave open which phonological properties produce them and how local effects aggregate.
  • The study uses established acoustic signatures including F3 lowering, F2 lowering, VOT and f0-onset separation, nasal-vowel lengthening, and geminate closure ratios.

3. METHOD

The method samples typologically diverse multilingual chart music, identifies phonological markers, aligns and filters vocals, and compares markers with matched within-song controls. Effects are summarized statistically across acoustic dimensions and used for artist-grouped language classification.

  • Nine languages span six families and three prosodic types, with rare, acoustically measurable phonological markers selected from typological resources.
  • 1,000–2,100 songs per language are sampled from Spotify national daily charts and processed through identification, separation, transcription, alignment, and deduplication.
  • The pipeline combines vocal separation, language-specific grapheme-to-phoneme conversion, and two parallel CTC aligners whose outputs are averaged.
  • The alignment-agreement gate excludes tracks whose model disagreement exceeds 200 ms, selecting for transcript completeness rather than performer identity.
  • Alignment error compresses duration contrasts: pairing within true-duration bins recovers 34% of the true difference, so reported millisecond magnitudes are lower bounds.
  • Marker syllables are paired with same-song controls using phone-type matching and temporal proximity, with controls reused at most three times.
  • Paired Cohen’s d is computed per language × dimension using the maximum absolute feature effect, with permutation tests and Bonferroni correction across 50 cells.
  • Song-level 35-dimensional vectors are classified with a 200-tree Random Forest using 5-fold artist-grouped cross-validation and undersampling.

4. RESULTS

Across languages, phoneme-conditioned effects produce distinct acoustic profiles spanning multiple dimensions, while sub-segmental timing and marker–control contrasts support language identification. Speech-to-singing preservation is channel-dependent, and the aggregation hypothesis remains unsupported with nine languages.

  • Significant effects appeared in at least three acoustic dimensions for 9 of 10 marker sets, with Hindi as the exception.Timbre and voice quality generally had the largest effect sizes, while formant transitions were best preserved only for retroflexion.
  • d = −0.80 was the corpus’s largest single Table 3 effect, from Mandarin retroflex markers on MFCC-1 over 25,899 pairs.The tonal marker set’s largest effect was less than half as large, at −0.36; Mandarin tonal markers instead separated most strongly through melisma.
  • Korean ranked as the strongest consonant-marker language, followed by Arabic, Turkish, and English, with the ranking largely determined by timbre.Six marker sets were significant in all five modules; Japanese reached significance in three, including melisma avoidance and shortening.
  • Aspirated–lax and tense–lax Korean contrasts were nearly identical, whereas aspirated–tense differences were minimal.The singing pattern follows the contemporary speech partition in which F0 onset has largely replaced VOT as the primary laryngeal cue, while retaining approximately 30% of its speech magnitude.
  • No durational contrast reached half its speech baseline, while preservation varied by channel and some duration effects reversed sign.Mandarin tonal lengthening retained approximately 30% of its speech effect, French nasal vowels approximately 19%, and Japanese geminates remained shortened.
  • 737 and 533 were the Kruskal–Wallis H statistics for song-level nPVI and melisma rate across nine languages, with language explaining ε2 = 0.15 and 0.11 of rank variance.These measures used forced-aligned phoneme durations rather than notated values, revealing language information unavailable to notation-based rhythm measures.
  • 85.5% balanced accuracy identified nine languages from 35-dimensional song-level vectors with artist-grouped folds, versus 11.1% chance.Arabic and Turkish were most identifiable at 95% and 94%, Japanese least at 73%; block ablation located separability in marker–control contrasts.
  • Neither marker rate nor effect magnitude predicted per-language recall, leaving H2’s aggregation prediction unsupported at n = 9.The classifier can use distinctive combinations of small contrasts, while residual accompaniment remains an uncontrolled cue after source separation.

5. CONCLUSION

The study finds that phoneme-local acoustic effects yield language-specific vocal signatures, while the broader emergence of language-level conventions remains causally unresolved. Classification separates languages from performers, but the nine-language scope and corpus constraints limit generalization.

  • Classification: 85.5% balanced accuracy was achieved in nine-way language classification with artist-grouped folds, indicating separation beyond performer identity.The analysis used song-level vectors built from phoneme–control contrasts; grouping by artist reduced accuracy by 1.2 points versus song-level splits.
  • Acoustic channels: The dominant singing cues are voice quality and spectral shape rather than formant transitions, pitch, or timing alone.The findings describe voice quality as attenuated but not eliminated, with phoneme identity carried more by source-spectrum properties than vocal-tract resonance.
  • Interpretation: The study establishes separability, not the accumulation predicted by H2, because marker frequency and effect magnitude do not predict per-language accuracy at n=9.Thus, the classification result does not by itself show that frequent local effects accumulate into language-wide conventions.
  • Interpretation: Within-song matched pairs hold singer, genre, and melody constant, supporting phoneme-local effects as consequences of vocal-tract physics rather than musical tradition.At the language level, however, cumulative articulation and musical tradition cannot be disentangled with the available chart data.
  • Limitations: The nine-language chart corpus excludes click languages, ejective-rich inventories, and richer tone systems, while residual accompaniment and lyric-quality failures constrain measurement.The authors also note that the data tie each language to one or two countries and that acoustics were measured without a perceptual audibility test.

7. AI USAGE STATEMENT

Generative AI tools assisted with code development and language polishing. The authors produced the research design, experiments, analysis, and conclusions.

  • Generative AI tools were used for code development assistance and language polishing, while the authors produced the research design, experiments, analysis, and conclusions.
Loading 2608.30823v1…