Source-linked AI summary
Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis
Thanathai Lertpetchpun, Yoonjeong Lee, Thanapat Trachu, Jihwan Lee, Tiantian Feng, Dani Byrd, Shrikanth Narayanan
TL;DR
Accent control through speaker embeddings is effective but opaque because embeddings encode traits beyond accent. This paper probes their interaction with phonological rules in American–British English TTS, introduces PSR, and finds that rules strengthen accent control while embeddings can attenuate or override them.
Problem
Speaker embeddings make accent representation opaque because they also encode traits such as timbre, emotion, and background noise.
Method
The study applies linguistically motivated American-to-British rules and uses PSR to measure how speaker embeddings preserve or override those transformations.
Results
Combining phonological rules with speaker embeddings strengthens accent attributes and preserves naturalness, while embeddings can reinforce or override explicit transformations.
Takeaways & Limitations
Coarse knowledge-driven rules provide interpretable accent-control levers and probes for evaluating disentanglement in speech generation.
Takeaways & Limitations
The study is limited to American-to-British mapping, categorical substitutions, automated metrics, and noise from the specific phoneme-recognition model.
Abstract
from arXiv · showhide
Many spoken languages, including English, exhibit wide variation in dialects and accents, making accent control an important capability for flexible text-to-speech (TTS) models. Current TTS systems typically generate accented speech by conditioning on speaker embeddings associated with specific accents. While effective, this approach offers limited interpretability and controllability, as embeddings also encode traits such as timbre and emotion. In this study, we analyze the interaction between speaker embeddings and linguistically motivated phonological rules in accented speech synthesis. Using American and British English as a case study, we implement rules for flapping, rhoticity, and vowel correspondences. We propose the phoneme shift rate (PSR), a novel metric quantifying how strongly embeddings preserve or override rule-based transformations. Experiments show that combining rules with embeddings yields more authentic accents, while embeddings can attenuate or overwrite rules, revealing entanglement between accent and speaker identity. Our findings highlight rules as a lever for accent control and a framework for evaluating disentanglement in speech generation.
1. INTRODUCTION
The study targets opaque accent control in TTS by probing speaker embeddings with linguistically motivated phonological rules. It introduces PSR to measure whether embeddings preserve or attenuate rule-based transformations.
- Speaker embeddings control accent but also encode timbre, emotion, and background noise, limiting interpretability and controllability.
- The study uses flapping, rhoticity, and vowel correspondences as salient contrasts between American and British English.These are intentionally treated as big-stroke transformations rather than complete models of dialect variation.
- PSR quantifies how much speaker embeddings preserve or overwrite rule-driven phoneme mappings.It captures partial reinforcement or attenuation when synthesized outputs drift toward the embedding-associated pronunciation.
- The proposed rules provide a controlled, interpretable probe for accent control and embedding disentanglement.The study presents them as complementary to data-driven TTS rather than a replacement for it.
2. PHONOLOGICAL RULES
The phonological-rule system maps American English phoneme sequences toward British English using one-to-one substitutions for flapping, rhoticity, and vowel correspondences. It preserves phoneme count while deliberately modeling only salient cross-accent differences.
- The system defines three substitution-rule families for mapping American to British English phoneme sequences.The families target flapping, rhoticity, and systematic vowel correspondences.
- Flapping maps American intervocalic /t/ realized as [R] to British [t], as in water.
- Rhoticity removes or vocalizes British post-vocalic /r/ that is retained in American English, as in car.
- Vowel correspondences map lexical-set patterns including TRAP, BATH, and GOAT.
- One-to-one substitutions preserve phoneme character count, isolating accent differences to mappings and speaker embeddings.The same pipeline can support additional accents by defining corresponding rule sets.
3. APPLICATION TO SPEECH GENERATION TASKS
The speech-generation pipeline derives British phoneme sequences from American input with the proposed rules, then synthesizes both conditions using Kokoro TTS. Fixed durations and preserved phoneme counts reduce timing and normalization confounds.
- The pipeline obtains American phonemes, applies phonological rules to derive British phonemes, and synthesizes speech from both inputs.
- Kokoro TTS receives a phoneme sequence, speaker embedding, and fixed phoneme durations in each condition.Phoneme character count and duration remain preserved across conditions.
4. METRICS
The study evaluates accent strength, phoneme realization, and naturalness, while PSR directly measures rule–embedding interaction. Its design also recognizes that synthesized phonological categories may surface partially rather than categorically.
- Accent strength is evaluated with Vox-Profile and Wav2Vec2Phoneme-based phoneme recognition, alongside naturalness evaluation.
- Vox-Profile Accent Classifier: Vox-Profile measures target-accent strength using classifier probabilities and cosine similarity to group-level reference accent embeddings.Its classifier distinguishes North American English, British Isles English, and English with other language backgrounds.
- Phoneme Shift Rate (PSR): PSR is defined as N2/N1, where N1 counts rule substitutions and N2 counts substitutions still needed after synthesizing and transcribing the target output.PSR = 0 indicates perfect rule compliance, whereas PSR = 1 indicates complete override under perfect TTS and recognition accuracy.
- Phoneme Shift Rate (PSR): PSR captures gradient rule realization, including outputs that partially drift toward the embedding-associated pronunciation.
5. DATASETS AND EXPERIMENTAL SETUPS
The study uses Kokoro-82M with speaker embeddings and phoneme sequences to test accent control through phonological rules. Experiments compare embeddings alone, individual rules, and full rule sets.
- Experimental Setup: The pretrained Kokoro-82M TTS model takes a speaker embedding and phoneme sequence as inputs.The model supports eight languages and provides 28 English speaker embeddings: 20 American and 8 British.
- Experimental Setup: The experiments manipulate both speaker embeddings and phonological rules to control accent strength.
- Experimental Setup: The model is identified as Kokoro-82M v0.194 and is publicly available through its Hugging Face repository.
- Experimental Setup: Rule analyses compare speaker embeddings alone, embeddings with one transformation rule, and embeddings with the full rule set.
- Experimental Setup: Full-rule experiments remove one rule at a time to isolate each rule’s contribution to accent strength.This design evaluates the combined benefit of data-driven and rule-driven approaches.
6. RESULTS AND DISCUSSION
Rules improve accent-related measures while preserving naturalness, but their effects depend on the speaker embedding and phonological context. PSR and phoneme-level analyses expose when embeddings preserve, reinforce, or override rule transformations.
- Results Overview: Three experiment sets vary rule counts, inspect synthesized phoneme sequences, and assess how individual speaker embeddings shape accented output.
- Naturalness: UTMOS remains around 4.4 for North American and 3.7 for British settings, with or without rules.The authors interpret this as no naturalness degradation from integrating phonological rules with speaker embeddings.
- Accent Strength: 86.5% North American probability falls to 58.8% and British probability rises to 17.3% when all rules are applied with a North American embedding.
- Accent Strength: 67.8% British probability with a British embedding increases to 78.4% when all British English substitution rules are applied.
- Accent Similarity: British similarity rises from -0.05 to 0.21 with North American embeddings and from 0.67 to 0.85 with British embeddings after rules are applied.
- Individual Rules: Vowel correspondences have the largest individual effect, raising British probability to 77.8% and reducing PSR to 0.693.Rhoticity raises British embedding-space similarity to 0.78, while flapping has minimal standalone impact but contributes additively in combinations.
- Phoneme Shift Rate: With British embeddings, PSR falls from 0.775 without rules to 0.628 with all rules, indicating stronger preservation of rule-driven changes.
- Phoneme-Level Analysis: Vowel correspondences produce the most changes, while phoneme-context effects show one embedding can represent different accent tendencies.Under North American embeddings, N2 increases for flapping but is significantly lower than N1 for rhoticity and vowels.
7. CONCLUSION AND FUTURE WORK
Phonological rules provide a coarse, effective, and interpretable lever for accent control while revealing how embeddings reinforce or override explicit transformations. Future work should broaden recognition architectures and add human evaluations.
- Conclusion: Phonological rules preserve naturalness, strengthen accent attributes, and reveal how speaker embeddings interact with explicit transformations.
- Conclusion: Rules complement speaker embeddings by providing interpretable, linguistically grounded modifications for accent control.
- Future Work: The study relies on automated metrics and noise from the specific phoneme recognition model used.
- Future Work: Future work will evaluate broader recognition architectures and human judgments.