Source-linked AI summary
PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech
Venkata Pushpak Teja Menta
TL;DR
Indic TTS can achieve strong intelligibility while retaining accent mismatches in phonemic features such as retroflexion, aspiration, vowel length, and Tamil zha. PSP measures these dimensions alongside distributional and prosodic distances, revealing that Telugu substantially increases retroflex collapse and that diagnostic dimensions can move differently under training changes.
Problem
Indic TTS evaluation measures intelligibility and naturalness but lacks decomposable evidence about native-like pronunciation of phonemic accent features.
Method
PSP benchmarks accent through six per-dimension measures using native-speaker prototypes, acoustic probes, and corpus-level audio and prosodic distances.
Results
Telugu retroflex collapse reaches 33–50%, while Praxy R5→R6 leaves it at 40%, improves FAD from 534 to 355, and worsens PSD from 14.1 to 61.7.
Takeaways & Limitations
PSP functions as a per-dimension diagnostic, showing whether interventions affect acoustic similarity, phonological fidelity, or prosody rather than producing one overall score.
Takeaways & Limitations
The per-phoneme probes depend on Indic forced aligners whose accuracy is limited, particularly for Telugu and Tamil, creating a language-dependent noise floor.
Abstract
from arXiv · showhide
Standard text-to-speech (TTS) evaluation measures intelligibility (WER, CER) and overall naturalness (MOS, UTMOS) but does not quantify accent. A synthesiser may score well on all four yet sound non-native on features that are phonemic in the target language. For Indic languages, these features include retroflex articulation, aspiration, vowel length, and the Tamil retroflex approximant (letter zha). We present PSP, the Phoneme Substitution Profile, an interpretable, per-phonological-dimension accent benchmark for Indic TTS. PSP decomposes accent into six complementary dimensions: retroflex collapse rate (RR), aspiration fidelity (AF), vowel-length fidelity (LF), Tamil-zha fidelity (ZF), Frechet Audio Distance (FAD), and prosodic signature divergence (PSD). The first four are measured via forced alignment plus native-speaker-centroid acoustic probes over Wav2Vec2-XLS-R layer-9 embeddings; the latter two are corpus-level distributional distances. In this v1 we benchmark four commercial and open-source systems (ElevenLabs v3, Cartesia Sonic-3, Sarvam Bulbul, Indic Parler-TTS) on Hindi, Telugu, and Tamil pilot sets, with a fifth system (Praxy Voice) included on all three languages, plus an R5->R6 case study on Telugu. Three findings: (i) retroflex collapse grows monotonically with phonological difficulty Hindi < Telugu < Tamil (~1%, ~40%, ~68%); (ii) PSP ordering diverges from WER ordering -- commercial WER-leaders do not uniformly lead on retroflex or prosodic fidelity; (iii) no single system is Pareto-optimal across all six dimensions. We release native reference centroids (500 clips per language), 1000-clip embeddings for FAD, 500-clip prosodic feature matrices for PSD, 300-utterance golden sets per language, scoring code under MIT, and centroids under CC-BY. Formal MOS-correlation is deferred to v2; v1 reports five internal-consistency signals plus a native-audio sanity check.
I. Introduction · II. Related work
The paper argues that Indic TTS systems can be highly intelligible yet retain non-native accent mismatches, motivating PSP as an interpretable, decomposable, per-dimension accent benchmark. It situates PSP among existing quality, distributional, phoneme-level, accent-similarity, and Indic speech evaluation approaches.
- I. Introduction: WER below 5% for Hindi and Tamil can coexist with residual accent mismatch despite correct word pronunciation.Subjective listening reveals that intelligibility does not guarantee native-like pronunciation.
- I. Introduction: Indic accent is decomposed into systematic contrasts, including retroflexes, aspiration, vowel length, and Tamil’s retroflex approximant.The proposed approach measures per-feature substitution rates with acoustic probes against native-speaker prototypes.
- I. Introduction: PSP formalises six per-dimension accent measures: RR, AF, LF, ZF, FAD, and PSD.The measures combine phonological dimensions with distributional and prosodic distances.
- I. Introduction: Five independent internal-consistency signals support PSP’s validity, while formal MOS calibration is deferred to v2.The paper also cites a companion diagnostic-loop case study where findings reinforce metric validity.
- II. Related work: WER, CER, UTMOS, and MOS prediction networks evaluate intelligibility or overall quality, but neither targets accent specifically.These metrics remain established tools for TTS quality evaluation.
- II. Related work: FAD compares synthesised and reference embedding distributions but yields a single scalar rather than an interpretable phonological feature profile.nPVI contributes one of PSP’s 5-D PSD feature dimensions by characterising rhythmic class.
- II. Related work: PSR is English-specific, rule-based, and scalar, whereas PSP is Indic-first, acoustic-probe-based, and decomposed into named per-phonological-dimension rates.The paper presents PSP as a conceptual sibling of PSR for a different evaluation setting.
- II. Related work: Existing accent-similarity metrics and Indic speech benchmarks provide pairwise scalars or listening-test pipelines, whereas PSP adds automatic per-dimension accent evaluation for Indic TTS.The benchmark is intended to complement resources using MUSHRA, cross-lingual evaluation, and forced-alignment infrastructure.
III. The phoneme substitution profile
PSP defines accent fidelity per phonological dimension using forced alignment and acoustic centroids, then complements these probes with corpus-level distributional distances. For Indic languages, it instantiates retroflex, aspiration, vowel-length, Tamil-zha, FAD, and prosodic dimensions without relying on high-quality target-language ASR.
- Formal definition: PSP defines each dimension through native and substitute phoneme sets plus an acoustic embedding, measuring generated-phoneme fidelity against native-speaker and substitute-speaker centroids.Forced alignment identifies each phoneme’s time span, and fidelity uses rectified cosine similarity.
- Indic dimensions: Four per-phoneme probes cover retroflex substitution, aspiration, vowel length, and Tamil zha, with aspiration language coverage varying and zha restricted to Tamil.DRR contrasts retroflex with non-retroflex sets; DAF is Hindi-primary, Telugu-sparse, and Tamil N/A; DLF compares long/short duration ratios; DZF is Tamil-only.
- Indic dimensions: DFAD measures Fréchet distance between generated and native XLS-R layer-9 distributions, capturing timbre, co-articulation, and phoneme-frequency signals missed by per-phoneme probes.It is computed once per system-language pair rather than per utterance.
- Indic dimensions: DPSD measures Fréchet distance in a 5-D prosodic space spanning pitch range, log-F0 mean, speech rate, nPVI, and log-duration.Conjunct epenthesis detection (CERconj) is scaffolded but not evaluated.
- Rationale for acoustic probes: Acoustic-space probes measure acoustics rather than transcriptions and avoid dependence on high-quality target-language ASR, which remains error-prone for Indic languages.The rationale specifically uses cosine similarity of XLS-R embeddings instead of ASR hypothesis matching or rule-based transformations.
IV. Implementation … C. Released artifacts
The implementation builds native acoustic centroids and corpus-level embedding resources, then scores retroflex fidelity through forced alignment and layer-9 XLS-R representations. Released code is MIT-licensed, while centroids are available under CC-BY.
- A. Centroid construction: A. Centroid construction: N=500 native clips per language are sampled from IndicTTS for Telugu and Tamil and Rasa for Hindi.Only studio-recorded utterances with confirmed native speakers are selected.
- A. Centroid construction: A. Centroid construction: Sampling is uniform across speakers, with ≥20 distinct speakers for Telugu and Tamil and ≥40 for Hindi.A per-speaker cap of 25 clips limits voice identity from dominating the centroid.
- A. Centroid construction: A. Centroid construction: Corpus-level FAD and PSD use N=1000 utterance-level XLS-R embeddings per language.The passage specifies these additional embeddings for corpus-level metrics.
- A. Centroid construction: A. Centroid construction: Language-specific CTC aligners identify canonical grapheme matches before layer-9 Wav2Vec2-XLS-R-300M embedding spans are added to per-phoneme bags.Native centroids are computed as the mean of each phoneme’s embedding bag.
- B. Scoring pipeline: B. Scoring pipeline: Forced alignment over CTC emissions and grapheme sequences extracts layer-9 XLS-R embeddings for each expected-retroflex span.Per-position fidelity is averaged into utterance-level PSP-RR, while corpus-level scores weight utterances by expected retroflex count.
- C. Released artifacts: C. Released artifacts: Evaluation code is open-source under MIT, centroids are licensed CC-BY, and a public leaderboard is forthcoming.The repository is github.com/praxelhq/psp-eval.
V. Calibration
Calibration in v1 is limited to internal-consistency signals and a native-audio sanity check; formal MOS correlation is deferred to v2. The signals support PSP’s phonological validity while exposing language-specific noise in per-phoneme probes.
- Internal-consistency signals: ∼1% on Hindi, ∼40% on Telugu, and ∼68% on Tamil: retroflex collapse rises monotonically with established phonological difficulty.This gradient is independent of per-system claims and matches the community-established maturity hierarchy.
- Internal-consistency signals: Indic-focused Sarvam Bulbul and Parler-TTS consistently outperform ElevenLabs v3 and Cartesia Sonic-3 on per-phoneme PSP dimensions, especially for Telugu and Tamil.This holds even when Indic-focused systems do not have the lowest WER.
- Internal-consistency signals: WER and PSP ordering diverge: ElevenLabs v3 records Hindi WER 0.006 but second-place FAD, while PSD detects ElevenLabs Telugu narrow-pitch-range failure.Cartesia’s second-place Telugu WER coincides with worst-place retroflex and FAD, showing dissociation between intelligibility and phonological or distributional metrics.
- Internal-consistency signals: On Tamil, Parler-TTS wins four of five PSP dimensions while Sarvam wins FAD only, so no single system dominates every phonological sub-dimension.The per-dimension decomposition preserves distinctions that a single scalar would lose.
- Native-audio sanity check: FAD 32–44 and PSD 2–6 on native audio are 5–100× lower than corresponding commercial-TTS values, while Telugu and Tamil per-phoneme probes show 43–86% apparent collapse.Hindi native audio achieves perfect RR and AF scores of 1.0; the Telugu and Tamil apparent collapse is attributed to coarser aligners, allophonic variation, and τ = 0.5 threshold strictness.
VI. Experiments · A. Systems and test sets · B. Hindi results: a mature target
The experiments benchmark commercial and open-source TTS systems on Indic pilot sets using per-phoneme PSP dimensions plus corpus-level FAD and PSD. On Hindi, all systems show near-zero retroflex collapse and perfect aspiration, while FAD distinguishes systems in an ordering that diverges from WER.
- VI. Experiments: The benchmark covers Hindi, Telugu, and Tamil using 10-utterance pilot sets, with two commercial-system voices or one Praxy Voice R5 voice.Each utterance is scored on applicable per-phoneme PSP dimensions; FAD and PSD are computed once per system-language pair against native reference distributions.
- A. Systems and test sets: Indic Parler-TTS, Praxy Voice R5/R6, ElevenLabs v3, Cartesia Sonic-3, and Sarvam Bulbul comprise the evaluated open-source and commercial systems.Praxy Voice R5 and R6 are evaluated on Telugu only; R5 uses approximately 85 hours of training data, while R6 uses approximately 1,220 hours.
- B. Hindi results: a mature target: Hindi retroflex and aspiration collapse stays within 0–4.5% across all four systems.ElevenLabs, Cartesia, and Sarvam have zero collapses across 22 retroflex and 18 aspirated tokens, while Indic Parler-TTS has one retroflex-collapse outlier.
- B. Hindi results: a mature target: Hindi aspiration is perfect across all four systems.The passage characterizes modern Hindi TTS, commercial and open-source, as having largely solved core phonological articulation.
- B. Hindi results: a mature target: Sarvam leads Hindi FAD at 211.8, followed by ElevenLabs at 227.5, Indic Parler at 248.4, and Cartesia at 267.4.The reported spread between the closest and farthest systems is 56 points.
- B. Hindi results: a mature target: ElevenLabs has the lowest prior-benchmark Hindi WER but ranks second on FAD, while Cartesia is second-lowest on published WER but last on FAD.This dissociation illustrates that distributional accent properties measured by PSP are not captured by WER alone.
C. Telugu results: the real difficulty
Telugu exposes substantial retroflex and prosodic weaknesses that WER alone misses, with retroflex collapse reaching 50% and ElevenLabs showing unusually high PSD. Praxy’s speaker-reference conditioning improves several PSP dimensions without changing LLM-WER.
- Telugu results: Telugu retroflex collapse ranges from 33% for Sarvam and Parler to 50% for Cartesia, while ElevenLabs and both Praxy checkpoints reach 40%.Telugu FAD values broadly follow retroflex-collapse ordering, with a 2.5× spread.
- Flatness failure: ElevenLabs Telugu PSD is 154 versus 11 for Sarvam and 10 for Parler, revealing a prosodic failure that WER and retroflex collapse miss.Its log-F0 range is 0.87 versus native 1.44, and nPVI is 92 versus native 107.
- Praxy R5 →R6 training delta: A 10× training-data increase from 85 hr to 1,220 hr leaves Praxy Telugu retroflex collapse unchanged at 40% →40%.The passage attributes this to LoRA-ont3 leaving the acoustic generator frozen.
- Voice-prompt recovery: Praxy R6 voice prompting reduces retroflex collapse from 40% to 33% with a Cartesia reference or 26.7% with a Sarvam reference.With Sarvam prompting, the 26.7% result falls below every measured commercial system.
- Voice-prompt recovery: Voice prompting reduces PSD from 61.7 to 26.5 with Cartesia or 13.1 with Sarvam, while LLM-WER remains 0.033–0.034.FAD changes from 355 to 291 with the Sarvam reference, closing the R6-to-Sarvam gap by 61%.
- Voice-prompt recovery: The reference-conditioned improvements support PSP’s thesis and motivate Praxy’s BYOR mode, which uses a 9 s Telugu speaker reference clip.The token path remains LoRA-adapted while users provide the speaker reference.
D. Tamil results: the hardest Indic language · E. Cross-language synthesis
Tamil is the hardest target, with severe retroflex, Tamil-zha, and vowel-length failures across systems. Across languages, PSP reveals difficulty-dependent degradation and metric orderings that diverge from intelligibility, with no universal system winner.
- D. Tamil results: the hardest Indic language: 64–70% retroflex collapse and 85.7% Tamil-zha collapse make Tamil the study’s most severe target; length fidelity falls to 0.13–0.30 versus a ∼1.90 native prior.Only 1 in 7 Tamil-zha tokens survives for three of four systems, and no system preserves the long/short vowel contrast.
- D. Tamil results: the hardest Indic language: Indic Parler-TTS wins four of five Tamil PSP dimensions—RR, ZF, LF, and PSD—while Sarvam leads FAD, so no system dominates overall.The per-dimension decomposition exposes complementary strengths that a single aggregate score would conceal.
- E. Cross-language synthesis: Cartesia’s FAD grows 51% from Hindi to Tamil, whereas Sarvam and Parler maintain or slightly improve FAD; ElevenLabs’ PSD explodes as language difficulty rises.The cross-language pattern indicates that Indic-first systems generalise better than Western-built commercial systems on these distributional measures.
- E. Cross-language synthesis: 33% retroflex collapse is shared by Parler-TTS and Sarvam on Telugu, while Cartesia reaches 50%, despite ElevenLabs, Cartesia, and Sarvam all achieving sub-5% LLM-WER.Telugu metrics produce different winners: Sarvam leads FAD, Parler leads PSD, and Praxy R6 reaches 100% intent-preservation.
- E. Cross-language synthesis: 33–50% Telugu retroflex collapse contrasts with 0–4.5% Hindi collapse for the same four commercial systems, showing a language-specific accent gap that WER does not surface.Hindi is essentially native-quality on PSP’s phonological dimensions, whereas Telugu is not.
- E. Cross-language synthesis: On Hindi, ElevenLabs leads WER at 0.006 but ranks second on FAD, while Sarvam has FAD 211.8 yet ranks third on WER.Cartesia illustrates the opposite mismatch: second on WER at 0.025 but last on FAD at 267.4.
- E. Cross-language synthesis: Praxy R6 transfers to Tamil without retraining, reaching FAD 276, PSD 71, 69% retroflex collapse, and 71% Tamil-zha collapse under native-language reference prompting.On Hindi, vanilla Chatterbox routing recovers LLM-WER 0.025, intent-preservation 1.0, 0% retroflex collapse, and 0% aspiration collapse.
- E. Cross-language synthesis: ∼1% Hindi, ∼40% Telugu, and ∼68% Tamil mean retroflex collapse across four commercial systems follows the difficulty ranking Hindi < Telugu < Tamil.This monotonic trajectory is presented as a metric-validity signal independent of system comparison.
VII. Discussion and limitations
PSP is intended as a per-dimension diagnostic that guides targeted interventions, while its interpretation is constrained by alignment noise, prototype centroids, language-dependent applicability, shared corpora, and unnormalised PSD features.
- Intended workflow: An R5→R6 data scale-up closed FAD on Telugu by 34% but opened PSD by ∼4×, indicating a prosodic-conditioning intervention rather than retraining.The workflow uses cell changes to route the next intervention; the case study pointed to voice-prompt recovery at inference time.
- Forced-alignment dependency: Forced-alignment accuracy is the dominant native-audio noise-floor source, especially for Telugu and Tamil, whose best public aligners are community fine-tunes.Hindi alignment uses larger and cleaner training data than Telugu/Tamil alignment.
- Other limitations: Per-phoneme probes are coarser than MFA-trained native acoustic models and should be interpreted as relative system rankings on Telugu and Tamil.Absolute interpretation is supported for Hindi and for FAD / PSD across all three languages.
- Other limitations: Tamil AF is inapplicable and Telugu AF is sparse because of their phonological realities, while both conditions affect metric interpretation.Tamil has no phonemic aspirated stops, and aspirated forms are used infrequently in modern Telugu speech.
- Threats to validity: Centroid and test corpora share IndicTTS / Rasa speaker pools, and PSD currently uses unnormalised features spanning disparate scales.A z-scored PSD variant and truly disjoint native sets are planned for v2; nPVI has order 102 and log-F0 has order 100.
VIII. Conclusion
PSP is an interpretable, six-dimension benchmark for evaluating Indic TTS accent by phonological dimension. The v1 release provides scoring code, native-speaker references, distributional reference data, and held-out golden test sets.
- Benchmark and release: PSP evaluates Indic TTS accent with six-dimension scoring code and native-speaker centroids for Telugu, Hindi, and Tamil.The benchmark is designed as an interpretable per-phonological-dimension measure.
- Benchmark and release: The release includes 1000-clip FAD reference embeddings, 500-clip PSD reference feature matrices, and 300-utterance held-out golden test sets per language.These resources support corpus-level distance measures and held-out evaluation.
- Benchmark and release: The v1 preprint benchmarks four commercial and open-source systems, plus in-progress Praxy Voice on Telugu, and reports five internal-consistency signals.Praxy Voice is included only on Telugu in this release.
Ethics and Reproducibility
PSP is released with the code, data artifacts, and scripts needed to reproduce v1 results, while commercial audio must be regenerated by users. The benchmark is framed as an engineering measure of phonological similarity, not a prescriptive judgement on human accents.
- Reproducibility: The full scoring and bootstrap code, native centroid pickles, 300-utterance held-out test texts, and benchmark_results.json are released under MIT.Commercially generated audio from ElevenLabs, Cartesia, and Sarvam is not redistributable, so users regenerate it under their own accounts using the provided scripts.
- Ethics: PSP treats “native-like” similarity as an engineering target for TTS developers, not a value judgement that ranks human accents.The benchmark is intended to support native-listener intelligibility and naturalness, while recognizing non-native accents and L2 speech as legitimate varieties.