Source-linked AI summary

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

Srija Anand, Ashwin Sankar, Ishvinder Sethi, Aaditya Pareek, Kartik Rajput, Gaurav Yadav, Nikhil Narasimhan, Adish Pandya, Deepon Halder, Mohammed Safi Ur Rahman Khan, Praveen S, Shobhit Banga, Mitesh M Khapra

arXiv:2604.21481v2cs.CL

TL;DR

Evaluating multilingual TTS is difficult because linguistic diversity and multidimensional speech perception create high variance in pairwise judgments. The paper addresses this with a controlled, multidimensional evaluation framework and finds that expressiveness and intelligibility primarily drive listener preference once basic robustness is established.

  • Problem

    Linguistic diversity and multidimensional speech perception make rigorous, scalable evaluation of multilingual TTS challenging.

  • Method

    The study evaluates 7 TTS systems using 5,357 sentences across 10 Indic languages, over 120K pairwise comparisons from 1900+ native raters, six perceptual axes, and Bradley-Terry modeling.

  • Results

    Expressiveness and intelligibility are the strongest predictors of overall preference, while GEMINI 2.5 PRO TTS ranks first overall and in 9 of 10 languages.

  • Takeaways & Limitations

    Multilingual TTS comparison benefits from broad sentence coverage, multidimensional judgments, and reliability analysis alongside aggregate leaderboard rankings.

Abstract

from arXiv · show

Crowdsourced pairwise evaluation has emerged as a scalable approach for assessing foundation models. However, applying it to Text to Speech(TTS) introduces high variance due to linguistic diversity and multidimensional nature of speech perception. We present a controlled multidimensional pairwise evaluation framework for multilingual TTS that combines linguistic control with perceptually grounded annotation. Using 5K+ native and code-mixed sentences across 10 Indic languages, we evaluate 7 state-of-the-art TTS systems and collect over 120K pairwise comparisons from over 1900 native raters. In addition to overall preference, raters provide judgments across 6 perceptual dimensions: intelligibility, expressiveness, voice quality, liveliness, noise, and hallucinations. Using Bradley-Terry modeling, we construct a multilingual leaderboard, interpret human preference using SHAP analysis and analyze leaderboard reliability alongside model strengths and trade-offs across perceptual dimensions.

1. Introduction

Multilingual TTS evaluation in India is difficult because speech is linguistically diverse and commonly includes code-mixing and other deployment-relevant forms. The paper addresses limitations of absolute-rating studies with a controlled, multidimensional pairwise framework.

  • India’s linguistic diversity and widespread bilingualism produce speech containing code-mixing, domain vocabulary, numerals, and alphanumeric expressions.
  • Existing TTS evaluations remain limited in scale, language coverage, or diagnostic depth despite advances in neural synthesis.
  • MOS-style absolute ratings are sensitive to individual rater calibration, whereas pairwise comparisons support direct judgments and Bradley-Terry rankings.
  • Fine-grained perceptual feedback is needed to identify the factors driving human preferences between systems.
  • The framework combines a 10-language benchmark, evaluation of 7 TTS systems, over 120K comparisons, and feedback across six perceptual dimensions.

2. Related Work

Prior TTS studies commonly use subjective listening tests or pairwise preference methods, but multilingual evaluations across Indic languages remain relatively scarce. Existing approaches also trade off scalability against perceptual detail.

  • MOS, CMOS, and MUSHRA remain standard subjective TTS evaluation methods but often provide aggregate scores that obscure perceptual factors.
  • Multidimensional extensions capture richer perceptual signals but increase evaluation complexity and limit scalability across multiple systems.
  • Large-scale multilingual evaluations across Indic languages remain relatively scarce.

3. Evaluation Framework

The evaluation framework combines a multilingual, deployment-oriented benchmark with trained native raters, a two-stage multidimensional annotation protocol, and Bradley-Terry ranking with bootstrap uncertainty.

  • Benchmark Construction: The benchmark contains 5,357 sentences across 10 Indian languages and 16 deployment-relevant domains.
  • Benchmark Construction: Three subsets test normalized text, raw symbols, and code-mixed multilingual usage across varied utterance durations.
  • Rater Recruitment: Raters pass auditory screening, justify choices using study criteria, and receive training before annotation.
  • Annotation Protocol: In the first annotation phase, raters lock holistic preferences between two anonymous randomized samples before rating the same pair on six perceptual axes.
  • Ranking Methodology: Pairwise preferences are converted into a leaderboard with maximum-likelihood Bradley-Terry modeling and 500 bootstrap refits for 95% confidence intervals.

4. Results

Across large-scale multilingual evaluation, GEMINI 2.5 PRO TTS leads overall, across nearly all languages and domains, while rankings vary across input conditions and perceptual dimensions. Preference is driven most strongly by expressiveness and intelligibility, and leaderboard ordering stabilizes before score estimates become precise.

  • Overall and Language-wise Rankings: GEMINI 2.5 PRO TTS ranks first overall at 1128.53 ± 3, followed by ELEVEN LABS V3 at 1056.28 ± 2 and SONIC 3 at 1050.83 ± 3.ELEVEN LABS V3 and SONIC 3 are statistically indistinguishable.
  • Overall and Language-wise Rankings: GEMINI 2.5 PRO TTS ranks first in 9 of 10 languages, while INDIC F5 remains at or near the bottom across languages.Marathi is the exception, where GEMINI 2.5 PRO TTS is near parity with ELEVEN LABS V3.
  • Domain Sensitivity: GEMINI 2.5 PRO TTS ranks first in all 16 domains, but SPEECH 2.8 HD leads Stress Test and systems tie in Tongue Twisters.ELEVEN LABS V3, SONIC 3, and BULBUL V3 BETA show smaller rank differences that vary by domain.
  • Input-Type Sensitivity: GEMINI 2.5 PRO TTS remains first across normalized, symbolic, and code-mixed inputs, although BULBUL V3 BETA performs relatively better on symbolic inputs.Condition-wise evaluation exposes complementary strengths and weaknesses despite modest overall rank changes.
  • Perceptual Dimensions: GEMINI 2.5 PRO TTS performs consistently across six perceptual axes, whereas other systems show trade-offs in expressiveness, liveliness, intelligibility, voice quality, and robustness.Higher hallucination and noise-axis scores indicate fewer artifacts and cleaner audio.
  • Preference Drivers: Expressiveness and intelligibility are the strongest predictors of overall preference, followed by liveliness and voice quality; hallucinations and noise contribute less.The lower contribution of hallucinations and noise is attributed to most systems already performing relatively strongly on those axes.
  • Leaderboard Reliability: 100–200 raters typically produce stable rankings, but non-trivial confidence intervals mean nearby score differences require caution.With 5 systems, approximately 200 raters exceed ρ ≥0.95; with 7 systems, the threshold is reached at around 100 raters.

5. Conclusion

The paper presents a controlled multidimensional pairwise evaluation framework for multilingual TTS and combines leaderboard construction with diagnostic preference analysis and reliability assessment.

  • The framework evaluates multilingual TTS using 5.3K sentences across 10 Indic languages and over 120K pairwise judgments from 1900+ vetted native raters.
  • Bradley-Terry modeling is used to construct a multilingual leaderboard from the pairwise judgments.
  • Six perceptual axes provide fine-grained diagnostic analysis, and axis-level judgments strongly predict overall preference across languages.
  • Once basic robustness to noise and hallucinations is ensured, expressiveness and intelligibility drive listener choice.
  • Stable rankings can be obtained with moderate rater counts when sentence coverage is sufficiently broad.

7. Generative AI Use Disclosure

The authors disclose that generative AI tools were used only for language polishing and editing during manuscript preparation.

  • Generative AI tools assisted with clarity, grammar, and conciseness but did not generate results, analyses, figures, or scientific conclusions.
Loading 2604.21481v2…