Source-linked AI summary

Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios

Changhao Pan, Rui Yang, Han Wang, Zhuan Zhou, Xuming He, Wenxiang Guo, Ziyue Jiang, Ruiqi Li, Yu Zhang, Chenyuhao Wen, Ke Lei, Xiang Yin, Jingyu Lu, Zhiyuan Zhu, Zhou Zhao

arXiv:2605.28618v1eess.AS

TL;DR

Long-form speech models remain difficult to evaluate across complex, diverse scenarios. SwanBench-Speech introduces a holistic benchmark and finds substantial gaps in consistency, coherence, and expressive hierarchy, especially in highly expressive settings.

  • Problem

    Existing long-form speech evaluations cover limited domains or single-speaker settings, leaving performance in complex scenarios underexplored.

  • Method

    SwanBench-Speech evaluates long-form TTS with 1,101 instances across 17 scenarios using seven complementary, human-aligned metric dimensions.

  • Results

    Current models rival human recordings in fidelity and accuracy but show substantial gaps in reverb consistency, prosodic coherence, and expressive hierarchy, especially in highly expressive scenarios.

  • Takeaways & Limitations

    SwanBench-Speech provides a standardized testbed for analyzing and advancing robust, immersive long-form speech synthesis.

  • Takeaways & Limitations

    The benchmark is limited to Chinese and English, uses prompt speech from 20 speakers, and lacks robust automated assessment of deep semantic emotional and stylistic transitions.

Abstract

from arXiv · show

Recent advances in speech generation have enabled high-fidelity synthesis, yet systematic evaluation of models under long-context conditions remains largely underexplored. A comprehensive evaluation benchmark for long-form speech is indispensable for two reasons: 1) existing test scenarios are often confined to limited domains, creating a significant gap with the diverse downstream applications; 2) existing metrics overlook critical long-text factors such as consistency and coherence, failing to generalize reliably. To this end, we propose Swanbench-Speech, a comprehensive benchmark that decomposes long-form speech quality into specific, disentangled dimensions. SwanBench-Speech has three key properties. 1) Rich speech scenarios: Focusing on long-form speech generation and dialog generation, SwanBench-Speech covers acoustics, semantics, and expressiveness challenges, and consists of 1,101 samples spanning 17 common speech scenarios; 2) Comprehensive evaluation dimensions: Along the acoustics, semantics, and expressiveness axes, SwanBench-Speech defines an automated evaluation protocol with seven metrics to provide a comprehensive, accurate, and standardized assessment; 3) Valuable Insights: Through extensive experiments, we reveal that current models still struggle in highly expressive scenarios and exhibit a notable gap in consistency and hierarchy compared to real recordings.

1 Introduction

SwanBench-Speech addresses the limited coverage and inadequate evaluation of long-form speech by benchmarking diverse generation and dialog scenarios across acoustics, semantics, and expressiveness. Its hierarchical metrics and experiments reveal persistent deficits in consistency, coherence, and expressive hierarchy, especially in highly expressive scenarios.

  • Motivation: Existing long-form speech evaluations cover limited domains or single-speaker settings, despite downstream applications involving multi-speaker interactions and rich semantic contexts.This mismatch prevents thorough assessment of long-form generation capabilities.
  • Benchmark: SwanBench-Speech contains 1,101 test samples spanning 17 scenarios across long-form speech generation and dialog generation.The benchmark organizes coverage around acoustics, semantics, and expressiveness.
  • Evaluation: The benchmark introduces a hierarchical automatic protocol with Acoustic Consistency, Prosodic Coherence, and Expressive Hierarchy alongside Fidelity and Accuracy.User studies validate these metrics as scalable proxies for human perception while they quantify temporal stability and expressive dynamics.
  • Findings: Current models rival human recordings in fidelity and accuracy but show substantial gaps in reverb consistency, prosodic coherence, and expressive hierarchy.Performance deteriorates in highly expressive scenarios, highlighting challenges in modeling long-term dependencies and dynamic stylistic variations.

2 Related Work

Related work addresses long-form TTS through prosody modeling, efficient long-sequence representations, and AR/NAR strategies for speaker transitions. However, existing objective metrics and benchmarks remain insufficient for evaluating long-context prosodic coherence, emotional richness, transition quality, and broader long-form behavior.

  • Long-form TTS: Long-form speech and dialogue generation must maintain prosodic coherence, model long sequences, and manage speaker transitions.Prior work explores joint style modeling, cross-sentence memory, and compact multi-resolution representations to address these challenges.
  • Long-form TTS: Recent approaches use flow matching in non-autoregressive models and speaker tokens in autoregressive models to improve long-context dialogue generation.Earlier systems combined autoregressive and non-autoregressive components, while newer work develops both paradigms further.
  • Evaluation: Existing metrics remain insufficient for evaluating prosodic coherence, emotional richness, and transition quality in long-form speech.Current evaluation relies on signal-based, MOS prediction, distributional, and accuracy metric families, which are nearly saturated on recent state-of-the-art systems.
  • Evaluation: Follow-up benchmarks increase difficulty through harder texts or controllability but remain sentence-level.This limits their coverage of long-context evaluation needs.

3 SwanBench-Speech

SwanBench-Speech is a hierarchical benchmark for long-form speech generation that contains 1,101 samples across 17 downstream applications. It organizes evaluation around acoustics, semantics, and expressiveness through seven objective metrics, supported by curated and refined data sources.

  • Benchmark scope: 1,101 samples span 17 downstream applications organized around acoustics, semantics, and expressiveness challenges.The benchmark is described as hierarchical and covers three core challenges across the downstream scenarios.
  • Acoustics Challenge: The acoustics challenge evaluates timbre consistency, reverb consistency, and sound fidelity across six scenarios including customer service, podcasts, and audiobooks.The six scenarios are customer service, podcast, chat, debate, audiobook, and interview.
  • Semantics Challenge: The semantics challenge measures content accuracy and prosodic coherence in information-dense scenarios such as lessons, presentations, seminars, and news.Its five scenarios are lesson, popular science, presentation, seminar, and news.
  • Expressiveness Challenge: The expressiveness challenge assesses expressive richness and expressive hierarchy in highly expressive scenarios including drama, talk shows, hosting, and sportscasts.Expressive richness captures sentence-level emotional impact, while expressive hierarchy captures paragraph-level expressive dynamics.
  • Data construction and refinement: The benchmark combines online text corpora, online audio media, and GPT-5-generated samples, followed by automated refinement and manual verification.Refinement includes semantic de-duplication, content-quality filtering, manual review, and dataset replenishment.
  • Metric validation: 0.82 SRCC indicates that the prosodic coherence metric aligns effectively with perceived coherence in a subjective test of 50 audio pairs rated by 10 evaluators.Evaluators compared clips synthesized from identical text on a scale from -2 to 2.

4 Experiments

Experiments evaluate SwanBench-Speech across broad model sets, dimensions, scenarios, and input lengths. The protocol also uses Real Speech and Real Dialogue as reference upper bounds and specifies dedicated evaluators for consistency, transcription, and expressiveness.

  • Models Evaluated: Sixteen systems are evaluated: ten open-source models and six closed-source flagship systems for single-speaker long-form speech.The open-source set includes ZipVoice, SparkTTS, CosyVoice2-0.5B, CosyVoice3-0.5B, GLM-TTS, MegaTTS3, IndexTTS2, FishSpeech-1.5, F5TTS, and VibeVoice.
  • Evaluation Models: WavLM TDCNN extracts speaker embeddings, Paraformer or WhisperX performs forced alignment, FunASR Nano computes WER, and Gemini3-pro evaluates expressiveness metrics.Paraformer is used for Chinese data, WhisperX for English data, and Gemini3-pro uses prompt enhancement.
  • Per-Dimension Evaluation: Real Speech and Real Dialogue serve as reference baselines and the topological upper bound for audio quality.Both baselines are derived from the source dataset.
  • Per-Scenario Evaluation: Models are evaluated across three core categories covering 17 scenarios, with performance calculated using the benchmark’s evaluation protocol.Figure 3 visualizes each model’s results across the three categories.
  • Evaluations On Generated Length: Five representative models are tested across increasing input lengths using 100 samples spanning acoustics, semantics, and expressiveness.The evaluated models are MegaTTS3, F5TTS, Cosyvoice2, SparkTTS, and VibeVoice; results appear in Figure 4.

5 Insights and Discussions

The discussion identifies persistent gaps between generated and real speech, with scenario difficulty, architecture choice, and training-data design shaping long-form quality. It highlights the need to balance robustness, expressiveness, semantic coherence, and temporal continuity rather than relying on scale alone.

  • Gap to Ground-Truth Audio: Proprietary models still outperform open-source models overall, although leading open-source systems match or surpass them on several evaluation dimensions.VibeVoice and SoulX-Podcast are identified as the strongest open-source models, while Minimax-Speech-02-hd and Gemini-2.5-pro-preview-tts lead proprietary systems.
  • Gap to Ground-Truth Audio: Real recordings expose persistent and systematic gaps in current long-form speech generation.
  • Impact of Scenarios: Acoustic challenge scenarios particularly impair acoustic-field consistency because frequent speaker transitions disrupt reverberation unity, while timbre consistency remains stable.The same transitions cause minor fidelity degradation, indicating greater robustness in timbre than in reverberation consistency.
  • AR v.s. NAR: NAR models offer stronger robustness and efficiency for long-text synthesis but often oversmooth rhythms and miss vocal dynamics and emotional nuances.The AR–NAR trade-off therefore contrasts parallel-generation robustness with the expressiveness needed for extended narration.
  • AR v.s. NAR: Coarse-to-fine architectures effectively reconcile long-range semantic coherence with local generation stability.
  • Data Quality v.s. Data Quantity: Fragmented short-form training data induces discourse-coherence bias; SparkTTS, trained on segments averaging less than 10 seconds, degrades in content accuracy and prosodic coherence as length increases.The discussion argues that data quality and temporal continuity should be prioritized over raw quantity, including curriculum learning from sentence- to paragraph-level training.

6 Conclusion

SwanBench-Speech is a holistic benchmark for long-form TTS models, covering 1,101 instances across 17 downstream scenarios. It enables precise automatic assessment through seven complementary, disentangled, human-aligned metric dimensions and benchmarks over 20 models.

  • 6 Conclusion: SwanBench-Speech evaluates long-form TTS models across 1,101 carefully curated instances spanning 17 downstream scenarios.The benchmark is designed to address three core challenges in long-form generation.
  • 6 Conclusion: Seven complementary metric dimensions provide a precise, automatic, disentangled, and human-aligned evaluation protocol.The protocol is intended to facilitate standardized assessment of long-form speech generation.
  • 6 Conclusion: Benchmarking over 20 models supports an in-depth analysis of long-form TTS performance.The conclusion describes extensive benchmarking conducted with more than 20 models.

Limitations

The benchmark is limited by its narrow language coverage and preliminary semantic evaluation, particularly for emotional and stylistic transitions in long-form speech.

  • Language coverage: SwanBench-Speech covers only Chinese and English, leaving low-resource languages and diverse dialects or accents underexplored.This restricts the benchmark’s linguistic scope.
  • Semantic evaluation: Its semantic evaluation remains preliminary because the metrics prioritize acoustic coherence without robustly assessing emotional and stylistic transitions through deep semantic understanding.The limitation concerns automated evaluation of long-form textual meaning and expressive continuity.

Ethical considerations … B Statistics of SwanBench-Speech

SwanBench-Speech constructs and refines a 1,101-sample benchmark spanning 17 scenarios and three dimensions—Acoustics, Semantics, and Expressiveness—while specifying ethical safeguards and release conditions. Its appendix details scenario design, data collection, filtering, human review, evaluation resources, and benchmark statistics.

  • Ethical considerations; Appendix Contents; A.4 Instructions for Use; B Statistics of SwanBench-Speech: Users must avoid infringing voice-actor rights and prohibited audio sources; the test set will be released under CC BY-NC-SA 4.0 for free non-commercial use with code publicly available.Additional voice profiles must follow their associated licenses, and the appendix also provides sections on evaluation protocols, validation, experiments, limitations, social impact, and benchmark statistics.
  • A.1 Explanation of Scenarios: The benchmark organizes long-form speech challenges into Acoustics, Semantics, and Expressiveness, covering 1,101 audio samples across 17 downstream scenarios.Examples include audiobooks, podcasts, talk shows, and news broadcasting.
  • Scenarios for Acoustics Challenges: Acoustic evaluation targets audio quality, timbre consistency, speaker transitions, and acoustic-environment consistency across six scenarios.These requirements include artifact-free audio, stable speaker identity, accurate switching, and unified acoustic scenes.
  • Scenarios for Semantics Challenges: Semantic evaluation separates content accuracy from prosodic coherence across five scenarios, testing omissions, repetitions, hallucinations, stress, intonation, and paragraph-level rhythm.News and popular science emphasize correctness, while lessons, seminars, and presentations add naturalness demands.
  • Scenarios for Expressiveness Challenges: Expressiveness is decomposed into Richness and Hierarchy, assessing sustained emotional quality alongside dynamic variation and alignment with semantic scenarios.Sportcast and live streaming stress Richness, whereas speech, hosting, talk shows, and drama require both dimensions.
  • A.2 Details of Data Collection; Online Text Corpora; Online Audio Media: Long-form texts are collected from online resources, cleaned with clean-text, and annotated for scenario, topic, and speaker identity.Audio materials are crawled from multiple platforms, denoised, filtered using a DNS-MOS threshold of 3.5, speaker-separated, transcribed, and proofread.
  • LLM Generation: GPT-5 generates structured cases for selected scenarios, with three undergraduate annotators double-checking samples at $0.20 per instance and total collection expenditure of $220.The generated cases undergo repetition, quality, privacy, social, and ethical checks to mitigate degradation and privacy infringement.

B.1 Categorical Statistics · B.2 Distributional Statistics

SwanBench-Speech is statistically characterized across categorical dimensions and text-length distributions, combining balanced language coverage, diverse speaker configurations, and three core challenges. Its texts are approximately normally distributed, concentrated between 80 and 500 units, with language-specific means supporting minute-level speech evaluation.

  • B.1 Categorical Statistics: B.1 Categorical Statistics: SwanBench-Speech analyzes 1,101 samples across language, speaker configuration, core challenges, scenarios, and content topics.The categorical analysis covers Chinese and English, single-, dual-, and multi-speaker settings, plus acoustics, semantics, and expressiveness.
  • B.1 Categorical Statistics: B.1 Categorical Statistics: 49.3% Chinese and 50.7% English samples provide a strictly balanced language ratio.The benchmark focuses on Chinese and English because both application ecosystems are relatively mature for long-form speech generation.
  • B.1 Categorical Statistics: B.1 Categorical Statistics: 101 multi-speaker samples involving 3 or 4 speakers extend evaluation beyond single-speaker speech and dual-speaker dialogue.These samples are included to evaluate multi-talker generation capabilities.
  • B.1 Categorical Statistics: B.1 Categorical Statistics: Acoustics is the largest of the three core challenges, accounting for 34.5% of samples.The dataset is described as relatively evenly distributed across acoustics, semantics, and expressiveness.
  • B.2 Distributional Statistics: B.2 Distributional Statistics: Text lengths approximately follow normal distributions and primarily fall within [80, 500].Chinese length is measured in characters and English length in words, excluding punctuation and other non-phonetic elements.
  • B.2 Distributional Statistics: B.2 Distributional Statistics: Mean text lengths are 271.8 for Chinese and 174.6 for English.These language-specific statistics quantify the benchmark’s long-form text distribution.
  • B.2 Distributional Statistics: B.2 Distributional Statistics: The benchmark selects 200 to 400 words to represent minute-level speech quality in extended application scenarios.This range prioritizes common use cases such as live streaming, customer service, and talk shows over only very long audiobook synthesis.
  • B.2 Distributional Statistics: B.2 Distributional Statistics: Privacy and ethical filtering selectively anonymizes private individuals while retaining public figures and checks for hate speech, violence, sexual content, and severe bias.The filtering procedure uses structured JSON output to support realistic evaluation data preparation.

C Details of Evaluation Protocol … D.3 Validation of Prosodic Coherence

The evaluation protocol measures long-form speech across acoustic consistency, fidelity, content accuracy, prosody, and expressiveness, complemented by subjective user studies. Validation shows strong human alignment for timbre consistency and establishes interpretable thresholds for timbre and prosodic quality.

  • C.1 Timbre Consistency: Timbre consistency uses 3-second windows with 2-second stride, WavLM speaker embeddings, pairwise cosine similarity, and their average score.For multi-speaker audio, 3D-Speaker verifies speaker turns, forced alignment separates speaker streams, and speaker-specific similarity averages are aggregated.
  • C.2 Reverb Consistency: Reverb Consistency is the standard deviation of windowed SRMR scores after discarding windows containing more than 60% nonspeech frames; lower values indicate greater stability.The metric assumes stable acoustic environments, while acknowledging that scenarios such as Outdoor Live Streaming may require dynamic shifts.
  • C.3 Sound Fidelity / C.4 Content Accuarcy: Sound Fidelity uses reference-free SQUIM-PESQ, while Content Accuracy uses CER for Chinese and WER for English after rigorous transcript normalization.Normalization removes punctuation, standardizes whitespace, converts Traditional Chinese to Simplified, and filters non-ASCII English characters; FunASR-Nano reports WER of 1.76% and CER of 2.56% on clean benchmarks.
  • C.5 Prosodic Coherence / C.6 Expressive Richness / C.7 Expressive Hierarchy: Prosodic evaluation uses SpeechJudge ratings from 1.0 to 5.0 across Prosodic Coherence & Flow, Rhythmic Hierarchy & Layering, and Overall Naturalness.Expressive Richness averages LALM scores over non-overlapping 10-second chunks, whereas Expressive Hierarchy evaluates the entire audio sequence across dimensions including Emotional Variation.
  • D User Study / D.1 Validation of Timbre Consistency / D.2 Validation of Sound Fidelity: Subjective evaluation recruits 10 expert listeners, balanced by gender and spanning audio engineering, live streaming, and signal-processing research backgrounds, using MOS tests.The user studies randomly select 50 test samples and instruct listeners to focus on the targeted attribute while disregarding unrelated acoustic, semantic, or expressive factors.
  • D.1 Validation of Timbre Consistency: SRCC=0.75, PLCC=0.77, and KRCC=0.59 show close alignment between objective Timbre Consistency and human MOS judgments.Scores below 0.85 indicate significant drift, scores in [0.85, 0.90] indicate generally acceptable performance, and scores below 0.93 are described as superior maintenance comparable to ground truth.
  • D.1 Validation of Timbre Consistency: The timbre metric may misclassify periodic timbre variations because global averaging can overlook rhythmic fluctuations, motivating future temporal modeling.This limitation is identified for edge cases such as looping patterns that are perceptually inconsistent despite receiving favorable global scores.
  • D.3 Validation of Prosodic Coherence: A Score Divergence > 1 indicates a substantial perceptually obvious prosodic gap, while Score ≥ 4 reflects competent basic prosody and Score ≥4.5 is virtually indistinguishable from ground truth.The Prosodic Coherence validation uses a SpeechJudge-based human preference test to assess the model’s evaluation performance.

D.4 Validation of Expressiveness … F.1 Inference Speed

The paper validates expressiveness evaluation against human ratings, details reproducible inference and voice-selection procedures, and shows that non-autoregressive models generate long-form speech faster than autoregressive models.

  • D.4 Validation of Expressiveness: A 200-sample subjective evaluation finds Gemini3-Pro significantly outperforms other models across both expressiveness metrics, while Qwen3-Omni models outperform GPT-4o.Listeners followed the same criteria as the LALM prompts; the comparison included four MOS predictors and eight flagship LALMs.
  • D.4 Validation of Expressiveness: High inter-rater correlation confirms the reliability and validity of the expressiveness evaluation protocol.The study also reports that Gemini 3 Pro produced inconsistent scores for only 11 instances across repeated trials.
  • E.1 Computational Resources and Environments: Open-source inference and evaluation use 8 NVIDIA GeForce RTX 4090 GPUs, an Intel Xeon Gold 6530 CPU, Ubuntu 22.04, and specified software dependencies.The pipeline uses Python 3.10, Py-Torch 2.8.0, Torchaudio 2.8.0, and Transformers 4.57.3.
  • E.2 Selected Voice: Open-source evaluation uses 25 reference audio prompts spanning diverse datasets and over 20 representative timbres, while closed-source models use official expressive voices.The open-source references cover multiple dimensions including language; closed-source voice specifications are provided separately.
  • E.3 Synthesis Strategy: Models follow official or default synthesis configurations, including repository-specific adjustments, zero-shot open-source generation, designated closed-source voices, and 24kHz audio resampling.Closed-source models do not manually adjust emotion, pitch, or speaking rate.
  • F Supplementary Experiment: Inference speed is evaluated with Real Time Factor (RTF), defined using generation time and generated-audio duration.Computational efficiency results are summarized for the evaluated open-source models.
  • F.1 Inference Speed: Non-autoregressive models show a significant generation-speed advantage over autoregressive models, consistent with their parallel decoding mechanism.The comparison concerns mono-speaker long-form speech generation and is summarized in Tables 9 and 10.

F.2 Ablation on Window Size … G More Analysis Based on SwanBench-Speech

The ablations identify suitable consistency-evaluation windows and show that longer generation increasingly degrades most acoustic, prosodic, and expressive dimensions. The benchmark also includes dedicated 3- and 4-speaker dialogue cases for evaluating multi-speaker synthesis.

  • F.2 Ablation on Window Size: Window sizes ≤2s misalign timbre-consistency results with human perception, while sizes ≥4s average out transient timbre mutations.Real data shows lower consistency than CosyVoice3 at ≤2s; larger windows reduce the real–synthetic discrepancy.
  • F.2 Ablation on Window Size: A 3s window and 2s stride are selected for timbre consistency because stride has no significant effect and larger strides improve efficiency.The selected setting balances evaluation efficiency with consistency assessment after the window-size and stride ablations.
  • F.2 Ablation on Window Size: Table 11 evaluates window settings for timbre consistency using CosyVoice3 and OpenAI-tts-1-hd in single-speaker settings.The table accompanies the timbre-consistency window ablation.
  • F.2 Ablation on Window Size: A 1s reverb-consistency window is unstable, whereas windows ≥4s reduce inter-model differences by overlooking small-scale acoustic-field mutations.VibeVoice has a mean reverb score of 9.25 with an excessively high standard deviation at 1s.
  • F.3 Ablation on Generated Length: Generation-length analysis tracks Timbre Consistency and Timbre Similarity in addition to the original six dimensions.The analysis extends Figure 4 to examine how these two metrics evolve as generation length increases.
  • F.2 Ablation on Window Size: Table 12 evaluates window settings for reverb consistency using VibeVoice and Gemini-2.5-pro-preview-tts in two-speaker settings.The table accompanies the reverb-consistency window ablation.
  • F.3 Ablation on Generated Length: Nearly all metrics decay as generation duration increases, with the strongest degradation in Reverb Consistency, Prosodic Coherence, and Expressive Hierarchy.These trends indicate difficulty maintaining acoustic-field stability and capturing long-term dependencies, while Timbre Similarity and Timbre Consistency remain relatively stable.
  • F.4 Multi-Speaker Dialogue Generation: The benchmark adds 101 test cases for 3- and 4-speaker dialogues to support multi-speaker long-form speech research.These cases evaluate ElevenLabs Multilingual V2, Gemini-2.5-pro-preview-tts, and OpenAI-tts-1-hd; results are reported in Table 13.

G.1 Detailed analysis on each metric … I Social Impacts

The benchmark analysis finds that current models generally match real speech in acoustic fidelity but remain weaker in prosodic coherence, expressive richness, and paragraph-level hierarchy, especially in demanding scenarios and languages. The paper also identifies limitations involving reproducibility, language coverage, reference-voice sensitivity, instruction following, and responsible deployment.

  • G.1 Detailed analysis on each metric: 0.96 vs. 0.93 and 0.95 vs. 0.92 show small real-versus-synthetic gaps in single- and two-speaker timbre consistency, while dialogue reverb consistency remains substantially weaker.Single-speaker reverb consistency is comparable to human recordings, but open average dialogue reverb consistency is 3.45 and most models lag real data.
  • G.1 Detailed analysis on each metric: Generated speech closely matches real data in sound fidelity, whereas content accuracy is strong for CosyVoice2 and MegaTTS3 but weaker for SparkTTS in long-form generation.Sound fidelity constraints appear largely resolved, while content accuracy remains especially relevant for autoregressive end-to-end architectures.
  • G.1 Detailed analysis on each metric: Approximately 1.5 points and nearly 1.0 point separate open- and closed-source models from real data in Expressive Richness, while Expressive Hierarchy remains difficult.Closed-source models outperform open-source models in prosodic coherence and expressive metrics, and models score lower on Expressive Hierarchy than Expressive Richness in single-speaker tasks.
  • G.2 Analysis based on the scenarios: Most metrics degrade in high-expressiveness scenarios, with sportscast, host, and talk-show settings showing the most severe declines.These scenario-level results indicate difficulty modeling highly dynamic prosody over long-form speech.
  • G.3 Analysis based on the Languages: Target language significantly affects most models’ performance: ElevenLabs Multilingual V2 scores 1.79 vs. 2.87 in Expressive Richness, while Seed-TTS-Podcast scores 4.19 vs. 3.49.The compared values correspond to Chinese versus English results, respectively.
  • H Future Works: Future work must improve reproducibility by developing open-source expressiveness evaluators, expand beyond English and Chinese, diversify reference voices, and evaluate instruction-following in long-context settings.The stated limitations concern closed-source evaluators, limited language coverage, reference-voice sensitivity, and predominantly zero-shot evaluation.
  • I Social Impacts: The paper warns that stronger speech-generation capabilities can increase misuse risks, requiring ethical alignment, oversight, and responsible deployment.The authors report rigorous ethical review and anonymization of text data as mitigation measures.
  • I Social Impacts: The evaluation prompts assess prosodic flow, rhythmic hierarchy, naturalness, emotional arcs, vocal dynamics, scene fit, and layered emotional resonance through structured JSON outputs.Separate prompts target Prosody Coherence, Expressive Hierarchy, and Expressive Richness in long-form audio.
Loading 2605.28618v1…