Source-linked AI summary

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

Yu Zhang, Ruiqi Li, Changhao Pan, Ke Lei, Xiang Yin, Cheng Yang

arXiv:2608.02023v2eess.AScs.SD

TL;DR

Creators need multi-speaker speech and audio generation that supports both reference-free voice design and reference-audio conditioning, while controlling styles, scenes, and fine-grained content. SwanTale addresses this with SwanData-Caption and a unified generation model, achieving leading zero-shot and instruct results, best expressiveness, and complex multi-speaker speech-and-audio generation.

  • Problem

    Existing systems support zero-shot synthesis from reference audio, but lack unified support for designing voices without recordings alongside controlled multi-speaker speech and acoustic scenes.

  • Method

    SwanTale combines fine-grained SwanData-Caption supervision with a unified model for zero-shot and instruct generation across multi-speaker speech and audio.

  • Results

    SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex multi-speaker speech-and-audio generation.

  • Takeaways & Limitations

    The model provides a unified generator for designing voices from descriptions, reusing them through reference audio, and jointly generating speech, environments, and audio effects.

  • Takeaways & Limitations

    Long-form instruct generation remains challenging for complex multi-speaker scenes longer than two minutes containing audio effects, alongside precise local style control.

Abstract

from arXiv · show

Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/#swantale.

1 Introduction

SwanTale addresses multi-speaker expressive speech and audio generation for both zero-shot and instruct settings, motivated by media-production needs for designed voices, natural-language control, acoustic scenes, and voice reuse. The paper contributes SwanData-Caption and SwanTale, and reports leading performance across key zero-shot and instruct metrics with strong expressiveness and complex instruct-generation support.

  • Motivation: Media production often requires voices without reference recordings, natural-language control of speaker styles, acoustic scenes, and reusable designed voices.These needs arise in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production.
  • Task formulation: Zero-shot synthesis uses speech content with reference audio for speaker identity, whereas instruct synthesis uses captions describing environments, speaker styles, and fine-grained content.Dialogue tasks may also include speaker-turn labels, and recent systems use local style descriptions for emotion and speaking rate.
  • Challenges: The introduction identifies data scarcity and task compatibility as central challenges for combining diverse, high-quality captioned audio with zero-shot and instruct generation.Expressive and clean audio collection is costly, while multi-level natural-language caption annotation is also expensive.
  • Data contribution: SwanData-Caption converts diverse speech-centered media audio into multi-level, multi-style captions using targeted synthetic subsets plus reliable speech spans and speaker-attributed text anchors.Its coverage design explicitly targets special styles such as pronunciation-challenging text.
  • Model and evaluation: SwanTale combines SwanVAE, reward-conditioned quality control, Engram conditioning, Unified MoE, curriculum learning, and GRPO post-training for multi-speaker expressive speech and audio generation.Experiments cover zero-shot monologue and dialogue TTS, instruction following, acoustic quality, and heterogeneous instruct generation; SwanTale leads on multiple key metrics and achieves the best expressiveness scores in both tasks.

2 Data Pipeline: SwanData-Caption

SwanData-Caption is a four-stage pipeline that designs broad speech-and-audio coverage, preprocesses speech-centered data, annotates unified fine-grained captions, and refines the resulting data. Its captions provide structured supervision by representing environments, speaking styles, fine-grained delivery, and local audio effects for controllable multi-speaker generation.

  • Coverage design: Its coverage spans speech, audio effects, and background music across varied speakers, recording conditions, scenes, sound fields, persistent effects, delivery changes, and local effects.Representative media include short dramas, advertisements, and animations, alongside internal, real-world, speech-centered, and media-style data.
  • Coverage design: Three targeted synthetic subsets each contain 100k utterances covering elderly speech, short Chinese and English utterances, and challenging pronunciation targets.A phoneme-aware TTS teacher is used to maintain pronunciation accuracy; the short-utterance subset averages 1.5 seconds and the elderly-speech subset about 10 seconds.
  • SwanData-Speech preprocessing: Speech-centered preprocessing separates vocals from residual backgrounds, performs coarse diarization and ASR, and aligns transcripts while preserving environmental sound fields and local audio effects.The pipeline uses Ultimate Vocal Remover, 3D-Speaker, Seed-ASR 2.0, SenseVoice ASR, and SwanAligner; caption annotations provide speaker discrimination and suitable punctuation.
  • Pipeline overview: SwanData-Caption converts raw audio and ASR transcripts into unified, fine-grained captions for structured supervision of controllable audio generation.The pipeline is organized into coverage design, SwanData-Speech preprocessing, caption annotation, and data refinement.
  • Caption annotation: Each caption organizes scene and acoustic context in Environment, actual speaking participants and their styles in Speakers, and local emotion, delivery, and audio effects in Content.A style-persona library provides soft priors for animation, short drama and film/TV drama, and advertisement/digital-human content to enrich speaker descriptions.

3 Method: SwanTale

SwanTale is a unified multi-speaker speech and audio generation model built on SwanVAE acoustic latents, a flow-based Transformer, Engram conditioning, reward-conditioned quality control, and Unified MoE. Its design supports instruct and zero-shot generation across heterogeneous speech, environmental audio, effects, singing, and music through adaptive modeling capacity.

  • Method overview: SwanTale comprises SwanVAE, a flow-based Transformer with content, caption, speaker-turn, Engram, and reward-conditioned controls, Unified MoE, curriculum training, GRPO post-training, and an inference procedure.The method section presents these components as the model’s main architectural and training stages.
  • SwanVAE: SwanVAE produces 25 Hz continuous acoustic latents by balancing reconstruction fidelity, representation compactness, and downstream flow-model learnability.SwanTale uses the globally normalized posterior mean µϕ as its deterministic acoustic representation.
  • Conditioning: Engram uses content-dependent gating within each caption, opening its memory path gradually during training and applying it more strongly to structured markers than free-form language.The gate has two branches sharing tables and WV while keeping separate WK, with averaged gated outputs.
  • Task coverage: SwanTale uses one network for instruct and zero-shot tasks spanning multi-speaker expressive speech, general audio, occasional singing voice, and music.The model must preserve speech content, prosody, and speaker continuity while handling persistent environmental sounds and transient audio effects.
  • Unified MoE: Unified MoE combines a task-level shared path, frame-level dynamic expert routing, and a diffusiontime-aware computation budget for heterogeneous speech and scene audio.This adaptive capacity avoids imposing the same fixed computation on every acoustic region.

4 Experiments

Experiments evaluate SwanTale and SwanVAE across reconstruction, zero-shot, instruct, scene-quality, and hard-caption benchmarks. SwanTale leads expressive and perceptual metrics while revealing weaknesses in content accuracy, English descriptive style, and role-play.

  • Evaluation Setup: The evaluation covers four SwanVAE audio domains, zero-shot speech, natural-language instruction following, scene quality, and difficult multi-speaker speech-and-audio instructions.SwanBench-Scene uses 180 instructions rated by five annotators across four 1–5 dimensions, while SwanBench-Caption contains 64 cases scored on three dimensions.
  • SwanVAE: SwanVAE ranks strongly across domains: it leads speech PESQ and MCD, singing-voice PESQ, STOI, and MCD, general-audio ViSQOL, and remains competitive on music.On music, EnCodec leads both metrics; SwanVAE ranks second on ViSQOL and third on LSD, using one checkpoint across all domains.
  • Zero-Shot Task: SwanTale ranks first in Timbre Consistency, Expressive Richness, and Expressive Hierarchy for monologue and dialogue, but competitors lead content accuracy and Sound Fidelity.Relative to SwanVoice, monologue scores reach 0.95 Timbre Consistency, 0.086 Content Error, and 3.75, 3.90, 3.70 for SpeechJudge, Expressive Richness, and Expressive Hierarchy; dialogue reaches 0.94, 0.120, 3.92, 3.66, and 3.85.
  • InstructTTSEval: SwanTale ranks first on Chinese APS (86.1), ties for first on English APS (84.2), and ranks second on Chinese DSD (80.1), while English DSD and RP remain weaker.The results indicate balanced Chinese–English performance within tasks, but weaker cross-system competitiveness on English DSD and role-to-voice mapping in RP.
  • SwanBench-Scene: SwanTale achieves the highest overall Mean MOS (4.22) and leads all four overall scene dimensions, with highest Mean MOS for advertising (3.88), comic drama (4.45), and general scenes (4.34).Its only second-place dimension score is Audio Fullness in comic drama (4.60).
  • SwanBench-Caption: Removing Unified MoE lowers Instruction Accuracy from 3.39 to 3.02, Acoustic Quality from 4.31 to 4.09, and Overall Expressiveness from 3.82 to 3.56; a 32B caption encoder raises them to 3.70, 4.34, and 3.98.The ablations support benefits from Unified MoE and increased caption-encoder capacity for complex speech-and-audio instructions.

5 Conclusion · Appendix

SwanTale unifies multi-speaker speech and audio generation for instruct and zero-shot tasks through dedicated data supervision and integrated modeling and training methods. Experiments show strong expressiveness and complex multi-speaker instruct capabilities, while highlighting challenges in background music, long-form scenes, and precise local style control.

  • 5 Conclusion: SwanTale unifies multi-speaker speech and audio generation across instruct and zero-shot tasks.
  • 5 Conclusion: SwanData-Caption supplies supervision through targeted data coverage, speech-aware preprocessing, multi-level caption annotation, and quality filtering.
  • 5 Conclusion: SwanTale combines SwanVAE, a flow-based Transformer, reward-conditioned quality control, Engram conditioning, Unified MoE, curriculum learning, and GRPO post-training.
  • 5 Conclusion: SwanTale leads on multiple key zero-shot and instruction-following metrics and achieves the best expressiveness scores in both tasks.
  • 5 Conclusion: The model supports complex instruct generation involving multi-speaker speech and audio within the same model.
  • 5 Conclusion: Complex background music generation remains difficult when music must change type or transition in response to different emotions.
  • 5 Conclusion: Long-form instruct generation remains challenging for complex multi-speaker scenes longer than two minutes that also contain audio effects.
  • 5 Conclusion: Precise local style control remains difficult, including continuous emotional changes for a specified context.

A Caption Style Matrices

The section condenses hierarchical annotation guides into style matrices for animation, drama, and advertisement/digital-human media. It also specifies how annotators select matrices, list speakers, and describe stable and utterance-level characteristics.

  • Scope: Three condensed matrices cover animation, short drama and film/TV drama, and advertisement and digital-human content.The source guides are hierarchical, with media-specific organization of audience, topics, roles, dialogue, persona, voice style, and expressiveness.
  • Annotation procedure: Annotators select a matrix using the scene trigger and list only speakers who actually speak, ordered by first utterance.Each Speakers entry begins with perceived gender and age range.
  • Annotation procedure: Each speaker receives three to five stable, discriminative characteristics covering supported role or persona, stable timbre, and habitual delivery.The matrix supplies candidate descriptions for these characteristics, while utterance-level changes capture emotion and other variable properties.

B SwanVerifier · B.1 Motivation

SwanVerifier addresses recurring errors in perceived-age and gender labels by checking their acoustic plausibility directly from waveforms. It abstains on ambiguity while leaving detailed persona, role, and expressive-style annotation to other processes.

  • B.1 Motivation: Perceived age and gender are difficult to caption automatically in noisy media audio and stylized performance recordings.The challenging cases include child or elderly speech, role-playing voices, advertisements, animation, and game-style dubbing.
  • B.1 Motivation: Because demographic labels recur throughout training data, systematic errors can create repeated supervision rather than isolated captioning mistakes.
  • B SwanVerifier: SwanTale uses the Speakers field to provide stable, controllable speaker attributes.
  • B SwanVerifier: SwanVerifier provides a waveform-grounded check for coarse demographic attributes.
  • B SwanVerifier: It tests whether demographic labels are acoustically plausible and abstains when the waveform or prediction is ambiguous.
  • B.1 Motivation: Caption generation is outside SwanVerifier’s scope.
  • B.1 Motivation: Detailed persona, role, and expressive style remain handled by caption annotation and human audit when automatic evidence is insufficient.

B.2 Overview

SwanVerifier is a compact WavLM-based audio tagger used by SwanData-Caption for demographic consistency checks. Its broader acoustic predictions remain auxiliary evidence, while performance-related variation is handled during final captioning and auditing.

  • Tagger architecture: SwanVerifier resamples waveforms to 16 kHz, encodes frame-level speech representations, and pools them through attribute-specific heads for utterance-level predictions.It is built on a pretrained WavLM encoder.
  • Demographic checks: The primary demographic checks use age-group labels: Child, Teenager, Youth-Adult, Middle-aged, and Elderly.These labels are part of the caption inventory used for consistency checks.
  • Demographic checks: Perceived gender is checked using the male and female labels in the caption inventory.Gender is one of the two primary demographic heads used for consistency checks.
  • Auxiliary evidence: Emotion, pitch, pitch standard deviation, and speaking speed are treated as auxiliary evidence rather than grounds for automatic demographic correction.Transient affect, emphasis, hesitation, and scene-specific performance remain in Content and are reviewed during final captioning and auditing.

B.3 Problem Setup · B.4 Backbone Encoding and Prediction Heads

B.3 defines captioned speaker inventories, verification segments, and selective demographic consistency checks, while B.4 describes WavLM encoding and independently auditable age and gender prediction heads.

  • B.3 Problem Setup: A caption c contains a speaker inventory for the speakers represented in the input.The setup introduces the caption as the container for speaker information.
  • B.3 Problem Setup: Each speaker k has a normalized description d_k, while x_k denotes a vocal segment attributed to that speaker.These descriptions and segments connect caption metadata with verification inputs.
  • B.3 Problem Setup: Automatic demographic checking applies only to single-speaker, acoustically separable segments; overlap, cross-talk, and uncertain attribution bypass hard verification.Those cases are deferred to later auditing before any automatic repair.
  • B.3 Problem Setup: SwanVerifier is a selective consistency check rather than a complete speaker profiler.Its scope is limited by selective verification and the policy against inserting missing labels.
  • B.3 Problem Setup: When present in d_k, normalized age and perceived-gender labels are denoted y_a,k and y_g,k.Missing caption labels are not inferred or inserted.
  • B.3 Problem Setup: SwanVerifier estimates age and gender class distributions and compares sufficiently confident predictions with corresponding caption labels.The verifier uses p_a(· | x_k) and p_g(· | x_k) for the two attributes.
  • B.4 Backbone Encoding and Prediction Heads: WavLM maps each vocal segment to frame-level hidden states, with T_k denoting the number of valid acoustic frames.These hidden states provide the backbone features for subsequent prediction heads.
  • B.4 Backbone Encoding and Prediction Heads: For age and gender, attention-pooling heads form utterance-level representations and separate classifiers produce attribute-specific output distributions.Separate heads allow the two predictions to be calibrated and audited independently.

B.5 Training Objective

The training objective uses cross-entropy for demographic prediction, compensates for age-class imbalance, and evaluates auxiliary terms only when their labels are available.

  • Demographic heads are trained with cross-entropy over normalized labels.
  • The demographic loss combines age and gender cross-entropy terms, weighted by λa and λg, with age weighting wya,k.
  • Each loss term is evaluated only when its label is available; the emotion head excludes unknown labels, while other auxiliary outputs provide secondary signals.

B.6 Training and Scope

SwanVerifier is a fine-tuned WavLM-based internal filtering component for demographic checking in the data pipeline, not a standalone tagging benchmark. It excludes samples lacking an unambiguous acoustic subject and should not be interpreted as a comparison with dedicated speaker-attribute systems.

  • Training and scope: SwanVerifier fine-tunes an existing WavLM-based tagger on labeled speech mapped to its verifier taxonomy, excluding samples without an unambiguous acoustic subject.It is used for supervised demographic checking in the data pipeline.
  • Training and scope: SwanVerifier is an internal filtering component rather than a new tagging benchmark, so its results characterize the pipeline verifier instead of comparing dedicated speaker-attribute systems.The reported results should not be read as a benchmark comparison.

B.7 Evaluation · B.8 Inference Procedure

The held-out evaluation reports SwanVerifier’s utterance-level accuracy, while inference applies confidence-based abstention and routes ambiguous cases to manual or human-audited handling. Automatic demographic corrections are limited to confident age- and gender-label evidence from speaker-attributed vocal segments.

  • B.7 Evaluation: SwanVerifier is evaluated on a held-out labeled split using utterance-level accuracy.Table 12 reports the accuracy in percentage terms.
  • B.7 Evaluation: Emotion accuracy is computed only for samples with a non-unknown emotion label.
  • B.7 Evaluation: Held-out accuracy characterizes the filtering component but does not establish that every prediction is safe for automatic correction.
  • B.8 Inference Procedure: Inference therefore uses confidence-based abstention and sends ambiguous cases to manual audit.
  • B.8 Inference Procedure: SwanVerifier compares speaker-attributed vocal segments with normalized demographic tokens only when attribute confidence exceeds a threshold.
  • B.8 Inference Procedure: Confident matches leave captions unchanged, whereas confident mismatches trigger repair, re-captioning, or removal.
  • B.8 Inference Procedure: Low-confidence predictions, unresolved overlap, and missing demographic tokens produce no automatic decision.
  • B.8 Inference Procedure: Automatic decisions target systematic age- and gender-label errors while excluding inferred persona, role identity, and transient delivery as demographic evidence.Caption-level checks and human listening audits address richer attributes.
Loading 2608.02023v2…