Source-linked AI summary
Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction
Kenichi Fujita, Yusuke Ijima
TL;DR
Direction-following TTS must modify a reference reading according to a direction while preserving content and speaker identity, but suitable paired training triplets are scarce. The paper constructs pseudo-triplets using impression-controlled synthesis and LLM-generated relative directions, then trains an embedding-space style refiner. Pseudo data alone enables stable identity-preserving refinement, while combining pseudo and recorded data provides the most favorable balance of direction alignment and speaker preservation.
Problem
Direction-following TTS lacks large-scale triplets pairing a reference utterance, modification direction, and corresponding modified utterance for the same script.
Method
The paper generates pseudo-triplets with impression-controllable TTS and LLM directions derived from estimated impression differences, then refines speech embeddings conditionally on the direction.
Results
Pseudo data alone enables stable speaker-preserving modification with reasonable direction alignment, while combining pseudo and recorded data achieves the best balance between stability and expressiveness.
Takeaways & Limitations
Scalable pseudo-triplet construction supports direction-following TTS, and recorded data complements it by improving direction alignment.
Abstract
from arXiv · showhide
Voice actors often re-read the same script while modifying their delivery in response to performance directions. We study this setting as direction-following TTS, where a system generates a new utterance that reflects a given direction relative to a reference utterance while preserving speaker identity and linguistic content. A key challenge is the lack of training data capturing such relative modifications. To address this, we propose a scalable pseudo-triplet construction pipeline that generates~(reference utterance, direction text, modified utterance) triplets. It generates controlled style variations using an impression-controllable TTS model and uses an LLM to produce natural language directions from estimated impression differences. Experimental results demonstrate that pseudo-triplets alone enable stable speaker-preserving modification, and that combining pseudo and recorded data further improves direction alignment while maintaining speaker similarity. Audio examples are available on our demo page https://ntt-hilab-gensp.github.io/IS2026pseudo/
1. Introduction
Direction-following TTS generates a new reading of fixed content that follows a natural-language performance direction relative to a reference while preserving speaker identity. The paper addresses scarce paired direction data with pseudo-triplets from controlled synthesis and LLM-generated directions.
- Direction-following TTS transforms a reference utterance according to a direction while preserving its speaker identity and linguistic content.It differs from conventional style-conditioned TTS by modeling a relative transformation between pre-modification and post-modification utterances.
- Training requires triplets containing a pre-modification utterance, direction text, and corresponding post-modification utterance for the same script.
- Large speech corpora usually provide only one reading per script, while smaller multi-recording corpora generally lack modification directions.This limits coverage of relative stylistic changes and robust speaker-preserving transformation across diverse speakers.
- The proposed pipeline uses impression-controllable TTS to generate paired readings and an LLM to describe their estimated impression differences as natural-language directions.The resulting pseudo-triplets train a transformation from pre-modification to post-modification utterances.
2. Data design for direction-following
The data design combines scalable pseudo speech generation with impression-based filtering and relative direction annotation. Directions are generated from the difference between estimated pre- and post-modification voice impressions.
- Pseudo data and recorded data provide complementary sources: synthesis offers scalability, while interactive recordings capture authentic human responses.
- 2.1. Pseudo paired utterance generation: For each script and reference, an impression-controllable zero-shot TTS model synthesizes baseline and style-modified utterances as pre-modification and post-modification speech.
- 2.1. Pseudo paired utterance generation: The impression representation contains 13 continuous antonym-based voice-impression dimensions, including pitch, clarity, calmness, emotionality, and fluency.Two added axes are fluent–hesitant and emotional–neutral.
- 2.2. Filtering and impression estimation: Generated pairs are filtered using speaker-embedding similarity and speaking-rate differences, while near-identical embeddings are removed because they provide little directional guidance.
- 2.2. Filtering and impression estimation: The impression estimator computes pre- and post-modification vectors, and their difference is defined as ∆I = Ipost − Ipre.
- 2.2. Filtering and impression estimation: An LLM generates a plausible performance direction from the relative impression difference rather than labeling either utterance with an absolute style.The generated direction describes the transformation from pre-modification to post-modification speech.
3. Direction-conditioned style refiner
The direction-conditioned style refiner predicts speaker-dependent embedding modifications while keeping the backbone TTS model fixed. Rectified flow matching models multiple plausible shifts instead of collapsing them into a conditional mean.
- The style refiner modifies speech embeddings, and inference adds the predicted modification to the pre-modification embedding.Keeping the backbone fixed focuses learning on direction-dependent style refinement.
- The target modification is ∆ = epost − epre, treated as a local additive approximation without assuming global linearity.Different speakers may require different embedding shifts for the same direction.
- Rectified flow matching models a direction-conditioned stochastic vector field because deterministic regression would otherwise collapse multiple plausible shifts into a conditional mean.
- The refiner predicts velocity conditioned on the pre-modification embedding, direction representation, and time, matching the target velocity ∆ − x0.Auxiliary losses encourage directional consistency and modification-magnitude alignment.
4. Experimental setup
The experiments compare scalable pseudo data, recorded human responses, and their combination across speaker similarity and direction alignment. The setup also varies pseudo-data speaker diversity and includes objective and subjective evaluation.
- 4.1.1. Pseudo data: Pseudo data were generated from 1,600 Japanese speakers by applying impression control to randomly selected dimensions and sampling control values from −2 to +2.
- 4.1.1. Pseudo data: Filtering retained pairs with ECAPA-TDNN cosine similarity from 0.80–0.95 and speaking-rate ratios from 0.85–1.15, yielding 74,619 utterance pairs.
- 4.1.1. Pseudo data: Up to five LLM-generated performance directions were paired with each valid speech pair to form pseudo triplets.
- 4.1.1. Pseudo data: The final pseudo dataset contained 350,617 triplets and 127.6 hours of unique speech, with subsets covering 200, 400, and 800 speakers.
- 4.1.2. Recorded data: Recorded data comprised 8.9 hours and 6,899 utterance pairs from interactive sessions with two professional voice actors.
- 4.2. Training and evaluation conditions: The comparison trains separate models with pseudo-only, recorded-only, combined, and different pseudo-speaker-scale configurations.
5. Results
Objective and subjective evaluations show that pseudo data improves speaker-identity robustness, while recorded data supports stronger direction alignment. Combining both sources offers the most favorable balance between stability and expressive modification.
- Evaluation setup: Evaluations covered naturalness, speaker similarity, and direction alignment for two seen and two unseen speakers.The objective evaluation generated 15,000 direction-conditioned utterances per speaker.
- Objective evaluation: Pseudo conditions produced more robust speaker similarity than Recorded for unseen speakers, while increasing pseudo-data diversity reduced extreme speaker drift.The Full condition remained more stable than Recorded while balancing robustness and variability.
- Objective evaluation: Direction refinement preserved naturalness, with scores comparable to recorded utterances across configurations.Source recorded utterances scored 2.67±0.01 for seen speakers and 2.97±0.01 for unseen speakers.
- Objective evaluation: Full and Recorded achieved higher direction alignment than Pseudo conditions, while Full retained improved speaker robustness relative to Recorded.Increasing pseudo-data scale further stabilized speaker identity preservation without substantially degrading alignment.
- Subjective evaluation: Subjective results showed lower speaker similarity for Recorded, whereas Full and Pseudo-all achieved higher similarity scores for seen and unseen speakers.Pseudo-all slightly outperformed Full, likely because its more conservative modulation produced smaller changes between utterances.
- Subjective evaluation: Recorded achieved the highest subjective direction alignment, followed by Full and Pseudo-all; Full balanced expressive variation and identity stability.Participants rated alignment on a 5-point scale and assigned a score of 1 when speaker identity changed significantly.
- Analysis of modulation: Recorded pairs showed larger absolute F0 changes than pseudo-generated pairs, helping explain Pseudo-all’s more conservative modulation.Mean ln F0 changes were 0.14±0.15 for Recorded and 0.05 ± 0.08 for Pseudo.
- Overall findings: Pseudo-all alone enabled stable identity-preserving refinement with reasonable direction alignment, while Full provided the most favorable overall trade-off.Subjective alignment results were broadly consistent with the LLM-based objective evaluation.
6. Conclusion
The paper introduces direction-following TTS and addresses scarce paired direction data with pseudo triplets. Evaluations show complementary roles for pseudo and recorded data, with their combination balancing identity stability and expressive alignment.
- Conclusion: Direction-following TTS models relative performance modification using pseudo triplets built through impression-controlled synthesis and LLM-based direction generation.The triplets address the scarcity of paired direction data.
- Conclusion: Pseudo data improves speaker robustness, recorded data enhances direction alignment, and combining them achieves the best balance between stability and expressiveness.Future work will investigate integrating human expressiveness with pseudo data.
- Conclusion: Pseudo data alone enables stable identity-preserving refinement with reasonable direction alignment, supporting the effectiveness of pseudo-triplet construction.This conclusion is reported alongside the complementary benefits of recorded data.