Source-linked AI summary

Seamless: Multilingual Expressive and Streaming Speech Translation

Seamless Communication, Loïc Barrault, Yu-An Chung, Mariano Coria Meglioli, David Dale, Ning Dong, Mark Duppenthaler, Paul-Ambroise Duquenne, Brian Ellis, Hady Elsahar, Justin Haaheim, John Hoffman, Min-Jae Hwang, Hirofumi Inaguma, Christopher Klaiber, Ilia Kulikov, Pengwei Li, Daniel Licht, Jean Maillard, Ruslan Mavlyutov, Alice Rakotoarison, Kaushik Ram Sadagopan, Abinesh Ramakrishnan, Tuan Tran, Guillaume Wenzek, Yilin Yang, Ethan Ye, Ivan Evtimov, Pierre Fernandez, Cynthia Gao, Prangthip Hansanti, Elahe Kalbassi, Amanda Kallet, Artyom Kozhevnikov, Gabriel Mejia Gonzalez, Robin San Roman, Christophe Touret, Corinne Wong, Carleigh Wood, Bokai Yu, Pierre Andrews, Can Balioglu, Peng-Jen Chen, Marta R. Costa-jussà, Maha Elbayad, Hongyu Gong, Francisco Guzmán, Kevin Heffernan, Somya Jain, Justine Kao, Ann Lee, Xutai Ma, Alex Mourachko, Benjamin Peloquin, Juan Pino, Sravya Popuri, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, Anna Sun, Paden Tomasello, Changhan Wang, Jeff Wang, Skyler Wang, Mary Williamson

arXiv:2312.05187v1cs.CLcs.SDeess.AS

TL;DR

Speech translation can preserve semantic content while losing vocal style and prosody, limiting natural cross-lingual communication. This work introduces multilingual models for expressive, streaming translation and reports improved coverage, toxicity mitigation, and alignment capabilities, while acknowledging misuse risks.

  • Problem

    Speech translation may preserve semantic meaning while losing defining elements such as vocal style and prosody, limiting natural cross-lingual communication.

  • Method

    The paper develops SeamlessM4T v2 as a multilingual foundation, adds SONAR encoders and a UnitY2 char-to-unit aligner, and applies inference-time toxicity mitigation.

  • Results

    The models support nearly 100 input languages, the aligner covers 36 target languages, and toxicity is reduced by up to 80% on ETOX and 35% on MuTox versus SeamlessM4T-Large.

  • Takeaways & Limitations

    Multilingual coverage, expressive communication, and streaming capabilities are presented as tools for more ordinary real-time communication across language barriers.

  • Takeaways & Limitations

    The models remain vulnerable to unintended malicious uses such as voice phishing and deepfakes, so watermarking alone is not sufficient for safe deployment.

Abstract

from arXiv · show

Large-scale automatic speech translation systems today lack key features that help machine-mediated communication feel seamless when compared to human-to-human dialogue. In this work, we introduce a family of models that enable end-to-end expressive and multilingual translations in a streaming fashion. First, we contribute an improved version of the massively multilingual and multimodal SeamlessM4T model-SeamlessM4T v2. This newer model, incorporating an updated UnitY2 framework, was trained on more low-resource language data. SeamlessM4T v2 provides the foundation on which our next two models are initiated. SeamlessExpressive enables translation that preserves vocal styles and prosody. Compared to previous efforts in expressive speech research, our work addresses certain underexplored aspects of prosody, such as speech rate and pauses, while also preserving the style of one's voice. As for SeamlessStreaming, our model leverages the Efficient Monotonic Multihead Attention mechanism to generate low-latency target translations without waiting for complete source utterances. As the first of its kind, SeamlessStreaming enables simultaneous speech-to-speech/text translation for multiple source and target languages. To ensure that our models can be used safely and responsibly, we implemented the first known red-teaming effort for multimodal machine translation, a system for the detection and mitigation of added toxicity, a systematic evaluation of gender bias, and an inaudible localized watermarking mechanism designed to dampen the impact of deepfakes. Consequently, we bring major components from SeamlessExpressive and SeamlessStreaming together to form Seamless, the first publicly available system that unlocks expressive cross-lingual communication in real-time. The contributions to this work are publicly released and accessible at https://github.com/facebookresearch/seamless_communication

1. Introduction

Speech translation can preserve more than semantic meaning by retaining vocal identity, style, and prosody while operating across languages in real time. The Seamless family combines improved multilingual foundations, expressive translation, streaming translation, evaluation, and responsible AI safeguards.

  • Motivation: Speech translation may lose indexical and pragmatical features that help make human communication natural, including cues about a speaker’s personhood.The paper distinguishes semantic accuracy from preservation of vocal characteristics and socially situated communication.
  • Prior Work: Existing work developed expressive and streaming speech-to-speech translation separately, including approaches for transferring vocal style and style qualities.The related methods include speech language models, flow matching, diffusion models, and cascaded S2ST systems.
  • Contributions: SeamlessM4T v2 provides the multilingual and multimodal foundation for SeamlessExpressive and SeamlessStreaming, with improved semantic accuracy and nearly 100 supported input languages.Its updated UnitY2 framework, non-auto-regressive unit decoder, hierarchical upsampling, 4.5M hours of unlabeled-audio pretraining, and additional automatically aligned supervision support the foundation model.
  • Contributions: SeamlessExpressive preserves vocal style and prosody, while SeamlessStreaming generates low-latency translations without waiting for complete source utterances.The evaluation framework includes semantic, speech-quality, prosodic-consistency, and latency measures such as XSTS, MOS, PCP, and Ending Offset.
  • Responsible AI: The authors address responsible use through red-teaming, added-toxicity detection and mitigation, gender-bias evaluation, and the inaudible localized SeamlessWM watermark.These safeguards are accompanied by a metric card compiling evaluation and Responsible AI metrics.
  • Unified System: Combining expressive and streaming components, Seamless is presented as the first publicly available system for expressive cross-lingual communication in real time.The paper also publicly releases its models, data, code, metadata, evaluation tools, and responsible AI tools.

2. Beyond Words: Expressive and Streaming Speech-to-Speech Translation

Speech translation research has lagged in coverage, expressivity, and real-time speech-to-speech capability, despite users’ need for synchronous and natural communication. The paper responds with a unified direction centered on multilingual, expressive, streaming systems and systematic evaluation.

  • Motivation: Speech translation has lagged behind text translation in language coverage and performance, while speech’s paralinguistic features increase both its difficulty and its potential value.The envisioned system would translate expressively and in real time without requiring users to wait for completed sentences.
  • User Needs: Interviews with 34 participants from diverse immigrant backgrounds examined how people use translation technologies for essential information gathering and communication.The study focuses on individuals who depend on translation technologies in everyday life rather than primarily recreational users.
  • User Needs: Participants described consecutive speech translation as an imperfect workaround in time-sensitive situations and wanted reliable real-time systems for synchronous social communication.Real-time translation was especially relevant for face-to-face and digitally mediated conversations, including practical interactions such as ordering food or scheduling appointments.
  • Expressive Translation: Naturalness was commonly associated with preserving prosody and vocal style, which participants linked to personality, self-expression, and communicating intent beyond words.Faithful reproduction of vocal style and prosody was reported by 29 of 34 participants as a central conception of naturalness.
  • Implications: Participants viewed language coverage, expressivity, and streaming together as tools that could support everyday social integration without erasing individual identity.The paper connects this need to ordinary activities such as communicating with shopkeepers or arranging medical appointments.
  • Streaming Translation: The paper identifies streaming S2ST gaps including limited prior focus on speech-to-speech, ad hoc dependence on offline models, and cascaded-system errors, storage, and computation costs.Direct S2ST models are presented as a potential way to alleviate limitations of cascaded streaming approaches, especially as training scale increases.
  • Research Directions: The proposed research agenda develops datasets and foundational models for end-to-end multilingual real-time translation, expands language coverage and directions, and maintains systematic quality and safety evaluation.The stated goals include broader vocal-style and expressive preservation and translation from and into English for expressive and streaming systems.

3. SeamlessM4T v2

SeamlessM4T v2 is a multilingual, multimodal foundation model redesigned around UnitY2, expanded data, and improved alignment to support speech and text translation. It achieves strong results across translation and ASR tasks, including notable gains for low-resource languages.

  • Architecture: UnitY2 replaces autoregressive unit decoding with a non-autoregressive T2U decoder and hierarchical upsampling from subwords to characters to units.The architecture also uses an unsupervised multilingual character-to-unit aligner because forced aligners are unavailable for many low-resource languages.
  • Data and alignment: SeamlessM4T v2 expands training resources through improved Sonar speech encoders, 76-language coverage, 114,800 additional aligned-data hours, and 351K S2TT plus 145K S2ST training hours.The updated SeamlessAlign data doubled language coverage from 37 to 76 languages and improved data quantity, quality, and low-resource representation.
  • Architecture: 3x faster S2ST inference follows from non-autoregressive T2U decoding, which decouples generation from output-length prediction and reduces hallucination or truncation risks with partial input.The new NAR T2U architecture with pre-trained alignment also produced faster convergence, with the model converging in less than an epoch.
  • Evaluation: Performance gains extend across S2ST, ASR, and resource levels, while T2TT slightly declines by 1.6 chrF2++ but remains on par with similarly sized NLLB models.Low-resource languages improve by averages of 2.8 BLEU, 4.3 ASR-BLEU, and -7.5 WER across three tasks; medium-resource languages improve by 3.0, 4.5, and -4.7 respectively.

4. SeamlessExpressive

SeamlessExpressive combines prosody-aware speech-to-unit translation with a textless acoustic generator to preserve speech rate, pauses, vocal style, and semantic translation quality. Its evaluations show gains in expressive metrics and content translation across language directions, with data selection and generator choices shaping the trade-offs.

  • Expressive modeling: SeamlessExpressive uses SeamlessM4T v2 as its semantic backbone, Prosody UnitY2 for prosody-aware unit generation, and PRETSSEL for cross-lingual expressivity preservation.Prosody UnitY2 incorporates expressivity embeddings, while PRETSSEL disentangles semantic and expressivity components through unsupervised speech reconstruction pretraining.
  • Expressive speech-to-unit translation: Prosody UnitY2 improves ASR-BLEU over UnitY2 by +4.73 on mExpresso and +1.59 on mDRAL for five X–eng directions, but trails it by 1.34 on FLEURS.For five eng–X directions, the gains are +4.14 BLEU on mExpresso and +9.07 BLEU on mDRAL; both models are comparable on FLEURS eng–X data.
  • Expressive speech generator: PRETSSEL improves vocal style similarity and AutoPCP while keeping ASR-BLEU similar to the comparison model.On mDRAL, gains are +0.22 vocal style similarity and +0.56 AutoPCP for X–eng, and +0.27 vocal style similarity and +0.4 AutoPCP for eng–X.
  • Combining into SeamlessExpressive: The combined SeamlessExpressive model improves content translation and preserves rhythm, tone, vocal style, pauses, and speech rate, with consistent prosody gains across language directions.Figure 10 evaluates Pause, Rate, and AutoPCP by language on mDRAL test data.
  • Generator comparison: PRETSSEL offers a smaller and faster alternative to Unit Voicebox, with 65M versus 329M parameters and a reported real-time factor of 0.014 versus 0.089.Unit Voicebox provides higher vocal style similarity, while AutoPCP performance is similar across language pairs; a cascade using style-preserving TTS is more sensitive to noisy source speech.
  • Data filtering and ablations: Data choices affect semantic and prosodic outcomes differently: higher semantic-quality data raises ASR-BLEU by up to 8%, while prosodic filtering improves rate and pause metrics without hurting ASR-BLEU.Large parallel data improves content translation, while prosody-aligned parallel data maintains prosody preservation.

5. SeamlessStreaming

SeamlessStreaming extends SeamlessM4T v2 into simultaneous multilingual speech-to-speech and speech-to-text translation. It uses EMMA-based policies and staged fine-tuning to balance latency with translation quality across languages and settings.

  • SeamlessStreaming supports 101 source speech languages, 36 target speech languages, 96 target text languages, and streaming ASR for 96 languages.
  • Efficient Monotonic Multihead Attention: EMMA lets the decoder choose whether to generate the next token or consume more source context during streaming translation.Each attention head operates an individual simultaneous policy based on monotonic alignment estimated during training.
  • Efficient Monotonic Multihead Attention: EMMA uses a closed-form monotonic-attention estimation designed to improve numerical stability and avoid bias from a denominator used in prior work.
  • Policy regularization: Two regularization losses constrain latency and alignment variance because infinite-lookback attention can otherwise learn a trivial offline policy.Latency regularization uses expected delays, while variance regularization targets uncertainty in monotonic alignment.
  • Streaming fine-tuning: Streaming fine-tuning uses two stages: simultaneous speech-to-text training followed by text-to-unit training, initialized from SeamlessM4T v2 components.The speech encoder is frozen in stage one, and the speech-to-text portion is frozen in stage two.
  • Results: SeamlessStreaming achieves ASR with less than 10 WER degradation from SeamlessM4T v2 while providing much lower latency.
  • Results: High-resource languages show smaller quality drops and lower latency than distant language groups, while zero-shot settings show large quality drops alongside very small average lagging.The zero-shot pattern indicates an over-generation issue under that setting.

6. Seamless

Seamless combines SeamlessExpressive and SeamlessStreaming into a unified real-time expressive speech-to-speech system. Its broader language coverage introduces quality trade-offs, and partial context reduces expressivity preservation.

  • Seamless combines SeamlessExpressive and SeamlessStreaming to provide real-time expressive speech-to-speech translation.
  • Architecture: SeamlessStreaming generates streaming text and discrete units, which PRETSSEL converts into expressive speech using partial source context.
  • Language coverage: Seamless-36 extends SeamlessExpressive’s coverage to SeamlessStreaming’s language coverage, whereas Seamless-6 retains the smaller high-resource-language coverage.
  • Evaluation: Only automatic evaluations were conducted for Seamless.
  • Results: Expanding PRETSSEL from six to 36 languages lowers non-English target-language performance on ASR-BLEU and vocal style similarity, although broad support remains feasible.English output remains of similar quality overall.
  • Results: Seamless has a small ASR-BLEU degradation from SeamlessStreaming and a lower ending offset, while Seamless-6 performs better than Seamless-36 on six expressive target directions.
  • Expressivity preservation: Partial context causes degradation in vocal style similarity and AutoPCP compared with offline PRETSSEL measurements.

7. Automatic and Human Evaluation

The evaluation introduces automatic measures for prosodic preservation and assesses translation quality, robustness, expressivity, and speech quality across models. Results show stronger robustness and expressivity in several enhanced models, alongside trade-offs in some MOS dimensions.

  • Automatic Expressivity Metrics: AutoPCP evaluates sentence-level prosodic preservation, while the rhythm toolkit measures speech rate and pauses.Speech-rate and pause similarity moderately correlate with overall human judgments of prosody preservation.
  • Robustness Automatic Evaluation: 53.2% and 50.5% average improvements over Whisper-Large-v2 were observed for X–eng S2TT and ASR under varied noise conditions.SeamlessM4T-Large v2 also improved over SeamlessM4T-Large by 14.9% and 14.5% on the same tasks.
  • Robustness Automatic Evaluation: 66.4% average improvement in CoefVarMS and 31.5% in chrFMS over Whisper-Large-v2 demonstrated stronger robustness to vocal style variations.The comparison covered X–eng S2TT and ASR tasks, and SeamlessM4T-Large v2 also outperformed the earlier SeamlessM4T-Large model.
  • Human Evaluation: PRETSSEL improved PCP scores for rhythm, emotion, and OEI, but reduced clarity of speech and sound quality MOS scores.Across languages and datasets, the reported PCP gains were δ=0.58, δ=0.41, and δ=0.44, while MOS declined by δ=-0.49 and δ=-0.79 for clarity and sound quality.
  • Human Evaluation: Unit Voicebox improved clarity of speech and sound quality MOS scores without improving expressivity preservation.The reported gains were δ=0.26 for clarity of speech and δ=0.55 for sound quality.
  • Analysis of Expressive Models: Expressivity-preserving models retained acoustic characteristics more than SeamlessM4T v2, while sensitivity to such features could also increase noise-related quality penalties.PRETSSEL models preserved HNR most strongly; HNR and SNR were positively associated with sound quality, whereas shimmer and jitter related positively to naturalness.

8. Responsible AI

The paper combines red-teaming, toxicity detection and mitigation, gender-bias evaluation, and watermarking to assess and strengthen the safety of its multilingual models. These efforts identify toxicity and gender-related limitations while reporting mitigation and detection results.

  • Toxicity: Toxicity is the most prevalent critical-error category, but approximately 25% of toxicity instances are added toxicity, 48% deleted toxicity, and the remainder vary in intensity.The analysis distinguishes toxicity introduced by translation from toxicity removed or changed in intensity.
  • Red-teaming: The authors introduce red-teaming for multimodal machine translation and quantify successful challenges for SeamlessM4T v2 and SeamlessExpressive.The red-teaming methodology covers multilingual conditional generative AI systems.
  • Toxicity mitigation: MinTox mitigates added toxicity at inference time by detecting toxicity and filtering problematic multi-token expressions from beam-search hypotheses when input toxicity is absent.The workflow leaves translations unchanged when no output toxicity is detected and does not mitigate cases where toxicity is already present in the input.
  • Toxicity mitigation: SeamlessM4T v2 with MinTox reduces added toxicity by up to 80% in S2TT versus the same model without mitigation and by up to 90% versus SeamlessM4T-Large.Across the reported detection metrics, the model's added-toxicity prevalence is below 0.1% with ETOX and below 3.5% with MuTox.
  • Gender bias: Gender evaluation finds stronger robustness across metrics and tasks, but masculine overgeneralization is not consistently improved and increases by 0.2% in S2TT.Neutral English inputs can be translated into masculine forms, while gender-inflected source variants can produce different adjectives in translation.

9. Social Impact & Conclusion

The Seamless family combines multilingual, expressive, and streaming translation with safety mechanisms, supporting possible real-time and asynchronous cross-lingual communication applications. The authors also identify performance variability, misuse risks, and future needs in language coverage, fairness, and multimodal translation.

  • Conclusion: Seamless combines expressive and streaming models into a publicly available system for real-time expressive cross-lingual communication.The family is designed for end-to-end multilingual, expressive, and streaming translation.
  • Potential Applications: Real-time applications include expressive multilingual dialogue in audio or video calls, AR/VR interactions, online streaming, and wearable devices.The proposed uses include live speech translation, multilingual captions, and translated audio rendered through local or wearable interfaces.
  • Potential Applications: The models could also translate long-form audio, video, and voice messages while preserving multilingual expressivity.Suggested pipelines cover lectures, podcasts, video dubbing, and audio notes.
  • Social Impact: The authors connect these capabilities to broader possibilities for cross-lingual communication, while cautioning that Seamless is not a panacea for social integration.The discussion frames these effects as potential consequences rather than guaranteed outcomes.
  • Limitations: Performance may vary by race, accent, or gender, and ASR errors can degrade downstream expressive and streaming performance.Some users may need to alter their speech patterns to use the systems fully.
  • Future Work: Future work should improve low-resource language coverage and performance gaps while addressing diverse users and visual communication modalities.The authors specifically identify sign languages, gestures, facial expressions, and lip movements as areas needing further attention.

Contribution Statements

The paper documents contributions spanning data, modeling, evaluation, responsible AI, engineering, management, and editorial coordination. It emphasizes that the project involved extensive collaborative work across these areas.

  • Project Coordination: The project also relied on coordination, research management, data licensing, technical support, user research, design, and product roles.The statements recognize both research and operational contributions across the project.
  • Research and Modeling: Contributors worked on multimodal data alignment, speech and text modeling, model compression, benchmarking, and multilingual inference.Named contributions include UnitY2, speech encoders, expressive decoding, streaming models, vocoders, and evaluation.
  • Expressive Modeling: The team developed expressive components involving prosody modeling, PRETSSEL training, controllable TTS data, and vocoder research.These contributions supported the expressive modeling and speech-generation work described in the project.
  • Streaming Modeling: Streaming contributions covered streaming TTS and encoders, multilingual modeling and inference, ASR, and the core streaming algorithm.Several contributors held technical-lead or research roles in streaming development.
  • Responsible AI: Responsible AI work included toxicity classification and mitigation, gender-bias research, red teaming, watermarking, and metric cards.The contribution list identifies research, engineering, annotation, and editorial roles for these activities.

B. Model Card - SeamlessM4T v2

SeamlessM4T-Large v2 is a research-oriented multilingual and multimodal translation model supporting speech, text, and speech synthesis across many language directions. Its model card specifies broad capabilities, evaluation settings, and deployment boundaries.

  • Model Description: SeamlessM4T-Large v2 uses a multitask-UnitY2 architecture with Conformer and Transformer components plus a non-autoregressive T2U decoder.The model card identifies the speech encoder, text encoder-decoder, and T2U architecture.
  • Capabilities: The model supports ASR for 96 languages and speech-to-speech translation from 100 source speech languages into 35 target speech languages.These are listed among the model’s intended capabilities.
  • Capabilities: It supports speech-to-text translation from 100 source speech languages into 95 target text languages and text-to-text translation across 95 source and 95 target languages.The card also lists text-to-speech translation from 95 source text languages into 35 target speech languages.
  • Capabilities: The model provides text-to-speech synthesis for 36 languages.This capability is listed separately from translation tasks.
  • Limitations: The model is intended for research rather than production or certified translation, and it is not intended for domain-specific medical or legal inputs.The card also cautions that testing was limited across domains and language variations may not be captured.
  • Evaluation: Evaluation covers Fleurs, Flores, CoVoST2, CVSS, HolisticBias, and Multilingual HolisticBias using task-specific translation, recognition, and bias metrics.Reported metrics include BLEU, spBLEU, BLaSER 2.0, ASR-BLEU, chrF, and WER.

C. Model Card - SeamlessExpressive

SeamlessExpressive is a research model for expressive multilingual speech-to-speech translation that preserves prosody and vocal style. Its model card describes supported language directions, evaluation measures, and important scope limitations.

  • Model Description: SeamlessExpressive uses a Prosody UnitY2 model with PRETSSEL and two HiFi-GAN mel-vocoders operating at 16 kHz and 24 kHz.The card provides this as the model type and component configuration.
  • Capabilities: SeamlessExpressive-M2M translates from five source languages into English and from English into five target languages.The model card identifies it as a multilingual expressive speech-to-speech translation model.
  • Capabilities: The model preserves prosodic rhythm, speech rate, pauses, and vocal style during speech-to-speech translation.These capabilities are explicitly listed in the intended-use description.
  • Limitations: The suite is intended for research, not production deployment or certified translations, and it is not intended for domain-specific inputs such as medical content.The models were trained on general-domain data.
  • Evaluation: Evaluation measures content preservation, vocal-style and prosody preservation, speech-rate correlation, pause alignment, speech quality, and human-rated expressive consistency.Automatic measures include ASR-BLEU and AutoPCP, alongside PCP and MOS protocols.
  • Responsible Use: The authors note that toxic, biased, or false outputs remain a concern and recommend additional integrity mitigations for added toxicity.This recommendation applies when using the model in a research application.

D. Model Card - SeamlessStreaming

SeamlessStreaming is a multilingual streaming translation model intended for simultaneous speech recognition and speech translation across many languages. It is released as a research model with documented evaluation measures and safety considerations.

  • Method: The model uses Efficient Monotonic Multihead Attention as its simultaneous translation algorithm.
  • Intended Use: SeamlessStreaming supports streaming automatic speech recognition across 96 languages and simultaneous speech translation from 101 source languages.
  • Scope: SeamlessStreaming is intended for researchers and the machine translation research community, not production deployment.
  • Metrics: Latency is measured with Average Lagging and Length-Adaptive Average Lagging for text output, and Ending Offset for speech output.
  • Evaluation Data: Fleurs provides the evaluation data because it contains an n-way parallel speech and text dataset covering 101 languages.
  • Limitations and Risks: Researchers are advised to add integrity mitigations for added toxicity, and low-resource-language deployment may increase exposure to misinformation or online scams.

H. Metric Card

This metric card describes the evaluation framework and inference-speed analysis for SeamlessM4T models. It covers model size, translation data, per-language evaluation, and the comparison between autoregressive and non-autoregressive T2U decoding.

  • Evaluation Metrics: Evaluation combines automatic and human metrics for translation quality, speech quality, expressivity, and latency.The framework includes BLEU, ASR-BLEU, XSTS, MOS, PCP, and latency measures described across the evaluation setup.
  • Model Comparison: SeamlessM4T v2 uses UnitY2 with a non-autoregressive T2U component, while the earlier model uses autoregressive T2U decoding.
  • Experimental Setup: The inference-speed experiment translated 1000 Fleurs S2ST audios using one A100 GPU, 96 CPUs, batch size 1, and beam width 5.
  • Inference Speed: SeamlessM4T v2 is significantly faster than SeamlessM4T v1 for S2ST inference because its T2U conversion time is independent of input length.In the earlier model, T2U decoding scales linearly with generated S2TT text length and accounts for 70% of inference time.
  • Data and Resources: The reported resources include multilingual speech encoders, automatically aligned data, ASR and S2TT data, and S2ST training data with per-language statistics.

I.3 Detailed results

The detailed-results section organizes per-language results for Fleurs speech-to-text and speech-to-speech translation in both X–eng and eng–X directions. Results are reported with BLEU for S2TT and ASR-BLEU for S2ST.

  • Results Organization: Per-language scores are reported across the evaluated Fleurs tasks, with CSV files shared in the Seamless Communication repository.
  • S2TT Results: Fleurs S2TT results cover both X–eng and eng–X language directions.The results are provided in Tables 65–68.
  • S2ST Results: Fleurs S2ST results cover both X–eng and eng–X language directions using ASR-BLEU.The results are provided in Tables 69–71.

J.1 Data

This section details evaluation datasets, expressive-speech resources, model components, and the equations and implementation used for monotonic alignment. It also identifies the latency and BLEU analyses for SeamlessStreaming.

  • Evaluation Data: mExpresso, mDRAL, and FLEURS provide empirical evaluation results, with descriptive statistics broken down by language pair.
  • Expressive Speech Modeling: Unit Voicebox was pretrained on speech and XLS-R units and finetuned on multilingual emotion data for Prosody UnitY2 training.
  • Monotonic Alignment: The monotonic-alignment derivation rewrites alignment computation through transition matrices so α can be computed by matrix multiplication.
  • Implementation: The implementation section provides a PyTorch code snippet for monotonic alignment using an extension probability matrix and iterative target-position updates.
  • Streaming Results: SeamlessStreaming evaluation compares BLEU and latency for S2TT and S2ST in both eng–X and X–eng directions.

L.1 PRETSSEL extension data

Table 77 reports the duration of the PRETSSEL-36 extended pretraining dataset in hours.

  • Table 77 organizes PRETSSEL-36 extended pretraining data duration in hours.

M. Automatic and Human Evaluation

The evaluation combines automatic metrics with human judgments of semantic and expressive similarity across translated speech. It examines content, emotion, rhythm, and overall expressive intent, while documenting a voice-overlap quality issue in the evaluation setup.

  • Automatic evaluation uses acoustic-correlate correlations and correlations between human PCP and automatic expressivity metrics.
  • Human Evaluation: Human evaluation asks listeners to compare audio pairs across languages for semantic similarity and expressive similarity.
  • Human Evaluation: Overall expressive intent combines rhythm, emotion, emphasis, and intonation to assess whether cross-lingual utterances convey equivalent expressive information.
  • Human Evaluation: Semantic ratings range from completely different meanings to completely similar meanings, including paraphrases and exact translations.
  • Human Evaluation: Rhythm evaluation considers speed, pacing changes, pauses, and word lengthening or shortening.
  • Limitations: A quality-assurance issue delayed detection of strong voice overlap because the second participant’s audio was not replayed during re-enactment.
Loading 2312.05187v1…