Source-linked AI summary
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Ailin Huang, Boyong Wu, Bruce Wang, Chao Yan, Chen Hu, Chengli Feng, Fei Tian, Feiyu Shen, Jingbei Li, Mingrui Chen, Peng Liu, Ruihang Miao, Wang You, Xi Chen, Xuerui Yang, Yechang Huang, Yuxiang Zhang, Zheng Gong, Zixin Zhang, Hongyu Zhou, Jianjian Sun, Brian Li, Chengting Feng, Changyi Wan, Hanpeng Hu, Jianchang Wu, Jiangjie Zhen, Ranchen Ming, Song Yuan, Xuelin Zhang, Yu Zhou, Bingxin Li, Buyun Ma, Hongyuan Wang, Kang An, Wei Ji, Wen Li, Xuan Wen, Xiangwen Kong, Yuankai Ma, Yuanwei Liang, Yun Mou, Bahtiyar Ahmidi, Bin Wang, Bo Li, Changxin Miao, Chen Xu, Chenrun Wang, Dapeng Shi, Deshan Sun, Dingyuan Hu, Dula Sai, Enle Liu, Guanzhe Huang, Gulin Yan, Heng Wang, Haonan Jia, Haoyang Zhang, Jiahao Gong, Junjing Guo, Jiashuai Liu, Jiahong Liu, Jie Feng, Jie Wu, Jiaoren Wu, Jie Yang, Jinguo Wang, Jingyang Zhang, Junzhe Lin, Kaixiang Li, Lei Xia, Li Zhou, Liang Zhao, Longlong Gu, Mei Chen, Menglin Wu, Ming Li, Mingxiao Li, Mingliang Li, Mingyao Liang, Na Wang, Nie Hao, Qiling Wu, Qinyuan Tan, Ran Sun, Shuai Shuai, Shaoliang Pang, Shiliang Yang, Shuli Gao, Shanshan Yuan, Siqi Liu, Shihong Deng, Shilei Jiang, Sitong Liu, Tiancheng Cao, Tianyu Wang, Wenjin Deng, Wuxun Xie, Weipeng Ming, Wenqing He, Wen Sun, Xin Han, Xin Huang, Xiaomin Deng, Xiaojia Liu, Xin Wu, Xu Zhao, Yanan Wei, Yanbo Yu, Yang Cao, Yangguang Li, Yangzhen Ma, Yanming Xu, Yaoyu Wang, Yaqiang Shi, Yilei Wang, Yizhuang Zhou, Yinmin Zhong, Yang Zhang, Yaoben Wei, Yu Luo, Yuanwei Lu, Yuhe Yin, Yuchu Luo, Yuanhao Ding, Yuting Yan, Yaqi Dai, Yuxiang Yang, Zhe Xie, Zheng Ge, Zheng Sun, Zhewei Huang, Zhichao Chang, Zhisheng Guan, Zidong Yang, Zili Zhang, Binxing Jiao, Daxin Jiang, Heung-Yeung Shum, Jiansheng Chen, Jing Li, Shuchang Zhou, Xiangyu Zhang, Xinhao Zhang, Yibo Zhu
TL;DR
Open-source speech systems face challenges in unified understanding and generation, affordable speech-data acquisition, dynamic control, and intelligent interaction. Step-Audio addresses these gaps with a unified multimodal model, generative data engine, controllable speech synthesis, and cognitive capabilities, achieving state-of-the-art evaluations and a 9.3-point average improvement on selected open-source benchmarks.
Problem
Open-source speech systems struggle with separated understanding and generation, laborious speech-data acquisition, limited dynamic control, and constrained dialogue intelligence.
Method
Step-Audio combines a 130B-parameter unified speech-text model, generative speech-data engine, instruction-based voice control, and post-training for speech and AQTA tasks.
Results
9.3 points average improvement over the best open-source metrics, with state-of-the-art results across LLaMA Question, TrivialQA, ComplexBench, and nine StepEval-Audio-360 dimensions.
Takeaways & Limitations
Step-Audio provides an open-source framework for cross-modal speech-text interaction with fine-grained emotional, dialectal, and prosodic control.
Takeaways & Limitations
The current implementation does not yet provide native trimodal understanding or eliminate intermediate cross-modal conversions in AQAA scenarios.
Abstract
from arXiv · showhide
Real-time speech interaction, serving as a fundamental interface for human-machine collaboration, holds immense potential. However, current open-source models face limitations such as high costs in voice data collection, weakness in dynamic control, and limited intelligence. To address these challenges, this paper introduces Step-Audio, the first production-ready open-source solution. Key contributions include: 1) a 130B-parameter unified speech-text multi-modal model that achieves unified understanding and generation, with the Step-Audio-Chat version open-sourced; 2) a generative speech data engine that establishes an affordable voice cloning framework and produces the open-sourced lightweight Step-Audio-TTS-3B model through distillation; 3) an instruction-driven fine control system enabling dynamic adjustments across dialects, emotions, singing, and RAP; 4) an enhanced cognitive architecture augmented with tool calling and role-playing abilities to manage complex tasks effectively. Based on our new StepEval-Audio-360 evaluation benchmark, Step-Audio achieves state-of-the-art performance in human evaluations, especially in terms of instruction following. On open-source benchmarks like LLaMA Question, shows 9.3% average performance improvement, demonstrating our commitment to advancing the development of open-source multi-modal language technologies. Our code and models are available at https://github.com/stepfun-ai/Step-Audio.
1 Introduction
Step-Audio addresses open-source speech interaction limitations with unified speech-text modeling, generative data, fine-grained voice control, and enhanced intelligence. It reports state-of-the-art results in human evaluations and open-source benchmarks.
- Existing open-source speech systems separate understanding and generation, rely on laborious speech-data acquisition, and face limited dynamic control and intelligence.
- A 130B-parameter unified model integrates speech comprehension and generation for recognition, understanding, dialogue, voice cloning, audio editing, and synthesis.The Step-Audio-Chat variant is open source.
- A generative data engine produces audio for training the resource-efficient Step-Audio-TTS-3B model, reducing reliance on traditional manual data collection.The released TTS model emphasizes instruction-following controllable speech synthesis.
- Instruction-based voice control supports multiple emotions, dialects, and vocal styles including RAP, singing, and a cappella humming.
- Tool calling and role-playing enhancements improve agent performance on complex tasks.
- 9.3 points average improvement over the best open-source metrics is reported on LLaMA Question, TrivialQA, and ComplexBench.StepEval-Audio-360 evaluates nine dimensions of end-to-end speech dialogue.
- 29.8% and 27.1% improvements are reported for IF (Instruction Following) and MOS (Mean Opinion Score) against open-source SoTA models.The comparison covers generation-control dimensions including emotion understanding, speech-rate control, RAP vocals, and role-playing.
- Human evaluators compare Step-Audio, GLM-4-Voice, and Qwen2-Audio across nine dimensions using Likert scales for naturalness and task completion.The passage reports Step-Audio as SoTA across all evaluated dimensions, with particularly strong language and singing performance.
2 Related Work
Related speech systems progressed from cascaded pipelines to integrated and parallel end-to-end architectures, but persistent trade-offs remain in latency, capability preservation, emotional nuance, and conversational naturalness.
- Cascaded ASR-LLM-TTS systems suffer latency buildup, error propagation, and disjointed optimization.
- Adapter-based systems connect speech encoders to LLMs but still require separate TTS modules for audio output.
- Llama-Omni and Freeze-Omni integrate speech decoders with language models and improve latency, but remain limited in emotional nuance and natural conversational flow.
- Moshi and Mini-Omni reduce latency through parallel generation and compressed speech tokens, while facing challenges in preserving linguistic capabilities as speech-token bandwidth scales.
- Emotion-aware systems have explored sentiment analysis, but multimodal integration remains nascent and bidirectional emotional resonance is limited.The naturalness gap also persists because LLM outputs tend toward verbose, text-optimized responses.
3 Architecture
Step-Audio unifies speech understanding and synthesis through dual-codebook tokenization, a large multimodal language model, a controllable speech decoder, and a latency-oriented inference pipeline.
- The AQTA plus TTS design supports real-time voice dialogue while retaining flexible control over output timbre and pitch.The design addresses limited pure-voice dialogue data and controllability requirements.
- The architecture comprises a speech tokenizer, an LLM modeling text and speech tokens, and a speech decoder generating waveform output.
- 3 Architecture: Dual-codebook tokenization interleaves linguistic and semantic tokens at a 2:3 temporal ratio to represent structured linguistic and coarse acoustic information.The linguistic and semantic tokenizers operate at 16.7 Hz and 25 Hz, respectively.
- 3 Architecture: A 130B-parameter Step-1-based LLM receives audio-contextualized continual pretraining, while the decoder generates stylized waveforms from text or audio tokens.
- 3.3 Speech Decoder: The speech decoder combines a 3-billion-parameter language model, flow matching, and a mel-to-wave vocoder for waveform generation incorporating history and instructions.
- 3.4 Real-time Inference: The real-time pipeline coordinates VAD, streaming tokenization, the language model, speech decoding, state transitions, and text-based context management.
- 3.4 Real-time Inference: Approximately 40% of speculative responses are committed, reducing per-response latency by approximately 500ms versus non-speculative methods.
- 3.4 Real-time Inference: Text transcription compresses historical context to an average text-to-audio token ratio of 1:14, enabling longer conversations with minimal quality impact.
4 Pretrain
Step-Omni pretraining combines large-scale audio, text, and image data with staged multimodal training, efficiency-oriented infrastructure, and dual-codebook tokenizer exploration.
- The multimodal dataset includes audio continuation, TTS, ASR, and audio-text alternating data totaling large-scale token and hour counts.The audio resources comprise 1.1T continuation tokens, 113B TTS tokens, 105B ASR tokens, and 350B alternating-data tokens.
- Step-Audio is part of Step-Omni, a unified pretrained model for speech, image, and text built from a pretrained text model and image encoder.The training process is divided into three stages.
- Stage 1 adds 5,120 audio tokens and an image encoder while using a low 2e-5 backbone learning rate and higher embedding and LM-head rates.
- Stage 2 trains on 1.2T tokens with audio continuation and audio-text interleaved data in a 1:1 ratio, while audio, text, and image data follow a 2:1:1 ratio.
- Stage 3 adds ASR and TTS data in a 1:1:1:1 ratio after 800B Stage 2 tokens, with audio, text, and image data adjusted to 4:3:3.
- Disaggregated data processing and model placement improve training efficiency by reducing processing interference and pipeline bubbles from heterogeneous submodels.
- Single-codebook semantic tokens provide low next-token perplexity and good semantic coherence but discard acoustic information, harming audio restoration quality.
5.1 TTS
Step-Audio addresses scarce, costly TTS data through a synthetic data-driven framework that combines text rewriting, target-speaker audio generation, and emotion/style editing. The resulting system supports instruction-controlled speech across languages, dialects, emotions, and vocal styles.
- Synthetic data engine: The framework uses Step-2 to generate diverse text, Step-Audio to create target-speaker audio, and Audio-Edit to produce emotional and stylistic data.The pipeline is designed to reduce dependence on manually collected high-quality speech data.
- Emotion and speaking styles: The Audio-Edit model converts emotion and style descriptions into comparative training pairs for nuanced paralinguistic control while preserving speaker consistency.The approach addresses the difficulty of defining emotion categories, intensities, and speaking styles.
- Singing and RAP: The dataset includes more than 10,000 hours of timestamped singing and RAP tracks, with separated vocals, removed silence, and aligned lyrics.These processing steps support training for singing and RAP generation.
- Language and dialect: Dual codes from native-speaker audio are combined with target-speaker prompt audio to regenerate multilingual and dialect-specific speech.This approach targets native-speaker quality while requiring only a small quantity of high-quality seed data.
- Quality assessment: Synthetic data quality is assessed with ASR, VAD, speaker diarization, emotion consistency, and DNS metrics.The multi-dimensional checks target reliability, validity, robustness, and practical utility of generated data.
- Instruction control: Instruction tags separately control language, dialect, vocal attributes, style, emotion, and speech speed in the chat-based TTS format.The system uses descriptive tags for attributes such as dialect and style, and comparative tags for emotion and speed hierarchies.
5.2 AQTA
AQTA training combines diversified speech-text supervision with preference optimization to improve conversational response quality. The process constructs heterogeneous multimodal data, trains a reward model, and applies PPO to obtain Step-Audio-Chat.
- RLHF: Figure 6 depicts iterative response collection, manual and LLM evaluation, reward-model pair selection, and PPO training of Step-Audio-Chat.The procedure operationalizes RLHF for the final conversational model.
- Data construction: AQTA data pairs audio inputs with textual outputs, while TQTA, TAQTA, and other multimodal formats add text-speech consistency and training diversity.TAQTA uses text as both input and loss-bearing output; other formats include audio-audio and vision-audio-text interactions.
- Data processing: SFT processing filters concise single-turn inputs, conversationalizes outputs, and retains only final-turn speech and responses for multi-turn loss calculation.These choices focus training on realistic spoken inputs and the most recent conversational response.
- Preference optimization: Preference data uses human ratings and LLM-as-a-Judge scores to form chosen/rejected response pairs based on instruction following, naturalness, and safety.The preference pipeline uses real user audio prompts and multiple sampled responses.
- Reward modeling: The reward model is pretrained on TQTA preferences, fine-tuned on AQTA preferences, and reaches 70.51% pair-wise accuracy on a human preference test set.Training uses Bradley-Terry loss across the two stages.
- Observed limitation: AQTA preference-only reward modeling produced “deaf hacking,” rewarding “I didn’t hear clearly” regardless of audio clarity.The authors constructed corrective data and plan rule-based rewards to mitigate this bias.
6 Evaluation
The evaluation covers benchmark design, ASR, TTS, and real-time voice chat. Step-Audio reports strong results across speech recognition, synthesis, instruction following, factuality, relevance, and overall chat quality.
- Benchmark design: StepEval-Audio-360 evaluates language, emotion, reasoning, creativity, instruction following, role-playing, safety, demographics, environments, and prosody.Its indicators combine automatic scripts, LLM evaluation, and human assessment, with quarterly updates and user feedback.
- ASR: 25.5 to 18.4: Dual-Code reduces CER on ASR tasks while keeping the audio training-data amount fixed.The passage attributes the improvement to the Dual-Code approach.
- ASR: 4.64 average CER: Step-Audio Pretrain achieves the best result among audio-token speech models and reaches 2.05 average CER on clean test sets.Its clean-set result is close to Qwen2-Audio’s 2.06 average CER, which the authors connect to preserved semantic information from dual-codebook compression.
- ASR: 5.89 average CER: Step-Audio-Chat maintains strong instruction-following performance in the reported speech-content transcription evaluation.The compared 3B model struggled to follow the evaluation instruction effectively.
- TTS: Step-Audio-TTS-3B achieves open-source SoTA CER and WER on SEED test sets while remaining highly competitive in speaker similarity.Scaling to 130B parameters substantially improves both CER and WER in the reported comparison.
- Voice chat: Step-Audio-Chat scores 66.4% in factuality, 75.2% in relevance, and 4.11 in overall chat score on StepEval-Audio-360.The chat score uses a 1-to-5 scale and is assessed by GPT-4o from conversation text.
7 Conclusion
Step-Audio combines multimodal pretraining, task-specific post-training, RLHF, fine-grained speech control, and streaming engineering for real-time voice interaction. Benchmark results across ASR, TTS, and AQTA demonstrate strong speech-dialogue capabilities.
- Conclusion: Step-Audio uses a dual-codebook tokenizer and 3.3T multimodal tokens to align textual and acoustic modalities during pretraining.Post-training adds task-specific SFT for TTS and ASR plus diversified SFT and RLHF for AQTA.
- Conclusion: The framework supports fine-grained emotional modulation, dialect adaptation, and prosodic pattern generation while using speculative streaming and full-duplex coordination for fluid dialogue.These capabilities are described as engineering components of the real-time voice-interaction system.
- Conclusion: Evaluations across ASR, TTS, and AQTA demonstrate Step-Audio’s reported capabilities in speech dialogue.The conclusion presents these benchmark outcomes as evidence for the framework’s performance.
8 Future Work
Step-Audio’s demonstrated scope is currently limited to speech–text cross-modal integration within an initial trimodal-system implementation. Future work targets native vision–speech–text understanding, more efficient voice dialogue, and deeper tool calling.
- Native trimodal understanding should incorporate vision, speech, and text.
- Pure voice dialogue efficiency should improve by eliminating intermediate cross-modal conversions in AQAA scenarios.
- Deep-thinking-enhanced tool calls should strengthen intelligent interaction with external knowledge bases.