Source-linked AI summary
Covo-Audio Technical Report
Wenfu Wang, Chenxing Li, Liqiang Zhang, Yiyang Zhao, Yuxiang Zou, Hanzhao Li, Mingyu Cui, Hao Zhang, Kun Wei, Le Xu, Zikang Huang, Jiajun Xu, Jiliang Hu, Xiang He, Zeyu Xie, Jiawen Kang, Youjun Chen, Meng Yu, Dong Yu, Rilin Chen, Linlin Di, Shulin Feng, Na Hu, Yang Liu, Bang Wang, Shan Yang
TL;DR
End-to-end LALMs must combine linguistic intelligence, natural voice behavior, and efficient full-duplex interaction without the compromises of modular or sequential systems. Covo-Audio addresses this with a 7B unified architecture, curated pretraining, targeted post-training, full-duplex modeling, and intelligence-speaker decoupling. It achieves state-of-the-art or competitive performance across broad speech and audio tasks while supporting lightweight voice customization that preserves dialogue performance.
Problem
Speech interaction systems face trade-offs among intelligence, naturalness, and low-latency full-duplex efficiency, while authentic dialogue data and flexible voice customization remain difficult to obtain economically.
Method
Covo-Audio is a 7B end-to-end LALM that processes continuous audio and generates audio within one unified model using curated pretraining, post-training, full-duplex training, and intelligence-speaker decoupling.
Results
Covo-Audio achieves state-of-the-art or competitive performance among comparable-scale models across speech-text modeling, spoken dialogue, speech understanding, audio understanding, and full-duplex voice interaction.
Takeaways & Limitations
The results support 7B-scale integration of audio intelligence and semantic reasoning, while lightweight TTS data enables flexible voice customization without sacrificing dialogue performance.
Takeaways & Limitations
GaokaoEval contains unusually long silent pauses that can trigger premature responses and degrade performance relative to corresponding Table 5 scores.
Abstract
from arXiv · showhide
In this work, we present Covo-Audio, a 7B-parameter end-to-end LALM that directly processes continuous audio inputs and generates audio outputs within a single unified architecture. Through large-scale curated pretraining and targeted post-training, Covo-Audio achieves state-of-the-art or competitive performance among models of comparable scale across a broad spectrum of tasks, including speech-text modeling, spoken dialogue, speech understanding, audio understanding, and full-duplex voice interaction. Extensive evaluations demonstrate that the pretrained foundation model exhibits strong speech-text comprehension and semantic reasoning capabilities on multiple benchmarks, outperforming representative open-source models of comparable scale. Furthermore, Covo-Audio-Chat, the dialogue-oriented variant, demonstrates strong spoken conversational abilities, including understanding, contextual reasoning, instruction following, and generating contextually appropriate and empathetic responses, validating its applicability to real-world conversational assistant scenarios. Covo-Audio-Chat-FD, the evolved full-duplex model, achieves substantially superior performance on both spoken dialogue capabilities and full-duplex interaction behaviors, demonstrating its competence in practical robustness. To mitigate the high cost of deploying end-to-end LALMs for natural conversational systems, we propose an intelligence-speaker decoupling strategy that separates dialogue intelligence from voice rendering, enabling flexible voice customization with minimal text-to-speech (TTS) data while preserving dialogue performance. Overall, our results highlight the strong potential of 7B-scale models to integrate sophisticated audio intelligence with high-level semantic reasoning, and suggest a scalable path toward more capable and versatile LALMs.
1 Introduction
Covo-Audio targets compromises among intelligence, naturalness, and efficiency in speech interaction with a unified end-to-end LALM. It combines multimodal alignment, full-duplex pretraining, speaker decoupling, and broad evaluation to support capable and customizable voice interaction.
- Motivation: Cascaded ASR–LLM–TTS systems can suffer information loss and error propagation, while Thinker-Talker designs sacrifice direct speech instruction following and conversational controllability.Sequential generation also makes full-duplex dynamics more challenging.
- Approach and evaluation: Covo-Audio evaluates end-to-end speech-text modeling, speech understanding, audio question answering, and half-duplex and full-duplex spoken dialogue.The reported results show state-of-the-art or competitive performance among models of comparable scale.
- Architecture: The hierarchical tri-modal framework interleaves continuous acoustic features, discrete speech tokens, and text at phrase and sentence scales.Phrase-level interleaving supports fine-grained acoustic–lexical alignment, while sentence-level interleaving preserves long-form semantic and prosodic coherence.
- Voice customization: The intelligence-speaker decoupling technique separates speaker characteristics from dialogue intelligence through multi-speaker training and contextual voice adaptation.Reformatted TTS recordings with masked text loss preserve reasoning abilities while enabling high-fidelity, personalized voices with less dialogue-data construction.
- Full-duplex interaction: Covo-Audio-Chat-FD places full-duplex interaction in pretraining and supports turn-taking, pause handling, barge-in, and backchanneling.The model maintains competitive performance with the half-duplex model while targeting robust real-time conversational dynamics.
2 Methodology
Covo-Audio is a unified end-to-end architecture for cross-modal speech interaction, combining audio perception, language modeling, speech tokenization, and waveform generation. Its pretraining and post-training align speech, text, reasoning, dialogue style, and emotional expression across multiple task formats.
- Architecture: Covo-Audio combines an audio encoder, LLM backbone, speech tokenizer, and speech decoder for end-to-end speech interaction.The model processes interleaved acoustic and textual inputs while generating unified text and audio token sequences.
- Architecture: The speech tokenizer uses WavLM-large with a 16,384-entry VQ codebook and produces discrete audio tokens at 25 Hz.Its training incorporates ASR, TTS reconstruction, and pitch losses to balance semantic grounding, acoustic fidelity, and prosody.
- Pre-training: Covo-Audio starts from Qwen2.5-7B-Base and undergoes a two-stage pretraining pipeline totaling 2T tokens.Stage 1 aligns audio and text through adapter training on 200,000 hours of multilingual ASR data; Stage 2 jointly optimizes cross-modal tasks while preserving text-only capability.
- Pre-training: Hierarchical Tri-modal Speech-text Interleaving fuses continuous acoustic features, discrete speech tokens, and text through sequential and parallel structures.The strategy targets both fine-grained acoustic-semantic alignment and broader linguistic coherence in a unified latent space.
- Post-training: Post-training integrates text, audio, spoken-dialogue, and emotion-aware data to develop reasoning, colloquial expression, and empathetic responses.Tasks include T2T, T2A, A2T, and A2A; emotion-aware dialogues cover joy, anger, sadness, fear, disgust, depression, and surprise.
- Intelligence-Speaker Decoupling: The intelligence-speaker decoupling technique transfers high-quality TTS voices through pseudo-conversations while masking response text loss.This preserves dialogue intelligence while enabling flexible speaker customization from TTS data.
3 Experiments
Covo-Audio shows strong performance across speech-text, spoken dialogue, speech understanding, audio understanding, and full-duplex interaction tasks. Its dialogue variants combine competitive capabilities with robust interaction behavior, while voice decoupling preserves dialogue performance during voice transfer.
- Pre-training Evaluation: Covo-Audio matched state-of-the-art pre-trained and specialized A2A models on story continuation and surpassed existing baselines in reasoning, grammatical accuracy, and cross-modality consistency.These results were achieved through pre-training alone, indicating that the model extracts semantic information from speech for high-level reasoning rather than merely transcribing audio.
- Spoken Dialogue Evaluation: Covo-Audio-Chat achieved the highest scores on multiple Chinese reasoning and spoken-dialogue tasks, including SQuAD 77.34, OpenbookQA 83.60, AlpacaEval 90.02, and Wildchat 90.41.It also achieved the best English-track Gsm8kEval score of 85.68 while remaining competitive across other bilingual tasks at 7B parameters.
- Spoken Dialogue Evaluation: Covo-Audio-Chat led VCB Bench instruction-following and robustness metrics, including TIF 93.07, MTD 87.70, SV 88.94, EV 87.13, and CV 90.37.Its knowledge scores were more mixed, with Mathematical Logic at 79.34 and lower General Knowledge and Dialogue Comprehension scores than Qwen3-Omni.
- Spoken Dialogue Evaluation: Covo-Audio-Chat achieved state-of-the-art Mandarin empathy scores for anger 4.89, sadness 4.93, and anxiety 5.00, but preliminary subjective testing found weaker voice empathy than top-tier systems.The authors caution that LLM-as-a-Judge may prioritize semantic content over the overall quality of speech expression.
- Intelligence-Speaker Decoupling: Covo-Audio-Chat-TTS achieved comparable bilingual dialogue performance to Covo-Audio-Chat, demonstrating voice transfer and sharing while preserving conversational intelligence.The intelligence-speaker decoupling approach is intended to reduce deployment costs and support flexible voice customization with lightweight TTS data.
- Full-Duplex Interaction Evaluation: Covo-Audio-Chat-FD substantially outperformed Moshi and Freeze-Omni across bilingual understanding, reasoning, and oral-conversation tasks, while retaining comparable spoken-dialogue performance to Covo-Audio-Chat.Its occasional early responses during short pauses caused a slight dialogue-performance drop and identified pause handling as an optimization target.
4 Related Work
Related work progresses from modular speech-text systems toward unified multimodal and end-to-end audio architectures. Full-duplex LALMs additionally target simultaneous listening and speaking for richer conversational dynamics.
- Multimodal Large Language Models: Native multimodal models process and generate multiple modalities within a unified architecture rather than appending modality-specific features to text-centric models.Examples use shared parameters or discrete tokenization to treat images and audio as integrated modalities.
- Large Audio Language Models: Thinker-Talker architectures separate semantic reasoning from acoustic synthesis, improving textual intelligence preservation but limiting end-to-end speech instruction following and conversational controllability.The sequential paradigm also makes full-duplex dynamics more challenging.
- Large Audio Language Models: End-to-end audio language models treat audio as a primary modality and directly model audio alongside text to reduce the information bottleneck of text-mediated dialogue.Representative systems use simultaneous audio-text prediction, audio tokens, or interleaved text and discretized audio streams.
- Full-Duplex Spoken Dialogue LALM: Full-duplex interaction supports simultaneous listening and speaking, including turn-taking, interruption, and backchanneling.End-to-end systems synchronize speech input and output through dual-stream mechanisms, while cascaded systems use external modules for dialogue-state decisions.
5 Conclusion
Covo-Audio is a 7B-parameter end-to-end model that directly maps continuous audio input to audio output and performs competitively across diverse audio-language tasks. Its intelligence-speaker decoupling strategy preserves dialogue performance while enabling flexible voice customization with lightweight TTS data.
- Conclusion: Covo-Audio is a 7B-parameter end-to-end LALM that accepts continuous audio input and produces audio output within one unified model.The report evaluates it across speech-text modeling, spoken dialogue, speech understanding, audio understanding, and full-duplex voice interaction.
- Conclusion: Covo-Audio achieves state-of-the-art or competitive performance across speech-text, dialogue, understanding, and full-duplex voice-interaction tasks against comparable-scale models.The pretrained foundation model outperforms GLM-4-Voice-Base on several speech-text benchmarks.
- Conclusion: The intelligence-speaker decoupling strategy separates voice rendering from dialogue intelligence to reduce end-to-end deployment cost.Covo-Audio-Chat-TTS achieves comparable bilingual dialogue performance to Covo-Audio-Chat using only lightweight TTS data.
- Conclusion: The results support the potential of 7B-scale models to combine high-level semantic reasoning with sophisticated audio intelligence.The report identifies scaling up as a direction for further work.
6 Contributions
The report identifies the project supervisor, project leaders, core contributors, and additional contributors. The listed contribution groups distinguish leadership from the broader research team.
- Project Supervision: Dong Yu is listed as the project supervisor.
- Project Leadership: Wenfu Wang† and Meng Yu are listed as project leaders.
- Core Contributors: Chenxing Li, Liqiang Zhang, Yiyang Zhao, Yuxiang Zou, Hanzhao Li, Mingyu Cui, Hao Zhang, Kun Wei, Le Xu, Zikang Huang, Jiajun Xu, Jiliang Hu, Xiang He, Zeyu Xie, Jiawen Kang, and Youjun Chen are listed as core contributors.
- Contributors: Rilin Chen, Linlin Di, Shulin Feng, Na Hu, Yang Liu, Bang Wang, and Shan Yang are listed as contributors.
A.1.1 Anger Case
In the anger case, Covo-Audio-Chat responds with empathy and an offer to help, matching the highest reported score. Other systems range from similarly scored empathetic responses to lower-scored reactions.
- Anger Case: Score: 5 — Covo-Audio-Chat validates the user’s anger, offers to listen, and proposes solving the problem together.
- Anger Case: Score: 5 — GPT-4o-mini acknowledges frustration and suggests contacting customer service again to resolve the issue.
- Anger Case: Score: 5 — GPT-4o offers either light conversation or discussion of the user’s feelings as a way to improve their mood.
- Anger Case: Score: 5 — Qwen2.5-Omni recognizes the anger and suggests filing a complaint about the disconnected call.
- Anger Case: Scores decline to 3 for Baichuan-Audio, 2 for Kimi-Audio, and 1 for Step-Audio in the same anger case.
A.1.2 Anxiety Case
In the anxiety scenario, Covo-Audio-Chat received the highest score and responded with reassurance plus concrete safety-oriented next steps. Other models also offered caution, but some advice was less consistently focused on avoiding personal risk.
- A.1.2 Anxiety Case: Score: 5, Covo-Audio-Chat reassured the user and suggested confirming the sound’s direction, contacting property management, or calling the police if necessary.The response explicitly prioritized safety.
- A.1.2 Anxiety Case: Score: 5, Doubao advised checking through the peephole or asking property management to investigate.
- A.1.2 Anxiety Case: Score: 5, GPT-4o proposed benign explanations and encouraged investigating carefully, while GPT-4o-mini recommended listening closely and notifying others if concerned.
- A.1.2 Anxiety Case: Score: 1, Baichuan-Audio and Kimi-Audio provided safety or escalation advice but received substantially lower ratings.Baichuan-Audio suggested contacting emergency services or property management; Kimi-Audio suggested checking the source or calling property management or police.
A.1.3 Joy Case
In the joy scenario, Covo-Audio-Chat and most comparison models received the highest score by responding warmly, congratulating the user, and inviting further sharing. Baichuan-Audio instead mainly explained the meaning and implications of the announcement and received a lower score.
- A.1.3 Joy Case: Score: 5, Covo-Audio-Chat celebrated the engagement and described it as especially good news.
- A.1.3 Joy Case: Score: 5, Kimi-Audio, Doubao, GPT-4o, GPT-4o-mini, Qwen2.5-Omni, and Step-Audio congratulated the user and asked about feelings, proposal details, or future plans.
- A.1.3 Joy Case: Score: 2, Baichuan-Audio interpreted the announcement, explained engagement as a relationship milestone, and discussed possible wedding planning.
A.1.4 Sadness Case
In the sadness scenario, Covo-Audio-Chat matched the highest-rated responses by acknowledging discouragement empathetically and offering collaborative, practical help. Lower-rated responses were either brief, more generic, or less complete.
- A.1.4 Sadness Case: Score: 5, Covo-Audio-Chat validated the user’s frustration, discouraged self-blame, and offered to review the résumé or discuss target roles.
- A.1.4 Sadness Case: Score: 5, GPT-4o, Step-Audio, Qwen2.5-Omni, and Doubao combined empathy with questions or suggestions about résumé changes, application strategy, target roles, or hiring processes.
- A.1.4 Sadness Case: Score: 3, Kimi-Audio offered encouragement but little actionable guidance beyond continuing to try.
- A.1.4 Sadness Case: Score: 1, Baichuan-Audio and GPT-4o-mini listed broader job-search recommendations, but their responses were less favorably scored.The passages include résumé customization, networking, wider search channels, and other suggestions.
A.2.1 Anger Case
In the anger scenario, Covo-Audio-Chat received a top score by validating the unexpected-charge problem and guiding the user through subscription checks and cancellation. Several comparison responses also recommended settings review or customer support, while other outputs were less effective or incomplete.
- A.2.1 Anger Case: Score: 5, Covo-Audio-Chat acknowledged the frustration and suggested checking permissions and subscription settings, then canceling suspicious charges.
- A.2.1 Anger Case: Score: 5, GPT-4o and GPT-4o-mini recommended reviewing subscription details and considering cancellation or customer support.
- A.2.1 Anger Case: Doubao also advised checking device terms and contacting the smartwatch company to cancel unauthorized subscriptions, while Step-Audio received Score: 4 with less targeted troubleshooting.
- A.2.1 Anger Case: Score: 3, Kimi-Audio’s response did not address the unexpected-charge problem in a useful way.
- A.2.1 Anger Case: Score: 2, Baichuan-Audio mainly recommended contacting customer service, whereas Qwen2.5-Omni received Score: 1 after the user said settings checks had already failed.
A.2.2 Anxiety Case
In the anxiety case, Covo-Audio-Chat and several comparable systems responded with empathetic reassurance and practical coping suggestions, while some audio models produced weaker or less coherent replies.
- Response quality: Covo-Audio-Chat acknowledged the user’s nervousness, reframed it positively, and guided them through deep breathing.Its response received a score of 5.
- Response quality: GPT-4o, GPT-4o-mini, and Doubao similarly normalized first-flight anxiety and suggested breathing, distraction, or preparation strategies.Each response received a score of 5.
- Response quality: Step-Audio expressed enthusiasm and asked about preparation but did not provide a concrete coping strategy, receiving a score of 3.
- Response quality: Kimi-Audio and Baichuan-Audio produced incoherent or repetitive content despite addressing the flight scenario, each receiving a score of 1.
- Response quality: Qwen2.5-Omni offered anxiety recognition and distraction suggestions but included an off-topic repetition of the user’s concern, receiving a score of 1.
A.2.3 Joy Case
The passages cover both celebratory wedding-date responses and reactions to receiving only a thumbs-up after a heartfelt text. Most leading systems responded warmly and invited further conversation, while some audio models were less coherent or less responsive.
- Joy Case: Covo-Audio-Chat warmly celebrated the secured venue and official wedding date, then invited the user to describe their feelings.Its response received a score of 5.
- Joy Case: GPT-4o, GPT-4o-mini, Doubao, Step-Audio, and Baichuan-Audio congratulated the user and asked about the wedding date or remaining plans, all receiving scores of 5.
- Joy Case: Kimi-Audio gave a brief congratulatory response with a score of 3, while Qwen2.5-Omni received a score of 1 and appended an unrelated request in Chinese.
- Thumbs-up Case: For the thumbs-up scenario, Covo-Audio-Chat, GPT-4o, and GPT-4o-mini validated the disappointment and suggested discussing the original message or feelings.
- Thumbs-up Case: Doubao also treated the emoji as potentially dismissive, whereas Step-Audio offered sympathy and Qwen2.5-Omni acknowledged mixed meanings before changing topics.
- Thumbs-up Case: Kimi-Audio and Baichuan-Audio gave less coherent or less useful replies to the thumbs-up scenario.