Search arXiv⌕ Search

arXiv · 2505.17589

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Abstract

In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming speech synthesis and human-parity quality. Despite these advancements, CosyVoice 2 exhibits limitations in language coverage, domain diversity, data volume, text formats, and post-training techniques. In this paper, we present CosyVoice 3, an improved model designed for zero-shot multilingual speech synthesis in the wild, surpassing its predecessor in content consistency, speaker similarity, and prosody naturalness. Key features of CosyVoice 3 include: 1) A novel speech tokenizer to improve prosody naturalness, developed via supervised multi-task training, including automatic speech recognition, speech emotion recognition, language identification, audio event detection, and speaker analysis. 2) A new differentiable reward model for post-training applicable not only to CosyVoice 3 but also to other LLM-based speech synthesis models. 3) Dataset Size Scaling: Training data is expanded from ten thousand hours to one million hours, encompassing 9 languages and 18 Chinese dialects across various domains and text formats. 4) Model Size Scaling: Model parameters are increased from 0.5 billion to 1.5 billion, resulting in enhanced performance on our multilingual benchmark due to the larger model capacity. These advancements contribute significantly to the progress of speech synthesis in the wild. We encourage readers to listen to the demo at https://funaudiollm.github.io/cosyvoice3.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhihao Du, Changfeng Gao, Yuxuan Wang, Fan Yu, Tianyu Zhao, Hao Wang, Xiang Lv, Hui Wang, Chongjia Ni, Xian Shi, Keyu An, Guanrou Yang, Yabin Li, Yanni Chen, Zhifu Gao, Qian Chen, Yue Gu, Mengzhe Chen, Yafeng Chen, Shiliang Zhang, Wen Wang, Jieping Ye. 2025-05-27. CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training. https://arxiv.org/abs/2505.17589

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Retain-Free Machine Unlearning for Speech Emotion Recognition

Speech Emotion Recognition (SER) infers a speaker's emotional state from speech and is increasingly deployed in human-computer interaction, education, and healthcare. Because speech also carries sensitive personal information, speakers may ask that some of their recordings be deleted, which requires removing the influence of those samples from an already trained SER model. Most machine unlearning methods can meet this request only with access to the remaining training data alongside the samples to be forgotten; this is impractical when the remaining data cannot be redistributed or has itself been deleted, and it adds storage and computation as the data grows. To this end, we propose a retain-free unlearning method that updates a pre-trained SER model using only the forget set. Our key idea is to synthesise adversarial samples from the forget set as a surrogate for the unavailable remaining data, and to constrain each parameter update by its estimated importance so that forgetting does not erase general knowledge. The experiments over several emotional-speech corpora and self-supervised backbones show that our method drives forget-set performance down to near chance while retaining much of the model's utility on the remaining and unseen-speaker data, narrowing the gap to methods that rely on the remaining set.

cs.SD↗

SHINE: Sequential Hierarchical Integration Network for EEG and MEG

How natural speech is represented in the brain constitutes a major challenge for cognitive neuroscience. Reconstructing the speech envelope and Mel spectrogram from EEG and MEG provides a time-resolved way to study its temporal and spectral structure. Speech-related neural activity spans sensors and temporal scales; extracting these representations while adapting the use of context to each acoustic target is a central problem in speech reconstruction. We propose SHINE, a Sequential Hierarchical Integration Network for EEG and MEG. A residual sensor adapter unifies input dimensions, intermediate dilated-block states retain temporal depth, and a target- and time-dependent gate fuses local hierarchical and attention-enhanced context predictions. Across two EEG and two MEG datasets, SHINE has the highest mean envelope and mean-Mel Pearson correlations among nine local baseline implementations on all eight dataset-metric combinations. SHINE also placed second in the speech-detection Extended Track of the NeurIPS 2025 PNPL Competition. Code will be released at https://github.com/xuxiran/SHINE.

cs.SD↗

Self-Distilled Pronunciation and Accent Control for Neural Text-to-Speech

Text-to-speech that reads raw text has no lexicon: a rare word is read as guessed, and a native Japanese listener accepts a word only if its reading and pitch accent are both right. A known remedy installs a reading-and-accent channel into a released model, but it needs many recordings. This paper removes the recordings: the frozen backbone reads a sentence containing a common word it already says correctly, and that output serves as the teacher for the same sentence with the word replaced by an annotated reading with its pitch accent. Screened raters judged the tag right on 0.80 to 0.93 of unseen difficult words on four backbones spanning autoregressive, diffusion, and encoder-decoder synthesis; plain kana, which cannot express an accent, got 0.38 to 0.60. On words needing no edit, naturalness is non-inferior on one backbone; on the other three, listeners prefer the unedited rendition by 0.19 to 0.26.

cs.SD↗