CosyVoice 3: Scaling Towards In-the-Wild Speech Generation
Explore CosyVoice 3's speech generation architecture, training approach, multilingual capabilities, and evaluation in this technical research overview.

Executive Summary
CosyVoice 3 represents a significant leap forward in zero-shot multilingual speech synthesis, designed specifically for real-world applications. Developed by Alibaba's Speech Team at Tongyi Lab, this model addresses the limitations of its predecessor CosyVoice 2 through massive scaling in both data (from 10K to 1M hours) and model parameters (from 0.5B to 1.5B), while introducing novel techniques for improved prosody naturalness and content consistency.
Key Innovations Overview
Diagram source
mindmap
root((CosyVoice 3))
Speech Tokenizer
Multi-task Training
MinMo Integration
FSQ Module
25Hz Token Rate
Post-training
DiffRO Method
Multi-task Rewards
Token-level Optimization
Data Scaling
1M Hours Total
9 Languages
18 Chinese Dialects
Real-world Audio
Model Scaling
1.5B Parameters
DiT Architecture
Enhanced CFMArchitecture Deep Dive
1. Multi-Task Speech Tokenizer
The foundation of CosyVoice 3's improved performance lies in its novel speech tokenizer, which builds upon the MinMo multimodal LLM rather than the SenseVoice-Large ASR model used in CosyVoice 2.
Diagram source
graph TD
A[Speech Input X] --> B[Voice Encoder1<br/>12 Transformer Blocks + RoPE]
B --> C[Intermediate Representations H]
C --> D[FSQ Module<br/>Finite Scalar Quantization]
D --> E[Voice Encoder2]
E --> F[MinMo LLM]
F --> G[Text Token Predictions]
D --> H[Speech Tokens μ<br/>25 Hz Rate]
I[Multi-task Training] --> F
I --> J[ASR]
I --> K[Language ID]
I --> L[Emotion Recognition]
I --> M[Audio Event Detection]
I --> N[Speaker Analysis]
style D fill:#e1f5fe
style I fill:#f3e5f5FSQ Quantization Process
The Finite Scalar Quantization (FSQ) module operates through a sophisticated two-step process:
- Dimensionality Reduction: Projects intermediate representations H into a D-dimensional low-rank space
- Bounded Quantization: Quantizes each dimension into the range [-K, K] using bounded round operations
Mathematical Formulation:
H̄ = ROUND(Proj_down(H))
Ĥ = Proj_up(H̄)
μᵢ = Σ(j=0 to D-1) h̄ᵢ,ⱼ × (2K + 1)ʲ
2. Differentiable Reward Optimization (DiffRO)
CosyVoice 3 introduces DiffRO, a novel post-training technique that optimizes speech tokens directly rather than synthesized audio, addressing computational challenges in traditional RL approaches.
Diagram source
sequenceDiagram
participant LLM as Language Model
participant GS as Gumbel-Softmax
participant T2T as Token2Text Model
participant Reward as Reward Calculator
LLM->>GS: Predicted Token Probabilities
GS->>T2T: Sampled Speech Tokens μ̃
T2T->>Reward: ASR Posterior Probability
Reward->>LLM: Gradient Signal
Note over LLM,Reward: Direct token optimization<br/>bypasses CFM/VocoderMulti-Task Reward (MTR) Mechanism
DiffRO extends beyond basic ASR rewards to include multiple downstream tasks:
- Speech Emotion Recognition (SER): Controls emotional expression
- MOS Score Prediction: Maintains audio quality
- Audio Event Detection (AED): Handles environmental sounds
- Speaker Analysis: Preserves speaker characteristics
Training Pipeline
The CosyVoice 3 training process follows a sophisticated multi-stage approach designed to maximize performance while maintaining stability.
Diagram source
graph LR
A[Large-scale Pretraining<br/>1M Hours] --> B[DiffRO Post-training<br/>Selected Data]
B --> C[Zero-shot LM & CFM]
C --> D[Continual Pretraining<br/>Text2Token LM]
D --> E[Speaker Fine-tuning<br/>Multi-speaker Data]
F[Text-based LLM<br/>Initialization] --> A
G[Emotional, Instructed,<br/>Multi-lingual Data] --> C
style A fill:#e8f5e8
style B fill:#fff3e0
style C fill:#e3f2fdStage Breakdown
- Initialization: Leverage pre-trained text-based LLMs for semantic understanding
- Large-scale Pretraining: Train on the massive 1M-hour multilingual dataset
- Post-training with DiffRO: Optimize performance using reward-based learning
- Continual Pretraining: Transfer capabilities to specialized models
- Speaker Fine-tuning: Enhance individual speaker quality and consistency
Dataset Scaling Analysis
CosyVoice 3's impressive performance stems significantly from its unprecedented dataset scale and diversity.
Diagram source
pie title Language Distribution (1M Hours Total)
"Chinese" : 45.2
"English" : 32.1
"Japanese" : 8.7
"Russian" : 6.4
"German" : 3.8
"Korean" : 2.1
"Others" : 1.7Chinese Dialect Coverage
Diagram source
pie title Chinese Dialect Distribution
"Sichuan" : 18.93
"Hubei" : 14.48
"Cantonese" : 8.61
"Wuzhong" : 7.34
"Shan1xi" : 8.54
"Suhang" : 6.64
"Shanghai" : 6.27
"Others" : 29.19Data Processing Pipeline
The multilingual data pipeline ensures high-quality training material through six critical steps:
Diagram source
graph TD
A[Raw Audio Data] --> B[Speech Detection &<br/>Segmentation]
B --> C[Noise Reduction<br/>MossFormer2]
C --> D[ASR Transcription<br/>Multi-model Validation]
D --> E[Punctuation Adjustment<br/>Montreal Forced Aligner]
E --> F[Volume Standardization<br/>0.6 Peak Normalization]
F --> G[Length Ratio Filtering<br/>Remove 1% smallest, 5% largest]
G --> H[Clean Training Data]
style C fill:#ffebee
style D fill:#e8f5e8
style G fill:#fff3e0Performance Benchmarks
SEED-TTS-Eval Results
CosyVoice 3 demonstrates substantial improvements over its predecessor and competitive models:
| Model | test-zh CER (%) | test-en WER (%) | test-hard CER (%) |
|---|---|---|---|
| CosyVoice 2 | 1.45 | 2.57 | 6.83 |
| CosyVoice 3-0.5B | 1.16 | 2.02 | 6.08 |
| CosyVoice 3-1.5B | 1.12 | 2.21 | 5.83 |
| CosyVoice 3-1.5B+RL | 0.71 | 1.45 | 5.66 |
Key Improvements:
- 44% relative improvement in Chinese content consistency
- 51% relative improvement in English content consistency
- 26% relative improvement on challenging test cases
CV3-Eval Multilingual Benchmark
CosyVoice 3 is the only system capable of handling all languages in the comprehensive CV3-Eval benchmark:
Diagram source
graph LR
A[CV3-Eval Benchmark] --> B[Multilingual Voice Cloning<br/>9 Languages × 500 Samples]
A --> C[Cross-lingual Transfer<br/>zh, en, ja, ko]
A --> D[Emotion Cloning<br/>Happy, Sad, Angry]
A --> E[Subjective Evaluation<br/>Expressive & Accent Cloning]
style B fill:#e8f5e8
style C fill:#fff3e0
style D fill:#f3e5f5
style E fill:#e3f2fdAdvanced Features
1. Pronunciation Inpainting
CosyVoice 3 addresses mispronunciations through mixed word-phoneme modeling:
Diagram source
graph TD
A[Raw Text Input] --> B{Contains Polyphonic<br/>Characters/Words?}
B -->|Yes| C[Replace with Phonemes<br/>Mixed Vocabulary]
B -->|No| D[Standard Processing]
C --> E[Enhanced Pronunciation<br/>Control]
D --> E
F[Auxiliary Training Set] --> G[Chinese Pinyin<br/>Replacement]
F --> H[English CMU Dict<br/>Phonemes]
G --> C
H --> C2. Self-Training for Text Normalization
The system eliminates hand-crafted rules through LLM-based text normalization:
Three-pronged Approach:
- Rule-based TN → Audio synthesis via CosyVoice 2
- Qwen-Max TN → Audio synthesis on normalized text
- Inverse TN → Raw text generation from existing pairs
3. Instructed Speech Generation
Extended from 1,500 to 5,000 hours of instruction-following data, supporting 100+ speaking styles:
Categories:
- Emotions: Happy, sad, angry, fearful, surprised, etc.
- Characteristics: Fast, slow, loud, soft, authoritative, etc.
- Roles: Warrior, poet, merchant, detective, etc.
- Dialects: 10 Chinese regional variants
- Accents: Indian English, Russian English, etc.
Technical Innovations Deep Dive
Model Architecture Enhancements
Diffusion Transformer (DiT) Integration
CosyVoice 3 adopts the DiT architecture for its Conditional Flow Matching (CFM) model:
Diagram source
graph TD
A[Speech Tokens<br/>25 Hz] --> B[Interpolation<br/>Rate Matching]
B --> C[DiT Backbone<br/>300M Parameters]
C --> D[Mel Features<br/>Generation]
D --> E[Vocoder<br/>Audio Output]
F[Text Encoder<br/>Removed] -.->|Simplified| C
G[Length Regularization<br/>Removed] -.->|Simplified| C
style C fill:#e1f5fe
style F fill:#ffebee,stroke-dasharray: 5 5
style G fill:#ffebee,stroke-dasharray: 5 5Scaling Impact Analysis
The transition from 0.5B to 1.5B parameters yields measurable improvements:
Diagram source
graph LR
A[Model Size] --> B[0.5B Parameters]
A --> C[1.5B Parameters]
B --> D[Good Performance<br/>Most Languages]
C --> E[Enhanced Performance<br/>Complex Scenarios]
D --> F[WER: 2.02% EN]
E --> G[WER: 1.45% EN<br/>with RL]
style C fill:#e8f5e8
style E fill:#e8f5e8
style G fill:#c8e6c9Speaker Fine-tuning Innovations
Monolingual to Polyglot Transformation
CosyVoice 3 can transform monolingual speakers into polyglots through targeted training:
Diagram source
sequenceDiagram
participant MS as Monolingual Speaker
participant AD as Auxiliary Dataset
participant LI as Language Instruction
participant PS as Polyglot Speaker
MS->>AD: Studio-quality samples
AD->>LI: "You are Speaker X. Please speak German."
LI->>PS: Cross-lingual capability
Note over MS,PS: Supports 9 languages<br/>18 Chinese dialectsCapability Transfer Mechanism
The fine-tuning process preserves pre-trained capabilities while adapting to specific speakers:
- Partial Speaker ID Labeling: Mix labeled and unlabeled data
- Instruction Masking: Randomly mask speaker/style prompts
- Catastrophic Forgetting Prevention: Maintain instruction coverage
Performance Analysis & Ablations
Speech Tokenizer Comparison
Diagram source
graph TD
A[Tokenizer Comparison] --> B[Supervised Semantic<br/>CosyVoice 3]
A --> C[Self-supervised<br/>HuBERT, W2v-BERT]
A --> D[Unsupervised<br/>SoundStream]
B --> E[Best Content Consistency<br/>Maintains Speaker Similarity]
C --> F[Good Speaker Similarity<br/>Language Limitations]
D --> G[Poor Content Consistency<br/>High Error Rates]
style B fill:#c8e6c9
style E fill:#c8e6c9DiffRO Impact Assessment
Relative improvements from DiffRO post-training:
- Korean: 68.7% WER reduction (CosyVoice 3-0.5B)
- Cross-lingual scenarios: 50%+ improvements in half of conditions
- Low-resource languages: Particularly significant gains
- Trade-off consideration: Slight speaker similarity reduction
Future Directions & Limitations
Current Limitations
- Acoustic Control: Cannot control timbre through textual instructions
- Singing Synthesis: Limited performance for singing voice generation
- Emotional Speech ASR: Evaluation challenges due to ASR model bias toward standard pronunciations
Potential Improvements
Diagram source
graph TD
A[Future Enhancements] --> B[Timbre Control<br/>Natural Language]
A --> C[Singing Data Integration<br/>Tokenizer + LM]
A --> D[Tens of Millions Hours<br/>Dataset Expansion]
A --> E[Improved Reward Balance<br/>DiffRO Enhancement]
style A fill:#e3f2fd
style B fill:#fff3e0
style C fill:#f3e5f5
style D fill:#e8f5e8
style E fill:#ffebeeConclusion
CosyVoice 3 represents a paradigm shift in speech synthesis, moving from controlled laboratory conditions to robust real-world applications. Through innovative multi-task tokenization, differentiable reward optimization, and unprecedented data scaling, it achieves state-of-the-art performance across multiple languages and domains.
The model's success demonstrates the importance of:
- Supervised semantic tokenization for better content-prosody balance
- Reward-based post-training for targeted performance improvements
- Massive multilingual datasets for robust generalization
- Architectural scaling combined with training innovations
For enthusiasts and researchers, CosyVoice 3 provides a comprehensive blueprint for building production-ready speech synthesis systems that can handle the complexity and diversity of real-world applications.
Demo: Listen to CosyVoice 3 samples at https://funaudiollm.github.io/cosyvoice3
Research Paper: arXiv:2505.17589v2
Try Our Voice Clone Demo
Hear your words come to life
Choose a voice and try a short preview.
Listen to sample voices
Hear examples before choosing a voice. Generated results can vary with the script and reference sample.
Looking for another voice?
Explore the library and listen to a sample before you create.
Morgan Freeman
Stephen Hawking
Christiano Ronaldo
Donald Trump
Kokoro
Disney XD Announcer
Cute Japanese Girl
Vin
Adam Stone
Transform Your Content with AI Voice Technology Today
Try a short voice preview, then create speech and save your audio in a workspace built for your next project.
Generate Your Voice Now