Skip to content
VocalCopyCat

CosyVoice 3: Scaling Towards In-the-Wild Speech Generation

Explore CosyVoice 3's speech generation architecture, training approach, multilingual capabilities, and evaluation in this technical research overview.

CosyVoice 3text to speech AIspeech synthesisvoice cloningmultilingual TTSzero shot voice generationAlibaba speech AIDiffRO training
By Randy WakeUpdated 5 min read
CosyVoice 3: Scaling Towards In-the-Wild Speech Generation
CosyVoice 3: Scaling Towards In-the-Wild Speech Generation

Executive Summary

CosyVoice 3 represents a significant leap forward in zero-shot multilingual speech synthesis, designed specifically for real-world applications. Developed by Alibaba's Speech Team at Tongyi Lab, this model addresses the limitations of its predecessor CosyVoice 2 through massive scaling in both data (from 10K to 1M hours) and model parameters (from 0.5B to 1.5B), while introducing novel techniques for improved prosody naturalness and content consistency.

Key Innovations Overview

Diagram source
mindmap
  root((CosyVoice 3))
    Speech Tokenizer
      Multi-task Training
      MinMo Integration
      FSQ Module
      25Hz Token Rate
    Post-training
      DiffRO Method
      Multi-task Rewards
      Token-level Optimization
    Data Scaling
      1M Hours Total
      9 Languages
      18 Chinese Dialects
      Real-world Audio
    Model Scaling
      1.5B Parameters
      DiT Architecture
      Enhanced CFM

Architecture Deep Dive

1. Multi-Task Speech Tokenizer

The foundation of CosyVoice 3's improved performance lies in its novel speech tokenizer, which builds upon the MinMo multimodal LLM rather than the SenseVoice-Large ASR model used in CosyVoice 2.

Diagram source
graph TD
    A[Speech Input X] --> B[Voice Encoder1<br/>12 Transformer Blocks + RoPE]
    B --> C[Intermediate Representations H]
    C --> D[FSQ Module<br/>Finite Scalar Quantization]
    D --> E[Voice Encoder2]
    E --> F[MinMo LLM]
    F --> G[Text Token Predictions]
    
    D --> H[Speech Tokens μ<br/>25 Hz Rate]
    
    I[Multi-task Training] --> F
    I --> J[ASR]
    I --> K[Language ID]
    I --> L[Emotion Recognition]
    I --> M[Audio Event Detection]
    I --> N[Speaker Analysis]
    
    style D fill:#e1f5fe
    style I fill:#f3e5f5

FSQ Quantization Process

The Finite Scalar Quantization (FSQ) module operates through a sophisticated two-step process:

  1. Dimensionality Reduction: Projects intermediate representations H into a D-dimensional low-rank space
  2. Bounded Quantization: Quantizes each dimension into the range [-K, K] using bounded round operations

Mathematical Formulation:

H̄ = ROUND(Proj_down(H))
Ĥ = Proj_up(H̄)
μᵢ = Σ(j=0 to D-1) h̄ᵢ,ⱼ × (2K + 1)ʲ

2. Differentiable Reward Optimization (DiffRO)

CosyVoice 3 introduces DiffRO, a novel post-training technique that optimizes speech tokens directly rather than synthesized audio, addressing computational challenges in traditional RL approaches.

Diagram source
sequenceDiagram
    participant LLM as Language Model
    participant GS as Gumbel-Softmax
    participant T2T as Token2Text Model
    participant Reward as Reward Calculator
    
    LLM->>GS: Predicted Token Probabilities
    GS->>T2T: Sampled Speech Tokens μ̃
    T2T->>Reward: ASR Posterior Probability
    Reward->>LLM: Gradient Signal
    
    Note over LLM,Reward: Direct token optimization<br/>bypasses CFM/Vocoder

Multi-Task Reward (MTR) Mechanism

DiffRO extends beyond basic ASR rewards to include multiple downstream tasks:

  • Speech Emotion Recognition (SER): Controls emotional expression
  • MOS Score Prediction: Maintains audio quality
  • Audio Event Detection (AED): Handles environmental sounds
  • Speaker Analysis: Preserves speaker characteristics

Training Pipeline

The CosyVoice 3 training process follows a sophisticated multi-stage approach designed to maximize performance while maintaining stability.

Diagram source
graph LR
    A[Large-scale Pretraining<br/>1M Hours] --> B[DiffRO Post-training<br/>Selected Data]
    B --> C[Zero-shot LM & CFM]
    C --> D[Continual Pretraining<br/>Text2Token LM]
    D --> E[Speaker Fine-tuning<br/>Multi-speaker Data]
    
    F[Text-based LLM<br/>Initialization] --> A
    
    G[Emotional, Instructed,<br/>Multi-lingual Data] --> C
    
    style A fill:#e8f5e8
    style B fill:#fff3e0
    style C fill:#e3f2fd

Stage Breakdown

  1. Initialization: Leverage pre-trained text-based LLMs for semantic understanding
  2. Large-scale Pretraining: Train on the massive 1M-hour multilingual dataset
  3. Post-training with DiffRO: Optimize performance using reward-based learning
  4. Continual Pretraining: Transfer capabilities to specialized models
  5. Speaker Fine-tuning: Enhance individual speaker quality and consistency

Dataset Scaling Analysis

CosyVoice 3's impressive performance stems significantly from its unprecedented dataset scale and diversity.

Diagram source
pie title Language Distribution (1M Hours Total)
    "Chinese" : 45.2
    "English" : 32.1
    "Japanese" : 8.7
    "Russian" : 6.4
    "German" : 3.8
    "Korean" : 2.1
    "Others" : 1.7

Chinese Dialect Coverage

Diagram source
pie title Chinese Dialect Distribution
    "Sichuan" : 18.93
    "Hubei" : 14.48
    "Cantonese" : 8.61
    "Wuzhong" : 7.34
    "Shan1xi" : 8.54
    "Suhang" : 6.64
    "Shanghai" : 6.27
    "Others" : 29.19

Data Processing Pipeline

The multilingual data pipeline ensures high-quality training material through six critical steps:

Diagram source
graph TD
    A[Raw Audio Data] --> B[Speech Detection &<br/>Segmentation]
    B --> C[Noise Reduction<br/>MossFormer2]
    C --> D[ASR Transcription<br/>Multi-model Validation]
    D --> E[Punctuation Adjustment<br/>Montreal Forced Aligner]
    E --> F[Volume Standardization<br/>0.6 Peak Normalization]
    F --> G[Length Ratio Filtering<br/>Remove 1% smallest, 5% largest]
    G --> H[Clean Training Data]
    
    style C fill:#ffebee
    style D fill:#e8f5e8
    style G fill:#fff3e0

Performance Benchmarks

SEED-TTS-Eval Results

CosyVoice 3 demonstrates substantial improvements over its predecessor and competitive models:

Modeltest-zh CER (%)test-en WER (%)test-hard CER (%)
CosyVoice 21.452.576.83
CosyVoice 3-0.5B1.162.026.08
CosyVoice 3-1.5B1.122.215.83
CosyVoice 3-1.5B+RL0.711.455.66

Key Improvements:

  • 44% relative improvement in Chinese content consistency
  • 51% relative improvement in English content consistency
  • 26% relative improvement on challenging test cases

CV3-Eval Multilingual Benchmark

CosyVoice 3 is the only system capable of handling all languages in the comprehensive CV3-Eval benchmark:

Diagram source
graph LR
    A[CV3-Eval Benchmark] --> B[Multilingual Voice Cloning<br/>9 Languages × 500 Samples]
    A --> C[Cross-lingual Transfer<br/>zh, en, ja, ko]
    A --> D[Emotion Cloning<br/>Happy, Sad, Angry]
    A --> E[Subjective Evaluation<br/>Expressive & Accent Cloning]
    
    style B fill:#e8f5e8
    style C fill:#fff3e0
    style D fill:#f3e5f5
    style E fill:#e3f2fd

Advanced Features

1. Pronunciation Inpainting

CosyVoice 3 addresses mispronunciations through mixed word-phoneme modeling:

Diagram source
graph TD
    A[Raw Text Input] --> B{Contains Polyphonic<br/>Characters/Words?}
    B -->|Yes| C[Replace with Phonemes<br/>Mixed Vocabulary]
    B -->|No| D[Standard Processing]
    C --> E[Enhanced Pronunciation<br/>Control]
    D --> E
    
    F[Auxiliary Training Set] --> G[Chinese Pinyin<br/>Replacement]
    F --> H[English CMU Dict<br/>Phonemes]
    G --> C
    H --> C

2. Self-Training for Text Normalization

The system eliminates hand-crafted rules through LLM-based text normalization:

Three-pronged Approach:

  1. Rule-based TN → Audio synthesis via CosyVoice 2
  2. Qwen-Max TN → Audio synthesis on normalized text
  3. Inverse TN → Raw text generation from existing pairs

3. Instructed Speech Generation

Extended from 1,500 to 5,000 hours of instruction-following data, supporting 100+ speaking styles:

Categories:

  • Emotions: Happy, sad, angry, fearful, surprised, etc.
  • Characteristics: Fast, slow, loud, soft, authoritative, etc.
  • Roles: Warrior, poet, merchant, detective, etc.
  • Dialects: 10 Chinese regional variants
  • Accents: Indian English, Russian English, etc.

Technical Innovations Deep Dive

Model Architecture Enhancements

Diffusion Transformer (DiT) Integration

CosyVoice 3 adopts the DiT architecture for its Conditional Flow Matching (CFM) model:

Diagram source
graph TD
    A[Speech Tokens<br/>25 Hz] --> B[Interpolation<br/>Rate Matching]
    B --> C[DiT Backbone<br/>300M Parameters]
    C --> D[Mel Features<br/>Generation]
    D --> E[Vocoder<br/>Audio Output]
    
    F[Text Encoder<br/>Removed] -.->|Simplified| C
    G[Length Regularization<br/>Removed] -.->|Simplified| C
    
    style C fill:#e1f5fe
    style F fill:#ffebee,stroke-dasharray: 5 5
    style G fill:#ffebee,stroke-dasharray: 5 5

Scaling Impact Analysis

The transition from 0.5B to 1.5B parameters yields measurable improvements:

Diagram source
graph LR
    A[Model Size] --> B[0.5B Parameters]
    A --> C[1.5B Parameters]
    
    B --> D[Good Performance<br/>Most Languages]
    C --> E[Enhanced Performance<br/>Complex Scenarios]
    
    D --> F[WER: 2.02% EN]
    E --> G[WER: 1.45% EN<br/>with RL]
    
    style C fill:#e8f5e8
    style E fill:#e8f5e8
    style G fill:#c8e6c9

Speaker Fine-tuning Innovations

Monolingual to Polyglot Transformation

CosyVoice 3 can transform monolingual speakers into polyglots through targeted training:

Diagram source
sequenceDiagram
    participant MS as Monolingual Speaker
    participant AD as Auxiliary Dataset
    participant LI as Language Instruction
    participant PS as Polyglot Speaker
    
    MS->>AD: Studio-quality samples
    AD->>LI: "You are Speaker X. Please speak German."
    LI->>PS: Cross-lingual capability
    
    Note over MS,PS: Supports 9 languages<br/>18 Chinese dialects

Capability Transfer Mechanism

The fine-tuning process preserves pre-trained capabilities while adapting to specific speakers:

  1. Partial Speaker ID Labeling: Mix labeled and unlabeled data
  2. Instruction Masking: Randomly mask speaker/style prompts
  3. Catastrophic Forgetting Prevention: Maintain instruction coverage

Performance Analysis & Ablations

Speech Tokenizer Comparison

Diagram source
graph TD
    A[Tokenizer Comparison] --> B[Supervised Semantic<br/>CosyVoice 3]
    A --> C[Self-supervised<br/>HuBERT, W2v-BERT]
    A --> D[Unsupervised<br/>SoundStream]
    
    B --> E[Best Content Consistency<br/>Maintains Speaker Similarity]
    C --> F[Good Speaker Similarity<br/>Language Limitations]
    D --> G[Poor Content Consistency<br/>High Error Rates]
    
    style B fill:#c8e6c9
    style E fill:#c8e6c9

DiffRO Impact Assessment

Relative improvements from DiffRO post-training:

  • Korean: 68.7% WER reduction (CosyVoice 3-0.5B)
  • Cross-lingual scenarios: 50%+ improvements in half of conditions
  • Low-resource languages: Particularly significant gains
  • Trade-off consideration: Slight speaker similarity reduction

Future Directions & Limitations

Current Limitations

  1. Acoustic Control: Cannot control timbre through textual instructions
  2. Singing Synthesis: Limited performance for singing voice generation
  3. Emotional Speech ASR: Evaluation challenges due to ASR model bias toward standard pronunciations

Potential Improvements

Diagram source
graph TD
    A[Future Enhancements] --> B[Timbre Control<br/>Natural Language]
    A --> C[Singing Data Integration<br/>Tokenizer + LM]
    A --> D[Tens of Millions Hours<br/>Dataset Expansion]
    A --> E[Improved Reward Balance<br/>DiffRO Enhancement]
    
    style A fill:#e3f2fd
    style B fill:#fff3e0
    style C fill:#f3e5f5
    style D fill:#e8f5e8
    style E fill:#ffebee

Conclusion

CosyVoice 3 represents a paradigm shift in speech synthesis, moving from controlled laboratory conditions to robust real-world applications. Through innovative multi-task tokenization, differentiable reward optimization, and unprecedented data scaling, it achieves state-of-the-art performance across multiple languages and domains.

The model's success demonstrates the importance of:

  • Supervised semantic tokenization for better content-prosody balance
  • Reward-based post-training for targeted performance improvements
  • Massive multilingual datasets for robust generalization
  • Architectural scaling combined with training innovations

For enthusiasts and researchers, CosyVoice 3 provides a comprehensive blueprint for building production-ready speech synthesis systems that can handle the complexity and diversity of real-world applications.


Demo: Listen to CosyVoice 3 samples at https://funaudiollm.github.io/cosyvoice3

Research Paper: arXiv:2505.17589v2

Try Our Voice Clone Demo

Hear your words come to life

Choose a voice and try a short preview.

77 / 120 input characters
Continue with 2,000 welcome credits

Listen to sample voices

Hear examples before choosing a voice. Generated results can vary with the script and reference sample.

Looking for another voice?

Explore the library and listen to a sample before you create.

Morgan Freeman avatar

Morgan Freeman

Morgan Freeman voice sample0:00 --:--
Stephen Hawking avatar

Stephen Hawking

Stephen Hawking voice sample0:00 --:--
Christiano Ronaldo avatar

Christiano Ronaldo

Christiano Ronaldo voice sample0:00 --:--
Donald Trump avatar

Donald Trump

Donald Trump voice sample0:00 --:--
Kokoro avatar

Kokoro

Kokoro voice sample0:00 --:--
Disney XD Announcer avatar

Disney XD Announcer

Disney XD Announcer voice sample0:00 --:--
Cute Japanese Girl avatar

Cute Japanese Girl

Cute Japanese Girl voice sample0:00 --:--
Vin avatar

Vin

Vin voice sample0:00 --:--
Adam Stone avatar

Adam Stone

Adam Stone voice sample0:00 --:--

Transform Your Content with AI Voice Technology Today

Try a short voice preview, then create speech and save your audio in a workspace built for your next project.

Generate Your Voice Now

Pricing Options

Credits are billed per UTF-8 byte after text normalization. Library voices use 1 credit per byte; custom voices and cloning use 5. Creating a saved voice costs 10,000 credits.

Starter Package
Start with a small prepaid balance for your next voiceover.
$5one-time

100,000 credits

  • 100,000 prepaid credits
  • Library speech: 1 credit per normalized UTF-8 byte
  • Custom voices and cloning: 5 credits per byte
  • Projects, saved voices, REST API and MCP access
Creator Package
Keep creating with a larger balance for regular voice projects.
$35one-time

1,750,000 credits

  • 1,750,000 prepaid credits
  • Library speech: 1 credit per normalized UTF-8 byte
  • Custom voices and cloning: 5 credits per byte
  • Projects, saved voices, REST API and MCP access
Premium Package
Get our best credit rate for a busy creative workflow.
$100one-time

10,000,000 credits

  • 10,000,000 prepaid credits
  • Library speech: 1 credit per normalized UTF-8 byte
  • Custom voices and cloning: 5 credits per byte
  • Projects, saved voices, REST API and MCP access

Every package. Every creative tool.

Library voicesVoice cloning & saved voicesParagraph projects & downloadsREST API & MCP access

Latest Posts