Skip to content
VocalCopyCat

VoxCPM2 Research: Continuous Speech and Voice Control

Understand VoxCPM2's continuous audio latents, voice design and cloning modes, and the evaluation questions behind its multilingual speech research.

VoxCPM2 researchcontinuous speech generationtokenizer-free TTScontrollable voice cloning
By Randy Wake6 min read
VoxCPM2 conceptual diagram where voice design, controlled cloning, and continuation are alternative inputs to continuous audio latents, local diffusion, and a 48 kHz waveform.
VoxCPM2 conceptual diagram where voice design, controlled cloning, and continuation are alternative inputs to continuous audio latents, local diffusion, and a 48 kHz waveform.

VoxCPM2 explores speech generation without an external discrete speech tokenizer, using continuous audio latents and a combination of autoregressive modeling and diffusion. Its research also unifies several tasks that users often encounter as separate products: designing a voice, controlling a cloned voice, and continuing from a reference recording.

Research reviewed September 17, 2026.

The useful question is not whether “tokenizer-free” sounds more advanced. It is what the representation and input modes allow, what they cost, and how to evaluate the resulting speech. The proposed tests below focus on those distinctions rather than an overall model ranking.

Understand what continuous means here

The June 5 VoxCPM2 technical report describes hierarchical diffusion-autoregressive generation in a continuous audio latent space. Its AudioVAE encodes audio at 16 kHz and reconstructs at 48 kHz. The reported model has a 2B backbone and supports a multilingual set covering 30 languages and nine Chinese dialects.

“Tokenizer-free” does not mean the system has no internal representation or that raw sound passes directly through an ordinary text model. It distinguishes the speech modeling approach from an external discrete speech-token pipeline.

For evaluation, resist treating an output sample rate as a listening score. A file’s technical format tells you how it is delivered. Whether the voice sounds clear, stable, and appropriate remains an audible question.

Similarly, language coverage identifies where to investigate. It does not establish that every supported language, accent, or speaking style performs equally well.

Separate the three generation modes

The official repository documents voice design from a description, reference-based cloning with style control, and continuation cloning using reference audio with its transcript. It presents these as different input arrangements for the model.

That distinction changes the evaluation target:

ModeWhat should be judged
Voice designWhether the result follows the requested character
Controlled cloningWhether identity survives the requested delivery change
ContinuationWhether new speech fits the supplied performance context

A continuation result that resembles the reference closely does not prove that the model can change emotion independently. A successful designed narrator does not prove similarity to a particular person.

The voice cloning versus text-to-speech guide provides a useful starting distinction. VoxCPM2 adds another layer: different ways of conditioning a speech generator can serve different editorial goals.

Ask what a unified model simplifies

A shared model can make experiments easier to organize because related tasks use the same underlying system. That is an interpretation of the design, not a guarantee that every task is equally strong.

For a studio creating fictional characters, design could produce candidate voices before a team commits to one. For localization, controlled cloning might be tested for identity preservation across languages. For an updated narration, continuation could be examined for consistency with an approved passage.

These workflows have different failure costs. A surprising new character voice may be welcome during exploration. The same surprise in a revised training module can be a continuity problem.

Write down the desired behavior before choosing the mode. Otherwise, a mode can appear successful simply because the evaluation changed to fit whatever it produced.

Read benchmark tables without flattening the task

The report presents public zero-shot and instruction-following evaluations as well as an internal multilingual evaluation. These are author-reported results with distinct datasets and scoring procedures. The full technical report supplies the experimental detail needed to interpret them.

Keep public and internal tests separate in notes. Also distinguish correct words from speaker similarity and subjective naturalness. A single column cannot answer all three questions.

An internal average can hide a difficult language pair. A speaker-similarity score can miss an awkward phrase. A preference result can depend on the text, reference, and listening task offered to reviewers.

Use the tables to choose a hypothesis worth testing. For example: “Does controlled cloning preserve this narrator’s identity when the target language changes?” That is more actionable than copying a headline that a model is competitive in general.

Build an identity-and-style experiment

Choose one consenting reference speaker and a short script with several sentence types. Generate a neutral version, then ask for one clear delivery change. Keep the reference, text, and other settings fixed.

Ask separate questions:

  • Are the words complete and correct?
  • Does the voice still resemble the reference?
  • Is the requested delivery change audible?
  • Did timing or pronunciation become less stable?
  • Does the result remain consistent across repeated runs?

Use an instruction such as a calm explanation followed by a more urgent final sentence only if the chosen mode supports that form of control. Do not assume a paper’s general capability description specifies every interface detail.

Preserve all planned runs. A model that occasionally produces an excellent result can still require a different workflow from one that produces acceptable results consistently.

Test multilingual output with qualified listeners

For a cross-language pilot, begin with material whose meaning and intended pronunciation are already reviewed. Avoid using an unverified machine translation as the only source text; a translation error and a synthesis error require different fixes.

Have a fluent reviewer assess each language independently. Ask about names, word endings, rhythm, and whether the accent fits the audience. Ask a separate identity question when that is part of the project.

A small proposed test could include an introduction, a numbered instruction, a question, and a brief emotional change. Use equivalent content across languages while allowing natural phrasing. Exact word-for-word translation is not necessary for a fair task.

The multilingual voice-over workflow explains how to keep the language versions and approvals organized. Add the chosen VoxCPM2 mode and reference treatment to that record.

Keep representation claims connected to outcomes

A continuous-latent approach is an architectural choice, not a universal verdict against discrete models. To compare approaches meaningfully, hold the task and review procedure steady while documenting implementation differences.

Record the checkpoint, generation mode, reference transcript when used, language, sampling settings, and output format. Measure preparation and correction effort alongside synthesis time. If a mode needs a carefully transcribed reference, include that labor in the workflow assessment.

Inspect the final exported file after any editing or encoding. An accepted raw generation can still acquire a bad join or lose a word during assembly.

VoxCPM2 is valuable research because it makes representation and conditioning choices explicit. The strongest practical conclusion will describe which mode worked for which material, with what limits, rather than treating “continuous” as a substitute for listening.

Try Our Voice Clone Demo

Hear your words come to life

Choose a voice and try a short preview.

77 / 120 input characters
Continue with 2,000 welcome credits

Listen to sample voices

Hear examples before choosing a voice. Generated results can vary with the script and reference sample.

Looking for another voice?

Explore the library and listen to a sample before you create.

Morgan Freeman avatar

Morgan Freeman

Morgan Freeman voice sample0:00 --:--
Stephen Hawking avatar

Stephen Hawking

Stephen Hawking voice sample0:00 --:--
Christiano Ronaldo avatar

Christiano Ronaldo

Christiano Ronaldo voice sample0:00 --:--
Donald Trump avatar

Donald Trump

Donald Trump voice sample0:00 --:--
Kokoro avatar

Kokoro

Kokoro voice sample0:00 --:--
Disney XD Announcer avatar

Disney XD Announcer

Disney XD Announcer voice sample0:00 --:--
Cute Japanese Girl avatar

Cute Japanese Girl

Cute Japanese Girl voice sample0:00 --:--
Vin avatar

Vin

Vin voice sample0:00 --:--
Adam Stone avatar

Adam Stone

Adam Stone voice sample0:00 --:--

Transform Your Content with AI Voice Technology Today

Try a short voice preview, then create speech and save your audio in a workspace built for your next project.

Generate Your Voice Now

Pricing Options

Credits are billed per UTF-8 byte after text normalization. Library voices use 1 credit per byte; custom voices and cloning use 5. Creating a saved voice costs 10,000 credits.

Starter Package
Start with a small prepaid balance for your next voiceover.
$5one-time

100,000 credits

  • 100,000 prepaid credits
  • Library speech: 1 credit per normalized UTF-8 byte
  • Custom voices and cloning: 5 credits per byte
  • Projects, saved voices, REST API and MCP access
Creator Package
Keep creating with a larger balance for regular voice projects.
$35one-time

1,750,000 credits

  • 1,750,000 prepaid credits
  • Library speech: 1 credit per normalized UTF-8 byte
  • Custom voices and cloning: 5 credits per byte
  • Projects, saved voices, REST API and MCP access
Premium Package
Get our best credit rate for a busy creative workflow.
$100one-time

10,000,000 credits

  • 10,000,000 prepaid credits
  • Library speech: 1 credit per normalized UTF-8 byte
  • Custom voices and cloning: 5 credits per byte
  • Projects, saved voices, REST API and MCP access

Every package. Every creative tool.

Library voicesVoice cloning & saved voicesParagraph projects & downloadsREST API & MCP access

Latest Posts