Skip to content
VocalCopyCat

Fish Audio S2 Research: Expressive Speech and Dialogue

Examine Fish Audio S2's dual-autoregressive design, expressive control, and dialogue generation, plus the limits to check before trusting benchmark claims.

Fish Audio S2 researchexpressive speech generationdual autoregressive TTSdialogue synthesis
By Randy Wake6 min read
Fish Audio S2 diagram showing text and style cues flowing through slow temporal autoregression, fast acoustic autoregression, and generated speech.
Fish Audio S2 diagram showing text and style cues flowing through slow temporal autoregression, fast acoustic autoregression, and generated speech.

Fish Audio S2 is a research case study in controlling how speech is delivered while keeping generation efficient enough for interactive use. Its distinctive combination is a two-level autoregressive architecture, natural-language performance cues, and support for dialogue involving multiple speakers.

Research reviewed September 17, 2026.

The most useful way to assess S2 is to separate expressive success from word accuracy and operational speed. A lively demonstration can make a system feel convincing while leaving unanswered questions about consistent speaker identity, repeatability, and the conditions behind a latency result.

Identify what the architecture separates

The S2 technical report, first submitted March 9 and revised March 11, 2026, describes a Dual-AR design. A larger temporal model predicts the sequence over time, while a smaller depth-wise model supplies acoustic detail within each step. The report identifies a Qwen3-4B backbone and a four-layer acoustic decoder.

This division makes the architecture worth examining beyond its parameter count. A speech representation can contain multiple pieces of information for each moment. Treating every piece as another item along one long timeline is not the only way to generate it.

For readers, the practical question is whether the separation delivers a useful balance among speech detail, generation cost, and stability. Architecture provides a hypothesis about that balance. It does not remove the need to examine actual outputs under comparable conditions.

Treat expressive tags as scoped instructions

Fish Audio’s official release article describes fine-grained performance control through natural-language tags and presents examples of emotional and nonverbal delivery. It also discusses multi-speaker, multi-turn generation. Those features make S2 relevant to dialogue and character narration rather than only isolated sentences.

A good experiment gives each instruction a clear scope. Consider a fictional exchange:

Speaker A: “You found the missing recording?”
Speaker B: “Yes. It was in the archive all along.”

Create a neutral baseline first. Then request surprise from A and quiet relief from B. Ask reviewers whether the change occurs in the intended turn, whether the words remain correct, and whether the speakers remain distinguishable.

This is a proposed evaluation script, not a claim about a particular output. If an instruction changes the entire exchange when only one turn should change, record that as a control issue.

Keep personality distinct from speaker identity

A model can make a voice more animated without preserving the same vocal identity consistently. Conversely, it can maintain a recognizable speaker while failing to deliver the requested performance. These dimensions deserve separate scores.

Build a small grid with one speaker across several styles and several speakers using one style. That structure helps reveal whether the system is changing the intended factor or collapsing different voices into a similar expressive pattern.

For example, test a restrained explanation, an excited discovery, and a disappointed response. Use content that makes each performance plausible. An emotionally contradictory sentence introduces another variable.

The guide to natural AI narration discusses delivery in a production context. For research evaluation, add an identity question: would a listener recognize the same fictional or consenting reference speaker across the variations?

Read speed claims with the serving setup attached

The report evaluates its optimized inference path on an NVIDIA H200 and reports an RTF of 0.195 with first-audio performance around 100 ms. It discusses serving optimizations and reference-context caching. These are author-reported system measurements, not independent results from every device or deployment. The report’s inference section gives the relevant context.

RTF compares generation time with the duration of the audio produced. It does not by itself tell you how soon a listener can hear the first word. Nor does a warm cached request establish the experience of loading a new voice for the first time.

For a local trial, separate fresh-reference requests from repeated-reference requests. Measure a short reply and a longer exchange. Keep concurrency fixed before increasing it, and record playback gaps as well as total generation time.

Test dialogue as a sequence of commitments

Dialogue quality includes more than alternating voices. Each turn should answer the previous one, arrive at an appropriate pace, and preserve the assigned speaker. If the script contains a laugh or hesitation, its placement should support the exchange rather than obscure a word.

Create a turn-level review sheet:

CheckExample failure to mark
Speaker assignmentA line uses the other speaker’s voice
Text completionThe final clause disappears
Turn transitionAn unintended gap interrupts the exchange
Performance scopeOne cue changes later neutral turns
Nonverbal eventA sound overlaps important speech

Review the complete sequence before isolating an error. A pause that sounds excessive alone may be appropriate after a question. A smooth join can still be wrong if it erases an intended interruption.

Retain a plain version of the dialogue without performance cues. That baseline helps identify whether the extra control improved the scene.

Distinguish a public release from unrestricted use

The official Fish Speech repository identifies the code and associated weights with the Fish Audio Research License. The release’s “open-source” wording should therefore not be treated as shorthand for unrestricted commercial permission.

Record the exact model and license revision used in an evaluation. A hosted service and downloaded weights may have different terms and operational arrangements. This article does not infer permissions beyond the named release materials.

Version identity matters technically too. A current repository may describe a model variant or serving integration that differs from the report’s original experiment. Keep the paper version, checkpoint, and implementation commit together when documenting a result.

That practice makes it possible to revisit an experiment without guessing which “S2” someone meant.

Judge the amount of usable dialogue

A useful pilot asks how much accepted dialogue the team obtains for a given amount of review and correction. Count rejected turns, speaker swaps, pronunciation corrections, and edits needed to preserve conversational timing.

Do not hide these behind a single preference score. A beautiful one-off character voice may suit a short creative piece while requiring too much supervision for a large lesson library.

Finish with the voice-over quality assurance checklist, applying it to the exported dialogue as well as individual turns. Preserve the accepted script, instructions, reference material, and issue log.

S2’s research contribution is a concrete approach to combining expressive control with an efficient generation structure. Its value for a particular project depends on whether that control remains predictable across the actual voices, turns, and listening conditions the project requires.

Try Our Voice Clone Demo

Hear your words come to life

Choose a voice and try a short preview.

77 / 120 input characters
Continue with 2,000 welcome credits

Listen to sample voices

Hear examples before choosing a voice. Generated results can vary with the script and reference sample.

Looking for another voice?

Explore the library and listen to a sample before you create.

Morgan Freeman avatar

Morgan Freeman

Morgan Freeman voice sample0:00 --:--
Stephen Hawking avatar

Stephen Hawking

Stephen Hawking voice sample0:00 --:--
Christiano Ronaldo avatar

Christiano Ronaldo

Christiano Ronaldo voice sample0:00 --:--
Donald Trump avatar

Donald Trump

Donald Trump voice sample0:00 --:--
Kokoro avatar

Kokoro

Kokoro voice sample0:00 --:--
Disney XD Announcer avatar

Disney XD Announcer

Disney XD Announcer voice sample0:00 --:--
Cute Japanese Girl avatar

Cute Japanese Girl

Cute Japanese Girl voice sample0:00 --:--
Vin avatar

Vin

Vin voice sample0:00 --:--
Adam Stone avatar

Adam Stone

Adam Stone voice sample0:00 --:--

Transform Your Content with AI Voice Technology Today

Try a short voice preview, then create speech and save your audio in a workspace built for your next project.

Generate Your Voice Now

Pricing Options

Credits are billed per UTF-8 byte after text normalization. Library voices use 1 credit per byte; custom voices and cloning use 5. Creating a saved voice costs 10,000 credits.

Starter Package
Start with a small prepaid balance for your next voiceover.
$5one-time

100,000 credits

  • 100,000 prepaid credits
  • Library speech: 1 credit per normalized UTF-8 byte
  • Custom voices and cloning: 5 credits per byte
  • Projects, saved voices, REST API and MCP access
Creator Package
Keep creating with a larger balance for regular voice projects.
$35one-time

1,750,000 credits

  • 1,750,000 prepaid credits
  • Library speech: 1 credit per normalized UTF-8 byte
  • Custom voices and cloning: 5 credits per byte
  • Projects, saved voices, REST API and MCP access
Premium Package
Get our best credit rate for a busy creative workflow.
$100one-time

10,000,000 credits

  • 10,000,000 prepaid credits
  • Library speech: 1 credit per normalized UTF-8 byte
  • Custom voices and cloning: 5 credits per byte
  • Projects, saved voices, REST API and MCP access

Every package. Every creative tool.

Library voicesVoice cloning & saved voicesParagraph projects & downloadsREST API & MCP access

Latest Posts