Fish Audio S2 Research: Expressive Speech and Dialogue
Examine Fish Audio S2's dual-autoregressive design, expressive control, and dialogue generation, plus the limits to check before trusting benchmark claims.

Fish Audio S2 is a research case study in controlling how speech is delivered while keeping generation efficient enough for interactive use. Its distinctive combination is a two-level autoregressive architecture, natural-language performance cues, and support for dialogue involving multiple speakers.
Research reviewed September 17, 2026.
The most useful way to assess S2 is to separate expressive success from word accuracy and operational speed. A lively demonstration can make a system feel convincing while leaving unanswered questions about consistent speaker identity, repeatability, and the conditions behind a latency result.
Identify what the architecture separates
The S2 technical report, first submitted March 9 and revised March 11, 2026, describes a Dual-AR design. A larger temporal model predicts the sequence over time, while a smaller depth-wise model supplies acoustic detail within each step. The report identifies a Qwen3-4B backbone and a four-layer acoustic decoder.
This division makes the architecture worth examining beyond its parameter count. A speech representation can contain multiple pieces of information for each moment. Treating every piece as another item along one long timeline is not the only way to generate it.
For readers, the practical question is whether the separation delivers a useful balance among speech detail, generation cost, and stability. Architecture provides a hypothesis about that balance. It does not remove the need to examine actual outputs under comparable conditions.
Treat expressive tags as scoped instructions
Fish Audio’s official release article describes fine-grained performance control through natural-language tags and presents examples of emotional and nonverbal delivery. It also discusses multi-speaker, multi-turn generation. Those features make S2 relevant to dialogue and character narration rather than only isolated sentences.
A good experiment gives each instruction a clear scope. Consider a fictional exchange:
Speaker A: “You found the missing recording?”
Speaker B: “Yes. It was in the archive all along.”
Create a neutral baseline first. Then request surprise from A and quiet relief from B. Ask reviewers whether the change occurs in the intended turn, whether the words remain correct, and whether the speakers remain distinguishable.
This is a proposed evaluation script, not a claim about a particular output. If an instruction changes the entire exchange when only one turn should change, record that as a control issue.
Keep personality distinct from speaker identity
A model can make a voice more animated without preserving the same vocal identity consistently. Conversely, it can maintain a recognizable speaker while failing to deliver the requested performance. These dimensions deserve separate scores.
Build a small grid with one speaker across several styles and several speakers using one style. That structure helps reveal whether the system is changing the intended factor or collapsing different voices into a similar expressive pattern.
For example, test a restrained explanation, an excited discovery, and a disappointed response. Use content that makes each performance plausible. An emotionally contradictory sentence introduces another variable.
The guide to natural AI narration discusses delivery in a production context. For research evaluation, add an identity question: would a listener recognize the same fictional or consenting reference speaker across the variations?
Read speed claims with the serving setup attached
The report evaluates its optimized inference path on an NVIDIA H200 and reports an RTF of 0.195 with first-audio performance around 100 ms. It discusses serving optimizations and reference-context caching. These are author-reported system measurements, not independent results from every device or deployment. The report’s inference section gives the relevant context.
RTF compares generation time with the duration of the audio produced. It does not by itself tell you how soon a listener can hear the first word. Nor does a warm cached request establish the experience of loading a new voice for the first time.
For a local trial, separate fresh-reference requests from repeated-reference requests. Measure a short reply and a longer exchange. Keep concurrency fixed before increasing it, and record playback gaps as well as total generation time.
Test dialogue as a sequence of commitments
Dialogue quality includes more than alternating voices. Each turn should answer the previous one, arrive at an appropriate pace, and preserve the assigned speaker. If the script contains a laugh or hesitation, its placement should support the exchange rather than obscure a word.
Create a turn-level review sheet:
| Check | Example failure to mark |
|---|---|
| Speaker assignment | A line uses the other speaker’s voice |
| Text completion | The final clause disappears |
| Turn transition | An unintended gap interrupts the exchange |
| Performance scope | One cue changes later neutral turns |
| Nonverbal event | A sound overlaps important speech |
Review the complete sequence before isolating an error. A pause that sounds excessive alone may be appropriate after a question. A smooth join can still be wrong if it erases an intended interruption.
Retain a plain version of the dialogue without performance cues. That baseline helps identify whether the extra control improved the scene.
Distinguish a public release from unrestricted use
The official Fish Speech repository identifies the code and associated weights with the Fish Audio Research License. The release’s “open-source” wording should therefore not be treated as shorthand for unrestricted commercial permission.
Record the exact model and license revision used in an evaluation. A hosted service and downloaded weights may have different terms and operational arrangements. This article does not infer permissions beyond the named release materials.
Version identity matters technically too. A current repository may describe a model variant or serving integration that differs from the report’s original experiment. Keep the paper version, checkpoint, and implementation commit together when documenting a result.
That practice makes it possible to revisit an experiment without guessing which “S2” someone meant.
Judge the amount of usable dialogue
A useful pilot asks how much accepted dialogue the team obtains for a given amount of review and correction. Count rejected turns, speaker swaps, pronunciation corrections, and edits needed to preserve conversational timing.
Do not hide these behind a single preference score. A beautiful one-off character voice may suit a short creative piece while requiring too much supervision for a large lesson library.
Finish with the voice-over quality assurance checklist, applying it to the exported dialogue as well as individual turns. Preserve the accepted script, instructions, reference material, and issue log.
S2’s research contribution is a concrete approach to combining expressive control with an efficient generation structure. Its value for a particular project depends on whether that control remains predictable across the actual voices, turns, and listening conditions the project requires.
Try Our Voice Clone Demo
Hear your words come to life
Choose a voice and try a short preview.
Listen to sample voices
Hear examples before choosing a voice. Generated results can vary with the script and reference sample.
Looking for another voice?
Explore the library and listen to a sample before you create.
Morgan Freeman
Stephen Hawking
Christiano Ronaldo
Donald Trump
Kokoro
Disney XD Announcer
Cute Japanese Girl
Vin
Adam Stone
Transform Your Content with AI Voice Technology Today
Try a short voice preview, then create speech and save your audio in a workspace built for your next project.
Generate Your Voice Now