Skip to main content
Fish Audio TTS gives you direct control over how speech sounds through inline tags written in the text itself. This guide covers every tag category the model understands, where to place them, and how to combine them. The examples below work in the Fish Audio Text to Speech node. Type them straight into the text field, replacing the sample script with your own. For node setup, workflow templates, and pricing, see Fish Audio: Text to Speech, Voice Cloning, and Speech to Text.

How tags work

Fish Audio S2 models interpret natural-language cues written in square brackets inside your text:
Three properties make this system flexible:
  • Position is meaning. A tag applies from where it appears until the next tag or the end of the sentence. [whispering] I didn't want to go inside whispers the whole line. I didn't want to go [whispering] inside starts whispering at “inside”.
  • Free-form descriptions work. You are not limited to a fixed tag set. If you can describe it to a voice actor, write it in brackets: [tired, end of a long shift], [voice rough from crying, trying to sound normal], [professional broadcast tone].
  • Tags work in your language. Write cues in the same language as your script: [低声说] 不要让他们听见, [ため息をついて] もう一度やり直そう.
The legacy s1 model uses the same emotion names but wraps them in parentheses with a fixed tag set: (happy) What a beautiful day! The current s2.1-pro model accepts bracket cues and free-form descriptions.

Emotion tags

Basic emotions

Advanced emotions

Tone markers

Tone markers control volume, intensity, and emphasis. Place [emphasis] right before the word or phrase you want to stress:

Audio effects

Audio effect tags add natural human sounds. Follow the tag with matching text where suggested:

Pauses and special effects

Plain pause words also shape rhythm without any tags: “um”, “uh”, “嗯”, “啊”.

Combining tags

Pair a physical tag with an emotion tag. Physical tags like [panting] or [whispering] register on their own, but pairing them with an emotion produces more consistent, natural results:
Layer emotions for complex expressions. Keep combinations to three tags or fewer per sentence:
Write emotion transitions as a natural progression across sentences:
Add background effects for atmosphere:
Modulate intensity with descriptive modifiers:

Best practices

Do:
  • Use one primary emotion per sentence
  • Match emotions to context logically
  • Add text after sound effects, for example “Ha ha” after [laughing]
  • Space out emotional changes for realism
  • Test different combinations on the same voice
Don’t:
  • Overuse tags in short text
  • Mix conflicting emotions in one sentence
  • Write bracket descriptions so long they interrupt the flow
  • Leave a descriptive tag without text after it
  • Place sentence-level emotion cues far from the sentence they control
If the output sounds unnatural, space out emotional changes, lower the intensity, or try a different voice: the same tag lands differently on different voices.

Multi-speaker dialogue

Connect two or more voices to the Fish Audio Text to Speech node, then mark speaker changes in the text with @Voice1, @Voice2, and so on. Every connected voice must appear in the text at least once, or the node raises an error. Voice connections are covered in Fish Audio: Text to Speech, Voice Cloning, and Speech to Text.

Phoneme control

Phoneme tags specify exact pronunciations for names, homographs, acronyms, and technical terms. Wrap the replacement pronunciation in <|phoneme_start|> and <|phoneme_end|> tags. Phoneme tags are preserved by text normalization, so keep normalize enabled. English (CMU Arpabet, replaces one word):
Chinese (tone-number pinyin, replaces one character or syllable, useful for polyphonic characters and names):
Japanese (OpenJTalk-style romaji with pitch accent digits):
Phoneme control combines with other tags in the same text:

Parameter tuning

The defaults are well tuned. Adjust only when you need to: The seed controls whether the node reruns; results are non-deterministic regardless of seed value.

Voice cloning quality

For the Fish Audio Instant Voice Clone node, reference recording quality decides the result. Recording limits are covered in Fish Audio: Text to Speech, Voice Cloning, and Speech to Text:
  • Record 2 to 3 clips of 15 to 20 seconds each, under 270 seconds in total
  • One speaker only, steady volume, consistent tone and emotion
  • Small pauses between sentences, about half a second
  • Quiet room, microphone about a hand’s width from the mouth
  • A phone recorder or headset mic works; avoid background music, TV, and echo
If the cloned voice sounds robotic, record more material, 30 to 60 seconds total, and speak naturally with pauses. Only clone voices you have permission to use: your own voice, or someone who gave you explicit permission.

References