text field, replacing the sample script with your own. For node setup, workflow templates, and pricing, see Fish Audio: Text to Speech, Voice Cloning, and Speech to Text.
How tags work
Fish Audio S2 models interpret natural-language cues written in square brackets inside your text:- Position is meaning. A tag applies from where it appears until the next tag or the end of the sentence.
[whispering] I didn't want to go insidewhispers the whole line.I didn't want to go [whispering] insidestarts whispering at “inside”. - Free-form descriptions work. You are not limited to a fixed tag set. If you can describe it to a voice actor, write it in brackets:
[tired, end of a long shift],[voice rough from crying, trying to sound normal],[professional broadcast tone]. - Tags work in your language. Write cues in the same language as your script:
[低声说] 不要让他们听见,[ため息をついて] もう一度やり直そう.
s1 model uses the same emotion names but wraps them in parentheses with a fixed tag set: (happy) What a beautiful day! The current s2.1-pro model accepts bracket cues and free-form descriptions.
Emotion tags
Basic emotions
Advanced emotions
Tone markers
Tone markers control volume, intensity, and emphasis. Place[emphasis] right before the word or phrase you want to stress:
Audio effects
Audio effect tags add natural human sounds. Follow the tag with matching text where suggested:Pauses and special effects
Plain pause words also shape rhythm without any tags: “um”, “uh”, “嗯”, “啊”.
Combining tags
Pair a physical tag with an emotion tag. Physical tags like[panting] or [whispering] register on their own, but pairing them with an emotion produces more consistent, natural results:
Best practices
Do:- Use one primary emotion per sentence
- Match emotions to context logically
- Add text after sound effects, for example “Ha ha” after
[laughing] - Space out emotional changes for realism
- Test different combinations on the same voice
- Overuse tags in short text
- Mix conflicting emotions in one sentence
- Write bracket descriptions so long they interrupt the flow
- Leave a descriptive tag without text after it
- Place sentence-level emotion cues far from the sentence they control
Multi-speaker dialogue
Connect two or more voices to the Fish Audio Text to Speech node, then mark speaker changes in the text with@Voice1, @Voice2, and so on. Every connected voice must appear in the text at least once, or the node raises an error. Voice connections are covered in Fish Audio: Text to Speech, Voice Cloning, and Speech to Text.
Phoneme control
Phoneme tags specify exact pronunciations for names, homographs, acronyms, and technical terms. Wrap the replacement pronunciation in<|phoneme_start|> and <|phoneme_end|> tags. Phoneme tags are preserved by text normalization, so keep normalize enabled.
English (CMU Arpabet, replaces one word):
Parameter tuning
The defaults are well tuned. Adjust only when you need to:
The
seed controls whether the node reruns; results are non-deterministic regardless of seed value.
Voice cloning quality
For the Fish Audio Instant Voice Clone node, reference recording quality decides the result. Recording limits are covered in Fish Audio: Text to Speech, Voice Cloning, and Speech to Text:- Record 2 to 3 clips of 15 to 20 seconds each, under 270 seconds in total
- One speaker only, steady volume, consistent tone and emotion
- Small pauses between sentences, about half a second
- Quiet room, microphone about a hand’s width from the mouth
- A phone recorder or headset mic works; avoid background music, TV, and echo