Skip to main content
This node converts written text into spoken audio using Fish Audio text-to-speech models. It supports emotion cues embedded in the text ([happy], [whispering] on s2.1-pro; (happy) on s1) and multi-speaker dialogue using @Voice1/@Voice2 tags when multiple voices are connected. Two models are available: s2.1-pro, which supports up to five voices and multi-speaker dialogue, and s1, which uses a single optional voice.

Inputs

Common Inputs

s2.1-pro Inputs

These inputs appear when the s2.1-pro model is selected.

s1 Inputs

These inputs appear when the s1 model is selected. Note: The text input must not be empty. Speaker tags (@Voice1, @Voice2, etc.) are case-insensitive and must refer to a connected voice; tagging a voice that is not connected raises an error. When two or more voices are connected, the text must reference every connected voice at least once, or the node reports the missing tags. On s2.1-pro, connecting 0 voices uses the default voice, 1 voice uses that voice alone, and 2 or more voices enable multi-speaker dialogue. On s1, a single optional voice is used and leaving it unconnected uses the default voice. Emotion cues can be placed in the text: [happy] and [whispering] on s2.1-pro, and (happy) on s1.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 6cc005ae76fc7b60d9399b1b0a3c5de40a6eff47cd6f0f0b73b4212c0270ae29