Skip to main content
Generate speech, music, sound effects and multi-speaker dialogue from a single prompt with ByteDance Seed Audio 1.0. Describe the voice(s), emotion, ambience, background music and sound effects in the prompt, and include the lines to speak. Optionally pick a built-in preset voice, clone voices from up to 3 reference clips (tagged @Audio1-3 in the prompt), or derive a voice from a character image. Up to 2 minutes of audio per run. The multilingual model supports 20 languages and timestamp-based timing control.

Inputs

Parameter Constraints

  • Reference mode dependencies: The reference_mode parameter determines which other inputs are required:
    • “text only”: No additional inputs required. The prompt must not contain @AudioN tags.
    • “audio reference”: Requires at least one of reference_audio_1, reference_audio_2, or reference_audio_3 to be connected. Reference clips must be connected in order without gaps. Each clip is limited to 30 seconds maximum duration. If @AudioN tags are used in the prompt, the highest tag number must not exceed the number of connected reference clips.
    • “image reference”: Requires reference_image to be connected. @AudioN tags are not used; the prompt should contain only the text to synthesize.
    • “preset voice”: Requires a preset voice to be selected. The whole prompt is read in the selected voice; @AudioN tags are not used as references, and tags such as @Audio2 or higher are rejected.
  • Audio reference ordering: In “audio reference” mode, reference audio inputs must be connected sequentially starting from reference_audio_1 without gaps. For example, you can connect reference_audio_1 and reference_audio_2, but not reference_audio_1 and reference_audio_3 without reference_audio_2.
  • Maximum audio tags: In “audio reference” mode, up to 3 reference clips can be connected (@Audio1, @Audio2, @Audio3), and the highest @AudioN tag in the prompt cannot exceed the number of connected reference audio inputs.
  • Model differences: The seed-audio-1.0-multilingual model supports 20 languages (English, Chinese, Japanese, Korean, Mexican & Castilian Spanish, Indonesian, German, Brazilian Portuguese, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish, Swedish) plus per-sentence timing control using timestamps in the format [5.5s:8.0s]. The seed-audio-1.0 model supports English and Chinese only, without timing control.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): e86e4edde424b4427d864350a9d3b082e271fbd2b1e335175637a9cc3ad51163