Skip to main content
Generate speech, music, sound effects and multi-speaker dialogue from a single prompt with ByteDance Seed Audio 1.0. Describe the voice(s), emotion, ambience, background music and sound effects in the prompt, and include the lines to speak. Optionally pick a built-in preset voice, clone voices from up to 3 reference clips (tagged @Audio1-3 in the prompt), or derive a voice from a character image. Up to 2 minutes of audio per run. The multilingual model supports 20 languages and timestamp-based timing control.

Inputs

Parameter Constraints

  • Reference mode dependencies: The reference_mode parameter determines which other inputs are required:
    • “text only”: No additional inputs required. The prompt must not contain @AudioN tags.
    • “audio reference”: Requires at least one of reference_audio_1, reference_audio_2, or reference_audio_3 to be connected. Reference clips must be connected in order without gaps (e.g., _1, then _2, then _3). Each clip is limited to 30 seconds maximum duration. The prompt must reference connected clips using @Audio1, @Audio2, @Audio3 tags.
    • “image reference”: Requires reference_image to be connected. The prompt must not contain @AudioN tags.
    • “preset voice”: Requires preset_voice to be selected. The prompt must not contain @AudioN tags (the entire prompt is read in the selected voice).
  • Audio reference ordering: When using “audio reference” mode, reference audio inputs must be connected sequentially starting from reference_audio_1 without gaps. For example, you can connect _1 and _2, but not _1 and _3 without _2.
  • Maximum audio tags: The prompt can reference up to 3 audio clips (@Audio1, @Audio2, @Audio3) when in “audio reference” mode. The highest numbered tag must not exceed the number of connected reference audio inputs.
  • Model differences: The “seed-audio-1.0-multilingual” model supports 20 languages (English, Chinese, Japanese, Korean, Mexican & Castilian Spanish, Indonesian, German, Brazilian Portuguese, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish, Swedish) plus per-sentence timing control using timestamps in the format [5.5s:8.0s]. The “seed-audio-1.0” model supports English and Chinese only, without timing control.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): e86e4edde424b4427d864350a9d3b082e271fbd2b1e335175637a9cc3ad51163