- Small-SFX: Sound effects and short ambiance, up to 2:00. Small enough to run on CPU.
- Small-Music: Short music loops, on-device-friendly, up to 2:00.
- Medium: Longer tracks with stronger structure and musicality, up to ~6:20. Requires a GPU.
Available workflows
Stable Audio 3.0 Medium
Input a short text idea, optional duration, seed, and category. Generate stereo audio (music, SFX, or instruments) using Stable Audio 3 with optional AI-driven text expansion.
Run on Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “Stable Audio 3.0 Medium” in Template Library
- Text idea: Enter a short description of the sound, music, or effect you want (e.g. “upbeat electronic dance track with heavy bass”)
- Duration: Set the desired clip length in seconds (default varies)
- Seed: Control reproducibility by adjusting the seed value
- Category: Choose a reprompt preset: Music, Instrument, SFX, or One-shot
- Enable reprompt: Toggle
use_reprompton to let Qwen expand your short idea into a detailed prompt before generation - Click Run (
Ctrl/Cmd + Enter) to generate. The audio will be saved toComfyUI/output/audio/
Stable Audio 3.0 Medium Base
Input a short text description of a sound, music, or effect. The workflow expands your prompt with Qwen and generates a stereo audio clip from Stable Audio 3.
Run on Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “Stable Audio 3.0 Medium Base” in Template Library
- Text prompt: Enter a detailed description of the audio you want
- Duration: Set the clip length in seconds
- Seed: Control reproducibility
- Click Run (
Ctrl/Cmd + Enter) to generate
Model download
When loading the workflow, ComfyUI will prompt you with download links for any missing models. To set up manually, download the files below and place them in the correct folders.Checkpoints
stable_audio_3_medium.safetensors
For the Medium workflow. Place in models/checkpoints/
stable_audio_3_medium_base.safetensors
For the Medium Base workflow. Place in models/checkpoints/
Text encoders
t5gemma_b_b_ul2.safetensors
Required for all Stable Audio 3 workflows. Place in models/text_encoders/
qwen3.5_2b_bf16.safetensors
Required for the Medium workflow (Qwen reprompt). Place in models/text_encoders/