Key capabilities
- Up to 30 seconds: Generates long clips with coherent motion and audio in a single pass
- Native audio: Every output includes a synchronized audio track by default
- Text, image, and reference input: Pure text prompts, first-frame animation, or up to 10 reference images, 5 reference videos, and 5 reference audio clips
- Reference tags: Mention connected media directly in the prompt as
@Image1,@Video1, or@Audio1 - Flexible output: 480p, 720p, or 1080p resolution with adaptive, 16:9, 9:16, 1:1, 4:3, or 3:4 aspect ratios
- Bilingual prompts: Prompt in English or Chinese
Wan 3.0 text-to-video workflow
Generate a video from a text prompt alone, with synchronized audio included by default.Run on Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “Wan 3.0” in Template Library
Prompting tips
- Duration: Set 2 to 30 seconds, or
autoto let the model choose a duration that fits the prompt - Resolution and ratio: The workflow ships at 720p with
adaptiveratio; switch to 1080p or a fixed ratio like16:9for specific formats - Audio: Audio is generated by default. Disable the
audiooption to get silent video - Prompt extend: Prompt enhancement is on by default and rewrites your prompt for better results. Turn it off in advanced settings if you need exact wording
Wan 3.0 image-to-video workflow
Animate a single image into a video, with an optional last frame that the output transitions toward.Run on Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “Wan 3.0” in Template Library
Prompting tips
- First frame: The
first_frameinput is required. Thelast_frameinput is optional, and the model generates the motion between the two - Adaptive ratio: With
adaptive, the output dimensions follow the input image - Prompt: Describe what moves and how, plus the audio you want. Keep the subjects and composition unchanged unless you want the model to alter them
Wan 3.0 reference-to-video workflow
Generate a video conditioned on reference images, videos, and audio, with each reference assigned a job through tags in the prompt.Run on Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “Wan 3.0” in Template Library
Prompting tips
- Reference by tag: Refer to each input in the prompt as
@Image1,@Video1, or@Audio1, numbered per type in input order. The sample prompt assigns the shoe design to@Image1and the character to@Image2 - Limits: Up to 10 reference images, 5 reference videos, and 5 reference audio clips. Each reference video or audio clip must be 15 seconds or shorter, and the combined duration of all reference videos must stay within 15 seconds (same for audio). Total reference video plus output duration must not exceed 30 seconds
- Assign a job: State what each reference controls (product design, character identity, style, motion, voice) and what must stay unchanged. Explicit assignments produce more consistent results
- Prompt or reference required: Provide a prompt or at least one reference input, or the node returns an error
Get started
- Update ComfyUI to the latest version, or use the latest Comfy Desktop
- Go to Template Library, search for
Wan 3.0 - Load one of the three workflows (text-to-video, image-to-video, or reference-to-video)
- Log in with your Comfy account and ensure you have enough credits
- Connect the input materials (first frame or reference images) and run
FAQs
What is Wan 3.0?
What is Wan 3.0?
Wan 3.0 is Alibaba’s latest video generation model, available in ComfyUI through Partner Nodes. It generates up to 30 seconds of video with synchronized audio in a single pass, and accepts text prompts, reference images, reference videos, and reference audio.
How long can Wan 3.0 videos be?
How long can Wan 3.0 videos be?
Up to 30 seconds per clip. The duration input accepts 2 to 30 seconds, or
auto to let the model choose a duration that fits the prompt.Does Wan 3.0 generate audio?
Does Wan 3.0 generate audio?
Yes. Every output includes a synchronized audio track by default, with dialogue, sound effects, and music generated together with the video. Disable the
audio option for silent video.Can Wan 3.0 use reference images, videos, or audio?
Can Wan 3.0 use reference images, videos, or audio?
Yes, in the reference-to-video workflow. Connect up to 10 reference images, 5 reference videos, and 5 reference audio clips, then refer to them in the prompt as
@Image1, @Video1, or @Audio1.What is the difference between Wan 3.0 and Wan 2.7?
What is the difference between Wan 3.0 and Wan 2.7?
Wan 3.0 is the newest generation of Alibaba’s video model. It extends the multimodal pipeline with reference tags for images, videos, and audio, native synchronized audio by default, and up to 30 seconds of output per clip. See the Wan 2.7 page for the previous generation.
Does Wan 3.0 support Chinese prompts?
Does Wan 3.0 support Chinese prompts?
Yes. The prompt input supports both English and Chinese.