Skip to main content
Vidu Q4 Preview is a video generation model from Vidu, available in ComfyUI through two partner nodes. It generates clips from 3 to 16 seconds with native audio, so the picture, the dialogue, and the sound effects all come from the model in a single pass. Output runs from 540p up to 4K. In ComfyUI, Vidu Q4 Image-to-Video Generation animates a single first frame with an optional prompt, and Vidu Q4 Reference-to-Video Generation builds a clip from up to 15 reference images, optional reference audio, and a prompt. Both are cloud API nodes: they need a Comfy account with credits and download no local model. See Partner Node pricing for the per-second rates. Image-to-video keeps the aspect ratio of the input image. Reference-to-video takes an explicit aspect_ratio, and both nodes expose duration, resolution, and an audio toggle.

Available workflows

Image to Video

Animate one image. The image is the first frame of the clip and the output keeps its aspect ratio, so crop the image when you need a different shape. The prompt is optional and describes what happens in the shot. Vidu Q4 Preview Image to Video workflow preview

Run on Comfy Cloud

Open in Comfy Cloud

Download Workflow

Download JSON or search “Vidu Q4 Preview: Image to Video” in Template Library
Input material

model_turquoise_yellow_outfit.png

Load into the LoadImage node that feeds the Vidu Q4 Image-to-Video node

Reference to Video

Build a clip from reference images, optional reference audio, and a prompt. Connect up to 15 reference images and up to 3 audio clips, then address them in the prompt by order: image 1, image 2, and so on. The template uses two reference images to keep a character and a prop consistent across the shot. Vidu Q4 Preview Reference to Video workflow preview

Run on Comfy Cloud

Open in Comfy Cloud

Download Workflow

Download JSON or search “Vidu Q4 Preview: Reference to Video” in Template Library
Input materials

clay_man_smiley_sweater.png

Load into the second reference image slot · image 2

yellow_fuzzy_smiley.png

Load into the first reference image slot · image 1

Workflow overview

Both templates use a small graph:
  • LoadImage: provides the first frame (image to video) or the reference images (reference to video)
  • Vidu4ImageToVideoNode / Vidu4ReferenceVideoNode: the core nodes, configured with the Vidu Q4 Preview model
  • SaveVideo: writes the finished clip

Steps to run

  1. Load the input: set the first frame for image to video, or the reference images for reference to video
  2. Select the model: keep Vidu Q4 Preview on the node
  3. Choose a resolution and duration: from 540p to 4K, and 3 to 16 seconds
  4. Toggle audio: leave it on to keep dialogue and sound effects, off for a silent clip
  5. Click Queue or press Ctrl+Enter to generate

Node controls

Prompting tips

  • Describe the shot, the line, and the sound. The model renders picture and audio together, so write the action, the spoken line in quotes, and the ambience in the same prompt.
  • Label references by their order. Reference images are addressed as image 1, image 2. Say what each one should keep, for example keeping a character’s face and outfit from image 1.
  • Write dialogue for a voice reference. Reference audio carries the voice, not the words, so write the line in the prompt and point it at a voice, such as image 1 says "Welcome back!" in the voice from audio 1.
  • Keep the aspect ratio inside the allowed range. Reference images must sit between 1:5 and 5:1.
  • Expect variation between runs. The seed does not lock the result; the same settings can still produce a different take.

Get started

  1. Update ComfyUI to the latest version
  2. Go to Template Library, search for Vidu Q4 Preview