Inputs
Note: When an image is provided, it is resized so its total pixel count stays close to 1,048,576 (1024 × 1024), and only its RGB channels are used. The resized image is passed to the CLIP tokenizer together with the prompt. When both
image and vae are provided, the node also encodes the image into reference latents and attaches them to the conditioning output.
Outputs
This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub
Source fingerprint (SHA-256):
ec6980a63eab0d6c95be3abea00b2bf3018d30a1267f0b39a21be29a3e9228fe